AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,860 results
Model Releases

Soft-Prompt Tuning for Fair and Efficient LLM Benchmark Evaluation

DGX agent

arXiv:2606.12117v1 Announce Type: cross Abstract: Benchmark scores often misrepresent a large language model's (LLM's) knowledge, because they rely, e.g., on the model's ability to follow specific for

model-releasesarxiv-cs-ai
11 Jun 2026
Model Releases

Initial impressions of Claude Fable 5

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

I didn't have early access to today's Claude Fable 5 release, but I've spent the past ~5.5 hours putting it through its paces. My initial impressions are that this is something of a beast. It's slow,

model-releasessimon-willison
9 Jun 2026
Model Releases

Benchmark and optimize LLMs on-device with AI Edge Portal

DGX agent

LLMs have become more powerful at smaller sizes, but deploying them to edge devices like smartphones remains a massive challenge. Today, developers have to optimize across a sprawling combination of a

model-releasesgoogle-cloud-ai
20 May 2026
Model Releases

CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization

DGX agent

arXiv:2605.08873v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has emerged as a powerful algorithm for improving the reasoning capabilities of language models, but often fai

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Benchmarking System Dynamics AI Assistants: Cloud Versus Local LLMs on CLD Extraction and Discussion

DGX agent

arXiv:2604.18566v1 Announce Type: cross Abstract: We present a systematic evaluation of large language model families -- spanning both proprietary cloud APIs and locally-hosted open-source models -- o

model-releasesarxiv-cs-lg
21 Apr 2026
Safety

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization

DGX agent

arXiv:2509.23542v2 Announce Type: replace Abstract: The LLM-as-a-judge paradigm is widely used in both evaluating free-text model responses and reward modeling for model alignment and fine-tuning. Rec

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Instance-Adaptive Parametrization for Amortized Variational Inference

DGX agent

arXiv:2604.06796v1 Announce Type: cross Abstract: Latent variable models, including variational autoencoders (VAE), remain a central tool in modern deep generative modeling due to their scalability an

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

DeepSeek launches V4-Pro, its most advanced model that rivals Kimi K3 on some benchmarks but is priced much lower, at 0.435/1M input and 0.87/1M output tokens (The Information)

DGX agent

The Information: DeepSeek launches V4-Pro, its most advanced model that rivals Kimi K3 on some benchmarks but is priced much lower, at 0.435/1M input and 0.87/1M output tokens — Chinese AI developer D

model-releasestechmeme
13 Aug 2026
Model Releases

How Can Driving World Models Do Counterfactual Prediction?

DGX agent

arXiv:2608.11601v1 Announce Type: new Abstract: Driving world models are often interpreted as counterfactual simulators for observed driving episodes: given a factual driving log, they are asked what

model-releasesarxiv-cs-cv
13 Aug 2026
Model Releases

Test-Time Hallucination Control in Large Vision-Language Models

DGX agent

arXiv:2608.11474v1 Announce Type: new Abstract: Object Hallucination in large vision-language models (LVLMs), where models generate non-factual content about input images, remains a critical barrier t

model-releasesarxiv-cs-cv
13 Aug 2026
Applications

Assessing Reliability of BERT-Based Models on Question Answering Tasks

DGX agent

arXiv:2608.10806v1 Announce Type: new Abstract: Reliability estimation of large language models is in many cases as crucial as their accuracy, as reliable models are more trustworthy, robust, and suit

applicationsarxiv-cs-cl
12 Aug 2026
Applications

FACT: Failure-Aware Causal Training for World-Action Models

DGX agent

arXiv:2608.10232v1 Announce Type: cross Abstract: Recent world-action models (WAMs) show that co-training policies with future prediction can provide physical priors for action generation. Building on

applicationsarxiv-cs-ai
12 Aug 2026
Model Releases

From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models

DGX agent

arXiv:2608.10444v1 Announce Type: cross Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This pro

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models

DGX agent

arXiv:2506.03922v4 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchma

model-releasesarxiv-cs-ai
12 Aug 2026
Safety

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

DGX agent

arXiv:2608.10333v1 Announce Type: new Abstract: LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured

safetyarxiv-cs-lg
12 Aug 2026
Model Releases

Neural Introspection Gating for Adaptive KV-Cache Reuse in Vision-Language-Action Models

DGX agent

arXiv:2608.10824v1 Announce Type: cross Abstract: Vision-Language-Action(VLA) models map camera images and language instructions directly to motor commands through a single autoregressive transformer.

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

RLMOpt: Adaptive Prompt Optimization via Recursive Language Models

DGX agent

arXiv:2608.10471v1 Announce Type: new Abstract: Prompt optimizers automate the search for prompts that improve language-model performance, but existing methods rely on a predefined optimization proced

model-releasesarxiv-cs-ai
12 Aug 2026
Safety

SapiensID 2.0: Aligning Human Recognition Foundation Models with Human Perception

DGX agent

arXiv:2608.10497v1 Announce Type: new Abstract: While foundation models have significantly advanced human recognition across diverse modalities, they predominantly rely on static, geometric feature ex

safetyarxiv-cs-cv
12 Aug 2026
Model Releases

Situation Graph Prediction for User Perspective Modeling

DGX agent

arXiv:2602.13319v2 Announce Type: replace Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited by a data

model-releasesarxiv-cs-ai
12 Aug 2026
Agents

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

DGX agent

arXiv:2608.10538v1 Announce Type: new Abstract: Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essenti

agentsarxiv-cs-ai
12 Aug 2026
Model Releases

VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?

DGX agent

arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world

model-releasesarxiv-cs-cl
12 Aug 2026
Safety

A Structural Dynamics Graph World Model: Unified Modeling, Constrained Rollout, and Interpretable Calibration

DGX agent

arXiv:2608.08689v1 Announce Type: new Abstract: The state evolution of a complex system arises jointly from object laws, relational propagation, domain conservation, and unmodeled error. Forcing all s

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

Can Gemma and Qwen models catch hallucinations by looking at their own logprobs?

DGX agent

Hi! I'm really obsessed with LLM hallucinations for the last 6 days 😭 I started by designing system prompts to attack hallucinations but failed, obviously. Now I tried reading logprobs and... I think

model-releasesr-localllama
11 Aug 2026
Model Releases

CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift

DGX agent

arXiv:2608.07809v1 Announce Type: new Abstract: A world model is only useful for physical AI if it changes what the agent does, and only safe if it declines to do so when it is wrong. We study both ha

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models

DGX agent

arXiv:2608.08061v1 Announce Type: new Abstract: The key question in moral judgement is not simply whether someone chooses the 'right' answer, but how they decide what matters most when moral principle

model-releasesarxiv-cs-ai
11 Aug 2026
Research

Evaluating Dedicated Monolingual and Joint Multilingual Causal Models for Dravidian Languages

DGX agent

arXiv:2608.07727v1 Announce Type: new Abstract: Dravidian languages, mainly Tamil, Telugu, Kannada, and Malayalam make up only a small part of the data used to train multilingual language models, so i

researcharxiv-cs-cl
11 Aug 2026
Model Releases

Evaluating Generative Time-Series Models on Data with Point Masses

DGX agent

arXiv:2608.09692v1 Announce Type: cross Abstract: Many of the series that generative time-series models are benchmarked on place a large probability mass on a single value --- it does not rain, no rid

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model

DGX agent

arXiv:2606.14574v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments. While existing benchmarks

model-releasesarxiv-cs-ai
11 Aug 2026
Safety

Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space

DGX agent

arXiv:2608.07786v1 Announce Type: new Abstract: Open-weight large language models (LLMs) are increasingly developed through complex, multi-stage pipelines, leading to intricate lineage relationships t

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs

DGX agent

arXiv:2608.06967v1 Announce Type: new Abstract: Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide

model-releasesarxiv-cs-cl
10 Aug 2026
Model Releases

Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation

DGX agent

arXiv:2608.06718v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paraling

model-releasesarxiv-cs-cl
10 Aug 2026
Safety

How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots

DGX agent

arXiv:2608.06898v1 Announce Type: cross Abstract: Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public l

safetyarxiv-cs-cl
10 Aug 2026
Model Releases

Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows. Muse Glimmer delivers strong pe…

DGX agent

Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows. Muse Glimmer delivers strong performance on key agentic use cases and benchmarks compared w

model-releasesclem-delangue--x
10 Aug 2026
Model Releases

Lost in Interpolation: Why Predictive Feedback Fails in Diffusion Language Models

DGX agent

arXiv:2608.06529v1 Announce Type: new Abstract: Soft-masking accelerates the convergence of Masked Diffusion Language Models (MDLMs). Existing formulations build this blend with linear interpolation (

model-releasesarxiv-cs-cl
10 Aug 2026
Model Releases

Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE Models

DGX agent

arXiv:2608.06690v1 Announce Type: cross Abstract: Most language-model access controls regulate behavior while leaving the same computation available to every request. We study a different systems ques

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agents

DGX agent

arXiv:2608.06968v1 Announce Type: cross Abstract: Language-model agents act on state encodings of their environment, yet these are treated as interchangeable interfaces. Using pretrained language mode

model-releasesarxiv-cs-ai
10 Aug 2026
Tutorials

When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction

DGX agent

arXiv:2505.16170v4 Announce Type: replace Abstract: We study the internal mechanisms that govern when LLMs choose to retract wrong answers, i.e., spontaneously and immediately acknowledge errors in th

tutorialsarxiv-cs-cl
10 Aug 2026
Tutorials

Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models

DGX agent

arXiv:2608.07261v1 Announce Type: new Abstract: Large language models (LLMs) can solve complex multi-hop problems yet exhibit puzzling failures on simple two-hop queries: although a model may correctl

tutorialsarxiv-cs-cl
10 Aug 2026
Safety

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN

DGX agent

arXiv:2608.07267v1 Announce Type: new Abstract: Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies t

safetyarxiv-cs-ai
10 Aug 2026
Model Releases

GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

DGX agent

arXiv:2608.05948v1 Announce Type: new Abstract: Physics engines facilitate large-scale training and evaluation for embodied intelligence, while generative video world models are emerging as implicit s

model-releasesarxiv-cs-ai
7 Aug 2026
Model Releases

Layer-wise Positional Bias in Short-Context Language Modeling

DGX agent

arXiv:2601.04098v2 Announce Type: replace-cross Abstract: Transformer language models systematically prefer tokens at specific input positions regardless of semantic relevance---a phenomenon known as

model-releasesarxiv-cs-ai
7 Aug 2026
Model Releases

OpenAI puts the brakes on a new model because it’s supposedly too powerful

DGX agent

OpenAI says it is pausing 'internal activities' around an in-development AI model, Astra, because it doesn't yet meet new security standards the company is putting in place. The announcement follows i

model-releasesthe-verge-ai
7 Aug 2026
Model Releases

Above-ground Biomass Estimation with Geospatial Foundation Models

DGX agent

arXiv:2608.04792v1 Announce Type: new Abstract: Accurate estimation of Above-Ground Biomass (AGB) from satellite imagery is essential for the large-scale monitoring of carbon stocks, yet it remains a

model-releasesarxiv-cs-lg
6 Aug 2026
Model Releases

Enhancing Trustworthy Clinical Diagnosis Decision-Making in Large Language Models via Etiology-Aware Attention Supervision

DGX agent

arXiv:2508.00285v2 Announce Type: replace Abstract: Objective: Large Language Models (LLMs) have demonstrated strong capabilities in medical text understanding and generation. However, their trustwort

model-releasesarxiv-cs-cl
6 Aug 2026
Local Ai

Representational separation between unitary and channel quantum generative models via shared classical randomness at shallow depth

DGX agent

arXiv:2608.05110v1 Announce Type: cross Abstract: Near-term quantum hardware limits circuit depth and often imposes geometrically local connectivity for quantum generative models, restricting the outp

local-aiarxiv-cs-ai
6 Aug 2026
Model Releases

Sources: Alibaba plans to ask heavy commercial users of its next Qwen open model for a share of revenue; Moonshot's Kimi K3 requires up to a 30% revenue share (Reuters)

DGX agent

Reuters: Sources: Alibaba plans to ask heavy commercial users of its next Qwen open model for a share of revenue; Moonshot's Kimi K3 requires up to a 30% revenue share — Chinese technology giant Aliba

model-releasestechmeme
6 Aug 2026
Model Releases

AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction

DGX agent

arXiv:2608.03416v1 Announce Type: new Abstract: Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different

model-releasesarxiv-cs-ai
5 Aug 2026
Research

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

DGX agent

arXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how m

researcharxiv-cs-cl
5 Aug 2026
← Previous
1…4243444546…1248
Next →