AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,620 results
29 May 2026

GroundAct: Can LLM Agents Ground Actions in Environmental States?

Model ReleasesDGX agent

arXiv:2508.05614v2 Announce Type: replace-cross Abstract: LLM agents achieve 85-96% success on tasks where instructions fully specify the action, but drop to 29-53% when action feasibility depends on

GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human

Model ReleasesDGX agent

arXiv:2605.28882v1 Announce Type: cross Abstract: With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important. However,

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

Model ReleasesDGX agent

arXiv:2605.29218v1 Announce Type: new Abstract: Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limi


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

GUITestScape: Towards Open-set Evaluation on Exploratory GUI Testing

Model ReleasesDGX agent

arXiv:2605.29532v1 Announce Type: cross Abstract: Exploratory GUI testing is a particularly demanding setting for MLLM agents: without predefined test scripts, an agent must autonomously navigate an a

Hallucination Detection-Guided Preference Optimization for Clinical Summarization

Model ReleasesDGX agent

arXiv:2605.28910v1 Announce Type: cross Abstract: Large language models (LLMs) have shown promise on summarization tasks, but they often produce hallucinations, which are unsupported or incorrect stat

Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching

Model ReleasesDGX agent

arXiv:2605.29055v1 Announce Type: new Abstract: Hallucination remains a major reliability barrier for production LLM systems, particularly in multi-agent pipelines where unsupported claims can propaga

Hands-on with Gemini Spark beta rolling out to AI Ultra subs: planned a birthday party from emails and calendar, but called a live-in boyfriend a 'close friend' (Reece Rogers/Wired)

Model ReleasesDGX agent

Reece Rogers / Wired: Hands-on with Gemini Spark beta rolling out to AI Ultra subs: planned a birthday party from emails and calendar, but called a live-in boyfriend a “close friend” — Google's new AI

Hear the architects of Gemini reflect on their journey to continue pushing the frontier of AI, on this episode of Release Notes. @JeffDean, …

Model ReleasesDGX agent

Hear the architects of Gemini reflect on their journey to continue pushing the frontier of AI, on this episode of Release Notes. @JeffDean, @koraykv, @OriolVinyalsML, and @NoamShazeer sit down on came

HEART-Bench: Do LLM Agents Exhibit Human-like Psychology?

Model ReleasesDGX agent

arXiv:2605.30058v1 Announce Type: new Abstract: While LLM agents have demonstrated remarkable task-oriented abilities such as planning, reasoning, and action, few works have treated them as complete h

Here's an extended edit of the quote that includes a following fragment where Andrew Macdonald called the trade 'harder to justify' - full, …

Model ReleasesDGX agent

Here's an extended edit of the quote that includes a following fragment where Andrew Macdonald called the trade 'harder to justify' - full, unedited transcript is here: https://gist.github.com/simonw/

Hista and Numca: Estimate State Value Effectively for LLM Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.29782v1 Announce Type: cross Abstract: Reinforcement learning (RL) refines large language models (LLMs) by directly optimizing model behavior through reward signals. While accurate state va

How Braintrust turns customer requests into code with Codex

Model ReleasesDGX agent

Braintrust leverages OpenAI's Codex model to automatically convert customer requests and natural language specifications into functional code, streamlining the software development process. This appli

How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions

Model ReleasesDGX agent

arXiv:2605.29442v1 Announce Type: cross Abstract: AI coding agents increasingly act directly within software environments, yet existing analyses of their failures rely on benchmark trajectories that m

How Far Ahead Do LLMs Plan? Uncovering the Latent Horizon in Chain-of-Thought Reasoning

Model ReleasesDGX agent

arXiv:2602.02103v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) reasoning has become a central mechanism for eliciting multi-step reasoning in Large Language Models (LLMs). Yet recent

How LoRA Remembers? A Parametric Memory Law for LLM Finetuning

Model ReleasesDGX agent

arXiv:2605.30260v1 Announce Type: cross Abstract: Large Language Models (LLMs) must continuously learn and update knowledge to remain effective in dynamic real-world environments. While Low-Rank Adapt

How lucky are you to have been born when and where you are? Had Opus 4.8 in Claude Code whip up a new visualization of all humans who ever l…

Model ReleasesDGX agent

How lucky are you to have been born when and where you are? Had Opus 4.8 in Claude Code whip up a new visualization of all humans who ever lived. In addition to being neat, it is an interesting test o

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

Model ReleasesDGX agent

arXiv:2605.30096v1 Announce Type: cross Abstract: Large language models (LLMs) can autonomously conduct multi-stage cyber attacks, but the consistency of their offensive behavior under repeated trials

I GOT THE DOMAIN! I FINALLY GOT IT!!!!!!!!!!1 🥳🎉 Paint​.NET is now at https://paint.net! Well, it will be just as soon as I push all the b…

Model ReleasesDGX agent

I GOT THE DOMAIN! I FINALLY GOT IT!!!!!!!!!!1 🥳🎉 Paint​.NET is now at https://paint.net! Well, it will be just as soon as I push all the buttons to migrate content and set up redirects from getpaint​.

- I still get nonstop questions about use cases, even though I think thinking about things in the use case framework is the wrong way of thi…

Model ReleasesDGX agent

- I still get nonstop questions about use cases, even though I think thinking about things in the use case framework is the wrong way of thinking about it - Enterprises are starting to see actual busi

iLoRA: Bayesian Low-Rank Adaptation with Latent Interaction Graphs for Microbiome Diagnosis

Model ReleasesDGX agent

arXiv:2605.30179v1 Announce Type: cross Abstract: Parameter-efficient adaptation has made LLMs practical for domain prediction, but standard LoRA still relies on a static low-rank update and does not

I’m going to let you in on a secret. I made a mistake once.* In August 2024.* It’s true. I actually got something wrong. I predicted that th…

Model ReleasesDGX agent

I’m going to let you in on a secret. I made a mistake once.* In August 2024.* It’s true. I actually got something wrong. I predicted that there would be a collapse of the AI bubble (which in my judgem

I'm suspicious of that that whole story about Uber blowing their AI budget and being disappointed in the results - I dug into it and it appe…

Model ReleasesDGX agent

Simon Willison expresses skepticism about reports claiming Uber overspent on AI and was disappointed with the results, indicating he investigated the story and found issues with its accuracy or framin

Improving Adversarial Robustness of Attribution via Implicit Regularization

Model ReleasesDGX agent

arXiv:2605.29983v1 Announce Type: cross Abstract: The adversarial robustness of attributions is a fundamental requirement for reliable explainability in deep learning, yet existing approaches typicall

Improving Full Waveform Inversion in Large Model Era

Model ReleasesDGX agent

arXiv:2603.00377v2 Announce Type: replace Abstract: Full Waveform Inversion (FWI) is a highly nonlinear and ill-posed problem that aims to recover subsurface velocity maps from surface-recorded seismi

In conversation with OpenAI’s @markchen90, Terence reflects on a future where AI reduces the cognitive friction of research, helps preserve …

Model ReleasesDGX agent

In conversation with OpenAI’s @markchen90, Terence reflects on a future where AI reduces the cognitive friction of research, helps preserve the paths behind discovery, and expands what mathematicians

Indexing the Unreadable: LLM-Native Recursive Construction and Search of Service Taxonomies

Model ReleasesDGX agent

arXiv:2605.29270v1 Announce Type: new Abstract: The era of the Internet of Agents (IoA) is taking shape: LLM agents are expected to fulfill user goals by orchestrating fast-growing populations of Mode

Inferring the Size of Large Language Models From Popular Text Memorization

Model ReleasesDGX agent

arXiv:2605.29223v1 Announce Type: new Abstract: The parameter counts of the most widely used large language models (LLMs) are often withheld by their developers, leaving model size -- a primary refere

Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles

Model ReleasesDGX agent

arXiv:2605.29473v1 Announce Type: cross Abstract: Language models are increasingly being deployed for conversational support in informal caregiving contexts, where interactions often extend beyond inf

InsightEval: An Expert-Curated Benchmark for Assessing Insight Discovery in LLM-Driven Data Agents

Model ReleasesDGX agent

arXiv:2511.22884v2 Announce Type: replace Abstract: Data analysis has become an indispensable part of scientific research. To discover the latent knowledge and insights hidden within massive datasets,

Interesting that the GPT-5 Pro series models have consistently been the best models for single-shot attempts at the hardest problems since l…

Model ReleasesDGX agent

Interesting that the GPT-5 Pro series models have consistently been the best models for single-shot attempts at the hardest problems since last summer. There has been no real competition in all that t

Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate

Model ReleasesDGX agent

arXiv:2605.29889v1 Announce Type: cross Abstract: Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases

Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives

Model ReleasesDGX agent

arXiv:2505.21627v4 Announce Type: replace-cross Abstract: State-of-the-art large language models require specialized hardware and substantial energy to operate. As a consequence, cloud-based services

I've largely switched over to using GPT-5.5 in recent weeks, which I like nearly as much as Opus 4.6 and 4.7, and is *very* reasonably price…

Model ReleasesDGX agent

Jeremy Howard expresses positive views on GPT-5.5, stating he has recently switched to using it as his primary model and finds it nearly comparable to Anthropic's Opus 4.6 and 4.7 while offering signi

Just dropped 🧑‍🍳 Fixed version of DeepSeek-V4-Pro-NVFP4 by @NVIDIAAI https://huggingface.co/nvidia/DeepSeek-V4-Pro-NVFP4

Model ReleasesDGX agent

NVIDIA has released a corrected version of DeepSeek-V4-Pro-NVFP4, a quantized model variant optimized for NVIDIA hardware using NV-FP4 (NVIDIA's 4-bit floating-point format). The model is available on

K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance

Model ReleasesDGX agent

arXiv:2605.29523v1 Announce Type: new Abstract: Large Language Models (LLMs) have advanced financial automation through Retrieval-Augmented Generation (RAG), yet hallucinations remain a critical barri

KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing

Model ReleasesDGX agent

arXiv:2605.29524v1 Announce Type: cross Abstract: Relay and reseller APIs increasingly intermediate access to large language models (LLMs), but users have no direct way to verify that a claimed endpoi

Kernel-based potential mean-field games with unbiased random Fourier U-statistics

Model ReleasesDGX agent

arXiv:2605.29371v1 Announce Type: cross Abstract: We study the subclass of potential mean-field games in which the running interaction cost and the terminal target cost are both expressed through repr

Kernel Renormalization in Bayesian Deep Neural Networks: the Equivalent Wishart Ansatz in the Proportional Regime

Model ReleasesDGX agent

arXiv:2605.29684v1 Announce Type: new Abstract: The scaling limit where both the size of the training set P and the width N of a deep neural network grow at the same rate, the so-called proportional-w

Knowledge Offloading: Decomposing LLMs into Sparse Backbones and Memory Modules

Model ReleasesDGX agent

arXiv:2605.29075v1 Announce Type: new Abstract: LLMs encode both general capabilities and domain-specific knowledge in a single set of parameters. We ask whether this capacity can be reorganized: keep

Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models

Model ReleasesDGX agent

arXiv:2605.29459v1 Announce Type: new Abstract: Large language models route every input through a learned embedding table of shape |V| x d_model, consuming hundreds of millions to billions of trainabl

Label-Free Reinforcement Learning via Cross-Model Entropy

Model ReleasesDGX agent

arXiv:2605.29009v1 Announce Type: cross Abstract: Post-training large language models with reinforcement learning is bottlenecked by the reward signal. Existing approaches require either ground-truth

Label Over Logic? How Source Cues Bias Human Fallacy Judgments More Than LLMs

Model ReleasesDGX agent

arXiv:2605.29928v1 Announce Type: cross Abstract: As AI-generated and AI-assisted content floods online spaces, source labels attached to such content can distort human reasoning judgments, with downs

Large language models reorganize representational geometry during in-context learning

Model ReleasesDGX agent

arXiv:2605.28854v1 Announce Type: new Abstract: Large language models (LLMs) exhibit remarkable flexibility: they can adapt to novel tasks from in-context examples without any parameter updates, a cap

Latent Performance Profiling of Large Language Models

Model ReleasesDGX agent

arXiv:2605.30018v1 Announce Type: new Abstract: Large language models (LLMs) frequently achieve impressive scores on standardized benchmarks, yet accuracy alone offers a limited view of their capabili

Leak@k: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding

Model ReleasesDGX agent

arXiv:2511.04934v3 Announce Type: replace Abstract: Unlearning in large language models (LLMs) is critical for regulatory compliance and for building ethical generative AI systems that avoid producing

Learning Design Skills as Memory Policies for Agentic Photonic Inverse Design

Model ReleasesDGX agent

arXiv:2605.29421v1 Announce Type: new Abstract: Photonic crystal fiber (PCF) inverse design remains challenging because candidate geometries must satisfy coupled optical targets under expensive electr

Learning Representations from 3D Gaussian Splats

Model ReleasesDGX agent

arXiv:2605.29549v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) is a recent approach for scene rendering. Although primarily designed for view synthesis, its potential for scene understan

Learning to Extrapolate to New Tasks: A Relational Approach to Task Extrapolation

Model ReleasesDGX agent

arXiv:2605.30132v1 Announce Type: new Abstract: Modern learning systems excel at interpolation but struggle to generalize to unseen tasks outside the training distribution's support. This failure occu

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2602.10388v3 Announce Type: replace-cross Abstract: The diversity of post-training data is critical for effective downstream performance in large language models (LLMs). Many existing approaches

Leveraging Routing Dynamics in Mixture-of-Experts Models for Efficient Language Adaptation

Model ReleasesDGX agent

arXiv:2605.29714v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models are widely used to scale language models, yet their expert routing behavior and adaptation in a multilingual setting rem

libhmm: A Modern C++20 Library for Hidden Markov Models with Correct MLE Emission M-Steps

Model ReleasesDGX agent

arXiv:2605.29208v1 Announce Type: cross Abstract: We describe libhmm, a C++20 library for Hidden Markov Model parameter estimation, sequence decoding, and model selection. libhmm addresses two gaps in

LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents

Model ReleasesDGX agent

arXiv:2605.29559v1 Announce Type: new Abstract: Mastering terminal environments requires language agents capable of multi-step planning, feedback-grounded execution, and dynamic state adaptation. Howe

LiveSVG: Zero-Shot SVG Animation via Video Generation

Model ReleasesDGX agent

arXiv:2605.30174v1 Announce Type: new Abstract: We introduce LiveSVG, a zero-shot approach for generating Scalable Vector Graphics (SVG) animations using video diffusion models. Current SVG animation

llama.cpp now has an official website: https://llama.app Our goal is to make local AI accessible to everyone, and improving the user experie…

Model ReleasesDGX agent

llama.cpp now has an official website: https://llama.app Our goal is to make local AI accessible to everyone, and improving the user experience is a big part of that. On the new landing page you’ll fi

LLM-Evolved Domain-Independent Heuristics for Symbolic AI Planning

Model ReleasesDGX agent

arXiv:2605.29649v1 Announce Type: new Abstract: Heuristic search is the dominant paradigm in symbolic AI planning, and the strongest heuristics are the result of decades of work by planning researcher

LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video Generation

Model ReleasesDGX agent

arXiv:2510.26412v3 Announce Type: replace-cross Abstract: Recent advances in text-to-video generation have achieved impressive performance on short clips, yet evaluating long-form generation under com

LogDx-CI: Benchmarking Log Reduction Tools for LLM Root-Cause Diagnosis

Model ReleasesDGX agent

arXiv:2605.28876v1 Announce Type: cross Abstract: CI failure logs are large (median 5k lines, max 200k in this corpus) and noisy. Coding agents that try to debug them depend on an upstream tool to red

Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference

Model ReleasesDGX agent

arXiv:2510.24606v2 Announce Type: replace Abstract: The quadratic cost of attention limits the scalability of long-context LLMs, especially under limited hardware memory budgets. While attention is of

Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selection

Model ReleasesDGX agent

arXiv:2605.30274v1 Announce Type: cross Abstract: Document-level translation remains one of the most challenging tasks for large language models, which are constrained by limited context windows that

LoopFM: Learning frOm HistOrical RePresentations of Foundation Model for Recommendation

Model ReleasesDGX agent

arXiv:2605.29280v1 Announce Type: cross Abstract: Knowledge distillation (KD) transfers a single scalar prediction from a large foundation model (FM) to compact vertical models (VMs), suffering from d

← Previous
1…194195196197198…377
Next →