AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,585 results
20 May 2026

Sequential Consensus for Multi-Agent LLM Debates: A Wald-SPRT compute governor with calibration-based failure detection

Model ReleasesDGX agent

arXiv:2605.19193v1 Announce Type: new Abstract: Multi-agent LLM debate improves factuality and reasoning, but most recipes pick a fixed round count, over-spending on easy items and under-spending on h

Sharper Bounds for Chebyshev Moment Matching, with Applications

Model ReleasesDGX agent

arXiv:2408.12385v3 Announce Type: replace-cross Abstract: We study the problem of approximately recovering a probability distribution given noisy measurements of its Chebyshev polynomial moments. This

Simply Stabilizing the Loop via Fully Looped Transformer

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.18797v1 Announce Type: cross Abstract: Scaling model performance typically requires increasing model size. Looped Transformer offers a compelling alternative by iteratively reusing the same

SLoW: Select Low-frequency Words! Automatic Dictionary Selection for Translation on Large Language Models

Model ReleasesDGX agent

arXiv:2507.18902v2 Announce Type: replace Abstract: There are more than 7,000 languages around the world, and current Large Language Models (LLMs) only support hundreds of languages. Dictionary-based

Smooth Piecewise Cutting for Neural Operator to Handle Discontinuities and Sharp Transitions

Model ReleasesDGX agent

arXiv:2605.19823v1 Announce Type: cross Abstract: Neural operators have achieved strong performance in learning solution operators of partial differential equations (PDEs), but their inherently contin

Sonar-TS: Search-Then-Verify Natural Language Querying for Time Series Databases

Model ReleasesDGX agent

arXiv:2602.17001v2 Announce Type: replace Abstract: Natural Language Querying for Time Series Databases (NLQ4TSDB) aims to assist non-expert users retrieve meaningful events, intervals, and summaries

SpecX: A Large-Scale Benchmark for Multi-Modal Spectroscopy and Cross-Paradigm Evaluation

Model ReleasesDGX agent

arXiv:2605.18791v1 Announce Type: cross Abstract: Existing spectral benchmarks are limited in scale, modality alignment, and evaluation scope, and typically focus on either specialized models or multi

STAR-PolyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision

Model ReleasesDGX agent

arXiv:2605.19338v1 Announce Type: cross Abstract: Frontier AI models and multi-agent systems have led to significant improvements in mathematical reasoning. However, for problems requiring extended, l

STAR: Semantic-Tuned and Tail-Adaptive Retriever for Graph-Augmented Generation

Model ReleasesDGX agent

arXiv:2605.18765v1 Announce Type: cross Abstract: To augment Large Language Models (LLMs) for multi-hop question answering, a mainstream solution within Graph Retrieval Augmented Generation (GraphRAG)

Stochastic Gradient Variational Inference with Price's Gradient Estimator from Bures-Wasserstein to Parameter Space

Model ReleasesDGX agent

arXiv:2602.18718v2 Announce Type: replace-cross Abstract: For approximating a target distribution given only its unnormalized log-density, stochastic gradient-based variational inference (VI) algorith

Streamlined Constraint Reasoning via CNN Pattern Recognition on Enumerated Solutions

Model ReleasesDGX agent

arXiv:2605.19895v1 Announce Type: new Abstract: Constraint programming practitioners accelerate hard problems through a layered set of techniques applied in order of risk. Standard hardening (symmetry

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding

Model ReleasesDGX agent

arXiv:2605.19866v1 Announce Type: new Abstract: Vision-Language Models (VLMs) parse documents end-to-end but frequently break down on layouts unlike those seen in training. We attribute this to a two-

Subagents running locally and simultaneously on MacBook Pro M5 with Codex CLI + @lmstudio to review code and find bugs using Qwen 3.6 Powere…

Model ReleasesDGX agent

Subagents running locally and simultaneously on MacBook Pro M5 with Codex CLI + @lmstudio to review code and find bugs using Qwen 3.6 Powered by the updated MLX engine with batching in beta in the app

SynGR: Unleashing the Potential of Cross-Modal Synergy for Generative Recommendation

Model ReleasesDGX agent

arXiv:2605.18920v1 Announce Type: cross Abstract: Generative Recommendation (GR) has emerged as a promising paradigm by formulating item recommendation as a sequence-to-sequence generation task over i

Synthesis and Evaluation of Long-term History-aware Medical Dialogue

Model ReleasesDGX agent

arXiv:2605.19766v1 Announce Type: cross Abstract: An effective healthcare agent must be able to recall and reason over a patient's longitudinal medical history. However, the absence of datasets with r

Synthetic Data Generation for Brain-Computer Interfaces: Overview, Benchmarking, and Future Directions

Model ReleasesDGX agent

arXiv:2603.12296v2 Announce Type: replace-cross Abstract: Deep learning has achieved transformative performance across diverse domains, largely driven by large-scale and high-quality training data. In

TADA! Tuning Audio Diffusion Models through Activation Steering

Model ReleasesDGX agent

arXiv:2602.11910v2 Announce Type: replace-cross Abstract: Audio diffusion models can synthesize high-fidelity music from text, yet achieving fine-grained control over specific musical attributes remai

Tail Annealing for Heavy-Tailed Flow Matching

Model ReleasesDGX agent

arXiv:2605.20068v1 Announce Type: cross Abstract: Standard generative models struggle with heavy-tailed data: Lipschitz architectures cannot produce power-law tails from Gaussian noise, and interpolat

Taming the Thinker: Conditional Entropy Shaping for Adaptive LLM Reasoning

Model ReleasesDGX agent

arXiv:2605.19358v1 Announce Type: new Abstract: Entropy-based deep reasoning has emerged as a promising direction for improving the reasoning capabilities of Large Language Models (LLMs), but existing

Target-Aligned Reinforcement Learning

Model ReleasesDGX agent

arXiv:2603.29501v2 Announce Type: replace-cross Abstract: Many value-based deep reinforcement learning algorithms rely on target networks - lagged copies of the online network - to stabilize training.

Targeted Downstream-Agnostic Attack

Model ReleasesDGX agent

arXiv:2605.19446v1 Announce Type: cross Abstract: Recently, pre-trained encoders have gained widespread use due to their strong capability in representation extraction. However, they are vulnerable to

TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards

Model ReleasesDGX agent

arXiv:2605.19320v1 Announce Type: new Abstract: Faithful text rendering remains a persistent weakness of large text-to-image generative models, as it requires both semantic instruction following and f

The Annotation Scarcity Paradox in Low-Resource NLP Evaluation: A Decade of Acceleration and Emerging Constraints

Model ReleasesDGX agent

arXiv:2605.19066v1 Announce Type: new Abstract: Over the past decade, low-resource natural language processing (NLP) has experienced explosive growth, propelled by cross-lingual transfer, massively mu

The Evaluation Game: Beyond Static LLM Benchmarking

Model ReleasesDGX agent

arXiv:2605.19377v1 Announce Type: cross Abstract: As jailbreaks, adversarially crafted inputs that bypass safety constraints, continue to be discovered in Large Language Models, practitioners increasi

The frontier is still jagged though (here is Gemini 3.5 Flash messing up counting letters in words) https://x.com/RRiscio37389/status/205719…

Model ReleasesDGX agent

The frontier is still jagged though (here is Gemini 3.5 Flash messing up counting letters in words) https://x.com/RRiscio37389/status/2057193260745883962?s=20 @emollick Proof: https://gemini.google.co

The future of biology shouldn’t stay behind black-box APIs. Especially when it touches personal health. Whether you’re @bryan_johnson measur…

Model ReleasesDGX agent

The future of biology shouldn’t stay behind black-box APIs. Especially when it touches personal health. Whether you’re @bryan_johnson measuring every biomarker, or @sytses openly sharing and analyzing

The Growing Pains of Frontier Models: When Leaderboards Stop Separating and What to Measure Next

Model ReleasesDGX agent

arXiv:2605.18840v1 Announce Type: cross Abstract: Leaderboards rank frontier models on independent axes but do not reveal whether capabilities reinforce or trade off across releases -- and at the fron

The proof came from a general-purpose reasoning model, not a system built specifically to solve math problems or this problem in particular,…

Model ReleasesDGX agent

The proof came from a general-purpose reasoning model, not a system built specifically to solve math problems or this problem in particular, and represents an important milestone for the math and AI c

The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility

Model ReleasesDGX agent

arXiv:2605.19537v1 Announce Type: new Abstract: Progress in LLMs is increasingly measured through standardized benchmarks, where state-of-the-art improvements are often separated by fractions of a per

The World Won't Stay Still: Programmable Evolution for Agent Benchmarks

Model ReleasesDGX agent

arXiv:2603.05910v2 Announce Type: replace Abstract: LLM-powered tool-calling agents fulfill user requests by interacting with environments, querying data, and invoking tools in a multi-turn process. Y

Theory-optimal Quantization Based on Flatness

Model ReleasesDGX agent

arXiv:2605.18800v1 Announce Type: cross Abstract: Post-training quantization has emerged as a widely adopted technique for compressing and accelerating the inference of Large Language Models (LLMs). T

This result points to something larger: AI systems are becoming capable of holding together long, difficult chains of reasoning, connecting …

Model ReleasesDGX agent

This result points to something larger: AI systems are becoming capable of holding together long, difficult chains of reasoning, connecting ideas across distant fields, and surfacing paths researchers

TideGS: Scalable Training of Over One Billion 3D Gaussian Splatting Primitives via Out-of-Core Optimization

Model ReleasesDGX agent

arXiv:2605.20150v1 Announce Type: new Abstract: Training 3D Gaussian Splatting (3DGS) at billion-primitive scale is fundamentally memory-bound: each Gaussian primitive carries a large attribute vector

Time-optimal neural feedback control of nilpotent systems as a binary classification problem

Model ReleasesDGX agent

arXiv:2503.17581v2 Announce Type: replace-cross Abstract: A computational method for the synthesis of time-optimal feedback control laws for linear nilpotent systems is proposed. The method is based o

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

Model ReleasesDGX agent

arXiv:2605.19196v1 Announce Type: new Abstract: Deep research agents increasingly automate complex information-seeking tasks, producing evidence-grounded reports via multi-step reasoning, tool use, an

To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents

Model ReleasesDGX agent

arXiv:2605.18882v1 Announce Type: cross Abstract: LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, six models

Today, we share a breakthrough on the planar unit distance problem, a famous open question first posed by Paul Erdős in 1946. For nearly 80 …

Model ReleasesDGX agent

Today, we share a breakthrough on the planar unit distance problem, a famous open question first posed by Paul Erdős in 1946. For nearly 80 years, mathematicians believed the best possible solutions l

Toto 2.0: Time Series Forecasting Enters the Scaling Era

Model ReleasesDGX agent

arXiv:2605.20119v1 Announce Type: cross Abstract: We show that time series foundation models scale: a single training recipe produces reliable forecast-quality improvements from 4M to 2.5B parameters.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs

Model ReleasesDGX agent

arXiv:2605.19528v1 Announce Type: new Abstract: 3D localization in Multimodal Large Language Models (MLLMs), including 3D object detection and 3D visual grounding, is fundamentally limited by camera i

Towards Consistent Detection of Cognitive Distortions: LLM-Based Annotation and Dataset-Agnostic Evaluation

Model ReleasesDGX agent

arXiv:2511.01482v2 Announce Type: replace Abstract: Text-based automated Cognitive Distortion detection is a challenging task due to its subjective nature, with low agreement scores observed even amon

Towards Trust Calibration in Socially Interactive Agents: Investigating Gendered Multimodal Behaviors Generation with LLMs

Model ReleasesDGX agent

arXiv:2605.19798v1 Announce Type: new Abstract: As Socially Interactive Agents (SIAs) become increasingly integrated into daily life, the ability to calibrate user trust to an agent's actual capabilit

Training Neural Networks with Optimal Double-Bayesian Learning

Model ReleasesDGX agent

arXiv:2605.20009v1 Announce Type: cross Abstract: Backpropagation with gradient descent is a common optimization strategy employed by most neural network architectures in machine learning. However, fi

Trajectory Planning and Control near the Limits: an Open Experimental Benchmark on the RoboRacer Platform

Model ReleasesDGX agent

arXiv:2605.19881v1 Announce Type: new Abstract: We present a modular framework to benchmark new and existing methods for trajectory planning and control in high-acceleration maneuvers that push autono

TravExplorer: Cross-Floor Embodied Exploration via Traversability-Aware 3-D Planning

Model ReleasesDGX agent

arXiv:2605.19958v1 Announce Type: new Abstract: Zero-shot Object Navigation (ZSON) has shown promise for open-vocabulary target search in unseen environments, yet most existing systems remain tied to

.@trq212 is a builder's builder. After research at MIT, exiting his company, raising millions for another, and exploring research threads as…

Model ReleasesDGX agent

.@trq212 is a builder's builder. After research at MIT, exiting his company, raising millions for another, and exploring research threads as an SPC member, he's now on the team building Claude Code. H

Trust or Abstain? A Self-Aware RAG Approach

Model ReleasesDGX agent

arXiv:2605.18792v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) improves large language models (LLMs) by incorporating external evidence, but it also introduces knowledge confli

TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing

Model ReleasesDGX agent

arXiv:2605.18859v1 Announce Type: cross Abstract: LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single user reque

Two research papers describe how Google's Co-Scientist and nonprofit FutureHouse's AI tools can succeed at drug-retargeting tasks by forming hypotheses (John Timmer/Ars Technica)

Model ReleasesDGX agent

John Timmer / Ars Technica: Two research papers describe how Google's Co-Scientist and nonprofit FutureHouse's AI tools can succeed at drug-retargeting tasks by forming hypotheses — Both tools generat

Unified Deployment-Aware Evaluation of Open Reasoning Language Models

Model ReleasesDGX agent

arXiv:2604.07035v2 Announce Type: replace Abstract: Open reasoning language models are often compared under mixed sample sizes, partially standardized prompts, and accuracy-centered summaries, which m

Unlocking the Potential of Continual Model Merging: An ODE Perspective

Model ReleasesDGX agent

arXiv:2605.19409v1 Announce Type: cross Abstract: Continual Model Merging (CMM) enables rapid customization of foundation models across sequentially arriving tasks, offering a scalable alternative to

Very interesting results from this NanoGPT-Bench eval. There is so much talk about self-improving agents. But can coding agents do real AI R…

Model ReleasesDGX agent

Very interesting results from this NanoGPT-Bench eval. There is so much talk about self-improving agents. But can coding agents do real AI R&D? @IntologyAI reports that Codex, Claude Code, and Autores

ViroGym: Realistic Large-Scale Benchmarks for Evaluating Viral Proteins

Model ReleasesDGX agent

arXiv:2603.06740v2 Announce Type: replace-cross Abstract: Protein language models (pLMs) have shown strong potential for zero-shot prediction of missense variant effects, yet systematic benchmarking o

Vision Harnessing Agent for Open Ad-hoc Segmentation

Model ReleasesDGX agent

arXiv:2605.19410v1 Announce Type: new Abstract: Segmentation has become easy when the concept is known, requiring retrieval of a learned visual grounding from text. It remains hard for open ad-hoc con

wait… did Cohere just release Command A+ models under Apache 2.0 for the first time ever?! 🙊 welcome to Europe! 🤗

Model ReleasesDGX agent

wait… did Cohere just release Command A+ models under Apache 2.0 for the first time ever?! 🙊 welcome to Europe! 🤗 Introducing: Cohere Command A+ We’ve created our most powerful LLM yet, optimized it t

WARC-Bench: Web Archive Based Benchmark for GUI Subtask Executions

Model ReleasesDGX agent

arXiv:2510.09872v2 Announce Type: replace-cross Abstract: Training web agents to navigate complex, real-world websites requires them to master extit{subtasks} - short-horizon interactions on multiple

We partnered with artists, designers, and builders to create new AI tools that solve real problems in their creative workflows. Here’s what’…

Model ReleasesDGX agent

We partnered with artists, designers, and builders to create new AI tools that solve real problems in their creative workflows. Here’s what’s new: — Introducing Google Pics in @GoogleWorkspace: A bran

We're building Gemini for Science with and for the scientific community. In collaboration with 100+ institutions and a trusted tester commun…

Model ReleasesDGX agent

We're building Gemini for Science with and for the scientific community. In collaboration with 100+ institutions and a trusted tester community that ranges from PhD students to Nobel laureates, we wan

We're excited to be an official shoutout at the Google I/O Developer Keynote 🔥 @llama_index is building the document infrastructure for AI …

Model ReleasesDGX agent

We're excited to be an official shoutout at the Google I/O Developer Keynote 🔥 @llama_index is building the document infrastructure for AI agents, and we plan to integrate even more heavily with both

We’re expanding our partnership with @SpaceX, and will be scaling up on GB200 capacity in Colossus 2 throughout June. Appreciate @elonmusk a…

Model ReleasesDGX agent

We’re expanding our partnership with @SpaceX, and will be scaling up on GB200 capacity in Colossus 2 throughout June. Appreciate @elonmusk and the team helping us find good homes for the Claudes. In t

What Do Evolutionary Coding Agents Evolve?

Model ReleasesDGX agent

arXiv:2605.20086v1 Announce Type: cross Abstract: Recent work pairs LLMs with evolutionary search to iteratively generate, modify, and select code using task-specific feedback. These systems have prod

← Previous
1…230231232233234…377
Next →