AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,585 results
Model Releases

Learning Multi-Indicator Weights for Data Selection: A Joint Task-Model Adaptation Framework with Efficient Proxies

DGX agent

arXiv:2605.09665v1 Announce Type: cross Abstract: Data selection is a key component of efficient instruction tuning for large language models, as recent work has shown that data quality often matters

model-releasesarxiv-cs-ai
12 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Learning to Perceive 'Where': Spatial Pretext Tasks for Robust Self-Supervised Learning

DGX agent

arXiv:2605.09963v1 Announce Type: new Abstract: Existing self-supervised learning (SSL) methods primarily learn object-invariant representations but often neglect the spatial structure and relationshi

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Learning to Sparsify Stochastic Linear Bandits

DGX agent

arXiv:2605.10151v1 Announce Type: new Abstract: This paper addresses the problem of learning to sparsify stochastic linear bandits, where a decision-maker sequentially selects actions from a high-dime

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

LegalCiteBench: Evaluating Citation Reliability in Legal Language Models

DGX agent

arXiv:2605.10186v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly integrated into legal drafting and research workflows, where incorrect citations or fabricated precedent

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models

DGX agent

arXiv:2510.08592v3 Announce Type: replace-cross Abstract: Test-Time Scaling (TTS) improves LLM reasoning by exploring multiple candidate responses and then operating over this set to find the best out

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LEVI: Stronger Search Architectures Can Substitute for Larger LLMs in Evolutionary Search

DGX agent

arXiv:2605.09764v1 Announce Type: cross Abstract: LLM-guided evolutionary methods such as AlphaEvolve have proven effective in domains like math, systems research, and algorithmic discovery, but their

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LightAVSeg: Lightweight Audio-Visual Segmentation

DGX agent

arXiv:2605.08805v1 Announce Type: new Abstract: Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-mo

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities

DGX agent

arXiv:2605.10810v1 Announce Type: new Abstract: We introduce an automatically generated benchmark for predicting hidden text in technical papers. A paper supplies visible context X and a hidden contin

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

LimeCross: Context-Conditioned Layered Image Editing with Structural Consistency

DGX agent

arXiv:2605.10319v1 Announce Type: new Abstract: Layered image assets are widely used in real-world creative workflows, enabling non-destructive iteration and flexible re-composition. Recent advances i

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering

DGX agent

arXiv:2605.09384v1 Announce Type: cross Abstract: The reasoning gap between large and compact vision-language models (VLMs) limits the deployment of medical AI on portable clinical devices. Compact VL

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LiteParse is the best open-source, model-free document parser for AI agents. Run it over over 50+ document types, and it will parse dense pa…

DGX agent

LiteParse is the best open-source, model-free document parser for AI agents. Run it over over 50+ document types, and it will parse dense pages with complex text layouts and tables, and it will extrac

model-releasesjerry-liu--x
12 May 2026
Model Releases

LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments

DGX agent

arXiv:2605.10779v1 Announce Type: cross Abstract: The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content s

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?

DGX agent

arXiv:2605.08985v1 Announce Type: new Abstract: Visual encoding constitutes a major computational bottleneck in Multimodal Large Language Models (MLLMs), especially for high-resolution image inputs. T

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

LLiMba: Sardinian on a Single GPU -- Adapting a 3B Language Model to a Vanishing Romance Language

DGX agent

arXiv:2605.09015v1 Announce Type: new Abstract: Sardinian, a Romance language with roughly one million speakers, has minimal presence in modern NLP. Commercial services do not support it, and current

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

llm 0.32a2

DGX agent

Release: llm 0.32a2 A bunch of useful stuff in this LLM alpha, but the most important detail is this one: Most reasoning-capable OpenAI models now use the /v1/responses endpoint instead of /v1/chat/co

model-releasessimon-willison
12 May 2026
Model Releases

LLM Agents Already Know When to Call Tools -- Even Without Reasoning

DGX agent

arXiv:2605.09252v1 Announce Type: new Abstract: Tool-augmented LLM agents tend to call tools indiscriminately, even when the model can answer directly. Each unnecessary call wastes API fees and latenc

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LLM Jaggedness Unlocks Scientific Creativity

DGX agent

arXiv:2605.10574v1 Announce Type: new Abstract: As artificial intelligence advances, models are not improving uniformly. Instead, progress unfolds in a jagged fashion, with capabilities growing uneven

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LLM Translation of Compiler Intermediate Representation

DGX agent

arXiv:2605.08247v1 Announce Type: cross Abstract: GCC and LLVM underpin much of modern software infrastructure, relying on distinct Intermediate Representations (IRs) to drive optimizations and code g

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight

DGX agent

arXiv:2605.08321v1 Announce Type: cross Abstract: LLMs are increasingly capable of persuasion, which raises the question of how to protect users against manipulation. In a preregistered user study (N=

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LLM4Branch: Large Language Model for Discovering Efficient Branching Policies of Integer Programs

DGX agent

arXiv:2605.10401v1 Announce Type: new Abstract: Efficient branching policies are essential for accelerating Mixed Integer Linear Programming (MILP) solvers. Their design has long relied on hand-crafte

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LLMs with in-context learning for Algorithmic Theoretical Physics

DGX agent

arXiv:2605.08212v1 Announce Type: cross Abstract: There is an increasing number of algorithmic computations in theoretical physics. These, while conceptually simple, can nevertheless be time-consuming

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems

DGX agent

arXiv:2605.08305v1 Announce Type: cross Abstract: Large Language Model (LLM) systems have been the frontier of AI in many application domains, leading to new challenges and opportunities for hyperpara

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Locking Pretrained Weights via Deep Low-Rank Residual Distillation

DGX agent

arXiv:2605.10777v1 Announce Type: new Abstract: The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling the

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Log analysis is necessary for credible evaluation of AI agents

DGX agent

arXiv:2605.08545v1 Announce Type: new Abstract: Agent benchmarks typically report only final outcomes: pass or fail. This threatens evaluation credibility in three ways. First, scores may be inflated

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories

DGX agent

arXiv:2509.20909v2 Announce Type: replace Abstract: Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, where benchmark contamination can make memor

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models

DGX agent

arXiv:2605.08787v1 Announce Type: new Abstract: Recent advances in 3D medical vision-language models have enabled joint reasoning over volumetric images and text, showing strong performance in medical

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models

DGX agent

arXiv:2605.09879v1 Announce Type: new Abstract: While reasoning has become a central capability of large language models (LLMs), the reasoning patterns required for different scenarios are often misal

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

M^2E-UAV: A Benchmark and Analysis for Onboard Motion-on-Motion Event-Based Tiny UAV Detection

DGX agent

arXiv:2605.10496v1 Announce Type: new Abstract: Tiny UAV detection from an onboard event camera is difficult when the observer and target move at the same time. In this motion-on-motion regime, ego-mo

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

MaD Physics: Evaluating information seeking under constraints in physical environments

DGX agent

arXiv:2605.10820v1 Announce Type: new Abstract: Scientific discovery is fundamentally a resource-constrained process that requires navigating complex trade-offs between the quality and quantity of mea

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

DGX agent

arXiv:2605.08437v1 Announce Type: cross Abstract: Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to judge such argumen

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MAGS-SLAM: Monocular Multi-Agent Gaussian Splatting SLAM for Geometrically and Photometrically Consistent Reconstruction

DGX agent

arXiv:2605.10760v1 Announce Type: new Abstract: Collaborative photorealistic 3D reconstruction from multiple agents enables rapid large-scale scene capture for virtual production and cooperative multi

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs

DGX agent

arXiv:2605.08498v1 Announce Type: cross Abstract: We introduce MathConstraint, a hard, adaptive benchmark for evaluating the combinatorial reasoning capabilities of LLMs. We combine constraint satisfa

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics

DGX agent

arXiv:2602.02561v2 Announce Type: replace-cross Abstract: While the ecosystem of Lean and Mathlib has enjoyed celebrated success in formal mathematical reasoning with the help of large language models

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MBP-KT: Learning Global Collaborative Information from Meta-Behavioral Pattern for Enhanced Knowledge Tracing

DGX agent

arXiv:2605.08697v1 Announce Type: new Abstract: The emerging collaborative information-based knowledge tracing (KT) has been a promising way to enhance modeling of learners' knowledge states. The core

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MC-RFM: Geometry-Aware Few-Shot Adaptation via Mixed-Curvature Riemannian Flow Matching

DGX agent

arXiv:2605.08557v1 Announce Type: cross Abstract: Parameter-efficient adaptation of pretrained vision models is commonly performed through linear probes, prompts, low-rank updates, or lightweight resi

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MC^2: Monte Carlo Correction for Fast Elliptic PDE Solving

DGX agent

arXiv:2605.09288v1 Announce Type: cross Abstract: Partial differential equation (PDE) solvers underpin scientific computing, but real-world deployment is bounded by compute. Classical Monte Carlo solv

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments

DGX agent

arXiv:2605.09131v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has unified the interface between Large Language Models (LLMs) and external tools, yet a fundamental gap remains in how

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MDGYM: Benchmarking AI Agents on Molecular Simulations

DGX agent

arXiv:2605.08941v1 Announce Type: new Abstract: The promise of AI-driven scientific discovery hinges on whether AI agents can autonomously design and execute the computational workflows that underpin

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MDrive: Benchmarking Closed-Loop Cooperative Driving for End-to-End Multi-agent Systems

DGX agent

arXiv:2605.10904v1 Announce Type: new Abstract: Vehicle-to-Everything (V2X) communication has emerged as a promising paradigm for autonomous driving, enabling connected agents to share complementary p

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

DGX agent

arXiv:2510.22170v2 Announce Type: replace Abstract: Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure o

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare

DGX agent

arXiv:2605.08445v1 Announce Type: new Abstract: AI models are increasingly deployed in live clinical environments where they must perform reliably across complex, high-stakes workflows that standard t

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks

DGX agent

arXiv:2507.23511v3 Announce Type: replace-cross Abstract: While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. Th

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models

DGX agent

arXiv:2605.10002v1 Announce Type: new Abstract: Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet inco

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Medical Incident Causal Factors and Preventive Measures Generation Using Tag-based Example Selection in Few-shot Learning

DGX agent

arXiv:2605.10025v1 Announce Type: cross Abstract: In high-stakes domains such as healthcare, the reliability of Large Language Models (LLMs) is critical, particularly when generating clinical insights

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies

DGX agent

arXiv:2605.09661v1 Announce Type: cross Abstract: Large language models (LLMs) have saturated standard medical benchmarks that test factual recall, yet their ability to perform higher-order reasoning,

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Meet physics-intern🧑‍🎓, our agentic framework for theoretical physics. It takes Gemini 3.1 Pro from 17.7% to 31.4% on CritPt, a new SOTA o…

DGX agent

Meet physics-intern🧑‍🎓, our agentic framework for theoretical physics. It takes Gemini 3.1 Pro from 17.7% to 31.4% on CritPt, a new SOTA on one of the hardest benchmarks for LLMs. Theoretical physics

model-releasesclem-delangue--x
12 May 2026
Model Releases

MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents

DGX agent

arXiv:2605.09530v1 Announce Type: cross Abstract: As LLM-powered agents are increasingly deployed in edge-cloud environments, personalized memory has become a key enabler of long-term adaptation and u

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs

DGX agent

arXiv:2605.08374v1 Announce Type: new Abstract: Episodic memory allows LLM agents to accumulate and retrieve experience, but current methods treat each memory independently, i.e., evaluating retrieval

model-releasesarxiv-cs-ai
12 May 2026
← Previous
1…331332333334335…471
Next →