AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,577 results
Model Releases

LimeCross: Context-Conditioned Layered Image Editing with Structural Consistency

DGX agent

arXiv:2605.10319v1 Announce Type: new Abstract: Layered image assets are widely used in real-world creative workflows, enabling non-destructive iteration and flexible re-composition. Recent advances i

model-releasesarxiv-cs-cv
12 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering

DGX agent

arXiv:2605.09384v1 Announce Type: cross Abstract: The reasoning gap between large and compact vision-language models (VLMs) limits the deployment of medical AI on portable clinical devices. Compact VL

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LiteParse is the best open-source, model-free document parser for AI agents. Run it over over 50+ document types, and it will parse dense pa…

DGX agent

LiteParse is the best open-source, model-free document parser for AI agents. Run it over over 50+ document types, and it will parse dense pages with complex text layouts and tables, and it will extrac

model-releasesjerry-liu--x
12 May 2026
Model Releases

LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments

DGX agent

arXiv:2605.10779v1 Announce Type: cross Abstract: The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content s

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?

DGX agent

arXiv:2605.08985v1 Announce Type: new Abstract: Visual encoding constitutes a major computational bottleneck in Multimodal Large Language Models (MLLMs), especially for high-resolution image inputs. T

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

LLiMba: Sardinian on a Single GPU -- Adapting a 3B Language Model to a Vanishing Romance Language

DGX agent

arXiv:2605.09015v1 Announce Type: new Abstract: Sardinian, a Romance language with roughly one million speakers, has minimal presence in modern NLP. Commercial services do not support it, and current

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

llm 0.32a2

DGX agent

Release: llm 0.32a2 A bunch of useful stuff in this LLM alpha, but the most important detail is this one: Most reasoning-capable OpenAI models now use the /v1/responses endpoint instead of /v1/chat/co

model-releasessimon-willison
12 May 2026
Model Releases

LLM Agents Already Know When to Call Tools -- Even Without Reasoning

DGX agent

arXiv:2605.09252v1 Announce Type: new Abstract: Tool-augmented LLM agents tend to call tools indiscriminately, even when the model can answer directly. Each unnecessary call wastes API fees and latenc

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LLM Jaggedness Unlocks Scientific Creativity

DGX agent

arXiv:2605.10574v1 Announce Type: new Abstract: As artificial intelligence advances, models are not improving uniformly. Instead, progress unfolds in a jagged fashion, with capabilities growing uneven

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LLM Translation of Compiler Intermediate Representation

DGX agent

arXiv:2605.08247v1 Announce Type: cross Abstract: GCC and LLVM underpin much of modern software infrastructure, relying on distinct Intermediate Representations (IRs) to drive optimizations and code g

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight

DGX agent

arXiv:2605.08321v1 Announce Type: cross Abstract: LLMs are increasingly capable of persuasion, which raises the question of how to protect users against manipulation. In a preregistered user study (N=

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LLM4Branch: Large Language Model for Discovering Efficient Branching Policies of Integer Programs

DGX agent

arXiv:2605.10401v1 Announce Type: new Abstract: Efficient branching policies are essential for accelerating Mixed Integer Linear Programming (MILP) solvers. Their design has long relied on hand-crafte

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LLMs with in-context learning for Algorithmic Theoretical Physics

DGX agent

arXiv:2605.08212v1 Announce Type: cross Abstract: There is an increasing number of algorithmic computations in theoretical physics. These, while conceptually simple, can nevertheless be time-consuming

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems

DGX agent

arXiv:2605.08305v1 Announce Type: cross Abstract: Large Language Model (LLM) systems have been the frontier of AI in many application domains, leading to new challenges and opportunities for hyperpara

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Locking Pretrained Weights via Deep Low-Rank Residual Distillation

DGX agent

arXiv:2605.10777v1 Announce Type: new Abstract: The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling the

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Log analysis is necessary for credible evaluation of AI agents

DGX agent

arXiv:2605.08545v1 Announce Type: new Abstract: Agent benchmarks typically report only final outcomes: pass or fail. This threatens evaluation credibility in three ways. First, scores may be inflated

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories

DGX agent

arXiv:2509.20909v2 Announce Type: replace Abstract: Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, where benchmark contamination can make memor

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models

DGX agent

arXiv:2605.08787v1 Announce Type: new Abstract: Recent advances in 3D medical vision-language models have enabled joint reasoning over volumetric images and text, showing strong performance in medical

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models

DGX agent

arXiv:2605.09879v1 Announce Type: new Abstract: While reasoning has become a central capability of large language models (LLMs), the reasoning patterns required for different scenarios are often misal

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

M^2E-UAV: A Benchmark and Analysis for Onboard Motion-on-Motion Event-Based Tiny UAV Detection

DGX agent

arXiv:2605.10496v1 Announce Type: new Abstract: Tiny UAV detection from an onboard event camera is difficult when the observer and target move at the same time. In this motion-on-motion regime, ego-mo

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

MaD Physics: Evaluating information seeking under constraints in physical environments

DGX agent

arXiv:2605.10820v1 Announce Type: new Abstract: Scientific discovery is fundamentally a resource-constrained process that requires navigating complex trade-offs between the quality and quantity of mea

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

DGX agent

arXiv:2605.08437v1 Announce Type: cross Abstract: Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to judge such argumen

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MAGS-SLAM: Monocular Multi-Agent Gaussian Splatting SLAM for Geometrically and Photometrically Consistent Reconstruction

DGX agent

arXiv:2605.10760v1 Announce Type: new Abstract: Collaborative photorealistic 3D reconstruction from multiple agents enables rapid large-scale scene capture for virtual production and cooperative multi

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs

DGX agent

arXiv:2605.08498v1 Announce Type: cross Abstract: We introduce MathConstraint, a hard, adaptive benchmark for evaluating the combinatorial reasoning capabilities of LLMs. We combine constraint satisfa

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics

DGX agent

arXiv:2602.02561v2 Announce Type: replace-cross Abstract: While the ecosystem of Lean and Mathlib has enjoyed celebrated success in formal mathematical reasoning with the help of large language models

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MBP-KT: Learning Global Collaborative Information from Meta-Behavioral Pattern for Enhanced Knowledge Tracing

DGX agent

arXiv:2605.08697v1 Announce Type: new Abstract: The emerging collaborative information-based knowledge tracing (KT) has been a promising way to enhance modeling of learners' knowledge states. The core

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MC-RFM: Geometry-Aware Few-Shot Adaptation via Mixed-Curvature Riemannian Flow Matching

DGX agent

arXiv:2605.08557v1 Announce Type: cross Abstract: Parameter-efficient adaptation of pretrained vision models is commonly performed through linear probes, prompts, low-rank updates, or lightweight resi

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MC^2: Monte Carlo Correction for Fast Elliptic PDE Solving

DGX agent

arXiv:2605.09288v1 Announce Type: cross Abstract: Partial differential equation (PDE) solvers underpin scientific computing, but real-world deployment is bounded by compute. Classical Monte Carlo solv

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments

DGX agent

arXiv:2605.09131v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has unified the interface between Large Language Models (LLMs) and external tools, yet a fundamental gap remains in how

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MDGYM: Benchmarking AI Agents on Molecular Simulations

DGX agent

arXiv:2605.08941v1 Announce Type: new Abstract: The promise of AI-driven scientific discovery hinges on whether AI agents can autonomously design and execute the computational workflows that underpin

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MDrive: Benchmarking Closed-Loop Cooperative Driving for End-to-End Multi-agent Systems

DGX agent

arXiv:2605.10904v1 Announce Type: new Abstract: Vehicle-to-Everything (V2X) communication has emerged as a promising paradigm for autonomous driving, enabling connected agents to share complementary p

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

DGX agent

arXiv:2510.22170v2 Announce Type: replace Abstract: Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure o

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare

DGX agent

arXiv:2605.08445v1 Announce Type: new Abstract: AI models are increasingly deployed in live clinical environments where they must perform reliably across complex, high-stakes workflows that standard t

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks

DGX agent

arXiv:2507.23511v3 Announce Type: replace-cross Abstract: While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. Th

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models

DGX agent

arXiv:2605.10002v1 Announce Type: new Abstract: Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet inco

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Medical Incident Causal Factors and Preventive Measures Generation Using Tag-based Example Selection in Few-shot Learning

DGX agent

arXiv:2605.10025v1 Announce Type: cross Abstract: In high-stakes domains such as healthcare, the reliability of Large Language Models (LLMs) is critical, particularly when generating clinical insights

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies

DGX agent

arXiv:2605.09661v1 Announce Type: cross Abstract: Large language models (LLMs) have saturated standard medical benchmarks that test factual recall, yet their ability to perform higher-order reasoning,

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Meet physics-intern🧑‍🎓, our agentic framework for theoretical physics. It takes Gemini 3.1 Pro from 17.7% to 31.4% on CritPt, a new SOTA o…

DGX agent

Meet physics-intern🧑‍🎓, our agentic framework for theoretical physics. It takes Gemini 3.1 Pro from 17.7% to 31.4% on CritPt, a new SOTA on one of the hardest benchmarks for LLMs. Theoretical physics

model-releasesclem-delangue--x
12 May 2026
Model Releases

MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents

DGX agent

arXiv:2605.09530v1 Announce Type: cross Abstract: As LLM-powered agents are increasingly deployed in edge-cloud environments, personalized memory has become a key enabler of long-term adaptation and u

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs

DGX agent

arXiv:2605.08374v1 Announce Type: new Abstract: Episodic memory allows LLM agents to accumulate and retrieve experience, but current methods treat each memory independently, i.e., evaluating retrieval

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology

DGX agent

arXiv:2605.09152v1 Announce Type: new Abstract: Deciphering animal intent is a fundamental challenge in computational ethology, largely because of semantic aliasing, the phenomenon where identical ext

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

MESD: A Risk-Sensitive Metric for Explanation Fairness Across Intersectional Subgroups

DGX agent

arXiv:2603.13452v2 Announce Type: replace Abstract: Fairness in machine learning is predominantly evaluated through outcome-oriented metrics, such as Demographic parity, which measure whether predicti

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon

DGX agent

arXiv:2605.09708v1 Announce Type: cross Abstract: We present Metal-Sci, a 10-task benchmark of scientific Apple Silicon Metal compute kernels spanning six optimization regimes (stencils, all-pairs in

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

DGX agent

arXiv:2605.10067v1 Announce Type: cross Abstract: Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing ap

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MicroFuse: Protein-to-Genome Expert Fusion for Microbial Operon Reasoning

DGX agent

arXiv:2605.08815v1 Announce Type: new Abstract: Predicting microbial operon co-membership requires integrating two complementary biological signals: protein-scale molecular identity and genome-context

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph

DGX agent

arXiv:2605.10120v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) show remarkable potential for scientific reasoning, yet their performance in specialized domains such as micr

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MIDUS: Memory-Infused Depth Up-Scaling

DGX agent

arXiv:2512.13751v2 Announce Type: replace-cross Abstract: Expanding pre-trained language models offers a practical way to increase capacity without training larger models from scratch. Depth Up-Scalin

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models

DGX agent

arXiv:2510.09592v2 Announce Type: replace Abstract: Real-time Spoken Language Models (SLMs) struggle to leverage Chain-of-Thought (CoT) reasoning due to the prohibitive latency of generating the entir

model-releasesarxiv-cs-cl
12 May 2026
← Previous
1…330331332333334…471
Next →