AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments

DGX agent

arXiv:2605.10779v1 Announce Type: cross Abstract: The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content s

model-releasesarxiv-cs-cl
12 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?

DGX agent

arXiv:2605.08985v1 Announce Type: new Abstract: Visual encoding constitutes a major computational bottleneck in Multimodal Large Language Models (MLLMs), especially for high-resolution image inputs. T

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

LLiMba: Sardinian on a Single GPU -- Adapting a 3B Language Model to a Vanishing Romance Language

DGX agent

arXiv:2605.09015v1 Announce Type: new Abstract: Sardinian, a Romance language with roughly one million speakers, has minimal presence in modern NLP. Commercial services do not support it, and current

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LLM Agents Already Know When to Call Tools -- Even Without Reasoning

DGX agent

arXiv:2605.09252v1 Announce Type: new Abstract: Tool-augmented LLM agents tend to call tools indiscriminately, even when the model can answer directly. Each unnecessary call wastes API fees and latenc

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LLM Jaggedness Unlocks Scientific Creativity

DGX agent

arXiv:2605.10574v1 Announce Type: new Abstract: As artificial intelligence advances, models are not improving uniformly. Instead, progress unfolds in a jagged fashion, with capabilities growing uneven

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LLM Translation of Compiler Intermediate Representation

DGX agent

arXiv:2605.08247v1 Announce Type: cross Abstract: GCC and LLVM underpin much of modern software infrastructure, relying on distinct Intermediate Representations (IRs) to drive optimizations and code g

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight

DGX agent

arXiv:2605.08321v1 Announce Type: cross Abstract: LLMs are increasingly capable of persuasion, which raises the question of how to protect users against manipulation. In a preregistered user study (N=

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LLM4Branch: Large Language Model for Discovering Efficient Branching Policies of Integer Programs

DGX agent

arXiv:2605.10401v1 Announce Type: new Abstract: Efficient branching policies are essential for accelerating Mixed Integer Linear Programming (MILP) solvers. Their design has long relied on hand-crafte

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LLMs with in-context learning for Algorithmic Theoretical Physics

DGX agent

arXiv:2605.08212v1 Announce Type: cross Abstract: There is an increasing number of algorithmic computations in theoretical physics. These, while conceptually simple, can nevertheless be time-consuming

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems

DGX agent

arXiv:2605.08305v1 Announce Type: cross Abstract: Large Language Model (LLM) systems have been the frontier of AI in many application domains, leading to new challenges and opportunities for hyperpara

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Locking Pretrained Weights via Deep Low-Rank Residual Distillation

DGX agent

arXiv:2605.10777v1 Announce Type: new Abstract: The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling the

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Log analysis is necessary for credible evaluation of AI agents

DGX agent

arXiv:2605.08545v1 Announce Type: new Abstract: Agent benchmarks typically report only final outcomes: pass or fail. This threatens evaluation credibility in three ways. First, scores may be inflated

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories

DGX agent

arXiv:2509.20909v2 Announce Type: replace Abstract: Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, where benchmark contamination can make memor

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models

DGX agent

arXiv:2605.08787v1 Announce Type: new Abstract: Recent advances in 3D medical vision-language models have enabled joint reasoning over volumetric images and text, showing strong performance in medical

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models

DGX agent

arXiv:2605.09879v1 Announce Type: new Abstract: While reasoning has become a central capability of large language models (LLMs), the reasoning patterns required for different scenarios are often misal

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

M^2E-UAV: A Benchmark and Analysis for Onboard Motion-on-Motion Event-Based Tiny UAV Detection

DGX agent

arXiv:2605.10496v1 Announce Type: new Abstract: Tiny UAV detection from an onboard event camera is difficult when the observer and target move at the same time. In this motion-on-motion regime, ego-mo

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

MaD Physics: Evaluating information seeking under constraints in physical environments

DGX agent

arXiv:2605.10820v1 Announce Type: new Abstract: Scientific discovery is fundamentally a resource-constrained process that requires navigating complex trade-offs between the quality and quantity of mea

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

DGX agent

arXiv:2605.08437v1 Announce Type: cross Abstract: Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to judge such argumen

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MAGS-SLAM: Monocular Multi-Agent Gaussian Splatting SLAM for Geometrically and Photometrically Consistent Reconstruction

DGX agent

arXiv:2605.10760v1 Announce Type: new Abstract: Collaborative photorealistic 3D reconstruction from multiple agents enables rapid large-scale scene capture for virtual production and cooperative multi

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs

DGX agent

arXiv:2605.08498v1 Announce Type: cross Abstract: We introduce MathConstraint, a hard, adaptive benchmark for evaluating the combinatorial reasoning capabilities of LLMs. We combine constraint satisfa

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics

DGX agent

arXiv:2602.02561v2 Announce Type: replace-cross Abstract: While the ecosystem of Lean and Mathlib has enjoyed celebrated success in formal mathematical reasoning with the help of large language models

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MBP-KT: Learning Global Collaborative Information from Meta-Behavioral Pattern for Enhanced Knowledge Tracing

DGX agent

arXiv:2605.08697v1 Announce Type: new Abstract: The emerging collaborative information-based knowledge tracing (KT) has been a promising way to enhance modeling of learners' knowledge states. The core

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MC-RFM: Geometry-Aware Few-Shot Adaptation via Mixed-Curvature Riemannian Flow Matching

DGX agent

arXiv:2605.08557v1 Announce Type: cross Abstract: Parameter-efficient adaptation of pretrained vision models is commonly performed through linear probes, prompts, low-rank updates, or lightweight resi

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MC^2: Monte Carlo Correction for Fast Elliptic PDE Solving

DGX agent

arXiv:2605.09288v1 Announce Type: cross Abstract: Partial differential equation (PDE) solvers underpin scientific computing, but real-world deployment is bounded by compute. Classical Monte Carlo solv

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments

DGX agent

arXiv:2605.09131v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has unified the interface between Large Language Models (LLMs) and external tools, yet a fundamental gap remains in how

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MDGYM: Benchmarking AI Agents on Molecular Simulations

DGX agent

arXiv:2605.08941v1 Announce Type: new Abstract: The promise of AI-driven scientific discovery hinges on whether AI agents can autonomously design and execute the computational workflows that underpin

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MDrive: Benchmarking Closed-Loop Cooperative Driving for End-to-End Multi-agent Systems

DGX agent

arXiv:2605.10904v1 Announce Type: new Abstract: Vehicle-to-Everything (V2X) communication has emerged as a promising paradigm for autonomous driving, enabling connected agents to share complementary p

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

DGX agent

arXiv:2510.22170v2 Announce Type: replace Abstract: Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure o

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare

DGX agent

arXiv:2605.08445v1 Announce Type: new Abstract: AI models are increasingly deployed in live clinical environments where they must perform reliably across complex, high-stakes workflows that standard t

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks

DGX agent

arXiv:2507.23511v3 Announce Type: replace-cross Abstract: While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. Th

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models

DGX agent

arXiv:2605.10002v1 Announce Type: new Abstract: Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet inco

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Medical Incident Causal Factors and Preventive Measures Generation Using Tag-based Example Selection in Few-shot Learning

DGX agent

arXiv:2605.10025v1 Announce Type: cross Abstract: In high-stakes domains such as healthcare, the reliability of Large Language Models (LLMs) is critical, particularly when generating clinical insights

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MedMeta: A Benchmark for LLMs in Synthesizing Meta-Analysis Conclusion from Medical Studies

DGX agent

arXiv:2605.09661v1 Announce Type: cross Abstract: Large language models (LLMs) have saturated standard medical benchmarks that test factual recall, yet their ability to perform higher-order reasoning,

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents

DGX agent

arXiv:2605.09530v1 Announce Type: cross Abstract: As LLM-powered agents are increasingly deployed in edge-cloud environments, personalized memory has become a key enabler of long-term adaptation and u

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs

DGX agent

arXiv:2605.08374v1 Announce Type: new Abstract: Episodic memory allows LLM agents to accumulate and retrieve experience, but current methods treat each memory independently, i.e., evaluating retrieval

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology

DGX agent

arXiv:2605.09152v1 Announce Type: new Abstract: Deciphering animal intent is a fundamental challenge in computational ethology, largely because of semantic aliasing, the phenomenon where identical ext

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

MESD: A Risk-Sensitive Metric for Explanation Fairness Across Intersectional Subgroups

DGX agent

arXiv:2603.13452v2 Announce Type: replace Abstract: Fairness in machine learning is predominantly evaluated through outcome-oriented metrics, such as Demographic parity, which measure whether predicti

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon

DGX agent

arXiv:2605.09708v1 Announce Type: cross Abstract: We present Metal-Sci, a 10-task benchmark of scientific Apple Silicon Metal compute kernels spanning six optimization regimes (stencils, all-pairs in

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

DGX agent

arXiv:2605.10067v1 Announce Type: cross Abstract: Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing ap

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MicroFuse: Protein-to-Genome Expert Fusion for Microbial Operon Reasoning

DGX agent

arXiv:2605.08815v1 Announce Type: new Abstract: Predicting microbial operon co-membership requires integrating two complementary biological signals: protein-scale molecular identity and genome-context

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph

DGX agent

arXiv:2605.10120v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) show remarkable potential for scientific reasoning, yet their performance in specialized domains such as micr

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MIDUS: Memory-Infused Depth Up-Scaling

DGX agent

arXiv:2512.13751v2 Announce Type: replace-cross Abstract: Expanding pre-trained language models offers a practical way to increase capacity without training larger models from scratch. Depth Up-Scalin

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models

DGX agent

arXiv:2510.09592v2 Announce Type: replace Abstract: Real-time Spoken Language Models (SLMs) struggle to leverage Chain-of-Thought (CoT) reasoning due to the prohibitive latency of generating the entir

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Minimizing Worst-Case Weighted Latency for Multi-Robot Persistent Monitoring: Theory and RL-Based Solutions

DGX agent

arXiv:2605.09633v1 Announce Type: new Abstract: We study multi-robot persistent monitoring on weighted graphs, where node weights encode monitoring priorities and edge weights encode travel distances.

model-releasesarxiv-cs-ro
12 May 2026
Model Releases

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?

DGX agent

arXiv:2605.08816v1 Announce Type: new Abstract: In the animal kingdom, mirror self-recognition is a canonical probe of higher-order cognition, emerging only in some species. We ask whether an analogou

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI

DGX agent

arXiv:2605.08678v1 Announce Type: new Abstract: Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonst

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection

DGX agent

arXiv:2605.10833v1 Announce Type: cross Abstract: Industrial anomaly detection is critical for manufacturing quality control, yet existing datasets mainly focus on static images or sparse views, which

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Model Capacity Determines Grokking through Competing Memorisation and Generalisation Speeds

DGX agent

arXiv:2605.09724v1 Announce Type: new Abstract: Existing accounts of grokking explain the phenomena in terms of mechanistic frameworks such as circuit efficiency or lazy-to-rich transitions. However,

model-releasesarxiv-cs-lg
12 May 2026
← Previous
1…256257258259260…361
Next →