AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,569 results
12 May 2026

jina-embeddings-v5-omni: Text-Geometry-Preserving Multimodal Embeddings via Frozen-Tower Composition

Model ReleasesDGX agent

arXiv:2605.08384v1 Announce Type: new Abstract: In this work, we introduce frozen-encoder model composition, a novel approach to multimodal embedding models. We build on the VLM-style architecture, in

jNO: A JAX Library for Neural Operator and Foundation Model Training

Model ReleasesDGX agent

arXiv:2605.10159v1 Announce Type: new Abstract: jNO (jax Neural Operators) is a JAX-native library for neural operators and foundation models with unified support for both data-driven and physics-info

JODA: Composable Joint Dynamics for Articulated Objects

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.09954v1 Announce Type: cross Abstract: Articulated objects used in simulation and embodied AI are typically specified by geometry and kinematic structure, but lack the fine-grained dynamica

K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs

Model ReleasesDGX agent

arXiv:2605.09635v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in K-12 education, yet existing benchmarks such as C-Eval, CMMLU, GaokaoBench, and EduEval mainly eva

KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation

Model ReleasesDGX agent

arXiv:2605.09572v1 Announce Type: cross Abstract: Sign language production from symbolic notation offers a scalable route to accessible sign animation. We present KANMultiSign, a multi-scale sequence

KARMA-MV: A Benchmark for Causal Question Answering on Music Videos

Model ReleasesDGX agent

arXiv:2605.08175v1 Announce Type: cross Abstract: While significant progress has been made in Video Question Answering and cross-modal understanding, causal reasoning about how visual dynamics drive m

KEPIL: Knowledge-Enhanced Prompt-Image Learning for Prompt-Robust Disease Detection

Model ReleasesDGX agent

arXiv:2605.09132v1 Announce Type: new Abstract: Vision--language models (VLMs) show promise for clinical decision support in radiology because they enable joint reasoning over radiological images and

Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation

Model ReleasesDGX agent

arXiv:2601.11258v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) face the 'knowledge cutoff' challenge, where their frozen parametric memory prevents direct internalization of ne

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving

Model ReleasesDGX agent

arXiv:2605.09735v1 Announce Type: cross Abstract: Static-graph LLM decoders provide predictable launches, fixed tensor shapes, and low submission overhead, but online decoding exposes highly irregular

Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes

Model ReleasesDGX agent

arXiv:2605.09751v1 Announce Type: new Abstract: Trainable input embedding tables are a standard component of modern language models. We ask whether they are actually necessary at the input interface.

Large Language Models as Students Who Think Aloud: Overly Coherent, Verbose, and Confident

Model ReleasesDGX agent

arXiv:2602.01015v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly embedded in AI-based tutoring systems. Can they faithfully model novice reasoning and metacognitive ju

Latent Geometry Beyond Search: Amortizing Planning in World Models

Model ReleasesDGX agent

arXiv:2605.08732v1 Announce Type: cross Abstract: Modern vision-based world models can represent observations as compact yet expressive latent manifolds, but fast goal-oriented planning in these space

Lattice Deduction Transformers

Model ReleasesDGX agent

arXiv:2605.08605v1 Announce Type: cross Abstract: We introduce the Lattice Deduction Transformer (LDT), a recurrent transformer that approximates logically sound deduction by projecting its latent sta

Layer Collapse in Diffusion Language Models

Model ReleasesDGX agent

arXiv:2605.06366v2 Announce Type: replace Abstract: Diffusion language models (DLMs) have recently emerged as competitive alternatives to autoregressive (AR) language models, yet differences in their

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models

Model ReleasesDGX agent

arXiv:2605.09806v1 Announce Type: cross Abstract: Large reasoning models, such as OpenAI o1 and DeepSeek-R1, tend to become increasingly verbose as their reasoning capabilities improve. These inflated

LEAF-SQL: Level-wise Exploration with Adaptive Fine-graining for Text-to-SQL Skeleton Prediction

Model ReleasesDGX agent

arXiv:2605.09295v1 Announce Type: new Abstract: Text-to-SQL translates natural language questions into executable SQL queries, enabling intuitive database access for non-experts. While large language

Learning Agile Striker Skills for Humanoid Soccer Robots from Noisy Sensory Input

Model ReleasesDGX agent

arXiv:2512.06571v3 Announce Type: replace Abstract: Learning fast and robust ball-kicking skills is a critical capability for humanoid soccer robots, yet it remains a challenging problem due to the ne

Learning Confidence Ellipsoids and Applications to Robust Subspace Recovery

Model ReleasesDGX agent

arXiv:2512.16875v4 Announce Type: replace-cross Abstract: We study the problem of finding confidence ellipsoids for an arbitrary distribution in high dimensions. Given samples from a distribution D an

Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs

Model ReleasesDGX agent

arXiv:2602.21198v2 Announce Type: replace-cross Abstract: Embodied LLMs endow robots with high-level task reasoning, but they cannot reflect on what went wrong or why, turning deployment into a sequen

Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining

Model ReleasesDGX agent

arXiv:2605.10504v1 Announce Type: new Abstract: A causal-decoder block is hierarchical: lower layers build the residual basis that upper layers attend over. We identify a failure mode in GPT pretraini

Learning Multi-Indicator Weights for Data Selection: A Joint Task-Model Adaptation Framework with Efficient Proxies

Model ReleasesDGX agent

arXiv:2605.09665v1 Announce Type: cross Abstract: Data selection is a key component of efficient instruction tuning for large language models, as recent work has shown that data quality often matters

Learning to Perceive 'Where': Spatial Pretext Tasks for Robust Self-Supervised Learning

Model ReleasesDGX agent

arXiv:2605.09963v1 Announce Type: new Abstract: Existing self-supervised learning (SSL) methods primarily learn object-invariant representations but often neglect the spatial structure and relationshi

Learning to Sparsify Stochastic Linear Bandits

Model ReleasesDGX agent

arXiv:2605.10151v1 Announce Type: new Abstract: This paper addresses the problem of learning to sparsify stochastic linear bandits, where a decision-maker sequentially selects actions from a high-dime

LegalCiteBench: Evaluating Citation Reliability in Legal Language Models

Model ReleasesDGX agent

arXiv:2605.10186v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly integrated into legal drafting and research workflows, where incorrect citations or fabricated precedent

Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models

Model ReleasesDGX agent

arXiv:2510.08592v3 Announce Type: replace-cross Abstract: Test-Time Scaling (TTS) improves LLM reasoning by exploring multiple candidate responses and then operating over this set to find the best out

LEVI: Stronger Search Architectures Can Substitute for Larger LLMs in Evolutionary Search

Model ReleasesDGX agent

arXiv:2605.09764v1 Announce Type: cross Abstract: LLM-guided evolutionary methods such as AlphaEvolve have proven effective in domains like math, systems research, and algorithmic discovery, but their

LightAVSeg: Lightweight Audio-Visual Segmentation

Model ReleasesDGX agent

arXiv:2605.08805v1 Announce Type: new Abstract: Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-mo

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities

Model ReleasesDGX agent

arXiv:2605.10810v1 Announce Type: new Abstract: We introduce an automatically generated benchmark for predicting hidden text in technical papers. A paper supplies visible context X and a hidden contin

LimeCross: Context-Conditioned Layered Image Editing with Structural Consistency

Model ReleasesDGX agent

arXiv:2605.10319v1 Announce Type: new Abstract: Layered image assets are widely used in real-world creative workflows, enabling non-destructive iteration and flexible re-composition. Recent advances i

LiteMedCoT-VL: Parameter-Efficient Adaptation for Medical Visual Question Answering

Model ReleasesDGX agent

arXiv:2605.09384v1 Announce Type: cross Abstract: The reasoning gap between large and compact vision-language models (VLMs) limits the deployment of medical AI on portable clinical devices. Compact VL

LiteParse is the best open-source, model-free document parser for AI agents. Run it over over 50+ document types, and it will parse dense pa…

Model ReleasesDGX agent

LiteParse is the best open-source, model-free document parser for AI agents. Run it over over 50+ document types, and it will parse dense pages with complex text layouts and tables, and it will extrac

LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments

Model ReleasesDGX agent

arXiv:2605.10779v1 Announce Type: cross Abstract: The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a new category of safety risk beyond content s

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?

Model ReleasesDGX agent

arXiv:2605.08985v1 Announce Type: new Abstract: Visual encoding constitutes a major computational bottleneck in Multimodal Large Language Models (MLLMs), especially for high-resolution image inputs. T

LLiMba: Sardinian on a Single GPU -- Adapting a 3B Language Model to a Vanishing Romance Language

Model ReleasesDGX agent

arXiv:2605.09015v1 Announce Type: new Abstract: Sardinian, a Romance language with roughly one million speakers, has minimal presence in modern NLP. Commercial services do not support it, and current

llm 0.32a2

Model ReleasesDGX agent

Release: llm 0.32a2 A bunch of useful stuff in this LLM alpha, but the most important detail is this one: Most reasoning-capable OpenAI models now use the /v1/responses endpoint instead of /v1/chat/co

LLM Agents Already Know When to Call Tools -- Even Without Reasoning

Model ReleasesDGX agent

arXiv:2605.09252v1 Announce Type: new Abstract: Tool-augmented LLM agents tend to call tools indiscriminately, even when the model can answer directly. Each unnecessary call wastes API fees and latenc

LLM Jaggedness Unlocks Scientific Creativity

Model ReleasesDGX agent

arXiv:2605.10574v1 Announce Type: new Abstract: As artificial intelligence advances, models are not improving uniformly. Instead, progress unfolds in a jagged fashion, with capabilities growing uneven

LLM Translation of Compiler Intermediate Representation

Model ReleasesDGX agent

arXiv:2605.08247v1 Announce Type: cross Abstract: GCC and LLVM underpin much of modern software infrastructure, relying on distinct Intermediate Representations (IRs) to drive optimizations and code g

LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight

Model ReleasesDGX agent

arXiv:2605.08321v1 Announce Type: cross Abstract: LLMs are increasingly capable of persuasion, which raises the question of how to protect users against manipulation. In a preregistered user study (N=

LLM4Branch: Large Language Model for Discovering Efficient Branching Policies of Integer Programs

Model ReleasesDGX agent

arXiv:2605.10401v1 Announce Type: new Abstract: Efficient branching policies are essential for accelerating Mixed Integer Linear Programming (MILP) solvers. Their design has long relied on hand-crafte

LLMs with in-context learning for Algorithmic Theoretical Physics

Model ReleasesDGX agent

arXiv:2605.08212v1 Announce Type: cross Abstract: There is an increasing number of algorithmic computations in theoretical physics. These, while conceptually simple, can nevertheless be time-consuming

LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems

Model ReleasesDGX agent

arXiv:2605.08305v1 Announce Type: cross Abstract: Large Language Model (LLM) systems have been the frontier of AI in many application domains, leading to new challenges and opportunities for hyperpara

Locking Pretrained Weights via Deep Low-Rank Residual Distillation

Model ReleasesDGX agent

arXiv:2605.10777v1 Announce Type: new Abstract: The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling the

Log analysis is necessary for credible evaluation of AI agents

Model ReleasesDGX agent

arXiv:2605.08545v1 Announce Type: new Abstract: Agent benchmarks typically report only final outcomes: pass or fail. This threatens evaluation credibility in three ways. First, scores may be inflated

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories

Model ReleasesDGX agent

arXiv:2509.20909v2 Announce Type: replace Abstract: Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, where benchmark contamination can make memor

Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.08787v1 Announce Type: new Abstract: Recent advances in 3D medical vision-language models have enabled joint reasoning over volumetric images and text, showing strong performance in medical

M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2605.09879v1 Announce Type: new Abstract: While reasoning has become a central capability of large language models (LLMs), the reasoning patterns required for different scenarios are often misal

M^2E-UAV: A Benchmark and Analysis for Onboard Motion-on-Motion Event-Based Tiny UAV Detection

Model ReleasesDGX agent

arXiv:2605.10496v1 Announce Type: new Abstract: Tiny UAV detection from an onboard event camera is difficult when the observer and target move at the same time. In this motion-on-motion regime, ego-mo

MaD Physics: Evaluating information seeking under constraints in physical environments

Model ReleasesDGX agent

arXiv:2605.10820v1 Announce Type: new Abstract: Scientific discovery is fundamentally a resource-constrained process that requires navigating complex trade-offs between the quality and quantity of mea

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

Model ReleasesDGX agent

arXiv:2605.08437v1 Announce Type: cross Abstract: Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to judge such argumen

MAGS-SLAM: Monocular Multi-Agent Gaussian Splatting SLAM for Geometrically and Photometrically Consistent Reconstruction

Model ReleasesDGX agent

arXiv:2605.10760v1 Announce Type: new Abstract: Collaborative photorealistic 3D reconstruction from multiple agents enables rapid large-scale scene capture for virtual production and cooperative multi

MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs

Model ReleasesDGX agent

arXiv:2605.08498v1 Announce Type: cross Abstract: We introduce MathConstraint, a hard, adaptive benchmark for evaluating the combinatorial reasoning capabilities of LLMs. We combine constraint satisfa

MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics

Model ReleasesDGX agent

arXiv:2602.02561v2 Announce Type: replace-cross Abstract: While the ecosystem of Lean and Mathlib has enjoyed celebrated success in formal mathematical reasoning with the help of large language models

MBP-KT: Learning Global Collaborative Information from Meta-Behavioral Pattern for Enhanced Knowledge Tracing

Model ReleasesDGX agent

arXiv:2605.08697v1 Announce Type: new Abstract: The emerging collaborative information-based knowledge tracing (KT) has been a promising way to enhance modeling of learners' knowledge states. The core

MC-RFM: Geometry-Aware Few-Shot Adaptation via Mixed-Curvature Riemannian Flow Matching

Model ReleasesDGX agent

arXiv:2605.08557v1 Announce Type: cross Abstract: Parameter-efficient adaptation of pretrained vision models is commonly performed through linear probes, prompts, low-rank updates, or lightweight resi

MC^2: Monte Carlo Correction for Fast Elliptic PDE Solving

Model ReleasesDGX agent

arXiv:2605.09288v1 Announce Type: cross Abstract: Partial differential equation (PDE) solvers underpin scientific computing, but real-world deployment is bounded by compute. Classical Monte Carlo solv

MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environments

Model ReleasesDGX agent

arXiv:2605.09131v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has unified the interface between Large Language Models (LLMs) and external tools, yet a fundamental gap remains in how

MDGYM: Benchmarking AI Agents on Molecular Simulations

Model ReleasesDGX agent

arXiv:2605.08941v1 Announce Type: new Abstract: The promise of AI-driven scientific discovery hinges on whether AI agents can autonomously design and execute the computational workflows that underpin

MDrive: Benchmarking Closed-Loop Cooperative Driving for End-to-End Multi-agent Systems

Model ReleasesDGX agent

arXiv:2605.10904v1 Announce Type: new Abstract: Vehicle-to-Everything (V2X) communication has emerged as a promising paradigm for autonomous driving, enabling connected agents to share complementary p

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

Model ReleasesDGX agent

arXiv:2510.22170v2 Announce Type: replace Abstract: Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure o

← Previous
1…264265266267268…377
Next →