AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

86,510Total entries
1Added by human
86,509Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,082 results
20 May 2026

To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents

Model ReleasesDGX agent

arXiv:2605.18882v1 Announce Type: cross Abstract: LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, six models

UCCI: Calibrated Uncertainty for Cost-Optimal LLM Cascade Routing

HardwareDGX agent

arXiv:2605.18796v1 Announce Type: cross Abstract: LLM cascades and model routing promise lower inference cost by sending easy queries to a small model and escalating hard ones to a large model, but mo

VL-DPO: Vision-Language-Guided Finetuning for Preference-Aligned Autonomous Driving

AgentsDGX agent

arXiv:2605.20082v1 Announce Type: cross Abstract: The rapid growth of autonomous driving datasets has enabled the scaling of powerful motion forecasting models. While large-scale pretraining provides

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making

SafetyDGX agent

arXiv:2602.07008v2 Announce Type: replace Abstract: Reliable models should not only predict correctly, but also justify decisions with acceptable evidence. Yet conventional supervised learning typical

19 May 2026

Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation

Model ReleasesDGX agent

arXiv:2602.16990v2 Announce Type: replace Abstract: Most recommendation benchmarks evaluate how well a model imitates user behavior. In financial advisory, however, observed actions can be noisy or sh

Detecting Verbatim LLM Copy-Paste in Homework

Model ReleasesDGX agent

arXiv:2605.16336v1 Announce Type: cross Abstract: Large language models (LLMs) have made fluent essay writing, code drafting, and quiz answering instantly available to students at every level, from se

Disentangled Latent Dynamics Manifold Fusion for Solving Parameterized PDEs

Model ReleasesDGX agent

arXiv:2603.12676v2 Announce Type: replace Abstract: Generalizing neural surrogate models across different PDE parameters remains difficult because changes in PDE coefficients often make learning harde

EfficientTDMPC: Improved MPC Objectives for Sample-Efficient Continuous Control

Model ReleasesDGX agent

arXiv:2605.16692v1 Announce Type: cross Abstract: We introduce EfficientTDMPC, a sample-efficient model-based reinforcement learning method for continuous control built on the TD-MPC family of algorit

Episodic-Semantic Memory Architecture for Long-Horizon Scientific Agents

Model ReleasesDGX agent

arXiv:2605.17625v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into persistent scientific collaborators, context window saturation has emerged as a critical bottleneck. Scienti

GLT-PEFT: Gated Lie-Tucker Parameter-Efficient Fine-Tuning for Alzheimer's Disease Diagnosis with Hippocampal Segmentation Pretraining

Model ReleasesDGX agent

arXiv:2605.16769v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) has emerged as a promising paradigm for adapting pretrained models under limited data conditions. However, most e

Language Models as Efficient Reward Function Searchers for Custom-Environment Multi-Objective Reinforcement

ResearchDGX agent

arXiv:2409.02428v4 Announce Type: replace-cross Abstract: Achieving the effective design and improvement of reward functions in reinforcement learning (RL) tasks with complex custom environments and m

LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models

HardwareDGX agent

arXiv:2605.17289v1 Announce Type: cross Abstract: Unstructured sparsity is now natively accelerated by recent GPU kernels and dataflow hardware, shifting the bottleneck from inference execution to the

Learning-Zone Energy: Online Data Selection for Efficient RL Post-Training

Model ReleasesDGX agent

arXiv:2605.17003v1 Announce Type: cross Abstract: Reinforcement Learning (RL) post-training has emerged as the dominant paradigm for eliciting mathematical reasoning in Large Language Models (LLMs), y

Measuring Changes in Instructor Class Design and Student Learning After the Release of Large Language Models (LLMs)

SafetyDGX agent

arXiv:2605.16284v1 Announce Type: cross Abstract: Student use of Generative AI (GenAI) products in completing their classwork, with or without their professors' knowledge and/or approval, has resulted

MiniGPT: Rebuilding GPT from First Principles

Model ReleasesDGX agent

arXiv:2605.17398v1 Announce Type: new Abstract: This paper presents MiniGPT, a compact from-scratch implementation of GPT-style autoregressive language modeling in PyTorch. The aim is to rebuild the c

Monocular Depth Perception Enhancement Based on Joint Shading/Contrast Model and Motion Parallax (JSM)

ResearchDGX agent

arXiv:2605.17252v1 Announce Type: new Abstract: Stereoscopic 3D displays adopt a binocular depth cue to provide depth perception. However, users should be equipped with expensive special devices to ap

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction

Model ReleasesDGX agent

arXiv:2605.17360v1 Announce Type: new Abstract: Real-time duplex interaction is essential for multimodal AI systems operating in real-world scenarios, where models must continuously process streaming

OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

Model ReleasesDGX agent

arXiv:2605.18577v1 Announce Type: new Abstract: Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emer

OpenJarvis: Personal AI, On Personal Devices

Model ReleasesDGX agent

arXiv:2605.17172v1 Announce Type: cross Abstract: Personal AI stacks, like OpenClaw and Hermes Agent, are becoming central to daily work, yet they route nearly every query (often over sensitive local

PARALLAX: Separating Genuine Hallucination Detection from Benchmark Construction Artifacts

Model ReleasesDGX agent

arXiv:2605.17028v1 Announce Type: cross Abstract: Large language models (LLMs) hallucinate with confidence: their outputs can be fluent, authoritative, and simply wrong. In medical, legal, and scienti

Position: AI Evaluations Should be Grounded on a Theory of Capability

Model ReleasesDGX agent

arXiv:2509.19590v2 Announce Type: replace Abstract: Evaluations of generative models are now ubiquitous, and their outcomes critically shape public and scientific expectations of AI's capabilities. Ye

Post-Trained MoE Can Skip Half Experts via Self-Distillation

Model ReleasesDGX agent

arXiv:2605.18643v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) scales language models efficiently through sparse expert activation, and its dynamic variant further reduces computation by a

Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control

Model ReleasesDGX agent

arXiv:2605.18414v1 Announce Type: cross Abstract: Large language models increasingly operate as autonomous agents that select and invoke tools from large registries. We identify a critical gap: when u

Provable Knowledge Acquisition and Extraction in One-Layer Transformers

Model ReleasesDGX agent

arXiv:2508.00901v4 Announce Type: replace-cross Abstract: Large language models may encounter factual knowledge during pre-training yet fail to reliably use that knowledge after fine-tuning. Despite g

Provably Shorter Scratchpads in Hybrid DeltaNet-Attention Decoders

Model ReleasesDGX agent

arXiv:2605.16640v1 Announce Type: new Abstract: We investigate the expressive power of hybrid recurrent-attention decoders, a class of architectures used in recent open-source language models such as

RT-Splatting: Joint Reflection-Transmission Modeling with Gaussian Splatting

ResearchDGX agent

arXiv:2605.18263v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) enables real-time novel view synthesis with high visual quality. However, existing methods struggle with semi-transparent s

SSTL: Self-Sensing Tendon Loop for Hysteresis Modeling and Compensation in Tendon-Sheath Mechanisms

ResearchDGX agent

arXiv:2605.16870v1 Announce Type: new Abstract: Flexible endoscopic robots enable minimally invasive access through natural orifices, but their control accuracy is limited by configuration-dependent h

StrLoRA: Towards Streaming Continual Visual Instruction Tuning for MLLMs

Model ReleasesDGX agent

arXiv:2605.16353v1 Announce Type: cross Abstract: Continual Visual Instruction Tuning (CVIT) enables Multimodal Large Language Models to incrementally acquire new abilities. However, existing CVIT met

STT-Arena: A More Realistic Environment for Tool-Using with Spatio-Temporal Dynamics

Model ReleasesDGX agent

arXiv:2605.18548v1 Announce Type: cross Abstract: Large language models (LLMs) deployed in real-world agentic applications must be capable of replanning and adapting when mid-task disruptions invalida

SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in the Gaming Vertical Domain

Model ReleasesDGX agent

arXiv:2605.17946v1 Announce Type: new Abstract: Multimodal large language models are increasingly used as agent backbones that understand multimodal inputs, plan retrieval actions, invoke external too

SwordBench: Evaluating Orthogonality of Steering Image Representations

Model ReleasesDGX agent

arXiv:2605.16372v1 Announce Type: cross Abstract: Steering or intervening on model representations at inference time to correct predictions is essential for AI interpretability and safety, yet existin

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2605.17467v1 Announce Type: new Abstract: Large language model-driven multi-agent systems (LLM-MAS) excel at complex tasks, yet unreliable agents remain a key bottleneck to system-level reliabil

Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding

AgentsDGX agent

arXiv:2605.17823v1 Announce Type: cross Abstract: When humans view scenes without a specific task (free-viewing), they initially direct their eye movements toward the scene center and then fixate on p

18 May 2026

Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time

SafetyDGX agent

arXiv:2605.15220v1 Announce Type: cross Abstract: Data mixing decides how to combine different sources or types of data and is a consequential problem throughout language model training. In pretrainin

Antidistillation Fingerprinting

ResearchDGX agent

arXiv:2602.03812v2 Announce Type: replace-cross Abstract: Model distillation enables efficient emulation of frontier large language models (LLMs), creating a need for robust mechanisms to detect when

CAX-Agent: A Lightweight Agent Harness for Reliable APDL Automation

Model ReleasesDGX agent

arXiv:2605.15218v1 Announce Type: new Abstract: Large language models deployed for MAPDL finite-element simulation face practical reliability challenges: without structured execution control, tool enc

ColPackAgent: Agent-Skill-Guided Hard-Particle Monte Carlo Workflows for Colloidal Packing

Model ReleasesDGX agent

arXiv:2605.15625v1 Announce Type: new Abstract: We introduce ColPackAgent, an agent framework that autonomously runs Monte Carlo simulations of colloidal packing through a Model Context Protocol (MCP)

ELDOR: A Dataset and Benchmark for Illegal Gold Mining in the Amazon Rainforest

Model ReleasesDGX agent

arXiv:2605.15397v1 Announce Type: new Abstract: Illegal gold mining in the Amazon rainforest causes deforestation, water contamination, and long-term ecosystem disruption, yet remains difficult to mon

IHF-Harmony: Multi-Modality Magnetic Resonance Images Harmonization using Invertible Hierarchy Flow Model

ResearchDGX agent

arXiv:2602.21536v2 Announce Type: replace Abstract: Retrospective MRI harmonization is limited by poor scalability across modalities and reliance on traveling subject datasets. To address these challe

Learning Bilevel Policies over Symbolic World Models for Long-Horizon Planning

SafetyDGX agent

arXiv:2605.15975v1 Announce Type: new Abstract: We tackle the challenge of building embodied AI agents that can reliably solve long-horizon planning problems. Imitation learning from demonstrations ha

PBT-Bench: Benchmarking AI Agents on Property-Based Testing

Model ReleasesDGX agent

arXiv:2605.15229v1 Announce Type: cross Abstract: Existing code benchmarks measure whether an agent can produce any test that reproduces a known bug, or whether it can produce a patch that fixes a des

Prospective multi-pathogen disease forecasting using autonomous LLM-guided tree search

AgentsDGX agent

arXiv:2605.16238v1 Announce Type: new Abstract: Probabilistic forecasting of infectious diseases is crucial for public health but relies on labor-intensive manual model curation by expert modeling tea

Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation

Model ReleasesDGX agent

arXiv:2605.15239v1 Announce Type: new Abstract: Safety alignment often improves robustness to harmful queries at the cost of reasoning ability, a tradeoff known as the safety tax. A common cause is di

SGR: A Stepwise Reasoning Framework for LLMs with External Subgraph Generation

Model ReleasesDGX agent

arXiv:2605.16117v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated strong capabilities across diverse NLP applications, such as translation, text generation, and question a

SurvivalPFN: Amortizing Survival Prediction via In-Context Bayesian Inference

Model ReleasesDGX agent

arXiv:2605.15488v1 Announce Type: new Abstract: Survival analysis provides a powerful statistical framework for modeling time-to-event outcomes in the presence of censoring. However, selecting an appr

Unlocking Dense Metric Depth Estimation in VLMs

Model ReleasesDGX agent

arXiv:2605.15876v1 Announce Type: new Abstract: Vision-Language Models (VLMs) excel at 2D tasks such as grounding and captioning, yet remain limited in 3D understanding. A key limitation is their text

15 May 2026

Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models

ResearchDGX agent

arXiv:2605.15082v1 Announce Type: cross Abstract: We study a prototypical situation when a learned predictor can discover useful low-dimensional structure in data, while using fewer samples than are n

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment

Model ReleasesDGX agent

arXiv:2605.14311v1 Announce Type: cross Abstract: Test-Time Scaling (TTS), which samples multiple candidate actions and ranks them via a Critic Model, has emerged as a promising paradigm for generalis

Cognitive-Uncertainty Guided Knowledge Distillation for Accurate Classification of Student Misconceptions

Model ReleasesDGX agent

arXiv:2605.14752v1 Announce Type: cross Abstract: Accurately identifying student misconceptions is crucial for personalized education but faces three challenges: (1) data scarcity with long-tail distr

FU-MPC: Frontier- and Uncertainty-Aware Model Predictive Control for Efficient and Accurate UAV Exploration with Motorized LiDAR

ResearchDGX agent

arXiv:2605.14920v1 Announce Type: new Abstract: Efficient UAV exploration in unknown environments requires rapid coverage expansion while maintaining accurate and reliable localization, since safe nav

L2R: Low-Rank and Lipschitz-Controlled Routing for Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2601.21349v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) models scale neural networks by conditionally activating a small subset of experts, where the router plays a central

ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World

Model ReleasesDGX agent

arXiv:2605.15081v1 Announce Type: cross Abstract: The development of high-quality text embeddings is increasingly drifting toward an exclusionary future, defined by three critical barriers: prohibitiv

Modeling Bounded Rationality in Drug Shortage Pharmacists Using Attention-Guided Dynamic Decomposition

AgentsDGX agent

arXiv:2605.14111v1 Announce Type: new Abstract: Hospital pharmacists make high-stakes decisions to mitigate drug shortages under uncertainty, time pressure, and patient risk. Interviews revealed that

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries

SafetyDGX agent

arXiv:2605.14605v1 Announce Type: cross Abstract: Model providers increasingly release open weights or allow users to fine-tune foundation models through APIs. Although these models are safety-aligned

PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts

Model ReleasesDGX agent

arXiv:2605.14002v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) embedded in agentic frameworks have transformed information retrieval from static, long context question answering into op

Randomized Atomic Feature Models for Physics-Informed Identification of Dynamic Systems

ResearchDGX agent

arXiv:2605.14351v1 Announce Type: cross Abstract: We present a physics-informed framework for system identification based on randomized stable atomic features. Impulse responses are represented as ran

Scaling Laws from Sequential Feature Recovery: A Solvable Hierarchical Model

ResearchDGX agent

arXiv:2605.14567v1 Announce Type: cross Abstract: We propose a simple mechanism by which scaling laws emerge from feature learning in multi-layer networks. We study a high-dimensional hierarchical tar

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture

Model ReleasesDGX agent

arXiv:2605.14448v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have emerged as a powerful backbone for multimodal embeddings. Recent methods introduce chain-of-thought (CoT

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation

SafetyDGX agent

arXiv:2602.02994v2 Announce Type: replace Abstract: Reinforcement learning has emerged as a principled post-training paradigm for Temporal Video Grounding (TVG) due to its on-policy optimization, yet

14 May 2026

Backbone is All You Need: Assessing Vulnerabilities of Frozen Foundation Models in Synthetic Image Forensics

ResearchDGX agent

arXiv:2605.13381v1 Announce Type: new Abstract: As AI-generated synthetic images become increasingly realistic, Vision Transformers (ViTs) have emerged as a cornerstone of modern deepfake detection. H

← Previous
1…283284285286287…1035
Next →