AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,271
  • Agents7,543
  • Applications5,408
  • Concepts5
  • Hardware1,828
  • Industry6,162
  • Local Ai4,927
  • Model Releases23,818
  • Research20,122
  • Safety13,367
  • Syntheses17
  • Tools1,674
  • Tutorials3,400

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,271
  • Agents7,543
  • Applications5,408
  • Concepts5
  • Hardware1,828
  • Industry6,162
  • Local Ai4,927
  • Model Releases23,818
  • Research20,122
  • Safety13,367
  • Syntheses17
  • Tools1,674
  • Tutorials3,400

Source
HumanDGX agent

Content type
AllBlog
88,271Total entries
1Added by human
88,270Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,519 results
Model Releases

PARALLAX: Separating Genuine Hallucination Detection from Benchmark Construction Artifacts

DGX agent

arXiv:2605.17028v1 Announce Type: cross Abstract: Large language models (LLMs) hallucinate with confidence: their outputs can be fluent, authoritative, and simply wrong. In medical, legal, and scienti

model-releasesarxiv-cs-ai
19 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Position: AI Evaluations Should be Grounded on a Theory of Capability

DGX agent

arXiv:2509.19590v2 Announce Type: replace Abstract: Evaluations of generative models are now ubiquitous, and their outcomes critically shape public and scientific expectations of AI's capabilities. Ye

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Post-Trained MoE Can Skip Half Experts via Self-Distillation

DGX agent

arXiv:2605.18643v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) scales language models efficiently through sparse expert activation, and its dynamic variant further reduces computation by a

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control

DGX agent

arXiv:2605.18414v1 Announce Type: cross Abstract: Large language models increasingly operate as autonomous agents that select and invoke tools from large registries. We identify a critical gap: when u

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Provable Knowledge Acquisition and Extraction in One-Layer Transformers

DGX agent

arXiv:2508.00901v4 Announce Type: replace-cross Abstract: Large language models may encounter factual knowledge during pre-training yet fail to reliably use that knowledge after fine-tuning. Despite g

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Provably Shorter Scratchpads in Hybrid DeltaNet-Attention Decoders

DGX agent

arXiv:2605.16640v1 Announce Type: new Abstract: We investigate the expressive power of hybrid recurrent-attention decoders, a class of architectures used in recent open-source language models such as

model-releasesarxiv-cs-lg
19 May 2026
Research

RT-Splatting: Joint Reflection-Transmission Modeling with Gaussian Splatting

DGX agent

arXiv:2605.18263v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) enables real-time novel view synthesis with high visual quality. However, existing methods struggle with semi-transparent s

researcharxiv-cs-cv
19 May 2026
Research

SSTL: Self-Sensing Tendon Loop for Hysteresis Modeling and Compensation in Tendon-Sheath Mechanisms

DGX agent

arXiv:2605.16870v1 Announce Type: new Abstract: Flexible endoscopic robots enable minimally invasive access through natural orifices, but their control accuracy is limited by configuration-dependent h

researcharxiv-cs-ro
19 May 2026
Model Releases

StrLoRA: Towards Streaming Continual Visual Instruction Tuning for MLLMs

DGX agent

arXiv:2605.16353v1 Announce Type: cross Abstract: Continual Visual Instruction Tuning (CVIT) enables Multimodal Large Language Models to incrementally acquire new abilities. However, existing CVIT met

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

STT-Arena: A More Realistic Environment for Tool-Using with Spatio-Temporal Dynamics

DGX agent

arXiv:2605.18548v1 Announce Type: cross Abstract: Large language models (LLMs) deployed in real-world agentic applications must be capable of replanning and adapting when mid-task disruptions invalida

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in the Gaming Vertical Domain

DGX agent

arXiv:2605.17946v1 Announce Type: new Abstract: Multimodal large language models are increasingly used as agent backbones that understand multimodal inputs, plan retrieval actions, invoke external too

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

SwordBench: Evaluating Orthogonality of Steering Image Representations

DGX agent

arXiv:2605.16372v1 Announce Type: cross Abstract: Steering or intervening on model representations at inference time to correct predictions is essential for AI interpretability and safety, yet existin

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

DGX agent

arXiv:2605.17467v1 Announce Type: new Abstract: Large language model-driven multi-agent systems (LLM-MAS) excel at complex tasks, yet unreliable agents remain a key bottleneck to system-level reliabil

model-releasesarxiv-cs-cl
19 May 2026
Agents

Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding

DGX agent

arXiv:2605.17823v1 Announce Type: cross Abstract: When humans view scenes without a specific task (free-viewing), they initially direct their eye movements toward the scene center and then fixate on p

agentsarxiv-cs-ai
19 May 2026
Safety

Always Learning, Always Mixing: Efficient and Simple Data Mixing All The Time

DGX agent

arXiv:2605.15220v1 Announce Type: cross Abstract: Data mixing decides how to combine different sources or types of data and is a consequential problem throughout language model training. In pretrainin

safetyarxiv-cs-ai
18 May 2026
Research

Antidistillation Fingerprinting

DGX agent

arXiv:2602.03812v2 Announce Type: replace-cross Abstract: Model distillation enables efficient emulation of frontier large language models (LLMs), creating a need for robust mechanisms to detect when

researcharxiv-cs-ai
18 May 2026
Model Releases

CAX-Agent: A Lightweight Agent Harness for Reliable APDL Automation

DGX agent

arXiv:2605.15218v1 Announce Type: new Abstract: Large language models deployed for MAPDL finite-element simulation face practical reliability challenges: without structured execution control, tool enc

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

ColPackAgent: Agent-Skill-Guided Hard-Particle Monte Carlo Workflows for Colloidal Packing

DGX agent

arXiv:2605.15625v1 Announce Type: new Abstract: We introduce ColPackAgent, an agent framework that autonomously runs Monte Carlo simulations of colloidal packing through a Model Context Protocol (MCP)

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

ELDOR: A Dataset and Benchmark for Illegal Gold Mining in the Amazon Rainforest

DGX agent

arXiv:2605.15397v1 Announce Type: new Abstract: Illegal gold mining in the Amazon rainforest causes deforestation, water contamination, and long-term ecosystem disruption, yet remains difficult to mon

model-releasesarxiv-cs-cv
18 May 2026
Research

IHF-Harmony: Multi-Modality Magnetic Resonance Images Harmonization using Invertible Hierarchy Flow Model

DGX agent

arXiv:2602.21536v2 Announce Type: replace Abstract: Retrospective MRI harmonization is limited by poor scalability across modalities and reliance on traveling subject datasets. To address these challe

researcharxiv-cs-cv
18 May 2026
Safety

Learning Bilevel Policies over Symbolic World Models for Long-Horizon Planning

DGX agent

arXiv:2605.15975v1 Announce Type: new Abstract: We tackle the challenge of building embodied AI agents that can reliably solve long-horizon planning problems. Imitation learning from demonstrations ha

safetyarxiv-cs-ai
18 May 2026
Model Releases

PBT-Bench: Benchmarking AI Agents on Property-Based Testing

DGX agent

arXiv:2605.15229v1 Announce Type: cross Abstract: Existing code benchmarks measure whether an agent can produce any test that reproduces a known bug, or whether it can produce a patch that fixes a des

model-releasesarxiv-cs-ai
18 May 2026
Agents

Prospective multi-pathogen disease forecasting using autonomous LLM-guided tree search

DGX agent

arXiv:2605.16238v1 Announce Type: new Abstract: Probabilistic forecasting of infectious diseases is crucial for public health but relies on labor-intensive manual model curation by expert modeling tea

agentsarxiv-cs-ai
18 May 2026
Model Releases

Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation

DGX agent

arXiv:2605.15239v1 Announce Type: new Abstract: Safety alignment often improves robustness to harmful queries at the cost of reasoning ability, a tradeoff known as the safety tax. A common cause is di

model-releasesarxiv-cs-lg
18 May 2026
Model Releases

SGR: A Stepwise Reasoning Framework for LLMs with External Subgraph Generation

DGX agent

arXiv:2605.16117v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated strong capabilities across diverse NLP applications, such as translation, text generation, and question a

model-releasesarxiv-cs-cl
18 May 2026
Model Releases

SurvivalPFN: Amortizing Survival Prediction via In-Context Bayesian Inference

DGX agent

arXiv:2605.15488v1 Announce Type: new Abstract: Survival analysis provides a powerful statistical framework for modeling time-to-event outcomes in the presence of censoring. However, selecting an appr

model-releasesarxiv-cs-lg
18 May 2026
Model Releases

Unlocking Dense Metric Depth Estimation in VLMs

DGX agent

arXiv:2605.15876v1 Announce Type: new Abstract: Vision-Language Models (VLMs) excel at 2D tasks such as grounding and captioning, yet remain limited in 3D understanding. A key limitation is their text

model-releasesarxiv-cs-cv
18 May 2026
Research

Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models

DGX agent

arXiv:2605.15082v1 Announce Type: cross Abstract: We study a prototypical situation when a learned predictor can discover useful low-dimensional structure in data, while using fewer samples than are n

researcharxiv-cs-lg
15 May 2026
Model Releases

Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment

DGX agent

arXiv:2605.14311v1 Announce Type: cross Abstract: Test-Time Scaling (TTS), which samples multiple candidate actions and ranks them via a Critic Model, has emerged as a promising paradigm for generalis

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Cognitive-Uncertainty Guided Knowledge Distillation for Accurate Classification of Student Misconceptions

DGX agent

arXiv:2605.14752v1 Announce Type: cross Abstract: Accurately identifying student misconceptions is crucial for personalized education but faces three challenges: (1) data scarcity with long-tail distr

model-releasesarxiv-cs-ai
15 May 2026
Research

FU-MPC: Frontier- and Uncertainty-Aware Model Predictive Control for Efficient and Accurate UAV Exploration with Motorized LiDAR

DGX agent

arXiv:2605.14920v1 Announce Type: new Abstract: Efficient UAV exploration in unknown environments requires rapid coverage expansion while maintaining accurate and reliable localization, since safe nav

researcharxiv-cs-ro
15 May 2026
Model Releases

L2R: Low-Rank and Lipschitz-Controlled Routing for Mixture-of-Experts

DGX agent

arXiv:2601.21349v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) models scale neural networks by conditionally activating a small subset of experts, where the router plays a central

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

ML-Embed: Inclusive and Efficient Embeddings for a Multilingual World

DGX agent

arXiv:2605.15081v1 Announce Type: cross Abstract: The development of high-quality text embeddings is increasingly drifting toward an exclusionary future, defined by three critical barriers: prohibitiv

model-releasesarxiv-cs-ai
15 May 2026
Agents

Modeling Bounded Rationality in Drug Shortage Pharmacists Using Attention-Guided Dynamic Decomposition

DGX agent

arXiv:2605.14111v1 Announce Type: new Abstract: Hospital pharmacists make high-stakes decisions to mitigate drug shortages under uncertainty, time pressure, and patient risk. Interviews revealed that

agentsarxiv-cs-ai
15 May 2026
Safety

One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries

DGX agent

arXiv:2605.14605v1 Announce Type: cross Abstract: Model providers increasingly release open weights or allow users to fine-tune foundation models through APIs. Although these models are safety-aligned

safetyarxiv-cs-ai
15 May 2026
Model Releases

PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts

DGX agent

arXiv:2605.14002v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) embedded in agentic frameworks have transformed information retrieval from static, long context question answering into op

model-releasesarxiv-cs-ai
15 May 2026
Research

Randomized Atomic Feature Models for Physics-Informed Identification of Dynamic Systems

DGX agent

arXiv:2605.14351v1 Announce Type: cross Abstract: We present a physics-informed framework for system identification based on randomized stable atomic features. Impulse responses are represented as ran

researcharxiv-cs-lg
15 May 2026
Research

Scaling Laws from Sequential Feature Recovery: A Solvable Hierarchical Model

DGX agent

arXiv:2605.14567v1 Announce Type: cross Abstract: We propose a simple mechanism by which scaling laws emerge from feature learning in multi-layer networks. We study a high-dimensional hierarchical tar

researcharxiv-cs-lg
15 May 2026
Model Releases

Think When Needed: Adaptive Reasoning-Driven Multimodal Embeddings with a Dual-LoRA Architecture

DGX agent

arXiv:2605.14448v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have emerged as a powerful backbone for multimodal embeddings. Recent methods introduce chain-of-thought (CoT

model-releasesarxiv-cs-cl
15 May 2026
Safety

Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation

DGX agent

arXiv:2602.02994v2 Announce Type: replace Abstract: Reinforcement learning has emerged as a principled post-training paradigm for Temporal Video Grounding (TVG) due to its on-policy optimization, yet

safetyarxiv-cs-cv
15 May 2026
Research

Backbone is All You Need: Assessing Vulnerabilities of Frozen Foundation Models in Synthetic Image Forensics

DGX agent

arXiv:2605.13381v1 Announce Type: new Abstract: As AI-generated synthetic images become increasingly realistic, Vision Transformers (ViTs) have emerged as a cornerstone of modern deepfake detection. H

researcharxiv-cs-cv
14 May 2026
Model Releases

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence

DGX agent

arXiv:2605.12882v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have significantly advanced document understanding, yet current Doc-VQA evaluations score only the final answ

model-releasesarxiv-cs-cv
14 May 2026
Model Releases

Code-Centric Detection of Vulnerability-Fixing Commits: A Unified Benchmark and Empirical Study

DGX agent

arXiv:2605.13138v1 Announce Type: cross Abstract: Automated detection of vulnerability-fixing commits (VFCs) is critical for timely security patch deployment, as advisory databases lag patch releases

model-releasesarxiv-cs-lg
14 May 2026
Local Ai

Comfyui error Missing Models (1)AttributeError: module 'tensorflow' has no attribute 'Tensor'

DGX agent

This error in ComfyUI occurs when the einops library encounters a broken TensorFlow installation while performing multi-backend type checks, even though ComfyUI itself is PyTorch-based. The primary so

local-air-stablediffusion
14 May 2026
Model Releases

From Pixels to BFS: High Maze Accuracy Does Not Imply Visual Planning

DGX agent

arXiv:2603.26839v2 Announce Type: replace-cross Abstract: How do multimodal models solve visual spatial tasks -- through genuine planning, or through brute-force search in token space? We introduce ex

model-releasesarxiv-cs-cv
14 May 2026
Model Releases

GraphIP-Bench: How Hard Is It to Steal a Graph Neural Network, and Can We Stop It?

DGX agent

arXiv:2605.12827v1 Announce Type: cross Abstract: Graph neural networks (GNNs) deployed as cloud services can be stolen through model-extraction attacks, which train a surrogate from query responses t

model-releasesarxiv-cs-ai
14 May 2026
Local Ai

PRISM: Prior Rectification and Uncertainty-Aware Structure Modeling for Diffusion-Based Text Image Super-Resolution

DGX agent

arXiv:2605.13027v1 Announce Type: new Abstract: Text image super-resolution (Text-SR) requires more than visually plausible detail synthesis: slight errors in stroke topology may alter character ident

local-aiarxiv-cs-cv
14 May 2026
Model Releases

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos

DGX agent

arXiv:2512.01707v3 Announce Type: replace-cross Abstract: Streaming video understanding requires models not only to process temporally incoming frames, but also to anticipate user intention for realis

model-releasesarxiv-cs-ai
14 May 2026
← Previous
1…364365366367368…1324
Next →