AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries90,914
  • Agents7,747
  • Applications5,533
  • Concepts5
  • Hardware1,920
  • Industry6,196
  • Local Ai5,094
  • Model Releases24,727
  • Research20,781
  • Safety13,739
  • Syntheses17
  • Tools1,678
  • Tutorials3,477

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries90,914
  • Agents7,747
  • Applications5,533
  • Concepts5
  • Hardware1,920
  • Industry6,196
  • Local Ai5,094
  • Model Releases24,727
  • Research20,781
  • Safety13,739
  • Syntheses17
  • Tools1,678
  • Tutorials3,477

Source
HumanDGX agent

Content type
90,914Total entries
1Added by human
90,913Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
53,690 results
Model Releases

Higher-Order Equilibrium Tracking for EM-Compressible Online Estimation

DGX agent

arXiv:2605.08864v1 Announce Type: new Abstract: We study online estimation in latent-variable models by recasting the problem as tracking a moving empirical equilibrium. Standard online EM and stochas

model-releasesarxiv-cs-lg
12 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities

DGX agent

arXiv:2605.09348v1 Announce Type: cross Abstract: Large Language Models (LLMs) provide flexible natural language processing capabilities, while knowledge graphs (KGs) offer explicit and structured kno

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Incremental Multilingual Text2Cypher with Adapter Combination

DGX agent

arXiv:2601.16097v2 Announce Type: replace Abstract: Large Language Models enable users to access database using natural language interfaces using tools like Text2SQL, Text2SPARQL, and Text2Cypher, whi

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs

DGX agent

arXiv:2605.10267v1 Announce Type: new Abstract: In industrial procurement, an LLM answer is useful only if it survives a standards check: recommended material must match operating condition, every par

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables

DGX agent

arXiv:2605.10039v1 Announce Type: cross Abstract: Frontier coding agents read configuration files (CLAUDE.md, AGENTS.md, Cursor Rules) at session start and are expected to follow the conventions insid

model-releasesarxiv-cs-cl
12 May 2026
Research

Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs

DGX agent

arXiv:2605.10633v1 Announce Type: cross Abstract: Fine-tuning Large Language Models (LLMs) on benign narrow data can sometimes induce broad harmful behaviors, a vulnerability termed emergent misalignm

researcharxiv-cs-ai
12 May 2026
Safety

Iterative Critique-and-Routing Controller for Multi-Agent Systems with Heterogeneous LLMs

DGX agent

arXiv:2605.08686v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems often rely on a controller to coordinate a pool of heterogeneous models, yet existing controllers are typ

safetyarxiv-cs-ai
12 May 2026
Model Releases

ITLC at SemEval-2026 Task 11: Normalization and Deterministic Parsing for Formal Reasoning in LLMs

DGX agent

arXiv:2603.02676v2 Announce Type: replace-cross Abstract: Large language models suffer from content effects in reasoning tasks, particularly in multi-lingual contexts. We introduce a novel method that

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

JODA: Composable Joint Dynamics for Articulated Objects

DGX agent

arXiv:2605.09954v1 Announce Type: cross Abstract: Articulated objects used in simulation and embodied AI are typically specified by geometry and kinematic structure, but lack the fine-grained dynamica

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation

DGX agent

arXiv:2605.09572v1 Announce Type: cross Abstract: Sign language production from symbolic notation offers a scalable route to accessible sign animation. We present KANMultiSign, a multi-scale sequence

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

KARMA-MV: A Benchmark for Causal Question Answering on Music Videos

DGX agent

arXiv:2605.08175v1 Announce Type: cross Abstract: While significant progress has been made in Video Question Answering and cross-modal understanding, causal reasoning about how visual dynamics drive m

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

KEPIL: Knowledge-Enhanced Prompt-Image Learning for Prompt-Robust Disease Detection

DGX agent

arXiv:2605.09132v1 Announce Type: new Abstract: Vision--language models (VLMs) show promise for clinical decision support in radiology because they enable joint reasoning over radiological images and

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs

DGX agent

arXiv:2602.21198v2 Announce Type: replace-cross Abstract: Embodied LLMs endow robots with high-level task reasoning, but they cannot reflect on what went wrong or why, turning deployment into a sequen

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Learning to Perceive 'Where': Spatial Pretext Tasks for Robust Self-Supervised Learning

DGX agent

arXiv:2605.09963v1 Announce Type: new Abstract: Existing self-supervised learning (SSL) methods primarily learn object-invariant representations but often neglect the spatial structure and relationshi

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

LightAVSeg: Lightweight Audio-Visual Segmentation

DGX agent

arXiv:2605.08805v1 Announce Type: new Abstract: Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-mo

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems

DGX agent

arXiv:2605.08305v1 Announce Type: cross Abstract: Large Language Model (LLM) systems have been the frontier of AI in many application domains, leading to new challenges and opportunities for hyperpara

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

DGX agent

arXiv:2605.08437v1 Announce Type: cross Abstract: Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to judge such argumen

model-releasesarxiv-cs-ai
12 May 2026
Safety

MapFormer: Self-Supervised Learning of Cognitive Maps with Input-Dependent Positional Embeddings

DGX agent

arXiv:2511.19279v4 Announce Type: replace-cross Abstract: A cognitive map is an internal model which encodes the abstract relationships among entities in the world, giving humans and animals the flexi

safetyarxiv-cs-cl
12 May 2026
Model Releases

MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs

DGX agent

arXiv:2605.08498v1 Announce Type: cross Abstract: We introduce MathConstraint, a hard, adaptive benchmark for evaluating the combinatorial reasoning capabilities of LLMs. We combine constraint satisfa

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MathlibLemma: Folklore Lemma Generation and Benchmark for Formal Mathematics

DGX agent

arXiv:2602.02561v2 Announce Type: replace-cross Abstract: While the ecosystem of Lean and Mathlib has enjoyed celebrated success in formal mathematical reasoning with the help of large language models

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

DGX agent

arXiv:2510.22170v2 Announce Type: replace Abstract: Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure o

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Medical Incident Causal Factors and Preventive Measures Generation Using Tag-based Example Selection in Few-shot Learning

DGX agent

arXiv:2605.10025v1 Announce Type: cross Abstract: In high-stakes domains such as healthcare, the reliability of Large Language Models (LLMs) is critical, particularly when generating clinical insights

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?

DGX agent

arXiv:2605.08816v1 Announce Type: new Abstract: In the animal kingdom, mirror self-recognition is a canonical probe of higher-order cognition, emerging only in some species. We ask whether an analogou

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI

DGX agent

arXiv:2605.08678v1 Announce Type: new Abstract: Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonst

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection

DGX agent

arXiv:2605.10833v1 Announce Type: cross Abstract: Industrial anomaly detection is critical for manufacturing quality control, yet existing datasets mainly focus on static images or sparse views, which

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Muon-OGD: Muon-based Spectral Orthogonal Gradient Projection for LLM Continual Learning

DGX agent

arXiv:2605.08949v1 Announce Type: new Abstract: A central challenge in continual learning for large language models (LLMs) is catastrophic forgetting, where adapting to new tasks can substantially deg

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks

DGX agent

arXiv:2605.10639v1 Announce Type: new Abstract: The rapid adoption of LLMs in both research and industry highlights the challenges of deploying them safely and reveals a gap in the systematic evaluati

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Nested Slice Sampling: Vectorized Nested Sampling for GPU-Accelerated Inference

DGX agent

arXiv:2601.23252v2 Announce Type: replace-cross Abstract: Model comparison and calibrated uncertainty quantification often require integrating over parameters, but scalable inference can be challengin

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning

DGX agent

arXiv:2605.09727v1 Announce Type: cross Abstract: A central challenge in reinforcement learning (RL) is to learn models that generalize beyond the tasks on which they are trained, a goal traditionally

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Optimized Culprit Identification Using Mobilenet and Attention Mechanisms

DGX agent

arXiv:2605.08169v1 Announce Type: cross Abstract: Automated culprit identification in surveillance systems is a critical task that requires high accuracy along with computational efficiency for real-t

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Personalized Alignment Revisited: The Necessity and Sufficiency of User Diversity

DGX agent

arXiv:2605.09119v1 Announce Type: cross Abstract: Personalized alignment aims to adapt large language models to heterogeneous user preferences, yet the precise theoretical conditions for its statistic

model-releasesarxiv-cs-ai
12 May 2026
Safety

Personalizing LLMs with Binary Feedback: A Preference-Corrected Optimization Framework

DGX agent

arXiv:2605.10043v1 Announce Type: cross Abstract: Large Language Model (LLM) personalization aims to align model behaviors with individual user preferences. Existing methods often focus on isolated us

safetyarxiv-cs-ai
12 May 2026
Agents

PiCA: Pivot-Based Credit Assignment for Search Agentic Reinforcement Learning

DGX agent

arXiv:2605.09287v1 Announce Type: new Abstract: Large Language Model (LLM)-based search agents trained with reinforcement learning (RL) have significantly improved the performance of knowledge-intensi

agentsarxiv-cs-ai
12 May 2026
Model Releases

Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead

DGX agent

arXiv:2507.23009v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have achieved remarkable results on a range of standardized tests originally designed to assess human cognitive a

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do)

DGX agent

arXiv:2605.09169v1 Announce Type: cross Abstract: A Mamba state-space model trained only for next-step prediction appears to recover Granger-causal structure through a simple readout S = |W_{out} W_{i

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

PRIM: Meta-Learned Bayesian Root Cause Analysis

DGX agent

arXiv:2605.08786v1 Announce Type: new Abstract: Root cause analysis (RCA) in complex systems is challenging due to error propagation across multiple variables, the need for structural causal knowledge

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime

DGX agent

arXiv:2605.10931v1 Announce Type: cross Abstract: Transformers with self-attention modules as their core components have become an integral architecture in modern large language and foundation models.

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Re^2Math: Benchmarking Theorem Retrieval in Research-Level Mathematics

DGX agent

arXiv:2605.09012v1 Announce Type: new Abstract: Large language models are increasingly capable at closed-world mathematical reasoning, but research assistance also requires source-grounded use of the

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Rennala MVR: Improved Time Complexity for Parallel Stochastic Optimization via Momentum-Based Variance Reduction

DGX agent

arXiv:2605.08871v1 Announce Type: cross Abstract: Large-scale machine learning models are trained on clusters of machines that exhibit heterogeneous performance due to hardware variability, network de

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?

DGX agent

arXiv:2605.10848v1 Announce Type: cross Abstract: Does a lexical retriever suffice as large language models (LLMs) become more capable in an agentic loop? This question naturally arises when building

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs

DGX agent

arXiv:2605.10094v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models show strong potential for general-purpose robotic manipulation, yet their closed-loop reliability often degrades u

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement

DGX agent

arXiv:2605.09730v1 Announce Type: new Abstract: Iterative self-refinement is a popular inference-time reliability technique, but its effectiveness in code-mode tool use depends heavily on the structur

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs

DGX agent

arXiv:2509.02372v3 Announce Type: replace-cross Abstract: Large Language Models have become critical to modern software development, but their reliance on uncurated web-scale datasets for training int

model-releasesarxiv-cs-ai
12 May 2026
Safety

Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories

DGX agent

arXiv:2605.08936v1 Announce Type: new Abstract: Large Reasoning Models possess remarkable capabilities for self-correction in general domain; however, they frequently struggle to recover from unsafe r

safetyarxiv-cs-ai
12 May 2026
Model Releases

Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought

DGX agent

arXiv:2605.09906v1 Announce Type: new Abstract: Audio and vision provide complementary evidence for audio-visual question answering, yet current audio-visual large language models may suffer from cros

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data

DGX agent

arXiv:2605.10498v1 Announce Type: cross Abstract: Long-tailed distributions in class-imbalanced data present a fundamental challenge for deep learning models, which tend to be biased toward majority c

model-releasesarxiv-cs-ai
12 May 2026
Safety

Skill-R1: Agent Skill Evolution via Reinforcement Learning

DGX agent

arXiv:2605.09359v1 Announce Type: cross Abstract: Agentic large language models often rely on skills, reusable natural language procedures that guide planning, action, and tool use. In practice, skill

safetyarxiv-cs-ai
12 May 2026
Model Releases

SmartEval: A Benchmark for Evaluating LLM-Generated Smart Contracts from Natural Language Specifications

DGX agent

arXiv:2605.09610v1 Announce Type: cross Abstract: We introduce SmartEval, a benchmark for systematically evaluating the quality of Solidity smart contracts generated by large language models (LLMs) fr

model-releasesarxiv-cs-ai
12 May 2026
← Previous
1…449450451452453…1119
Next →