AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
23 Apr 2026

IMPACT-CYCLE: A Contract-Based Multi-Agent System for Claim-Level Supervisory Correction of Long-Video Semantic Memory

Model ReleasesDGX agent

arXiv:2604.20136v1 Announce Type: cross Abstract: Correcting errors in long-video understanding is disproportionately costly: existing multimodal pipelines produce opaque, end-to-end outputs that expo

Improving Molecular Force Fields with Minimal Temporal Information

ResearchDGX agent

arXiv:2604.19806v1 Announce Type: cross Abstract: Accurate prediction of energy and forces for 3D molecular systems is one of fundamental challenges at the core of AI for Science applications. Many po

Infection-Reasoner: A Compact Vision-Language Model for Wound Infection Classification with Evidence-Grounded Clinical Reasoning

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2604.19937v1 Announce Type: cross Abstract: Assessing chronic wound infection from photographs is challenging because visual appearance varies across wound etiologies, anatomical locations, and

Inference Headroom Ratio: A Diagnostic and Control Framework for Inference Stability Under Constraint

ResearchDGX agent

arXiv:2604.19760v1 Announce Type: new Abstract: We present a simulation-based evaluation of the Inference Headroom Ratio (IHR), a dimensionless diagnostic quantity for characterizing inference stabili

Information Aggregation with AI Agents

ResearchDGX agent

arXiv:2604.20050v1 Announce Type: cross Abstract: Can Large Language Models (AI agents) aggregate dispersed private information through trading and reason about the knowledge of others by observing pr

Interval POMDP Shielding for Imperfect-Perception Agents

SafetyDGX agent

arXiv:2604.20728v1 Announce Type: new Abstract: Autonomous systems that rely on learned perception can make unsafe decisions when sensor readings are misclassified. We study shielding for this setting

Is Four Enough? Automated Reasoning Approaches and Dual Bounds for Condorcet Dimensions of Elections

ResearchDGX agent

arXiv:2604.19851v1 Announce Type: cross Abstract: In an election where n voters rank m candidates, a Condorcet winning set is a committee of k candidates such that for any outside candidate, a majorit

ItemRAG: Item-Based Retrieval-Augmented Generation for LLM-Based Recommendation

ResearchDGX agent

arXiv:2511.15141v2 Announce Type: replace-cross Abstract: Recently, large language models (LLMs) have been widely used as recommender systems, owing to their reasoning capability and effectiveness in

IVY-FAKE: A Unified Explainable Framework and Benchmark for Image and Video AIGC Detection

Model ReleasesDGX agent

arXiv:2506.00979v5 Announce Type: replace-cross Abstract: The rapid development of Artificial Intelligence Generated Content (AIGC) techniques has enabled the creation of high-quality synthetic conten

JTPRO: A Joint Tool-Prompt Reflective Optimization Framework for Language Agents

AgentsDGX agent

arXiv:2604.19821v1 Announce Type: new Abstract: Large language model (LLM) agents augmented with external tools often struggle as number of tools grow large and become domain-specific. In such setting

Knowledge Capsules: Structured Nonparametric Memory Units for LLMs

Model ReleasesDGX agent

arXiv:2604.20487v1 Announce Type: cross Abstract: Large language models (LLMs) encode knowledge in parametric weights, making it costly to update or extend without retraining. Retrieval-augmented gene

KoALa-Bench: Evaluating Large Audio Language Models on Korean Speech Understanding and Faithfulness

Model ReleasesDGX agent

arXiv:2604.19782v1 Announce Type: cross Abstract: Recent advances in large audio language models (LALMs) have enabled multilingual speech understanding. However, benchmarks for evaluating LALMs remain

KOCO-BENCH: Can Large Language Models Leverage Domain Knowledge in Software Development?

Model ReleasesDGX agent

arXiv:2601.13240v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel at general programming but struggle with domain-specific software development, necessitating domain special

Language Models Learn Universal Representations of Numbers and Here's Why You Should Care

TutorialsDGX agent

arXiv:2510.26285v2 Announce Type: replace-cross Abstract: Prior work has shown that large language models (LLMs) often converge to accurate input embedding for numbers, based on sinusoidal representat

LaplacianFormer:Rethinking Linear Attention with Laplacian Kernel

HardwareDGX agent

arXiv:2604.20368v1 Announce Type: cross Abstract: The quadratic complexity of softmax attention presents a major obstacle for scaling Transformers to high-resolution vision tasks. Existing linear atte

Large Language Models Meet Biomedical Knowledge Graphs for Mechanistically Grounded Therapeutic Prioritization

Model ReleasesDGX agent

arXiv:2604.19815v1 Announce Type: new Abstract: Drug repurposing is often framed as a candidate identification task, but existing approaches provide limited guidance for distinguishing biologically pl

Large Language Models Outperform Humans in Fraud Detection and Resistance to Motivated Investor Pressure

Model ReleasesDGX agent

arXiv:2604.20652v1 Announce Type: new Abstract: Large language models trained on human feedback may suppress fraud warnings when investors arrive already persuaded of a fraudulent opportunity. We test

LayerTracer: A Joint Task-Particle and Vulnerable-Layer Analysis framework for Arbitrary Large Language Model Architectures

Model ReleasesDGX agent

arXiv:2604.20556v1 Announce Type: cross Abstract: Currently, Large Language Models (LLMs) feature a diversified architectural landscape, including traditional Transformer, GateDeltaNet, and Mamba. How

LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning

Local AiDGX agent

arXiv:2603.06870v2 Announce Type: replace Abstract: Long-horizon execution in Large Language Models (LLMs) remains unstable even when high-level strategies are provided. Evaluating on controlled algor

Learning Multi-Modal Whole-Body Control for Real-World Humanoid Robots

ApplicationsDGX agent

arXiv:2408.07295v4 Announce Type: replace-cross Abstract: A major challenge in humanoid robotics is designing a unified interface for commanding diverse whole-body behaviors, from precise footstep seq

Learning to Evolve: A Self-Improving Framework for Multi-Agent Systems via Textual Parameter Graph Optimization

Model ReleasesDGX agent

arXiv:2604.20714v1 Announce Type: new Abstract: Designing and optimizing multi-agent systems (MAS) is a complex, labor-intensive process of 'Agent Engineering.' Existing automatic optimization methods

Learning to Solve the Quadratic Assignment Problem with Warm-Started MCMC Finetuning

ApplicationsDGX agent

arXiv:2604.20109v1 Announce Type: cross Abstract: The quadratic assignment problem (QAP) is a fundamental NP-hard task that poses significant challenges for both traditional heuristics and modern lear

Learning When Not to Decide: A Framework for Overcoming Factual Presumptuousness in AI Adjudication

Model ReleasesDGX agent

arXiv:2604.19895v1 Announce Type: new Abstract: A well-known limitation of AI systems is presumptuousness: the tendency of AI systems to provide confident answers when information may be lacking. This

Lightweight LLM Agent Memory with Small Language Models

AgentsDGX agent

arXiv:2604.07798v3 Announce Type: replace Abstract: Although LLM agents can leverage tools for complex tasks, they still need memory to maintain cross-turn consistency and accumulate reusable informat

LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent

Model ReleasesDGX agent

arXiv:2604.17931v2 Announce Type: replace Abstract: Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research remains

LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals

AgentsDGX agent

arXiv:2411.10109v2 Announce Type: replace Abstract: Machine learning can predict human behavior well when substantial structured data and well-defined outcomes are available, but these models are typi

LLM Agents Predict Social Media Reactions but Do Not Outperform Text Classifiers: Benchmarking Simulation Accuracy Using 120K+ Personas of 1511 Humans

SafetyDGX agent

arXiv:2604.19787v1 Announce Type: cross Abstract: Social media platforms mediate how billions form opinions and engage with public discourse. As autonomous AI agents increasingly participate in these

LLM-guided phase diagram construction through high-throughput experimentation

Model ReleasesDGX agent

arXiv:2604.20304v1 Announce Type: cross Abstract: Constructing phase diagrams for multicomponent alloys requires extensive experimental measurements and is a time-consuming task. Here we investigate w

LLMs Can Get 'Brain Rot': A Pilot Study on Twitter/X

SafetyDGX agent

arXiv:2510.13928v2 Announce Type: replace-cross Abstract: We propose and test the LLM Brain Rot Hypothesis: continual exposure to junk web text induces lasting cognitive decline in large language mode

Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images

Local AiDGX agent

arXiv:2510.04225v2 Announce Type: replace-cross Abstract: The rapid growth of AI-generated imagery has blurred the boundary between real and synthetic content, raising practical concerns for digital i

Location-Aware Pretraining for Medical Difference Visual Question Answering

ResearchDGX agent

arXiv:2603.04950v2 Announce Type: replace-cross Abstract: Differential medical VQA models compare multiple images to identify clinically meaningful changes and rely on vision encoders to capture fine-

MambaLiteUNet: Cross-Gated Adaptive Feature Fusion for Robust Skin Lesion Segmentation

Model ReleasesDGX agent

arXiv:2604.20286v1 Announce Type: cross Abstract: Recent segmentation models have demonstrated promising efficiency by aggressively reducing parameter counts and computational complexity. However, the

Measuring Creativity in the Age of Generative AI: Distinguishing Human and AI-Generated Creative Performance in Hiring and Talent Systems

ResearchDGX agent

arXiv:2604.19799v1 Announce Type: cross Abstract: Generative AI is rapidly transforming how organizations create value and evaluate talent. While large language models enhance baseline output quality,

Measuring the Machine: Evaluating Generative AI as Pluralist Sociotechical Systems

Model ReleasesDGX agent

arXiv:2604.20545v1 Announce Type: new Abstract: In measurement theory, instruments do not simply record reality; they help constitute what is observed. The same holds for generative AI evaluation: ben

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills

SafetyDGX agent

arXiv:2604.20441v1 Announce Type: new Abstract: Background: Agent skills are increasingly deployed as modular, reusable capability units in AI agent systems. Medical research agent skills require safe

Membership Inference for Contrastive Pre-training Models with Text-only PII Queries

SafetyDGX agent

arXiv:2603.14222v2 Announce Type: replace-cross Abstract: Contrastive pretraining models such as CLIP and CLAP, serve as the ubiquitous perceptual backbones for modern multimodal large models, yet the

Memory-Augmented LLM-based Multi-Agent System for Automated Feature Generation on Tabular Data

AgentsDGX agent

arXiv:2604.20261v1 Announce Type: new Abstract: Automated feature generation extracts informative features from raw tabular data without manual intervention and is crucial for accurate, generalizable

Meta Additive Model: Interpretable Sparse Learning With Auto Weighting

ApplicationsDGX agent

arXiv:2604.20111v1 Announce Type: cross Abstract: Sparse additive models have attracted much attention in high-dimensional data analysis due to their flexible representation and strong interpretabilit

Meta-Tool: Efficient Few-Shot Tool Adaptation for Small Language Models

Model ReleasesDGX agent

arXiv:2604.20148v1 Announce Type: cross Abstract: Can small language models achieve strong tool-use performance without complex adaptation mechanisms? This paper investigates this question through Met

MetaboNet: The Largest Publicly Available Consolidated Dataset for Type 1 Diabetes Management

Model ReleasesDGX agent

arXiv:2601.11505v2 Announce Type: replace-cross Abstract: Progress in Type 1 Diabetes (T1D) algorithm development is limited by the fragmentation and lack of standardization across existing T1D manage

MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models

Model ReleasesDGX agent

arXiv:2604.19809v1 Announce Type: new Abstract: We introduce MIRROR, a benchmark comprising eight experiments across four metacognitive levels that evaluates whether large language models can use self

MirrorBench: Evaluating Self-centric Intelligence in MLLMs by Introducing a Mirror

Model ReleasesDGX agent

arXiv:2604.14785v2 Announce Type: replace Abstract: Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated remarkable advances in perception and reasoning, suggesting their poten

Mitigating Prompt-Induced Cognitive Biases in General-Purpose AI for Software Engineering

Model ReleasesDGX agent

arXiv:2604.16756v2 Announce Type: replace-cross Abstract: Prompt-induced cognitive biases are changes in a general-purpose AI (GPAI) system's decisions caused solely by biased wording in the input (e.

MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings

ResearchDGX agent

arXiv:2604.19902v1 Announce Type: cross Abstract: We present MMCORE, a unified framework designed for multimodal image generation and editing. MMCORE leverages a pre-trained Vision-Language Model (VLM

Model Capability Assessment and Safeguards for Biological Weaponization

Model ReleasesDGX agent

arXiv:2604.19811v1 Announce Type: cross Abstract: AI leaders and safety reports increasingly warn that advances in model reasoning may enable biological misuse, including by low-expertise users, while

Mol-Debate: Multi-Agent Debate Improves Structural Reasoning in Molecular Design

AgentsDGX agent

arXiv:2604.20254v1 Announce Type: new Abstract: Text-guided molecular design is a key capability for AI-driven drug discovery, yet it remains challenging to map sequential natural-language instruction

MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation

ResearchDGX agent

arXiv:2604.20468v1 Announce Type: cross Abstract: Industrial robot applications require increasingly flexible systems that non-expert users can easily adapt for varying tasks and environments. However

More Is Different: Toward a Theory of Emergence in AI-Native Software Ecosystems

AgentsDGX agent

arXiv:2604.19827v1 Announce Type: cross Abstract: Software engineering faces a fundamental challenge: multi-agent AI systems fail in ways that defy explanation by traditional theories. While individua

Mythos and the Unverified Cage: Z3-Based Pre-Deployment Verification for Frontier-Model Sandbox Infrastructure

Model ReleasesDGX agent

arXiv:2604.20496v1 Announce Type: cross Abstract: The April 2026 Claude Mythos sandbox escape exposed a critical weakness in frontier AI containment: the infrastructure surrounding advanced models rem

Neural posterior estimation of the neutrino direction in IceCube using transformer-encoded normalizing flows on the sphere

ResearchDGX agent

arXiv:2604.19846v1 Announce Type: cross Abstract: IceCube is a cubic-kilometer-scale neutrino detector located at the geographic South Pole. A precise directional reconstruction of IceCube neutrinos i

NeuroSymActive: Differentiable Neural-Symbolic Reasoning with Active Exploration for Knowledge Graph Question Answering

SafetyDGX agent

arXiv:2602.15353v2 Announce Type: replace-cross Abstract: Large pretrained language models and neural reasoning systems have advanced many natural language tasks, yet they remain challenged by knowled

No More Marching: Learning Humanoid Locomotion for Short-Range SE(2) Targets

ApplicationsDGX agent

arXiv:2508.14098v2 Announce Type: replace-cross Abstract: Humanoids operating in real-world workspaces must frequently execute task-driven, short-range movements to SE(2) target poses. To be practical

Normalizing Flows with Iterative Denoising

ResearchDGX agent

arXiv:2604.20041v1 Announce Type: cross Abstract: Normalizing Flows (NFs) are a classical family of likelihood-based methods that have received revived attention. Recent efforts such as TARFlow have s

OISMA: On-the-fly In-memory Stochastic Multiplication Architecture for Matrix-Multiplication Workloads

ResearchDGX agent

arXiv:2508.08822v2 Announce Type: replace-cross Abstract: Artificial intelligence (AI) models are currently driven by a significant upscaling of their complexity, with massive matrix-multiplication wo

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model

Model ReleasesDGX agent

arXiv:2604.20806v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) have made substantial advances in reasoning tasks at the Olympiad level. Nevertheless, current Olympiad-level mul

On-Meter Graph Machine Learning: A Case Study of PV Power Forecasting for Grid Edge Intelligence

ApplicationsDGX agent

arXiv:2604.19800v1 Announce Type: cross Abstract: This paper presents a detailed study of how graph neural networks can be used on edge intelligent meters in a microgrid to forecast photovoltaic power

On the Existence of Universal Simulators of Attention

ResearchDGX agent

arXiv:2506.18739v2 Announce Type: replace-cross Abstract: Previous work on the learnability of transformers extemdash focused primarily on examining their ability to approximate specific algorithmic p

On the Stability and Generalization of First-order Bilevel Minimax Optimization

ResearchDGX agent

arXiv:2604.20115v1 Announce Type: cross Abstract: Bilevel optimization and bilevel minimax optimization have recently emerged as unifying frameworks for a range of machine-learning tasks, including hy

ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence

Model ReleasesDGX agent

arXiv:2604.20719v1 Announce Type: cross Abstract: Omnimodal Notation Processing (ONP) represents a unique frontier for omnimodal AI due to the rigorous, multi-dimensional alignment required across aud

Onyx: Cost-Efficient Disk-Oblivious ANN Search

ResearchDGX agent

arXiv:2604.20401v1 Announce Type: cross Abstract: Approximate nearest neighbor (ANN) search in AI systems increasingly handles sensitive data on third-party infrastructure. Trusted execution environme

← Previous
1…313314315316317…354
Next →