AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
15 Apr 2026

Should There be a Teacher In-the-Loop? A Study of Generative AI Personalized Tasks Middle School

ResearchDGX agent

arXiv:2602.15876v1 Announce Type: cross Abstract: Adapting instruction to the fine-grained needs of individual students is a powerful application of recent advances in large language models. These gen

Siamese Foundation Models for Crystal Structure Prediction

TutorialsDGX agent

arXiv:2503.10471v2 Announce Type: replace-cross Abstract: Predicting crystal structures from chemical compositions is a fundamental challenge in materials discovery, complicated by complex 3D geometri

Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2603.01045v2 Announce Type: replace-cross Abstract: Large language models are increasingly deployed in multi-agent systems to overcome context limitations by distributing information across agen

Simulation as Supervision: Mechanistic Pretraining for Scientific Discovery

SafetyDGX agent

arXiv:2507.08977v4 Announce Type: replace-cross Abstract: Scientific modeling faces a tradeoff between the interpretability of mechanistic theory and the predictive power of machine learning. While ex

SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents

Model ReleasesDGX agent

arXiv:2604.12040v1 Announce Type: cross Abstract: We present SIR-Bench, a benchmark of 794 test cases for evaluating autonomous security incident response agents that distinguishes genuine forensic in

SmellNet: A Large-scale Dataset for Real-world Smell Recognition

ApplicationsDGX agent

arXiv:2506.00239v5 Announce Type: replace Abstract: The ability of AI to sense and identify various substances based on their smell alone can have profound impacts on allergen detection (e.g. smelling

SOAR: Self-Correction for Optimal Alignment and Refinement in Diffusion Models

SafetyDGX agent

arXiv:2604.12617v1 Announce Type: cross Abstract: The post-training pipeline for diffusion models currently has two stages: supervised fine-tuning (SFT) on curated data and reinforcement learning (RL)

Social Learning Strategies for Evolved Virtual Soft Robots

TutorialsDGX agent

arXiv:2604.12482v1 Announce Type: cross Abstract: Optimizing the body and brain of a robot is a coupled challenge: the morphology determines what control strategies are effective, while the control pa

Socrates Loss: Unifying Confidence Calibration and Classification by Leveraging the Unknown

Model ReleasesDGX agent

arXiv:2604.12245v1 Announce Type: cross Abstract: Deep neural networks, despite their high accuracy, often exhibit poor confidence calibration, limiting their reliability in high-stakes applications.

SpanKey: Dynamic Key Space Conditioning for Neural Network Access Control

ResearchDGX agent

arXiv:2604.12254v1 Announce Type: cross Abstract: SpanKey is a lightweight way to gate inference without encrypting weights or chasing leaderboard accuracy on gated inference. The idea is to condition

Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks

Model ReleasesDGX agent

arXiv:2604.12102v1 Announce Type: new Abstract: We introduce compute-grounded reasoning (CGR), a design paradigm for spatial-aware research agents in which every answerable sub-problem is resolved by

SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration

ResearchDGX agent

arXiv:2604.12247v1 Announce Type: cross Abstract: Speculative decoding has emerged as a promising approach to accelerate autoregressive inference in large language models (LLMs). Self-draft methods, w

SpecBranch: Speculative Decoding via Hybrid Drafting and Rollback-Aware Branch Parallelism

ApplicationsDGX agent

arXiv:2506.01979v4 Announce Type: replace-cross Abstract: Recently, speculative decoding (SD) has emerged as a promising technique to accelerate LLM inference by employing a small draft model to propo

StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback

SafetyDGX agent

arXiv:2510.20093v2 Announce Type: replace-cross Abstract: Although recent advancements in diffusion models have significantly enriched the quality of generated images, challenges remain in synthesizin

Synthetic POMDPs to Challenge Memory-Augmented RL: Memory Demand Structure Modeling

ResearchDGX agent

arXiv:2508.04282v3 Announce Type: replace Abstract: Recent benchmarks for memory-augmented reinforcement learning (RL) have introduced partially observable Markov decision process (POMDP) environments

Technical Report -- A Context-Sensitive Multi-Level Similarity Framework for First-Order Logic Arguments: An Axiomatic Study

ResearchDGX agent

arXiv:2604.12534v1 Announce Type: new Abstract: Similarity in formal argumentation has recently gained attention due to its significance in problems such as argument aggregation in semantics and enthy

TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs

SafetyDGX agent

arXiv:2604.12232v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed across diverse domains, yet their vulnerability to jailbreak attacks, where adversarial inputs

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment

SafetyDGX agent

arXiv:2604.12116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as tool-augmented agents capable of executing system-level operations. While existing benchmarks

The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break

Model ReleasesDGX agent

arXiv:2604.11978v1 Announce Type: new Abstract: Large language model (LLM) agents perform strongly on short- and mid-horizon tasks, but often break down on long-horizon tasks that require extended, in

The Non-Optimality of Scientific Knowledge: Path Dependence, Lock-In, and The Local Minimum Trap

ResearchDGX agent

arXiv:2604.11828v1 Announce Type: new Abstract: Science is widely regarded as humanity's most reliable method for uncovering truths about the natural world. Yet the trajectory of scientific discovery

The Second Challenge on Cross-Domain Few-Shot Object Detection at NTIRE 2026: Methods and Results

ResearchDGX agent

arXiv:2604.11998v1 Announce Type: cross Abstract: Cross-domain few-shot object detection (CD-FSOD) remains a challenging problem for existing object detectors and few-shot learning approaches, particu

The Stackelberg Speaker: Optimizing Persuasive Communication in Social Deduction Games

SafetyDGX agent

arXiv:2510.09087v2 Announce Type: replace Abstract: Large language model (LLM) agents have shown remarkable progress in social deduction games (SDGs). However, existing approaches primarily focus on i

Thermodynamic Liquid Manifold Networks: Physics-Bounded Deep Learning for Solar Forecasting in Autonomous Off-Grid Microgrids

AgentsDGX agent

arXiv:2604.11909v1 Announce Type: cross Abstract: The stable operation of autonomous off-grid photovoltaic systems requires solar forecasting algorithms that respect atmospheric thermodynamics. Contem

Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training

SafetyDGX agent

arXiv:2509.25758v2 Announce Type: replace Abstract: The remarkable capabilities of modern large reasoning models are largely unlocked through post-training techniques such as supervised fine-tuning (S

TimeSAF: Towards LLM-Guided Semantic Asynchronous Fusion for Time Series Forecasting

TutorialsDGX agent

arXiv:2604.12648v1 Announce Type: cross Abstract: Despite the recent success of large language models (LLMs) in time-series forecasting, most existing methods still adopt a Deep Synchronous Fusion str

Topology-Aware Reasoning over Incomplete Knowledge Graph with Graph-Based Soft Prompting

ResearchDGX agent

arXiv:2604.12503v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown remarkable capabilities across various tasks but remain prone to hallucinations in knowledge-intensive scenari

Towards grounded autonomous research: an end-to-end LLM mini research loop on published computational physics

AgentsDGX agent

arXiv:2604.12198v1 Announce Type: cross Abstract: Recent autonomous LLM agents have demonstrated end-to-end automation of machine-learning research. Real-world physical science is intrinsically harder

Towards Long-horizon Agentic Multimodal Search

Model ReleasesDGX agent

arXiv:2604.12890v1 Announce Type: cross Abstract: Multimodal deep search agents have shown great potential in solving complex tasks by iteratively collecting textual and visual evidence. However, mana

Towards Platonic Representation for Table Reasoning: A Foundation for Permutation-Invariant Retrieval

SafetyDGX agent

arXiv:2604.12133v1 Announce Type: new Abstract: Historical approaches to Table Representation Learning (TRL) have largely adopted the sequential paradigms of Natural Language Processing (NLP). We argu

Transferable Expertise for Autonomous Agents via Real-World Case-Based Learning

Model ReleasesDGX agent

arXiv:2604.12717v1 Announce Type: new Abstract: LLM-based autonomous agents perform well on general reasoning tasks but still struggle to reliably use task structure, key constraints, and prior experi

TRUST Agents: A Collaborative Multi-Agent Framework for Fake News Detection, Explainable Verification, and Logic-Aware Claim Reasoning

Model ReleasesDGX agent

arXiv:2604.12184v1 Announce Type: new Abstract: TRUST Agents is a collaborative multi-agent framework for explainable fact verification and fake news detection. Rather than treating verification as a

Turbo-DDCM: Fast and Flexible Zero-Shot Diffusion-Based Image Compression

ResearchDGX agent

arXiv:2511.06424v2 Announce Type: replace-cross Abstract: While zero-shot diffusion-based compression methods have seen significant progress in recent years, they remain notoriously slow and computati

Understanding or Memorizing? A Case Study of German Definite Articles in Language Models

Model ReleasesDGX agent

arXiv:2601.09313v2 Announce Type: replace-cross Abstract: Language models perform well on grammatical agreement, but it is unclear whether this reflects rule-based generalization or memorization. We s

Unveiling the Surprising Efficacy of Navigation Understanding in End-to-End Autonomous Driving

Local AiDGX agent

arXiv:2604.12208v1 Announce Type: cross Abstract: Global navigation information and local scene understanding are two crucial components of autonomous driving systems. However, our experimental result

Variation in Verification: Understanding Verification Dynamics in Large Language Models

Model ReleasesDGX agent

arXiv:2509.17995v2 Announce Type: replace-cross Abstract: Recent advances have shown that scaling test-time computation enables large language models (LLMs) to solve increasingly complex problems acro

Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices

Model ReleasesDGX agent

arXiv:2512.06443v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed on edge devices. To meet strict resource constraints, real-world deployment has pushed

VFA: Relieving Vector Operations in Flash Attention with Global Maximum Pre-computation

HardwareDGX agent

arXiv:2604.12798v1 Announce Type: cross Abstract: FlashAttention-style online softmax enables exact attention computation with linear memory by streaming score tiles through on-chip memory and maintai

VISTA: Validation-Informed Trajectory Adaptation via Self-Distillation

ResearchDGX agent

arXiv:2604.12044v1 Announce Type: cross Abstract: Deep learning models may converge to suboptimal solutions despite strong validation accuracy, masking an optimization failure we term Trajectory Devia

Visual Preference Optimization with Rubric Rewards

Model ReleasesDGX agent

arXiv:2604.13029v1 Announce Type: cross Abstract: The effectiveness of Direct Preference Optimization (DPO) depends on preference data that reflect the quality differences that matter in multimodal ta

WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces

SafetyDGX agent

arXiv:2603.05295v3 Announce Type: replace Abstract: We introduce WebChain, the largest open-source dataset of human-annotated trajectories on real-world websites, designed to accelerate reproducible r

WebFactory: Automated Compression of Foundational Language Intelligence into Grounded Web Agents

AgentsDGX agent

arXiv:2603.05044v2 Announce Type: replace Abstract: Current paradigms for training GUI agents are fundamentally limited by a reliance on either unsafe, non-reproducible live web interactions or costly

When Does Data Augmentation Help? Evaluating LLM and Back-Translation Methods for Hausa and Fongbe NLP

Model ReleasesDGX agent

arXiv:2604.12540v1 Announce Type: cross Abstract: Data scarcity limits NLP development for low-resource African languages. We evaluate two data augmentation methods -- LLM-based generation (Gemini 2.5

When to Forget: A Memory Governance Primitive

AgentsDGX agent

arXiv:2604.12007v1 Announce Type: new Abstract: Agent memory systems accumulate experience but currently lack a principled operational metric for memory quality governance -- deciding which memories t

Why Did Apple Fall: Evaluating Curiosity in Large Language Models

TutorialsDGX agent

arXiv:2510.20635v2 Announce Type: replace-cross Abstract: Curiosity serves as a pivotal conduit for human beings to discover and learn new knowledge. Recent advancements of large language models (LLMs

WiseOWL: A Methodology for Evaluating Ontological Descriptiveness and Semantic Correctness for Ontology Reuse and Ontology Recommendations

SafetyDGX agent

arXiv:2604.12025v1 Announce Type: new Abstract: The Semantic Web standardizes concept meaning for humans and machines, enabling machine-operable content and consistent interpretation that improves adv

X-VC: Zero-shot Streaming Voice Conversion in Codec Space

Model ReleasesDGX agent

arXiv:2604.12456v1 Announce Type: cross Abstract: Zero-shot voice conversion (VC) aims to convert a source utterance into the voice of an unseen target speaker while preserving its linguistic content.

14 Apr 2026

3D-Anchored Lookahead Planning for Persistent Robotic Scene Memory via World-Model-Based MCTS

ResearchDGX agent

arXiv:2604.11302v1 Announce Type: cross Abstract: We present 3D-Anchored Lookahead Planning (3D-ALP), a System 2 reasoning engine for robotic manipulation that combines Monte Carlo Tree Search (MCTS)

A Benchmark for Gap and Overlap Analysis as a Test of KG Task Readiness

Model ReleasesDGX agent

arXiv:2604.10853v1 Announce Type: new Abstract: Task-oriented evaluation of knowledge graph (KG) quality increasingly asks whether an ontology-based representation can answer the competency questions

A collaborative agent with two lightweight synergistic models for autonomous crystal materials research

AgentsDGX agent

arXiv:2604.11540v1 Announce Type: new Abstract: Current large language models require hundreds of billions of parameters yet struggle with domain-specific reasoning and tool coordination in materials

A Compact and Efficient 1.251 Million Parameter Machine Learning CNN Model PD36-C for Plant Disease Detection: A Case Study

Model ReleasesDGX agent

arXiv:2604.11332v1 Announce Type: cross Abstract: Deep learning has markedly advanced image based plant disease diagnosis as improved hardware and dataset quality have enabled increasingly accurate ne

A Comparative Theoretical Analysis of Entropy Control Methods in Reinforcement Learning

SafetyDGX agent

arXiv:2604.09676v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a key approach for enhancing reasoning in large language models (LLMs), yet scalable training is often hindered

A Diffusion-Contrastive Graph Neural Network with Virtual Nodes for Wind Nowcasting in Unobserved Regions

TutorialsDGX agent

arXiv:2604.10328v1 Announce Type: cross Abstract: Accurate weather nowcasting remains one of the central challenges in atmospheric science, with critical implications for climate resilience, energy se

A Dual Cross-Attention Graph Learning Framework For Multimodal MRI-Based Major Depressive Disorder Detection

ResearchDGX agent

arXiv:2604.10116v1 Announce Type: cross Abstract: Major depressive disorder (MDD) is a prevalent mental disorder associated with complex neurobiological changes that cannot be fully captured using a s

A Dual-Positive Monotone Parameterization for Multi-Segment Bids and a Validity Assessment Framework for Reinforcement Learning Agent-based Simulation of Electricity Markets

SafetyDGX agent

arXiv:2604.10252v1 Announce Type: new Abstract: Reinforcement learning agent-based simulation (RL-ABS) has become an important tool for electricity market mechanism analysis and evaluation. In the mod

A Hybrid Intelligent Framework for Uncertainty-Aware Condition Monitoring of Industrial Systems

Model ReleasesDGX agent

arXiv:2604.09932v1 Announce Type: cross Abstract: Hybrid approaches that combine data-driven learning with physics-based insight have shown promise for improving the reliability of industrial conditio

A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs

ResearchDGX agent

arXiv:2604.09752v1 Announce Type: cross Abstract: During the deployment of Large Language Models (LLMs), the autoregressive decoding phase on heterogeneous NPU platforms (e.g., Ascend 910B) faces seve

A Mamba-Based Multimodal Network for Multiscale Blast-Induced Rapid Structural Damage Assessment

SafetyDGX agent

arXiv:2604.11709v1 Announce Type: new Abstract: Accurate and rapid structural damage assessment (SDA) is crucial for post-disaster management, helping responders prioritise resources, plan rescues, an

A Mathematical Explanation of Transformers

ResearchDGX agent

arXiv:2510.03989v2 Announce Type: replace-cross Abstract: The Transformer architecture has revolutionized the field of sequence modeling and underpins the recent breakthroughs in large language models

A mathematical theory of evolution for self-designing AIs

SafetyDGX agent

arXiv:2604.05142v2 Announce Type: replace Abstract: As artificial intelligence systems (AIs) become increasingly produced by recursive self-improvement, a form of evolution may emerge, with the traits

A Mechanistic Analysis of Looped Reasoning Language Models

TutorialsDGX agent

arXiv:2604.11791v1 Announce Type: cross Abstract: Reasoning has become a central capability in large language models. Recent research has shown that reasoning performance can be improved by looping an

← Previous
1…330331332333334…354
Next →