AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
20 May 2026

Beyond Mode Collapse: Distribution Matching for Diverse Reasoning

SafetyDGX agent

arXiv:2605.19461v1 Announce Type: new Abstract: On-policy reinforcement learning methods like GRPO suffer from mode collapse: they exhibit reduced solution diversity, concentrating probability mass on

Beyond Prediction Accuracy: Target-Space Recovery Profiles for Evaluating Model-Brain Alignment

Model ReleasesDGX agent

arXiv:2605.20127v1 Announce Type: cross Abstract: Artificial vision models are often evaluated against the human visual cortex by measuring how accurately their internal representations predict brain

Beyond Rational Illusion: Behaviorally Realistic Strategic Classification

TutorialsDGX agent

arXiv:2605.19674v1 Announce Type: new Abstract: Strategic classification(SC) studies the interaction between decision models and agents who strategically manipulate their features for favorable outcom


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

BLINKG: A Benchmark for LLM-Integrated Knowledge Graph Generation

Model ReleasesDGX agent

arXiv:2605.19518v1 Announce Type: new Abstract: Generating Knowledge Graphs (KGs) remains one of the most time-consuming and labor-intensive tasks for knowledge engineers, as they need to identify sem

Block-Based Double Decoders

ResearchDGX agent

arXiv:2605.18807v1 Announce Type: cross Abstract: Encoder-decoder models offer substantial inference-time savings over decoder-only models, but their pretraining objectives suffer from sparse supervis

Block-Sphere Vector Quantization

ResearchDGX agent

arXiv:2605.19972v1 Announce Type: cross Abstract: Vector quantization is a fundamental primitive for scalable machine learning systems, enabling memory-efficient storage, fast retrieval, and compresse

Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay

SafetyDGX agent

arXiv:2605.19352v1 Announce Type: cross Abstract: Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the

Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models

ResearchDGX agent

arXiv:2605.19929v1 Announce Type: cross Abstract: Low-bit post-training quantization (PTQ) is a pivotal technique for deploying Vision-Language Models (VLMs) on resource-constrained devices. However,

Bridge: Retrieval-Augmented Spatiotemporal Modeling for Urban Delivery Demand

ApplicationsDGX agent

arXiv:2605.19172v1 Announce Type: cross Abstract: Forecasting urban delivery demand becomes substantially more challenging when newly added service regions lack historical records. Existing spatiotemp

BuildArena: A Physics-Aligned Interactive Benchmark of LLMs for Engineering Construction

Model ReleasesDGX agent

arXiv:2510.16559v5 Announce Type: replace Abstract: Engineering construction automation aims to transform natural language specifications into physically viable structures, requiring complex integrate

CADENet: Condition-Adaptive Asynchronous Dual-Stream Enhancement Network for Adverse Weather Perception in Autonomous Driving

SafetyDGX agent

arXiv:2605.19837v1 Announce Type: cross Abstract: Adverse weather (rain, fog, sand, and snow) degrades camera-based object detection in autonomous vehicles. Existing enhancement-then-detect approaches

Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses

SafetyDGX agent

arXiv:2605.19229v1 Announce Type: new Abstract: Survey research faces mounting structural challenges: declining response rates, sample bias, block-wise missingness among at-risk respondents, and AI-as

Can LLMs Emulate Human Belief Dynamics?

Model ReleasesDGX agent

arXiv:2605.18781v1 Announce Type: cross Abstract: Can LLMs simulate how humans form and change beliefs in social networks? We put this to the test by replicating an established study on belief dynamic

CANINE: Coaching Visually Impaired Users for Interactive Navigation with a Robot Guide Dog

TutorialsDGX agent

arXiv:2605.19501v1 Announce Type: cross Abstract: Robot guide dogs offer navigation assistance that greatly expands the independent mobility of the visually impaired, but their effective use requires

CaptchaMind: Training CAPTCHA Solvers via Reinforcement Learning with Explicit Reasoning Supervision

Model ReleasesDGX agent

arXiv:2605.19538v1 Announce Type: cross Abstract: CAPTCHAs are widely deployed as human verification mechanisms and frequently block intelligent agents from completing end-to-end automation in real-wo

Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination

Model ReleasesDGX agent

arXiv:2605.19250v1 Announce Type: new Abstract: Modality-conflict hallucination occurs when multimodal large language models (MLLMs) prioritize erroneous textual premises over contradictory visual evi

Chunking German Legal Code

Model ReleasesDGX agent

arXiv:2605.19806v1 Announce Type: cross Abstract: This paper investigates chunking strategies for retrieval-augmented generation on German statutory law, using the German Civil Code as a structured be

ClusterRAG: Cluster-Based Collaborative Filtering for Personalized Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2605.18769v1 Announce Type: cross Abstract: Personalized Retrieval-Augmented Generation (RAG) relies on accurately selecting user-relevant documents. In practice, existing RAG approaches often s

COBALT: Crowdsourcing Robot Learning via Cloud-Based Teleoperation with Smartphones

HardwareDGX agent

arXiv:2605.19138v1 Announce Type: cross Abstract: The scarcity of large-scale, high-quality demonstration data remains a bottleneck in scaling imitation learning for robotic manipulation. We present C

CogScale: Scalable Benchmark for Sequence Processing

Model ReleasesDGX agent

arXiv:2605.19758v1 Announce Type: new Abstract: The ability to maintain and manipulate information over time is a fundamental aspect of living beings and Artificial Intelligence. While modern models h

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning

SafetyDGX agent

arXiv:2507.15698v2 Announce Type: replace-cross Abstract: Process Reward Models (PRMs) play a central role in evaluating and guiding multi-step reasoning in large language models (LLMs), especially fo

Component-Aware Structure-Preserving Style Transfer for Satellite Sim2Real 6D Pose Estimation

ResearchDGX agent

arXiv:2605.19624v1 Announce Type: cross Abstract: Monocular 6D pose estimation for non-cooperative satellites depends heavily on annotated training data, yet real satellite images with reliable pose l

Composition of Memory Experts for Diffusion World Models

ApplicationsDGX agent

arXiv:2605.18813v1 Announce Type: cross Abstract: World models aim to predict plausible futures consistent with past observations, a capability central to planning and decision-making in reinforcement

Compositional Literary Primitives in Instruction-Tuned LLMs: Cross-Architectural SAE Features for Self, Style, and Affect

Model ReleasesDGX agent

arXiv:2605.18808v1 Announce Type: cross Abstract: We characterize a compositional architecture of literary primitives in two instruction-tuned large language models (Llama 3.1 8B-Instruct and Gemma 2

Concept-Guided Noisy Negative Suppression for Zero-Shot Classification and Grounding of Chest X-Ray Findings

SafetyDGX agent

arXiv:2605.19374v1 Announce Type: cross Abstract: Vision-language alignment using chest X-rays and radiology reports has emerged as an advanced paradigm for zero-shot classification and grounding of c

Conflict-Free Replicated Data Types for Neural Network Model Merging: A Two-Layer Architecture Enabling CRDT-Compliant Model Merging Across 26 Strategies

ApplicationsDGX agent

arXiv:2605.19373v1 Announce Type: cross Abstract: All 26 neural network merge strategies we tested including weight averaging, SLERP, TIES, DARE, Fisher merging, and evolutionary approaches -- fail th

Conflict-Resilient Multi-Agent Reasoning via Signed Graph Modeling

Model ReleasesDGX agent

arXiv:2605.19418v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) have demonstrated strong reasoning and decision-making capabilities that consistently surpass those of single LLM ag

ContextFlow: Hierarchical Task-State Alignment for Long-Horizon Embodied Agents

SafetyDGX agent

arXiv:2605.19314v1 Announce Type: cross Abstract: Long-horizon embodied agents increasingly delegate navigation, search, approach, and manipulation to specialist executors. As these executors become s

ContextRAG: Extraction-Free Hierarchical Graph Construction for Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2605.19735v1 Announce Type: cross Abstract: Graph-structured retrieval-augmented generation (RAG) systems can improve answer quality on multi-hop questions, but many current systems rely on larg

Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations

Model ReleasesDGX agent

arXiv:2603.17305v2 Announce Type: replace Abstract: We propose CRAFT, a red-teaming alignment framework that leverages model reasoning capabilities and hidden representations to improve robustness aga

CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning

SafetyDGX agent

arXiv:2605.20075v1 Announce Type: cross Abstract: Chain-of-thought (CoT) is a standard approach for eliciting reasoning capabilities from large language models (LLMs). However, the common CoT paradigm

Counterfactual Likelihood Tests for Indirect Influence in Private Reasoning Channels

ResearchDGX agent

arXiv:2605.19092v1 Announce Type: cross Abstract: Reasoning systems increasingly separate intermediate computation into private and public channels, creating evaluation cases that look similar in tran

CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation

ResearchDGX agent

arXiv:2605.18916v1 Announce Type: cross Abstract: We investigate Counterfactual Video Foley Generation, which aims to adopt a sound-source identity that contradicts the visual evidence while remaining

CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering

Model ReleasesDGX agent

arXiv:2605.19075v1 Announce Type: cross Abstract: Grounded multi-video question answering over real-world news events requires systems to surface query-relevant evidence across heterogeneous video arc

CriterAlign: Criterion-Centric Rationale Alignment for Code Preference Judging

SafetyDGX agent

arXiv:2605.19665v1 Announce Type: cross Abstract: Pairwise human preference prediction is central to evaluating code-generation systems, where quality often depends on task-specific trade-offs beyond

Cross-Subject Intracranial EEG Reconstruction from Scalp Recordings Using Multi-Scale Cross-Attention Transformers

ResearchDGX agent

arXiv:2605.18897v1 Announce Type: cross Abstract: Intracranial EEG (iEEG) provides high-fidelity neural recordings essential for clinical and brain-computer interface applications, but acquiring these

CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing

Model ReleasesDGX agent

arXiv:2605.19484v1 Announce Type: cross Abstract: While GUI agents have made significant progress in web navigation and basic operating system tasks, their capabilities in professional creative workfl

D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting

ResearchDGX agent

arXiv:2605.18810v1 Announce Type: cross Abstract: Speculative decoding accelerates LLM inference by having a small drafter propose tokens that a larger target model verifies in parallel. Recent diffus

DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models

ApplicationsDGX agent

arXiv:2605.18868v1 Announce Type: cross Abstract: While vision and multimodal foundation models underpin critical tasks from perception to complex reasoning, they remain highly vulnerable to adversari

Data-driven Acceleration of MPC with Guarantees

SafetyDGX agent

arXiv:2511.13588v2 Announce Type: replace-cross Abstract: Model Predictive Control (MPC) is a powerful framework for optimal control but can be too slow for low-latency applications. We present a data

Data-Free Client Contribution Estimation via Logit Maximization for Federated Learning

ResearchDGX agent

arXiv:2605.18892v1 Announce Type: cross Abstract: Federated learning (FL) enables collaborative learning of computer vision models, where privacy and regulatory constraints prevent centralizing data a

Decentralized autonomous organization and blockchain-based incentivization framework for community-based facilities management

AgentsDGX agent

arXiv:2605.18773v1 Announce Type: cross Abstract: Traditional facility management often relies on centralized decision-making structures that limit stakeholder participation, leading to misalignment w

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows

Model ReleasesDGX agent

arXiv:2605.19099v1 Announce Type: new Abstract: We introduce DecisionBench, a benchmark substrate for emergent delegation in long-horizon agentic workflows. The substrate fixes a task suite (GAIA, tau

Deep Neural Network for Musical Instrument Recognition using MFCCs

ResearchDGX agent

arXiv:2105.00933v3 Announce Type: replace-cross Abstract: The task of efficient automatic music classification is of vital importance and forms the basis for various advanced applications of AI in the

Deep Tech to Space: Space Data Centers and AI Revolution at the Edge

ResearchDGX agent

arXiv:2605.19892v1 Announce Type: cross Abstract: Dramatic cost reductions driven by private sector innovations have led to a rapid increase in the number of satellites in orbit and a corresponding su

DEFLECT: Delay-Robust Execution via Flow-matching Likelihood-Estimated Counterfactual Tuning for VLA Policies

SafetyDGX agent

arXiv:2605.19294v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies are typically deployed with asynchronous inference: the robot executes a previously predicted action chunk while

Detecting Fluent Optimization-Based Adversarial Prompts via Sequential Entropy Changes

Model ReleasesDGX agent

arXiv:2605.19966v1 Announce Type: cross Abstract: Optimization-based adversarial suffixes can jailbreak aligned large language models (LLMs) while remaining fluent, weakening static and windowed perpl

Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution

TutorialsDGX agent

arXiv:2605.19228v1 Announce Type: cross Abstract: Large Language Models have achieved strong performance on reasoning tasks with objective answers by generating step-by-step solutions, but diagnosing

Digital Voices of Survival: From Social Media Disclosures to Support Provisions for Domestic Violence Victims

ResearchDGX agent

arXiv:2509.12288v2 Announce Type: replace-cross Abstract: Domestic Violence (DV) is a pervasive public health problem characterized by patterns of coercive and abusive behavior within intimate relatio

Dimensional Balance Improves Large Scale Spatiotemporal Prediction Performance

SafetyDGX agent

arXiv:2605.18793v1 Announce Type: cross Abstract: Accurate spatiotemporal pattern analysis is critical in fields such as urban traffic, meteorology, and public health monitoring. However, existing met

Discoverable Agent Knowledge -- A Formal Framework for Agentic KG Affordances (Extended Version)

AgentsDGX agent

arXiv:2605.19186v1 Announce Type: new Abstract: Two decades ago, the Semantic Web Services community was asked how agents with different ontological commitments could discover, compose, and invoke web

Disentangling generalization and memorization in large language models using chess

Model ReleasesDGX agent

arXiv:2601.16823v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) exhibit remarkable capabilities, yet it remains unclear to what extent these reflect sophisticated recall or genu

Distilling Linearized Behavior for Effective Task Arithmetic

Model ReleasesDGX agent

arXiv:2605.18993v1 Announce Type: cross Abstract: Task vector composition has emerged as a promising paradigm for editing pre-trained models, enabling model merging through addition and unlearning thr

Distribution-Free Uncertainty Quantification for Continuous AI Agent Evaluation

Model ReleasesDGX agent

arXiv:2605.19779v1 Announce Type: new Abstract: We adapt split conformal prediction and adaptive conformal inference (ACI) to continuous AI agent evaluation, providing distribution-free coverage guara

Distributional AGI Safety

SafetyDGX agent

arXiv:2512.16856v2 Announce Type: replace Abstract: AI safety and alignment research has predominantly been focused on methods for safeguarding individual AI systems, resting on the assumption of an e

Distributional Energy-Based Models for Uncertainty-Aware Structured LLM Reasoning

Model ReleasesDGX agent

arXiv:2605.18871v1 Announce Type: cross Abstract: When Large Language Models produce structured outputs such as travel plans, code solutions, or multi-step proofs, individual reasoning steps may appea

DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model

Model ReleasesDGX agent

arXiv:2602.23622v2 Announce Type: replace-cross Abstract: Significant progress has been made in the field of Instruction-based Image Editing Models (IIEMs). However, while these models demonstrate pla

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs

Model ReleasesDGX agent

arXiv:2605.18915v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are vulnerable to jailbreak attacks, which can elicit harmful responses from MLLMs. Many MLLMs support multi-

Does Code Cleanliness Affect Coding Agents? A Controlled Minimal-Pair Study

Model ReleasesDGX agent

arXiv:2605.20049v1 Announce Type: cross Abstract: As autonomous coding agents see rapid adoption, their evaluation has primarily focused on task completion rates holding the target codebase fixed. Thi

Does Your Wildfire Prediction Model Actually Work, or Just Score Well?

ResearchDGX agent

arXiv:2605.18911v1 Announce Type: cross Abstract: Wildfire prediction is important for early warning and resource allocation, yet existing Earth foundation models (Earth FMs) are pretrained for genera

← Previous
1…226227228229230…358
Next →