AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
28 May 2026

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning

SafetyDGX agent

arXiv:2605.28023v1 Announce Type: cross Abstract: Visual captioning requires models to capture visual content faithfully while minimizing both omission and hallucination. As the dominant paradigm for

Verifiable Benchmarking of Long-Horizon Spatial Biology

Model ReleasesDGX agent

arXiv:2605.28065v1 Announce Type: new Abstract: AI agents are increasingly useful for biological data analysis, but existing benchmarks mostly test broad biological knowledge, executable workflows, or

Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs

SafetyDGX agent

arXiv:2605.28565v1 Announce Type: cross Abstract: Users of search-augmented LLMs rely on citations as evidence that responses are grounded in real sources, and rarely verify the cited pages themselves


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

VeriTrip: A Verifiable Benchmark for Travel Planning Agents over Unstructured Web Corpora

Model ReleasesDGX agent

arXiv:2605.28683v1 Announce Type: new Abstract: Existing benchmarks have laid the foundation for travel planning agents by establishing API-centric paradigms. However, as the capabilities of Autonomou

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild

Model ReleasesDGX agent

arXiv:2605.27882v1 Announce Type: cross Abstract: LLM-based agents score well on search benchmarks, yet real users consistently find results unsatisfying, revealing a persistent evaluation-experience

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer

ResearchDGX agent

arXiv:2605.28229v1 Announce Type: cross Abstract: With the rapid development of pre-training technologies, adapting large-scale Vision-Language Models (VLMs) for video understanding ie image-to-video

Visualizing Latent Phase Structures in Locomotion Policies: A Multi-Environment Study with Temporal Feature Extension

SafetyDGX agent

arXiv:2605.28186v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) has been shown to achieve high performance on locomotion control tasks in MuJoCo benchmarks such as HalfCheetah, Ant

VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs

Model ReleasesDGX agent

arXiv:2605.28422v1 Announce Type: cross Abstract: Latent reasoning enables reasoning over continuous hidden states rather than explicit tokens, avoiding the language bottleneck and inference overhead

Voluntary Collusion with Secret Tools in Competing LLM Agents

SafetyDGX agent

arXiv:2605.27593v1 Announce Type: new Abstract: Even when a tool is explicitly described as unfair and harmful to others, ostensibly safety-aligned LLM agents still voluntarily engage in secret collus

VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization

Model ReleasesDGX agent

arXiv:2511.11896v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently shown strong potential in vulnerability detection (VD). However, accurately detecting vulnerabiliti

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

Model ReleasesDGX agent

arXiv:2605.27851v1 Announce Type: new Abstract: Safety benchmark scores provide incomplete evidence of deployment readiness: aligned language models often adhere to rigid rules even when a situational

When Does Memory Help Multi-Trajectory Inference for Tool-Use LLM Agents?

AgentsDGX agent

arXiv:2605.28224v1 Announce Type: new Abstract: Multi-trajectory inference for tool-use LLM agents - generating multiple reasoning attempts and selecting among them - benefits from transferring knowle

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference

Local AiDGX agent

arXiv:2605.27435v1 Announce Type: cross Abstract: Deploying large language models (LLMs) on mobile devices increasingly relies on heterogeneous execution, yet no prior study has systematically charact

When prompt perturbations break your A/B test: A valid statistical test for generative surveying

ResearchDGX agent

arXiv:2605.27463v1 Announce Type: cross Abstract: Generative surveying -- where collections of LLM-based personas provide feedback on messages -- has emerged as a cheap and scalable alternative to tra

When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?

SafetyDGX agent

arXiv:2605.27932v1 Announce Type: cross Abstract: Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly underst

Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models

Local AiDGX agent

arXiv:2605.27997v1 Announce Type: cross Abstract: Large language models frequently generate toxic, hateful, or harmful content, yet existing mitigation methods rely on costly retraining or output-leve

Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR

SafetyDGX agent

arXiv:2605.28295v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) trains reasoning models without labeled trajectories, relying on grouped rollouts to expose the po

Who Uses AI? Platform Selection and the Measurement of Occupational AI Exposure

SafetyDGX agent

arXiv:2605.21743v2 Announce Type: replace Abstract: Conversation logs from AI platforms are increasingly used to measure occupational exposure to artificial intelligence, but the users observed in the

Whose Name Comes Up? III: Persona Prompting Effects in LLM-Based Scholar Recommendation

Model ReleasesDGX agent

arXiv:2605.28187v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as scholar recommenders, shaping who is seen as an expert in academia. Existing audits remain Engli

Why LLMs Fail at Causal Discovery and How Interventional Agents Escape

Model ReleasesDGX agent

arXiv:2605.27567v1 Announce Type: new Abstract: Causal discovery is a cornerstone of scientific reasoning, yet whether large language models can perform it reliably remains an open question. Recent be

Worker Disagreement Reveals Sharp Directions in Local SGD

ResearchDGX agent

arXiv:2605.27739v1 Announce Type: cross Abstract: Deep neural network training often exhibits highly anisotropic loss geometry, where a few sharp dominant Hessian directions coexist with a large flatt

You Are in Control of Your State: Why Human Outcomes Are Controllable Through Causal State Intervention

ApplicationsDGX agent

arXiv:2605.27580v1 Announce Type: new Abstract: A central puzzle for the behavioural sciences and for human-facing artificial intelligence is the persistence of within-person variability. The same ind

You Live More Than Once: Towards Hierarchical Skill Meta-Evolving

Model ReleasesDGX agent

arXiv:2605.28390v1 Announce Type: new Abstract: Test-time skill evolving is regarded as a new paradigm for enhancing deployed agentic systems. Existing works mainly focus on hard-coded skill evolving

Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training

ResearchDGX agent

arXiv:2605.28008v1 Announce Type: new Abstract: Large language models (LLMs) can now solve complex problems through long chain-of-thought (CoT) reasoning, but the trade-off between performance and tok

ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay

Model ReleasesDGX agent

arXiv:2605.28069v1 Announce Type: new Abstract: Adaptive context compression is vital for scaling Large Language Models (LLMs) to complex, multi-turn agent tasks. However, rule-based compression metho

27 May 2026

2-ASP(Q) programs with weak constraints: Complexity and efficient implementation

ResearchDGX agent

arXiv:2605.27338v1 Announce Type: new Abstract: ASP(Q) extends Answer Set Programming (ASP) with Quantifiers over answer sets. In this paper we focus on the class of ASP(Q) programs with two quantifie

A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks

Model ReleasesDGX agent

arXiv:2605.26747v1 Announce Type: new Abstract: Large Language Models (LLMs) have brought huge improvements to Artificial Intelligence (AI), which can be applied to general-purpose tasks. However, the

A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection

Model ReleasesDGX agent

arXiv:2605.26533v1 Announce Type: cross Abstract: Automated industrial inspection requires both precise defect localization and structured maintenance report generation; in current practice these task

A Physics-Informed Hierarchical Neural Network for Microwave Scattering Analysis of 3D PEC Targets

ResearchDGX agent

arXiv:2508.03774v5 Announce Type: replace-cross Abstract: Accurate modeling of scattering from three-dimensional (3D) perfectly electrically conducting (PEC) targets at microwave frequencies constitut

A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration

SafetyDGX agent

arXiv:2605.26174v1 Announce Type: cross Abstract: Production language-model systems answer a request by partitioning it across an invisible orchestration of worker agents that recompose one integrated

Adapting Actively on the Fly: Relevance-Guided Online Meta-Learning with Latent Concepts for Geospatial Discovery

ApplicationsDGX agent

arXiv:2602.17605v2 Announce Type: replace-cross Abstract: In environmental monitoring, data collection is often costly, sparse, and shaped by urgent public-health needs. This is particularly true for

Adaptive Multi-prompt Contrastive Network for Few-shot Out-of-distribution Detection

ApplicationsDGX agent

arXiv:2506.17633v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection attempts to distinguish outlier samples to prevent models trained on the in-distribution (ID) dataset from

Advancing Creative Physical Intelligence in Large Multimodal Models

Model ReleasesDGX agent

arXiv:2605.26396v1 Announce Type: new Abstract: Large multimodal models (LMMs) have rapidly advanced in perception and reasoning; however, it remains unclear whether these capabilities generalize to d

Adversarial Training for Robust Coverage Network under Worst-case Facility Losses

AgentsDGX agent

arXiv:2605.26763v1 Announce Type: cross Abstract: The Maximal Covering Location-Interdiction Problem (MCLIP) is a classic bi-level optimization problem, which is fundamental to resilient infrastructur

AgentSociety: Incentivizing Agentic Social Intelligence

Model ReleasesDGX agent

arXiv:2605.26203v1 Announce Type: cross Abstract: The success of deployed agents relies on their ability to handle open-ended user requests using their inherent capabilities, not only in solving reque

AGORA: Adapter-Grounded Observation-Action Retention for Inference-Free Prompt Compression in LLM Agents

Model ReleasesDGX agent

arXiv:2605.26596v1 Announce Type: new Abstract: The token-level extractive compressors widely used for general LM context are structurally inappropriate for LLM agents: across 17 (env, backbone, metho

AI Agent for Reverse-Engineering Legacy Finite-Difference Code and Translating to Devito

AgentsDGX agent

arXiv:2601.18381v2 Announce Type: replace Abstract: To facilitate the transformation of legacy finite difference implementations into the Devito environment, this study develops an integrated AI agent

AI-Driven Contribution Evaluation and Conflict Resolution: A Framework & Design for Group Workload Investigation

SafetyDGX agent

arXiv:2511.07667v2 Announce Type: replace Abstract: The equitable assessment of individual contribution in teams remains a persistent challenge, where conflict and disparity in workload can result in

AI evaluation may bias perceptions: The importance of context in interpreting academic writing

Model ReleasesDGX agent

arXiv:2605.26662v1 Announce Type: cross Abstract: This paper examines how estimates of AI use in scientific writing can be biased when evaluation methods ignore contextual differences across countries

Algorithmic Monocultures in Hiring

ResearchDGX agent

arXiv:2605.27371v1 Announce Type: cross Abstract: Many employers screen job applicants with algorithms built by the same few algorithm vendors. We hypothesize that algorithmic monoculture leads to the

Aligning Few-Step Generative Models by Amortizing Sample-based Variational Inference

SafetyDGX agent

arXiv:2605.26552v1 Announce Type: cross Abstract: Aligning a few-step generative model is challenging, since existing alignment frameworks typically rely on restrictive assumptions: a tractable likeli

Alignment Makes Language Models Normative, Not Descriptive

SafetyDGX agent

arXiv:2603.17218v2 Announce Type: replace-cross Abstract: Post-training alignment optimizes language models to match human preference signals, but this objective is not equivalent to modeling observed

Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases

SafetyDGX agent

arXiv:2605.27355v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we

Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelines

SafetyDGX agent

arXiv:2605.26442v1 Announce Type: cross Abstract: Much of the alignment tuning literature is organized around optimization objectives, while the construction of alignment data is often treated implici

An End-to-End Learning Approach for Solving Capacitated Location-Routing Problems

Model ReleasesDGX agent

arXiv:2511.02525v2 Announce Type: replace-cross Abstract: The capacitated location-routing problems (CLRPs) are classical problems in combinatorial optimization, which require simultaneously making lo

An In-Vitro Study on Cross-Lingual Generalization in Language Models

ResearchDGX agent

arXiv:2605.26683v1 Announce Type: cross Abstract: Cross-lingual transfer in language models is difficult to study in natural corpora because lexical overlap, morphology, data imbalance, and tokenizati

An investigation of AI integration in sound designer workflows and experiences

TutorialsDGX agent

arXiv:2605.27174v1 Announce Type: cross Abstract: Artificial intelligence is increasingly being integrated into professional audio production workflows, yet a gap persists between the tools developers

Anchor: Mitigating Artifact Drift in Agent Benchmark Generation

Model ReleasesDGX agent

arXiv:2605.26321v1 Announce Type: new Abstract: AI agents are beginning to complete valuable, long-horizon business operations tasks, but training and evaluation environments for enterprise work still

AnchorDiff: Training-Free Concept Grounding for MM-DiTs via Anchor-Based Graph Propagation

ResearchDGX agent

arXiv:2605.26460v1 Announce Type: cross Abstract: Multi-Modal Diffusion Transformers (MM-DiTs) encode rich representations for training-free concept grounding, but existing attention-based methods oft

Annotator Positionality as Signal: Psychometric Weighting for Anti-Autistic Ableism Detection

SafetyDGX agent

arXiv:2605.26397v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in decision-making tasks where they can amplify or suppress perspectives, raising concerns in high-

Aperiodic and Low-Frequency Spectral Bias in Reconstruction based EEG Foundation Models

SafetyDGX agent

arXiv:2605.26434v1 Announce Type: cross Abstract: EEG foundation models, pre-trained on large-scale unlabelled EEG data, have emerged as a promising direction towards learning generalizable EEG repres

APEX-Searcher: Refining Credit Assignment with Subgoaling for Agentic Retrieval-Augmented Generation

AgentsDGX agent

arXiv:2603.13853v3 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) connects large language models (LLMs) to external knowledge, but single-round retrieval is often insuffic

Assessing Per-Sample Membership Inference Vulnerability without Retraining

ResearchDGX agent

arXiv:2602.15919v2 Announce Type: replace-cross Abstract: Recent work in the privacy literature shows that sample-targeted membership inference attacks (MIAs) significantly outperform untargeted appro

AssetGen: Deployable 3D Asset Generation at Interactive Speed

HardwareDGX agent

arXiv:2605.26137v1 Announce Type: cross Abstract: While 3D generation is progressing rapidly, recent work has often focused on obtaining high-resolution assets, leaving user experience and deployabili

Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models

SafetyDGX agent

arXiv:2506.09532v5 Announce Type: replace-cross Abstract: We present Athena-PRM, a multimodal process reward model (PRM) designed to evaluate the reward score for each step in solving complex reasonin

Auditing and Fixing Economic Validity in Tabular Foundation Models for Discrete Choice

SafetyDGX agent

arXiv:2605.26559v1 Announce Type: cross Abstract: Tabular foundation models achieve strong accuracy on choice prediction tasks, but their predictions often violate the economic logic those tasks requi

Augment Engineering: A Methodology for Multi-Tool AI Orchestration Across Professional Domains

ApplicationsDGX agent

arXiv:2605.26146v1 Announce Type: cross Abstract: Organizations increasingly deploy separate purpose-built AI tools across professional domains, often hiring domain specialists for each, recreating th

AutoDFT: A Closed-Loop Multi-Agent Framework for Autonomous DFT Calculations

Model ReleasesDGX agent

arXiv:2605.26179v1 Announce Type: cross Abstract: Density functional theory (DFT) serves as the basis for computational discovery in materials science and chemistry, yet each calculation demands exten

Automatic Layer Selection for Hallucination Detection

TutorialsDGX agent

arXiv:2605.26366v1 Announce Type: new Abstract: Recent studies on hallucination detection have shown that hallucination-related signals are more strongly encoded in intermediate layers than in the fin

BatteryMFormer: Multi-level Learning for Battery Degradation Trajectory Forecasting

Local AiDGX agent

arXiv:2605.27044v1 Announce Type: new Abstract: Early battery degradation trajectory forecasting (BDTF), which predicts the full-life state-of-health trajectory from early operational data, is critica

← Previous
1…203204205206207…358
Next →