AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlog
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
Safety

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning

DGX agent

arXiv:2605.28023v1 Announce Type: cross Abstract: Visual captioning requires models to capture visual content faithfully while minimizing both omission and hallucination. As the dominant paradigm for

safetyarxiv-cs-ai
28 May 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Verifiable Benchmarking of Long-Horizon Spatial Biology

DGX agent

arXiv:2605.28065v1 Announce Type: new Abstract: AI agents are increasingly useful for biological data analysis, but existing benchmarks mostly test broad biological knowledge, executable workflows, or

model-releasesarxiv-cs-ai
28 May 2026
Safety

Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs

DGX agent

arXiv:2605.28565v1 Announce Type: cross Abstract: Users of search-augmented LLMs rely on citations as evidence that responses are grounded in real sources, and rarely verify the cited pages themselves

safetyarxiv-cs-ai
28 May 2026
Model Releases

VeriTrip: A Verifiable Benchmark for Travel Planning Agents over Unstructured Web Corpora

DGX agent

arXiv:2605.28683v1 Announce Type: new Abstract: Existing benchmarks have laid the foundation for travel planning agents by establishing API-centric paradigms. However, as the capabilities of Autonomou

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild

DGX agent

arXiv:2605.27882v1 Announce Type: cross Abstract: LLM-based agents score well on search benchmarks, yet real users consistently find results unsatisfying, revealing a persistent evaluation-experience

model-releasesarxiv-cs-ai
28 May 2026
Research

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer

DGX agent

arXiv:2605.28229v1 Announce Type: cross Abstract: With the rapid development of pre-training technologies, adapting large-scale Vision-Language Models (VLMs) for video understanding ie image-to-video

researcharxiv-cs-ai
28 May 2026
Safety

Visualizing Latent Phase Structures in Locomotion Policies: A Multi-Environment Study with Temporal Feature Extension

DGX agent

arXiv:2605.28186v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) has been shown to achieve high performance on locomotion control tasks in MuJoCo benchmarks such as HalfCheetah, Ant

safetyarxiv-cs-ai
28 May 2026
Model Releases

VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs

DGX agent

arXiv:2605.28422v1 Announce Type: cross Abstract: Latent reasoning enables reasoning over continuous hidden states rather than explicit tokens, avoiding the language bottleneck and inference overhead

model-releasesarxiv-cs-ai
28 May 2026
Safety

Voluntary Collusion with Secret Tools in Competing LLM Agents

DGX agent

arXiv:2605.27593v1 Announce Type: new Abstract: Even when a tool is explicitly described as unfair and harmful to others, ostensibly safety-aligned LLM agents still voluntarily engage in secret collus

safetyarxiv-cs-ai
28 May 2026
Model Releases

VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization

DGX agent

arXiv:2511.11896v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently shown strong potential in vulnerability detection (VD). However, accurately detecting vulnerabiliti

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

DGX agent

arXiv:2605.27851v1 Announce Type: new Abstract: Safety benchmark scores provide incomplete evidence of deployment readiness: aligned language models often adhere to rigid rules even when a situational

model-releasesarxiv-cs-ai
28 May 2026
Agents

When Does Memory Help Multi-Trajectory Inference for Tool-Use LLM Agents?

DGX agent

arXiv:2605.28224v1 Announce Type: new Abstract: Multi-trajectory inference for tool-use LLM agents - generating multiple reasoning attempts and selecting among them - benefits from transferring knowle

agentsarxiv-cs-ai
28 May 2026
Local Ai

When NPUs Are Not Always Faster: A Stage-Level Analysis of Mobile LLM Inference

DGX agent

arXiv:2605.27435v1 Announce Type: cross Abstract: Deploying large language models (LLMs) on mobile devices increasingly relies on heterogeneous execution, yet no prior study has systematically charact

local-aiarxiv-cs-ai
28 May 2026
Research

When prompt perturbations break your A/B test: A valid statistical test for generative surveying

DGX agent

arXiv:2605.27463v1 Announce Type: cross Abstract: Generative surveying -- where collections of LLM-based personas provide feedback on messages -- has emerged as a cheap and scalable alternative to tra

researcharxiv-cs-ai
28 May 2026
Safety

When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?

DGX agent

arXiv:2605.27932v1 Announce Type: cross Abstract: Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly underst

safetyarxiv-cs-ai
28 May 2026
Local Ai

Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models

DGX agent

arXiv:2605.27997v1 Announce Type: cross Abstract: Large language models frequently generate toxic, hateful, or harmful content, yet existing mitigation methods rely on costly retraining or output-leve

local-aiarxiv-cs-ai
28 May 2026
Safety

Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR

DGX agent

arXiv:2605.28295v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) trains reasoning models without labeled trajectories, relying on grouped rollouts to expose the po

safetyarxiv-cs-ai
28 May 2026
Safety

Who Uses AI? Platform Selection and the Measurement of Occupational AI Exposure

DGX agent

arXiv:2605.21743v2 Announce Type: replace Abstract: Conversation logs from AI platforms are increasingly used to measure occupational exposure to artificial intelligence, but the users observed in the

safetyarxiv-cs-ai
28 May 2026
Model Releases

Whose Name Comes Up? III: Persona Prompting Effects in LLM-Based Scholar Recommendation

DGX agent

arXiv:2605.28187v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as scholar recommenders, shaping who is seen as an expert in academia. Existing audits remain Engli

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Why LLMs Fail at Causal Discovery and How Interventional Agents Escape

DGX agent

arXiv:2605.27567v1 Announce Type: new Abstract: Causal discovery is a cornerstone of scientific reasoning, yet whether large language models can perform it reliably remains an open question. Recent be

model-releasesarxiv-cs-ai
28 May 2026
Research

Worker Disagreement Reveals Sharp Directions in Local SGD

DGX agent

arXiv:2605.27739v1 Announce Type: cross Abstract: Deep neural network training often exhibits highly anisotropic loss geometry, where a few sharp dominant Hessian directions coexist with a large flatt

researcharxiv-cs-ai
28 May 2026
Applications

You Are in Control of Your State: Why Human Outcomes Are Controllable Through Causal State Intervention

DGX agent

arXiv:2605.27580v1 Announce Type: new Abstract: A central puzzle for the behavioural sciences and for human-facing artificial intelligence is the persistence of within-person variability. The same ind

applicationsarxiv-cs-ai
28 May 2026
Model Releases

You Live More Than Once: Towards Hierarchical Skill Meta-Evolving

DGX agent

arXiv:2605.28390v1 Announce Type: new Abstract: Test-time skill evolving is regarded as a new paradigm for enhancing deployed agentic systems. Existing works mainly focus on hard-coded skill evolving

model-releasesarxiv-cs-ai
28 May 2026
Research

Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training

DGX agent

arXiv:2605.28008v1 Announce Type: new Abstract: Large language models (LLMs) can now solve complex problems through long chain-of-thought (CoT) reasoning, but the trade-off between performance and tok

researcharxiv-cs-ai
28 May 2026
Model Releases

ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay

DGX agent

arXiv:2605.28069v1 Announce Type: new Abstract: Adaptive context compression is vital for scaling Large Language Models (LLMs) to complex, multi-turn agent tasks. However, rule-based compression metho

model-releasesarxiv-cs-ai
28 May 2026
Research

2-ASP(Q) programs with weak constraints: Complexity and efficient implementation

DGX agent

arXiv:2605.27338v1 Announce Type: new Abstract: ASP(Q) extends Answer Set Programming (ASP) with Quantifiers over answer sets. In this paper we focus on the class of ASP(Q) programs with two quantifie

researcharxiv-cs-ai
27 May 2026
Model Releases

A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks

DGX agent

arXiv:2605.26747v1 Announce Type: new Abstract: Large Language Models (LLMs) have brought huge improvements to Artificial Intelligence (AI), which can be applied to general-purpose tasks. However, the

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection

DGX agent

arXiv:2605.26533v1 Announce Type: cross Abstract: Automated industrial inspection requires both precise defect localization and structured maintenance report generation; in current practice these task

model-releasesarxiv-cs-ai
27 May 2026
Research

A Physics-Informed Hierarchical Neural Network for Microwave Scattering Analysis of 3D PEC Targets

DGX agent

arXiv:2508.03774v5 Announce Type: replace-cross Abstract: Accurate modeling of scattering from three-dimensional (3D) perfectly electrically conducting (PEC) targets at microwave frequencies constitut

researcharxiv-cs-ai
27 May 2026
Safety

A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration

DGX agent

arXiv:2605.26174v1 Announce Type: cross Abstract: Production language-model systems answer a request by partitioning it across an invisible orchestration of worker agents that recompose one integrated

safetyarxiv-cs-ai
27 May 2026
Applications

Adapting Actively on the Fly: Relevance-Guided Online Meta-Learning with Latent Concepts for Geospatial Discovery

DGX agent

arXiv:2602.17605v2 Announce Type: replace-cross Abstract: In environmental monitoring, data collection is often costly, sparse, and shaped by urgent public-health needs. This is particularly true for

applicationsarxiv-cs-ai
27 May 2026
Applications

Adaptive Multi-prompt Contrastive Network for Few-shot Out-of-distribution Detection

DGX agent

arXiv:2506.17633v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection attempts to distinguish outlier samples to prevent models trained on the in-distribution (ID) dataset from

applicationsarxiv-cs-ai
27 May 2026
Model Releases

Advancing Creative Physical Intelligence in Large Multimodal Models

DGX agent

arXiv:2605.26396v1 Announce Type: new Abstract: Large multimodal models (LMMs) have rapidly advanced in perception and reasoning; however, it remains unclear whether these capabilities generalize to d

model-releasesarxiv-cs-ai
27 May 2026
Agents

Adversarial Training for Robust Coverage Network under Worst-case Facility Losses

DGX agent

arXiv:2605.26763v1 Announce Type: cross Abstract: The Maximal Covering Location-Interdiction Problem (MCLIP) is a classic bi-level optimization problem, which is fundamental to resilient infrastructur

agentsarxiv-cs-ai
27 May 2026
Model Releases

AgentSociety: Incentivizing Agentic Social Intelligence

DGX agent

arXiv:2605.26203v1 Announce Type: cross Abstract: The success of deployed agents relies on their ability to handle open-ended user requests using their inherent capabilities, not only in solving reque

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

AGORA: Adapter-Grounded Observation-Action Retention for Inference-Free Prompt Compression in LLM Agents

DGX agent

arXiv:2605.26596v1 Announce Type: new Abstract: The token-level extractive compressors widely used for general LM context are structurally inappropriate for LLM agents: across 17 (env, backbone, metho

model-releasesarxiv-cs-ai
27 May 2026
Agents

AI Agent for Reverse-Engineering Legacy Finite-Difference Code and Translating to Devito

DGX agent

arXiv:2601.18381v2 Announce Type: replace Abstract: To facilitate the transformation of legacy finite difference implementations into the Devito environment, this study develops an integrated AI agent

agentsarxiv-cs-ai
27 May 2026
Safety

AI-Driven Contribution Evaluation and Conflict Resolution: A Framework & Design for Group Workload Investigation

DGX agent

arXiv:2511.07667v2 Announce Type: replace Abstract: The equitable assessment of individual contribution in teams remains a persistent challenge, where conflict and disparity in workload can result in

safetyarxiv-cs-ai
27 May 2026
Model Releases

AI evaluation may bias perceptions: The importance of context in interpreting academic writing

DGX agent

arXiv:2605.26662v1 Announce Type: cross Abstract: This paper examines how estimates of AI use in scientific writing can be biased when evaluation methods ignore contextual differences across countries

model-releasesarxiv-cs-ai
27 May 2026
Research

Algorithmic Monocultures in Hiring

DGX agent

arXiv:2605.27371v1 Announce Type: cross Abstract: Many employers screen job applicants with algorithms built by the same few algorithm vendors. We hypothesize that algorithmic monoculture leads to the

researcharxiv-cs-ai
27 May 2026
Safety

Aligning Few-Step Generative Models by Amortizing Sample-based Variational Inference

DGX agent

arXiv:2605.26552v1 Announce Type: cross Abstract: Aligning a few-step generative model is challenging, since existing alignment frameworks typically rely on restrictive assumptions: a tractable likeli

safetyarxiv-cs-ai
27 May 2026
Safety

Alignment Makes Language Models Normative, Not Descriptive

DGX agent

arXiv:2603.17218v2 Announce Type: replace-cross Abstract: Post-training alignment optimizes language models to match human preference signals, but this objective is not equivalent to modeling observed

safetyarxiv-cs-ai
27 May 2026
Safety

Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases

DGX agent

arXiv:2605.27355v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we

safetyarxiv-cs-ai
27 May 2026
Safety

Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelines

DGX agent

arXiv:2605.26442v1 Announce Type: cross Abstract: Much of the alignment tuning literature is organized around optimization objectives, while the construction of alignment data is often treated implici

safetyarxiv-cs-ai
27 May 2026
Model Releases

An End-to-End Learning Approach for Solving Capacitated Location-Routing Problems

DGX agent

arXiv:2511.02525v2 Announce Type: replace-cross Abstract: The capacitated location-routing problems (CLRPs) are classical problems in combinatorial optimization, which require simultaneously making lo

model-releasesarxiv-cs-ai
27 May 2026
Research

An In-Vitro Study on Cross-Lingual Generalization in Language Models

DGX agent

arXiv:2605.26683v1 Announce Type: cross Abstract: Cross-lingual transfer in language models is difficult to study in natural corpora because lexical overlap, morphology, data imbalance, and tokenizati

researcharxiv-cs-ai
27 May 2026
Tutorials

An investigation of AI integration in sound designer workflows and experiences

DGX agent

arXiv:2605.27174v1 Announce Type: cross Abstract: Artificial intelligence is increasingly being integrated into professional audio production workflows, yet a gap persists between the tools developers

tutorialsarxiv-cs-ai
27 May 2026
Model Releases

Anchor: Mitigating Artifact Drift in Agent Benchmark Generation

DGX agent

arXiv:2605.26321v1 Announce Type: new Abstract: AI agents are beginning to complete valuable, long-horizon business operations tasks, but training and evaluation environments for enterprise work still

model-releasesarxiv-cs-ai
27 May 2026
← Previous
1…254255256257258…448
Next →