AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,115
  • Agents7,313
  • Applications5,228
  • Concepts5
  • Hardware1,762
  • Industry6,105
  • Local Ai4,756
  • Model Releases22,759
  • Research19,333
  • Safety12,889
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,115
  • Agents7,313
  • Applications5,228
  • Concepts5
  • Hardware1,762
  • Industry6,105
  • Local Ai4,756
  • Model Releases22,759
  • Research19,333
  • Safety12,889
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
85,115Total entries
1Added by human
85,114Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,688 results
Model Releases

BlueFin: Benchmarking LLM Agents on Financial Spreadsheets

DGX agent

arXiv:2605.30907v1 Announce Type: cross Abstract: We present BlueFin, a benchmark that tasks large language model (LLM) agents with synthesis, manipulation, and comprehension tasks over spreadsheet wo

model-releasesarxiv-cs-ai
1 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies

DGX agent

arXiv:2512.19673v3 Announce Type: replace-cross Abstract: Existing reinforcement learning (RL) approaches treat large language models (LLMs) as a unified policy, overlooking their internal mechanisms.

model-releasesarxiv-cs-ai
1 Jun 2026
Safety

Boundary-Guided Policy Optimization for Memory-efficient RL of Diffusion Large Language Models

DGX agent

arXiv:2510.11683v3 Announce Type: replace-cross Abstract: A key challenge in applying reinforcement learning (RL) to diffusion large language models (dLLMs) is the intractability of their likelihood f

safetyarxiv-cs-ai
1 Jun 2026
Safety

Breaking Information Cocoons: A Hyperbolic Framework for Balancing Exploration and Exploitation in Recommender Systems

DGX agent

arXiv:2411.13865v4 Announce Type: replace-cross Abstract: Modern recommender systems often create information cocoons, restricting users' exposure to diverse content. The central challenge is to balan

safetyarxiv-cs-ai
1 Jun 2026
Model Releases

Breaking the Simplification Bottleneck in Amortized Neural Symbolic Regression

DGX agent

arXiv:2602.08885v5 Announce Type: replace-cross Abstract: Symbolic regression (SR) aims to discover interpretable analytical expressions that accurately describe observed data. Amortized SR promises t

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Calibrated Preference Learning: The Case of Label Ranking

DGX agent

arXiv:2605.30447v1 Announce Type: cross Abstract: Calibration, the alignment of predicted probabilities with true outcome frequencies, is essential for reliable decision-making. While extensively stud

model-releasesarxiv-cs-ai
1 Jun 2026
Local Ai

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects

DGX agent

arXiv:2510.14904v3 Announce Type: replace-cross Abstract: Dense Video Object Captioning (DVOC) is the task of jointly detecting, tracking, and captioning object trajectories in a video, requiring the

local-aiarxiv-cs-ai
1 Jun 2026
Research

Certified Circuits: Stability Guarantees for Mechanistic Circuits

DGX agent

arXiv:2602.22968v3 Announce Type: replace Abstract: Understanding how neural networks arrive at their predictions is essential for debugging, auditing, and deployment. Mechanistic interpretability pur

researcharxiv-cs-ai
1 Jun 2026
Model Releases

Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

DGX agent

arXiv:2503.08679v5 Announce Type: replace Abstract: Recent studies indicate that when faced with explicit biases in prompts, models often omit mentioning these biases in their Chain-of-Thought (CoT) o

model-releasesarxiv-cs-ai
1 Jun 2026
Research

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS

DGX agent

arXiv:2605.30748v1 Announce Type: cross Abstract: We present Chatterbox-Flash, a zero-shot text-to-speech model obtained by fine-tuning a pretrained autoregressive TTS decoder into a block-diffusion d

researcharxiv-cs-ai
1 Jun 2026
Agents

Choosing the Lens: Strategic Perspective Activation in Context-Dependent Argumentation

DGX agent

arXiv:2605.31581v1 Announce Type: new Abstract: The same arguments often need to be evaluated under different external regimes. An agent with influence over the regime has a strategic lever that stand

agentsarxiv-cs-ai
1 Jun 2026
Research

Circuit-Inspired High-Order Neural Networks with Unified Neural Dynamics Modeling for PDE Solving and Visual Perception

DGX agent

arXiv:2603.23977v2 Announce Type: replace-cross Abstract: Deep networks often rely on architectural heuristics to shape representation evolution, limiting their ability to model data governed by intri

researcharxiv-cs-ai
1 Jun 2026
Local Ai

CobSeg: Coherence Boundary Modeling for Dialogue Topic Segmentation

DGX agent

arXiv:2605.30668v1 Announce Type: cross Abstract: Dialogue topic segmentation is critical in many human-AI collaborative applications which requires identifying heterogeneous boundary cues, including

local-aiarxiv-cs-ai
1 Jun 2026
Model Releases

CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models

DGX agent

arXiv:2605.30394v1 Announce Type: cross Abstract: This paper introduces Code Bench, a benchmark capable of evaluating Large Language Models (LLMs) concise code generation abilities in 60 programming l

model-releasesarxiv-cs-ai
1 Jun 2026
Safety

COFT: Counterfactual-Conformal Decoding for Fair Chain-of-Thought Reasoning in Large Language Models

DGX agent

arXiv:2605.30641v1 Announce Type: cross Abstract: Large language models (LLMs) can reveal and amplify societal biases during chain-of-thought (CoT) generation. We present COFT (Chain of Fair Thought),

safetyarxiv-cs-ai
1 Jun 2026
Agents

COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation

DGX agent

arXiv:2605.31264v1 Announce Type: new Abstract: LLM agents are increasingly expected not only to complete isolated tasks, but also to carry bounded representations of human expertise, judgment, and in

agentsarxiv-cs-ai
1 Jun 2026
Agents

Comparing LLM-Based Conversational and Graphical Interfaces for Industrial Decision Tasks: An Exploratory Mixed-Methods Study

DGX agent

arXiv:2605.31224v1 Announce Type: cross Abstract: The use of Generative AI Conversational User Interfaces (CUI) as a new way to access and analyze data is growing in all sectors, and the industrial on

agentsarxiv-cs-ai
1 Jun 2026
Safety

COMPASS: Cognitive MCTS-Guided Process Alignment for Safe Search Agents

DGX agent

arXiv:2605.30838v1 Announce Type: new Abstract: LLM-powered search agents enable multi-step reasoning and tool use. However, these capabilities introduce retrieval-induced safety degradation, as harmf

safetyarxiv-cs-ai
1 Jun 2026
Model Releases

Conditional Coverage Diagnostics for Conformal Prediction

DGX agent

arXiv:2512.11779v2 Announce Type: replace-cross Abstract: Evaluating conditional coverage remains one of the most persistent challenges in assessing the reliability of predictive systems. Although con

model-releasesarxiv-cs-ai
1 Jun 2026
Safety

ConSensus: Multi-Agent Collaboration for Multimodal Sensing

DGX agent

arXiv:2601.06453v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly grounded in sensor data to perceive and reason about human physiology and the physical world. However,

safetyarxiv-cs-ai
1 Jun 2026
Model Releases

ConTrans: Learning Text-enhanced Local-global Temporal Representations for Zero-shot Temporal Action Localization

DGX agent

arXiv:2605.30689v1 Announce Type: cross Abstract: Zero-shot Temporal Action Localization (ZS-TAL) aims to detect and locate previously unseen actions in untrimmed videos. However, existing approaches

model-releasesarxiv-cs-ai
1 Jun 2026
Research

Controllable Lung Nodule Synthesis via Histogram-Regularized Latent Diffusion Models

DGX agent

arXiv:2605.30631v1 Announce Type: cross Abstract: While automated diagnosis systems have achieved remarkable success in computed tomography (CT)-based lung cancer screening, their development remains

researcharxiv-cs-ai
1 Jun 2026
Research

Correcting Split Selection in Online Decision Trees via Anytime-Valid Inference

DGX agent

arXiv:2605.31239v1 Announce Type: cross Abstract: Bagging-based ensembles, most notably Adaptive Random Forests, are among the strongest performers for learning from data streams. A common denominator

researcharxiv-cs-ai
1 Jun 2026
Safety

Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents

DGX agent

arXiv:2605.30590v1 Announce Type: cross Abstract: Two clinical AI systems can score nearly identically on coverage-based rubrics yet behave radically differently when their patient inputs change: one

safetyarxiv-cs-ai
1 Jun 2026
Model Releases

Counterfactual Trace Auditing of LLM Agent Skills

DGX agent

arXiv:2605.11946v2 Announce Type: replace Abstract: Large Language Model agents are increasingly augmented with agent skills. Current evaluation methods for skills remain limited. Most deployed benchm

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs

DGX agent

arXiv:2605.30611v1 Announce Type: cross Abstract: Scientific figures are among the most effective means of communicating complex research ideas, yet producing publication-quality illustrations remains

model-releasesarxiv-cs-ai
1 Jun 2026
Safety

Cross-Modal Attention Calibration for LVLM Hallucination Mitigation

DGX agent

arXiv:2501.01926v3 Announce Type: replace-cross Abstract: Large vision-language models (LVLMs) have shown remarkable capabilities in visual-language understanding. Despite their success, LVLMs still s

safetyarxiv-cs-ai
1 Jun 2026
Applications

D^3: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training

DGX agent

arXiv:2605.31164v1 Announce Type: cross Abstract: Training data plays a central role in large language models (LLMs) optimization, motivating extensive research on data scheduling strategies. Most exi

applicationsarxiv-cs-ai
1 Jun 2026
Research

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning

DGX agent

arXiv:2605.30859v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has become pivotal for improving model capabilities yet suffers from rollout efficiency bottlenecks due to the long-tail r

researcharxiv-cs-ai
1 Jun 2026
Safety

dashi: A Python library for Dataset Shift Characterization to Support Trustworthy AI Development and Deployment

DGX agent

arXiv:2605.31360v1 Announce Type: cross Abstract: The Artificial Intelligence (AI) life cycle requires a thorough understanding of the underlying data dynamics for robust, safe and cost-effective AI d

safetyarxiv-cs-ai
1 Jun 2026
Research

De-attribute to Forget for LLM Unlearning

DGX agent

arXiv:2605.30919v1 Announce Type: cross Abstract: The rapid development of large language models (LLMs) has raised concerns on the use of inappropriate data for training, which has led to a growing in

researcharxiv-cs-ai
1 Jun 2026
Research

DEM: A Distilled Explanation Model for Interpretable Anomaly Detection in Physiological Sensor Networks

DGX agent

arXiv:2605.31007v1 Announce Type: cross Abstract: Anomaly detection in physiological sensor data from Wireless Body Area Networks (WBANs) can be caused by sensor faults, network disruptions, or missin

researcharxiv-cs-ai
1 Jun 2026
Model Releases

DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation

DGX agent

arXiv:2605.31286v1 Announce Type: cross Abstract: Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable manipulation skills across diverse object

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, Payload Framing, and Turn-Budget Sensitivity

DGX agent

arXiv:2605.30686v1 Announce Type: cross Abstract: ReAct agents that interleave chain-of-thought reasoning with tool calls are increasingly deployed for real tasks such as scheduling, file retrieval, a

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Design and Evaluation of Multi-Agent AI Oracle Systems for Prediction Market Resolution

DGX agent

arXiv:2605.30802v1 Announce Type: cross Abstract: Prediction markets aggregate collective intelligence to forecast uncertain events, but their utility depends on reliable outcome resolution. Existing

model-releasesarxiv-cs-ai
1 Jun 2026
Tutorials

Developing a Culturally Grounded, AI-Augmented UX Research Point of View (POV): An Exemplar Case Study from Telemedicine Dementia Care

DGX agent

arXiv:2605.31147v1 Announce Type: cross Abstract: User Experience Research (UXR) Points of View (POVs) distil complex and often fragmented research evidence into actionable perspectives that guide how

tutorialsarxiv-cs-ai
1 Jun 2026
Research

Developing a UXR Point of View for Cognitive Accessibility in Mobile Learning with Generative AI

DGX agent

arXiv:2605.31149v1 Announce Type: cross Abstract: This study investigates how UX research (UXR) principles, combined with Large Language Model (LLM)-supported analysis, can be used to improve the qual

researcharxiv-cs-ai
1 Jun 2026
Tutorials

Developing an AI-Powered UX Research Point of View for Digital Health in A Regulatory Context: An Exemplar Case from MSM and Transgender HIV Care in Nigeria

DGX agent

arXiv:2605.31138v1 Announce Type: cross Abstract: User Experience Research (UXR) in a legal and regulatory contexts presents unique challenges that require specialised approaches to protect vulnerable

tutorialsarxiv-cs-ai
1 Jun 2026
Safety

Diagnosing Failure Modes of Shared-State Collaboration in Resource-Constrained Visual Agents

DGX agent

arXiv:2605.31354v1 Announce Type: new Abstract: Modular visual reasoning systems increasingly rely on shared working memory for multi-step collaboration, yet the failure dynamics of intermediate state

safetyarxiv-cs-ai
1 Jun 2026
Safety

Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory

DGX agent

arXiv:2602.00521v2 Announce Type: replace Abstract: While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offer

safetyarxiv-cs-ai
1 Jun 2026
Safety

Differentially Private Preference Data Synthesis for Large Language Model Alignment

DGX agent

arXiv:2605.30808v1 Announce Type: cross Abstract: Preference alignment is a crucial post-training step for large language models (LLMs) to ensure their outputs align with human values. However, post-t

safetyarxiv-cs-ai
1 Jun 2026
Safety

DISCO: Mitigating Bias in Deep Learning with Conditional Distance Correlation

DGX agent

arXiv:2506.11653v3 Announce Type: replace-cross Abstract: Dataset bias often leads deep learning models to exploit spurious correlations instead of task-relevant signals. We introduce the Standard Ant

safetyarxiv-cs-ai
1 Jun 2026
Research

Discovering Differences in Strategic Behavior Between Humans and LLMs

DGX agent

arXiv:2602.10324v2 Announce Type: replace Abstract: As Large Language Models (LLMs) are increasingly deployed in social and strategic scenarios, it becomes critical to understand where and why their b

researcharxiv-cs-ai
1 Jun 2026
Safety

Distilling LLM Feedback for Lean Theorem Proving

DGX agent

arXiv:2605.30861v1 Announce Type: new Abstract: Post-training for reasoning models typically combines supervised fine-tuning with reinforcement learning from verifiable rewards, most commonly with GRP

safetyarxiv-cs-ai
1 Jun 2026
Applications

Do Large Language Models Encode Institutional Experience? Evidence from Cross-Linguistic Moral Reasoning Under Ambiguity

DGX agent

arXiv:2605.30934v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit systematic differences in moral reasoning across languages, yet the source of this variation remains unclear. We

applicationsarxiv-cs-ai
1 Jun 2026
Safety

DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs

DGX agent

arXiv:2605.31432v1 Announce Type: cross Abstract: Simultaneous speech-to-text translation (SimulST) generates translations while speech is still unfolding, requiring a streaming policy that decides wh

safetyarxiv-cs-ai
1 Jun 2026
Safety

Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?

DGX agent

arXiv:2605.31041v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated promising capability in autonomous driving, highlighting the potential of unified multimodal arc

safetyarxiv-cs-ai
1 Jun 2026
Research

Domain Adaptation and Reasoning Frameworks in Language Models: A Controlled Experiment with Historical Cosmology

DGX agent

arXiv:2605.30415v1 Announce Type: cross Abstract: We investigate how domain adaptation reshapes explanatory behavior in language models using historical cosmology as a controlled setting. In Phase 1,

researcharxiv-cs-ai
1 Jun 2026
← Previous
1…233234235236237…452
Next →