AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
1 Jun 2026

Correcting Split Selection in Online Decision Trees via Anytime-Valid Inference

ResearchDGX agent

arXiv:2605.31239v1 Announce Type: cross Abstract: Bagging-based ensembles, most notably Adaptive Random Forests, are among the strongest performers for learning from data streams. A common denominator

Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents

SafetyDGX agent

arXiv:2605.30590v1 Announce Type: cross Abstract: Two clinical AI systems can score nearly identically on coverage-based rubrics yet behave radically differently when their patient inputs change: one

Counterfactual Trace Auditing of LLM Agent Skills

Model ReleasesDGX agent

arXiv:2605.11946v2 Announce Type: replace Abstract: Large Language Model agents are increasingly augmented with agent skills. Current evaluation methods for skills remain limited. Most deployed benchm


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs

Model ReleasesDGX agent

arXiv:2605.30611v1 Announce Type: cross Abstract: Scientific figures are among the most effective means of communicating complex research ideas, yet producing publication-quality illustrations remains

Cross-Modal Attention Calibration for LVLM Hallucination Mitigation

SafetyDGX agent

arXiv:2501.01926v3 Announce Type: replace-cross Abstract: Large vision-language models (LVLMs) have shown remarkable capabilities in visual-language understanding. Despite their success, LVLMs still s

D^3: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training

ApplicationsDGX agent

arXiv:2605.31164v1 Announce Type: cross Abstract: Training data plays a central role in large language models (LLMs) optimization, motivating extensive research on data scheduling strategies. Most exi

DARTS: Distribution-Aware Active Rollout Trajectory Shaping for Accelerating LLM Reinforcement Learning

ResearchDGX agent

arXiv:2605.30859v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has become pivotal for improving model capabilities yet suffers from rollout efficiency bottlenecks due to the long-tail r

dashi: A Python library for Dataset Shift Characterization to Support Trustworthy AI Development and Deployment

SafetyDGX agent

arXiv:2605.31360v1 Announce Type: cross Abstract: The Artificial Intelligence (AI) life cycle requires a thorough understanding of the underlying data dynamics for robust, safe and cost-effective AI d

De-attribute to Forget for LLM Unlearning

ResearchDGX agent

arXiv:2605.30919v1 Announce Type: cross Abstract: The rapid development of large language models (LLMs) has raised concerns on the use of inappropriate data for training, which has led to a growing in

DEM: A Distilled Explanation Model for Interpretable Anomaly Detection in Physiological Sensor Networks

ResearchDGX agent

arXiv:2605.31007v1 Announce Type: cross Abstract: Anomaly detection in physiological sensor data from Wireless Body Area Networks (WBANs) can be caused by sensor faults, network disruptions, or missin

DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation

Model ReleasesDGX agent

arXiv:2605.31286v1 Announce Type: cross Abstract: Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable manipulation skills across diverse object

Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, Payload Framing, and Turn-Budget Sensitivity

Model ReleasesDGX agent

arXiv:2605.30686v1 Announce Type: cross Abstract: ReAct agents that interleave chain-of-thought reasoning with tool calls are increasingly deployed for real tasks such as scheduling, file retrieval, a

Design and Evaluation of Multi-Agent AI Oracle Systems for Prediction Market Resolution

Model ReleasesDGX agent

arXiv:2605.30802v1 Announce Type: cross Abstract: Prediction markets aggregate collective intelligence to forecast uncertain events, but their utility depends on reliable outcome resolution. Existing

Developing a Culturally Grounded, AI-Augmented UX Research Point of View (POV): An Exemplar Case Study from Telemedicine Dementia Care

TutorialsDGX agent

arXiv:2605.31147v1 Announce Type: cross Abstract: User Experience Research (UXR) Points of View (POVs) distil complex and often fragmented research evidence into actionable perspectives that guide how

Developing a UXR Point of View for Cognitive Accessibility in Mobile Learning with Generative AI

ResearchDGX agent

arXiv:2605.31149v1 Announce Type: cross Abstract: This study investigates how UX research (UXR) principles, combined with Large Language Model (LLM)-supported analysis, can be used to improve the qual

Developing an AI-Powered UX Research Point of View for Digital Health in A Regulatory Context: An Exemplar Case from MSM and Transgender HIV Care in Nigeria

TutorialsDGX agent

arXiv:2605.31138v1 Announce Type: cross Abstract: User Experience Research (UXR) in a legal and regulatory contexts presents unique challenges that require specialised approaches to protect vulnerable

Diagnosing Failure Modes of Shared-State Collaboration in Resource-Constrained Visual Agents

SafetyDGX agent

arXiv:2605.31354v1 Announce Type: new Abstract: Modular visual reasoning systems increasingly rely on shared working memory for multi-step collaboration, yet the failure dynamics of intermediate state

Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory

SafetyDGX agent

arXiv:2602.00521v2 Announce Type: replace Abstract: While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offer

Differentially Private Preference Data Synthesis for Large Language Model Alignment

SafetyDGX agent

arXiv:2605.30808v1 Announce Type: cross Abstract: Preference alignment is a crucial post-training step for large language models (LLMs) to ensure their outputs align with human values. However, post-t

DISCO: Mitigating Bias in Deep Learning with Conditional Distance Correlation

SafetyDGX agent

arXiv:2506.11653v3 Announce Type: replace-cross Abstract: Dataset bias often leads deep learning models to exploit spurious correlations instead of task-relevant signals. We introduce the Standard Ant

Discovering Differences in Strategic Behavior Between Humans and LLMs

ResearchDGX agent

arXiv:2602.10324v2 Announce Type: replace Abstract: As Large Language Models (LLMs) are increasingly deployed in social and strategic scenarios, it becomes critical to understand where and why their b

Distilling LLM Feedback for Lean Theorem Proving

SafetyDGX agent

arXiv:2605.30861v1 Announce Type: new Abstract: Post-training for reasoning models typically combines supervised fine-tuning with reinforcement learning from verifiable rewards, most commonly with GRP

Do Large Language Models Encode Institutional Experience? Evidence from Cross-Linguistic Moral Reasoning Under Ambiguity

ApplicationsDGX agent

arXiv:2605.30934v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit systematic differences in moral reasoning across languages, yet the source of this variation remains unclear. We

DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs

SafetyDGX agent

arXiv:2605.31432v1 Announce Type: cross Abstract: Simultaneous speech-to-text translation (SimulST) generates translations while speech is still unfolding, requiring a streaming policy that decides wh

Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?

SafetyDGX agent

arXiv:2605.31041v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated promising capability in autonomous driving, highlighting the potential of unified multimodal arc

Domain Adaptation and Reasoning Frameworks in Language Models: A Controlled Experiment with Historical Cosmology

ResearchDGX agent

arXiv:2605.30415v1 Announce Type: cross Abstract: We investigate how domain adaptation reshapes explanatory behavior in language models using historical cosmology as a controlled setting. In Phase 1,

Dreaming Of Others: Latent Teammate Modeling In World Models For Multi-Agent Reinforcement Learning

AgentsDGX agent

arXiv:2605.31361v1 Announce Type: cross Abstract: In cooperative multi-agent reinforcement learning (MARL), agents must coordinate with partners whose internal policies and intentions are not directly

DRIFT: Joint Channel Estimation and Prediction Towards Pilotless 6G Non-Terrestrial Networks

ResearchDGX agent

arXiv:2605.31065v1 Announce Type: cross Abstract: Non-terrestrial networks (NTNs) are expected to play a pivotal role in sixth-generation (6G) systems by enabling ubiquitous connectivity and massive c

DTBench: A Synthetic Benchmark for Document-to-Table Extraction

Model ReleasesDGX agent

arXiv:2602.13812v3 Announce Type: replace-cross Abstract: Document-to-table (Doc2Table) extraction derives structured tables from unstructured documents under a target schema, enabling reliable and ve

DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training

ResearchDGX agent

arXiv:2512.13996v2 Announce Type: replace Abstract: Sparse Mixture-of-Experts architectures are essential for scaling model capacity efficiently, yet the standard Top-k routing imposes a rigid sparsit

Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models

SafetyDGX agent

arXiv:2509.24319v4 Announce Type: replace-cross Abstract: Large language models can express values in two main ways: (1) intrinsic expression, reflecting the model's inherent values learned during tra

DynaTree: Dynamic Agentic Retrieval Tree for Time-Sensitive News Retrieval

Model ReleasesDGX agent

arXiv:2605.31377v1 Announce Type: cross Abstract: Agentic Retrieval-Augmented Generation improves retrieval by integrating planning, tool use, and iterative reasoning, but existing agentic RAG methods

EchoRL: Reinforcement Learning via Rollout Echoing

SafetyDGX agent

arXiv:2605.31228v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards is an effective route for post-training to strengthen the reasoning capability of large language models

Effective Reasoning Chains Reduce Intrinsic Dimensionality

Model ReleasesDGX agent

arXiv:2602.09276v2 Announce Type: replace-cross Abstract: Chain-of-thought (CoT) reasoning and its variants have substantially improved the performance of language models on complex reasoning tasks, y

EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs

Model ReleasesDGX agent

arXiv:2605.30637v1 Announce Type: new Abstract: Clinical decision-making (CDM) is central to real-world clinical workflows, where clinicians infer diagnoses, select treatments, or anticipate future he

EMCEE: Improving Multilingual Capability of LLMs via Bridging Knowledge and Reasoning with Extracted Synthetic Multilingual Context

ResearchDGX agent

arXiv:2503.05846v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have achieved impressive progress across a wide range of tasks, yet their heavy reliance on English-centric train

Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion

Model ReleasesDGX agent

arXiv:2605.31170v1 Announce Type: cross Abstract: Monitoring autonomous language model agents currently relies mostly on surface behavior. But what happens when agent populations invent new languages

Empirical Characterization of Inference-Time Elicited Probability Transformations in Large Language Models

ResearchDGX agent

arXiv:2603.19262v2 Announce Type: replace-cross Abstract: Large language models increasingly rely on inference-time procedures such as chain-of-thought reasoning, self-refinement, retrieval augmentati

Enhancing Regime Shift Detection Using Unstructured Data: A Study on the Treasury Market

SafetyDGX agent

arXiv:2605.30363v1 Announce Type: cross Abstract: Regime shifts in financial markets reorganise the joint dynamics of asset prices and macro variables, breaking any single-regime calibration. They are

Entropic Projection Alignment: Estimating, Explaining, and Improving Model Performance Under Distribution Shift

SafetyDGX agent

arXiv:2605.31250v1 Announce Type: cross Abstract: We propose a unified framework for addressing three key challenges of distribution shift: (1) estimating a model's performance on an unlabeled target

Envisioning Beyond the Few: Disentangled Semantics and Primitives for Few-Shot Atypical Layout-to-Image Generation

SafetyDGX agent

arXiv:2605.31266v1 Announce Type: cross Abstract: The layout-to-image (L2I) task enables fine-grained control over image generation via object categories and spatial layouts. However, existing L2I met

ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2605.31251v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have shown strong potential as embodied agents, yet embodied geo-localization remains underexplored due to th

EUDAIMONIA: Evaluating Undesirable Dynamics in AI

Model ReleasesDGX agent

arXiv:2605.30654v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as conversational partners for companionship, emotional disclosure, and interpersonal advice, but t

Evolutionary Algorithm for Reservoir Learning and Yielding

ResearchDGX agent

arXiv:2605.30372v1 Announce Type: cross Abstract: Reservoir computing, a type of recurrent neural network, is a promising approach for temporal learning as it separates dynamic processing from the tra

Exploring Autonomous Agentic Data Engineering for Model Specialization

Model ReleasesDGX agent

arXiv:2605.30407v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without hig

Extending the UXR Point of View Pyramid: A Generative AI-Augmented Methodology for Human-Centred AI Systems

SafetyDGX agent

arXiv:2605.31143v1 Announce Type: cross Abstract: Rising household debt and cost-of-living pressures in the United Kingdom have intensified the role of AI-driven financial technologies in mediating cr

FAM-Bench: A Multimodal Benchmark for Condition-Aware Food-as-Medicine Reasoning

Model ReleasesDGX agent

arXiv:2605.31410v1 Announce Type: new Abstract: Food-as-Medicine requires models to reason beyond what a dish is or what nutrition it contains: they must decide whether a concrete food choice is appro

FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection

Model ReleasesDGX agent

arXiv:2605.31349v1 Announce Type: cross Abstract: Hateful meme detection remains a formidable challenge for vision-language models, as existing benchmarks are structurally observational - confounding

Feature-Optimized Vision for Adaptive 3D Scene Reconstruction

SafetyDGX agent

arXiv:2605.31534v1 Announce Type: cross Abstract: Three-dimensional scene reconstruction depends on local image evidence that is both visually discriminative and geometrically useful. Fixed feature th

Federated Variational Preference Alignment with Gumbel-Softmax Prior for Personalized User Preferences

Local AiDGX agent

arXiv:2605.30873v1 Announce Type: cross Abstract: Federated Learning (FL) offers a privacy-preserving pathway for aligning Large Language Models (LLMs); however, existing frameworks typically enforce

FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs

Model ReleasesDGX agent

arXiv:2512.20732v2 Announce Type: replace-cross Abstract: As LLMs advance their reasoning capabilities about the physical world, the absence of rigorous benchmarks for evaluating their ability to gene

Fighting Numerical Hallucinations via Data-centric Compilation for Online Financial QA

AgentsDGX agent

arXiv:2605.31064v1 Announce Type: cross Abstract: Large Language Models (LLMs) have significantly advanced online data services, particularly in the domain of financial question answering (FinQA). How

Fine-grained Verification via Diagnostic Reasoning Supervision for Aspect Sentiment Triplet Extraction

Local AiDGX agent

arXiv:2605.31446v1 Announce Type: cross Abstract: Aspect Sentiment Triplet Extraction (ASTE) aims to identify aspect terms, opinion terms, and sentiment polarities as structured triplets, providing es

Fine-Tuning Improves Information Conveyance in Language Models

ResearchDGX agent

arXiv:2605.30844v1 Announce Type: cross Abstract: Fine-tuning is often believed to reduce uncertainty and diversity in large language models, but existing analyses overlook output length, a key confou

Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments

ResearchDGX agent

arXiv:2601.01075v2 Announce Type: replace-cross Abstract: Embodied systems experience the world as 'a symphony of flows': a combination of many continuous streams of sensory input coupled to self-moti

FOCUS: Forcing In-Context Object Localization through Visual Support Constraints and Policy Optimization

Model ReleasesDGX agent

arXiv:2605.31145v1 Announce Type: cross Abstract: In-context localization (ICL) seeks to localize a target object specified by a small set of support examples in a query image, operating on the fly wi

Formalizing and falsifying causal pathways of rare events

ResearchDGX agent

arXiv:2605.31254v1 Announce Type: new Abstract: Building on recent formalizations of root cause analysis for rare events (``outliers'') in structural equation models, we propose a formal definition of

From Evidence to Design: Developing an AI-Augmented UX Research Point of View for Digital Wellbeing in Emergency and Public Safety Contexts

SafetyDGX agent

arXiv:2605.31146v1 Announce Type: cross Abstract: This paper investigates how User Experience Research (UXR) methods can be combined with AI-supported analysis to develop clearer design direction for

From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves

ResearchDGX agent

arXiv:2602.24210v2 Announce Type: replace-cross Abstract: Large reasoning models (LRMs) produce reasoning traces (RTs) that often contain sensitive information. These leaky thoughts are difficult to c

From Out-of-Distribution Detection to Hallucination Detection: A Geometric View

SafetyDGX agent

arXiv:2602.07253v2 Announce Type: replace Abstract: Detecting hallucinations in large language models is a critical open problem with significant implications for safety and reliability. While existin

← Previous
1…183184185186187…358
Next →