AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
2 Jun 2026

Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories

Model ReleasesDGX agent

arXiv:2606.02060v1 Announce Type: new Abstract: Deep-research agents solve tasks through long trajectories of search, tool use, evidence inspection, and answer synthesis. Evaluation based on final ans

Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025

ResearchDGX agent

arXiv:2606.02255v1 Announce Type: cross Abstract: Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave unclear who p

Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2511.05613v2 Announce Type: replace-cross Abstract: Foundation models are increasingly central to high-stakes AI systems, and governance frameworks now depend on evaluations to assess their risk

Why Do Time Series Models Need Long Context Windows?

ApplicationsDGX agent

arXiv:2606.01999v1 Announce Type: cross Abstract: Modern deep learning models for forecasting groups of time series rely on increasingly longer observation windows. However, the benefit of increasing

Why Not Hyperparameter-Friendly Optimisation? A Monotonic Adaptive Norm Rescaling Approach For Long-Tailed Recognition

Model ReleasesDGX agent

arXiv:2606.02526v1 Announce Type: cross Abstract: Long-tailed recognition poses a significant challenge for deep learning. The two-stage decoupling paradigm, which separates representation learning fr

WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis

Model ReleasesDGX agent

arXiv:2606.01869v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly asked not only to write static interfaces, but to construct executable interactive worlds from natural lan

XAI-SOH-FL: Enhancing SOH-FL with Adaptive Aggregation and Explainable AI for Intrusion Detection in Heterogeneous IoT

Model ReleasesDGX agent

arXiv:2606.00134v1 Announce Type: cross Abstract: Intrusion Detection Systems (IDS) in Internet of Things (IoT) environments face significant challenges due to data heterogeneity, lack of labeled data

You Can Learn Tokenization End-to-End with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2602.13940v2 Announce Type: replace-cross Abstract: Tokenization is a hardcoded compression step which remains in the training pipeline of Large Language Models (LLMs), despite a general trend t

You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models

TutorialsDGX agent

arXiv:2603.00133v2 Announce Type: replace-cross Abstract: Generative models have been shown to 'memorize' certain training data, leading to verbatim or near-verbatim generating images, which may cause

Zamba2-VL Technical Report

Model ReleasesDGX agent

arXiv:2606.00390v1 Announce Type: cross Abstract: We present Zamba2-VL, a suite of vision-language models built on Zamba2, a hybrid language-model architecture combining Mamba2 state-space layers with

Zero-Shot Off-Policy Learning

Model ReleasesDGX agent

arXiv:2602.01962v2 Announce Type: replace-cross Abstract: Off-policy learning methods seek to derive an optimal policy directly from a fixed dataset of prior interactions. This objective presents sign

1 Jun 2026

A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents

AgentsDGX agent

arXiv:2602.08964v2 Announce Type: replace-cross Abstract: Understanding an agent's goals helps explain and predict its behaviour, yet there is no established methodology for reliably attributing goals

A Kinetic Energy Perspective of Flow Matching

Model ReleasesDGX agent

arXiv:2602.07928v2 Announce Type: replace-cross Abstract: Flow-based generative models can be viewed through a physics lens: sampling transports a particle from noise to data by integrating a learned

A Novel Global Context-aware Deep Neural Network for Enhanced Brain Tumor Segmentation using Magnetic Resonance Images

Model ReleasesDGX agent

arXiv:2605.30510v1 Announce Type: cross Abstract: Brain cancer's severity necessitates precise brain tumor segmentation, which is crucial for effective brain tumor diagnosis. Manual identification, bu

A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI

SafetyDGX agent

arXiv:2605.31021v1 Announce Type: new Abstract: Current alignment paradigms for generative artificial intelligence rely predominantly on monolithic benchmarking frameworks that reduce the plurality of

A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models

ResearchDGX agent

arXiv:2605.31080v1 Announce Type: cross Abstract: Blind and low-vision (BLV) audiences remain underserved by visual art descriptions, particularly across languages and in museum settings where privacy

A Unified and Reproducible Experimentation Framework for Speech Understanding

AgentsDGX agent

arXiv:2605.30899v1 Announce Type: cross Abstract: Speech foundation models and Speech LLMs have advanced speech understanding, yet deployment-oriented model selection is hindered by non-comparable eva

A Unified Framework for Gradient Aggregation in Multi-Objective Optimization

SafetyDGX agent

arXiv:2605.30452v1 Announce Type: cross Abstract: Many machine learning problems involve multiple inherent trade-offs that are best addressed by gradient-based multi-objective optimization (MOO) algor

Active Timepoint Selection for Learning Measure-Valued Trajectories

SafetyDGX agent

arXiv:2605.30625v1 Announce Type: cross Abstract: Inferring continuous probability paths from sparse snapshots is a fundamental challenge in domains like single-cell biology, where high-fidelity data

AI Loss of Control Incident Management: Response & Resilience

SafetyDGX agent

arXiv:2605.30406v1 Announce Type: cross Abstract: Recent research demonstrating AI systems exhibiting deception and shutdown resistance suggests that AI loss of control (LOC) is an urgent policy conce

AMix-2: Establishing Protein as a Native Modality in Large Language Models

Model ReleasesDGX agent

arXiv:2605.30963v1 Announce Type: cross Abstract: We present AMix-2, a protein-text foundation model that establishes protein as a native modality in large language models (LLMs), unifying protein und

An Odd Estimator for Shapley Values

Model ReleasesDGX agent

arXiv:2602.01399v2 Announce Type: replace-cross Abstract: The Shapley value is a ubiquitous framework for attribution in machine learning, encompassing feature importance, data valuation, and causal i

An Organization-Scoped LLM Agent Runtime Architecture for Regulated Cybersecurity Operations

Local AiDGX agent

arXiv:2605.30604v1 Announce Type: cross Abstract: Regulated cybersecurity workflows lack a runtime substrate that enforces organization-level scope across retrieval, tool calls, memory, findings, repo

AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music Editing

ResearchDGX agent

arXiv:2605.31053v1 Announce Type: cross Abstract: Controllable music editing is to modify high-level attributes while strictly preserving rhythmic and melodic structures. However, this task is challen

Annealed Softmax Greedy in Many-Armed Bayesian Bandits

SafetyDGX agent

arXiv:2605.31034v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and group-based policy optimization methods such as GRPO update a stochastic policy by sampling

Answer-Set-Programming-based Abstractions for Reinforcement Learning

AgentsDGX agent

arXiv:2605.31444v1 Announce Type: new Abstract: Reinforcement Learning (RL) enables autonomous agents to learn policies from experience, but realistic problems often involve enormous state spaces, mak

Appropriateness of Empathy in AI: A Signal-Cost Perspective

TutorialsDGX agent

arXiv:2605.31340v1 Announce Type: cross Abstract: The appropriateness of empathy in AI has emerged as a critical concern, as excessive empathy risks seeming manipulative while insufficient empathy app

Auto-Discovery-Bench: Diagnosing Structured State Tracking in Oracle-Guided Discovery

Model ReleasesDGX agent

arXiv:2502.15224v2 Announce Type: replace-cross Abstract: Interactive discovery requires agents to maintain and update structured beliefs over many rounds of feedback. Before evaluating agents in nois

Automatically Attacking Software Reverse Engineering AI Agents

AgentsDGX agent

arXiv:2605.30667v1 Announce Type: cross Abstract: Software tools for reverse engineering executable binary files, such as Ghidra, enable malware analysts to safely conduct robust static analysis witho

Autoregressive Visual Generation Needs a Prologue

ResearchDGX agent

arXiv:2605.06137v2 Announce Type: replace-cross Abstract: In this work, we propose Prologue, an approach to bridging the reconstruction-generation gap in autoregressive (AR) image generation. Instead

AutoSci: A Memory-Centric Agentic System for the Full Scientific Research Lifecycle

AgentsDGX agent

arXiv:2605.31468v1 Announce Type: new Abstract: Scientific research has traditionally been human-intensive, requiring researchers to coordinate literature, ideas, experiments, manuscripts, and review

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education

Model ReleasesDGX agent

arXiv:2605.31212v1 Announce Type: cross Abstract: AI systems are increasingly used to support educational content creation, yet it remains unclear whether they can generate outputs that faithfully rep

Benchmarking Machine Learning Uncertainty Quantification Methodologies for Predicting Turbine Gas Temperature Degradation

SafetyDGX agent

arXiv:2605.30585v1 Announce Type: cross Abstract: Effective prognostics and health management of modern engines relies on accurate turbine gas temperature predictions and robust uncertainty quantifica

Beyond Agreement: Scoring Panel-Surfaced Biomedical Entity Candidates for Curator Triage

Model ReleasesDGX agent

arXiv:2605.30826v1 Announce Type: cross Abstract: Biomedical NER is deceptively simple for modern LLMs: plausible biomedical mentions are easy to surface, but corpus-convention correctness depends on

Beyond Classification: Dynamic Adapter Routing for Continual Multimodal Retrieval

ResearchDGX agent

arXiv:2605.31229v1 Announce Type: cross Abstract: While retrieval is a core function of vision-language models, continually updating these models for retrieval tasks remains critically underexplored.

Beyond Memorization: Assessing Semantic Generalization in Large Language Models Using Phrasal Constructions

ApplicationsDGX agent

arXiv:2501.04661v3 Announce Type: replace-cross Abstract: The web-scale of pretraining data has created an important evaluation challenge: to disentangle linguistic competence on cases well-represente

Biases in the Blind Spot: Detecting What LLMs Fail to Mention

SafetyDGX agent

arXiv:2602.10117v5 Announce Type: replace-cross Abstract: Large Language Models (LLMs) often provide chain-of-thought (CoT) reasoning traces that appear plausible, but may hide internal biases. We cal

BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs

Model ReleasesDGX agent

arXiv:2605.30900v1 Announce Type: new Abstract: Current multimodal models handle static image recognition well, but intuitive physical reasoning remains a weakness. Predicting how objects will move an

BlueFin: Benchmarking LLM Agents on Financial Spreadsheets

Model ReleasesDGX agent

arXiv:2605.30907v1 Announce Type: cross Abstract: We present BlueFin, a benchmark that tasks large language model (LLM) agents with synthesis, manipulation, and comprehension tasks over spreadsheet wo

Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies

Model ReleasesDGX agent

arXiv:2512.19673v3 Announce Type: replace-cross Abstract: Existing reinforcement learning (RL) approaches treat large language models (LLMs) as a unified policy, overlooking their internal mechanisms.

Boundary-Guided Policy Optimization for Memory-efficient RL of Diffusion Large Language Models

SafetyDGX agent

arXiv:2510.11683v3 Announce Type: replace-cross Abstract: A key challenge in applying reinforcement learning (RL) to diffusion large language models (dLLMs) is the intractability of their likelihood f

Breaking Information Cocoons: A Hyperbolic Framework for Balancing Exploration and Exploitation in Recommender Systems

SafetyDGX agent

arXiv:2411.13865v4 Announce Type: replace-cross Abstract: Modern recommender systems often create information cocoons, restricting users' exposure to diverse content. The central challenge is to balan

Breaking the Simplification Bottleneck in Amortized Neural Symbolic Regression

Model ReleasesDGX agent

arXiv:2602.08885v5 Announce Type: replace-cross Abstract: Symbolic regression (SR) aims to discover interpretable analytical expressions that accurately describe observed data. Amortized SR promises t

Calibrated Preference Learning: The Case of Label Ranking

Model ReleasesDGX agent

arXiv:2605.30447v1 Announce Type: cross Abstract: Calibration, the alignment of predicted probabilities with true outcome frequencies, is essential for reliable decision-making. While extensively stud

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects

Local AiDGX agent

arXiv:2510.14904v3 Announce Type: replace-cross Abstract: Dense Video Object Captioning (DVOC) is the task of jointly detecting, tracking, and captioning object trajectories in a video, requiring the

Certified Circuits: Stability Guarantees for Mechanistic Circuits

ResearchDGX agent

arXiv:2602.22968v3 Announce Type: replace Abstract: Understanding how neural networks arrive at their predictions is essential for debugging, auditing, and deployment. Mechanistic interpretability pur

Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Model ReleasesDGX agent

arXiv:2503.08679v5 Announce Type: replace Abstract: Recent studies indicate that when faced with explicit biases in prompts, models often omit mentioning these biases in their Chain-of-Thought (CoT) o

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS

ResearchDGX agent

arXiv:2605.30748v1 Announce Type: cross Abstract: We present Chatterbox-Flash, a zero-shot text-to-speech model obtained by fine-tuning a pretrained autoregressive TTS decoder into a block-diffusion d

Choosing the Lens: Strategic Perspective Activation in Context-Dependent Argumentation

AgentsDGX agent

arXiv:2605.31581v1 Announce Type: new Abstract: The same arguments often need to be evaluated under different external regimes. An agent with influence over the regime has a strategic lever that stand

Circuit-Inspired High-Order Neural Networks with Unified Neural Dynamics Modeling for PDE Solving and Visual Perception

ResearchDGX agent

arXiv:2603.23977v2 Announce Type: replace-cross Abstract: Deep networks often rely on architectural heuristics to shape representation evolution, limiting their ability to model data governed by intri

CobSeg: Coherence Boundary Modeling for Dialogue Topic Segmentation

Local AiDGX agent

arXiv:2605.30668v1 Announce Type: cross Abstract: Dialogue topic segmentation is critical in many human-AI collaborative applications which requires identifying heterogeneous boundary cues, including

CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models

Model ReleasesDGX agent

arXiv:2605.30394v1 Announce Type: cross Abstract: This paper introduces Code Bench, a benchmark capable of evaluating Large Language Models (LLMs) concise code generation abilities in 60 programming l

COFT: Counterfactual-Conformal Decoding for Fair Chain-of-Thought Reasoning in Large Language Models

SafetyDGX agent

arXiv:2605.30641v1 Announce Type: cross Abstract: Large language models (LLMs) can reveal and amplify societal biases during chain-of-thought (CoT) generation. We present COFT (Chain of Fair Thought),

COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation

AgentsDGX agent

arXiv:2605.31264v1 Announce Type: new Abstract: LLM agents are increasingly expected not only to complete isolated tasks, but also to carry bounded representations of human expertise, judgment, and in

Comparing LLM-Based Conversational and Graphical Interfaces for Industrial Decision Tasks: An Exploratory Mixed-Methods Study

AgentsDGX agent

arXiv:2605.31224v1 Announce Type: cross Abstract: The use of Generative AI Conversational User Interfaces (CUI) as a new way to access and analyze data is growing in all sectors, and the industrial on

COMPASS: Cognitive MCTS-Guided Process Alignment for Safe Search Agents

SafetyDGX agent

arXiv:2605.30838v1 Announce Type: new Abstract: LLM-powered search agents enable multi-step reasoning and tool use. However, these capabilities introduce retrieval-induced safety degradation, as harmf

Conditional Coverage Diagnostics for Conformal Prediction

Model ReleasesDGX agent

arXiv:2512.11779v2 Announce Type: replace-cross Abstract: Evaluating conditional coverage remains one of the most persistent challenges in assessing the reliability of predictive systems. Although con

ConSensus: Multi-Agent Collaboration for Multimodal Sensing

SafetyDGX agent

arXiv:2601.06453v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly grounded in sensor data to perceive and reason about human physiology and the physical world. However,

ConTrans: Learning Text-enhanced Local-global Temporal Representations for Zero-shot Temporal Action Localization

Model ReleasesDGX agent

arXiv:2605.30689v1 Announce Type: cross Abstract: Zero-shot Temporal Action Localization (ZS-TAL) aims to detect and locate previously unseen actions in untrimmed videos. However, existing approaches

Controllable Lung Nodule Synthesis via Histogram-Regularized Latent Diffusion Models

ResearchDGX agent

arXiv:2605.30631v1 Announce Type: cross Abstract: While automated diagnosis systems have achieved remarkable success in computed tomography (CT)-based lung cancer screening, their development remains

← Previous
1…182183184185186…358
Next →