AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
25 Jun 2026

AI Snitches Get Glitches: Towards Evading Agentic Surveillance

AgentsDGX agent

arXiv:2606.25836v1 Announce Type: new Abstract: To better assist users with completing challenging tasks, AI agents mediate communications, access data, and interact with different APIs. Many employer

AIChilles: Automatically Uncovering Hidden Weaknesses in AI-Evolved Systems

Model ReleasesDGX agent

arXiv:2606.15834v2 Announce Type: replace Abstract: The computer systems community has recently seen growing interest in AI-driven system evolution, where AI agents iteratively rewrite systems. Framew

An Approach for a Supporting Multi-LLM System for Automated Certification Based on the German IT-Grundschutz

ResearchDGX agent

arXiv:2606.25608v1 Announce Type: cross Abstract: This paper presents a novel approach to perform semi-automated BSI IT-Grundschutz certification using a MultiLarge Language Model system (MLS) with Hy


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Attractive and Repulsive Pattern Control in Sequence Generation

ResearchDGX agent

arXiv:2606.24911v1 Announce Type: cross Abstract: Variable-order Markov models preserve local symbolic syntax by adapting context length, but long continuations can enter recurring high-order 'tunnels

AutoRelAnnotator: Calibrated Model Cascades for Cost-Efficient Relevance Evaluation in Sponsored Search

ApplicationsDGX agent

arXiv:2606.25871v1 Announce Type: cross Abstract: How can we generate high-quality relevance annotations at scale without the cost and delays of human labeling? Relevance annotations are the backbone

BCoughBench: Benchmarking Respiratory Acoustic Foundation Models Under Body-Coupled Wearable Sensor Conditions

ResearchDGX agent

arXiv:2606.25116v1 Announce Type: cross Abstract: Respiratory acoustic foundation models (FMs) are benchmarked exclusively on smartphone recordings, yet clinical deployment increasingly targets body-c

Beyond Shapley: Efficient Computation of Asymmetric Shapley Values

ResearchDGX agent

arXiv:2606.25103v1 Announce Type: new Abstract: We address the problem of explainability in machine learning models through feature attribution methods. In particular, we consider a variant of Shapley

BrainAgent: A Large Language Model-Driven Multi-Agent Framework for Autonomous Brain Signal Understanding

Model ReleasesDGX agent

arXiv:2606.25400v1 Announce Type: new Abstract: Brain-Computer Interfaces (BCIs) and brain signal understanding are pivotal for clinical health and next-generation interactions. Despite this significa

Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent Ecosystem

AgentsDGX agent

arXiv:2606.26028v1 Announce Type: cross Abstract: As autonomous AI agents increasingly transact across organizational boundaries, a fundamental trust challenge emerges: how can an agent assess whether

CausalRAG2: Hierarchical Causal Knowledge Graph Design for RAG

Model ReleasesDGX agent

arXiv:2602.05143v2 Announce Type: replace Abstract: Retrieval augmented generation (RAG) has enhanced large language models by enabling access to external knowledge, with graph-based RAG emerging as a

Confidence Sequences for Online Statistical Model Checking of Markov Decision Processes

ResearchDGX agent

arXiv:2606.25797v1 Announce Type: new Abstract: Markov decision processes (MDPs) are a classic model of decision making under uncertainty, exhibiting both non-deterministic choice as well as probabili

Conformal Recovery-Deadline Certificates for Runtime Assurance of Adapting Controllers

SafetyDGX agent

arXiv:2606.25371v1 Announce Type: cross Abstract: Runtime assurance (RTA) protects a safety-critical system by switching from an advanced controller to a verified safe controller when a monitored cond

CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations

ResearchDGX agent

arXiv:2606.25403v1 Announce Type: cross Abstract: Accent conversion and controllability remain fundamental challenges in cross-lingual text-to-speech (TTS), particularly for low-resource and phonetica

Decoupling Reconnaissance and Exploitation: Measuring the Capability Boundaries of LLM-Based Web Penetration Testing

AgentsDGX agent

arXiv:2606.25332v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown promise for automated penetration testing, yet existing end-to-end black-box evaluations are highly susceptibl

Domain-Specific Agents for Cherenkov Telescope Array Control Software and Gamma-Ray Data Analysis

AgentsDGX agent

arXiv:2510.01299v3 Announce Type: replace-cross Abstract: We present domain-adapted large language model agents designed to support Cherenkov Telescope Array operation and data analysis. The agents co

Elo-Disentangled Player-Style Embeddings for Human Chess via Rating-Conditioned Residual Move Model

Model ReleasesDGX agent

arXiv:2606.25176v1 Announce Type: new Abstract: We study representation learning for individual human chess style: a per-player embedding learned from a player's move history such that inner products

EmotionAI: A Privacy-Preserving Computational Intelligence Pipeline for Speech-Emotion-Grounded Conversational Analysis

Local AiDGX agent

arXiv:2606.24941v1 Announce Type: cross Abstract: Reviewing recorded interviews for affective cues such as composure, hesitation and agitation is slow and subjective, and cloud services that could aut

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users

ResearchDGX agent

arXiv:2606.24910v1 Announce Type: cross Abstract: Voice control offers an intuitive alternative to manual drone piloting, yet most existing systems rely on rigid command vocabularies that fail to hand

Epistemic Bias Injection: Manipulating LLM Opinion via Selective Context Retrieval

Model ReleasesDGX agent

arXiv:2512.00804v3 Announce Type: replace-cross Abstract: When answering user queries, LLMs often retrieve knowledge from external sources stored in retrieval-augmented generation (RAG) databases. The

Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?

AgentsDGX agent

arXiv:2602.11988v2 Announce Type: replace-cross Abstract: A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md. Although this

Explainable Control Framework (XCF) based on Fuzzy Model-Agnostic Explanation and LLM Agent-Supported Interface

Local AiDGX agent

arXiv:2606.25941v1 Announce Type: cross Abstract: Increasing demand for precise and reliable control in complex scenarios has led to the development of increasingly sophisticated controllers, includin

Exploring Information Seeking Agent Consolidation

Model ReleasesDGX agent

arXiv:2602.00585v2 Announce Type: replace Abstract: Information-seeking agents have emerged as a powerful paradigm for knowledge-intensive tasks, yet today's systems remain specialized for the open we

Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness

ResearchDGX agent

arXiv:2606.18874v2 Announce Type: replace Abstract: AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and final claim

Failure Modes of Large Language Models on Research-Level Mathematics: A Taxonomy and an Empirical Characterisation

Model ReleasesDGX agent

arXiv:2606.24902v1 Announce Type: cross Abstract: The 'First Proof' benchmark [1] posed ten research-level mathematics questions to the strongest publicly available LLMs and found them consistently wr

FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming

Model ReleasesDGX agent

arXiv:2606.19887v2 Announce Type: replace-cross Abstract: Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks. Financial LLMs face regulatory compliance vio

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models

Model ReleasesDGX agent

arXiv:2606.25391v1 Announce Type: cross Abstract: Recent Large Audio Language Models (LALMs) have achieved remarkable progress in audio perceptual tasks across individual acoustic layers, including sp

Fuzzy Quantification over OWL Ontologies and Knowledge Graphs

ResearchDGX agent

arXiv:2606.25778v1 Announce Type: new Abstract: This paper presents a versatile framework for evaluating fuzzy quantification queries over both standard and fuzzy ontologies as well as knowledge graph

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice

Model ReleasesDGX agent

arXiv:2606.22327v2 Announce Type: replace Abstract: The explosive demand for interactive Large Language Model serving has highlighted the management of the Key-Value cache's dynamic memory footprint a

GUI agent: Guided Exploration of User-Sensitive Screens

SafetyDGX agent

arXiv:2606.25705v1 Announce Type: new Abstract: LLM agents are increasingly being used to automate tasks for users within an open GUI environment. They inevitably encounter screens containing user-sen

Helpful or Harmful? Evaluating LLM-Assisted Vulnerability Patching via a Human Study

ApplicationsDGX agent

arXiv:2606.25973v1 Announce Type: cross Abstract: Software vulnerability remediation is a cognitively demanding task that requires specialized security expertise often lacking in general developers. I

Heuresis: Search Strategies for Autonomous AI Research Agents Across Quality, Diversity and Novelty

SafetyDGX agent

arXiv:2606.25198v1 Announce Type: new Abstract: Autonomous AI Research promises to accelerate the scientific progress of machine learning. To realise this goal, current Large Language Model (LLM)-base

How Small Can 6G Reason? Scaling Tiny-to-Small Language Models for AI-Native Networks

Model ReleasesDGX agent

arXiv:2603.02156v2 Announce Type: replace-cross Abstract: Emerging 6G visions, reflected in ongoing standardization efforts within 3GPP, IETF, ETSI, ITU-T, and the O-RAN Alliance, increasingly charact

Improving Zero-Shot Offline RL via Behavioral Task Sampling

Model ReleasesDGX agent

arXiv:2604.25496v2 Announce Type: replace Abstract: Offline zero-shot reinforcement learning (RL) aims to learn agents that optimize unseen reward functions without additional environment interaction.

LibEvoBench: Probing Temporal Knowledge Stratification in Code Generation Models

Model ReleasesDGX agent

arXiv:2606.25402v1 Announce Type: cross Abstract: Large software projects often depend on older versions of libraries, even as APIs continue to evolve across releases. This creates a challenge for LLM

Lightweight PCGAE-Net: Parallel CrossGate Attention and Bottleneck AutoEncoder for Efficient 5G Channel Prediction

SafetyDGX agent

arXiv:2606.25401v1 Announce Type: cross Abstract: Accurate channel state information (CSI) prediction is essential for proactive beamforming and resource management in 5G massive MIMO systems, yet the

Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions

SafetyDGX agent

arXiv:2606.25396v1 Announce Type: new Abstract: AI companions powered by large language models increasingly interact with cognition-developing users, including children and adolescents, creating risks

Measurable Majorities Are Not Finitely Axiomatizable

ResearchDGX agent

arXiv:2606.25954v1 Announce Type: cross Abstract: This theoretical note studies the finite axiomatizability of strict majority reasoning in finite social decision frames. Moss and Pedersen (2026) intr

Offline Multi-agent Continual Cooperation via Skill Partition and Reuse

SafetyDGX agent

arXiv:2606.25389v1 Announce Type: new Abstract: Extracting skills from multi-agent offline dataset improves learning efficiency via sharing task-invariant coordination skills among tasks. In settings

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning

Model ReleasesDGX agent

arXiv:2606.25325v1 Announce Type: new Abstract: We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning tra

Phoneme-Level Mispronunciation Screening in Polish-Speaking Children with an Explainable Assistant

SafetyDGX agent

arXiv:2606.25181v1 Announce Type: cross Abstract: Early identification of speech sound errors in children is often limited by access to specialists, motivating lightweight screening tools that can ope

Plausible but Wrong: A case study on Agentic Failures in Astrophysical Workflows

AgentsDGX agent

arXiv:2604.25345v2 Announce Type: replace Abstract: Agentic AI systems are increasingly being integrated into scientific workflows, yet their behavior under realistic conditions remains insufficiently

Position Spaces and Graphs

SafetyDGX agent

arXiv:2606.25719v1 Announce Type: new Abstract: In this paper, we introduce position graphs, a graph-based reasoning framework based on the formalization of position spaces. This framework utilizes tw

Power-Flexible AI Data Centers: A New Paradigm for Grid-Responsive Compute

HardwareDGX agent

arXiv:2606.25098v1 Announce Type: cross Abstract: The rapid expansion of artificial intelligence (AI) infrastructure is driving unprecedented growth in electricity demand from data centers. Traditiona

Privacy-preserving federated tensor decomposition of single-cell immune data: recovering multicellular programs across institutions

ResearchDGX agent

arXiv:2606.24938v1 Announce Type: cross Abstract: Tensor decomposition of donor imes cell-type imes gene single-cell data recovers multicellular programs: coordinated axes of inter-individual transcri

Privacy Vulnerabilities of Attention Layers in Tabular Foundation Models and Protection of High-Risk Queries

ResearchDGX agent

arXiv:2606.26021v1 Announce Type: cross Abstract: Tabular foundation models are commonly assumed to present limited privacy concerns as they are often pre-trained on large collections of synthetic dat

Proactive Systems in HCI and AI: Concepts, Challenges, and Opportunities

AgentsDGX agent

arXiv:2606.25149v1 Announce Type: cross Abstract: The last few years have seen a significant rise in interest in highly autonomous and proactive systems, fueled by advances in AI. Systems that anticip

Probabilistic Agents in Deterministic Audits: Evaluating Multi-Agent Systems for Automated Audits Based on the German IT-Grundschutz

AgentsDGX agent

arXiv:2606.25622v1 Announce Type: cross Abstract: The NIS-2 Directive mandates robust Risk Management from thousands of small and medium enterprises. To ensure compliance, companies rely on establishe

Rate-Aware Quantum-Inspired Trajectory Learning for Interference-Limited Multi-UAV Networks

ResearchDGX agent

arXiv:2606.25480v1 Announce Type: cross Abstract: Unmanned aerial vehicle (UAV) can provide on-demand, high-capacity connectivity in disaster and normal situation. However, it faces a challenge of cur

ReviewGuard: Aligning LLM-Assisted Peer Review with Long-Term Scientific Impact

ResearchDGX agent

arXiv:2606.24892v1 Announce Type: cross Abstract: Peer review is central to scientific quality control, yet it can undervalue papers that later achieve substantial citation impact. While frontier larg

RWGBench: Evaluating Scholarly Positioning in Related Work Generation

Model ReleasesDGX agent

arXiv:2606.24894v1 Announce Type: cross Abstract: Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited. Existing R

SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety

Model ReleasesDGX agent

arXiv:2606.18936v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly embedded in AI for Science (AI4Science) workflows, from scientific question answering and literature a

SE-AGCNet: An End-to-End Framework for Joint Speech Enhancement and Loudness Control in Meeting Scenarios

ResearchDGX agent

arXiv:2606.25959v1 Announce Type: cross Abstract: Conventional audio pipelines typically treat speech enhancement (SE) and automatic gain control (AGC) as discrete modules, which often limits overall

Small Initialization Matters for Large Language Models

Model ReleasesDGX agent

arXiv:2606.17945v2 Announce Type: replace Abstract: Large language models provide a tractable system for asking how intelligence itself emerges, rather than only how LLMs can be engineered. Although p

SoK: AI Secure Code Generation: Progress, Pitfalls, and Paths Forward

AgentsDGX agent

arXiv:2606.25195v1 Announce Type: cross Abstract: The increasing use of AI systems for code generation raises a central security question: what can today's models and coding agents actually do to prod

Specifying AI-SDLC Processes: A Protocol Language for Human-Agent Boundaries

SafetyDGX agent

arXiv:2606.20615v2 Announce Type: replace Abstract: AI agents now participate as first-class team members across the software development lifecycle, yet no specification language exists for expressing

STEB: A Speech-to-Speech Translation Expressiveness Benchmark for Evaluating Beyond Translation Fidelity

Model ReleasesDGX agent

arXiv:2606.25529v1 Announce Type: cross Abstract: Speech-to-speech translation (S2ST) should preserve not only lexical meaning, but also expressive attributes: emotion, scenario style (e.g., news repo

Supervised Post-training of Speech Foundation Models for Robust Adaptation in Speech Deepfake Detection

Local AiDGX agent

arXiv:2606.25328v1 Announce Type: cross Abstract: Large speech foundation models have shown strong potential for speech deepfake detection, but direct fine-tuning is limited by a mismatch between self

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care

SafetyDGX agent

arXiv:2601.16529v3 Announce Type: replace Abstract: Large language models (LLMs) deployed in clinical decision support may acquiesce to patient requests for care that conflicts with evidence-based gui

Taxonomy of Risks on Automated Fact-Checking Systems Considering its Propagation

TutorialsDGX agent

arXiv:2606.25645v1 Announce Type: cross Abstract: In recent years, the posting of fake news including disinformation and misinformation on social networking services (SNS) has become a social problem.

The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing

SafetyDGX agent

arXiv:2606.25108v1 Announce Type: new Abstract: Autonomous AI systems are transitioning from advisory to autonomous roles for medication prescriptions. Recent United States bill H.R. 238 and Utah's pr

← Previous
1…124125126127128…358
Next →