AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
5 Aug 2026

Diversity is Not Ambiguity: Toward Accurate and Efficient Ambiguity Detection for Open-Domain QA

Model ReleasesDGX agent

arXiv:2608.03177v1 Announce Type: new Abstract: How can question answering (QA) systems determine whether a query is ambiguous? Ambiguity detection is essential in open-domain QA, as misclassification

DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning

SafetyDGX agent

arXiv:2608.03292v1 Announce Type: new Abstract: Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason over heterogeneo

Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.03791v1 Announce Type: new Abstract: Vision-Language Models (VLMs), like Large Language Models (LLMs), may memorize sensitive, copyrighted, or harmful knowledge from their pretraining corpo

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR

SafetyDGX agent

arXiv:2608.03119v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Vo

Don't Regenerate, Debug: A Domain-Specific Agent for Repairing Near-Miss Hardware Operators

AgentsDGX agent

arXiv:2608.02712v1 Announce Type: cross Abstract: Kernel generation for hardware accelerators such as GPUs and NPUs has become a proving ground for large language models (LLMs), and state-of-the-art s

dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

Model ReleasesDGX agent

arXiv:2608.02673v1 Announce Type: cross Abstract: Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language pr

Dr. AGENTONOMICS: A Didactic Experiment of AGENTONOMICS

AgentsDGX agent

arXiv:2608.03524v1 Announce Type: new Abstract: AGENTONOMICS is a framework that treats AI agents as economic entities that can be designed, managed, and governed through an integrated management arch

EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners

Model ReleasesDGX agent

arXiv:2608.03206v1 Announce Type: cross Abstract: Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only re

Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss

Model ReleasesDGX agent

arXiv:2608.03796v1 Announce Type: cross Abstract: Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained fro

Efficient unsupervised domain adaptation via self-supervised vision transformer and synergistic cross-domain alignment

Model ReleasesDGX agent

arXiv:2407.21311v2 Announce Type: replace-cross Abstract: Unsupervised domain adaptation (UDA) aims to mitigate domain shift, where the distribution of labeled source data differs from that of unlabel

Efficient Video Dataset Distillation via Cluster-Guided Prototype Blending

ResearchDGX agent

arXiv:2608.03269v1 Announce Type: cross Abstract: Video dataset distillation aims to compress a large video dataset into a compact surrogate set that preserves its training utility. Most existing appr

EFX Allocation In (Multi)Hypergraphs

ResearchDGX agent

arXiv:2608.03171v1 Announce Type: cross Abstract: We study fair allocations of indivisible goods among agents with heterogeneous monotone valuations. As fair we consider the allocations that are envy-

Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation

SafetyDGX agent

arXiv:2608.03044v1 Announce Type: cross Abstract: Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignm

Enactive Artificial Intelligence: A Decision-Centric Architecture for Complex Systems

ApplicationsDGX agent

arXiv:2608.03413v1 Announce Type: new Abstract: As artificial intelligence (AI) continues to evolve and mature, recent AI practices have moved beyond large language models (LLMs) and text or image gen

Enhancing Tabular Learners with Context-Aware Semantic Embeddings

Model ReleasesDGX agent

arXiv:2608.03565v1 Announce Type: new Abstract: While modern tabular learners excel at capturing statistical patterns, they frequently operate in a semantic vacuum, treating textual features as discre

Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning

SafetyDGX agent

arXiv:2608.03875v1 Announce Type: cross Abstract: Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-Language Mode

Equivariant Music Transformer

ResearchDGX agent

arXiv:2608.03920v1 Announce Type: cross Abstract: Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the representation s

Evading Chain-of-Thought Monitoring Through Model Poisoning

SafetyDGX agent

arXiv:2608.02820v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring is an increasingly important component of AI safety stacks but relies on the assumption that a model's reasoning tra

Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning

Model ReleasesDGX agent

arXiv:2608.03028v1 Announce Type: new Abstract: Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations

Evaluating LLM Trade-offs for Enterprise Automation: Lessons from Workflow Generation in a Production Enterprise Platform

Model ReleasesDGX agent

arXiv:2608.03311v1 Announce Type: cross Abstract: Enterprise compliance management requires rapid adaptation to evolving regulatory frameworks (e.g., DORA, AI RMF, FedRAMP) and tight remediation SLAs.

Evaluating LLMs in Database Scenarios: A Lifecycle Benchmark for Assessing Their Potential in Core Database Tasks

Model ReleasesDGX agent

arXiv:2608.03794v1 Announce Type: cross Abstract: Large Language Models (LLMs) are transforming database interaction paradigms, evolving from simple query translators to autonomous database administra

Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks

Model ReleasesDGX agent

arXiv:2608.02616v1 Announce Type: cross Abstract: We present the first independent, systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, across 42 synth

Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning

ResearchDGX agent

arXiv:2608.03161v1 Announce Type: new Abstract: Lecture videos distribute knowledge across speech, slide text, diagrams, equations, and presentation order, which transcript-only retrieval does not ful

Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation Gap

ApplicationsDGX agent

arXiv:2608.02699v1 Announce Type: new Abstract: When algorithms make or influence consequential decisions---about loan eligibility, hiring, or healthcare---EU law grants affected individuals a Right t

Externally Validated Breast Ultrasound Segmentation via Multi-task Learning with BI-RADS-Consistent Morphological Priors

Model ReleasesDGX agent

arXiv:2511.15968v2 Announce Type: replace-cross Abstract: External validation of breast ultrasound segmentation models remains limited because internal train--test splits do not capture domain shifts

FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact

ResearchDGX agent

arXiv:2608.03372v1 Announce Type: cross Abstract: AI systems rewrite information constantly: conversations become stored memories, documents become answers. The rewrite can keep a claim while washing

Fail-Fast, Restart-Smart: Early Failure Prediction and Restart for SWE Agentic Tasks

Model ReleasesDGX agent

arXiv:2608.03222v1 Announce Type: cross Abstract: Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates. F

Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement

ResearchDGX agent

arXiv:2608.03733v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-

FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection

Model ReleasesDGX agent

arXiv:2608.03096v1 Announce Type: cross Abstract: Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks rem

FinVerse: Financial Time-Series Benchmark

Model ReleasesDGX agent

arXiv:2608.03259v1 Announce Type: cross Abstract: As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has become incre

FlowForm: Synergizing Fluid Physics with Topological Consistency for Satellite Flood Synthesis

ResearchDGX agent

arXiv:2608.03822v1 Announce Type: cross Abstract: Developing robust flood assessment models requires high-quality paired satellite imagery, yet such data remain scarce for flood-specific image generat

Formal Verification of Agentic Systems over Operational Data

AgentsDGX agent

arXiv:2608.03609v1 Announce Type: new Abstract: Agentic systems driven by large language models (LLMs) are increasingly deployed in real-world workflows where they act on persistent operational data.

FOUND-AF: Benchmarking ECG Foundation Models for Atrial Fibrillation Detection

ResearchDGX agent

arXiv:2608.03597v1 Announce Type: new Abstract: Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia and is associated with increased risks of stroke, heart failure, and mortality.

FraQ: Efficient Coordinate-Space Recompression for Federated Low-Rank Adaptation

ResearchDGX agent

arXiv:2608.03605v1 Announce Type: new Abstract: Federated fine-tuning with Low-Rank Adaptation (LoRA) enables efficient collaborative adaptation of Large Language Models (LLMs) without centralizing pr

From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model

Model ReleasesDGX agent

arXiv:2508.00955v3 Announce Type: replace-cross Abstract: Adapting generative Multimodal Large Language Models (MLLMs) into universal embedding models typically demands resource-intensive contrastive

From Social Coding to Agentic Coding: Productivity and Relational Reconfiguration in Open-Source Communities

Model ReleasesDGX agent

arXiv:2608.03585v1 Announce Type: new Abstract: Open-source software communities are a form of digital public infrastructure that not only produces code, but also generates public knowledge and interp

From Wearable Data to Personalized and Actionable Health Insights

ResearchDGX agent

arXiv:2608.03251v1 Announce Type: cross Abstract: Commercial wearable devices continuously capture rich physiological data (e.g., heart rate, respiration), opening new possibilities for monitoring hea

GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement

ResearchDGX agent

arXiv:2604.01832v1 Announce Type: cross Abstract: We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge. The system integrates a

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks

Model ReleasesDGX agent

arXiv:2608.03764v1 Announce Type: new Abstract: Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-ev

GENESIS: Towards Explainable Causal Discovery

Model ReleasesDGX agent

arXiv:2608.03868v1 Announce Type: cross Abstract: Causal Discovery (CD) from observational data faces two fundamental challenges. First, purely statistical methods often lack the power to resolve stru

GenOS: Compositional Certificates for Semantic Robustness in AI Code Generation

SafetyDGX agent

arXiv:2608.03588v1 Announce Type: cross Abstract: AI coding agents are stochastic workflows: prompts are interpreted, artifacts are sampled, validators produce observations, and orchestrators commit o

Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls

Model ReleasesDGX agent

arXiv:2608.03071v1 Announce Type: new Abstract: Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting the right tool

GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models

ResearchDGX agent

arXiv:2608.03729v1 Announce Type: cross Abstract: Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large language mode

GraphCliff: Short-Long Range Gating for Modeling Critical Activity Changes Caused by Subtle Molecular Differences

ResearchDGX agent

arXiv:2511.03170v3 Announce Type: replace-cross Abstract: The quantitative structure-activity relationship assumes a smooth mapping between molecular structure and biological activity. However, activi

GROW: Group-Relative Advantage-Weighted On-Policy Reinforcement Learning of Autoregressive-Diffusion Text-to-Speech model

SafetyDGX agent

arXiv:2608.03215v1 Announce Type: cross Abstract: Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient methods typical

GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs

Model ReleasesDGX agent

arXiv:2608.03270v1 Announce Type: cross Abstract: GUI grounding maps natural-language instructions to click locations and is essential for reliable GUI agents. The task remains difficult on high-resol

HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize

SafetyDGX agent

arXiv:2601.03321v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have substantially advanced Radiology Report Generation (RRG), yet aligning them through reinforcemen

How Closely Do LLM Reviews Align with Human Peer Review?

Model ReleasesDGX agent

arXiv:2608.03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers

How Many Labels Are Enough? ALDA: Active Learning Deployment Advisor for Medical Image Classification

ResearchDGX agent

arXiv:2608.03511v1 Announce Type: cross Abstract: Active learning (AL) promises to reduce the cost of medical imaging projects by lowering the number of clinical labels required. However, practical de

Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks

AgentsDGX agent

arXiv:2608.03502v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous agents. Howe

HyperAgent: Planning and Acting over Tool-Schema Hypergraphs for Tool-Use LLM Agents

AgentsDGX agent

arXiv:2608.02650v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on external tools to complete complex real-world tasks. However, reliable tool-use planning remains

Hypercubes, Hyperplanes, and Constraint-Induced Complexity Collapse in Atomic Concept Learning

ResearchDGX agent

arXiv:2608.02930v1 Announce Type: new Abstract: We revisit higher-arity atomic concept learning through the geometry of hypercubes and hyperplanes of ground instances. Our starting point is the observ

HyperFL: Query-Adaptive Representation Learning for Software Fault Localization

Model ReleasesDGX agent

arXiv:2608.02967v1 Announce Type: cross Abstract: Software fault localization identifies the code locations responsible for reported issues and is a fundamental step toward automated debugging and pro

Implementing Causal Perception: Competing SCMs and Situated Fairness

SafetyDGX agent

arXiv:2608.03917v1 Announce Type: new Abstract: Causal perception occurs when agents with competing Structural Causal Models (SCMs) of the same system infer different probability distributions, includ

Improved Quantum Algorithms for Reinforcement Learning Under a Generative Model

SafetyDGX agent

arXiv:2608.02826v1 Announce Type: cross Abstract: Reinforcement learning is a subfield of machine learning that studies how an agent interacts with an environment in order to extract as large a reward

In-Context Collapse in Vision-Language Models and How to Mitigate it?

Model ReleasesDGX agent

arXiv:2608.02830v1 Announce Type: cross Abstract: Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and is widely as

In-Context Pure Exploration in Continuous Decision Spaces

Model ReleasesDGX agent

arXiv:2602.17976v2 Announce Type: replace-cross Abstract: In active sequential testing, also termed pure exploration, a learner is tasked with the goal to adaptively acquire information so as to ident

Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt Compilation

Model ReleasesDGX agent

arXiv:2608.02639v1 Announce Type: cross Abstract: Production prompts rarely carry a single instruction. One system message may require valid JSON, a word limit, three citations, and a fixed tone at th

Internalising the Identity Primitive: Cryptographic Individuality for an Autonomous Agent on a Public Blockchain

AgentsDGX agent

arXiv:2608.02986v1 Announce Type: cross Abstract: A software agent on a public blockchain accumulates authority and economic stakes, raising the engineering question of what makes it count as an indiv

Internalizing Academic Writing Workflows for Introduction Generation via Struct-Aware Policy Learning

Model ReleasesDGX agent

arXiv:2608.03138v1 Announce Type: cross Abstract: Generating a rigorous paper introduction with large language models (LLMs) remains challenging, since it requires coordinating background, gap identif

← Previous
1…3132333435…354
Next →