AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
5 Aug 2026

AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities

ResearchDGX agent

arXiv:2608.03742v1 Announce Type: cross Abstract: Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of v

AI Forensics Across White-, Grey-, and Black-Box Access: A Process Model and Research Agenda for Post-Incident Investigation of AI Systems

ResearchDGX agent

arXiv:2608.03520v1 Announce Type: cross Abstract: AI systems are increasingly involved in decisions and actions that may later require investigation. When an AI related incident occurs, investigators

AI Sandbox: Technical Report

TutorialsDGX agent

arXiv:2608.02679v1 Announce Type: cross Abstract: Collaborative AI experimentation across industry and academia requires platforms that enable rapid prototyping while preserving controlled access, ten


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

AI Security Leaderboard: Methodology, Results and Minimal Standard

Model ReleasesDGX agent

arXiv:2608.03070v1 Announce Type: cross Abstract: Frontier AI model developers increasingly rely on layered safeguards to prevent catastrophic misuse, but little public evidence exists on how much pro

AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction

Model ReleasesDGX agent

arXiv:2608.03416v1 Announce Type: new Abstract: Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different

Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach

Local AiDGX agent

arXiv:2608.03204v1 Announce Type: cross Abstract: Post-training reinforcement learning (RL) algorithms are commonly used to align large vision-language models (LVLMs) with human intent and the require

Approximate Speculative Decoding

Model ReleasesDGX agent

arXiv:2608.03447v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greedy verificat

Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution

AgentsDGX agent

arXiv:2608.01366v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are integral to complex intellectual tasks, yet output quality remains constrained by user-provided prompts. Iter

Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions

ResearchDGX agent

arXiv:2509.24457v1 Announce Type: cross Abstract: Objective speech-quality metrics are widely used to assess codec performance. However, for neural codecs, it is often unclear which metrics provide re

Assessing the Effect of Cross-Domain Mapping on Creativity in Humans and Large Language Models

ResearchDGX agent

arXiv:2603.19087v2 Announce Type: replace Abstract: Creativity is the ability to come up with novel ideas, a capacity crucial for human development and flourishing. Are large language models (LLMs) cr

Attribute-based Undetectable Watermarking for Generative AI Models

SafetyDGX agent

arXiv:2608.03174v1 Announce Type: cross Abstract: Generative AI systems increasingly produce content whose provenance is difficult to verify, motivating watermarking techniques for identifying model-g

Automated Visualization Code Synthesis via Multi-Path Reasoning and Feedback-Driven Optimization

Model ReleasesDGX agent

arXiv:2502.11140v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have become a cornerstone for automated visualization code generation, enabling users to create charts through na

Autoreflection: How Agentic Strange Loops Turn Human Culture into AI Infrastructure

AgentsDGX agent

arXiv:2608.03800v1 Announce Type: cross Abstract: An LLM-based agent is a loop that reads itself. Agentic frameworks externalize identity, memory, and disposition into editable files. The agent loads

AutoSND: From Execution Evidence to Structural Policies for Automated Network Dismantling Heuristic Discovery

ApplicationsDGX agent

arXiv:2608.03653v1 Announce Type: new Abstract: Network dismantling is fundamental to analyzing the robustness and vulnerability of complex systems, yet practical heuristics must balance effectiveness

Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs

Model ReleasesDGX agent

arXiv:2608.03450v1 Announce Type: cross Abstract: Reasoning in Multimodal Large Language Models (MLLMs) requires both fine-grained visual perception and rigorous logical deduction. Explicit text-based

BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL

SafetyDGX agent

arXiv:2608.02876v1 Announce Type: new Abstract: Tool-using agents do not merely consume observations: their actions determine what arrives next. In agentic text-to-SQL, a broad query can spend context

Before Reasoning Can Fail: Pre-Evidence Procedural Failures in Agentic RAG

AgentsDGX agent

arXiv:2608.02011v2 Announce Type: replace Abstract: Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candidate snipp

Behaviorally Adaptive Visual Diversion for Inclusive and Resilient Digital Assessment Delivery

ResearchDGX agent

arXiv:2608.03531v1 Announce Type: new Abstract: Institutions increasingly rely on browser lockdown, webcam monitoring, and behavioral analytics to secure high-stakes digital assessments, yet these mec

Beyond Average Performance: Dynamic Instance Clustering and Specialized Algorithm Design in LLM-Assisted Evolutionary Search

ApplicationsDGX agent

arXiv:2608.03129v1 Announce Type: new Abstract: Large Language Model-assisted Evolutionary Search (LES) has emerged as a powerful paradigm for automated algorithm design. However, existing LES methods

Beyond Either-Or Reasoning: Transduction and Induction as Cooperative Problem-Solving Paradigms

Model ReleasesDGX agent

arXiv:2505.14744v3 Announce Type: replace-cross Abstract: Traditionally, in Programming-by-example (PBE) the goal is to synthesize a program from a small set of input-output examples. Lately, PBE has

Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking

Model ReleasesDGX agent

arXiv:2608.03859v1 Announce Type: cross Abstract: Large language models (LLMs) pose challenges to academic integrity and peer review. Yet generative plagiarism detection remains an underexplored and l

Beyond the Hivemind: Escaping LLM Homogeneity via Meta-Persona Anchoring and Sequential Temperature Scaling

ResearchDGX agent

arXiv:2608.02618v1 Announce Type: new Abstract: Recent studies have identified an ``Artificial Hivemind'' effect in Large Language Models (LLMs) causing models to converge on a narrow, homogenized con

BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?

SafetyDGX agent

arXiv:2608.02867v1 Announce Type: cross Abstract: Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large language models (LLMs) across a variety of reason

BulkPR-Bench: Benchmarking Queue-Level Governance of Interacting Pull Requests

Model ReleasesDGX agent

arXiv:2608.02685v1 Announce Type: cross Abstract: Coding-agent benchmarks increasingly cover long-horizon, end-to-end, and interactive development, but typically retain one requested outcome or a fixe

CADET: Physics-Grounded Causal Auditing and Training-Free Deconfounding of End-to-End Driving Planners

Model ReleasesDGX agent

arXiv:2606.14438v3 Announce Type: replace-cross Abstract: End-to-end (E2E) autonomous-driving planners trained by imitation are prone to statistical shortcuts: they associate scene elements that merel

Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss?

Model ReleasesDGX agent

arXiv:2608.03983v1 Announce Type: cross Abstract: Optimizing compilers miss profitable transformations when their enabling semantics are absent from the analyzed program representation. We ask whether

Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design

Model ReleasesDGX agent

arXiv:2608.03501v1 Announce Type: new Abstract: AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research p

Can LLMs Test Terminal User Interfaces?

Model ReleasesDGX agent

arXiv:2608.03743v1 Announce Type: cross Abstract: Terminal User Interfaces (TUIs) combine the stateful, screen-oriented behaviour of GUIs with terminal deployment and are now common in developer tools

Can Training Logs Make Model Comparisons More Precise?

ResearchDGX agent

arXiv:2608.02705v1 Announce Type: cross Abstract: Comparing stochastically trained models requires estimating both a performance difference and its uncertainty from repeated runs. We study whether tra

CARE-Bench: Benchmarking Patient-Facing LLM Triage

Model ReleasesDGX agent

arXiv:2608.03731v1 Announce Type: new Abstract: Patient-facing medical LLMs and agents increasingly answer symptom questions before clinician contact, where the key safety question is what action the

CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Local AiDGX agent

arXiv:2608.03890v1 Announce Type: cross Abstract: A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize

CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting

AgentsDGX agent

arXiv:2608.03031v1 Announce Type: new Abstract: Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations b

ChartAnno: Evaluating MLLMs for Chart Annotation Generation

Model ReleasesDGX agent

arXiv:2608.03464v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have made significant progress in chart understanding, generation, and editing, but their ability to annotate e

Chat Debugging: An Exploratory Study of Human-AI Collaboration to Debug Analog Circuits

ResearchDGX agent

arXiv:2608.02955v1 Announce Type: cross Abstract: This research paper describes an exploratory study on the effectiveness of Chat Debugging: troubleshooting malfunctioning analog circuits on breadboar

ChiEngMixBench: Evaluating Large Language Models on Expert-Style Chinese-English Terminology Mixing

Model ReleasesDGX agent

arXiv:2601.16217v2 Announce Type: replace-cross Abstract: Large language models increasingly mediate multilingual professional communication, where useful generation requires adapting to community con

ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels

ResearchDGX agent

arXiv:2608.03507v1 Announce Type: cross Abstract: Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incomp

CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification

ResearchDGX agent

arXiv:2403.09281v3 Announce Type: cross Abstract: We propose CLIP-EBC, the first fully CLIP-based model for accurate crowd density estimation. While the CLIP model has demonstrated remarkable success

Compound and Parallel Modes of Tropical Convolutional Neural Networks

HardwareDGX agent

arXiv:2504.06881v2 Announce Type: replace-cross Abstract: Convolutional neural networks (CNNs) are foundational to many state-of-the-art computer vision systems, yet their reliance on multiplication-i

Computing Actual Causes for Neural Network Predictions under Structured Causal Inputs

ApplicationsDGX agent

arXiv:2608.03772v1 Announce Type: new Abstract: Explaining the predictions of neural networks is a central challenge in trustworthy AI. Existing explanation methods, such as those based on feature att

ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?

AgentsDGX agent

arXiv:2608.03874v1 Announce Type: new Abstract: Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these syst

Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution

SafetyDGX agent

arXiv:2608.03483v1 Announce Type: cross Abstract: Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.e., execution horizon) before replanning, turning replan

CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report Generation

Model ReleasesDGX agent

arXiv:2608.03079v1 Announce Type: cross Abstract: Breast core needle biopsy (CNB) is central to breast cancer diagnosis yet remains challenging because limited tissue sampling, lesion heterogeneity, a

Cross-Anesthetic ECoG State Decoding Fails at the Decision Threshold, Not the Representation

ResearchDGX agent

arXiv:2608.02646v1 Announce Type: cross Abstract: Decoders of anesthetic state from cortical activity fail across drug classes, most notoriously ketamine, but reported accuracy cannot say whether the

Cross-Layer Interaction under Weight-Space Ablation: A Closed-Form Attention Jacobian Bound and a Test on a Real Pretrained Model

ResearchDGX agent

arXiv:2608.03629v1 Announce Type: new Abstract: A companion paper studies when activation patching and weight-space ablation agree, inside an idealized model where a conditional computation is carried

CT-HEG: A Bidirectional, Timestamp-Attributed Event Graph for ICU In-Hospital Mortality Prediction - An Architectural Ablation Study

SafetyDGX agent

arXiv:2608.02663v1 Announce Type: cross Abstract: Accurate ICU mortality prediction requires modeling irregular clinical observations across heterogeneous entity types. Existing sequence models handle

CUADebug: Diagnosing and Repairing Computer-Use Agent Failures

Model ReleasesDGX agent

arXiv:2608.02643v1 Announce Type: cross Abstract: Computer-use agents (CUAs) operate real desktop and web interfaces through screenshots, mouse and keyboard actions, and stateful UI feedback, yet thei

Cura 1T: Specialized Model for Agentic Healthcare

Model ReleasesDGX agent

arXiv:2607.15314v2 Announce Type: replace Abstract: Healthcare AI agents handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR)

CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning

ApplicationsDGX agent

arXiv:2608.02833v1 Announce Type: cross Abstract: Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning, yet current

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning

SafetyDGX agent

arXiv:2608.03068v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as an effective method for enhancing the reasoning capabilities of large language models (LLMs). However, exis

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

Model ReleasesDGX agent

arXiv:2608.03451v1 Announce Type: new Abstract: Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured file

Decoupling Generation and Selection for Budget-Constrained Faithful Summarization

ResearchDGX agent

arXiv:2608.03655v1 Announce Type: cross Abstract: Abstractive summarization models remain vulnerable to factual inconsistency, redundancy, and weak length control. We propose a modular generation-and-

Deep Divide-and-Reduce in Symbolic Regression

ResearchDGX agent

arXiv:2608.02628v1 Announce Type: cross Abstract: Symbolic regression (SR) is the task of discovering underlying patterns from data and representing them using mathematical expressions. Current machin

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

SafetyDGX agent

arXiv:2608.01755v2 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reason

DenialRAG: Single-Document RAG Poisoning via Embedded Parametric Denial

Model ReleasesDGX agent

arXiv:2608.02678v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) systems are vulnerable to corpus poisoning: an attacker who inserts a crafted document into the retrieval corpus

Designing a Good Virtual Node: Addressable and Cardinality-Preserving Global Memory for Message Passing Architectures

ResearchDGX agent

arXiv:2608.02709v1 Announce Type: cross Abstract: Virtual nodes give message-passing neural networks a simple global communication route, but the standard node--VN--node pipeline compresses the graph

DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction

Model ReleasesDGX agent

arXiv:2608.03591v1 Announce Type: cross Abstract: Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to

DiffImaginE: Imagine to Verify Entity Types with Diffusio

TutorialsDGX agent

arXiv:2608.03025v1 Announce Type: new Abstract: Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual and visual ev

DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units

ResearchDGX agent

arXiv:2608.03127v1 Announce Type: cross Abstract: Hand motion carries the finest-grained information in human activity, yet the representations behind hand generation, understanding, and robot learnin

Distilled Roads: Generalisable Road Network Extraction Across Sensors, Resolutions, and Region

ResearchDGX agent

arXiv:2608.03407v1 Announce Type: cross Abstract: Road network segmentation from satellite imagery remains challenging due to large geographic variation in road appearance, occlusions, and domain shif

Distractor-Aware Truncation: Disentangling Context-Length Effects from Signal Loss in Long-Context LLM Benchmarks

Model ReleasesDGX agent

arXiv:2608.03297v1 Announce Type: new Abstract: A standard claim in the literature on retrieval-augmented and memory-augmented language models is that shorter context is better when the relevant infor

← Previous
1…3031323334…354
Next →