AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
31 Jul 2026

Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents

AgentsDGX agent

arXiv:2607.15715v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly used for complex information-extraction tasks, yet it remains unclear whether agentic components

Belief-Guided Decision Making with Uncertainty Gating in the Game of Go

SafetyDGX agent

arXiv:2607.26946v1 Announce Type: new Abstract: Recent advancements in Computer Go, driven by AlphaZero and MuZero, rely heavily on Monte Carlo Tree Search (MCTS) to correct the errors of the neural n

BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models

SafetyDGX agent

arXiv:2512.00807v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) inherit significant social biases from their training data, notably in gender representation. Current fairness interve


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses

SafetyDGX agent

arXiv:2607.26639v1 Announce Type: cross Abstract: A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% def

Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning

TutorialsDGX agent

arXiv:2508.16129v3 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities with reinforcement learning paradigm. Although se

Calibrate Globally, Measure Everywhere: Scaling LLM-Based Prevalence Measurement Across A/B Experiments

ApplicationsDGX agent

arXiv:2602.16111v2 Announce Type: replace-cross Abstract: Online media platforms track the share of impressions associated with content attributes, or prevalence, to evaluate trade-offs and set guardr

CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games

AgentsDGX agent

arXiv:2607.26393v1 Announce Type: new Abstract: Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents. These games require complex social skills such as reasoni

ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science

Model ReleasesDGX agent

arXiv:2607.26155v1 Announce Type: new Abstract: Clinical data-science agents must transform heterogeneous longitudinal records into auditable analyses, yet existing benchmarks largely isolate medical

Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction

Model ReleasesDGX agent

arXiv:2607.26385v1 Announce Type: cross Abstract: Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered 'no' by a conspiracy

Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech

ResearchDGX agent

arXiv:2607.26236v1 Announce Type: cross Abstract: AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue. Yet, existing

Decision-oriented joint optimization of evidence fusion based on event-conditioned credibility

ResearchDGX agent

arXiv:2504.04128v3 Announce Type: replace Abstract: In decision-level fusion tasks involving heterogeneous sources with unequal precision and potential anomalies, evidence deviating from the majority

Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM

AgentsDGX agent

arXiv:2607.26773v1 Announce Type: new Abstract: Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but gr

(EC)2: Event-Centric Explainability for Cybersecurity Through Multi-Agent LLM Investigations

AgentsDGX agent

arXiv:2607.26201v1 Announce Type: cross Abstract: Security operations centers rely on anomaly detection systems to flag suspicious events. Feature-level explanations for anomaly detectors offer limite

Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models

SafetyDGX agent

arXiv:2607.26588v1 Announce Type: new Abstract: The rapid development of large language models (LLMs) has renewed interest in agent-based modeling (ABM). However, current LLM-based ABM research faces

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

AgentsDGX agent

arXiv:2607.26120v1 Announce Type: new Abstract: Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric i

Evidence-Ledger Adjudication for Claim-Evidence Traceability

Model ReleasesDGX agent

arXiv:2607.26512v1 Announce Type: new Abstract: AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them. We study evidence-ledger adjudication: a

Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

Model ReleasesDGX agent

arXiv:2607.26367v1 Announce Type: new Abstract: An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model. We study this skill as an AI-agent t

Fantastic Adaptive Taxonomies and How to Use Them

Model ReleasesDGX agent

arXiv:2607.16387v2 Announce Type: replace-cross Abstract: An agent system's execution traces record how it fails, and procedures that improve such a system without changing model weights (trajectory s

FARI: Robust One-Step Inversion for Watermarking in Diffusion Models

HardwareDGX agent

arXiv:2607.26723v1 Announce Type: cross Abstract: Inversion-based watermarking is a promising approach to authenticate diffusion-generated images, yet practical use is bottlenecked by inversion that i

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

Model ReleasesDGX agent

arXiv:2607.26611v1 Announce Type: new Abstract: AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in us

FinCacheServe: Dependency-Consistent Answer Reuse for Cost-Efficient RAG Serving over Mutable Enterprise Documents

HardwareDGX agent

arXiv:2607.26076v1 Announce Type: cross Abstract: Retrieval-augmented generation services over mutable enterprise documents repeatedly execute semantically equivalent analysis requests. Answer reuse c

FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing

Model ReleasesDGX agent

arXiv:2508.02092v3 Announce Type: replace-cross Abstract: Large language models represent significant investments in computation, data, and engineering expertise, making them extraordinarily valuable

Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents

SafetyDGX agent

arXiv:2606.13097v2 Announce Type: replace-cross Abstract: Code-writing large language models (CodeLLMs) generate executable code policies for embodied agents by translating natural language goals and

GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure

HardwareDGX agent

arXiv:2607.26181v1 Announce Type: new Abstract: Functional verification dominates integrated circuit (IC) front-end engineering effort, and a single missed bug that escapes to silicon can trigger a co

Graph Is the Verifier: Agentic Reinforcement Learning for Interprocedural Vulnerability Detection

SafetyDGX agent

arXiv:2607.26656v1 Announce Type: cross Abstract: Real-world vulnerabilities often span multiple functions, yet most learning-based detectors classify each function in isolation: on a sample of real C

Guarding Organizations Against Malware Risk: A Novel Graph-Based Malware Detection Method

TutorialsDGX agent

arXiv:2607.26634v1 Announce Type: cross Abstract: Organizational digitalization expands cybersecurity risks, making cybersecurity an increasingly important research area in Information Systems (IS). A

GuidedRAG: Semantic Steering of Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2607.26071v1 Announce Type: cross Abstract: In this work, we propose GuidedRAG, a novel extension to traditional Retrieval-Augmented Generation (RAG) that introduces a dedicated selection stage

GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning

Model ReleasesDGX agent

arXiv:2607.26160v1 Announce Type: new Abstract: Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather t

How does downsampling affect needle electromyography signals? A generalisable workflow for understanding downsampling effects on high-frequency time series

ResearchDGX agent

arXiv:2601.10191v2 Announce Type: replace Abstract: Automated analysis of needle electromyography (nEMG) signals is emerging as a tool to support the detection of neuromuscular diseases (NMDs), yet th

Human diversity fuels collective creativity that large language models cannot simulate or sustain

ResearchDGX agent

arXiv:2607.26899v1 Announce Type: cross Abstract: Diverse human groups produce diverse ideas, the raw material of innovation. Generative AI challenges this engine twice over: everyday AI assistance ma

IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations

Model ReleasesDGX agent

arXiv:2607.26075v1 Announce Type: cross Abstract: We present IDP AutoOpt, an autonomous LLM agent that discovers high-performing configurations for intelligent document processing (IDP) pipelines. Tun

IFCMemoryBench: Evaluating Long-Term Memory of LLM-Based Agents in BIM Information Retrieval

Model ReleasesDGX agent

arXiv:2607.26072v1 Announce Type: cross Abstract: Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-domain or pers

Improving Item Discoverability in e-Commerce Search via Related Intent Generation

ResearchDGX agent

arXiv:2607.27172v1 Announce Type: cross Abstract: Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. In e-commerce market

MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair

Model ReleasesDGX agent

arXiv:2607.27080v1 Announce Type: cross Abstract: Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist. A malicious instr

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

Model ReleasesDGX agent

arXiv:2607.27109v2 Announce Type: cross Abstract: With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-gra

Model-Driven Requirements Configuration with Three-Valued Uncertainty Scoring

AgentsDGX agent

arXiv:2607.26220v1 Announce Type: cross Abstract: Context: Large Language Models (LLMs) offer natural-language flexibility for automated requirements elicitation but frequently generate structurally i

MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation

SafetyDGX agent

arXiv:2607.26698v1 Announce Type: cross Abstract: Cover song generation (CSG) should preserve the melodic and linguistic content of a reference song while recreating the remaining musical components.

Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges

AgentsDGX agent

arXiv:2607.26212v1 Announce Type: cross Abstract: Multi-Agent Debate (MAD) is a promising paradigm for improving the accuracy and robustness of Large Language Model (LLM)-based agentic systems. It ena

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

Model ReleasesDGX agent

arXiv:2607.26465v1 Announce Type: new Abstract: Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform seq

One Run Is Not an Idea: The Implementation Lottery in Automated Research

AgentsDGX agent

arXiv:2607.26587v1 Announce Type: cross Abstract: Automated research systems use experimental scores both to deliver artifacts and to decide which ideas to retain, transfer, and pursue. Yet one run sc

Optimizing Sensor Placement for Hydrogen Leak Detection in Enclosed Infrastructure: A Comparative Study Using CFD-informed Genetic Algorithm and DeepSets Neural Surrogate

SafetyDGX agent

arXiv:2607.26078v1 Announce Type: cross Abstract: Hydrogen infrastructure in enclosed environments, such as parking facilities for fuel cell vehicles, presents significant safety challenges due to hyd

Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork

AgentsDGX agent

arXiv:2607.27177v1 Announce Type: new Abstract: Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume t

Pramana: A Composable, Domain-Specific Backend for Empirical Networking Research

AgentsDGX agent

arXiv:2607.26352v1 Announce Type: cross Abstract: Networking research advances by turning hypotheses into empirical evidence, so accelerating it means reducing the lag between ideation (synthesizing a

Predict before you train: Scaling Laws for particle physics foundation models

ResearchDGX agent

arXiv:2607.23377v1 Announce Type: cross Abstract: The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architecture cannot be

Property-driven Causal Abstractions for Markov Decision Processes

ResearchDGX agent

arXiv:2607.26787v1 Announce Type: new Abstract: Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state variables and th

PUDA: An AI-Native Hardware Harness for Self-Driving Laboratories

AgentsDGX agent

arXiv:2607.26464v1 Announce Type: cross Abstract: Physical Unified Device Architecture (PUDA) is an AI-native hardware harness for self-driving laboratories (SDLs). Rather than building a human-center

Pushing the Frontier on Approximate EFX Allocations

ResearchDGX agent

arXiv:2406.12413v3 Announce Type: replace-cross Abstract: We study the problem of allocating a set of indivisible goods to a set of agents with additive valuation functions, aiming to achieve approxim

SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation

Model ReleasesDGX agent

arXiv:2607.26313v1 Announce Type: cross Abstract: Agentic systems act, so a defect in the evidence they retrieve becomes a wrong action with a currency cost. The most dangerous enterprise defects are

ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform

SafetyDGX agent

arXiv:2607.26566v1 Announce Type: cross Abstract: Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke them inter

Shared Symbolic Backbones for Physically Consistent Multi-Output Symbolic Regression

Model ReleasesDGX agent

arXiv:2607.26528v1 Announce Type: cross Abstract: Symbolic regression provides analytical expressions, but it is usually applied one output at a time. This is limiting in process systems, where state

SimpleWikiSearch: A Clean Offline Wikipedia Environment for Agentic Search

SafetyDGX agent

arXiv:2607.26070v1 Announce Type: cross Abstract: Large language model (LLM)-based agentic search systems are often evaluated as if the underlying LLM were the only component that matters, yet their m

SkillSight: Calibrating Generic Content Bias for Skill Retrieval

SafetyDGX agent

arXiv:2607.18785v2 Announce Type: replace Abstract: As large language model agents gain access to increasingly large skill libraries, retrieving the right skill becomes critical to reliable capability

State-Dependent Safety Failures in Multi-Turn Language Model Interaction

SafetyDGX agent

arXiv:2603.15684v2 Announce Type: replace-cross Abstract: Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Although

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

Model ReleasesDGX agent

arXiv:2607.26314v1 Announce Type: cross Abstract: Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisti

The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science

AgentsDGX agent

arXiv:2607.26064v1 Announce Type: cross Abstract: AI systems are becoming autonomous research agents that generate hypotheses, design experiments, and produce discoveries at scales beyond human oversi

The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty

SafetyDGX agent

arXiv:2607.26067v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for estimating item difficulty in educational assessment. However, it remains unclear whether such

The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem

SafetyDGX agent

arXiv:2607.26068v1 Announce Type: cross Abstract: Existing AI governance frameworks, including the EU AI Act and NIST AI RMF, address safety, transparency, and accountability but do not operationalize

The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making

TutorialsDGX agent

arXiv:2607.27179v1 Announce Type: cross Abstract: Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communication among

TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning

Model ReleasesDGX agent

arXiv:2607.26307v1 Announce Type: new Abstract: Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code through benc

UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks

Model ReleasesDGX agent

arXiv:2607.26724v1 Announce Type: new Abstract: Large language model (LLM) agents have been widely applied in automating data science tasks. However, existing methods typically rely on a limited set o

← Previous
1…3940414243…354
Next →