AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
5 Aug 2026

VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs

Model ReleasesDGX agent

arXiv:2608.03810v1 Announce Type: cross Abstract: Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

Model ReleasesDGX agent

arXiv:2608.03979v1 Announce Type: cross Abstract: We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense s

Vulnerabilities, Secrets and Misconfiguration in the Highest-Exposure Docker Hub Images

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.02669v1 Announce Type: cross Abstract: Docker Hub is the registry underneath most container deployments, and a flaw in a widely reused base image is inherited by every image built on it. Pr

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

Model ReleasesDGX agent

arXiv:2608.03499v1 Announce Type: new Abstract: Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served b

When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

Local AiDGX agent

arXiv:2608.03918v1 Announce Type: cross Abstract: Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence.

When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning

Local AiDGX agent

arXiv:2608.02940v1 Announce Type: new Abstract: A reproducible compression statistic can still select the wrong candidate. A dense pruning score with 0.906 split-half reliability predicted a 16.1% gai

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO

Model ReleasesDGX agent

arXiv:2608.03467v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) com- monly optimizes each correct completion as an independent learning signal. In GRPO, this comp

When Efficiency Becomes Fragility: Exploiting Dynamic Routing Vulnerabilities in Adaptive UAV Tracking

Local AiDGX agent

arXiv:2608.03902v1 Announce Type: new Abstract: Resource constraints on UAV platforms have driven a paradigm shift in aerial tracking, from pursuing performance toward balancing accuracy with efficien

When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs

Model ReleasesDGX agent

arXiv:2608.03506v1 Announce Type: new Abstract: Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples of

When Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation Boundary

Model ReleasesDGX agent

arXiv:2608.01679v2 Announce Type: replace Abstract: Persistent memory allows (self-evolving) LLM agents to adapt across tasks by consolidating heterogeneous interaction histories into reusable facts,

When Oracle Conditioning Misleads Deployment: Conditioning-Availability Bias in Echocardiographic Segmentation

SafetyDGX agent

arXiv:2608.03342v1 Announce Type: cross Abstract: Conditional segmentation models may be trained and evaluated with auxiliary signals cleaner than those available at deployment. We study this protocol

When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Coupling Diagnostic for Machine Collectives

Model ReleasesDGX agent

arXiv:2608.03722v1 Announce Type: new Abstract: Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capa

When Policies Change Probabilities: Modular Decision-Making for LLM Code Review

SafetyDGX agent

arXiv:2608.02677v1 Announce Type: cross Abstract: LLM code reviewers often estimate patch risk and make approval decisions in one prompt. A probability should depend on evidence; costs should determin

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models

Model ReleasesDGX agent

arXiv:2608.03201v1 Announce Type: new Abstract: Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs. We audit

When Search Teaches Style: Causal Internalization of Tactical Priors in AlphaZero

SafetyDGX agent

arXiv:2504.14636v3 Announce Type: replace-cross Abstract: AlphaZero is normally evaluated as one agent: a policy-value network fused with Monte Carlo tree search. That fusion hides a causal question.

When Should Graph Attention Be Sparse? Learning a Per-Edge Tsallis Index

ResearchDGX agent

arXiv:2608.02938v1 Announce Type: cross Abstract: Graph attention normalizes neighborhood scores with softmax, the maximum-entropy choice under Shannon statistics. But homophilic and heterophilic grap

When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation

SafetyDGX agent

arXiv:2608.03632v1 Announce Type: new Abstract: On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent s

Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

Model ReleasesDGX agent

arXiv:2606.15673v2 Announce Type: replace Abstract: Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and of

Where Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models

ApplicationsDGX agent

arXiv:2601.09445v2 Announce Type: replace-cross Abstract: In language models (LMs), intra-memory knowledge conflict arises when inconsistent information about the same subject is encoded within the mo

Where Reasoning Diverges: Localized Multi-Agent Debate for Multi-Hop Question Answering

Local AiDGX agent

arXiv:2608.01463v2 Announce Type: replace Abstract: Multi-agent debate commonly exchanges complete rationales even when disagreements concern only a few intermediate claims. We introduce Localized Mul

Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure

Model ReleasesDGX agent

arXiv:2608.02657v1 Announce Type: cross Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts

ZK-SR117: A Chunked Zero-Knowledge Attestation Design for Aggregated Fair-Lending Metrics, with a Control Mapping toward Full SR 11-7 Coverage

SafetyDGX agent

arXiv:2608.02664v1 Announce Type: cross Abstract: Deploying ML models in regulated decision-making (credit underwriting, fraud detection, loan approval) requires demonstrating fairness and robustness

3 Aug 2026

A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation

ResearchDGX agent

arXiv:2607.29077v1 Announce Type: new Abstract: Counterfactual explanations (CEs) enhance the interpretability of machine learning models by identifying the smallest change to an input required to obt

A Human-Centered Validation of the Explainability-Performance Coefficient

ResearchDGX agent

arXiv:2607.29614v1 Announce Type: cross Abstract: The rapid adoption of deep learning models in high-risk domains has intensified the need for trustworthy Explainable Artificial Intelligence (XAI). Ho

A robust association between LLM use and scientific productivity: Assessing stopping-time selection

Model ReleasesDGX agent

arXiv:2607.28968v1 Announce Type: cross Abstract: Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a st

A Unified Benchmark of Deep Learning Models for Multi-task 3D Brain Tumor Segmentation from Magnetic Resonance Imaging

Model ReleasesDGX agent

arXiv:2607.28858v1 Announce Type: cross Abstract: Automatic brain tumor segmentation from magnetic resonance imaging (MRI) has become a fundamental task in computer-assisted diagnosis, treatment plann

A user's guide to PINNs in geometric analysis: lessons from the asymptotic Plateau problem

TutorialsDGX agent

arXiv:2607.28733v1 Announce Type: cross Abstract: This proceedings contribution elaborates on the findings of arXiv:2605.26234v2: a joint work with Marco Usula, where we introduced a machine learning

ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency

SafetyDGX agent

arXiv:2607.29169v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies achieve strong performance in robotic manipulation but remain vulnerable to runtime disturbances that break the

ActionParty: Multi-Subject Action Binding in Generative Video Games

Model ReleasesDGX agent

arXiv:2604.02330v2 Announce Type: replace-cross Abstract: Recent advances in video diffusion have enabled the development of 'world models' capable of simulating interactive environments. However, the

Adaptive Policy Backbone via Shared Network

Model ReleasesDGX agent

arXiv:2509.22310v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has achieved impressive results across domains, yet learning an optimal policy typically requires extensive intera

Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning

SafetyDGX agent

arXiv:2607.28986v1 Announce Type: cross Abstract: Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only corpora and

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

Model ReleasesDGX agent

arXiv:2607.29626v1 Announce Type: new Abstract: As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly i

Agentic Harness for Real-World Compilers

Model ReleasesDGX agent

arXiv:2603.20075v2 Announce Type: replace-cross Abstract: Compilers are critical to modern computing, yet fixing compiler bugs is difficult. While recent large language model (LLM) advancements enable

AgenticRepair: Multi-Faceted Program Context Engineering for Agentic Vulnerability Repair

AgentsDGX agent

arXiv:2607.29422v1 Announce Type: cross Abstract: Automated vulnerability repair aims to reduce the time and effort required to patch security flaws from a vulnerability triage report. Recent agentic

Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth

SafetyDGX agent

arXiv:2607.28890v1 Announce Type: cross Abstract: Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that presumes h

AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent Profiles

AgentsDGX agent

arXiv:2602.10429v2 Announce Type: replace-cross Abstract: AIvilization v0 is a publicly deployed large-scale artificial society that couples a resource-constrained sandbox with a unified LLM-agent arc

AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction

Model ReleasesDGX agent

arXiv:2607.29549v1 Announce Type: new Abstract: Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers remains challen

An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents

Local AiDGX agent

arXiv:2607.28662v1 Announce Type: new Abstract: Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across docu

ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation

Model ReleasesDGX agent

arXiv:2607.29539v1 Announce Type: cross Abstract: Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has

AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance

AgentsDGX agent

arXiv:2512.05131v2 Announce Type: replace-cross Abstract: Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rather

AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis

AgentsDGX agent

arXiv:2607.15755v2 Announce Type: replace-cross Abstract: Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-agent

Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving

AgentsDGX agent

arXiv:2607.29031v1 Announce Type: cross Abstract: Existing autonomous-driving world models typically perform dense prediction of future videos, occupancy states, BEV representations, or agent motion.

Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search

Model ReleasesDGX agent

arXiv:2607.29055v1 Announce Type: cross Abstract: Multi-agent systems (MAS) are increasingly deployed to solve complex tasks. In case of incorrect or unsatisfactory outputs, users have to manually loc

BeatEdit: Symbolic Music Generation as Explicit Editing

ResearchDGX agent

arXiv:2607.11124v2 Announce Type: replace-cross Abstract: Music creation is fundamentally a process of revision. Yet symbolic music generation remains dominated by paradigms that produce complete sequ

Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives

Model ReleasesDGX agent

arXiv:2607.29064v1 Announce Type: cross Abstract: Police crash narratives contain information that may supplement structured crash databases, but manual review is labor-intensive and it remains unclea

Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

Model ReleasesDGX agent

arXiv:2607.28801v1 Announce Type: cross Abstract: Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial

Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions

Local AiDGX agent

arXiv:2607.28818v1 Announce Type: new Abstract: As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do

Beyond Component Testing: Validating Agentic AI Systems

SafetyDGX agent

arXiv:2607.29405v1 Announce Type: new Abstract: Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation. This behavior stretches val

Beyond Retrieval: Analytic Memory for Multimodal Agents

ResearchDGX agent

arXiv:2607.29440v1 Announce Type: new Abstract: Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across interactions.

CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents

SafetyDGX agent

arXiv:2607.29190v1 Announce Type: new Abstract: Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates gener

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

ApplicationsDGX agent

arXiv:2607.29252v1 Announce Type: cross Abstract: Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated

Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review

Model ReleasesDGX agent

arXiv:2607.28631v1 Announce Type: new Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and compari

Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery

Model ReleasesDGX agent

arXiv:2603.03322v2 Announce Type: replace-cross Abstract: Recent advancements in Large Language Model (LLM) agents have demonstrated remarkable potential in automatic knowledge discovery. However, rig

CENDRe: Concept Extraction with Natural Domain Representations

Local AiDGX agent

arXiv:2607.29621v1 Announce Type: cross Abstract: Convolutional neural networks (CNNs) are widely used for time-series classification, but their deployment in critical domains requires understanding t

CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning

Model ReleasesDGX agent

arXiv:2607.29172v1 Announce Type: cross Abstract: While robot foundation models are growing increasingly capable, the strongest models are typically trained on proprietary data and remain closed-sourc

Code Is the Body: Agent-Owned Software Bodies for Recursive Evolution and Descent

AgentsDGX agent

arXiv:2607.28691v1 Announce Type: cross Abstract: Personalized AI agents are often configurable without giving users control over the artifacts that determine their future behavior. We present OurArk,

CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

Model ReleasesDGX agent

arXiv:2607.19338v2 Announce Type: replace Abstract: Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect an

Combining Large Language Models and Symbolic Reasoning for Multi-Robot Temporal Planning through Explainable Knowledge Bases

Model ReleasesDGX agent

arXiv:2502.19135v2 Announce Type: replace Abstract: We present PLANTOR, a framework for generating and executing multi-robot task plans from natural-language task descriptions through LLM-assisted kno

Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation

SafetyDGX agent

arXiv:2604.05150v2 Announce Type: replace-cross Abstract: We study compiled AI, a paradigm in which large language models generate executable code artifacts during a compilation phase, after which wor

ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning

SafetyDGX agent

arXiv:2607.28647v1 Announce Type: cross Abstract: This paper presents ConnectED, a human-centered AI system that supports the full instructional lifecycle in Vietnamese education by linking curriculum

← Previous
1…3536373839…354
Next →