AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
Human
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
59,851 results
5 Aug 2026

Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents

Local AiDGX agent

arXiv:2608.03137v1 Announce Type: new Abstract: Large language model (LLM) agents must retain reusable information, control a bounded active context, and recover earlier evidence during long-horizon i

Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures

AgentsDGX agent

arXiv:2608.02645v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on external tools to perform multistage tasks. Existing agent frameworks typically assume that tool calls are a

Verifier-Guided Model Discovery for Physical Dynamical Systems with Pretrained Symbolic Transformers

Model ReleasesDGX agent
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.02662v1 Announce Type: cross Abstract: Reliable forecasting of nonlinear physical systems underpins scientific discovery and engineering decision-making. Yet high-fidelity simulations are p

VeriTrace: Human-Like Temporal Exploration Completes Agentic Action Space

Model ReleasesDGX agent

arXiv:2608.02878v1 Announce Type: new Abstract: Large language models have shown promise for automated Verilog RTL generation, yet state-of-the-art multi-agent systems plateau at ~95% accuracy on stan

VetScore: Risk-Weighted Fact Verification for Veterinary Long-Form QA with Citations

ResearchDGX agent

arXiv:2608.03675v1 Announce Type: new Abstract: Citation excerpts can be used to increase the reliability of generated outputs and their faithfulness to cited sources, which is especially important in

VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs

Model ReleasesDGX agent

arXiv:2608.03810v1 Announce Type: cross Abstract: Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events

VIBE: Vector Index Benchmark for Embeddings

Model ReleasesDGX agent

arXiv:2505.17810v2 Announce Type: replace Abstract: Approximate nearest neighbor (ANN) search is a performance-critical component of many machine learning pipelines, and rigorous benchmarking is essen

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent

Model ReleasesDGX agent

arXiv:2608.03979v1 Announce Type: cross Abstract: We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense s

Virtual Patients, Real Gains: Digital Twin-Based Simulated CT for Multitask Lung Nodule Analysis

ResearchDGX agent

arXiv:2502.21187v4 Announce Type: replace Abstract: AI-based lung cancer screening is constrained by scarce, annotated CT data, particularly for rare nodule presentations. We investigate whether physi

VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP

Model ReleasesDGX agent

arXiv:2608.03095v1 Announce Type: new Abstract: We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurati

VLC Fusion: Vision-Language Conditioned Sensor Fusion for Robust Object Detection

AgentsDGX agent

arXiv:2505.12715v2 Announce Type: replace Abstract: Although fusing multiple sensor modalities can enhance object detection performance, existing fusion approaches often overlook subtle variations in

Vulnerabilities, Secrets and Misconfiguration in the Highest-Exposure Docker Hub Images

Model ReleasesDGX agent

arXiv:2608.02669v1 Announce Type: cross Abstract: Docker Hub is the registry underneath most container deployments, and a flaw in a widely reused base image is inherited by every image built on it. Pr

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

Model ReleasesDGX agent

arXiv:2608.03499v1 Announce Type: new Abstract: Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served b

What is the Right Embedding Space for Contrastive Learning in REC?

SafetyDGX agent

arXiv:2505.22850v2 Announce Type: replace Abstract: Referring Expression Counting (REC) requires distinguishing visually similar objects described by fine-grained text cues. Existing methods tackle th

What Language Does and What the Evidence Supports: A Functional Role Taxonomy and Evidence Audit of Language Grounding in Embodied Agents

ResearchDGX agent

arXiv:2608.03099v1 Announce Type: new Abstract: Foundation models place language throughout embodied agents, but its presence does not show what it contributes or how well that contribution is grounde

When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills

Model ReleasesDGX agent

arXiv:2608.03700v1 Announce Type: cross Abstract: Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personaliz

When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

Local AiDGX agent

arXiv:2608.03918v1 Announce Type: cross Abstract: Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence.

When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings

Model ReleasesDGX agent

arXiv:2608.03994v1 Announce Type: new Abstract: We identify a previously overlooked failure mode of ALiBi positional encoding: its linear bias scaling underflows floating-point precision, which zeroes

When Classes Evolve: A Benchmark and Framework for Stage-Aware Class-Incremental Learning

Model ReleasesDGX agent

arXiv:2602.00573v2 Announce Type: replace-cross Abstract: Class-Incremental Learning (CIL) aims to sequentially learn new classes while mitigating catastrophic forgetting of previously learned knowled

When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning

Local AiDGX agent

arXiv:2608.02940v1 Announce Type: new Abstract: A reproducible compression statistic can still select the wrong candidate. A dense pruning score with 0.906 split-half reliability predicted a 16.1% gai

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO

Model ReleasesDGX agent

arXiv:2608.03467v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) com- monly optimizes each correct completion as an independent learning signal. In GRPO, this comp

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware

Model ReleasesDGX agent

arXiv:2608.03649v1 Announce Type: new Abstract: Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead,

When Efficiency Becomes Fragility: Exploiting Dynamic Routing Vulnerabilities in Adaptive UAV Tracking

Local AiDGX agent

arXiv:2608.03902v1 Announce Type: new Abstract: Resource constraints on UAV platforms have driven a paradigm shift in aerial tracking, from pursuing performance toward balancing accuracy with efficien

When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs

Model ReleasesDGX agent

arXiv:2608.03506v1 Announce Type: new Abstract: Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples of

When Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation Boundary

Model ReleasesDGX agent

arXiv:2608.01679v2 Announce Type: replace Abstract: Persistent memory allows (self-evolving) LLM agents to adapt across tasks by consolidating heterogeneous interaction histories into reusable facts,

When Oracle Conditioning Misleads Deployment: Conditioning-Availability Bias in Echocardiographic Segmentation

SafetyDGX agent

arXiv:2608.03342v1 Announce Type: cross Abstract: Conditional segmentation models may be trained and evaluated with auxiliary signals cleaner than those available at deployment. We study this protocol

When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Coupling Diagnostic for Machine Collectives

Model ReleasesDGX agent

arXiv:2608.03722v1 Announce Type: new Abstract: Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capa

When Policies Change Probabilities: Modular Decision-Making for LLM Code Review

SafetyDGX agent

arXiv:2608.02677v1 Announce Type: cross Abstract: LLM code reviewers often estimate patch risk and make approval decisions in one prompt. A probability should depend on evidence; costs should determin

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models

Model ReleasesDGX agent

arXiv:2608.03201v1 Announce Type: new Abstract: Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs. We audit

When Search Teaches Style: Causal Internalization of Tactical Priors in AlphaZero

SafetyDGX agent

arXiv:2504.14636v3 Announce Type: replace-cross Abstract: AlphaZero is normally evaluated as one agent: a policy-value network fused with Monte Carlo tree search. That fusion hides a causal question.

When Should Graph Attention Be Sparse? Learning a Per-Edge Tsallis Index

ResearchDGX agent

arXiv:2608.02938v1 Announce Type: cross Abstract: Graph attention normalizes neighborhood scores with softmax, the maximum-entropy choice under Shannon statistics. But homophilic and heterophilic grap

When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation

SafetyDGX agent

arXiv:2608.03632v1 Announce Type: new Abstract: On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent s

Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

Model ReleasesDGX agent

arXiv:2606.15673v2 Announce Type: replace Abstract: Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and of

Where Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models

ApplicationsDGX agent

arXiv:2601.09445v2 Announce Type: replace-cross Abstract: In language models (LMs), intra-memory knowledge conflict arises when inconsistent information about the same subject is encoded within the mo

Where Reasoning Diverges: Localized Multi-Agent Debate for Multi-Hop Question Answering

Local AiDGX agent

arXiv:2608.01463v2 Announce Type: replace Abstract: Multi-agent debate commonly exchanges complete rationales even when disagreements concern only a few intermediate claims. We introduce Localized Mul

Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't

Model ReleasesDGX agent

arXiv:2608.02829v1 Announce Type: new Abstract: Model families train every size from scratch. Can a pretrained large model be converted into a smaller sibling? We characterize the 1.4B->410M conversio

Word Recovery in Large Language Models Enables Character-Level Tokenization Robustness

ResearchDGX agent

arXiv:2603.10771v2 Announce Type: replace Abstract: Large language models (LLMs) trained with canonical tokenization exhibit surprising robustness to non-canonical inputs such as character-level token

WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament

Model ReleasesDGX agent

arXiv:2608.04008v1 Announce Type: new Abstract: Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewher

XiDepth: a Lightweight and Efficient Network for Self-supervised Monocular Depth Estimation

ApplicationsDGX agent

arXiv:2608.03666v1 Announce Type: new Abstract: Self-supervised monocular depth estimation has emerged as an appealing solution to design lightweight and effective models for deployment on computation

Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure

Model ReleasesDGX agent

arXiv:2608.02657v1 Announce Type: cross Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts

ZK-SR117: A Chunked Zero-Knowledge Attestation Design for Aggregated Fair-Lending Metrics, with a Control Mapping toward Full SR 11-7 Coverage

SafetyDGX agent

arXiv:2608.02664v1 Announce Type: cross Abstract: Deploying ML models in regulated decision-making (credit underwriting, fraud detection, loan approval) requires demonstrating fairness and robustness

4 Aug 2026

3D-CovDiffusion: 3D-Aware Diffusion Policy for Coverage Path Planning

SafetyDGX agent

arXiv:2510.03011v2 Announce Type: replace Abstract: Diffusion models have shown strong potential for robot skill learning, yet their role in coverage path planning remains underexplored. In industrial

3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering

ResearchDGX agent

arXiv:2608.01185v1 Announce Type: new Abstract: Recent 3D vision-language models (3D VLMs) construct geometry aware tokens by projecting 2D visual features into world coordinates, enabling spatial rea

A 2-Block Architecture for Real-Time EEG Gait Decoding: A Pilot Study

ResearchDGX agent

arXiv:2608.02083v1 Announce Type: new Abstract: Closed-loop lower-limb exoskeleton control via Electroencephalography (EEG) remains limited by motion artifacts, low signal-to-noise ratio, and binary g

A Benchmark Dataset for MLLM-Generated Image Detection: GPT Image2 & Nano Banana2

Model ReleasesDGX agent

arXiv:2608.01258v1 Announce Type: new Abstract: The realism of images generated by multimodal large language models (MLLMs), such as GPT Image2 and Nano Banana2, has improved rapidly in recent years.

A Comparative Analysis of MLP and Kolmogorov-Arnold Networks (KAN) for Faster-than-Nyquist (FTN) Signaling Detection

Model ReleasesDGX agent

arXiv:2608.02062v1 Announce Type: cross Abstract: Faster-than-Nyquist signaling improves spectral ef- ficiency by deliberately introducing inter-symbol interference. Classical sequence detectors such

A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models

ResearchDGX agent

arXiv:2509.22536v5 Announce Type: replace Abstract: The immense computational cost of training Large Language Models (LLMs) presents a major barrier to innovation. While FP8 training offers a promisin

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard)

SafetyDGX agent

arXiv:2608.00180v1 Announce Type: new Abstract: Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard. Training a safety guard with RL means optimizing two

A drone that learns to efficiently find non-uniformly distributed objects in agricultural fields: from simulation to the real world

AgentsDGX agent

arXiv:2505.09278v2 Announce Type: replace Abstract: Drones are promising for data collection in precision agriculture but are limited by battery capacity. Drone paths are usually planned using full co

A False Average: Chain-of-Thought Monitors Collapse Where They Are the Only Defense

AgentsDGX agent

arXiv:2608.00583v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring is meant to catch the reward hacks that look clean in the actions and betray themselves only in the reasoning. We sh

A Few Neurons Reveal When LLMs Misuse Tools: Sparse Detection and Selective Steering for Reliable Tool Use

Model ReleasesDGX agent

arXiv:2608.00218v1 Announce Type: new Abstract: Agentic LLMs exhibit three consequential tool-use failures: invalid arguments (validity), unnecessary calls (over-calling), and omitted calls when tools

A Fortran General-Purpose Transpiler: Proof of Concept

HardwareDGX agent

arXiv:2608.00130v1 Announce Type: cross Abstract: Fortran has been the cornerstone of high-performance computing for decades and remains unmatched in many domains. Yet the language faces an expertise

A Forward-Inverse Dynamic Game Framework for Enhanced Multi-Agent Trajectory Planning

SafetyDGX agent

arXiv:2608.01636v1 Announce Type: new Abstract: This paper studies feedback Nash equilibrium (FBNE) seeking for multi-agent trajectory planning in nonlinear dynamical systems with unknown agents' obje

A General-Purpose VLM Can Teach an Astronomy Foundation Model to Better Recognize Galaxy Morphology

Model ReleasesDGX agent

arXiv:2608.02300v1 Announce Type: new Abstract: Existing astronomy foundation models provide strong galaxy representations, but adapting them to new survey conditions and survey-specific morphology re

A Heuristic Perspective on Debiasing Language Models

SafetyDGX agent

arXiv:2608.00622v1 Announce Type: new Abstract: Language models (LMs) often acquire various biases during pre-training and may express them in interactions, potentially causing social harm. Existing m

A Large-Scale Multi-Dimensional Empirical Study of LLMs for Conversation Summarization

Model ReleasesDGX agent

arXiv:2606.15974v2 Announce Type: replace Abstract: Despite the significant advancement of LLMs in conversation summarization, their evaluation remains limited by insufficient scenarios, input lengths

A Multi-Objective AutoML-based Efficient Intrusion Detection System for EV Charging Networks

ResearchDGX agent

arXiv:2608.02274v1 Announce Type: cross Abstract: Electric Vehicle Charging Systems (EVCSs) are increasingly connected with Internet of Things (IoT) devices, which improves charging intelligence but a

A Physics-Chemistry-Informed Neural Network (PCINN) for Real-Time Spatial-ALD Coverage Prediction and Reliable Kinetics Inversion

ResearchDGX agent

arXiv:2608.00212v1 Announce Type: new Abstract: Spatial atomic layer deposition (SALD) is a leading atmospheric-pressure, high-throughput route to industrial ALD, but design and control are limited by

A reproducible and extensible framework for benchmarking competing risks survival models

ResearchDGX agent

arXiv:2608.00271v1 Announce Type: cross Abstract: A wide range of statistical and machine learning methods have been proposed for survival analysis with competing risks, where the occurrence of one ev

A Robotic System for Automated Manufacturing of Dielectric Elastomer Actuators

ApplicationsDGX agent

arXiv:2608.00369v1 Announce Type: new Abstract: This letter presents an automated robotic manufacturing system for soft capacitors which operate as actuators and sensors. Emphasis is placed on the two

← Previous
1…8788899091…998
Next →