AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
16 Jul 2026

Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows

SafetyDGX agent

arXiv:2607.13078v1 Announce Type: cross Abstract: LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public literature

Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents

AgentsDGX agent

arXiv:2607.13157v1 Announce Type: new Abstract: Agent memory is a systems problem for long-horizon agents. Practical deployments require retention of task state across extended conversations, recovery

OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets

ResearchDGX agent

arXiv:2607.13037v1 Announce Type: new Abstract: When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require a forget set, yet no tool can locate which


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

OvisOCR2 Technical Report

Model ReleasesDGX agent

arXiv:2607.13639v1 Announce Type: cross Abstract: We introduce OvisOCR2, a 0.8B document parsing model. OvisOCR2 is designed as an end-to-end parser: given a document page image, it generates a Markdo

Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot Ceilings

Model ReleasesDGX agent

arXiv:2607.13918v1 Announce Type: cross Abstract: Serial verification gates are a core reliability primitive in LLM harnesses: a candidate answer is returned only if k verifier calls all accept it. Un

PC-Diffuser: Path-Consistent Capsule CBF Safety Filtering for Diffusion-Based Trajectory Planner

SafetyDGX agent

arXiv:2603.10330v2 Announce Type: replace-cross Abstract: Autonomous driving in complex traffic requires planners that generalize beyond hand-crafted rules, motivating data-driven approaches that lear

PersGuard: Preventing Malicious Personalization in Text-to-Image Diffusion Models via Model Backdoors

ResearchDGX agent

arXiv:2502.16167v2 Announce Type: replace-cross Abstract: Diffusion models (DMs) have advanced text-to-image (T2I) synthesis, yet their personalization capabilities raise serious privacy and copyright

Policy of Thoughts: Scaling Test-Time Training for LLM Reasoning via Online Policy Evolution

Model ReleasesDGX agent

arXiv:2601.20379v2 Announce Type: replace Abstract: Large language models (LLMs) struggle with complex, long-horizon reasoning due to instability caused by their frozen policy assumption. Current test

Post-Disaster Affected Area Segmentation with a Vision Transformer (ViT)-based EVAP Model using Sentinel-2 and Formosat-5 Imagery

ResearchDGX agent

arXiv:2507.16849v3 Announce Type: replace-cross Abstract: We propose a vision transformer (ViT)-based deep learning framework to refine disaster-affected area segmentation from remote sensing imagery,

Price of Fairness in Bandits: A Tight Minimax Characterization

SafetyDGX agent

arXiv:2607.13402v1 Announce Type: cross Abstract: In bandit problems, standard regret-minimizing algorithms treat exploration as an amortized cost, which can expose early participants to unfair ex-ant

Privacy Preserving Recommender Systems Balancing Personalization with Privacy

SafetyDGX agent

arXiv:2607.13328v1 Announce Type: cross Abstract: Personalized recommendation systems are central to modern e-commerce and retail platforms, but they typically rely on centralized storage of detailed

Probabilistic Extension of Neuro-Symbolic AGI Robots based on Belnap's Typed Intensional FOL

ResearchDGX agent

arXiv:2607.13073v1 Announce Type: new Abstract: Neuro-symbolic AI based on IFOL_B is a way to combine neural learning and symbolic reasoning to overcome limitations of purely neural systems (like lack

Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities

SafetyDGX agent

arXiv:2607.13596v1 Announce Type: cross Abstract: When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by acknowledgi

RADAR: Closed-Loop Robotic Data Generation via Semantic Planning and Autonomous Causal Environment Reset

SafetyDGX agent

arXiv:2603.11811v2 Announce Type: replace-cross Abstract: The acquisition of large-scale physical interaction data, a critical prerequisite for modern robot learning, is severely bottlenecked by the p

RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar

Model ReleasesDGX agent

arXiv:2607.13189v1 Announce Type: cross Abstract: We present RAGthoven, our system for SemEval-2026 Task 1 (MWAHAHA), Subtask A (multilingual constrained humor generation in English, Spanish, and Chin

Reassessing Muon for Matrix Factorization

ResearchDGX agent

arXiv:2607.13246v1 Announce Type: cross Abstract: Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalizatio

Representation-Based Exploration for Language Models: From Test-Time to Post-Training

Model ReleasesDGX agent

arXiv:2510.11686v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) promises to expand the capabilities of language models, but it is unclear if current RL techniques promote the dis

Rethinking Multimodal Fusion for Time Series: Text Modalities Need Constrained Fusion

ResearchDGX agent

arXiv:2603.22372v3 Announce Type: replace-cross Abstract: Recent advances in multimodal learning have motivated the integration of auxiliary modalities such as text or vision into time series (TS) for

Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation

AgentsDGX agent

arXiv:2607.14006v1 Announce Type: cross Abstract: Penetration testing traditionally evaluates whether adversaries can exploit weaknesses in software, infrastructure, configurations, or operational con

Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants

Model ReleasesDGX agent

arXiv:2607.13039v1 Announce Type: cross Abstract: Safety evaluations for dual-use biology assistants often measure base-model capability, refusal behavior, or jailbreak success. These metrics miss a d

SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing

SafetyDGX agent

arXiv:2607.13594v1 Announce Type: new Abstract: LLM agents act on real-world environments through tool calls, and a single misjudged action can cause irreversible harm. The standard safeguard is a gua

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding

Model ReleasesDGX agent

arXiv:2607.13421v1 Announce Type: cross Abstract: Spatio-Temporal Video Grounding (STVG) aims to retrieve the visual trajectory of a specific object from a video stream as described by a natural langu

Self-Improvements in Modern Agentic Systems: A Survey

AgentsDGX agent

arXiv:2607.13104v1 Announce Type: new Abstract: Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or adaptation, fro

Self-Improving AI Coding Agents Through Accumulated Behavioral Rules: A Closed-Loop Framework

AgentsDGX agent

arXiv:2607.13091v1 Announce Type: cross Abstract: LLM-based coding agents repeat the same classes of mistakes across sessions because they lack a mechanism to retain corrections from human review feed

SemaDiff: Identifying Semantic-Changing Commits with Generated Code and Tests

Local AiDGX agent

arXiv:2607.13111v1 Announce Type: cross Abstract: Distinguishing semantic-preserving commits from changing ones remains an open challenge in software repository mining. While existing approaches detec

Semantic Anchoring for Robotic Action Representations

ApplicationsDGX agent

arXiv:2607.13597v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limited robot dem

Set-shifting Behavioral Test for Harnessed Agents

Model ReleasesDGX agent

arXiv:2607.13396v1 Announce Type: new Abstract: What happens to an LLM agent's tool choice when the reliable tool silently changes within an ongoing session? We borrow set-shifting from cognitive psyc

ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

SafetyDGX agent

arXiv:2607.13124v1 Announce Type: cross Abstract: Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compre

SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification

Model ReleasesDGX agent

arXiv:2607.13081v1 Announce Type: cross Abstract: We present nsfaguard, a guardrail framework for securing agentic AI systems against operational threats, such as prompt injection, sensitive informati

Social Simulations: from Agent-Based Modeling to Digital Twins

AgentsDGX agent

arXiv:2607.13693v1 Announce Type: cross Abstract: This book chapter covers the evolution of social simulation from classical agent-based models, in which agents interact according to explicitly define

Spectral-Informed Neural Networks Outperform Spectral Methods in High-dimensional PDEs

ResearchDGX agent

arXiv:2607.13566v1 Announce Type: cross Abstract: For low-dimensional problems (dleq3), spectral methods can achieve exceptionally high accuracy. For middle-dimensional problems (4 leq d lesssim 10),

SPINE: Bridging the Cyber-Physical Gap with Agentic AI

Model ReleasesDGX agent

arXiv:2607.13049v1 Announce Type: new Abstract: Foundation models have given robots a sophisticated brain for complex decision-making, yet deploying that intelligence into a physical platform still de

SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy

SafetyDGX agent

arXiv:2607.13175v1 Announce Type: cross Abstract: Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrop

STKAN: Kolmogorov-Arnold Networks for Spatio-Temporal Forecasting

ApplicationsDGX agent

arXiv:2607.13108v1 Announce Type: cross Abstract: Real-world traffic data exhibit heterogeneous spatial correlations and nonlinear temporal dynamics, posing substantial challenges for accurate spatio-

STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle

Model ReleasesDGX agent

arXiv:2607.13618v1 Announce Type: new Abstract: LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such tasks the fin

Symbiosis-Inspired Knowledge Distillation for Incremental Object Detection

Model ReleasesDGX agent

arXiv:2607.13452v1 Announce Type: cross Abstract: Incremental object detection (IOD) aims to extend detectors to new categories while retaining previously acquired knowledge. Existing methods often ad

Tabular Foundation Models for Discrete Choice Estimation

ResearchDGX agent

arXiv:2607.13314v1 Announce Type: cross Abstract: Tabular foundation models (TFMs) generate predictions on structured data via in-context learning, without task-specific estimation. We ask whether TFM

The Cafe in Amsterdam: When the Incumbent Becomes the Oracle

SafetyDGX agent

arXiv:2607.13393v1 Announce Type: cross Abstract: A field can reformulate its computations freely exactly where its demand is stated independently of any incumbent implementation, and finds itself una

The Dynamic Verifiable Multi-Agent Human Agentic Loyalty Loop (DVM-HALL) Model and the Net Human-Agent Score (NHAS) in Autonomous Commerce

Model ReleasesDGX agent

arXiv:2607.13998v1 Announce Type: cross Abstract: The rapid proliferation of Agentic Artificial Intelligence fundamentally disrupts traditional customer loyalty paradigms. As AI evolves from passive r

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators

ResearchDGX agent

arXiv:2607.13075v1 Announce Type: cross Abstract: Context can change whether a request is harmful without changing its topic or surface form. We ask whether residual-stream probes distinguish harmful

The Hitchhiker's Guide to Monoculture

TutorialsDGX agent

arXiv:2607.13077v1 Announce Type: cross Abstract: Large language models (LLMs) often produce homogeneous outputs, raising concerns that AI coding assistants may lead to convergence in the software art

The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI

Model ReleasesDGX agent

arXiv:2607.13044v1 Announce Type: cross Abstract: The European Patent Office (EPO) reported record filings in 2025, and the 2026 EPO Guidelines hold applicants strictly responsible for LLM-assisted co

The Refusal Residue: When Probes Catch Alignment Faking and When They Don't

Model ReleasesDGX agent

arXiv:2607.13346v1 Announce Type: cross Abstract: Alignment faking is dangerous because a model can appear compliant under monitoring while preserving behavior it would reveal when unmonitored. When n

The SIGReg Objective as Variational Free Energy: A Theoretical Active-Inference Account of JEPA World Models

SafetyDGX agent

arXiv:2607.13612v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) are the dominant design for latent world models, yet they are usually justified by empirical performa

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

ResearchDGX agent

arXiv:2607.08803v2 Announce Type: replace-cross Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine underst

Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases

ResearchDGX agent

arXiv:2607.13292v1 Announce Type: new Abstract: Autoformalization translates informal natural language into formal, machine-verifiable languages. While most work focuses on individual statements, real

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems

AgentsDGX agent

arXiv:2604.02668v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion. While prior

Traffic-Aware Randomized Smoothing for LLM-Based Network Intrusion Detection

Model ReleasesDGX agent

arXiv:2607.13801v1 Announce Type: cross Abstract: Large language model (LLM)-based intrusion detection systems (IDS) are increasingly studied for security monitoring, yet their robustness against feas

Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth

Model ReleasesDGX agent

arXiv:2607.14018v1 Announce Type: cross Abstract: We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initializ

TSSM: Triaxial State Space Model for Global Station Weather Forecasting with Temporal-Variable-Historical Modeling

Local AiDGX agent

arXiv:2607.13101v1 Announce Type: cross Abstract: Global Station Weather Forecasting (GSWF) is pivotal for localized and extreme weather prediction over key regions. Despite efforts to exploit look-ba

UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following

Model ReleasesDGX agent

arXiv:2607.13621v1 Announce Type: new Abstract: Language-guided human following is an important capability for embodied agents, but existing benchmarks typically assume that the target person is visib

Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems

SafetyDGX agent

arXiv:2607.13048v1 Announce Type: cross Abstract: Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) that provide rich semantic understanding at

Uniform Approximation of Functions with Asymmetric Growth and Decay by Deep Weighted Polynomials

ResearchDGX agent

arXiv:2506.21306v2 Announce Type: replace-cross Abstract: Functions that grow without bound on one side of the real line and decay to zero on the other cannot be approximated uniformly by ordinary pol

Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild

AgentsDGX agent

arXiv:2607.13881v1 Announce Type: cross Abstract: Human-object interaction detection (HOID) has traditionally been formulated as a supervised detection problem over predefined interaction categories.

UTS at ELOQUENT 2026 Voight-Kampff: structural shifts in AI writing bypass state-of-the-art detectors

ResearchDGX agent

arXiv:2607.13565v1 Announce Type: cross Abstract: We investigate which language model evasion attacks survive state-of-the-art adversarial fine-tuning, developing strategies that sweep the top 5 posit

Verifier-Guided Twelve-Tone Composition: A Generate-Verify-Repair Harness for Symbolic Music Generation

ApplicationsDGX agent

arXiv:2607.11334v2 Announce Type: replace Abstract: Large language models can produce superficially legal twelve-tone scores that collapse into degenerate textures. We introduce a neuro-symbolic harne

Verifying formulas for interventional distributions

ResearchDGX agent

arXiv:2607.13883v1 Announce Type: cross Abstract: We formalize verification in causal graphical models: deciding whether a given observational formula identifies a target interventional distribution.

WaterMoE: Expert-Routing-based Watermarking for High Fidelity and Efficiency

Model ReleasesDGX agent

arXiv:2607.13099v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable success but raise growing concerns about content provenance and misuse, motivating the need for

What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors

AgentsDGX agent

arXiv:2607.13162v1 Announce Type: cross Abstract: What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by

When Agents Disagree With Themselves: Behavioral Consistency as an Uncertainty Signal for LLM Agents

Model ReleasesDGX agent

arXiv:2602.11619v2 Announce Type: replace Abstract: Running the same LLM agent on identical inputs yields 2.3-4.2 distinct action sequences per 10 runs; this behavioral variance constitutes a training

← Previous
1…6869707172…358
Next →