AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
9 Jun 2026

TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2510.27544v2 Announce Type: replace Abstract: Temporal reasoning involves understanding how systems evolve over time through input-driven state transitions. A key aspect is temporal causal reaso

Test-Time Adaptive Composition for Machine Learning as a Service (MLaaS) in IoT Environments

ResearchDGX agent

arXiv:2606.07685v1 Announce Type: cross Abstract: The dynamic nature of Internet of Things (IoT) environments affects the long-term effectiveness of Machine Learning as a Service (MLaaS) compositions.

Testing the Black Box: Structural Barriers to Independent Evaluation of Consumer-Facing Health LLMs

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.08483v1 Announce Type: new Abstract: Background: Consumer-facing large language models are now a common source of health information, and they interpret and personalize responses rather tha

The ACUTE Protocol: Operationalizing Language Model Activations for Better Calibration, Utility, and Trust

SafetyDGX agent

arXiv:2606.07822v1 Announce Type: cross Abstract: As language models improve and become increasingly deployed to solve a variety of tasks, trustworthiness becomes essential. Calibration is a good prox

The AI Epistemic Deference Index: A Continuous Measure of Sycophancy

Model ReleasesDGX agent

arXiv:2606.07897v1 Announce Type: new Abstract: Current AI models frequently exhibit epistemic sycophancy, endorsing claims to agree with a user. Existing evaluations typically measure this either by

The CIFAR Synthetic Evidence Corpus for Detecting AI-Generated Evidence

TutorialsDGX agent

arXiv:2606.07916v1 Announce Type: new Abstract: The growing ability of generative models to produce realistic documents poses a direct challenge to evidentiary workflows in the justice system and the

The Confidence Trap: Calibration Attacks for Graph Neural Networks

SafetyDGX agent

arXiv:2606.08467v1 Announce Type: cross Abstract: While confidence calibration is essential for trustworthy decision-making in safety-critical applications, the robustness of calibrated GNNs to advers

The Cross-Architecture Substrate: A Domain-Transcendent, Calibration-Surviving Geometric Invariant of Modern Vision Encoders

SafetyDGX agent

arXiv:2606.07882v1 Announce Type: cross Abstract: Different vision neural networks -- trained to classify, contrast, reconstruct, or match images to text -- should have correspondingly different inter

The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models

SafetyDGX agent

arXiv:2601.15165v4 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) break the rigid left-to-right constraint of traditional LLMs, enabling token generation in arbitrary o

The Governance of Human-LLM Interaction: Safety Gating, Civility Steering, and Affective Default Lock-In

SafetyDGX agent

arXiv:2606.08172v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly mediate high-stakes interactions in finance, medicine, and mental-health support, yet users have limited con

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.07861v1 Announce Type: cross Abstract: Recent vision-language models (VLMs) excel at multimodal understanding and reasoning, yet their fine-grained visual perception remains underexplored.

The Montparnasse Algorithm for RNA Design

Model ReleasesDGX agent

arXiv:2606.07562v1 Announce Type: cross Abstract: RNA design consists of discovering a nucleotide sequence that optimizes predefined criteria, such as secondary structure. It is useful for synthetic b

The Token Not Taken: Sampling, State, and the Variability of AI Agent Outputs

AgentsDGX agent

arXiv:2606.08998v1 Announce Type: new Abstract: Agentic AI systems can behave differently across runs: the same request may produce a different plan, a different tool call, a different code edit, or a

The Topological Dual of a Dataset: A Logic-to-Topology Encoding for AlphaGeometry-Style Data

ResearchDGX agent

arXiv:2604.18050v2 Announce Type: replace Abstract: AlphaGeometry represents a milestone in neuro-symbolic reasoning, yet its architecture faces a log-linear scaling bottleneck within its symbolic ded

TheoremBench: Evaluating LLMs on Theorem Proving in Formal Mathematics

Model ReleasesDGX agent

arXiv:2606.09450v1 Announce Type: new Abstract: LLMs have recently achieved strong results on formal proving benchmarks. However, existing evaluations remain heavily concentrated on competition-style

Think Before You Act: Intention-Guided Reasoning for LLM-Based Location Prediction

SafetyDGX agent

arXiv:2606.08122v1 Announce Type: new Abstract: Predicting a user's next Point-of-Interest (POI) based on their historical check-in records is a fundamental task in location-based services. While rece

Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning

Model ReleasesDGX agent

arXiv:2601.04805v2 Announce Type: replace Abstract: Large reasoning models (LRMs) have attracted much attention due to their exceptional performance. However, their performance mainly stems from think

TianJi-Environ: An Autonomous AI Scientist for Atmospheric Environmental Research

AgentsDGX agent

arXiv:2606.07697v1 Announce Type: cross Abstract: As atmospheric environmental prediction continues to improve, interpretable validation of pollution mechanisms and feedback processes has become a mai

TimpaTeks: Automatic In-place Text Sequence Modification via Diffusion Language Model Steering

ResearchDGX agent

arXiv:2606.08408v1 Announce Type: cross Abstract: We extend activation steering to diffusion language models (DLMs) and study a novel problem that arose due to the inference mechanism of DLMs: Modifyi

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech

ResearchDGX agent

arXiv:2606.09019v1 Announce Type: cross Abstract: Codec-based autoregressive (AR) speech language models have achieved strong text-to-speech (TTS) quality by modeling speech as sequences of discrete a

To Nuke or Not to Nuke: LLMs' (Missing) Ethical Reasoning and Actions in a High-Stakes Decision-Making Simulation

AgentsDGX agent

arXiv:2606.08310v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as long-horizon agents with decision-making capacities. While LLMs can show ethical competence on

Topological Neural Operators

ResearchDGX agent

arXiv:2606.09806v1 Announce Type: cross Abstract: We introduce Topological Neural Operators (TNOs), a principled framework for operator learning on cell complexes that lifts neural operators (NOs) fro

Toward autocorrection of chemical process flowsheets using large language models

SafetyDGX agent

arXiv:2312.02873v2 Announce Type: replace-cross Abstract: The process engineering domain widely uses Process Flow Diagrams (PFDs) and Process and Instrumentation Diagrams (P&IDs) to represent process

Towards Long-Horizon Vessel Trajectory and Destination Forecasting with Reasoning Large Language Models

Model ReleasesDGX agent

arXiv:2606.08633v1 Announce Type: new Abstract: Long-horizon maritime trajectory prediction is important for shipping management, logistics planning, and maritime risk analysis, yet month-level foreca

TQA-Bench: Evaluating LLMs for Multi-Table Question Answering

Model ReleasesDGX agent

arXiv:2411.19504v2 Announce Type: replace Abstract: The advance of large language models (LLMs) has unlocked great opportunities in complex multi-modal data management tasks, particularly in question

TRACER: Token ReAssignment for Concept ERasure in Generative Recommendation

SafetyDGX agent

arXiv:2606.07688v1 Announce Type: cross Abstract: Generative recommendation formulates next-item prediction as autoregressive generation over semantic ID (SID) sequences derived from users' historical

Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning Model

TutorialsDGX agent

arXiv:2603.25184v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become essential for post-training large language models (LLMs) in reasoning tasks. While scaling rollouts can

Training-Free Intelligibility-Guided Observation Addition for Noisy ASR

ResearchDGX agent

arXiv:2602.20967v2 Announce Type: replace-cross Abstract: Automatic speech recognition (ASR) degrades severely in noisy environments. Although speech enhancement (SE) front-ends effectively suppress b

Training-Inference Kernel Contracts: Bounding Divergence in Post-Training and Deployment

SafetyDGX agent

arXiv:2606.07581v1 Announce Type: cross Abstract: A modern post-training pipeline often writes one symbol for its policy, pi_theta, while evaluating it through two different programs: a training kerne

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning

SafetyDGX agent

arXiv:2606.07631v1 Announce Type: cross Abstract: Emergent misalignment (EM) occurs when narrow finetuning causes a model to behave dangerously outside the finetuning task. Standard training signals c

Trajectory-Refined Distillation

Model ReleasesDGX agent

arXiv:2606.08432v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision alo

Transforming Police-Car Swerving for Mitigating Isolated Stop-and-Go Traffic Waves: A Practice-Oriented Jam-Absorption Driving Strategy

SafetyDGX agent

arXiv:2602.10234v3 Announce Type: replace-cross Abstract: Stop-and-go traffic waves, a major form of freeway congestion, impose severe and persistent adverse impacts, including reduced traffic efficie

Transition-Based Digital Twin Modelling for Alzheimer's Disease under Sparse Longitudinal Data

ResearchDGX agent

arXiv:2606.09671v1 Announce Type: cross Abstract: Alzheimer's disease (AD) progression is highly heterogeneous and is typically observed through sparse and irregular longitudinal data, posing challeng

Traxia: A Framework for Verifiable, Agent-Native Scientific Publishing

AgentsDGX agent

arXiv:2606.08256v1 Announce Type: new Abstract: Verifiability, attribution, and reproducibility are foundational requirements of scientific knowledge, yet current publishing infrastructure does not en

TRIAGE: Dialectical Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series with LLMs

ResearchDGX agent

arXiv:2606.09030v1 Announce Type: cross Abstract: Clinical early warning systems built on electronic health records, in which clinical observations are recorded as irregularly sampled medical time ser

TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders

Model ReleasesDGX agent

arXiv:2606.09323v1 Announce Type: new Abstract: Tabular encoders are usually evaluated inside task-specific end-to-end pipelines, so models from different training paradigms are difficult to compare d

Trustworthy Smart Fabs via Professional Proxies: Scaling Safe and Sustainable by Design (SSbD) through Industrial Data Spaces

AgentsDGX agent

arXiv:2606.09227v1 Announce Type: cross Abstract: The convergence of the 2026 European Union Safe and Sustainable by Design (SSbD) framework, Corporate Sustainability Due Diligence Directive (CSDDD),

TT-DAC-PS: Twin-Target Deterministic Actor-Critic with Policy Smoothing for Optimal Trade Execution

Model ReleasesDGX agent

arXiv:2606.08379v1 Announce Type: new Abstract: This study addresses the optimal execution of large stock sell programs by introducing TT-DAC-PS (Twin-Target Deterministic Actor-Critic with Policy Smo

Tyan-WP: A Wind Power Foundation Model for Ultra-Short-Term Probabilistic Forecasting

ResearchDGX agent

arXiv:2606.08630v1 Announce Type: cross Abstract: Global wind power capacity, especially in China, is booming, with new farms spanning diverse terrains and climates. The industry urgently needs accura

UA-DCM: Uncertainty-aware Causal Decision Making via Effect Bound Decomposition

ApplicationsDGX agent

arXiv:2601.22736v2 Announce Type: replace-cross Abstract: Causal inference from observational data can provide strong evidence for finding the best action in a decision-making scenario without having

Unambiguous Representations in Neural Networks: An Information-Theoretic Approach to Intentionality

ResearchDGX agent

arXiv:2512.11000v2 Announce Type: replace-cross Abstract: Representations pervade our daily experience, from letters representing sounds to bit strings encoding digital files. While such representatio

Understanding Benchmark Language Under Weakened Formal Semantics

Model ReleasesDGX agent

arXiv:2509.17455v2 Announce Type: replace-cross Abstract: State-of-the-art NLP benchmarks require interpretation of natural language that specifies conditions, procedures, and exceptions, often relyin

Understanding Quantization-Aware Training: Gradients at Quantized Weights Bias to the Low-Loss Basin

Local AiDGX agent

arXiv:2606.09012v1 Announce Type: cross Abstract: Post-training quantization (PTQ) converts a trained full-precision model into low-bit weights without task-level retraining, while quantization-aware

Unification of Closed-Open Industrial Detection Scenarios: New Large-Scale Benchmarks,Challenges and Baselines

Model ReleasesDGX agent

arXiv:2606.07953v1 Announce Type: new Abstract: Large-scale Visual-Language Models (LVLMs) have achieved remarkable success in natural visual tasks, yet their application to industrial defect detectio

Unified Energy for Invariant and Independent Decoding in Diffusion Language Models

ResearchDGX agent

arXiv:2606.09159v1 Announce Type: cross Abstract: Diffusion Language Models (DLMs) enable parallel text generation by iteratively denoising a full sequence, offering attractive flexibility compared to

Unifying Object-Centric World Models and Diffusion Policy: A Hierarchical Framework for Multi-Stage Robotic Tasks

SafetyDGX agent

arXiv:2606.08775v1 Announce Type: cross Abstract: Visual world models have shown great potential in learning complex system dynamics. Recent advancements leverage these models as transition functions

UniQL: Towards Dialect-Universal Benchmarking for Text-to-SQL

Model ReleasesDGX agent

arXiv:2606.08018v1 Announce Type: new Abstract: Existing text-to-SQL benchmarks are largely centered on SQLite, making it difficult to evaluate whether models can generalize across heterogeneous SQL d

Unsupervised Partner Design Enables Robust Ad-hoc Teamwork

Model ReleasesDGX agent

arXiv:2508.06336v2 Announce Type: replace-cross Abstract: We introduce Unsupervised Partner Design (UPD), a population-free multi-agent reinforcement learning method for robust ad-hoc teamwork. UPD ge

Unveiling Privacy Risks in Multi-modal Large Language Models: Task-specific Vulnerabilities and Mitigation Challenges

ResearchDGX agent

arXiv:2606.09125v1 Announce Type: cross Abstract: Privacy risks in text-only Large Language Models (LLMs) are well studied, particularly their tendency to memorize and leak sensitive information. Howe

UnWeaving the knots of GraphRAG -- turns out VectorRAG is almost enough

ResearchDGX agent

arXiv:2603.29875v3 Announce Type: replace-cross Abstract: One of the key problems in Retrieval-augmented generation (RAG) systems is that chunk-based retrieval pipelines represent the source chunks as

Variational Speculative Decoding: Rethinking Draft Training from Token Likelihood to Sequence Acceptance

ResearchDGX agent

arXiv:2602.05774v4 Announce Type: replace-cross Abstract: Speculative decoding accelerates inference for (M)LLMs, yet a training-decoding discrepancy persists: while existing methods optimize single g

VATS: Exploiting Implicit Authority in Error-Path Injection via Systematic Mutation

Model ReleasesDGX agent

arXiv:2606.07992v1 Announce Type: new Abstract: As the Model Context Protocol (MCP) standardizes tool-calling for autonomous agents, it introduces a critical, unexamined attack surface: the error-hand

VESTA: A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents

SafetyDGX agent

arXiv:2606.08531v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, a

VFEM: Visual Feature Empowered Multivariate Time Series Forecasting with Cross-Modal Fusion

ResearchDGX agent

arXiv:2510.03244v2 Announce Type: replace-cross Abstract: Large time series foundation models often adopt channel-independent architectures to handle varying data dimensions, but this design ignores c

Video Understanding by Design: How Datasets Shape Video Models

SafetyDGX agent

arXiv:2509.09151v2 Announce Type: replace-cross Abstract: Research in video understanding has advanced rapidly, driven by increasingly diverse datasets and more powerful model architectures. While exi

ViMax: Agentic Video Generation

AgentsDGX agent

arXiv:2606.07649v1 Announce Type: cross Abstract: Long-form video generation requires systematic narrative planning and visual consistency that current short-clip methods cannot provide. Existing meth

Vision-Based Early Fault Diagnosis and Self-Recovery for Strawberry Harvesting Robots

ResearchDGX agent

arXiv:2601.02085v3 Announce Type: replace-cross Abstract: Strawberry-harvesting robots faced challenges such as poor visual perception, gripper misalignment, empty grasp/misgrasp, and slippage, which

Vision Language Model Helps Private Information De-Identification in Vision Data

TutorialsDGX agent

arXiv:2606.09132v1 Announce Type: new Abstract: Visual Language Models (VLMs) have gained significant popularity due to their remarkable ability. While various methods exist to enhance privacy in text

Visual Prompting Meets Feature Reconstruction-Based Anomaly Detection with Dual-Teacher Supervision

ApplicationsDGX agent

arXiv:2606.09670v1 Announce Type: cross Abstract: Recent Anomaly Detection methods achieve perfect detection and segmentation scores on well-established datasets, such as MVTec. However, many of these

VisualLeakBench: Reproducible Action-Boundary Propagation Failures in Vision-Language Agents

Model ReleasesDGX agent

arXiv:2606.07595v1 Announce Type: cross Abstract: Vision-language agents increasingly consume screenshots, documents, and user interfaces before writing to memory, sending messages, or invoking extern

← Previous
1…149150151152153…358
Next →