AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlog
85,136Total entries
1Added by human
85,135Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,688 results
Safety

Testing the Black Box: Structural Barriers to Independent Evaluation of Consumer-Facing Health LLMs

DGX agent

arXiv:2606.08483v1 Announce Type: new Abstract: Background: Consumer-facing large language models are now a common source of health information, and they interpret and personalize responses rather tha

safetyarxiv-cs-ai
9 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

The ACUTE Protocol: Operationalizing Language Model Activations for Better Calibration, Utility, and Trust

DGX agent

arXiv:2606.07822v1 Announce Type: cross Abstract: As language models improve and become increasingly deployed to solve a variety of tasks, trustworthiness becomes essential. Calibration is a good prox

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

The AI Epistemic Deference Index: A Continuous Measure of Sycophancy

DGX agent

arXiv:2606.07897v1 Announce Type: new Abstract: Current AI models frequently exhibit epistemic sycophancy, endorsing claims to agree with a user. Existing evaluations typically measure this either by

model-releasesarxiv-cs-ai
9 Jun 2026
Tutorials

The CIFAR Synthetic Evidence Corpus for Detecting AI-Generated Evidence

DGX agent

arXiv:2606.07916v1 Announce Type: new Abstract: The growing ability of generative models to produce realistic documents poses a direct challenge to evidentiary workflows in the justice system and the

tutorialsarxiv-cs-ai
9 Jun 2026
Safety

The Confidence Trap: Calibration Attacks for Graph Neural Networks

DGX agent

arXiv:2606.08467v1 Announce Type: cross Abstract: While confidence calibration is essential for trustworthy decision-making in safety-critical applications, the robustness of calibrated GNNs to advers

safetyarxiv-cs-ai
9 Jun 2026
Safety

The Cross-Architecture Substrate: A Domain-Transcendent, Calibration-Surviving Geometric Invariant of Modern Vision Encoders

DGX agent

arXiv:2606.07882v1 Announce Type: cross Abstract: Different vision neural networks -- trained to classify, contrast, reconstruct, or match images to text -- should have correspondingly different inter

safetyarxiv-cs-ai
9 Jun 2026
Safety

The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models

DGX agent

arXiv:2601.15165v4 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) break the rigid left-to-right constraint of traditional LLMs, enabling token generation in arbitrary o

safetyarxiv-cs-ai
9 Jun 2026
Safety

The Governance of Human-LLM Interaction: Safety Gating, Civility Steering, and Affective Default Lock-In

DGX agent

arXiv:2606.08172v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly mediate high-stakes interactions in finance, medicine, and mental-health support, yet users have limited con

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models

DGX agent

arXiv:2606.07861v1 Announce Type: cross Abstract: Recent vision-language models (VLMs) excel at multimodal understanding and reasoning, yet their fine-grained visual perception remains underexplored.

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

The Montparnasse Algorithm for RNA Design

DGX agent

arXiv:2606.07562v1 Announce Type: cross Abstract: RNA design consists of discovering a nucleotide sequence that optimizes predefined criteria, such as secondary structure. It is useful for synthetic b

model-releasesarxiv-cs-ai
9 Jun 2026
Agents

The Token Not Taken: Sampling, State, and the Variability of AI Agent Outputs

DGX agent

arXiv:2606.08998v1 Announce Type: new Abstract: Agentic AI systems can behave differently across runs: the same request may produce a different plan, a different tool call, a different code edit, or a

agentsarxiv-cs-ai
9 Jun 2026
Research

The Topological Dual of a Dataset: A Logic-to-Topology Encoding for AlphaGeometry-Style Data

DGX agent

arXiv:2604.18050v2 Announce Type: replace Abstract: AlphaGeometry represents a milestone in neuro-symbolic reasoning, yet its architecture faces a log-linear scaling bottleneck within its symbolic ded

researcharxiv-cs-ai
9 Jun 2026
Model Releases

TheoremBench: Evaluating LLMs on Theorem Proving in Formal Mathematics

DGX agent

arXiv:2606.09450v1 Announce Type: new Abstract: LLMs have recently achieved strong results on formal proving benchmarks. However, existing evaluations remain heavily concentrated on competition-style

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Think Before You Act: Intention-Guided Reasoning for LLM-Based Location Prediction

DGX agent

arXiv:2606.08122v1 Announce Type: new Abstract: Predicting a user's next Point-of-Interest (POI) based on their historical check-in records is a fundamental task in location-based services. While rece

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning

DGX agent

arXiv:2601.04805v2 Announce Type: replace Abstract: Large reasoning models (LRMs) have attracted much attention due to their exceptional performance. However, their performance mainly stems from think

model-releasesarxiv-cs-ai
9 Jun 2026
Agents

TianJi-Environ: An Autonomous AI Scientist for Atmospheric Environmental Research

DGX agent

arXiv:2606.07697v1 Announce Type: cross Abstract: As atmospheric environmental prediction continues to improve, interpretable validation of pollution mechanisms and feedback processes has become a mai

agentsarxiv-cs-ai
9 Jun 2026
Research

TimpaTeks: Automatic In-place Text Sequence Modification via Diffusion Language Model Steering

DGX agent

arXiv:2606.08408v1 Announce Type: cross Abstract: We extend activation steering to diffusion language models (DLMs) and study a novel problem that arose due to the inference mechanism of DLMs: Modifyi

researcharxiv-cs-ai
9 Jun 2026
Research

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech

DGX agent

arXiv:2606.09019v1 Announce Type: cross Abstract: Codec-based autoregressive (AR) speech language models have achieved strong text-to-speech (TTS) quality by modeling speech as sequences of discrete a

researcharxiv-cs-ai
9 Jun 2026
Agents

To Nuke or Not to Nuke: LLMs' (Missing) Ethical Reasoning and Actions in a High-Stakes Decision-Making Simulation

DGX agent

arXiv:2606.08310v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as long-horizon agents with decision-making capacities. While LLMs can show ethical competence on

agentsarxiv-cs-ai
9 Jun 2026
Research

Topological Neural Operators

DGX agent

arXiv:2606.09806v1 Announce Type: cross Abstract: We introduce Topological Neural Operators (TNOs), a principled framework for operator learning on cell complexes that lifts neural operators (NOs) fro

researcharxiv-cs-ai
9 Jun 2026
Safety

Toward autocorrection of chemical process flowsheets using large language models

DGX agent

arXiv:2312.02873v2 Announce Type: replace-cross Abstract: The process engineering domain widely uses Process Flow Diagrams (PFDs) and Process and Instrumentation Diagrams (P&IDs) to represent process

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Towards Long-Horizon Vessel Trajectory and Destination Forecasting with Reasoning Large Language Models

DGX agent

arXiv:2606.08633v1 Announce Type: new Abstract: Long-horizon maritime trajectory prediction is important for shipping management, logistics planning, and maritime risk analysis, yet month-level foreca

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

TQA-Bench: Evaluating LLMs for Multi-Table Question Answering

DGX agent

arXiv:2411.19504v2 Announce Type: replace Abstract: The advance of large language models (LLMs) has unlocked great opportunities in complex multi-modal data management tasks, particularly in question

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

TRACER: Token ReAssignment for Concept ERasure in Generative Recommendation

DGX agent

arXiv:2606.07688v1 Announce Type: cross Abstract: Generative recommendation formulates next-item prediction as autoregressive generation over semantic ID (SID) sequences derived from users' historical

safetyarxiv-cs-ai
9 Jun 2026
Tutorials

Train at Moving Edge: Online-Verified Prompt Selection for Efficient RL Training of Large Reasoning Model

DGX agent

arXiv:2603.25184v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become essential for post-training large language models (LLMs) in reasoning tasks. While scaling rollouts can

tutorialsarxiv-cs-ai
9 Jun 2026
Research

Training-Free Intelligibility-Guided Observation Addition for Noisy ASR

DGX agent

arXiv:2602.20967v2 Announce Type: replace-cross Abstract: Automatic speech recognition (ASR) degrades severely in noisy environments. Although speech enhancement (SE) front-ends effectively suppress b

researcharxiv-cs-ai
9 Jun 2026
Safety

Training-Inference Kernel Contracts: Bounding Divergence in Post-Training and Deployment

DGX agent

arXiv:2606.07581v1 Announce Type: cross Abstract: A modern post-training pipeline often writes one symbol for its policy, pi_theta, while evaluating it through two different programs: a training kerne

safetyarxiv-cs-ai
9 Jun 2026
Safety

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning

DGX agent

arXiv:2606.07631v1 Announce Type: cross Abstract: Emergent misalignment (EM) occurs when narrow finetuning causes a model to behave dangerously outside the finetuning task. Standard training signals c

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Trajectory-Refined Distillation

DGX agent

arXiv:2606.08432v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision alo

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Transforming Police-Car Swerving for Mitigating Isolated Stop-and-Go Traffic Waves: A Practice-Oriented Jam-Absorption Driving Strategy

DGX agent

arXiv:2602.10234v3 Announce Type: replace-cross Abstract: Stop-and-go traffic waves, a major form of freeway congestion, impose severe and persistent adverse impacts, including reduced traffic efficie

safetyarxiv-cs-ai
9 Jun 2026
Research

Transition-Based Digital Twin Modelling for Alzheimer's Disease under Sparse Longitudinal Data

DGX agent

arXiv:2606.09671v1 Announce Type: cross Abstract: Alzheimer's disease (AD) progression is highly heterogeneous and is typically observed through sparse and irregular longitudinal data, posing challeng

researcharxiv-cs-ai
9 Jun 2026
Agents

Traxia: A Framework for Verifiable, Agent-Native Scientific Publishing

DGX agent

arXiv:2606.08256v1 Announce Type: new Abstract: Verifiability, attribution, and reproducibility are foundational requirements of scientific knowledge, yet current publishing infrastructure does not en

agentsarxiv-cs-ai
9 Jun 2026
Research

TRIAGE: Dialectical Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series with LLMs

DGX agent

arXiv:2606.09030v1 Announce Type: cross Abstract: Clinical early warning systems built on electronic health records, in which clinical observations are recorded as irregularly sampled medical time ser

researcharxiv-cs-ai
9 Jun 2026
Model Releases

TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders

DGX agent

arXiv:2606.09323v1 Announce Type: new Abstract: Tabular encoders are usually evaluated inside task-specific end-to-end pipelines, so models from different training paradigms are difficult to compare d

model-releasesarxiv-cs-ai
9 Jun 2026
Agents

Trustworthy Smart Fabs via Professional Proxies: Scaling Safe and Sustainable by Design (SSbD) through Industrial Data Spaces

DGX agent

arXiv:2606.09227v1 Announce Type: cross Abstract: The convergence of the 2026 European Union Safe and Sustainable by Design (SSbD) framework, Corporate Sustainability Due Diligence Directive (CSDDD),

agentsarxiv-cs-ai
9 Jun 2026
Model Releases

TT-DAC-PS: Twin-Target Deterministic Actor-Critic with Policy Smoothing for Optimal Trade Execution

DGX agent

arXiv:2606.08379v1 Announce Type: new Abstract: This study addresses the optimal execution of large stock sell programs by introducing TT-DAC-PS (Twin-Target Deterministic Actor-Critic with Policy Smo

model-releasesarxiv-cs-ai
9 Jun 2026
Research

Tyan-WP: A Wind Power Foundation Model for Ultra-Short-Term Probabilistic Forecasting

DGX agent

arXiv:2606.08630v1 Announce Type: cross Abstract: Global wind power capacity, especially in China, is booming, with new farms spanning diverse terrains and climates. The industry urgently needs accura

researcharxiv-cs-ai
9 Jun 2026
Applications

UA-DCM: Uncertainty-aware Causal Decision Making via Effect Bound Decomposition

DGX agent

arXiv:2601.22736v2 Announce Type: replace-cross Abstract: Causal inference from observational data can provide strong evidence for finding the best action in a decision-making scenario without having

applicationsarxiv-cs-ai
9 Jun 2026
Research

Unambiguous Representations in Neural Networks: An Information-Theoretic Approach to Intentionality

DGX agent

arXiv:2512.11000v2 Announce Type: replace-cross Abstract: Representations pervade our daily experience, from letters representing sounds to bit strings encoding digital files. While such representatio

researcharxiv-cs-ai
9 Jun 2026
Model Releases

Understanding Benchmark Language Under Weakened Formal Semantics

DGX agent

arXiv:2509.17455v2 Announce Type: replace-cross Abstract: State-of-the-art NLP benchmarks require interpretation of natural language that specifies conditions, procedures, and exceptions, often relyin

model-releasesarxiv-cs-ai
9 Jun 2026
Local Ai

Understanding Quantization-Aware Training: Gradients at Quantized Weights Bias to the Low-Loss Basin

DGX agent

arXiv:2606.09012v1 Announce Type: cross Abstract: Post-training quantization (PTQ) converts a trained full-precision model into low-bit weights without task-level retraining, while quantization-aware

local-aiarxiv-cs-ai
9 Jun 2026
Model Releases

Unification of Closed-Open Industrial Detection Scenarios: New Large-Scale Benchmarks,Challenges and Baselines

DGX agent

arXiv:2606.07953v1 Announce Type: new Abstract: Large-scale Visual-Language Models (LVLMs) have achieved remarkable success in natural visual tasks, yet their application to industrial defect detectio

model-releasesarxiv-cs-ai
9 Jun 2026
Research

Unified Energy for Invariant and Independent Decoding in Diffusion Language Models

DGX agent

arXiv:2606.09159v1 Announce Type: cross Abstract: Diffusion Language Models (DLMs) enable parallel text generation by iteratively denoising a full sequence, offering attractive flexibility compared to

researcharxiv-cs-ai
9 Jun 2026
Safety

Unifying Object-Centric World Models and Diffusion Policy: A Hierarchical Framework for Multi-Stage Robotic Tasks

DGX agent

arXiv:2606.08775v1 Announce Type: cross Abstract: Visual world models have shown great potential in learning complex system dynamics. Recent advancements leverage these models as transition functions

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

UniQL: Towards Dialect-Universal Benchmarking for Text-to-SQL

DGX agent

arXiv:2606.08018v1 Announce Type: new Abstract: Existing text-to-SQL benchmarks are largely centered on SQLite, making it difficult to evaluate whether models can generalize across heterogeneous SQL d

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Unsupervised Partner Design Enables Robust Ad-hoc Teamwork

DGX agent

arXiv:2508.06336v2 Announce Type: replace-cross Abstract: We introduce Unsupervised Partner Design (UPD), a population-free multi-agent reinforcement learning method for robust ad-hoc teamwork. UPD ge

model-releasesarxiv-cs-ai
9 Jun 2026
Research

Unveiling Privacy Risks in Multi-modal Large Language Models: Task-specific Vulnerabilities and Mitigation Challenges

DGX agent

arXiv:2606.09125v1 Announce Type: cross Abstract: Privacy risks in text-only Large Language Models (LLMs) are well studied, particularly their tendency to memorize and leak sensitive information. Howe

researcharxiv-cs-ai
9 Jun 2026
Research

UnWeaving the knots of GraphRAG -- turns out VectorRAG is almost enough

DGX agent

arXiv:2603.29875v3 Announce Type: replace-cross Abstract: One of the key problems in Retrieval-augmented generation (RAG) systems is that chunk-based retrieval pipelines represent the source chunks as

researcharxiv-cs-ai
9 Jun 2026
← Previous
1…191192193194195…452
Next →