AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,435 results
Safety

VeriGate: Verifier-Gated Step-Level Supervision for GRPO

DGX agent

arXiv:2605.30451v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) is an effective recipe for training reasoning models with verifier-based outcome rewards, but its supervision

safetyarxiv-cs-lg
1 Jun 2026
Safety

Vision-Language Models Suppress Female Representations Under Ambiguous Input

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2605.31556v1 Announce Type: cross Abstract: Alignment teaches vision-language models (VLMs) to avoid expressing demographic biases, and when gender is clearly visible they largely succeed. Far l

safetyarxiv-cs-ai
1 Jun 2026
Safety

Wall-OSS-0.5 Technical Report

DGX agent

arXiv:2605.30877v1 Announce Type: new Abstract: Large-scale Vision-Language-Action (VLA) pretraining is increasingly adopted as the foundation for robot policies, yet the evidence for pretrained VLAs

safetyarxiv-cs-ro
1 Jun 2026
Safety

What Am I Missing? Question-Answering as Hidden State Probing

DGX agent

arXiv:2605.31561v1 Announce Type: new Abstract: Test-time reasoning has become a significant field of study since the introduction of chain-of-thought reasoning in large language models (LLMs). Howeve

safetyarxiv-cs-cl
1 Jun 2026
Safety

When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?

DGX agent

arXiv:2605.30719v1 Announce Type: cross Abstract: We study when large language models (LLMs) can serve as effective black-box policy optimizers for reinforcement learning (RL) tasks, i.e., when can we

safetyarxiv-cs-ai
1 Jun 2026
Safety

Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems

DGX agent

arXiv:2506.00175v5 Announce Type: replace-cross Abstract: Modern AI systems are typically developed through multiple stages-pretraining, fine-tuning rounds, and subsequent adaptation or alignment, whe

safetyarxiv-cs-ai
1 Jun 2026
Safety

Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning

DGX agent

arXiv:2605.31261v1 Announce Type: cross Abstract: The family of linear recurrent neural networks has shown strong performance as recurrent memory units in partially observable reinforcement learning.

safetyarxiv-cs-ai
1 Jun 2026
Safety

World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry

DGX agent

arXiv:2604.01985v2 Announce Type: replace-cross Abstract: General-purpose world models promise scalable policy evaluation, optimization, and planning, yet achieving the required level of robustness re

safetyarxiv-cs-ai
1 Jun 2026
Safety

World2Act: Latent Action Post-Training from World Model Dynamics

DGX agent

arXiv:2603.10422v2 Announce Type: replace Abstract: World Models (WMs) offer a promising mechanism for post-training Vision-Language-Action (VLA) policies by providing dynamics priors that improve gen

safetyarxiv-cs-cv
1 Jun 2026
Safety

Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation

DGX agent

arXiv:2605.30833v1 Announce Type: cross Abstract: On-policy distillation transfers reasoning capabilities by training a student model on its own generated trajectories using token-level feedback from

safetyarxiv-cs-ai
1 Jun 2026
Safety

ZAPS-DA: Zero-Phase Action Policy Smoothing with Decoupled Actor for Continuous Control in Reinforcement Learning

DGX agent

arXiv:2605.30612v1 Announce Type: cross Abstract: Continuous control policies trained with off-policy reinforcement learning frequently exhibit high-frequency action jitter, rendering direct deploymen

safetyarxiv-cs-lg
1 Jun 2026
Safety

Zero Collapse: A Failure Mode of Policy Gradient Methods in Discontinuous Reward Environments

DGX agent

arXiv:2605.30896v1 Announce Type: new Abstract: Bidding in repeated auctions is a central challenge for reinforcement learning (RL), combining continuous control with the strategic complexities of dig

safetyarxiv-cs-lg
1 Jun 2026
Safety

A Fully Convolutional Approach to Denoising Structural Dynamics Data from X-Ray Photon Correlation Spectroscopy

DGX agent

arXiv:2605.29975v1 Announce Type: new Abstract: We present a fully convolutional denoising autoencoder (FC-DAE) for denoising two-time intensity-intensity correlation functions (C_2) in X-ray photon c

safetyarxiv-cs-lg
29 May 2026
Safety

A Geometric View of SRC: Learning Representations for Stable Residual Inference

DGX agent

arXiv:2605.29673v1 Announce Type: cross Abstract: Reconstruction-based inference assigns a class by comparing class-wise reconstruction residuals; Sparse Representation Classification (SRC) is a canon

safetyarxiv-cs-cv
29 May 2026
Safety

A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms

DGX agent

arXiv:2605.30313v1 Announce Type: new Abstract: Simulation-based RL for contemporary robot control is increasingly organized around GPU-resident simulation: physics, rollout collection, and learning a

safetyarxiv-cs-ro
29 May 2026
Safety

A Modular Architecture for Typologically Controlled Lexicon Generation

DGX agent

arXiv:2605.28824v1 Announce Type: new Abstract: Constructing artificial lexicons that are pronounceable, typologically plausible, and semantically structured remains an open challenge in computational

safetyarxiv-cs-cl
29 May 2026
Safety

A Predictive Law for On-Policy Self-Distillation From World Feedback

DGX agent

arXiv:2605.30070v1 Announce Type: cross Abstract: Moving beyond simple scalar rewards toward richer world feedback is a natural path to more scalable RL post-training. On-policy self-distillation (OPS

safetyarxiv-cs-ai
29 May 2026
Safety

ActTraitBench: Quantifying the Knowledge-Decision Gap in Large Language Models via Human-Grounded Behavioral Validation

DGX agent

arXiv:2605.29791v1 Announce Type: new Abstract: While Large Language Models (LLMs) can convincingly simulate personas in explicit self-reports, they often deviate in implicit behavioral decisions, rev

safetyarxiv-cs-cl
29 May 2026
Safety

Adaptive Interviewing for Persona Simulation in LLMs: Evidence-Grounded Reasoning Improves Decision Alignment

DGX agent

arXiv:2605.29458v1 Announce Type: cross Abstract: Accurately simulating the decisions of a specific individual remains challenging for large language models (LLMs), partly because persona information

safetyarxiv-cs-ai
29 May 2026
Safety

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

DGX agent

arXiv:2603.01006v2 Announce Type: replace-cross Abstract: REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher

safetyarxiv-cs-ai
29 May 2026
Safety

AIRGuard: Guarding Agent Actions with Runtime Authority Control

DGX agent

arXiv:2605.28914v1 Announce Type: cross Abstract: Tool-using language agents turn model decisions into external side effects: they read files, run scripts, call APIs, send messages, and invoke Model C

safetyarxiv-cs-ai
29 May 2026
Safety

AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing

DGX agent

arXiv:2605.29434v1 Announce Type: cross Abstract: Existing sentence-level watermarking methods enhance robustness to paraphrasing by anchoring watermarks in sentence semantics. However, their prefix-b

safetyarxiv-cs-ai
29 May 2026
Safety

Auditing Training Data in Generative Music Models via Black-Box Membership Inference

DGX agent

arXiv:2605.29202v1 Announce Type: new Abstract: Recent advances in text-to-music generation enable high-fidelity synthesis of structured musical audio, raising growing concerns about data provenance,

safetyarxiv-cs-lg
29 May 2026
Safety

BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation

DGX agent

arXiv:2605.28994v1 Announce Type: new Abstract: AI tools to support real world decision making must be able to build simulation models that inform their recommendations and render them interpretable.

safetyarxiv-cs-ai
29 May 2026
Safety

Behavior-Induced Mirror-Prox Temporal-Difference Learning for Faster Off-Policy Prediction

DGX agent

arXiv:2605.28849v1 Announce Type: new Abstract: Gradient temporal-difference methods provide stable off-policy prediction with linear function approximation, but their practical performance is strongl

safetyarxiv-cs-ai
29 May 2026
Safety

Beyond Bilingual Transfer: Multilingual Code-Switching in Instruction Tuning

DGX agent

arXiv:2605.29414v1 Announce Type: cross Abstract: Recent studies have shown that code-switching data (CSD), in which multiple languages are mixed within the same context, can improve cross-lingual tra

safetyarxiv-cs-ai
29 May 2026
Safety

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling

DGX agent

arXiv:2605.29697v1 Announce Type: new Abstract: In Agentic Search, trajectory-level outcome rewards fail to quantify the behavioral contributions of individual steps, while existing step-level reward

safetyarxiv-cs-ai
29 May 2026
Safety

BORA: Bridging Offline Reinforcement Learning and Online Residual Adaptation for Real-World Dexterous VLA Models

DGX agent

arXiv:2605.30226v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising paradigm for grounding visual-language understanding into real-world robotic manipulat

safetyarxiv-cs-ai
29 May 2026
Safety

Bridging the Sim-to-Real Gap in Reinforcement Learning-Based Industrial Dispatching through Execution Semantics

DGX agent

arXiv:2605.29078v1 Announce Type: new Abstract: Event-driven scheduling policies are increasingly deployed in industrial environments, where decisions are made under asynchronous and partially observe

safetyarxiv-cs-ai
29 May 2026
Safety

Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations

DGX agent

arXiv:2601.08064v2 Announce Type: replace Abstract: Confidence estimation (CE) indicates how reliable the answers of large language models are and impacts user trust and decision-making. Existing eval

safetyarxiv-cs-cl
29 May 2026
Safety

Causal Interventions on Continuous Variables: A Case Study on Verb Bias in Steering Vectors for In-Context Learning

DGX agent

arXiv:2605.29971v1 Announce Type: new Abstract: Causal interventions in language model representations have largely targeted discrete features, like grammatical number. However, language models must a

safetyarxiv-cs-cl
29 May 2026
Safety

Causal-JEPA: Learning World Models through Object-Level Latent Masking

DGX agent

arXiv:2602.11389v2 Announce Type: replace Abstract: World models require robust relational understanding to support prediction, reasoning, and control. While object-centric representations provide a u

safetyarxiv-cs-ai
29 May 2026
Safety

CB-SLICE: Concept-Based Interpretable Error Slice Discovery

DGX agent

arXiv:2605.29836v1 Announce Type: cross Abstract: Despite strong average-case performance, deep learning models often exhibit systematic errors on specific population groups, known as error slices. Id

safetyarxiv-cs-ai
29 May 2026
Safety

Certified Policy Optimisation for Nested Causal Bandits via PAC-Bayes Risk

DGX agent

arXiv:2605.29788v1 Announce Type: new Abstract: Critical sequential decisions are rarely single-timescale: a strategic decision causally shapes the context in which every subsequent tactical choice is

safetyarxiv-cs-ai
29 May 2026
Safety

Colored Noise Diffusion Sampling

DGX agent

arXiv:2605.30332v1 Announce Type: new Abstract: Diffusion models achieve state-of-the-art image synthesis, with their generative trajectories fundamentally exhibiting a spectral bias, resolving low-fr

safetyarxiv-cs-cv
29 May 2026
Safety

Comparative evaluation of photogrammetric reconstruction methods and 3D Gaussian Splatting for road surface roughness analysis

DGX agent

arXiv:2605.29452v1 Announce Type: new Abstract: Image-based 3D reconstruction offers a low-cost alternative to traditional sensor-based techniques for road surface assessment. This study compares four

safetyarxiv-cs-cv
29 May 2026
Safety

Crafting Desirable Climate Trajectories with RL Explored Socio-Environmental Simulations

DGX agent

arXiv:2410.07287v2 Announce Type: replace-cross Abstract: Climate change poses an existential threat, necessitating effective climate policies to enact impactful change. Decisions in this domain are i

safetyarxiv-cs-ai
29 May 2026
Safety

CRITIC-R1: Learning Structured Critics for Retrieval-Augmented Generation

DGX agent

arXiv:2605.29886v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) improves knowledge-intensive question answering by incorporating external evidence. However, existing RAG methods

safetyarxiv-cs-ai
29 May 2026
Safety

Cycle Consistency in Video Object-Centric Learning

DGX agent

arXiv:2605.30211v1 Announce Type: new Abstract: Self-supervised video Object-Centric Learning (OCL) aims to discover distinct objects and associate them across time, whereas self-supervised Multi-Obje

safetyarxiv-cs-cv
29 May 2026
Safety

DAMEL: Dual-Axis Multi-Expert Learning for Class-Imbalanced Learning

DGX agent

arXiv:2605.30135v1 Announce Type: cross Abstract: Various algorithms have been proposed to address the challenges posed by class-imbalanced learning from real-world data with long-tailed distributions

safetyarxiv-cs-ai
29 May 2026
Safety

DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation

DGX agent

arXiv:2605.29522v1 Announce Type: new Abstract: As scientific literature grows rapidly, automated survey generation has become a key capability for AI scientists and human researchers. However, existi

safetyarxiv-cs-ai
29 May 2026
Safety

Deja View: Looping Transformers for Multi-View 3D Reconstruction

DGX agent

arXiv:2605.30215v1 Announce Type: new Abstract: Recent feed-forward 3D reconstruction transformers have scaled to over a billion parameters, following the broader trend of increasing model capacity in

safetyarxiv-cs-cv
29 May 2026
Safety

Discovering Cooperative Pipelines: Autoresearch for Sequential Social Dilemmas

DGX agent

arXiv:2605.30003v1 Announce Type: cross Abstract: We study two-level autoresearch for cooperation: an outer-loop AI agent autonomously redesigns the inner-loop pipeline of an LLM policy-synthesis syst

safetyarxiv-cs-ai
29 May 2026
Safety

Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias

DGX agent

arXiv:2605.29152v1 Announce Type: new Abstract: Randomly initialized neural networks induce a prior over functions, but the predictor used in practice is produced only after training. We ask how much

safetyarxiv-cs-lg
29 May 2026
Safety

Draft-OPD: On-Policy Distillation for Speculative Draft Models

DGX agent

arXiv:2605.29343v1 Announce Type: new Abstract: Speculative decoding accelerates large language model inference by pairing a target model with a lightweight draft model whose proposed tokens are verif

safetyarxiv-cs-cl
29 May 2026
Safety

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

DGX agent

arXiv:2510.27607v3 Announce Type: replace Abstract: Augmenting vision-language-action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predictin

safetyarxiv-cs-cv
29 May 2026
Safety

DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

DGX agent

arXiv:2605.30350v1 Announce Type: cross Abstract: Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built

safetyarxiv-cs-lg
29 May 2026
Safety

Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure

DGX agent

arXiv:2602.08783v3 Announce Type: replace Abstract: Latent or continuous chain-of-thought methods replace explicit textual rationales with a number of internal latent steps, but these intermediate com

safetyarxiv-cs-ai
29 May 2026
← Previous
1…149150151152153…260
Next →