AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
9 Jun 2026

MC-CPO: Mastery-Conditioned Constrained Policy Optimization for Pedagogically Safe Intelligent Tutoring Systems

SafetyDGX agent

arXiv:2604.04251v2 Announce Type: replace Abstract: Intelligent tutoring systems increasingly rely on reinforcement learning to personalise instruction, yet optimising for observable engagement signal

Safe-RULE: Safe Reinforcement UnLEarning

Model ReleasesDGX agent

arXiv:2606.09559v1 Announce Type: cross Abstract: Offline safe reinforcement learning (Safe RL) enables policy learning without online interactions, making it suitable for safety-critical systems such

Semantic Quorum Assurance: Collective Certification for Non-Deterministic AI Infrastructure

SafetyDGX agent

arXiv:2606.08021v1 Announce Type: cross Abstract: As large language model (LLM) agents are integrated into autonomous cloud operations, distributed systems face a semantic reliability problem: propose

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
8 Jun 2026

Workflow-to-Skill: Skill Creation via Routing-Workflow-Semantics-Attachments Decomposition

SafetyDGX agent

arXiv:2606.06893v1 Announce Type: new Abstract: Large language model agents increasingly rely on Skills to encode procedural knowledge, yet high-quality Skills remain costly to hand-write. This paper

5 Jun 2026

UNIVID: Unified Vision-Language Model for Video Moderation

SafetyDGX agent

arXiv:2606.05748v1 Announce Type: cross Abstract: Global-scale video moderation faces a dual challenge: the need for fine-grained multi-modal reasoning and the demand for interpretable outputs to supp

4 Jun 2026

Distribution-Free Risk-Aware Planning and Control Under Uncertainty Using Conformal Spectral Risk Control

SafetyDGX agent

arXiv:2606.04185v1 Announce Type: new Abstract: Safe navigation in dynamic and uncertain environments often relies on accurate estimation of, or assumptions about, the true underlying uncertainty. How

From Agent Traces to Trust: Evidence Tracing and Execution Provenance in LLM Agents

SafetyDGX agent

arXiv:2606.04990v1 Announce Type: cross Abstract: Large language model (LLM)-based agents increasingly solve complex tasks by interacting with external tools, retrieval systems, memory modules, enviro

Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use

Model ReleasesDGX agent

arXiv:2603.03205v2 Announce Type: replace Abstract: Agentic language models operate in a fundamentally different safety regime than chat models: they must plan, call tools, and execute long-horizon ac

Reasoning over Boundaries: Enhancing Specification Alignment via Test-time Deliberation

Model ReleasesDGX agent

arXiv:2509.14760v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied in diverse real-world scenarios, each governed by bespoke behavioral and safety specifications

RSC: Decentralized Rigid Formation Flocking for Large-Scale Swarms via Hybrid Predictive Control and Online Reconfiguration

SafetyDGX agent

arXiv:2606.04248v1 Announce Type: new Abstract: Decentralized rigid formation flocking requires a swarm of autonomous agents to maintain a predetermined geometric configuration while moving, relying s

3 Jun 2026

Bridging Predictive Uncertainty and Safe Action: Sample-Conditioned Differentiable Planning for Autonomous Driving

SafetyDGX agent

arXiv:2606.03296v1 Announce Type: new Abstract: Complex, dynamic, and interactive driving environments pose significant challenges for autonomous driving, primarily due to the pervasive uncertainty of

DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair

Local AiDGX agent

arXiv:2606.03601v1 Announce Type: cross Abstract: While safety alignment and guardrails help large language models (LLMs) avoid harmful outputs, they can also induce overrefusal, i.e., unwarranted rej

Latent Activation Editing: Inference-Time Refinement of Learned Policies for Safer Multirobot Navigation

SafetyDGX agent

arXiv:2509.20623v2 Announce Type: replace Abstract: Reinforcement learning has enabled significant progress in complex domains such as coordinating and navigating multiple quadrotors. However, even we

OpenAI public policy agenda

SafetyDGX agent

OpenAI's public policy agenda outlines the organization's priorities and positions regarding AI regulation and governance. The document likely details OpenAI's recommendations for government policies,

Towards a Science of AI Agent Reliability

SafetyDGX agent

arXiv:2602.16666v3 Announce Type: replace Abstract: AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many age

When Models Refuse: Political Steerability and Feature Richness as Measures of Ideological Depth

SafetyDGX agent

arXiv:2508.21448v3 Announce Type: replace Abstract: Large language models (LLMs) sometimes refuse to follow benign instructions, such as declining to argue a political position or adopt a stated perso

2 Jun 2026

Infeasible optimization problems and the hierarchical augmented Lagrangian method in imitation learning

SafetyDGX agent

arXiv:2606.00730v1 Announce Type: new Abstract: Imitation learning (IL) is an effective approach to train complex robotics policies. Recent works have introduced hard constraints into imitation-learni

Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing

SafetyDGX agent

arXiv:2606.00033v1 Announce Type: cross Abstract: While mechanistic interpretability (MI) has produced important insights into neural network internals, the field has yet to establish a standardized s

Safe2Drive: Evaluating Safe Driving Behaviors of E2E Autonomous Driving Models

SafetyDGX agent

arXiv:2606.00191v1 Announce Type: cross Abstract: Recent end-to-end (E2E) autonomous driving policies achieve high driving scores in closed-loop simulations. Yet it remains unclear whether these polic

Tether-Aware Dynamic Collision Avoidance for USV-HROV Systems

SafetyDGX agent

arXiv:2606.01112v1 Announce Type: new Abstract: Heterogeneous marine robotic systems composed of an unmanned surface vehicle (USV) and a hybrid remotely operated vehicle (HROV) have shown great potent

1 Jun 2026

Stateful Online Monitoring Catches Distributed Agent Attacks

SafetyDGX agent

arXiv:2605.31593v1 Announce Type: cross Abstract: Language models can find thousands of severe software vulnerabilities, and agents are increasingly being misused for cyberattacks. To avoid detection,

TARIC: Memory-Augmented Traversability-Aware Outdoor VLN under Interrupted Semantic Cues

SafetyDGX agent

arXiv:2605.31121v1 Announce Type: cross Abstract: Outdoor vision-language navigation (VLN) in long-range, open-world environments is frequently disrupted by semantic-cue interruptions, where informati

29 May 2026

Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency

SafetyDGX agent

arXiv:2605.30208v1 Announce Type: cross Abstract: AI-assisted coding tools have altered software production. At Meta, significant lines of code per human-landed diff grew by 105.9% year over year and

LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback

SafetyDGX agent

arXiv:2605.30273v1 Announce Type: cross Abstract: Large language models (LLMs) show promise in generating supportive responses for mental health queries, but improving their usefulness, empathy, and s

28 May 2026

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving

Local AiDGX agent

arXiv:2605.27763v1 Announce Type: new Abstract: Safety evaluations of language models often treat serving configuration as fixed background infrastructure, but batch condition is an untested treatment

Chance-Constrained MPPI under State and Dynamic Object Prediction Uncertainty and the Evaluation of Collision Risk Calibration

SafetyDGX agent

arXiv:2605.28330v1 Announce Type: new Abstract: Chance-constrained Model Predictive Path Integral (MPPI) control is increasingly adopted for navigation in dynamic environments to explicitly bound coll

CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders

SafetyDGX agent

arXiv:2604.01604v2 Announce Type: replace Abstract: While modern LLMs are aligned to refuse harmful requests, it is essential to understand the underlying mechanistic basis of this refusal behavior fo

Grounded Cache Routing for Retrieval-Augmented Generation: When Is It Safe to Reuse an Answer?

SafetyDGX agent

arXiv:2605.27494v1 Announce Type: cross Abstract: Modern retrieval-augmented generation(RAG) deployments increasingly rely on caching to reduce token cost and time-to-first-token(TTFT). Prefix-level K

LACUNA: Safe Agents as Recursive Program Holes

SafetyDGX agent

arXiv:2605.28617v1 Announce Type: new Abstract: LLM agents increasingly act by writing code, yet a split persists between the runtime that drives the agent and the code the model writes. The runtime o

Models That Know How Evaluations Are Designed Score Safer

Model ReleasesDGX agent

arXiv:2605.28591v1 Announce Type: cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified tes

Trump loses more control over AI regulation as Illinois passes landmark law

SafetyDGX agent

Illinois' House of Representatives passed SB 315, a landmark bill requiring frontier AI companies like OpenAI and Anthropic to create, publish and annually update plans addressing severe or catastroph

27 May 2026

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models

SafetyDGX agent

arXiv:2605.06213v2 Announce Type: replace Abstract: Evaluating large language models (LLMs) today rests on fixed benchmarks that apply the same set of items to any model, producing ceiling and floor e

Bridging Control with Neural Network Verifier alpha-beta-CROWN: A Tutorial

SafetyDGX agent

arXiv:2605.26577v1 Announce Type: cross Abstract: Learning-based methods for synthesizing controllers have gained popularity due to their high expressiveness and strong empirical performance. However,

ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation

SafetyDGX agent

arXiv:2605.26542v1 Announce Type: cross Abstract: Tool-using agents increasingly operate in open-ended deployment environments, where they compose file systems, web APIs, code interpreters, and enterp

Modernising Reinforcement Learning-Based Navigation for Embodied Semantic Scene Graph Generation

SafetyDGX agent

arXiv:2603.25415v2 Announce Type: replace Abstract: Semantic world models enable embodied agents to reason about objects, relations, and spatial context beyond purely geometric representations. In Org

Provably Safe Motion Planning Under Unknown Disturbances

SafetyDGX agent

arXiv:2605.26625v1 Announce Type: new Abstract: We present a provably safe sampling-based motion planning algorithm for robotic systems affected by random disturbances of unknown distribution. We cons

Stochastic Decision Horizons for Constrained Reinforcement Learning

SafetyDGX agent

arXiv:2602.04599v2 Announce Type: replace Abstract: We propose stochastic decision horizons (SDH), a theoretically grounded framework for solving constrained RL problems with every-step constraint sat

26 May 2026

Auditing medical multi-agent AI reveals risks of false consensus

SafetyDGX agent

arXiv:2510.10185v2 Announce Type: replace-cross Abstract: Large language models are increasingly being assembled into medical multi-agent systems that emulate multidisciplinary consultation through sp

Energy Shields for Fairness

SafetyDGX agent

arXiv:2605.24926v1 Announce Type: new Abstract: Runtime fairness is not a one-time constraint but a dynamic property evaluated over a sequence of decisions. To ensure fairness at runtime, it is necess

First, do no harm: Breaking suicidogenic echo chambers in media recommendation

SafetyDGX agent

arXiv:2605.25258v1 Announce Type: cross Abstract: Recommender systems generally optimises user engagement, but this approach is dangerous in mental health contexts. When vulnerable users show signs of

On the Stability and Realizability of Recurrent Polynomial Surrogate Ternary Logic Gate Networks

SafetyDGX agent

arXiv:2605.24649v1 Announce Type: cross Abstract: Recurrent Neural Networks (RNNs) can learn to predict Signal Temporal Logic (STL) verdicts online from partial trajectories, but deploying them as run

Reason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous Driving

SafetyDGX agent

arXiv:2605.24004v1 Announce Type: new Abstract: Large language models (LLMs) are promising for autonomous driving, but semantics-only decision policies can yield physically unsafe behavior in dynamic

RECTOR: Priority-Aware Rule-Based Reranking for Compliance-Aware Autonomous Driving Trajectory Selection

Model ReleasesDGX agent

arXiv:2605.25095v1 Announce Type: new Abstract: Autonomous driving stacks must pick one trajectory from a multi-modal candidate set; choosing by model confidence ignores safety, traffic-law, and comfo

Referential Security as a New Paradigm for AI Evaluations

SafetyDGX agent

arXiv:2605.25673v1 Announce Type: cross Abstract: Security evaluations inherently depend on stable identifiers. Any finding, audit, or regulatory decision must remain attached to the specific artifact

SEIDM: A Safe and Efficient Intelligent Driver Model for Autonomous Driving Behavior

SafetyDGX agent

arXiv:2605.23915v1 Announce Type: cross Abstract: The Intelligent Driver Model (IDM) is a cornerstone of Adaptive Cruise Control (ACC), valued for its interpretable parameters and effectiveness in car

25 May 2026

RAG-Pull: Turning Retrieval into a Code-Injection Channel via Invisible Unicode Perturbations

SafetyDGX agent

arXiv:2510.11195v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) increases the reliability and trustworthiness of the LLM response and reduces hallucination by eliminatin

22 May 2026

Branch-Stochastic Model Predictive Control for Motion Planning under Multi-Modal Uncertainty with Scenario Clustering

SafetyDGX agent

arXiv:2605.22600v1 Announce Type: new Abstract: Motion planning for autonomous driving must account for multi-modal uncertainty in both the intentions and trajectories of surrounding vehicles. Handlin

Learning to Evolve: Multi-modal Interactive Fields for Robust Humanoid Navigation in Dynamic Environments

SafetyDGX agent

arXiv:2605.21935v1 Announce Type: new Abstract: Safe manipulation-oriented navigation for humanoid robots requires scene memory that remains reliable under locomotion-induced perceptual distortion, en

Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI

SafetyDGX agent

arXiv:2507.05660v3 Announce Type: replace-cross Abstract: Customizing Large Language Models (LLMs) on untrusted datasets poses severe risks of injecting toxic behaviors. In this work, we introduce Opt

Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning

SafetyDGX agent

arXiv:2605.22748v1 Announce Type: new Abstract: Autonomous systems have achieved superhuman performance in isolation or simulation, yet they remain brittle in shared, dynamic real-world spaces. This f

21 May 2026

Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving

SafetyDGX agent

arXiv:2605.21139v1 Announce Type: new Abstract: Current end-to-end autonomous driving models are fundamentally constrained by the behavioral cloning ceiling of imitation learning. While reinforcement

REFLECTOR: Internalizing Step-wise Reflection against Indirect Jailbreak

SafetyDGX agent

arXiv:2605.20654v1 Announce Type: new Abstract: While Large Language Models (LLMs) demonstrate remarkable capabilities, they remain susceptible to sophisticated, multi-step jailbreak attacks that circ

20 May 2026

Neural Configuration-Space Barriers for Manipulation Planning and Control

SafetyDGX agent

arXiv:2503.04929v3 Announce Type: replace-cross Abstract: Planning and control for high-dimensional robot manipulators in cluttered dynamic environments require computational efficiency and robust saf

Surviving the Unseen: Predictive Defense for Novel Multi-Turn Multimodal Attacks

SafetyDGX agent

arXiv:2605.18988v1 Announce Type: cross Abstract: The expansion of Multimodal Large Language Models (MLLMs) and their integration into autonomous agentic workflows has introduced a non-stationary atta

19 May 2026

Adversarial Fragility and Language Vulnerability in Clinical AI: A Systematic Audit of Diagnostic Collapse Under Imperceptible Perturbations and Cross-Lingual Drift in Low-Resource Healthcare Settings

SafetyDGX agent

arXiv:2605.16993v1 Announce Type: cross Abstract: Current clinical artificial intelligence (AI) systems are evaluated almost exclusively on clean, standardised, English-language inputs, conditions tha

Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression

SafetyDGX agent

arXiv:2605.17304v1 Announce Type: cross Abstract: LLM context is not just tokens; it is a set of commitments. Long-running conversations accumulate goals, constraints, decisions, preferences, tool res

Contrastive-SDXL: Annotation-Preserving Night-Time Augmentation for Pedestrian Detection

SafetyDGX agent

arXiv:2605.16406v1 Announce Type: new Abstract: Night-time pedestrian detection remains challenging because labelled night-time data are limited and large illumination differences make daytime-only tr

Pedestrian-Aware LLM-Driven Behavioral Planning for Autonomous Vehicles

SafetyDGX agent

arXiv:2605.16858v1 Announce Type: cross Abstract: Autonomous Vehicles (AVs) must make reliable decisions in dense urban environments where pedestrian behavior is variable, sometimes abnormal, and ofte

Prediction of Challenging Behaviors Associated with Profound Autism in a Classroom Setting Using Wearable Sensors

SafetyDGX agent

arXiv:2605.17618v1 Announce Type: new Abstract: Autism Spectrum Disorder (ASD) is characterized by challenges with social interaction and communication and by restricted or repetitive patterns of thou

SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training

SafetyDGX agent

arXiv:2605.18719v1 Announce Type: new Abstract: Diffusion models have been widely studied for removing unsafe content learned during pre-training. Existing methods require expensive supervised data, e

← Previous
1…2829303132…240
Next →