AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
17 Apr 2026

Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems

SafetyDGX agent

arXiv:2510.14133v2 Announce Type: replace Abstract: Agentic AI systems, which leverage multiple autonomous agents and large language models (LLMs), are increasingly used to address complex, multi-step

15 Apr 2026

Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design

SafetyDGX agent

arXiv:2604.12500v1 Announce Type: new Abstract: Specification gaming under Reinforcement Learning (RL) is known to cause LLMs to develop sycophantic, manipulative, or deceptive behavior, yet the condi

14 Apr 2026

Anthropic believes that good transparency legislation needs to ensure public safety and accountability for the companies developing this pow…

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
SafetyDGX agent

Anthropic believes that good transparency legislation needs to ensure public safety and accountability for the companies developing this powerful technology, not provide a get-out-of-jail-free card ag

Seeing No Evil: Blinding Large Vision-Language Models to Safety Instructions via Adversarial Attention Hijacking

Model ReleasesDGX agent

arXiv:2604.10299v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) rely on attention-based retrieval of safety instructions to maintain alignment during generation. Existing attack

11 Apr 2026

Given the extremely high rate of recidivism, this is important for community safety

SafetyDGX agent

Given the extremely high rate of recidivism, this is important for community safety Murder registry is a good idea from Elon. People should know if they are close to a murderer so proper precautions c

10 Apr 2026

ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis

Model ReleasesDGX agent

arXiv:2604.02022v2 Announce Type: replace Abstract: Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step interactions

Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation

SafetyDGX agent

arXiv:2604.06205v1 Announce Type: cross Abstract: The growth of online platforms and user content requires strong content moderation systems that can handle complex inputs from various media types. Wh

9 Apr 2026

Tesla V14.3 self-driving review. The point releases will bring polish. V15 will far exceed human levels of safety, even in completely unsupe…

SafetyDGX agent

Tesla V14.3 self-driving review. The point releases will bring polish. V15 will far exceed human levels of safety, even in completely unsupervised and complex situations. 600 miles in with FSD v14.3 a

11 Aug 2026

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

SafetyDGX agent

arXiv:2608.09542v1 Announce Type: cross Abstract: Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent

Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?

SafetyDGX agent

arXiv:2608.08077v1 Announce Type: new Abstract: Theory of Space framework (ToS) assesses the spatial understanding of curiosity-driven Vision-Language Models (VLMs) under partial observability. As AI

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

Model ReleasesDGX agent

arXiv:2608.08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, c

7 Aug 2026

From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs

Model ReleasesDGX agent

arXiv:2608.05560v1 Announce Type: cross Abstract: Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evaluations focus on harmful content or general risks, l

6 Aug 2026

Integrated Noise and Safety Management in UAM via A Unified Reinforcement Learning Framework

SafetyDGX agent

arXiv:2508.16440v2 Announce Type: replace-cross Abstract: Urban Air Mobility (UAM) envisions the widespread use of small aerial vehicles to transform transportation in dense urban environments. Howeve

Item Response Theory for AI Safety

Model ReleasesDGX agent

arXiv:2608.05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to tr

5 Aug 2026

Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech

SafetyDGX agent

arXiv:2608.02941v1 Announce Type: new Abstract: We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-C

4 Aug 2026

MonitorVLM-v2: A Deployed Vision-Language Framework for Real-Time Safety Violation Detection

SafetyDGX agent

arXiv:2608.00975v1 Announce Type: new Abstract: Large vision--language models (VLMs) can reason step by step about complex visual scenes, but this open-ended, autoregressive chain-of-thought (CoT) app

30 Jul 2026

A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment

SafetyDGX agent

arXiv:2607.26170v1 Announce Type: new Abstract: This study developed a hybrid computer vision method to quantify exposed skin from images for dermal exposure assessment. Using 170 indoor-painting imag

29 Jul 2026

Safety-Aware Cascaded Inference for Crop Damage Assessment with Controlled Error Trade-offs

SafetyDGX agent

arXiv:2607.25468v1 Announce Type: new Abstract: In picture-based agricultural insurance for smallholder farmers, missed damage detections carry substantially higher cost than false alarms: a farmer wh

28 Jul 2026

SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI

Model ReleasesDGX agent

arXiv:2607.22926v1 Announce Type: new Abstract: High-impact generative AI makes catastrophic misuse a lifecycle-control problem, not merely a prompt-filtering problem. SAGE is a safety-first, authoriz

27 Jul 2026

Tech industry leaders join to form Open Secure AI Alliance to promote safety and security

SafetyDGX agent

Nvidia Corp. today announced the launch of the Open Secure AI Alliance, a new organization founded by technology, cloud computing and cybersecurity leaders to build and share open artificial intellige

23 Jul 2026

Distributed Motion Planning with Safety Guarantees for Self-Reconfiguring Robotic Boats

SafetyDGX agent

arXiv:2607.20352v1 Announce Type: new Abstract: Aquatic self-reconfigurable robots must assemble into desired shapes while ensuring safe interactions among multiple agents. This paper proposes a hybri

NEXUS: Structured Runtime Safety for Tool-Using LLM Agents

Model ReleasesDGX agent

arXiv:2607.19356v1 Announce Type: new Abstract: Tool-using LLM agents increasingly execute high-impact actions, making runtime safety monitoring essential. We present NEXUS (Neural EXecution Utility a

16 Jul 2026

Adversarial Prompting Framework for AI Safety Assessment

SafetyDGX agent

arXiv:2607.13453v1 Announce Type: cross Abstract: Artificial Intelligence (AI), especially Generative AI (GenAI), adoption has increased in industries significantly in recent years. However, the use o

9 Jul 2026

Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

Model ReleasesDGX agent

arXiv:2607.07097v1 Announce Type: new Abstract: Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single 'pipe

When Certificates Fail: A Unified Safety Framework for Embedded Neural Interface Models

SafetyDGX agent

arXiv:2607.06630v1 Announce Type: new Abstract: Formal robustness certificates for embedded neural-interface models can pass while task accuracy collapses: at perturbation budget e=0.25, EEGNet classi

7 Jul 2026

How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs

SafetyDGX agent

arXiv:2607.03561v1 Announce Type: new Abstract: As AI models continue to develop powerful capabilities, it becomes critical that we are able to verify that their output is aligned with our intentions.

3 Jul 2026

Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment

Model ReleasesDGX agent

arXiv:2607.01239v1 Announce Type: cross Abstract: Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable. We identify and test a central structural

HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety

Model ReleasesDGX agent

arXiv:2607.02079v1 Announce Type: new Abstract: We present HaloGuard 1.0, an open-weights implementation of the constitutional-classifier paradigm for input safety. It achieves state-of-the-art perfor

Hardware-Enforced Semantic Coordination for Safety-Critical Real-Time Autonomous Systems

SafetyDGX agent

arXiv:2607.02376v1 Announce Type: new Abstract: Recent advances in agentic AI are producing increasingly complex autonomous systems that integrate large language models, world models, optimization eng

2 Jul 2026

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

Model ReleasesDGX agent

arXiv:2607.01153v1 Announce Type: cross Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an in

EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards

Model ReleasesDGX agent

arXiv:2607.00218v1 Announce Type: cross Abstract: Vision-language models (VLMs) are now proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genu

MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules

Model ReleasesDGX agent

arXiv:2607.00464v1 Announce Type: cross Abstract: Current molecular generation benchmarks emphasize task complexity, molecule novelty, and property alignment; they largely overlook a critical concern:

30 Jun 2026

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models

Model ReleasesDGX agent

arXiv:2501.14940v4 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption. Current LLM safety ben

DriftGuard: Safety-Aware Multi-Monitor Detection and Selective Adaptation for Evolving Toxicity Moderation

Local AiDGX agent

arXiv:2606.28725v1 Announce Type: new Abstract: Automated toxicity moderation systems operate in dynamic online environments where harmful behavior evolves through coded language, shifting targets, an

Safety from Honesty in a Disinterested AI Predictor

SafetyDGX agent

arXiv:2606.29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior th

26 Jun 2026

RedVox: Safety and Fairness Gaps in Speech Models Across Languages

Model ReleasesDGX agent

arXiv:2606.26968v1 Announce Type: new Abstract: Speech-capable models are increasingly deployed in real-world applications across languages. Yet their safety and fairness beyond English settings and u

The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report

Model ReleasesDGX agent

arXiv:2606.26529v1 Announce Type: cross Abstract: AI safety is evaluated by how reliably a model detects the hazards it is told to find, yet accidents often arise from the hazard no one specified. We

25 Jun 2026

The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems

SafetyDGX agent

arXiv:2606.26057v1 Announce Type: cross Abstract: AI agents are granted access to tools, APIs, and other infrastructure, making them active principals in those systems. The dominant approach places co

23 Jun 2026

JPPD: Joint Prediction_Planning Diffusion with Differentiable Safety Guidance for Dynamic Obstacle Avoidance in Intelligent Transportation Systems

SafetyDGX agent

arXiv:2606.20686v1 Announce Type: new Abstract: Shared-space transportation operation requires low-speed autonomous platforms to navigate safely and efficiently among pedestrians, service robots, micr

9 Jun 2026

Brain-Prompt Injection: A Route-Safety Audit for BCI-LLM Agents

SafetyDGX agent

arXiv:2606.09315v1 Announce Type: cross Abstract: BCI-to-agent pipelines turn decoded neural activity into an authorization channel for tool-use agents, exposing a new attack surface we call brain-pro

The Governance of Human-LLM Interaction: Safety Gating, Civility Steering, and Affective Default Lock-In

SafetyDGX agent

arXiv:2606.08172v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly mediate high-stakes interactions in finance, medicine, and mental-health support, yet users have limited con

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

Model ReleasesDGX agent

arXiv:2606.08044v1 Announce Type: cross Abstract: Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, as these eva

5 Jun 2026

Ensuring Interaction Safety in Multitask Exoskeleton Control: A Simulation-Trained Variable Impedance Framework

SafetyDGX agent

arXiv:2606.06370v1 Announce Type: new Abstract: Wearable exoskeletons can augment human phys ical capabilities during complex activities. However, ensuring adaptation across diverse tasks while guaran

4 Jun 2026

NVIDIA Nemotron 3.5 Content Safety Now Available on Vultr

Model ReleasesDGX agent

Deploy NVIDIA Nemotron 3.5 Content Safety on Vultr Cloud GPU with Day Zero support for multimodal AI moderation, custom policy enforcement, multilingual safety workflows, and scalable enterprise AI go

2 Jun 2026

Lost in Delusion: Examining LLM Safety Under User Delusions and Distress

SafetyDGX agent

arXiv:2606.00975v1 Announce Type: new Abstract: LLM chatbots increasingly serve as a first source of support for people in psychological distress, including those whose distress is entangled with delu

SkyShield: Occupancy as a Safety Interface for Low-Altitude UAV Autonomy

Model ReleasesDGX agent

arXiv:2606.00747v1 Announce Type: cross Abstract: For low-altitude Unmanned Aerial Vehicle (UAV) autonomy, 3D spatial understanding is not merely a perception objective, but the safety interface betwe

The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models

Model ReleasesDGX agent

arXiv:2605.05427v2 Announce Type: replace Abstract: Refusal rates are a poor proxy for LLM safety, i.e., a model may over-refuse benign prompts while still complying with harmful ones. We audit both f

1 Jun 2026

From Evidence to Design: Developing an AI-Augmented UX Research Point of View for Digital Wellbeing in Emergency and Public Safety Contexts

SafetyDGX agent

arXiv:2605.31146v1 Announce Type: cross Abstract: This paper investigates how User Experience Research (UXR) methods can be combined with AI-supported analysis to develop clearer design direction for

26 May 2026

D^2-Monitor: Dynamic Safety Monitoring for Diffusion LLMs via Hesitation-Aware Routing

Model ReleasesDGX agent

arXiv:2605.25893v1 Announce Type: new Abstract: Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monitoring

RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry

SafetyDGX agent

arXiv:2605.24817v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures have become an increasingly important paradigm for scaling Large Language Models (LLMs). As MoE models are incr

The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models

Model ReleasesDGX agent

arXiv:2605.25510v1 Announce Type: new Abstract: Children increasingly have access to Large Language Models (LLMs), which may expose them to responses that are developmentally inappropriate or require

25 May 2026

Test-Time Training Undermines Safety Guardrails

SafetyDGX agent

arXiv:2605.22984v1 Announce Type: cross Abstract: Test-Time Training (TTT) is an emerging paradigm that enables models to adapt their parameters during inference, improving performance on tasks such a

23 May 2026

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control

Model ReleasesDGX agent

arXiv:2602.07340v2 Announce Type: replace Abstract: Safety alignment of large language models remains brittle under domain shift and noisy preference supervision. Most existing robust alignment method

20 May 2026

Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South

Local AiDGX agent

arXiv:2605.19190v1 Announce Type: cross Abstract: Despite the global deployment of text-to-image (T2I) models, their safety frameworks are largely calibrated to a Western-centric default, creating sig

Multi-Pedestrian Safety Warning at Urban Intersections Use Case of Digital Twin

Local AiDGX agent

arXiv:2605.18823v1 Announce Type: new Abstract: Digital twins (DTs) for urban transportation systems have gained increasing attention; however, their systematic evaluation in safety-critical scenarios

19 May 2026

Differentiable Optimization Layered Safety-Critical Control for Risk-Aware Navigation via Conformal Prediction

SafetyDGX agent

arXiv:2605.16327v1 Announce Type: cross Abstract: Risk-aware navigation in unknown environments is a fundamental challenge for autonomous vehicles operating in complex urban systems. To address this i

Distributed 3D Leader-Follower Formation Control with Field-of-View Safety via Control Barrier Functions

SafetyDGX agent

arXiv:2605.17533v1 Announce Type: cross Abstract: This letter proposes a distributed 3D leader-follower formation (3D-LFF) control framework for multi-UAV systems that achieves formation tracking whil

DriveSafer: End-to-End Autonomous Driving with Safety Guidance

Model ReleasesDGX agent

arXiv:2605.16737v1 Announce Type: cross Abstract: End-to-End (E2E) autonomous driving models have shown growing capability in recent years, with performance improving on increasingly challenging bench

Multi-task Linear Regression without Eigenvalue Lower Bounds: Adaptivity, Robustness and Safety

SafetyDGX agent

arXiv:2605.17126v1 Announce Type: cross Abstract: We study the multi-task linear regression problem in the presence of contaminated tasks. We address the setting where the unknown parameters of a majo

Responsible Federated LLMs via Safety Filtering and Constitutional AI

SafetyDGX agent

arXiv:2502.16691v2 Announce Type: replace Abstract: Recent research has increasingly focused on training large language models (LLMs) using federated learning, known as FedLLM. However, responsible AI

← Previous
1…7891011…238
Next →