AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlog
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,234 results
Safety

Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems

DGX agent

arXiv:2510.14133v2 Announce Type: replace Abstract: Agentic AI systems, which leverage multiple autonomous agents and large language models (LLMs), are increasingly used to address complex, multi-step

safetyarxiv-cs-ai
17 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design

DGX agent

arXiv:2604.12500v1 Announce Type: new Abstract: Specification gaming under Reinforcement Learning (RL) is known to cause LLMs to develop sycophantic, manipulative, or deceptive behavior, yet the condi

safetyarxiv-cs-lg
15 Apr 2026
Safety

Anthropic believes that good transparency legislation needs to ensure public safety and accountability for the companies developing this pow…

DGX agent

Anthropic believes that good transparency legislation needs to ensure public safety and accountability for the companies developing this powerful technology, not provide a get-out-of-jail-free card ag

safetygary-marcus--x
14 Apr 2026
Model Releases

Seeing No Evil: Blinding Large Vision-Language Models to Safety Instructions via Adversarial Attention Hijacking

DGX agent

arXiv:2604.10299v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) rely on attention-based retrieval of safety instructions to maintain alignment during generation. Existing attack

model-releasesarxiv-cs-cl
14 Apr 2026
Safety

Given the extremely high rate of recidivism, this is important for community safety

DGX agent

Given the extremely high rate of recidivism, this is important for community safety Murder registry is a good idea from Elon. People should know if they are close to a murderer so proper precautions c

safetyelon-musk--x
11 Apr 2026
Model Releases

ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis

DGX agent

arXiv:2604.02022v2 Announce Type: replace Abstract: Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step interactions

model-releasesarxiv-cs-ai
10 Apr 2026
Safety

Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation

DGX agent

arXiv:2604.06205v1 Announce Type: cross Abstract: The growth of online platforms and user content requires strong content moderation systems that can handle complex inputs from various media types. Wh

safetyarxiv-cs-ai
10 Apr 2026
Safety

Tesla V14.3 self-driving review. The point releases will bring polish. V15 will far exceed human levels of safety, even in completely unsupe…

DGX agent

Tesla V14.3 self-driving review. The point releases will bring polish. V15 will far exceed human levels of safety, even in completely unsupervised and complex situations. 600 miles in with FSD v14.3 a

safetyelon-musk--x
9 Apr 2026
Safety

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

DGX agent

arXiv:2608.09542v1 Announce Type: cross Abstract: Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent

safetyarxiv-cs-ai
11 Aug 2026
Safety

Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?

DGX agent

arXiv:2608.08077v1 Announce Type: new Abstract: Theory of Space framework (ToS) assesses the spatial understanding of curiosity-driven Vision-Language Models (VLMs) under partial observability. As AI

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs

DGX agent

arXiv:2608.08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, c

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs

DGX agent

arXiv:2608.05560v1 Announce Type: cross Abstract: Timely anticipation of physical hazards is essential for real-world safety, yet existing MLLM evaluations focus on harmful content or general risks, l

model-releasesarxiv-cs-cl
7 Aug 2026
Safety

Integrated Noise and Safety Management in UAM via A Unified Reinforcement Learning Framework

DGX agent

arXiv:2508.16440v2 Announce Type: replace-cross Abstract: Urban Air Mobility (UAM) envisions the widespread use of small aerial vehicles to transform transportation in dense urban environments. Howeve

safetyarxiv-cs-lg
6 Aug 2026
Model Releases

Item Response Theory for AI Safety

DGX agent

arXiv:2608.05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to tr

model-releasesarxiv-cs-ai
6 Aug 2026
Safety

Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech

DGX agent

arXiv:2608.02941v1 Announce Type: new Abstract: We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-C

safetyarxiv-cs-cl
5 Aug 2026
Safety

MonitorVLM-v2: A Deployed Vision-Language Framework for Real-Time Safety Violation Detection

DGX agent

arXiv:2608.00975v1 Announce Type: new Abstract: Large vision--language models (VLMs) can reason step by step about complex visual scenes, but this open-ended, autoregressive chain-of-thought (CoT) app

safetyarxiv-cs-cv
4 Aug 2026
Safety

A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment

DGX agent

arXiv:2607.26170v1 Announce Type: new Abstract: This study developed a hybrid computer vision method to quantify exposed skin from images for dermal exposure assessment. Using 170 indoor-painting imag

safetyarxiv-cs-cv
30 Jul 2026
Safety

Safety-Aware Cascaded Inference for Crop Damage Assessment with Controlled Error Trade-offs

DGX agent

arXiv:2607.25468v1 Announce Type: new Abstract: In picture-based agricultural insurance for smallholder farmers, missed damage detections carry substantially higher cost than false alarms: a farmer wh

safetyarxiv-cs-cv
29 Jul 2026
Model Releases

SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI

DGX agent

arXiv:2607.22926v1 Announce Type: new Abstract: High-impact generative AI makes catastrophic misuse a lifecycle-control problem, not merely a prompt-filtering problem. SAGE is a safety-first, authoriz

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

Tech industry leaders join to form Open Secure AI Alliance to promote safety and security

DGX agent

Nvidia Corp. today announced the launch of the Open Secure AI Alliance, a new organization founded by technology, cloud computing and cybersecurity leaders to build and share open artificial intellige

safetysiliconangle
27 Jul 2026
Safety

Distributed Motion Planning with Safety Guarantees for Self-Reconfiguring Robotic Boats

DGX agent

arXiv:2607.20352v1 Announce Type: new Abstract: Aquatic self-reconfigurable robots must assemble into desired shapes while ensuring safe interactions among multiple agents. This paper proposes a hybri

safetyarxiv-cs-ro
23 Jul 2026
Model Releases

NEXUS: Structured Runtime Safety for Tool-Using LLM Agents

DGX agent

arXiv:2607.19356v1 Announce Type: new Abstract: Tool-using LLM agents increasingly execute high-impact actions, making runtime safety monitoring essential. We present NEXUS (Neural EXecution Utility a

model-releasesarxiv-cs-ai
23 Jul 2026
Safety

Adversarial Prompting Framework for AI Safety Assessment

DGX agent

arXiv:2607.13453v1 Announce Type: cross Abstract: Artificial Intelligence (AI), especially Generative AI (GenAI), adoption has increased in industries significantly in recent years. However, the use o

safetyarxiv-cs-ai
16 Jul 2026
Model Releases

Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

DGX agent

arXiv:2607.07097v1 Announce Type: new Abstract: Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single 'pipe

model-releasesarxiv-cs-ai
9 Jul 2026
Safety

When Certificates Fail: A Unified Safety Framework for Embedded Neural Interface Models

DGX agent

arXiv:2607.06630v1 Announce Type: new Abstract: Formal robustness certificates for embedded neural-interface models can pass while task accuracy collapses: at perturbation budget e=0.25, EEGNet classi

safetyarxiv-cs-lg
9 Jul 2026
Safety

How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs

DGX agent

arXiv:2607.03561v1 Announce Type: new Abstract: As AI models continue to develop powerful capabilities, it becomes critical that we are able to verify that their output is aligned with our intentions.

safetyarxiv-cs-ai
7 Jul 2026
Model Releases

Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment

DGX agent

arXiv:2607.01239v1 Announce Type: cross Abstract: Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable. We identify and test a central structural

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety

DGX agent

arXiv:2607.02079v1 Announce Type: new Abstract: We present HaloGuard 1.0, an open-weights implementation of the constitutional-classifier paradigm for input safety. It achieves state-of-the-art perfor

model-releasesarxiv-cs-cl
3 Jul 2026
Safety

Hardware-Enforced Semantic Coordination for Safety-Critical Real-Time Autonomous Systems

DGX agent

arXiv:2607.02376v1 Announce Type: new Abstract: Recent advances in agentic AI are producing increasingly complex autonomous systems that integrate large language models, world models, optimization eng

safetyarxiv-cs-ai
3 Jul 2026
Model Releases

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

DGX agent

arXiv:2607.01153v1 Announce Type: cross Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an in

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards

DGX agent

arXiv:2607.00218v1 Announce Type: cross Abstract: Vision-language models (VLMs) are now proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genu

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules

DGX agent

arXiv:2607.00464v1 Announce Type: cross Abstract: Current molecular generation benchmarks emphasize task complexity, molecule novelty, and property alignment; they largely overlook a critical concern:

model-releasesarxiv-cs-cl
2 Jul 2026
Model Releases

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models

DGX agent

arXiv:2501.14940v4 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption. Current LLM safety ben

model-releasesarxiv-cs-ai
30 Jun 2026
Local Ai

DriftGuard: Safety-Aware Multi-Monitor Detection and Selective Adaptation for Evolving Toxicity Moderation

DGX agent

arXiv:2606.28725v1 Announce Type: new Abstract: Automated toxicity moderation systems operate in dynamic online environments where harmful behavior evolves through coded language, shifting targets, an

local-aiarxiv-cs-cl
30 Jun 2026
Safety

Safety from Honesty in a Disinterested AI Predictor

DGX agent

arXiv:2606.29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior th

safetyarxiv-cs-ai
30 Jun 2026
Model Releases

RedVox: Safety and Fairness Gaps in Speech Models Across Languages

DGX agent

arXiv:2606.26968v1 Announce Type: new Abstract: Speech-capable models are increasingly deployed in real-world applications across languages. Yet their safety and fairness beyond English settings and u

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report

DGX agent

arXiv:2606.26529v1 Announce Type: cross Abstract: AI safety is evaluated by how reliably a model detects the hazards it is told to find, yet accidents often arise from the hazard no one specified. We

model-releasesarxiv-cs-ai
26 Jun 2026
Safety

The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems

DGX agent

arXiv:2606.26057v1 Announce Type: cross Abstract: AI agents are granted access to tools, APIs, and other infrastructure, making them active principals in those systems. The dominant approach places co

safetyarxiv-cs-lg
25 Jun 2026
Safety

JPPD: Joint Prediction_Planning Diffusion with Differentiable Safety Guidance for Dynamic Obstacle Avoidance in Intelligent Transportation Systems

DGX agent

arXiv:2606.20686v1 Announce Type: new Abstract: Shared-space transportation operation requires low-speed autonomous platforms to navigate safely and efficiently among pedestrians, service robots, micr

safetyarxiv-cs-ro
23 Jun 2026
Safety

Brain-Prompt Injection: A Route-Safety Audit for BCI-LLM Agents

DGX agent

arXiv:2606.09315v1 Announce Type: cross Abstract: BCI-to-agent pipelines turn decoded neural activity into an authorization channel for tool-use agents, exposing a new attack surface we call brain-pro

safetyarxiv-cs-ai
9 Jun 2026
Safety

The Governance of Human-LLM Interaction: Safety Gating, Civility Steering, and Affective Default Lock-In

DGX agent

arXiv:2606.08172v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly mediate high-stakes interactions in finance, medicine, and mental-health support, yet users have limited con

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

DGX agent

arXiv:2606.08044v1 Announce Type: cross Abstract: Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, as these eva

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Ensuring Interaction Safety in Multitask Exoskeleton Control: A Simulation-Trained Variable Impedance Framework

DGX agent

arXiv:2606.06370v1 Announce Type: new Abstract: Wearable exoskeletons can augment human phys ical capabilities during complex activities. However, ensuring adaptation across diverse tasks while guaran

safetyarxiv-cs-ro
5 Jun 2026
Model Releases

NVIDIA Nemotron 3.5 Content Safety Now Available on Vultr

DGX agent

Deploy NVIDIA Nemotron 3.5 Content Safety on Vultr Cloud GPU with Day Zero support for multimodal AI moderation, custom policy enforcement, multilingual safety workflows, and scalable enterprise AI go

model-releasesvultr
4 Jun 2026
Safety

Lost in Delusion: Examining LLM Safety Under User Delusions and Distress

DGX agent

arXiv:2606.00975v1 Announce Type: new Abstract: LLM chatbots increasingly serve as a first source of support for people in psychological distress, including those whose distress is entangled with delu

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

SkyShield: Occupancy as a Safety Interface for Low-Altitude UAV Autonomy

DGX agent

arXiv:2606.00747v1 Announce Type: cross Abstract: For low-altitude Unmanned Aerial Vehicle (UAV) autonomy, 3D spatial understanding is not merely a perception objective, but the safety interface betwe

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models

DGX agent

arXiv:2605.05427v2 Announce Type: replace Abstract: Refusal rates are a poor proxy for LLM safety, i.e., a model may over-refuse benign prompts while still complying with harmful ones. We audit both f

model-releasesarxiv-cs-ai
2 Jun 2026
Safety

From Evidence to Design: Developing an AI-Augmented UX Research Point of View for Digital Wellbeing in Emergency and Public Safety Contexts

DGX agent

arXiv:2605.31146v1 Announce Type: cross Abstract: This paper investigates how User Experience Research (UXR) methods can be combined with AI-supported analysis to develop clearer design direction for

safetyarxiv-cs-ai
1 Jun 2026
← Previous
1…910111213…297
Next →