AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Model Releases

Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

DGX agent

arXiv:2607.07097v1 Announce Type: new Abstract: Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single 'pipe

model-releasesarxiv-cs-ai
9 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

When Certificates Fail: A Unified Safety Framework for Embedded Neural Interface Models

DGX agent

arXiv:2607.06630v1 Announce Type: new Abstract: Formal robustness certificates for embedded neural-interface models can pass while task accuracy collapses: at perturbation budget e=0.25, EEGNet classi

safetyarxiv-cs-lg
9 Jul 2026
Safety

How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs

DGX agent

arXiv:2607.03561v1 Announce Type: new Abstract: As AI models continue to develop powerful capabilities, it becomes critical that we are able to verify that their output is aligned with our intentions.

safetyarxiv-cs-ai
7 Jul 2026
Model Releases

Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment

DGX agent

arXiv:2607.01239v1 Announce Type: cross Abstract: Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable. We identify and test a central structural

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safety

DGX agent

arXiv:2607.02079v1 Announce Type: new Abstract: We present HaloGuard 1.0, an open-weights implementation of the constitutional-classifier paradigm for input safety. It achieves state-of-the-art perfor

model-releasesarxiv-cs-cl
3 Jul 2026
Safety

Hardware-Enforced Semantic Coordination for Safety-Critical Real-Time Autonomous Systems

DGX agent

arXiv:2607.02376v1 Announce Type: new Abstract: Recent advances in agentic AI are producing increasingly complex autonomous systems that integrate large language models, world models, optimization eng

safetyarxiv-cs-ai
3 Jul 2026
Model Releases

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

DGX agent

arXiv:2607.01153v1 Announce Type: cross Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an in

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards

DGX agent

arXiv:2607.00218v1 Announce Type: cross Abstract: Vision-language models (VLMs) are now proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genu

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules

DGX agent

arXiv:2607.00464v1 Announce Type: cross Abstract: Current molecular generation benchmarks emphasize task complexity, molecule novelty, and property alignment; they largely overlook a critical concern:

model-releasesarxiv-cs-cl
2 Jul 2026
Model Releases

CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models

DGX agent

arXiv:2501.14940v4 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption. Current LLM safety ben

model-releasesarxiv-cs-ai
30 Jun 2026
Local Ai

DriftGuard: Safety-Aware Multi-Monitor Detection and Selective Adaptation for Evolving Toxicity Moderation

DGX agent

arXiv:2606.28725v1 Announce Type: new Abstract: Automated toxicity moderation systems operate in dynamic online environments where harmful behavior evolves through coded language, shifting targets, an

local-aiarxiv-cs-cl
30 Jun 2026
Safety

Safety from Honesty in a Disinterested AI Predictor

DGX agent

arXiv:2606.29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior th

safetyarxiv-cs-ai
30 Jun 2026
Model Releases

RedVox: Safety and Fairness Gaps in Speech Models Across Languages

DGX agent

arXiv:2606.26968v1 Announce Type: new Abstract: Speech-capable models are increasingly deployed in real-world applications across languages. Yet their safety and fairness beyond English settings and u

model-releasesarxiv-cs-cl
26 Jun 2026
Model Releases

The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report

DGX agent

arXiv:2606.26529v1 Announce Type: cross Abstract: AI safety is evaluated by how reliably a model detects the hazards it is told to find, yet accidents often arise from the hazard no one specified. We

model-releasesarxiv-cs-ai
26 Jun 2026
Safety

The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems

DGX agent

arXiv:2606.26057v1 Announce Type: cross Abstract: AI agents are granted access to tools, APIs, and other infrastructure, making them active principals in those systems. The dominant approach places co

safetyarxiv-cs-lg
25 Jun 2026
Safety

JPPD: Joint Prediction_Planning Diffusion with Differentiable Safety Guidance for Dynamic Obstacle Avoidance in Intelligent Transportation Systems

DGX agent

arXiv:2606.20686v1 Announce Type: new Abstract: Shared-space transportation operation requires low-speed autonomous platforms to navigate safely and efficiently among pedestrians, service robots, micr

safetyarxiv-cs-ro
23 Jun 2026
Safety

Brain-Prompt Injection: A Route-Safety Audit for BCI-LLM Agents

DGX agent

arXiv:2606.09315v1 Announce Type: cross Abstract: BCI-to-agent pipelines turn decoded neural activity into an authorization channel for tool-use agents, exposing a new attack surface we call brain-pro

safetyarxiv-cs-ai
9 Jun 2026
Safety

The Governance of Human-LLM Interaction: Safety Gating, Civility Steering, and Affective Default Lock-In

DGX agent

arXiv:2606.08172v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly mediate high-stakes interactions in finance, medicine, and mental-health support, yet users have limited con

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

DGX agent

arXiv:2606.08044v1 Announce Type: cross Abstract: Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, as these eva

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Ensuring Interaction Safety in Multitask Exoskeleton Control: A Simulation-Trained Variable Impedance Framework

DGX agent

arXiv:2606.06370v1 Announce Type: new Abstract: Wearable exoskeletons can augment human phys ical capabilities during complex activities. However, ensuring adaptation across diverse tasks while guaran

safetyarxiv-cs-ro
5 Jun 2026
Safety

Lost in Delusion: Examining LLM Safety Under User Delusions and Distress

DGX agent

arXiv:2606.00975v1 Announce Type: new Abstract: LLM chatbots increasingly serve as a first source of support for people in psychological distress, including those whose distress is entangled with delu

safetyarxiv-cs-cl
2 Jun 2026
Model Releases

SkyShield: Occupancy as a Safety Interface for Low-Altitude UAV Autonomy

DGX agent

arXiv:2606.00747v1 Announce Type: cross Abstract: For low-altitude Unmanned Aerial Vehicle (UAV) autonomy, 3D spatial understanding is not merely a perception objective, but the safety interface betwe

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models

DGX agent

arXiv:2605.05427v2 Announce Type: replace Abstract: Refusal rates are a poor proxy for LLM safety, i.e., a model may over-refuse benign prompts while still complying with harmful ones. We audit both f

model-releasesarxiv-cs-ai
2 Jun 2026
Safety

From Evidence to Design: Developing an AI-Augmented UX Research Point of View for Digital Wellbeing in Emergency and Public Safety Contexts

DGX agent

arXiv:2605.31146v1 Announce Type: cross Abstract: This paper investigates how User Experience Research (UXR) methods can be combined with AI-supported analysis to develop clearer design direction for

safetyarxiv-cs-ai
1 Jun 2026
Model Releases

D^2-Monitor: Dynamic Safety Monitoring for Diffusion LLMs via Hesitation-Aware Routing

DGX agent

arXiv:2605.25893v1 Announce Type: new Abstract: Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monitoring

model-releasesarxiv-cs-ai
26 May 2026
Safety

RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry

DGX agent

arXiv:2605.24817v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures have become an increasingly important paradigm for scaling Large Language Models (LLMs). As MoE models are incr

safetyarxiv-cs-cl
26 May 2026
Model Releases

The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models

DGX agent

arXiv:2605.25510v1 Announce Type: new Abstract: Children increasingly have access to Large Language Models (LLMs), which may expose them to responses that are developmentally inappropriate or require

model-releasesarxiv-cs-cl
26 May 2026
Safety

Test-Time Training Undermines Safety Guardrails

DGX agent

arXiv:2605.22984v1 Announce Type: cross Abstract: Test-Time Training (TTT) is an emerging paradigm that enables models to adapt their parameters during inference, improving performance on tasks such a

safetyarxiv-cs-ai
25 May 2026
Model Releases

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control

DGX agent

arXiv:2602.07340v2 Announce Type: replace Abstract: Safety alignment of large language models remains brittle under domain shift and noisy preference supervision. Most existing robust alignment method

model-releasesarxiv-cs-lg
23 May 2026
Local Ai

Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South

DGX agent

arXiv:2605.19190v1 Announce Type: cross Abstract: Despite the global deployment of text-to-image (T2I) models, their safety frameworks are largely calibrated to a Western-centric default, creating sig

local-aiarxiv-cs-ai
20 May 2026
Local Ai

Multi-Pedestrian Safety Warning at Urban Intersections Use Case of Digital Twin

DGX agent

arXiv:2605.18823v1 Announce Type: new Abstract: Digital twins (DTs) for urban transportation systems have gained increasing attention; however, their systematic evaluation in safety-critical scenarios

local-aiarxiv-cs-lg
20 May 2026
Safety

Differentiable Optimization Layered Safety-Critical Control for Risk-Aware Navigation via Conformal Prediction

DGX agent

arXiv:2605.16327v1 Announce Type: cross Abstract: Risk-aware navigation in unknown environments is a fundamental challenge for autonomous vehicles operating in complex urban systems. To address this i

safetyarxiv-cs-ai
19 May 2026
Safety

Distributed 3D Leader-Follower Formation Control with Field-of-View Safety via Control Barrier Functions

DGX agent

arXiv:2605.17533v1 Announce Type: cross Abstract: This letter proposes a distributed 3D leader-follower formation (3D-LFF) control framework for multi-UAV systems that achieves formation tracking whil

safetyarxiv-cs-ro
19 May 2026
Model Releases

DriveSafer: End-to-End Autonomous Driving with Safety Guidance

DGX agent

arXiv:2605.16737v1 Announce Type: cross Abstract: End-to-End (E2E) autonomous driving models have shown growing capability in recent years, with performance improving on increasingly challenging bench

model-releasesarxiv-cs-cv
19 May 2026
Safety

Multi-task Linear Regression without Eigenvalue Lower Bounds: Adaptivity, Robustness and Safety

DGX agent

arXiv:2605.17126v1 Announce Type: cross Abstract: We study the multi-task linear regression problem in the presence of contaminated tasks. We address the setting where the unknown parameters of a majo

safetyarxiv-cs-lg
19 May 2026
Safety

Responsible Federated LLMs via Safety Filtering and Constitutional AI

DGX agent

arXiv:2502.16691v2 Announce Type: replace Abstract: Recent research has increasingly focused on training large language models (LLMs) using federated learning, known as FedLLM. However, responsible AI

safetyarxiv-cs-cl
19 May 2026
Model Releases

UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages

DGX agent

arXiv:2601.12696v2 Announce Type: replace Abstract: Current guardian models are predominantly Western-centric and optimized for high-resource languages, leaving low-resource African languages vulnerab

model-releasesarxiv-cs-cl
19 May 2026
Safety

Whole-body motion planning and safety-critical control for aerial manipulation

DGX agent

arXiv:2511.02342v3 Announce Type: replace Abstract: Aerial manipulation combines the maneuverability of multirotors with the dexterity of robotic arms to perform complex tasks in cluttered spaces. Yet

safetyarxiv-cs-ro
18 May 2026
Safety

Quantifying LLM Safety Degradation Under Repeated Attacks Using Survival Analysis

DGX agent

arXiv:2605.12869v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in a wide range of applications, yet remain vulnerable to adversarial jailbreak attacks that ci

safetyarxiv-cs-ai
14 May 2026
Safety

Embodied AI in Action: Insights from SAE World Congress 2026 on Safety, Trust, Robotics, and Real-World Deployment

DGX agent

arXiv:2605.10653v1 Announce Type: new Abstract: Embodied artificial intelligence is rapidly moving from research into real-world systems such as autonomous vehicles, mobile robots, and industrial mach

safetyarxiv-cs-ro
12 May 2026
Model Releases

NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims

DGX agent

arXiv:2605.08192v1 Announce Type: cross Abstract: Frontier AI safety claims - published assertions that a highly capable general-purpose model is below a threshold of concern, adequately mitigated, or

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play

DGX agent

arXiv:2605.08427v1 Announce Type: new Abstract: Self-play red team is an established approach to improving AI safety in which different instances of the same model play attacker and defender roles in

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

DGX agent

arXiv:2605.07630v1 Announce Type: cross Abstract: When a phone-use agent avoids harm, does that show safety, or simply inability to act? Existing evaluations often cannot tell. A harmful outcome may b

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

THINKSAFE: Self-Generated Safety Alignment for Reasoning Models

DGX agent

arXiv:2601.23143v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieve remarkable performance by leveraging reinforcement learning (RL) on reasoning tasks to generate long chain-of-

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Enhancing Agent Safety Judgment: Controlled Benchmark Rewriting and Analogical Reasoning for Deceptive Out-of-Distribution Scenarios

DGX agent

arXiv:2605.03242v1 Announce Type: new Abstract: Tool-using agent systems powered by large language models (LLMs) are increasingly deployed across web, app, operating-system, and transactional environm

model-releasesarxiv-cs-ai
7 May 2026
Safety

Safety Must Precede the Deployment of Open-Ended AI

DGX agent

arXiv:2502.04512v3 Announce Type: replace Abstract: AI advancements have been significantly driven by a combination of foundation models and curiosity-driven learning aimed at increasing capability an

safetyarxiv-cs-ai
7 May 2026
Safety

Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment

DGX agent

arXiv:2605.01899v1 Announce Type: new Abstract: The growing capabilities of large language models (LLMs) have driven their widespread deployment across diverse domains, even in potentially high-risk s

safetyarxiv-cs-ai
6 May 2026
Model Releases

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety

DGX agent

arXiv:2605.01687v1 Announce Type: new Abstract: We present MultiBreak, a scalable and diverse multi-turn jailbreak benchmark to evaluate large language model (LLM) safety. Multi-turn jailbreaks mimic

model-releasesarxiv-cs-cl
5 May 2026
← Previous
1…89101112…255
Next →