AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
19 May 2026

UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages

Model ReleasesDGX agent

arXiv:2601.12696v2 Announce Type: replace Abstract: Current guardian models are predominantly Western-centric and optimized for high-resource languages, leaving low-resource African languages vulnerab

18 May 2026

Whole-body motion planning and safety-critical control for aerial manipulation

SafetyDGX agent

arXiv:2511.02342v3 Announce Type: replace Abstract: Aerial manipulation combines the maneuverability of multirotors with the dexterity of robotic arms to perform complex tasks in cluttered spaces. Yet

14 May 2026

Quantifying LLM Safety Degradation Under Repeated Attacks Using Survival Analysis

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
SafetyDGX agent

arXiv:2605.12869v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in a wide range of applications, yet remain vulnerable to adversarial jailbreak attacks that ci

12 May 2026

Embodied AI in Action: Insights from SAE World Congress 2026 on Safety, Trust, Robotics, and Real-World Deployment

SafetyDGX agent

arXiv:2605.10653v1 Announce Type: new Abstract: Embodied artificial intelligence is rapidly moving from research into real-world systems such as autonomous vehicles, mobile robots, and industrial mach

NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims

Model ReleasesDGX agent

arXiv:2605.08192v1 Announce Type: cross Abstract: Frontier AI safety claims - published assertions that a highly capable general-purpose model is below a threshold of concern, adequately mitigated, or

The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play

Model ReleasesDGX agent

arXiv:2605.08427v1 Announce Type: new Abstract: Self-play red team is an established approach to improving AI safety in which different instances of the same model play attacker and defender roles in

11 May 2026

Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

Model ReleasesDGX agent

arXiv:2605.07630v1 Announce Type: cross Abstract: When a phone-use agent avoids harm, does that show safety, or simply inability to act? Existing evaluations often cannot tell. A harmful outcome may b

THINKSAFE: Self-Generated Safety Alignment for Reasoning Models

Model ReleasesDGX agent

arXiv:2601.23143v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieve remarkable performance by leveraging reinforcement learning (RL) on reasoning tasks to generate long chain-of-

7 May 2026

Enhancing Agent Safety Judgment: Controlled Benchmark Rewriting and Analogical Reasoning for Deceptive Out-of-Distribution Scenarios

Model ReleasesDGX agent

arXiv:2605.03242v1 Announce Type: new Abstract: Tool-using agent systems powered by large language models (LLMs) are increasingly deployed across web, app, operating-system, and transactional environm

Safety Must Precede the Deployment of Open-Ended AI

SafetyDGX agent

arXiv:2502.04512v3 Announce Type: replace Abstract: AI advancements have been significantly driven by a combination of foundation models and curiosity-driven learning aimed at increasing capability an

6 May 2026

Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment

SafetyDGX agent

arXiv:2605.01899v1 Announce Type: new Abstract: The growing capabilities of large language models (LLMs) have driven their widespread deployment across diverse domains, even in potentially high-risk s

5 May 2026

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety

Model ReleasesDGX agent

arXiv:2605.01687v1 Announce Type: new Abstract: We present MultiBreak, a scalable and diverse multi-turn jailbreak benchmark to evaluate large language model (LLM) safety. Multi-turn jailbreaks mimic

TAIL-Safe: Task-Agnostic Safety Monitoring for Imitation Learning Policies

SafetyDGX agent

arXiv:2605.01195v1 Announce Type: new Abstract: Recent imitation learning (IL) algorithms such as flow-matching and diffusion policies demonstrate remarkable performance in learning complex manipulati

30 Apr 2026

One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety

SafetyDGX agent

arXiv:2604.25921v1 Announce Type: new Abstract: Large Language Models (LLMs) are trained to refuse harmful requests, yet they remain vulnerable to jailbreak attacks that exploit weaknesses in conversa

29 Apr 2026

Evaluating LLM Safety Under Repeated Inference via Accelerated Prompt Stress Testing

Model ReleasesDGX agent

arXiv:2602.11786v2 Announce Type: replace Abstract: Traditional benchmarks for large language models (LLMs), such as HELM and AIR-BENCH, primarily assess safety through breadth-oriented evaluation acr

28 Apr 2026

SkyfireAI lands $11M to bring AI autonomy to public safety and defense drones

SafetyDGX agent

Autonomous drone startup SkyfireAI Inc. today announced that it has raised 11 million in new funding to accelerate the development of its dual-use, artificial intelligence-native platform for autonomo

23 Apr 2026

Soft-Label Governance for Distributional Safety in Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2604.19752v1 Announce Type: cross Abstract: Multi-agent AI systems exhibit emergent risks that no single agent produces in isolation. Existing safety frameworks rely on binary classifications of

22 Apr 2026

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety

SafetyDGX agent

arXiv:2604.19083v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in cross-modal understanding and generation, yet their deployment is threate

21 Apr 2026

Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence

SafetyDGX agent

arXiv:2601.11886v2 Announce Type: replace Abstract: In high-stakes domains like medicine, it may be generally desirable for models to faithfully adhere to the context provided. But what happens if the

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction

SafetyDGX agent

arXiv:2506.01770v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have achieved tremendous success in various tasks, yet concerns about their safety and security have emerged. In

15 Apr 2026

AISafetyBenchExplorer: A Metric-Aware Catalogue of AI Safety Benchmarks Reveals Fragmented Measurement and Weak Benchmark Governance

Model ReleasesDGX agent

arXiv:2604.12875v1 Announce Type: new Abstract: The rapid expansion of large language model (LLM) safety evaluation has produced a substantial benchmark ecosystem, but not a correspondingly coherent m

Is Vibe Coding the Future? An Empirical Assessment of LLM Generated Codes for Construction Safety

Model ReleasesDGX agent

arXiv:2604.12311v1 Announce Type: cross Abstract: The emergence of vibe coding, a paradigm where non-technical users instruct Large Language Models (LLMs) to generate executable codes via natural lang

14 Apr 2026

Decision-Theoretic Safety Assessment of Persona-Driven Multi-Agent Systems in O-RAN

SafetyDGX agent

arXiv:2604.09682v1 Announce Type: cross Abstract: Autonomous network management in Open Radio Access Networks requires intelligent decision making across conflicting objectives, yet existing LLM based

Evaluating Reliability Gaps in Large Language Model Safety via Repeated Prompt Sampling

Model ReleasesDGX agent

arXiv:2604.09606v1 Announce Type: new Abstract: Traditional benchmarks for large language models (LLMs), such as HELM and AIR-BENCH, primarily assess safety risk through breadth-oriented evaluation ac

The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents

Model ReleasesDGX agent

arXiv:2604.10577v1 Announce Type: cross Abstract: Computer-use agents (CUAs) can now autonomously complete complex tasks in real digital environments, but when misled, they can also be used to automat

13 Apr 2026

Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations

SafetyDGX agent

arXiv:2506.09067v2 Announce Type: replace-cross Abstract: Generative medical vision-language models~(Med-VLMs) are primarily designed to generate complex textual information~(e.g., diagnostic reports)

12 Aug 2026

Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning

SafetyDGX agent

arXiv:2608.10403v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due

11 Aug 2026

From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs

SafetyDGX agent

arXiv:2608.09158v1 Announce Type: cross Abstract: Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency

Multi-Agent AI Safety as an Institutional Design Problem

SafetyDGX agent

arXiv:2608.09828v1 Announce Type: cross Abstract: AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent wo

7 Aug 2026

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

Model ReleasesDGX agent

arXiv:2608.06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet m

6 Aug 2026

Social Pressure Breaks Majority Voting in LLM Safety Panels

SafetyDGX agent

arXiv:2608.04415v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct

Traceable LLM-Generated Hazard Scenarios for Operational Safety Analysis of Aviation Systems Using ASRS Reports

SafetyDGX agent

arXiv:2608.04697v1 Announce Type: new Abstract: Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, aircraft operatio

4 Aug 2026

ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification

Model ReleasesDGX agent

arXiv:2608.01291v1 Announce Type: new Abstract: We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syrian, E

SafeBuild-Bench: A Temporal-Robust Construction Safety Benchmark with Graph-Enhanced Data Mining

Model ReleasesDGX agent

arXiv:2608.00068v1 Announce Type: new Abstract: Construction-safety models must handle concrete deployment risks, such as a worker standing near a scaffold edge without guardrails, rather than only re

28 Jul 2026

Diffusion-Guided Search via Exponential Tilting (DiffTilt): An Application to Falsification of Safety-Critical Systems

Model ReleasesDGX agent

arXiv:2607.23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation. Existing f

27 Jul 2026

Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security

SafetyDGX agent

Open source software is a critical pillar of the global economy. It underpins cloud computing, financial services, manufacturing, telecommunications, government and internet services by making technol

24 Jul 2026

Safety-oriented sidewalk and road segmentation for smartphone-based assistive navigation

SafetyDGX agent

arXiv:2607.21137v1 Announce Type: new Abstract: Independent sidewalk mobility is essential for blind and visually impaired pedestrians (BVIPs), yet smartphone-based assistive navigation requires perce

15 Jul 2026

Data Safety: Synthetic Data Quality Analysis Using CIFAKE Dataset

SafetyDGX agent

arXiv:2607.12165v1 Announce Type: new Abstract: Recently, the societal implementation of high-performance image classification models has expanded rapidly. While these models require vast amounts of t

9 Jul 2026

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

Model ReleasesDGX agent

arXiv:2602.05088v4 Announce Type: replace Abstract: Millions of people now use generative AI chatbots for psychological support. Despite their promise, the most pressing question in AI for mental heal

7 Jul 2026

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale

Model ReleasesDGX agent

arXiv:2607.02714v1 Announce Type: cross Abstract: There is no doubt that safety alignment is an essential step in LLM training. However, conceptually it does not distinguish between various domains an

5 Jul 2026

Chemical accidents rise as Trump administration proposes weakening safety rules

SafetyDGX agent

Chemical accidents involving releases of dangerous chemicals rose 57 percent between 2021 and 2025, from 83 to 131 incidents , with injuries or deaths rising from 60 to 89 over the same period . The E

3 Jul 2026

Neuro-Symbolic Safety Guidance for Vision-Language-Action Models via Constrained Flow Matching

Model ReleasesDGX agent

arXiv:2607.01378v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated promising generalization capabilities across robotic manipulation tasks, yet their real-world depl

Online Safety Monitoring for LLMs

SafetyDGX agent

arXiv:2607.02510v1 Announce Type: new Abstract: Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safet

1 Jul 2026

Off the Rails: Hijacking the Scoring Head in Generative End-to-End Driving Planners with Safety-Violating Adversarial Perturbations

SafetyDGX agent

arXiv:2606.30807v1 Announce Type: cross Abstract: Generative models have recently seen rapid adoption in End-to-End (E2E) autonomous driving (AD), with diffusion-based denoising and vocabulary-based r

OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation

Model ReleasesDGX agent

arXiv:2606.31993v1 Announce Type: new Abstract: While robotic manipulation capabilities have advanced rapidly, physical safety remains a major barrier to deploying household robots: task success is in

Revealing Safety-Critical Scenarios for UTM via Transformer

SafetyDGX agent

arXiv:2606.31114v1 Announce Type: new Abstract: Unmanned Traffic Management (UTM) systems are cloud-based platforms designed to manage and coordinate multiple aerial vehicles remotely. UTM systems are

26 Jun 2026

Necessary but Not Sufficient: Temperature Control and Reproducibility in LLM-as-Judge Safety Evaluations

Model ReleasesDGX agent

arXiv:2606.26185v1 Announce Type: new Abstract: LLM-as-judge ('grader') components are now standard in evaluation harnesses, including safety evaluations where a pass/fail verdict may gate downstream

25 Jun 2026

Do Thinking Tokens Help with Safety?

Model ReleasesDGX agent

arXiv:2606.25013v1 Announce Type: cross Abstract: Today's reasoning models use thinking tokens to attain stronger performance on benchmarks than their instruction-tuned counterparts. It is also genera

24 Jun 2026

Safe and Generalizable Hierarchical Multi-Agent RL via Constraint Manifold Control

SafetyDGX agent

arXiv:2606.24010v1 Announce Type: new Abstract: Multi-agent systems are widely used in safety-critical applications that require coordinated behavior under strict safety constraints. Existing approach

23 Jun 2026

Enhancing Road Safety: An IoT-Based Accident Detection and Prevention Mechanism

SafetyDGX agent

arXiv:2606.22381v1 Announce Type: cross Abstract: Road traffic accidents remain a critical global crisis, consistently serving as a primary driver of preventable mortality and severe injury. These inc

11 Jun 2026

SAFER-Nav: Enhancing Safety for Visual Robot Navigation via Segmentation-Aware Fine-Tuning

SafetyDGX agent

arXiv:2606.11636v1 Announce Type: new Abstract: Vision-based navigation models, particularly foundation models, generate viable trajectories from RGB observations alone. However, even state-of-the-art

10 Jun 2026

Anthropic secretly limiting Claude's usefulness for LLM development strengthens the argument that Anthropic is using AI safety to justify monopolistic behavior (Dean W. Ball/@deanwball)

Model ReleasesDGX agent

Dean W. Ball / @deanwball: Anthropic secretly limiting Claude's usefulness for LLM development strengthens the argument that Anthropic is using AI safety to justify monopolistic behavior — My last obs

9 Jun 2026

Enhancing AI Interpretability and Safety through Localised Architectures

SafetyDGX agent

arXiv:2606.07998v1 Announce Type: cross Abstract: Recent advances in generative AI, especially powerful Large Language Models (LLMs) and Large Reasoning Models (LRMs), raise concerns over the interpre

The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection

Model ReleasesDGX agent

arXiv:2606.09204v1 Announce Type: new Abstract: We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injection Paradox -- in which prompt injections embedde

8 Jun 2026

Apple refines Liquid Glass design and expands child safety tools at WWDC

SafetyDGX agent

Apple Inc.’s WWDC 2026 keynote today focused on making its software platform feel more polished, more responsive and more tightly managed across the company’s device ecosystem with refinements to the

6 Jun 2026

PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage

Model ReleasesDGX agent

arXiv:2606.05463v1 Announce Type: new Abstract: Patient safety event triage, determining whether a clinical event is reportable under jurisdiction-specific policy, is a high-stakes task typically perf

2 Jun 2026

Low-Resource Safety Failures Are Action Failures, Not Representation Failures

Model ReleasesDGX agent

arXiv:2606.01196v1 Announce Type: cross Abstract: Safety alignment learned in high-resource languages transfers poorly to low-resource languages. Models refuse harmful prompts in English but fail to r

29 May 2026

Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures

SafetyDGX agent

arXiv:2605.29629v1 Announce Type: new Abstract: Attack Success Rate (ASR) evaluates each jailbreak with a single yes/no label at the end of generation, telling us whether a failure happened but not ho

Beyond English and Evasion: A Human-Annotated Multi-Domain Benchmark for High-Stakes LLM Safety Evaluation in Chinese

Model ReleasesDGX agent

arXiv:2605.29667v1 Announce Type: new Abstract: When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break

28 May 2026

ICAN-Deploy: Identity-Stable Canary Deployment for Safety-Critical Embodied Agents

SafetyDGX agent

arXiv:2605.28097v1 Announce Type: new Abstract: Canary deployment routes a fraction of traffic to a new software version, monitors metrics, and rolls back on regression. Mainstream controllers (Argo R

← Previous
1…89101112…238
Next →