AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Model Releases

D^2-Monitor: Dynamic Safety Monitoring for Diffusion LLMs via Hesitation-Aware Routing

DGX agent

arXiv:2605.25893v1 Announce Type: new Abstract: Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monitoring

model-releasesarxiv-cs-ai
26 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry

DGX agent

arXiv:2605.24817v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures have become an increasingly important paradigm for scaling Large Language Models (LLMs). As MoE models are incr

safetyarxiv-cs-cl
26 May 2026
Model Releases

The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models

DGX agent

arXiv:2605.25510v1 Announce Type: new Abstract: Children increasingly have access to Large Language Models (LLMs), which may expose them to responses that are developmentally inappropriate or require

model-releasesarxiv-cs-cl
26 May 2026
Safety

Test-Time Training Undermines Safety Guardrails

DGX agent

arXiv:2605.22984v1 Announce Type: cross Abstract: Test-Time Training (TTT) is an emerging paradigm that enables models to adapt their parameters during inference, improving performance on tasks such a

safetyarxiv-cs-ai
25 May 2026
Model Releases

Revisiting Robustness for LLM Safety Alignment via Selective Geometry Control

DGX agent

arXiv:2602.07340v2 Announce Type: replace Abstract: Safety alignment of large language models remains brittle under domain shift and noisy preference supervision. Most existing robust alignment method

model-releasesarxiv-cs-lg
23 May 2026
Local Ai

Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South

DGX agent

arXiv:2605.19190v1 Announce Type: cross Abstract: Despite the global deployment of text-to-image (T2I) models, their safety frameworks are largely calibrated to a Western-centric default, creating sig

local-aiarxiv-cs-ai
20 May 2026
Local Ai

Multi-Pedestrian Safety Warning at Urban Intersections Use Case of Digital Twin

DGX agent

arXiv:2605.18823v1 Announce Type: new Abstract: Digital twins (DTs) for urban transportation systems have gained increasing attention; however, their systematic evaluation in safety-critical scenarios

local-aiarxiv-cs-lg
20 May 2026
Safety

Differentiable Optimization Layered Safety-Critical Control for Risk-Aware Navigation via Conformal Prediction

DGX agent

arXiv:2605.16327v1 Announce Type: cross Abstract: Risk-aware navigation in unknown environments is a fundamental challenge for autonomous vehicles operating in complex urban systems. To address this i

safetyarxiv-cs-ai
19 May 2026
Safety

Distributed 3D Leader-Follower Formation Control with Field-of-View Safety via Control Barrier Functions

DGX agent

arXiv:2605.17533v1 Announce Type: cross Abstract: This letter proposes a distributed 3D leader-follower formation (3D-LFF) control framework for multi-UAV systems that achieves formation tracking whil

safetyarxiv-cs-ro
19 May 2026
Model Releases

DriveSafer: End-to-End Autonomous Driving with Safety Guidance

DGX agent

arXiv:2605.16737v1 Announce Type: cross Abstract: End-to-End (E2E) autonomous driving models have shown growing capability in recent years, with performance improving on increasingly challenging bench

model-releasesarxiv-cs-cv
19 May 2026
Safety

Multi-task Linear Regression without Eigenvalue Lower Bounds: Adaptivity, Robustness and Safety

DGX agent

arXiv:2605.17126v1 Announce Type: cross Abstract: We study the multi-task linear regression problem in the presence of contaminated tasks. We address the setting where the unknown parameters of a majo

safetyarxiv-cs-lg
19 May 2026
Safety

Responsible Federated LLMs via Safety Filtering and Constitutional AI

DGX agent

arXiv:2502.16691v2 Announce Type: replace Abstract: Recent research has increasingly focused on training large language models (LLMs) using federated learning, known as FedLLM. However, responsible AI

safetyarxiv-cs-cl
19 May 2026
Model Releases

UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages

DGX agent

arXiv:2601.12696v2 Announce Type: replace Abstract: Current guardian models are predominantly Western-centric and optimized for high-resource languages, leaving low-resource African languages vulnerab

model-releasesarxiv-cs-cl
19 May 2026
Safety

Whole-body motion planning and safety-critical control for aerial manipulation

DGX agent

arXiv:2511.02342v3 Announce Type: replace Abstract: Aerial manipulation combines the maneuverability of multirotors with the dexterity of robotic arms to perform complex tasks in cluttered spaces. Yet

safetyarxiv-cs-ro
18 May 2026
Safety

Quantifying LLM Safety Degradation Under Repeated Attacks Using Survival Analysis

DGX agent

arXiv:2605.12869v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in a wide range of applications, yet remain vulnerable to adversarial jailbreak attacks that ci

safetyarxiv-cs-ai
14 May 2026
Safety

Embodied AI in Action: Insights from SAE World Congress 2026 on Safety, Trust, Robotics, and Real-World Deployment

DGX agent

arXiv:2605.10653v1 Announce Type: new Abstract: Embodied artificial intelligence is rapidly moving from research into real-world systems such as autonomous vehicles, mobile robots, and industrial mach

safetyarxiv-cs-ro
12 May 2026
Model Releases

NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims

DGX agent

arXiv:2605.08192v1 Announce Type: cross Abstract: Frontier AI safety claims - published assertions that a highly capable general-purpose model is below a threshold of concern, adequately mitigated, or

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play

DGX agent

arXiv:2605.08427v1 Announce Type: new Abstract: Self-play red team is an established approach to improving AI safety in which different instances of the same model play attacker and defender roles in

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

DGX agent

arXiv:2605.07630v1 Announce Type: cross Abstract: When a phone-use agent avoids harm, does that show safety, or simply inability to act? Existing evaluations often cannot tell. A harmful outcome may b

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

THINKSAFE: Self-Generated Safety Alignment for Reasoning Models

DGX agent

arXiv:2601.23143v2 Announce Type: replace Abstract: Large reasoning models (LRMs) achieve remarkable performance by leveraging reinforcement learning (RL) on reasoning tasks to generate long chain-of-

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Enhancing Agent Safety Judgment: Controlled Benchmark Rewriting and Analogical Reasoning for Deceptive Out-of-Distribution Scenarios

DGX agent

arXiv:2605.03242v1 Announce Type: new Abstract: Tool-using agent systems powered by large language models (LLMs) are increasingly deployed across web, app, operating-system, and transactional environm

model-releasesarxiv-cs-ai
7 May 2026
Safety

Safety Must Precede the Deployment of Open-Ended AI

DGX agent

arXiv:2502.04512v3 Announce Type: replace Abstract: AI advancements have been significantly driven by a combination of foundation models and curiosity-driven learning aimed at increasing capability an

safetyarxiv-cs-ai
7 May 2026
Safety

Disentangling Intent from Role: Adversarial Self-Play for Persona-Invariant Safety Alignment

DGX agent

arXiv:2605.01899v1 Announce Type: new Abstract: The growing capabilities of large language models (LLMs) have driven their widespread deployment across diverse domains, even in potentially high-risk s

safetyarxiv-cs-ai
6 May 2026
Model Releases

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety

DGX agent

arXiv:2605.01687v1 Announce Type: new Abstract: We present MultiBreak, a scalable and diverse multi-turn jailbreak benchmark to evaluate large language model (LLM) safety. Multi-turn jailbreaks mimic

model-releasesarxiv-cs-cl
5 May 2026
Safety

TAIL-Safe: Task-Agnostic Safety Monitoring for Imitation Learning Policies

DGX agent

arXiv:2605.01195v1 Announce Type: new Abstract: Recent imitation learning (IL) algorithms such as flow-matching and diffusion policies demonstrate remarkable performance in learning complex manipulati

safetyarxiv-cs-ro
5 May 2026
Safety

One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety

DGX agent

arXiv:2604.25921v1 Announce Type: new Abstract: Large Language Models (LLMs) are trained to refuse harmful requests, yet they remain vulnerable to jailbreak attacks that exploit weaknesses in conversa

safetyarxiv-cs-cl
30 Apr 2026
Model Releases

Evaluating LLM Safety Under Repeated Inference via Accelerated Prompt Stress Testing

DGX agent

arXiv:2602.11786v2 Announce Type: replace Abstract: Traditional benchmarks for large language models (LLMs), such as HELM and AIR-BENCH, primarily assess safety through breadth-oriented evaluation acr

model-releasesarxiv-cs-lg
29 Apr 2026
Safety

SkyfireAI lands $11M to bring AI autonomy to public safety and defense drones

DGX agent

Autonomous drone startup SkyfireAI Inc. today announced that it has raised 11 million in new funding to accelerate the development of its dual-use, artificial intelligence-native platform for autonomo

safetysiliconangle
28 Apr 2026
Model Releases

Soft-Label Governance for Distributional Safety in Multi-Agent Systems

DGX agent

arXiv:2604.19752v1 Announce Type: cross Abstract: Multi-agent AI systems exhibit emergent risks that no single agent produces in isolation. Existing safety frameworks rely on binary classifications of

model-releasesarxiv-cs-ai
23 Apr 2026
Safety

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety

DGX agent

arXiv:2604.19083v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in cross-modal understanding and generation, yet their deployment is threate

safetyarxiv-cs-ai
22 Apr 2026
Safety

Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence

DGX agent

arXiv:2601.11886v2 Announce Type: replace Abstract: In high-stakes domains like medicine, it may be generally desirable for models to faithfully adhere to the context provided. But what happens if the

safetyarxiv-cs-cl
21 Apr 2026
Safety

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction

DGX agent

arXiv:2506.01770v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have achieved tremendous success in various tasks, yet concerns about their safety and security have emerged. In

safetyarxiv-cs-lg
21 Apr 2026
Model Releases

AISafetyBenchExplorer: A Metric-Aware Catalogue of AI Safety Benchmarks Reveals Fragmented Measurement and Weak Benchmark Governance

DGX agent

arXiv:2604.12875v1 Announce Type: new Abstract: The rapid expansion of large language model (LLM) safety evaluation has produced a substantial benchmark ecosystem, but not a correspondingly coherent m

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

Is Vibe Coding the Future? An Empirical Assessment of LLM Generated Codes for Construction Safety

DGX agent

arXiv:2604.12311v1 Announce Type: cross Abstract: The emergence of vibe coding, a paradigm where non-technical users instruct Large Language Models (LLMs) to generate executable codes via natural lang

model-releasesarxiv-cs-ai
15 Apr 2026
Safety

Decision-Theoretic Safety Assessment of Persona-Driven Multi-Agent Systems in O-RAN

DGX agent

arXiv:2604.09682v1 Announce Type: cross Abstract: Autonomous network management in Open Radio Access Networks requires intelligent decision making across conflicting objectives, yet existing LLM based

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

Evaluating Reliability Gaps in Large Language Model Safety via Repeated Prompt Sampling

DGX agent

arXiv:2604.09606v1 Announce Type: new Abstract: Traditional benchmarks for large language models (LLMs), such as HELM and AIR-BENCH, primarily assess safety risk through breadth-oriented evaluation ac

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents

DGX agent

arXiv:2604.10577v1 Announce Type: cross Abstract: Computer-use agents (CUAs) can now autonomously complete complex tasks in real digital environments, but when misled, they can also be used to automat

model-releasesarxiv-cs-ai
14 Apr 2026
Safety

Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations

DGX agent

arXiv:2506.09067v2 Announce Type: replace-cross Abstract: Generative medical vision-language models~(Med-VLMs) are primarily designed to generate complex textual information~(e.g., diagnostic reports)

safetyarxiv-cs-ai
13 Apr 2026
Safety

Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning

DGX agent

arXiv:2608.10403v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due

safetyarxiv-cs-ai
12 Aug 2026
Safety

From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs

DGX agent

arXiv:2608.09158v1 Announce Type: cross Abstract: Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency

safetyarxiv-cs-ai
11 Aug 2026
Safety

Multi-Agent AI Safety as an Institutional Design Problem

DGX agent

arXiv:2608.09828v1 Announce Type: cross Abstract: AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent wo

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

DGX agent

arXiv:2608.06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet m

model-releasesarxiv-cs-ai
7 Aug 2026
Safety

Social Pressure Breaks Majority Voting in LLM Safety Panels

DGX agent

arXiv:2608.04415v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct

safetyarxiv-cs-cl
6 Aug 2026
Safety

Traceable LLM-Generated Hazard Scenarios for Operational Safety Analysis of Aviation Systems Using ASRS Reports

DGX agent

arXiv:2608.04697v1 Announce Type: new Abstract: Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, aircraft operatio

safetyarxiv-cs-ai
6 Aug 2026
Model Releases

ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification

DGX agent

arXiv:2608.01291v1 Announce Type: new Abstract: We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syrian, E

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

SafeBuild-Bench: A Temporal-Robust Construction Safety Benchmark with Graph-Enhanced Data Mining

DGX agent

arXiv:2608.00068v1 Announce Type: new Abstract: Construction-safety models must handle concrete deployment risks, such as a worker standing near a scaffold edge without guardrails, rather than only re

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Diffusion-Guided Search via Exponential Tilting (DiffTilt): An Application to Falsification of Safety-Critical Systems

DGX agent

arXiv:2607.23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation. Existing f

model-releasesarxiv-cs-lg
28 Jul 2026
Safety

Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security

DGX agent

Open source software is a critical pillar of the global economy. It underpins cloud computing, financial services, manufacturing, telecommunications, government and internet services by making technol

safetynvidia-blog
27 Jul 2026
← Previous
1…1011121314…297
Next →