AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Safety

TAIL-Safe: Task-Agnostic Safety Monitoring for Imitation Learning Policies

DGX agent

arXiv:2605.01195v1 Announce Type: new Abstract: Recent imitation learning (IL) algorithms such as flow-matching and diffusion policies demonstrate remarkable performance in learning complex manipulati

safetyarxiv-cs-ro
5 May 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety

DGX agent

arXiv:2604.25921v1 Announce Type: new Abstract: Large Language Models (LLMs) are trained to refuse harmful requests, yet they remain vulnerable to jailbreak attacks that exploit weaknesses in conversa

safetyarxiv-cs-cl
30 Apr 2026
Model Releases

Evaluating LLM Safety Under Repeated Inference via Accelerated Prompt Stress Testing

DGX agent

arXiv:2602.11786v2 Announce Type: replace Abstract: Traditional benchmarks for large language models (LLMs), such as HELM and AIR-BENCH, primarily assess safety through breadth-oriented evaluation acr

model-releasesarxiv-cs-lg
29 Apr 2026
Model Releases

Soft-Label Governance for Distributional Safety in Multi-Agent Systems

DGX agent

arXiv:2604.19752v1 Announce Type: cross Abstract: Multi-agent AI systems exhibit emergent risks that no single agent produces in isolation. Existing safety frameworks rely on binary classifications of

model-releasesarxiv-cs-ai
23 Apr 2026
Safety

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety

DGX agent

arXiv:2604.19083v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in cross-modal understanding and generation, yet their deployment is threate

safetyarxiv-cs-ai
22 Apr 2026
Safety

Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence

DGX agent

arXiv:2601.11886v2 Announce Type: replace Abstract: In high-stakes domains like medicine, it may be generally desirable for models to faithfully adhere to the context provided. But what happens if the

safetyarxiv-cs-cl
21 Apr 2026
Safety

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction

DGX agent

arXiv:2506.01770v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have achieved tremendous success in various tasks, yet concerns about their safety and security have emerged. In

safetyarxiv-cs-lg
21 Apr 2026
Model Releases

AISafetyBenchExplorer: A Metric-Aware Catalogue of AI Safety Benchmarks Reveals Fragmented Measurement and Weak Benchmark Governance

DGX agent

arXiv:2604.12875v1 Announce Type: new Abstract: The rapid expansion of large language model (LLM) safety evaluation has produced a substantial benchmark ecosystem, but not a correspondingly coherent m

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

Is Vibe Coding the Future? An Empirical Assessment of LLM Generated Codes for Construction Safety

DGX agent

arXiv:2604.12311v1 Announce Type: cross Abstract: The emergence of vibe coding, a paradigm where non-technical users instruct Large Language Models (LLMs) to generate executable codes via natural lang

model-releasesarxiv-cs-ai
15 Apr 2026
Safety

Decision-Theoretic Safety Assessment of Persona-Driven Multi-Agent Systems in O-RAN

DGX agent

arXiv:2604.09682v1 Announce Type: cross Abstract: Autonomous network management in Open Radio Access Networks requires intelligent decision making across conflicting objectives, yet existing LLM based

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

Evaluating Reliability Gaps in Large Language Model Safety via Repeated Prompt Sampling

DGX agent

arXiv:2604.09606v1 Announce Type: new Abstract: Traditional benchmarks for large language models (LLMs), such as HELM and AIR-BENCH, primarily assess safety risk through breadth-oriented evaluation ac

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents

DGX agent

arXiv:2604.10577v1 Announce Type: cross Abstract: Computer-use agents (CUAs) can now autonomously complete complex tasks in real digital environments, but when misled, they can also be used to automat

model-releasesarxiv-cs-ai
14 Apr 2026
Safety

Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations

DGX agent

arXiv:2506.09067v2 Announce Type: replace-cross Abstract: Generative medical vision-language models~(Med-VLMs) are primarily designed to generate complex textual information~(e.g., diagnostic reports)

safetyarxiv-cs-ai
13 Apr 2026
Safety

Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning

DGX agent

arXiv:2608.10403v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due

safetyarxiv-cs-ai
12 Aug 2026
Safety

From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs

DGX agent

arXiv:2608.09158v1 Announce Type: cross Abstract: Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency

safetyarxiv-cs-ai
11 Aug 2026
Safety

Multi-Agent AI Safety as an Institutional Design Problem

DGX agent

arXiv:2608.09828v1 Announce Type: cross Abstract: AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent wo

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

DGX agent

arXiv:2608.06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet m

model-releasesarxiv-cs-ai
7 Aug 2026
Safety

Social Pressure Breaks Majority Voting in LLM Safety Panels

DGX agent

arXiv:2608.04415v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to detect unsafe content. A common approach is to combine judgments from a panel of models to correct

safetyarxiv-cs-cl
6 Aug 2026
Safety

Traceable LLM-Generated Hazard Scenarios for Operational Safety Analysis of Aviation Systems Using ASRS Reports

DGX agent

arXiv:2608.04697v1 Announce Type: new Abstract: Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, aircraft operatio

safetyarxiv-cs-ai
6 Aug 2026
Model Releases

ArabicDialectSafety: A Dialect-Aware Benchmark for Arabic Content Safety Classification

DGX agent

arXiv:2608.01291v1 Announce Type: new Abstract: We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syrian, E

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

SafeBuild-Bench: A Temporal-Robust Construction Safety Benchmark with Graph-Enhanced Data Mining

DGX agent

arXiv:2608.00068v1 Announce Type: new Abstract: Construction-safety models must handle concrete deployment risks, such as a worker standing near a scaffold edge without guardrails, rather than only re

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Diffusion-Guided Search via Exponential Tilting (DiffTilt): An Application to Falsification of Safety-Critical Systems

DGX agent

arXiv:2607.23134v1 Announce Type: new Abstract: Discovering rare safety-critical failures in autonomous and cyber-physical systems is a fundamental challenge in verification and validation. Existing f

model-releasesarxiv-cs-lg
28 Jul 2026
Safety

Safety-oriented sidewalk and road segmentation for smartphone-based assistive navigation

DGX agent

arXiv:2607.21137v1 Announce Type: new Abstract: Independent sidewalk mobility is essential for blind and visually impaired pedestrians (BVIPs), yet smartphone-based assistive navigation requires perce

safetyarxiv-cs-cv
24 Jul 2026
Safety

Data Safety: Synthetic Data Quality Analysis Using CIFAKE Dataset

DGX agent

arXiv:2607.12165v1 Announce Type: new Abstract: Recently, the societal implementation of high-performance image classification models has expanded rapidly. While these models require vast amounts of t

safetyarxiv-cs-cv
15 Jul 2026
Model Releases

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

DGX agent

arXiv:2602.05088v4 Announce Type: replace Abstract: Millions of people now use generative AI chatbots for psychological support. Despite their promise, the most pressing question in AI for mental heal

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale

DGX agent

arXiv:2607.02714v1 Announce Type: cross Abstract: There is no doubt that safety alignment is an essential step in LLM training. However, conceptually it does not distinguish between various domains an

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Neuro-Symbolic Safety Guidance for Vision-Language-Action Models via Constrained Flow Matching

DGX agent

arXiv:2607.01378v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated promising generalization capabilities across robotic manipulation tasks, yet their real-world depl

model-releasesarxiv-cs-ro
3 Jul 2026
Safety

Online Safety Monitoring for LLMs

DGX agent

arXiv:2607.02510v1 Announce Type: new Abstract: Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safet

safetyarxiv-cs-ai
3 Jul 2026
Safety

Off the Rails: Hijacking the Scoring Head in Generative End-to-End Driving Planners with Safety-Violating Adversarial Perturbations

DGX agent

arXiv:2606.30807v1 Announce Type: cross Abstract: Generative models have recently seen rapid adoption in End-to-End (E2E) autonomous driving (AD), with diffusion-based denoising and vocabulary-based r

safetyarxiv-cs-cv
1 Jul 2026
Model Releases

OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation

DGX agent

arXiv:2606.31993v1 Announce Type: new Abstract: While robotic manipulation capabilities have advanced rapidly, physical safety remains a major barrier to deploying household robots: task success is in

model-releasesarxiv-cs-ro
1 Jul 2026
Safety

Revealing Safety-Critical Scenarios for UTM via Transformer

DGX agent

arXiv:2606.31114v1 Announce Type: new Abstract: Unmanned Traffic Management (UTM) systems are cloud-based platforms designed to manage and coordinate multiple aerial vehicles remotely. UTM systems are

safetyarxiv-cs-ai
1 Jul 2026
Model Releases

Necessary but Not Sufficient: Temperature Control and Reproducibility in LLM-as-Judge Safety Evaluations

DGX agent

arXiv:2606.26185v1 Announce Type: new Abstract: LLM-as-judge ('grader') components are now standard in evaluation harnesses, including safety evaluations where a pass/fail verdict may gate downstream

model-releasesarxiv-cs-lg
26 Jun 2026
Model Releases

Do Thinking Tokens Help with Safety?

DGX agent

arXiv:2606.25013v1 Announce Type: cross Abstract: Today's reasoning models use thinking tokens to attain stronger performance on benchmarks than their instruction-tuned counterparts. It is also genera

model-releasesarxiv-cs-cl
25 Jun 2026
Safety

Safe and Generalizable Hierarchical Multi-Agent RL via Constraint Manifold Control

DGX agent

arXiv:2606.24010v1 Announce Type: new Abstract: Multi-agent systems are widely used in safety-critical applications that require coordinated behavior under strict safety constraints. Existing approach

safetyarxiv-cs-ai
24 Jun 2026
Safety

Enhancing Road Safety: An IoT-Based Accident Detection and Prevention Mechanism

DGX agent

arXiv:2606.22381v1 Announce Type: cross Abstract: Road traffic accidents remain a critical global crisis, consistently serving as a primary driver of preventable mortality and severe injury. These inc

safetyarxiv-cs-cv
23 Jun 2026
Safety

SAFER-Nav: Enhancing Safety for Visual Robot Navigation via Segmentation-Aware Fine-Tuning

DGX agent

arXiv:2606.11636v1 Announce Type: new Abstract: Vision-based navigation models, particularly foundation models, generate viable trajectories from RGB observations alone. However, even state-of-the-art

safetyarxiv-cs-ro
11 Jun 2026
Safety

Enhancing AI Interpretability and Safety through Localised Architectures

DGX agent

arXiv:2606.07998v1 Announce Type: cross Abstract: Recent advances in generative AI, especially powerful Large Language Models (LLMs) and Large Reasoning Models (LRMs), raise concerns over the interpre

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection

DGX agent

arXiv:2606.09204v1 Announce Type: new Abstract: We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injection Paradox -- in which prompt injections embedde

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage

DGX agent

arXiv:2606.05463v1 Announce Type: new Abstract: Patient safety event triage, determining whether a clinical event is reportable under jurisdiction-specific policy, is a high-stakes task typically perf

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Low-Resource Safety Failures Are Action Failures, Not Representation Failures

DGX agent

arXiv:2606.01196v1 Announce Type: cross Abstract: Safety alignment learned in high-resource languages transfers poorly to low-resource languages. Models refuse harmful prompts in English but fail to r

model-releasesarxiv-cs-ai
2 Jun 2026
Safety

Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures

DGX agent

arXiv:2605.29629v1 Announce Type: new Abstract: Attack Success Rate (ASR) evaluates each jailbreak with a single yes/no label at the end of generation, telling us whether a failure happened but not ho

safetyarxiv-cs-ai
29 May 2026
Model Releases

Beyond English and Evasion: A Human-Annotated Multi-Domain Benchmark for High-Stakes LLM Safety Evaluation in Chinese

DGX agent

arXiv:2605.29667v1 Announce Type: new Abstract: When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break

model-releasesarxiv-cs-cl
29 May 2026
Safety

ICAN-Deploy: Identity-Stable Canary Deployment for Safety-Critical Embodied Agents

DGX agent

arXiv:2605.28097v1 Announce Type: new Abstract: Canary deployment routes a fraction of traffic to a new software version, monitors metrics, and rolls back on regression. Mainstream controllers (Argo R

safetyarxiv-cs-ro
28 May 2026
Safety

VLM-Based Advanced Rider Assistance System for Motorcycle Safety

DGX agent

arXiv:2605.27948v1 Announce Type: new Abstract: Motorcycles face disproportionately high crash risks compared to cars due to limited protection and heightened sensitivity to surface hazards, yet Advan

safetyarxiv-cs-ro
28 May 2026
Safety

FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents

DGX agent

arXiv:2605.27333v1 Announce Type: new Abstract: Finance LLM agents must simultaneously block prompt-induced unauthorized actions and approve legitimate multi-step business workflows. However, boundary

safetyarxiv-cs-cl
27 May 2026
Safety

Learning to Balance Motor Thermal Safety and Quadrupedal Locomotion Performance with Residual Policy

DGX agent

arXiv:2605.27046v1 Announce Type: new Abstract: Motor thermal management is often overlooked in the context of electrically-actuated robots, particularly legged robots, but motor overheating is a key

safetyarxiv-cs-ro
27 May 2026
Safety

Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts

DGX agent

arXiv:2605.24270v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) language models activate only a small subset of parameters for each token, making router behavior a central part of mode

safetyarxiv-cs-ai
26 May 2026
Safety

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

DGX agent

arXiv:2605.23989v1 Announce Type: new Abstract: Agentic AI systems -- Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions -- can execute complex tasks

safetyarxiv-cs-ai
26 May 2026
← Previous
1…910111213…255
Next →