AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Safety

When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent Systems

DGX agent

arXiv:2605.01133v1 Announce Type: cross Abstract: Large language model (LLM)-powered multi-agent systems (MAS) enable agents to communicate and share information, achieving strong performance on compl

safetyarxiv-cs-lg
5 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex

DGX agent

arXiv:2604.14858v2 Announce Type: replace Abstract: As agent systems move into increasingly diverse execution settings, trajectory-level safety evaluation and diagnosis require benchmarks that evolve

model-releasesarxiv-cs-ai
30 Apr 2026
Safety

Instantaneous Planning, Control and Safety for Navigation in Unknown Underwater Spaces

DGX agent

arXiv:2604.05310v2 Announce Type: replace Abstract: Navigating autonomous underwater vehicles (AUVs) in unknown environments is significantly challenging due to poor visibility, weak signal transmissi

safetyarxiv-cs-ro
29 Apr 2026
Model Releases

Evaluating whether AI models would sabotage AI safety research

DGX agent

arXiv:2604.24618v1 Announce Type: new Abstract: We evaluate the propensity of frontier models to sabotage or refuse to assist with safety research when deployed as AI research agents within a frontier

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

How Sensitive Are Safety Benchmarks to Judge Configuration Choices?

DGX agent

arXiv:2604.24074v1 Announce Type: new Abstract: Safety benchmarks such as HarmBench rely on LLM judges to classify model responses as harmful or safe, yet the judge configuration, namely the combinati

model-releasesarxiv-cs-cl
28 Apr 2026
Model Releases

Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models

DGX agent

arXiv:2603.21697v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) extend text-only LLMs with visual reasoning, but also introduce new safety failure modes under visual

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety

DGX agent

arXiv:2604.18487v1 Announce Type: new Abstract: The Adversarial Humanities Benchmark (AHB) evaluates whether model safety refusals survive a shift away from familiar harmful prompt forms. Starting fro

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation

DGX agent

arXiv:2602.07954v4 Announce Type: replace Abstract: As Large Language Models (LLMs) become increasingly deployed in Polish language applications, the need for efficient and accurate content safety cla

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration

DGX agent

arXiv:2604.16541v1 Announce Type: new Abstract: Recent advancements in Large Generative Models (LGMs) have revolutionized multi-modal generation. However, generating illustrated storybooks remains an

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

Structure-Aware Diversity Pursuit as an AI Safety Strategy against Homogenization

DGX agent

arXiv:2601.06116v2 Announce Type: replace-cross Abstract: Generative AI models reproduce the biases in the training data and can further amplify them through mode collapse. We refer to the resulting h

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

DGX agent

arXiv:2604.06173v1 Announce Type: cross Abstract: Legal QA benchmarks have predominantly focused on case law, overlooking the unique challenges of statute-centric regulatory reasoning. In statutory do

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding

DGX agent

arXiv:2604.07879v1 Announce Type: new Abstract: Diffusion-based image generation models have advanced rapidly but pose a safety risk due to their potential to generate Not-Safe-For-Work (NSFW) content

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity

DGX agent

arXiv:2608.09443v1 Announce Type: new Abstract: Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbidity depend on

model-releasesarxiv-cs-ai
11 Aug 2026
Safety

Learning Multi-Timescale Interventions under Safety and Resource Constraints

DGX agent

arXiv:2508.03875v2 Announce Type: replace Abstract: Many sequential decision problems offer qualitatively different ways of influencing the environment: some interventions act immediately, whereas oth

safetyarxiv-cs-lg
11 Aug 2026
Model Releases

Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

DGX agent

arXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection

DGX agent

arXiv:2509.23690v2 Announce Type: replace-cross Abstract: Safety hazards in the home are a leading cause of preventable domestic injuries, motivating an automated inspector that actively explores a ho

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models

DGX agent

arXiv:2608.03201v1 Announce Type: new Abstract: Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs. We audit

model-releasesarxiv-cs-ai
5 Aug 2026
Safety

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard)

DGX agent

arXiv:2608.00180v1 Announce Type: new Abstract: Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard. Training a safety guard with RL means optimizing two

safetyarxiv-cs-cl
4 Aug 2026
Model Releases

🛡️Introducing Shieldstral, Mistral’s 3B open-weights model for content safety that can be deployed on-device 🧵 http://mistral.ai/news/shie…

DGX agent

Mistral AI introduced Shieldstral, a 3‑billion‑parameter, open‑weights model designed for content‑safety tasks and capable of on‑device deployment. The announcement was shared via a tweet from @Mistra

model-releasesmistral-ai--x
4 Aug 2026
Safety

It’s time to panic about AI safety

DGX agent

When the phrase 'OpenAI hacked Hugging Face' has more or less entered mainstream culture, you know we have an AI problem. This week, we learned more about exactly how OpenAI's agent broke out of a san

safetythe-verge-ai
31 Jul 2026
Safety

Safety Verification of Wait-Only Non-Blocking Broadcast Protocols

DGX agent

arXiv:2403.18591v3 Announce Type: replace-cross Abstract: Broadcast protocols are programs designed to be executed by networks of processes. Each process runs the same protocol, and communication betw

safetyarxiv-cs-cl
31 Jul 2026
Model Releases

EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

DGX agent

arXiv:2607.26518v1 Announce Type: new Abstract: Reliable visual safety understanding in real-world scenarios demands more than just object recognition; it requires causal reasoning under epistemic unc

model-releasesarxiv-cs-cv
30 Jul 2026
Safety

We’re running out of reasons to ignore AI safety

DGX agent

Earlier this month, OpenAI gave several of its AI models a task: complete a test designed to measure their cybersecurity capabilities. It put the systems in a sandboxed environment without an internet

safetythe-verge-ai
29 Jul 2026
Local Ai

Distribution-Specific Curvature Control with Finite-Sample Guarantees for Open-Weight Safety

DGX agent

arXiv:2607.22929v1 Announce Type: new Abstract: A short fine-tuning run can undo the safety guards of an open-weight model---retraining a refusal-trained assistant to aid weapons development or produc

local-aiarxiv-cs-lg
28 Jul 2026
Model Releases

OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows

DGX agent

arXiv:2510.24411v3 Announce Type: replace Abstract: Computer-using agents powered by Vision-Language Models (VLMs) have demonstrated human-like capabilities in operating digital environments like mobi

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

HERMES: Heterogeneous Edge-Relational Multi-Head Embedded SSM Attention for Traffic Conflict Prediction at Signalized Intersections

DGX agent

arXiv:2607.20505v1 Announce Type: cross Abstract: Surrogate safety measures (SSMs) enable proactive traffic safety assessment, but many existing methods evaluate pairwise interactions independently or

safetyarxiv-cs-lg
24 Jul 2026
Model Releases

Geometry-Guided Constraint Learning for LLM Safety Classification

DGX agent

arXiv:2607.19366v1 Announce Type: new Abstract: Safety as Polytope (SaP) learns linear half-space constraints in LLM hidden space but requires per-category tuning of the constraint count K. We show th

model-releasesarxiv-cs-ai
23 Jul 2026
Safety

Sound Probabilistic Safety Bounds for Large Language Models

DGX agent

arXiv:2607.20286v1 Announce Type: cross Abstract: We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given pr

safetyarxiv-cs-ai
23 Jul 2026
Model Releases

Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric

DGX agent

arXiv:2607.12469v1 Announce Type: cross Abstract: Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

SoftVTBench: A Safety-Aware Visuo-Tactile Benchmark for Physically Constrained Robotic Manipulation of Deformable Objects

DGX agent

arXiv:2607.04234v1 Announce Type: cross Abstract: Deformable object manipulation poses challenges beyond task completion: successful execution must also maintain safe physical interaction, holding the

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer

DGX agent

arXiv:2512.11891v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in generalizing across diverse robotic manipulation tasks. However, de

model-releasesarxiv-cs-ro
3 Jul 2026
Model Releases

High-Speed Vision-Based Flight in Clutter with Safety-Shielded Reinforcement Learning

DGX agent

arXiv:2602.08653v2 Announce Type: replace Abstract: Quadrotor unmanned aerial vehicles (UAVs) are increasingly deployed in complex missions that demand reliable autonomous navigation and robust obstac

model-releasesarxiv-cs-ro
1 Jul 2026
Model Releases

Bridging the Gap Between Image Restoration and Navigational Safety in Hazy Conditions: A New Visibility Estimation Metric for Maritime Surveillance

DGX agent

arXiv:2606.30049v1 Announce Type: new Abstract: Visibility distance is critical to maritime navigational safety because it determines the effective observation range of shipborne and shore-based monit

model-releasesarxiv-cs-cv
30 Jun 2026
Safety

HJ-SafeDMP: Hamilton-Jacobi Reachability-Guided Dynamic Movement Primitives for Provably Safe Robot Motion

DGX agent

arXiv:2606.28995v1 Announce Type: new Abstract: Robots deployed in safety-critical environments must execute motions that are simultaneously robust to disturbances and provably safe from collisions. D

safetyarxiv-cs-ro
30 Jun 2026
Safety

Beyond Feedforward Networks: Reentry Neural Systems as the Fundamental Basis of Subjecthood and Intrinsic Safety of Next-Generation AGI

DGX agent

arXiv:2606.26406v1 Announce Type: cross Abstract: We propose a complete architectural blueprint for safe artificial general intelligence based on a closed reentry loop (D I cycle). In contrast to feed

safetyarxiv-cs-ai
26 Jun 2026
Model Releases

Distributed Model Predictive Control with Adaptive Safety Zones for Multi-Fleet Drone Operations

DGX agent

arXiv:2606.20651v1 Announce Type: cross Abstract: Autonomous drone swarms in space-constrained environments such as warehouses, inspection corridors, and urban delivery routes must share limited airsp

model-releasesarxiv-cs-ro
23 Jun 2026
Safety

For Robotaxis, Safety Must Be Built In, Not Bolted On

DGX agent

A car pulls up to the curb. The app says, “Your ride is here.” No one’s in the driver’s seat. For people who live in one of the dozens of cities now hosting robotaxi services, this is already a realit

safetynvidia-blog
10 Jun 2026
Model Releases

Claude Fable 5 and new AI safety fables

DGX agent

This article discusses Claude Fable 5, likely exploring Anthropic's latest version of their AI model and examining new fables or narratives related to AI safety concepts. The piece probably analyzes h

model-releasesinterconnects
9 Jun 2026
Model Releases

Stress-testing medical large language models reveals latent safety pathology beyond benchmark accuracy

DGX agent

arXiv:2606.07929v1 Announce Type: new Abstract: Large language models (LLMs) are entering clinical practice based on benchmark accuracy that may fail to detect safety-relevant failure modes. Here we p

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

An Analysis Focused on Womens Safety: Can VAD Models Be Enhanced by a Multi-modal Dataset?

DGX agent

arXiv:2605.25806v2 Announce Type: replace Abstract: Women's safety and security are paramount for a modern society. Crimes against women occur in daylight as well as in low-light conditions. Often, su

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

Cross-Generational Transfer of Adversarial Attacks Reveals Non-Monotonic Safety Alignment in LLMs

DGX agent

arXiv:2606.00813v1 Announce Type: cross Abstract: Safety alignment in LLMs does not improve monotonically across model generations. Studying four generations of Google's Gemma family (7B-31B) with qua

model-releasesarxiv-cs-cl
2 Jun 2026
Safety

Food Noise & False Safety: A Systematic Evaluation of How LLMs Fail to Adapt to Eating Disorder Queries with Clinician Feedback

DGX agent

arXiv:2606.02444v1 Announce Type: new Abstract: Recent evidence shows that people with eating disorders (EDs) are increasingly seeking guidance, advice, and emotional support from Large Language Model

safetyarxiv-cs-ai
2 Jun 2026
Safety

InFerActive: Interactive Tree-Based Exploration of LLM Sampling for Safety Evaluation

DGX agent

arXiv:2512.10234v2 Announce Type: replace-cross Abstract: Even LLMs that appear safe during evaluation can still produce harmful responses in deployment. Because stochastic sampling yields different r

safetyarxiv-cs-ai
2 Jun 2026
Safety

Market-Based Replanning for Safety-Critical UAV Swarms in Search and Rescue Missions

DGX agent

arXiv:2606.01970v1 Announce Type: new Abstract: Reliable autonomous UAV swarms in Search and Rescue (SAR) missions require fault-tolerant coordination capable of sustaining operations despite agent de

safetyarxiv-cs-ro
2 Jun 2026
Model Releases

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

DGX agent

arXiv:2605.29224v1 Announce Type: cross Abstract: AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating

model-releasesarxiv-cs-ai
29 May 2026
Safety

No Certificate for Alignment: Two Independent Impossibilities and the Pareto Frontier of Achievable Safety Guarantees

DGX agent

arXiv:2603.08761v2 Announce Type: replace-cross Abstract: We argue that formal certification of AI alignment over open-ended or unbounded input domains is impossible under standard assumptions in comp

safetyarxiv-cs-lg
28 May 2026
Safety

Safe Reinforcement Learning with Preference-based Constraint Inference

DGX agent

arXiv:2603.23565v2 Announce Type: replace-cross Abstract: Safe reinforcement learning (RL) is a standard paradigm for safety-critical decision making. However, real-world safety constraints can be com

safetyarxiv-cs-ai
25 May 2026
Local Ai

Boundary-targeted Membership Inference Attacks on Safety Classifiers

DGX agent

arXiv:2605.22373v1 Announce Type: cross Abstract: Safety classifiers are essential safeguards within generative AI systems, filtering harmful content or identifying at-risk users when interacting with

local-aiarxiv-cs-cl
22 May 2026
← Previous
1…1213141516…297
Next →