AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
28 May 2026

VLM-Based Advanced Rider Assistance System for Motorcycle Safety

SafetyDGX agent

arXiv:2605.27948v1 Announce Type: new Abstract: Motorcycles face disproportionately high crash risks compared to cars due to limited protection and heightened sensitivity to surface hazards, yet Advan

27 May 2026

FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents

SafetyDGX agent

arXiv:2605.27333v1 Announce Type: new Abstract: Finance LLM agents must simultaneously block prompt-induced unauthorized actions and approve legitimate multi-step business workflows. However, boundary

Learning to Balance Motor Thermal Safety and Quadrupedal Locomotion Performance with Residual Policy

Safety
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2605.27046v1 Announce Type: new Abstract: Motor thermal management is often overlooked in the context of electrically-actuated robots, particularly legged robots, but motor overheating is a key

26 May 2026

Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts

SafetyDGX agent

arXiv:2605.24270v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) language models activate only a small subset of parameters for each token, making router behavior a central part of mode

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

SafetyDGX agent

arXiv:2605.23989v1 Announce Type: new Abstract: Agentic AI systems -- Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions -- can execute complex tasks

25 May 2026

SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety

SafetyDGX agent

arXiv:2605.05704v2 Announce Type: replace-cross Abstract: Recent advances in foundation models have transformed LLMs from passive conversational systems into autonomous agents capable of reasoning and

22 May 2026

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

Model ReleasesDGX agent

arXiv:2605.22643v1 Announce Type: new Abstract: Background. Traditional safety benchmarks for language models evaluate generated text: whether a model outputs toxic language, reproduces bias, or follo

GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety

Model ReleasesDGX agent

arXiv:2605.20203v1 Announce Type: cross Abstract: As older adults increasingly use LLM-based chatbots for companionship and assistance, a safety gap is emerging. Older adults may face vulnerabilities

HealthCraft: A Reinforcement Learning Safety Environment for Emergency Medicine

Model ReleasesDGX agent

arXiv:2605.21496v1 Announce Type: cross Abstract: Frontier language models are being deployed into clinical workflows faster than the infrastructure to evaluate them safely. Static medical-QA benchmar

21 May 2026

Tech researchers are suing the Trump administration over the future of online safety

SafetyDGX agent

Since its earliest days back in office, the Trump administration has been going after researchers who study and try to counter hate speech, harassment, propaganda, and disinformation online. Now, some

The Download: online safety’s future and climate tech’s big pivot

SafetyDGX agent

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Tech researchers are suing the Trump administration over the f

20 May 2026

Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework

SafetyDGX agent

arXiv:2603.11768v2 Announce Type: replace Abstract: Long-term memory has emerged as a foundational component of autonomous Large Language Model (LLM) agents, enabling continuous adaptation, lifelong m

Measuring Safety Alignment Effects in Autonomous Security Agents

Model ReleasesDGX agent

arXiv:2605.19722v1 Announce Type: cross Abstract: Do stock safety-aligned language models and their uncensored or abliterated derivatives behave differently when run as autonomous security agents? Sin

19 May 2026

AgentWall: A Runtime Safety Layer for Local AI Agents

Model ReleasesDGX agent

arXiv:2605.16265v1 Announce Type: new Abstract: The safety of autonomous AI agents is increasingly recognized as a critical open problem. As agents transition from passive text generators to active ac

Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation

SafetyDGX agent

arXiv:2605.17268v1 Announce Type: new Abstract: We present the first systematic study of faithfulness in Vision-Language-Action (VLA) driving models, analyzing 300 Alpamayo-R1-10B inferences across 10

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents

Model ReleasesDGX agent

arXiv:2601.17887v2 Announce Type: replace Abstract: Long-term memory enables large language model (LLM) agents to support personalized and sustained interactions. However, most work on personalized ag

12 May 2026

Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing

Model ReleasesDGX agent

arXiv:2605.10146v1 Announce Type: new Abstract: Large language models (LLMs) increasingly rely on knowledge editing to support knowledge-intensive reasoning, but this flexibility also introduces criti

11 May 2026

SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints

SafetyDGX agent

arXiv:2512.23770v3 Announce Type: replace-cross Abstract: In safety-critical domains, reinforcement learning (RL) agents must often satisfy strict, zero-cost safety constraints while accomplishing tas

10 May 2026

Big AI Lobbyists: if you regulate us at all, we lose to China because they will never regulate ... Actual China: 'safety first, innovation second ... Development must be controllable and orderly.'

SafetyDGX agent

This post highlights a contradiction in AI industry arguments, contrasting Western tech company claims that regulation will disadvantage them competitively against China with evidence of China's own s

6 May 2026

Google, Microsoft and xAI agree to allow government safety checks of their AI models prior to release

SafetyDGX agent

Google LLC, Microsoft Corp. and xAI have agreed to share unreleased versions of their artificial intelligence models with the U.S. Department of Commerce to ensure the technologies do not pose a threa

Safety and accuracy follow different scaling laws in clinical large language models

Model ReleasesDGX agent

arXiv:2605.04039v1 Announce Type: new Abstract: Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation

Set-Based Training of Neural Barrier Certificates for Safety Verification of Dynamical Systems

SafetyDGX agent

arXiv:2605.02526v1 Announce Type: cross Abstract: Barrier certificates are scalar functions over the state space of dynamical systems that separate all unsafe states from all reachable states. The exi

5 May 2026

A decoupled diffusion planner that adapts to changing cost limits by using cost-conditioned generation for safety and reward gradients for performance

Model ReleasesDGX agent

arXiv:2605.02777v1 Announce Type: new Abstract: Offline safe reinforcement learning often requires policies to adapt at deployment time to safety budgets that vary across episodes or change within a s

Value Functions for Temporal Logic: Optimal Policies and Safety Filters

SafetyDGX agent

arXiv:2605.01051v1 Announce Type: cross Abstract: While Bellman equations for basic reach, avoid, and reach-avoid problems are well studied, the relationship between value optimality and policy optima

When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent Systems

SafetyDGX agent

arXiv:2605.01133v1 Announce Type: cross Abstract: Large language model (LLM)-powered multi-agent systems (MAS) enable agents to communicate and share information, achieving strong performance on compl

30 Apr 2026

Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex

Model ReleasesDGX agent

arXiv:2604.14858v2 Announce Type: replace Abstract: As agent systems move into increasingly diverse execution settings, trajectory-level safety evaluation and diagnosis require benchmarks that evolve

29 Apr 2026

Instantaneous Planning, Control and Safety for Navigation in Unknown Underwater Spaces

SafetyDGX agent

arXiv:2604.05310v2 Announce Type: replace Abstract: Navigating autonomous underwater vehicles (AUVs) in unknown environments is significantly challenging due to poor visibility, weak signal transmissi

28 Apr 2026

Evaluating whether AI models would sabotage AI safety research

Model ReleasesDGX agent

arXiv:2604.24618v1 Announce Type: new Abstract: We evaluate the propensity of frontier models to sabotage or refuse to assist with safety research when deployed as AI research agents within a frontier

How Sensitive Are Safety Benchmarks to Judge Configuration Choices?

Model ReleasesDGX agent

arXiv:2604.24074v1 Announce Type: new Abstract: Safety benchmarks such as HarmBench rely on LLM judges to classify model responses as harmful or safe, yet the judge configuration, namely the combinati

24 Apr 2026

Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2603.21697v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) extend text-only LLMs with visual reasoning, but also introduce new safety failure modes under visual

21 Apr 2026

Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety

Model ReleasesDGX agent

arXiv:2604.18487v1 Announce Type: new Abstract: The Adversarial Humanities Benchmark (AHB) evaluates whether model safety refusals survive a shift away from familiar harmful prompt forms. Starting fro

Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation

Model ReleasesDGX agent

arXiv:2602.07954v4 Announce Type: replace Abstract: As Large Language Models (LLMs) become increasingly deployed in Polish language applications, the need for efficient and accurate content safety cla

BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration

Model ReleasesDGX agent

arXiv:2604.16541v1 Announce Type: new Abstract: Recent advancements in Large Generative Models (LGMs) have revolutionized multi-modal generation. However, generating illustrated storybooks remains an

Structure-Aware Diversity Pursuit as an AI Safety Strategy against Homogenization

SafetyDGX agent

arXiv:2601.06116v2 Announce Type: replace-cross Abstract: Generative AI models reproduce the biases in the training data and can further amplify them through mode collapse. We refer to the resulting h

10 Apr 2026

Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

Model ReleasesDGX agent

arXiv:2604.06173v1 Announce Type: cross Abstract: Legal QA benchmarks have predominantly focused on case law, overlooking the unique challenges of statute-centric regulatory reasoning. In statutory do

FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding

Model ReleasesDGX agent

arXiv:2604.07879v1 Announce Type: new Abstract: Diffusion-based image generation models have advanced rapidly but pose a safety risk due to their potential to generate Not-Safe-For-Work (NSFW) content

11 Aug 2026

Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity

Model ReleasesDGX agent

arXiv:2608.09443v1 Announce Type: new Abstract: Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbidity depend on

Learning Multi-Timescale Interventions under Safety and Resource Constraints

SafetyDGX agent

arXiv:2508.03875v2 Announce Type: replace Abstract: Many sequential decision problems offer qualitatively different ways of influencing the environment: some interventions act immediately, whereas oth

5 Aug 2026

Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

Model ReleasesDGX agent

arXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive

HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection

Model ReleasesDGX agent

arXiv:2509.23690v2 Announce Type: replace-cross Abstract: Safety hazards in the home are a leading cause of preventable domestic injuries, motivating an automated inspector that actively explores a ho

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models

Model ReleasesDGX agent

arXiv:2608.03201v1 Announce Type: new Abstract: Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs. We audit

4 Aug 2026

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard)

SafetyDGX agent

arXiv:2608.00180v1 Announce Type: new Abstract: Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard. Training a safety guard with RL means optimizing two

🛡️Introducing Shieldstral, Mistral’s 3B open-weights model for content safety that can be deployed on-device 🧵 http://mistral.ai/news/shie…

Model ReleasesDGX agent

Mistral AI introduced Shieldstral, a 3‑billion‑parameter, open‑weights model designed for content‑safety tasks and capable of on‑device deployment. The announcement was shared via a tweet from @Mistra

31 Jul 2026

It’s time to panic about AI safety

SafetyDGX agent

When the phrase 'OpenAI hacked Hugging Face' has more or less entered mainstream culture, you know we have an AI problem. This week, we learned more about exactly how OpenAI's agent broke out of a san

Safety Verification of Wait-Only Non-Blocking Broadcast Protocols

SafetyDGX agent

arXiv:2403.18591v3 Announce Type: replace-cross Abstract: Broadcast protocols are programs designed to be executed by networks of processes. Each process runs the same protocol, and communication betw

30 Jul 2026

EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

Model ReleasesDGX agent

arXiv:2607.26518v1 Announce Type: new Abstract: Reliable visual safety understanding in real-world scenarios demands more than just object recognition; it requires causal reasoning under epistemic unc

29 Jul 2026

We’re running out of reasons to ignore AI safety

SafetyDGX agent

Earlier this month, OpenAI gave several of its AI models a task: complete a test designed to measure their cybersecurity capabilities. It put the systems in a sandboxed environment without an internet

28 Jul 2026

Distribution-Specific Curvature Control with Finite-Sample Guarantees for Open-Weight Safety

Local AiDGX agent

arXiv:2607.22929v1 Announce Type: new Abstract: A short fine-tuning run can undo the safety guards of an open-weight model---retraining a refusal-trained assistant to aid weapons development or produc

OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows

Model ReleasesDGX agent

arXiv:2510.24411v3 Announce Type: replace Abstract: Computer-using agents powered by Vision-Language Models (VLMs) have demonstrated human-like capabilities in operating digital environments like mobi

24 Jul 2026

HERMES: Heterogeneous Edge-Relational Multi-Head Embedded SSM Attention for Traffic Conflict Prediction at Signalized Intersections

SafetyDGX agent

arXiv:2607.20505v1 Announce Type: cross Abstract: Surrogate safety measures (SSMs) enable proactive traffic safety assessment, but many existing methods evaluate pairwise interactions independently or

23 Jul 2026

Geometry-Guided Constraint Learning for LLM Safety Classification

Model ReleasesDGX agent

arXiv:2607.19366v1 Announce Type: new Abstract: Safety as Polytope (SaP) learns linear half-space constraints in LLM hidden space but requires per-category tuning of the constraint count K. We show th

Sound Probabilistic Safety Bounds for Large Language Models

SafetyDGX agent

arXiv:2607.20286v1 Announce Type: cross Abstract: We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given pr

15 Jul 2026

Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric

Model ReleasesDGX agent

arXiv:2607.12469v1 Announce Type: cross Abstract: Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may

7 Jul 2026

SoftVTBench: A Safety-Aware Visuo-Tactile Benchmark for Physically Constrained Robotic Manipulation of Deformable Objects

Model ReleasesDGX agent

arXiv:2607.04234v1 Announce Type: cross Abstract: Deformable object manipulation poses challenges beyond task completion: successful execution must also maintain safe physical interaction, holding the

3 Jul 2026

VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer

Model ReleasesDGX agent

arXiv:2512.11891v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in generalizing across diverse robotic manipulation tasks. However, de

1 Jul 2026

High-Speed Vision-Based Flight in Clutter with Safety-Shielded Reinforcement Learning

Model ReleasesDGX agent

arXiv:2602.08653v2 Announce Type: replace Abstract: Quadrotor unmanned aerial vehicles (UAVs) are increasingly deployed in complex missions that demand reliable autonomous navigation and robust obstac

30 Jun 2026

Bridging the Gap Between Image Restoration and Navigational Safety in Hazy Conditions: A New Visibility Estimation Metric for Maritime Surveillance

Model ReleasesDGX agent

arXiv:2606.30049v1 Announce Type: new Abstract: Visibility distance is critical to maritime navigational safety because it determines the effective observation range of shipborne and shore-based monit

HJ-SafeDMP: Hamilton-Jacobi Reachability-Guided Dynamic Movement Primitives for Provably Safe Robot Motion

SafetyDGX agent

arXiv:2606.28995v1 Announce Type: new Abstract: Robots deployed in safety-critical environments must execute motions that are simultaneously robust to disturbances and provably safe from collisions. D

26 Jun 2026

Beyond Feedforward Networks: Reentry Neural Systems as the Fundamental Basis of Subjecthood and Intrinsic Safety of Next-Generation AGI

SafetyDGX agent

arXiv:2606.26406v1 Announce Type: cross Abstract: We propose a complete architectural blueprint for safe artificial general intelligence based on a closed reentry loop (D I cycle). In contrast to feed

23 Jun 2026

Distributed Model Predictive Control with Adaptive Safety Zones for Multi-Fleet Drone Operations

Model ReleasesDGX agent

arXiv:2606.20651v1 Announce Type: cross Abstract: Autonomous drone swarms in space-constrained environments such as warehouses, inspection corridors, and urban delivery routes must share limited airsp

← Previous
1…910111213…238
Next →