AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
9 Jun 2026

Your Model Already Knows: Attention-Guided Safety Filter for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.09749v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated impressive end-to-end performance across a variety of robotic manipulation tasks. However, these

6 Jun 2026

Data Flow Control: Data Safety Policies for AI Agents

Model ReleasesDGX agent

arXiv:2606.05679v1 Announce Type: cross Abstract: Agents increasingly generate SQL, orchestrate pipelines, and automate data analysis on behalf of users. While recent work improves query correctness,

4 Jun 2026

Safety by narrow control has shown to fail many times. Need more transparency on the absolute frontier, and openness close behind.

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

Safety by narrow control has shown to fail many times. Need more transparency on the absolute frontier, and openness close behind. I found another API that offers claude-oceanus-v1-p the pricing and t

AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety

Model ReleasesDGX agent

arXiv:2606.04867v1 Announce Type: new Abstract: As AI companion platforms such as Replika and Character.AI rapidly grow, concerns about unsafe human-AI interactions have intensified. This study introd

In policy paper, OpenAI diverges from White House on AI safety

Model ReleasesDGX agent

OpenAI Group PBC’s newly released proposal for how advanced artificial intelligence should be regulated differs slightly from the Trump administration’s executive order, also released this week. Relea

3 Jun 2026

Edge-Aware and Content-Adaptive Infrared Gas Leak Detection for Industrial Safety Monitoring

Local AiDGX agent

arXiv:2512.23234v3 Announce Type: replace-cross Abstract: Infrared gas leak detection is important for industrial safety and environmental monitoring, but automatic detection remains challenging becau

EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents

Model ReleasesDGX agent

arXiv:2606.03678v1 Announce Type: new Abstract: Generating safety-critical scenarios is essential for validating and improving autonomous driving systems, yet it inherently requires maximizing adversa

From 3D Perception to Safety Reasoning: A Graph-Based Framework for Real-Time Underground Mine Monitoring

Local AiDGX agent

arXiv:2606.03460v1 Announce Type: new Abstract: Underground coal mining requires personnel and heavy equipment to operate within shared, confined, and poorly illuminated spaces where hazards such as e

28 May 2026

glad to know Mythos' safety concerns have been addressed right as Anthropic also secured tens of billions in inference compute 👍

Model ReleasesDGX agent

glad to know Mythos' safety concerns have been addressed right as Anthropic also secured tens of billions in inference compute 👍 JUST IN: Anthropic announces it will roll out Claude Mythos “in the com

26 May 2026

StructBreak: Structural Cognitive Overload-Induced Safety Failures in MLLMs

Model ReleasesDGX agent

arXiv:2605.25534v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) excel at structural reasoning yet suffer from a sharp logical brittleness in structural consistency. We term th

Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models

SafetyDGX agent

arXiv:2605.24550v1 Announce Type: new Abstract: Fine-tuning-as-a-Service (FaaS) enables personalization of large language models (LLMs), but it can weaken safety-alignment under harmful fine-tuning at

25 May 2026

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

Model ReleasesDGX agent

arXiv:2602.12316v2 Announce Type: replace Abstract: Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely ev

22 May 2026

Steins;Gate Drive: Semantic Safety Arbitration over Structured Futures for Latency-Decoupled LLM Planning

Model ReleasesDGX agent

arXiv:2605.22456v1 Announce Type: new Abstract: Cloud-hosted LLM driver agents provide useful semantic judgments, but their inference latency exceeds stepwise vehicle-control windows. Learned world mo

20 May 2026

Adversarial Stress Testing of SPARK Humanoid Safety Filters

Model ReleasesDGX agent

arXiv:2605.19009v1 Announce Type: new Abstract: Humanoid robots are difficult to deploy safely because they have high-dimensional bodies, many collision constraints, and must operate near people and o

Jailbreaking on Text-to-Video Models via Scene Splitting Strategy

SafetyDGX agent

arXiv:2509.22292v2 Announce Type: replace-cross Abstract: Along with the rapid advancement of numerous Text-to-Video (T2V) models, growing concerns have emerged regarding their safety risks. While rec

Passive Construction Site Safety Monitoring via Persona-Scaffolded Adversarial Chain-of-Thought VLM Verification

Model ReleasesDGX agent

arXiv:2605.19869v1 Announce Type: cross Abstract: Construction remains the deadliest industry sector in the United States, with 1,055 fatal worker injuries recorded in 2023, and the majority preventab

19 May 2026

Conservative AI for Safety-Sensitive Medical Image Restoration: Residual-Bounded CT-CTA Enhancement for Intracranial Aneurysm-Relevant Signal Recovery

Local AiDGX agent

arXiv:2605.16458v1 Announce Type: cross Abstract: Image restoration models are increasingly applied to degraded medical scans, but in safety-sensitive settings they must improve image quality without

EvoQRE: Modeling Bounded Rationality in Safety-Critical Traffic Simulation via Evolutionary Quantal Response Equilibrium

Model ReleasesDGX agent

arXiv:2601.05653v2 Announce Type: replace Abstract: Existing traffic simulation frameworks for autonomous vehicles typically rely on imitation learning or game-theoretic approaches that solve for Nash

13 May 2026

🎉 We published a new AI safety study: shopping agents fall for whimsical attacks and lose money. A whimsical attack is an absurd scenario a…

Model ReleasesDGX agent

🎉 We published a new AI safety study: shopping agents fall for whimsical attacks and lose money. A whimsical attack is an absurd scenario a human would never try on another human. In one run, GPT-5.1

11 May 2026

CommandSwarm: Safety-Aware Natural Language-to-Behavior-Tree Generation for Robotic Swarms

Model ReleasesDGX agent

arXiv:2605.07764v1 Announce Type: new Abstract: Natural-language interfaces can make swarm robotics more accessible to non-expert operators, but they must translate ambiguous user intent into executab

Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with…

Model ReleasesDGX agent

Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with the full power of Bash access? We built exactly that. Meet

9 May 2026

Anthropic details how it improved Claude's safety training after finding agentic misalignment in older models, such as Opus 4 blackmailing engineers (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic details how it improved Claude's safety training after finding agentic misalignment in older models, such as Opus 4 blackmailing engineers — Last year, we released a case study on

6 May 2026

Self-Mined Hardness for Safety Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.03226v1 Announce Type: new Abstract: Safety fine-tuning of language models typically requires a curated adversarial dataset. We take a different approach: score each candidate prompt's diff

5 May 2026

ARMOR 2025: A Military-Aligned Benchmark for Evaluating Large Language Model Safety Beyond Civilian Contexts

Model ReleasesDGX agent

arXiv:2605.00245v1 Announce Type: new Abstract: Large language models (LLMs) are now being explored for defense applications that require reliable and legally compliant decision support. They also hol

Attention Is Where You Attack

Model ReleasesDGX agent

arXiv:2605.00236v1 Announce Type: cross Abstract: Safety-aligned large language models rely on RLHF and instruction tuning to refuse harmful requests, yet the internal mechanisms implementing safety b

27 Apr 2026

SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs

Model ReleasesDGX agent

arXiv:2604.22134v1 Announce Type: new Abstract: Large Language Models (LLMs) have been widely explored in educational scenarios. We identify a critical vulnerability in current educational LLMs, pedag

22 Apr 2026

SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2604.19638v1 Announce Type: new Abstract: Multimodal Large Language Models are increasingly adopted as autonomous agents in interactive environments, yet their ability to proactively address saf

21 Apr 2026

Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models

Model ReleasesDGX agent

arXiv:2604.16980v1 Announce Type: new Abstract: Background: Large language models (LLMs) are increasingly proposed for diagnostic support, but few evaluations use real-world multimodal inpatient data,

Driving risk emerges from the required two-dimensional joint evasive acceleration

SafetyDGX agent

arXiv:2604.17841v1 Announce Type: new Abstract: Most autonomous driving safety benchmarks use time-to-collision (TTC) to assess risk and guide safe behaviour. However, TTC-based methods treat risk as

Integrated Wheel Sensor Communication using ESP32 -- A Contribution towards a Digital Twin of the Road System

SafetyDGX agent

arXiv:2509.04061v2 Announce Type: replace Abstract: While current onboard state estimation methods are adequate for most driving and safety-related applications, they do not provide insights into the

19 Apr 2026

I wrapped my Ollama agent with deterministic safety checks — here's the setup (catches bad JSON, prompt injection, and refusals before they hit your app) apache 2.0 [GitHub: https://github.com/qaysSE/AG-X]

Local AiDGX agent

This GitHub project (AG-X) implements deterministic safety checks for Ollama agents, designed to catch malformed JSON responses, detect prompt injection attempts, and handle model refusals before they

14 Apr 2026

Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs

Model ReleasesDGX agent

arXiv:2604.11120v1 Announce Type: new Abstract: Personality imbuing customizes LLM behavior, but safety evaluations almost always study prompt-based personas alone. We show this is incomplete: prompti

STARS: Skill-Triggered Audit for Request-Conditioned Invocation Safety in Agent Systems

Model ReleasesDGX agent

arXiv:2604.10286v1 Announce Type: new Abstract: Autonomous language-model agents increasingly rely on installable skills and tools to complete user tasks. Static skill auditing can expose capability s

11 Aug 2026

ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents

Model ReleasesDGX agent

arXiv:2608.09476v1 Announce Type: cross Abstract: Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavio

10 Aug 2026

HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses

Model ReleasesDGX agent

arXiv:2608.06984v1 Announce Type: cross Abstract: Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts. However,

5 Aug 2026

Risky Business: Measuring The Faithfulness-Safety Tension

Model ReleasesDGX agent

arXiv:2608.03745v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring. However, monitoring relies on faithfulness, i.e., the model output str

30 Jul 2026

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

Model ReleasesDGX agent

arXiv:2409.07314v4 Announce Type: replace Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become satur

Risk-Aware Motion Planning with Learned Trajectory Primitives and Probabilistic Safety Assessment

Model ReleasesDGX agent

arXiv:2607.26802v1 Announce Type: new Abstract: This paper presents a radial basis function network (RBFN)-informed motion planning framework for safe and efficient urban autonomous driving. The propo

29 Jul 2026

AMPBench-MT: A Homology-Controlled Benchmark for Antimicrobial Peptide Potency, Spectrum, and Safety Prediction

Model ReleasesDGX agent

arXiv:2607.25518v1 Announce Type: new Abstract: Computational AMP discovery is often evaluated through AMP/non-AMP recognition, yet follow-up decisions depend on assay-derived evidence such as target-

24 Jul 2026

Is Your Safe Controller Actually Safe? A Critical Review of CBF Tautologies and Hidden Assumptions

SafetyDGX agent

arXiv:2603.06954v2 Announce Type: replace Abstract: This tutorial provides a critical review of the practical application of Control Barrier Functions (CBFs) in robotic safety. While the theoretical f

23 Jul 2026

OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills

Model ReleasesDGX agent

arXiv:2607.20121v1 Announce Type: new Abstract: LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra secur

10 Jul 2026

Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies

SafetyDGX agent

arXiv:2603.15136v2 Announce Type: replace-cross Abstract: Offline safe reinforcement learning (RL) seeks reward-maximizing policies from static datasets under strict safety constraints. Existing metho

9 Jul 2026

Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety

Model ReleasesDGX agent

arXiv:2607.07695v1 Announce Type: new Abstract: We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task

Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

Model ReleasesDGX agent

arXiv:2607.07229v1 Announce Type: new Abstract: Prior work has shown that chain-of-thought (CoT) reasoning is often unfaithful: a model's stated reasoning does not reliably reflect the process that pr

7 Jul 2026

Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives

Model ReleasesDGX agent

arXiv:2604.16870v2 Announce Type: replace-cross Abstract: AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP). These tool calls are the age

1 Jul 2026

Addressing Over-Refusal in LLMs with Competing Rewards

SafetyDGX agent

arXiv:2606.31748v1 Announce Type: new Abstract: Safety training on language models often induces over-refusal: improved safety on harmful prompts at the cost of increased refusal on harmless ones. Tho

ShardNet: Training Neural Controllers with Hard, Non-Convex Constraints

SafetyDGX agent

arXiv:2606.30935v1 Announce Type: cross Abstract: While neural network control policies are powerful, their deployment on safety critical systems depends on ensuring that they obey strict constraints.

25 Jun 2026

Safe Learning Control with Optimality and Stability Guarantees

SafetyDGX agent

arXiv:2501.15373v2 Announce Type: replace-cross Abstract: Merely pursuing performance may adversely affect safety, while a conservative policy for safe exploration will degrade the performance. How to

5 Jun 2026

Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming

Model ReleasesDGX agent

arXiv:2606.05233v1 Announce Type: cross Abstract: Recent computer-using-agent (CUA) red-teaming papers report prompt-injection attack success rates (ASR) of 42-98%, but these headline numbers cluster

2 Jun 2026

A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models

Model ReleasesDGX agent

arXiv:2606.00027v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or

SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation

Model ReleasesDGX agent

arXiv:2606.01481v1 Announce Type: new Abstract: With the rapid advancements in text-to-image diffusion models, generative video models (T2V models) like Sora can now produce short synthetic videos fro

18 May 2026

Graph-Regularized Sparse Autoencoders for LLM Safety Steering

Model ReleasesDGX agent

arXiv:2512.06655v3 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) are increasingly used to extract activation directions for inference-time steering, but their standard sparsity obj

15 May 2026

Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders: Safety Risks in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2605.13851v1 Announce Type: new Abstract: Multi-agent orchestration -- in which a hidden coordinator manages specialized worker agents -- is becoming the default architecture for enterprise AI d

14 May 2026

Children's English Reading Story Generation via Supervised Fine-Tuning of Compact LLMs with Controllable Difficulty and Safety

Model ReleasesDGX agent

arXiv:2605.13709v1 Announce Type: cross Abstract: Large Language Models (LLMs) are widely applied in educational practices, such as for generating children's stories. However, the generated stories ar

4 May 2026

LWiAI Podcast #243 - GPT 5.5, DeepSeek V4, AI safety sabotage

Model ReleasesDGX agent

This podcast episode from Last Week in AI discusses recent developments in large language models, including updates on GPT 5.5 and DeepSeek V4, while also covering concerns about potential sabotage or

29 Apr 2026

Safe-Support Q-Learning: Learning without Unsafe Exploration

SafetyDGX agent

arXiv:2604.25379v1 Announce Type: new Abstract: Ensuring safety during reinforcement learning (RL) training is critical in real-world applications where unsafe exploration can lead to devastating outc

28 Apr 2026

CLIN-LLM: A Safety-Constrained Hybrid Framework for Clinical Diagnosis and Treatment Generation

Model ReleasesDGX agent

arXiv:2510.22609v2 Announce Type: replace Abstract: Accurate symptom-to-disease classification and clinically grounded treatment recommendations remain challenging, particularly in heterogeneous patie

Control Barrier Functions Solved with Hierarchical Quadratic Programming for Safe Physical Human-Robot Interaction

SafetyDGX agent

arXiv:2604.23039v1 Announce Type: new Abstract: Physical human-robot interaction offers the potential to leverage human intelligence and robot physical capabilities to enable a range of exciting appli

16 Apr 2026

Capability-Aware Heterogeneous Control Barrier Functions for Decentralized Multi-Robot Safe Navigation

SafetyDGX agent

arXiv:2604.13245v1 Announce Type: new Abstract: Safe navigation for multi-robot systems requires enforcing safety without sacrificing task efficiency under decentralized decision-making. Existing dece

15 Apr 2026

RACF: A Resilient Autonomous Car Framework with Object Distance Correction

SafetyDGX agent

arXiv:2604.12418v1 Announce Type: cross Abstract: Autonomous vehicles are increasingly deployed in safety-critical applications, where sensing failures or cyberphysical attacks can lead to unsafe oper

← Previous
1…1112131415…238
Next →