AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Model Releases

Graph-Regularized Sparse Autoencoders for LLM Safety Steering

DGX agent

arXiv:2512.06655v3 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) are increasingly used to extract activation directions for inference-time steering, but their standard sparsity obj

model-releasesarxiv-cs-ai
18 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders: Safety Risks in Multi-Agent LLM Systems

DGX agent

arXiv:2605.13851v1 Announce Type: new Abstract: Multi-agent orchestration -- in which a hidden coordinator manages specialized worker agents -- is becoming the default architecture for enterprise AI d

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Children's English Reading Story Generation via Supervised Fine-Tuning of Compact LLMs with Controllable Difficulty and Safety

DGX agent

arXiv:2605.13709v1 Announce Type: cross Abstract: Large Language Models (LLMs) are widely applied in educational practices, such as for generating children's stories. However, the generated stories ar

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

Attention Is Where You Attack

DGX agent

arXiv:2605.00236v1 Announce Type: cross Abstract: Safety-aligned large language models rely on RLHF and instruction tuning to refuse harmful requests, yet the internal mechanisms implementing safety b

model-releasesarxiv-cs-ai
5 May 2026
Model Releases

LWiAI Podcast #243 - GPT 5.5, DeepSeek V4, AI safety sabotage

DGX agent

This podcast episode from Last Week in AI discusses recent developments in large language models, including updates on GPT 5.5 and DeepSeek V4, while also covering concerns about potential sabotage or

model-releaseslast-week-in-ai
4 May 2026
Safety

Safe-Support Q-Learning: Learning without Unsafe Exploration

DGX agent

arXiv:2604.25379v1 Announce Type: new Abstract: Ensuring safety during reinforcement learning (RL) training is critical in real-world applications where unsafe exploration can lead to devastating outc

safetyarxiv-cs-lg
29 Apr 2026
Model Releases

CLIN-LLM: A Safety-Constrained Hybrid Framework for Clinical Diagnosis and Treatment Generation

DGX agent

arXiv:2510.22609v2 Announce Type: replace Abstract: Accurate symptom-to-disease classification and clinically grounded treatment recommendations remain challenging, particularly in heterogeneous patie

model-releasesarxiv-cs-ai
28 Apr 2026
Safety

Control Barrier Functions Solved with Hierarchical Quadratic Programming for Safe Physical Human-Robot Interaction

DGX agent

arXiv:2604.23039v1 Announce Type: new Abstract: Physical human-robot interaction offers the potential to leverage human intelligence and robot physical capabilities to enable a range of exciting appli

safetyarxiv-cs-ro
28 Apr 2026
Safety

Driving risk emerges from the required two-dimensional joint evasive acceleration

DGX agent

arXiv:2604.17841v1 Announce Type: new Abstract: Most autonomous driving safety benchmarks use time-to-collision (TTC) to assess risk and guide safe behaviour. However, TTC-based methods treat risk as

safetyarxiv-cs-ro
21 Apr 2026
Safety

Integrated Wheel Sensor Communication using ESP32 -- A Contribution towards a Digital Twin of the Road System

DGX agent

arXiv:2509.04061v2 Announce Type: replace Abstract: While current onboard state estimation methods are adequate for most driving and safety-related applications, they do not provide insights into the

safetyarxiv-cs-ro
21 Apr 2026
Safety

Capability-Aware Heterogeneous Control Barrier Functions for Decentralized Multi-Robot Safe Navigation

DGX agent

arXiv:2604.13245v1 Announce Type: new Abstract: Safe navigation for multi-robot systems requires enforcing safety without sacrificing task efficiency under decentralized decision-making. Existing dece

safetyarxiv-cs-ro
16 Apr 2026
Safety

RACF: A Resilient Autonomous Car Framework with Object Distance Correction

DGX agent

arXiv:2604.12418v1 Announce Type: cross Abstract: Autonomous vehicles are increasingly deployed in safety-critical applications, where sensing failures or cyberphysical attacks can lead to unsafe oper

safetyarxiv-cs-ai
15 Apr 2026
Model Releases

Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers

DGX agent

arXiv:2603.28013v3 Announce Type: replace-cross Abstract: Multi-agent LLM systems are entering production -- processing documents, managing workflows, acting on behalf of users -- yet their resilience

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning

DGX agent

arXiv:2604.09452v1 Announce Type: cross Abstract: Safety guarantees are a prerequisite to the deployment of reinforcement learning (RL) agents in safety-critical tasks. Often, deployment environments

model-releasesarxiv-cs-ai
13 Apr 2026
Safety

According to Waymo's published data, their technology is preventing injuries & deaths. My view is that if this is true, and I have yet to se…

DGX agent

According to Waymo's published data, their technology is preventing injuries & deaths. My view is that if this is true, and I have yet to see a debunking of their data, then we safety advocates should

safetyethan-mollick--x
12 Apr 2026
Model Releases

ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces

DGX agent

arXiv:2604.05172v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly deployed to automate productivity tasks (e.g., email, scheduling, document management), but evalu

model-releasesarxiv-cs-ai
10 Apr 2026
Safety

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training

DGX agent

arXiv:2604.07754v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) raises significant ethical and safety concerns. While LLM alignment techniques are adopted to improve m

safetyarxiv-cs-cl
10 Apr 2026
Safety

Topological Feasibility Guarantees for Differentiable Predictive Control

DGX agent

arXiv:2608.10332v1 Announce Type: cross Abstract: Differentiable predictive control (DPC), a self-supervised learning approach for approximating explicit model predictive control (MPC) policies, offer

safetyarxiv-cs-lg
12 Aug 2026
Model Releases

InfoOps Bench: A live information operations safety benchmark

DGX agent

arXiv:2607.28503v3 Announce Type: replace Abstract: In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment

DGX agent

arXiv:2608.06110v1 Announce Type: new Abstract: This paper presents ECHO (Enhanced Care & Health Observer), a locally-deployable conversational health assistant for long-term chronic care management.

model-releasesarxiv-cs-ai
7 Aug 2026
Model Releases

Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load

DGX agent

arXiv:2608.05018v1 Announce Type: new Abstract: Short-term load forecasting (STLF) play a vital role in the electric power industry. It serves infrastructure that European and German law designate as

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity

DGX agent

arXiv:2608.02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is f

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Grasp Execution Without a Planner: Configuration-Space Grasp Distance Fields with Certified Safety & Guaranteed Quality

DGX agent

arXiv:2608.00600v1 Announce Type: new Abstract: Standard multifingered grasp execution architectures plan a collision-free trajectory to a selected grasp pose and track it with a feedback law. Executi

model-releasesarxiv-cs-ro
4 Aug 2026
Model Releases

White House invites AI companies to review its new AI safety framework

DGX agent

Cybersecurity chiefs at the White House have reportedly finalized the outline of a forthcoming framework that will enable artificial intelligence companies to voluntarily submit their latest frontier

model-releasessiliconangle
3 Aug 2026
Model Releases

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs

DGX agent

arXiv:2607.24392v1 Announce Type: cross Abstract: Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. W

model-releasesarxiv-cs-lg
28 Jul 2026
Model Releases

Anthropic launches Claude Opus 5 with efficiency, safety improvements

DGX agent

Anthropic PBC today rolled out a large language model called Claude Opus 5 to its chatbot service and developer platform. The company says the LLM approaches the output quality of its top-end Mythos 5

model-releasessiliconangle
25 Jul 2026
Model Releases

DocShield: Towards AI Document Safety via Evidence-Grounded Agentic Reasoning

DGX agent

arXiv:2604.02694v2 Announce Type: replace-cross Abstract: The rapid progress of generative AI has enabled increasingly realistic text-centric image forgeries, posing major challenges to document safet

model-releasesarxiv-cs-ai
23 Jul 2026
Safety

Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies

DGX agent

arXiv:2607.18325v1 Announce Type: new Abstract: Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and support decision-making during emergencies. Visi

safetyarxiv-cs-cv
23 Jul 2026
Safety

AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to …

DGX agent

AI agents are already being used to improve the capabilities of our next-generation models. We believe with GPT-Red that we have started to unlock a similar flywheel for safety, where today's models c

safetyopenai--x
15 Jul 2026
Safety

A Graph-Based Reinforcement Learning Approach with Frontier Potential Based Reward for Safe Cluttered Environment Exploration

DGX agent

arXiv:2504.11907v3 Announce Type: replace Abstract: Autonomous exploration of cluttered environments requires efficient exploration strategies that guarantee safety against potential collisions with u

safetyarxiv-cs-ro
7 Jul 2026
Safety

Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs

DGX agent

arXiv:2508.10031v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and

safetyarxiv-cs-ai
7 Jul 2026
Model Releases

Anthropic launches Claude Sonnet 5 AI model with coding, safety upgrades as Fable and Mythos controls lifted

DGX agent

Anthropic PBC today debuted Claude Sonnet 5, a midrange large language model that outperforms its predecessor in several areas. The LLM will be the default option in the consumer tiers of the company’

model-releasessiliconangle
1 Jul 2026
Model Releases

Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents

DGX agent

arXiv:2606.22528v2 Announce Type: replace Abstract: Modern LLM agents increasingly rely on context compaction, summarization, or eviction to keep long-running sessions within a token budget. We show t

model-releasesarxiv-cs-ai
30 Jun 2026
Safety

Lateral String Stability for Vehicle Platoons

DGX agent

arXiv:2606.29677v1 Announce Type: new Abstract: Connected and automated vehicle (CAV) platooning promises gains in energy efficiency and traffic throughput and, most critically, in safety. These safet

safetyarxiv-cs-ro
30 Jun 2026
Safety

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models

DGX agent

arXiv:2606.28153v1 Announce Type: cross Abstract: Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood. We provide evidence that attacks do not comprehensively

safetyarxiv-cs-ai
29 Jun 2026
Local Ai

Do Safety Guardrails Need to Reason? LeanGuard: A Fast and Light Approach for Robust Moderation

DGX agent

arXiv:2606.26686v1 Announce Type: new Abstract: In order to screen a prompt or a response, the recent guardrail methods generate a chain-of-thought (CoT) before they issue a verdict. This design follo

local-aiarxiv-cs-ai
26 Jun 2026
Safety

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy

DGX agent

arXiv:2606.26527v1 Announce Type: new Abstract: Transfer learning improves policy learning efficiency by reusing knowledge from source tasks, providing a feasible paradigm for safe and efficient auton

safetyarxiv-cs-lg
26 Jun 2026
Safety

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model

DGX agent

arXiv:2606.20698v1 Announce Type: new Abstract: Safe control is a prerequisite for real-world embodied intelligence, for which safe reinforcement learning has emerged as a promising paradigm. However,

safetyarxiv-cs-ro
23 Jun 2026
Safety

Runtime Enforcement of Hybrid System Properties

DGX agent

arXiv:2606.12022v1 Announce Type: cross Abstract: Runtime enforcement has emerged as a promising approach for ensuring the safety of autonomous and cyber-physical systems operating in uncertain and dy

safetyarxiv-cs-ai
11 Jun 2026
Model Releases

An Integrated Roadside Sensing and Communication Framework for Vulnerable Road User Safety at Signalized Intersections

DGX agent

arXiv:2606.07016v1 Announce Type: cross Abstract: Vulnerable road users (VRUs) account for approximately half of urban traffic deaths globally, with intersections concentrating a disproportionate shar

model-releasesarxiv-cs-cv
8 Jun 2026
Safety

Explainably Safe Reinforcement Learning

DGX agent

arXiv:2606.04634v1 Announce Type: new Abstract: Trust in a decision-making system requires both safety guarantees and the ability to interpret and understand its behavior. This is particularly importa

safetyarxiv-cs-lg
4 Jun 2026
Safety

MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs

DGX agent

arXiv:2511.07107v3 Announce Type: replace Abstract: Ensuring the safety of Large Language Models (LLMs) is critical for real-world deployment. However, current safety measures often fail to address im

safetyarxiv-cs-ai
4 Jun 2026
Model Releases

MultiTurnPSB: Evaluating Multi-Turn Jailbreak Attacks an dClassifier-Based Defenses for Medical AI Safety

DGX agent

arXiv:2606.02630v1 Announce Type: cross Abstract: Patient-facing medical chatbots are commonly evaluated on single-turn prompts, yet real users push back after refusals, add urgency, and invoke author

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning

DGX agent

arXiv:2606.02443v1 Announce Type: cross Abstract: Between the first visible sign of danger and the moment an accident occurs, there is often a window where intervention remains possible. Video-capable

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Quality-Diversity Evolution for Discovering Diverse Vulnerabilities in LLM Safety

DGX agent

arXiv:2606.00801v1 Announce Type: cross Abstract: Current approaches to LLM adversarial testing suffer from coverage gaps: manual red-teaming does not scale, LLM-as-attacker methods exhibit mode colla

model-releasesarxiv-cs-cl
2 Jun 2026
Safety

Robust Shielding for Safe Reinforcement Learning

DGX agent

arXiv:2606.00270v1 Announce Type: new Abstract: Shielding is an effective approach to formally guarantee the safety of reinforcement learning agents in Markov decision processes (MDPs). However, exist

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

Device Context Protocol: A Compact, Safety-First Architecture for LLM-Driven Control of Constrained Devices

DGX agent

arXiv:2605.26159v1 Announce Type: cross Abstract: Large language models are increasingly used as orchestrators of external tools via the Model Context Protocol (MCP), but MCP is built for software ser

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies

DGX agent

arXiv:2507.06513v3 Announce Type: replace Abstract: Advances in vision-based sensors and computer vision algorithms have significantly improved the analysis and understanding of traffic scenarios. To

model-releasesarxiv-cs-cv
27 May 2026
← Previous
1…1516171819…297
Next →