AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Model Releases

Children's English Reading Story Generation via Supervised Fine-Tuning of Compact LLMs with Controllable Difficulty and Safety

DGX agent

arXiv:2605.13709v1 Announce Type: cross Abstract: Large Language Models (LLMs) are widely applied in educational practices, such as for generating children's stories. However, the generated stories ar

model-releasesarxiv-cs-ai
14 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Attention Is Where You Attack

DGX agent

arXiv:2605.00236v1 Announce Type: cross Abstract: Safety-aligned large language models rely on RLHF and instruction tuning to refuse harmful requests, yet the internal mechanisms implementing safety b

model-releasesarxiv-cs-ai
5 May 2026
Safety

Safe-Support Q-Learning: Learning without Unsafe Exploration

DGX agent

arXiv:2604.25379v1 Announce Type: new Abstract: Ensuring safety during reinforcement learning (RL) training is critical in real-world applications where unsafe exploration can lead to devastating outc

safetyarxiv-cs-lg
29 Apr 2026
Model Releases

CLIN-LLM: A Safety-Constrained Hybrid Framework for Clinical Diagnosis and Treatment Generation

DGX agent

arXiv:2510.22609v2 Announce Type: replace Abstract: Accurate symptom-to-disease classification and clinically grounded treatment recommendations remain challenging, particularly in heterogeneous patie

model-releasesarxiv-cs-ai
28 Apr 2026
Safety

Control Barrier Functions Solved with Hierarchical Quadratic Programming for Safe Physical Human-Robot Interaction

DGX agent

arXiv:2604.23039v1 Announce Type: new Abstract: Physical human-robot interaction offers the potential to leverage human intelligence and robot physical capabilities to enable a range of exciting appli

safetyarxiv-cs-ro
28 Apr 2026
Safety

Driving risk emerges from the required two-dimensional joint evasive acceleration

DGX agent

arXiv:2604.17841v1 Announce Type: new Abstract: Most autonomous driving safety benchmarks use time-to-collision (TTC) to assess risk and guide safe behaviour. However, TTC-based methods treat risk as

safetyarxiv-cs-ro
21 Apr 2026
Safety

Integrated Wheel Sensor Communication using ESP32 -- A Contribution towards a Digital Twin of the Road System

DGX agent

arXiv:2509.04061v2 Announce Type: replace Abstract: While current onboard state estimation methods are adequate for most driving and safety-related applications, they do not provide insights into the

safetyarxiv-cs-ro
21 Apr 2026
Safety

Capability-Aware Heterogeneous Control Barrier Functions for Decentralized Multi-Robot Safe Navigation

DGX agent

arXiv:2604.13245v1 Announce Type: new Abstract: Safe navigation for multi-robot systems requires enforcing safety without sacrificing task efficiency under decentralized decision-making. Existing dece

safetyarxiv-cs-ro
16 Apr 2026
Safety

RACF: A Resilient Autonomous Car Framework with Object Distance Correction

DGX agent

arXiv:2604.12418v1 Announce Type: cross Abstract: Autonomous vehicles are increasingly deployed in safety-critical applications, where sensing failures or cyberphysical attacks can lead to unsafe oper

safetyarxiv-cs-ai
15 Apr 2026
Model Releases

Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers

DGX agent

arXiv:2603.28013v3 Announce Type: replace-cross Abstract: Multi-agent LLM systems are entering production -- processing documents, managing workflows, acting on behalf of users -- yet their resilience

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning

DGX agent

arXiv:2604.09452v1 Announce Type: cross Abstract: Safety guarantees are a prerequisite to the deployment of reinforcement learning (RL) agents in safety-critical tasks. Often, deployment environments

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces

DGX agent

arXiv:2604.05172v2 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly deployed to automate productivity tasks (e.g., email, scheduling, document management), but evalu

model-releasesarxiv-cs-ai
10 Apr 2026
Safety

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training

DGX agent

arXiv:2604.07754v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) raises significant ethical and safety concerns. While LLM alignment techniques are adopted to improve m

safetyarxiv-cs-cl
10 Apr 2026
Safety

Topological Feasibility Guarantees for Differentiable Predictive Control

DGX agent

arXiv:2608.10332v1 Announce Type: cross Abstract: Differentiable predictive control (DPC), a self-supervised learning approach for approximating explicit model predictive control (MPC) policies, offer

safetyarxiv-cs-lg
12 Aug 2026
Model Releases

InfoOps Bench: A live information operations safety benchmark

DGX agent

arXiv:2607.28503v3 Announce Type: replace Abstract: In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment

DGX agent

arXiv:2608.06110v1 Announce Type: new Abstract: This paper presents ECHO (Enhanced Care & Health Observer), a locally-deployable conversational health assistant for long-term chronic care management.

model-releasesarxiv-cs-ai
7 Aug 2026
Model Releases

Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load

DGX agent

arXiv:2608.05018v1 Announce Type: new Abstract: Short-term load forecasting (STLF) play a vital role in the electric power industry. It serves infrastructure that European and German law designate as

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity

DGX agent

arXiv:2608.02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is f

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Grasp Execution Without a Planner: Configuration-Space Grasp Distance Fields with Certified Safety & Guaranteed Quality

DGX agent

arXiv:2608.00600v1 Announce Type: new Abstract: Standard multifingered grasp execution architectures plan a collision-free trajectory to a selected grasp pose and track it with a feedback law. Executi

model-releasesarxiv-cs-ro
4 Aug 2026
Model Releases

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs

DGX agent

arXiv:2607.24392v1 Announce Type: cross Abstract: Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. W

model-releasesarxiv-cs-lg
28 Jul 2026
Model Releases

DocShield: Towards AI Document Safety via Evidence-Grounded Agentic Reasoning

DGX agent

arXiv:2604.02694v2 Announce Type: replace-cross Abstract: The rapid progress of generative AI has enabled increasingly realistic text-centric image forgeries, posing major challenges to document safet

model-releasesarxiv-cs-ai
23 Jul 2026
Safety

Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies

DGX agent

arXiv:2607.18325v1 Announce Type: new Abstract: Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and support decision-making during emergencies. Visi

safetyarxiv-cs-cv
23 Jul 2026
Safety

A Graph-Based Reinforcement Learning Approach with Frontier Potential Based Reward for Safe Cluttered Environment Exploration

DGX agent

arXiv:2504.11907v3 Announce Type: replace Abstract: Autonomous exploration of cluttered environments requires efficient exploration strategies that guarantee safety against potential collisions with u

safetyarxiv-cs-ro
7 Jul 2026
Safety

Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs

DGX agent

arXiv:2508.10031v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and

safetyarxiv-cs-ai
7 Jul 2026
Model Releases

Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents

DGX agent

arXiv:2606.22528v2 Announce Type: replace Abstract: Modern LLM agents increasingly rely on context compaction, summarization, or eviction to keep long-running sessions within a token budget. We show t

model-releasesarxiv-cs-ai
30 Jun 2026
Safety

Lateral String Stability for Vehicle Platoons

DGX agent

arXiv:2606.29677v1 Announce Type: new Abstract: Connected and automated vehicle (CAV) platooning promises gains in energy efficiency and traffic throughput and, most critically, in safety. These safet

safetyarxiv-cs-ro
30 Jun 2026
Safety

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models

DGX agent

arXiv:2606.28153v1 Announce Type: cross Abstract: Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood. We provide evidence that attacks do not comprehensively

safetyarxiv-cs-ai
29 Jun 2026
Local Ai

Do Safety Guardrails Need to Reason? LeanGuard: A Fast and Light Approach for Robust Moderation

DGX agent

arXiv:2606.26686v1 Announce Type: new Abstract: In order to screen a prompt or a response, the recent guardrail methods generate a chain-of-thought (CoT) before they issue a verdict. This design follo

local-aiarxiv-cs-ai
26 Jun 2026
Safety

Sample-efficient Transfer Reinforcement Learning via Adaptive Reward Shaping and Policy-Ratio Reweighting Strategy

DGX agent

arXiv:2606.26527v1 Announce Type: new Abstract: Transfer learning improves policy learning efficiency by reusing knowledge from source tasks, providing a feasible paradigm for safe and efficient auton

safetyarxiv-cs-lg
26 Jun 2026
Safety

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model

DGX agent

arXiv:2606.20698v1 Announce Type: new Abstract: Safe control is a prerequisite for real-world embodied intelligence, for which safe reinforcement learning has emerged as a promising paradigm. However,

safetyarxiv-cs-ro
23 Jun 2026
Safety

Runtime Enforcement of Hybrid System Properties

DGX agent

arXiv:2606.12022v1 Announce Type: cross Abstract: Runtime enforcement has emerged as a promising approach for ensuring the safety of autonomous and cyber-physical systems operating in uncertain and dy

safetyarxiv-cs-ai
11 Jun 2026
Model Releases

An Integrated Roadside Sensing and Communication Framework for Vulnerable Road User Safety at Signalized Intersections

DGX agent

arXiv:2606.07016v1 Announce Type: cross Abstract: Vulnerable road users (VRUs) account for approximately half of urban traffic deaths globally, with intersections concentrating a disproportionate shar

model-releasesarxiv-cs-cv
8 Jun 2026
Safety

Explainably Safe Reinforcement Learning

DGX agent

arXiv:2606.04634v1 Announce Type: new Abstract: Trust in a decision-making system requires both safety guarantees and the ability to interpret and understand its behavior. This is particularly importa

safetyarxiv-cs-lg
4 Jun 2026
Safety

MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs

DGX agent

arXiv:2511.07107v3 Announce Type: replace Abstract: Ensuring the safety of Large Language Models (LLMs) is critical for real-world deployment. However, current safety measures often fail to address im

safetyarxiv-cs-ai
4 Jun 2026
Model Releases

MultiTurnPSB: Evaluating Multi-Turn Jailbreak Attacks an dClassifier-Based Defenses for Medical AI Safety

DGX agent

arXiv:2606.02630v1 Announce Type: cross Abstract: Patient-facing medical chatbots are commonly evaluated on single-turn prompts, yet real users push back after refusals, add urgency, and invoke author

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

PaSBench-Video: A Streaming Video Benchmark for Proactive Safety Warning

DGX agent

arXiv:2606.02443v1 Announce Type: cross Abstract: Between the first visible sign of danger and the moment an accident occurs, there is often a window where intervention remains possible. Video-capable

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Quality-Diversity Evolution for Discovering Diverse Vulnerabilities in LLM Safety

DGX agent

arXiv:2606.00801v1 Announce Type: cross Abstract: Current approaches to LLM adversarial testing suffer from coverage gaps: manual red-teaming does not scale, LLM-as-attacker methods exhibit mode colla

model-releasesarxiv-cs-cl
2 Jun 2026
Safety

Robust Shielding for Safe Reinforcement Learning

DGX agent

arXiv:2606.00270v1 Announce Type: new Abstract: Shielding is an effective approach to formally guarantee the safety of reinforcement learning agents in Markov decision processes (MDPs). However, exist

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

Device Context Protocol: A Compact, Safety-First Architecture for LLM-Driven Control of Constrained Devices

DGX agent

arXiv:2605.26159v1 Announce Type: cross Abstract: Large language models are increasingly used as orchestrators of external tools via the Model Context Protocol (MCP), but MCP is built for software ser

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies

DGX agent

arXiv:2507.06513v3 Announce Type: replace Abstract: Advances in vision-based sensors and computer vision algorithms have significantly improved the analysis and understanding of traffic scenarios. To

model-releasesarxiv-cs-cv
27 May 2026
Safety

Guiding Neuro-Symbolic Scenario Generation with Spatio-Temporal Logic

DGX agent

arXiv:2605.19038v1 Announce Type: cross Abstract: The rapid advancement of autonomous driving (AD) technologies has outpaced the development of robust safety evaluation methods. Conventional testing r

safetyarxiv-cs-lg
20 May 2026
Safety

LPG: Balancing Efficiency and Policy Reasoning in Latent Policy Guardrails

DGX agent

arXiv:2605.17329v1 Announce Type: cross Abstract: Guardrails are a critical safety layer for modern AI systems, but their operating regime is changing. As LLMs are deployed as customized assistants, s

safetyarxiv-cs-ai
19 May 2026
Local Ai

LiSA: Lifelong Safety Adaptation via Conservative Policy Induction

DGX agent

arXiv:2605.14454v1 Announce Type: cross Abstract: As AI agents move from chat interfaces to systems that read private data, call tools, and execute multi-step workflows, guardrails become a last line

local-aiarxiv-cs-cl
15 May 2026
Model Releases

Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks

DGX agent

arXiv:2605.14604v1 Announce Type: new Abstract: This position paper argues that effective tutoring requires corrective friction: surfacing misconceptions and challenging them supportively to drive con

model-releasesarxiv-cs-ai
15 May 2026
Safety

Metaphor Is Not All Attention Needs

DGX agent

arXiv:2605.12128v1 Announce Type: new Abstract: Large language models are increasingly deployed in safety-critical applications, where their ability to resist harmful instructions is essential. Althou

safetyarxiv-cs-cl
13 May 2026
Safety

Safe and Real-Time Consistent Planning for Autonomous Vehicles in Partially Observed Environments via Parallel Consensus Optimization

DGX agent

arXiv:2409.10310v3 Announce Type: replace Abstract: Ensuring safety and driving consistency is a significant challenge for autonomous vehicles operating in partially observed environments. This work i

safetyarxiv-cs-ro
12 May 2026
Local Ai

Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off

DGX agent

arXiv:2605.08878v1 Announce Type: cross Abstract: Aligned large language models (LLMs) remain vulnerable to jailbreak attacks. Recent mechanistic studies have identified latent features and representa

local-aiarxiv-cs-ai
12 May 2026
Safety

LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment

DGX agent

arXiv:2601.19487v2 Announce Type: replace Abstract: Safety-aligned LLMs suffer from two failure modes: jailbreak (answering harmful inputs) and over-refusal (declining benign queries). Existing vector

safetyarxiv-cs-lg
5 May 2026
← Previous
1…1314151617…255
Next →