AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Safety

SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety

DGX agent

arXiv:2605.05704v2 Announce Type: replace-cross Abstract: Recent advances in foundation models have transformed LLMs from passive conversational systems into autonomous agents capable of reasoning and

safetyarxiv-cs-ai
25 May 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

DGX agent

arXiv:2605.22643v1 Announce Type: new Abstract: Background. Traditional safety benchmarks for language models evaluate generated text: whether a model outputs toxic language, reproduces bias, or follo

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety

DGX agent

arXiv:2605.20203v1 Announce Type: cross Abstract: As older adults increasingly use LLM-based chatbots for companionship and assistance, a safety gap is emerging. Older adults may face vulnerabilities

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

HealthCraft: A Reinforcement Learning Safety Environment for Emergency Medicine

DGX agent

arXiv:2605.21496v1 Announce Type: cross Abstract: Frontier language models are being deployed into clinical workflows faster than the infrastructure to evaluate them safely. Static medical-QA benchmar

model-releasesarxiv-cs-cl
22 May 2026
Safety

Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework

DGX agent

arXiv:2603.11768v2 Announce Type: replace Abstract: Long-term memory has emerged as a foundational component of autonomous Large Language Model (LLM) agents, enabling continuous adaptation, lifelong m

safetyarxiv-cs-ai
20 May 2026
Model Releases

Measuring Safety Alignment Effects in Autonomous Security Agents

DGX agent

arXiv:2605.19722v1 Announce Type: cross Abstract: Do stock safety-aligned language models and their uncensored or abliterated derivatives behave differently when run as autonomous security agents? Sin

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

AgentWall: A Runtime Safety Layer for Local AI Agents

DGX agent

arXiv:2605.16265v1 Announce Type: new Abstract: The safety of autonomous AI agents is increasingly recognized as a critical open problem. As agents transition from passive text generators to active ac

model-releasesarxiv-cs-ai
19 May 2026
Safety

Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation

DGX agent

arXiv:2605.17268v1 Announce Type: new Abstract: We present the first systematic study of faithfulness in Vision-Language-Action (VLA) driving models, analyzing 300 Alpamayo-R1-10B inferences across 10

safetyarxiv-cs-ai
19 May 2026
Model Releases

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents

DGX agent

arXiv:2601.17887v2 Announce Type: replace Abstract: Long-term memory enables large language model (LLM) agents to support personalized and sustained interactions. However, most work on personalized ag

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing

DGX agent

arXiv:2605.10146v1 Announce Type: new Abstract: Large language models (LLMs) increasingly rely on knowledge editing to support knowledge-intensive reasoning, but this flexibility also introduces criti

model-releasesarxiv-cs-ai
12 May 2026
Safety

SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints

DGX agent

arXiv:2512.23770v3 Announce Type: replace-cross Abstract: In safety-critical domains, reinforcement learning (RL) agents must often satisfy strict, zero-cost safety constraints while accomplishing tas

safetyarxiv-cs-ai
11 May 2026
Model Releases

Safety and accuracy follow different scaling laws in clinical large language models

DGX agent

arXiv:2605.04039v1 Announce Type: new Abstract: Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation

model-releasesarxiv-cs-cl
6 May 2026
Safety

Set-Based Training of Neural Barrier Certificates for Safety Verification of Dynamical Systems

DGX agent

arXiv:2605.02526v1 Announce Type: cross Abstract: Barrier certificates are scalar functions over the state space of dynamical systems that separate all unsafe states from all reachable states. The exi

safetyarxiv-cs-ai
6 May 2026
Model Releases

A decoupled diffusion planner that adapts to changing cost limits by using cost-conditioned generation for safety and reward gradients for performance

DGX agent

arXiv:2605.02777v1 Announce Type: new Abstract: Offline safe reinforcement learning often requires policies to adapt at deployment time to safety budgets that vary across episodes or change within a s

model-releasesarxiv-cs-lg
5 May 2026
Safety

Value Functions for Temporal Logic: Optimal Policies and Safety Filters

DGX agent

arXiv:2605.01051v1 Announce Type: cross Abstract: While Bellman equations for basic reach, avoid, and reach-avoid problems are well studied, the relationship between value optimality and policy optima

safetyarxiv-cs-lg
5 May 2026
Safety

When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent Systems

DGX agent

arXiv:2605.01133v1 Announce Type: cross Abstract: Large language model (LLM)-powered multi-agent systems (MAS) enable agents to communicate and share information, achieving strong performance on compl

safetyarxiv-cs-lg
5 May 2026
Model Releases

Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex

DGX agent

arXiv:2604.14858v2 Announce Type: replace Abstract: As agent systems move into increasingly diverse execution settings, trajectory-level safety evaluation and diagnosis require benchmarks that evolve

model-releasesarxiv-cs-ai
30 Apr 2026
Safety

Instantaneous Planning, Control and Safety for Navigation in Unknown Underwater Spaces

DGX agent

arXiv:2604.05310v2 Announce Type: replace Abstract: Navigating autonomous underwater vehicles (AUVs) in unknown environments is significantly challenging due to poor visibility, weak signal transmissi

safetyarxiv-cs-ro
29 Apr 2026
Model Releases

Evaluating whether AI models would sabotage AI safety research

DGX agent

arXiv:2604.24618v1 Announce Type: new Abstract: We evaluate the propensity of frontier models to sabotage or refuse to assist with safety research when deployed as AI research agents within a frontier

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

How Sensitive Are Safety Benchmarks to Judge Configuration Choices?

DGX agent

arXiv:2604.24074v1 Announce Type: new Abstract: Safety benchmarks such as HarmBench rely on LLM judges to classify model responses as harmful or safe, yet the judge configuration, namely the combinati

model-releasesarxiv-cs-cl
28 Apr 2026
Model Releases

Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models

DGX agent

arXiv:2603.21697v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) extend text-only LLMs with visual reasoning, but also introduce new safety failure modes under visual

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety

DGX agent

arXiv:2604.18487v1 Announce Type: new Abstract: The Adversarial Humanities Benchmark (AHB) evaluates whether model safety refusals survive a shift away from familiar harmful prompt forms. Starting fro

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation

DGX agent

arXiv:2602.07954v4 Announce Type: replace Abstract: As Large Language Models (LLMs) become increasingly deployed in Polish language applications, the need for efficient and accurate content safety cla

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration

DGX agent

arXiv:2604.16541v1 Announce Type: new Abstract: Recent advancements in Large Generative Models (LGMs) have revolutionized multi-modal generation. However, generating illustrated storybooks remains an

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

Structure-Aware Diversity Pursuit as an AI Safety Strategy against Homogenization

DGX agent

arXiv:2601.06116v2 Announce Type: replace-cross Abstract: Generative AI models reproduce the biases in the training data and can further amplify them through mode collapse. We refer to the resulting h

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA

DGX agent

arXiv:2604.06173v1 Announce Type: cross Abstract: Legal QA benchmarks have predominantly focused on case law, overlooking the unique challenges of statute-centric regulatory reasoning. In statutory do

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding

DGX agent

arXiv:2604.07879v1 Announce Type: new Abstract: Diffusion-based image generation models have advanced rapidly but pose a safety risk due to their potential to generate Not-Safe-For-Work (NSFW) content

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity

DGX agent

arXiv:2608.09443v1 Announce Type: new Abstract: Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbidity depend on

model-releasesarxiv-cs-ai
11 Aug 2026
Safety

Learning Multi-Timescale Interventions under Safety and Resource Constraints

DGX agent

arXiv:2508.03875v2 Announce Type: replace Abstract: Many sequential decision problems offer qualitatively different ways of influencing the environment: some interventions act immediately, whereas oth

safetyarxiv-cs-lg
11 Aug 2026
Model Releases

Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

DGX agent

arXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection

DGX agent

arXiv:2509.23690v2 Announce Type: replace-cross Abstract: Safety hazards in the home are a leading cause of preventable domestic injuries, motivating an automated inspector that actively explores a ho

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models

DGX agent

arXiv:2608.03201v1 Announce Type: new Abstract: Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs. We audit

model-releasesarxiv-cs-ai
5 Aug 2026
Safety

A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard)

DGX agent

arXiv:2608.00180v1 Announce Type: new Abstract: Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard. Training a safety guard with RL means optimizing two

safetyarxiv-cs-cl
4 Aug 2026
Safety

Safety Verification of Wait-Only Non-Blocking Broadcast Protocols

DGX agent

arXiv:2403.18591v3 Announce Type: replace-cross Abstract: Broadcast protocols are programs designed to be executed by networks of processes. Each process runs the same protocol, and communication betw

safetyarxiv-cs-cl
31 Jul 2026
Model Releases

EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

DGX agent

arXiv:2607.26518v1 Announce Type: new Abstract: Reliable visual safety understanding in real-world scenarios demands more than just object recognition; it requires causal reasoning under epistemic unc

model-releasesarxiv-cs-cv
30 Jul 2026
Local Ai

Distribution-Specific Curvature Control with Finite-Sample Guarantees for Open-Weight Safety

DGX agent

arXiv:2607.22929v1 Announce Type: new Abstract: A short fine-tuning run can undo the safety guards of an open-weight model---retraining a refusal-trained assistant to aid weapons development or produc

local-aiarxiv-cs-lg
28 Jul 2026
Model Releases

OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows

DGX agent

arXiv:2510.24411v3 Announce Type: replace Abstract: Computer-using agents powered by Vision-Language Models (VLMs) have demonstrated human-like capabilities in operating digital environments like mobi

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

HERMES: Heterogeneous Edge-Relational Multi-Head Embedded SSM Attention for Traffic Conflict Prediction at Signalized Intersections

DGX agent

arXiv:2607.20505v1 Announce Type: cross Abstract: Surrogate safety measures (SSMs) enable proactive traffic safety assessment, but many existing methods evaluate pairwise interactions independently or

safetyarxiv-cs-lg
24 Jul 2026
Model Releases

Geometry-Guided Constraint Learning for LLM Safety Classification

DGX agent

arXiv:2607.19366v1 Announce Type: new Abstract: Safety as Polytope (SaP) learns linear half-space constraints in LLM hidden space but requires per-category tuning of the constraint count K. We show th

model-releasesarxiv-cs-ai
23 Jul 2026
Safety

Sound Probabilistic Safety Bounds for Large Language Models

DGX agent

arXiv:2607.20286v1 Announce Type: cross Abstract: We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given pr

safetyarxiv-cs-ai
23 Jul 2026
Model Releases

Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric

DGX agent

arXiv:2607.12469v1 Announce Type: cross Abstract: Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

SoftVTBench: A Safety-Aware Visuo-Tactile Benchmark for Physically Constrained Robotic Manipulation of Deformable Objects

DGX agent

arXiv:2607.04234v1 Announce Type: cross Abstract: Deformable object manipulation poses challenges beyond task completion: successful execution must also maintain safe physical interaction, holding the

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

VLSA: Vision-Language-Action Models with Plug-and-Play Safety Constraint Layer

DGX agent

arXiv:2512.11891v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in generalizing across diverse robotic manipulation tasks. However, de

model-releasesarxiv-cs-ro
3 Jul 2026
Model Releases

High-Speed Vision-Based Flight in Clutter with Safety-Shielded Reinforcement Learning

DGX agent

arXiv:2602.08653v2 Announce Type: replace Abstract: Quadrotor unmanned aerial vehicles (UAVs) are increasingly deployed in complex missions that demand reliable autonomous navigation and robust obstac

model-releasesarxiv-cs-ro
1 Jul 2026
Model Releases

Bridging the Gap Between Image Restoration and Navigational Safety in Hazy Conditions: A New Visibility Estimation Metric for Maritime Surveillance

DGX agent

arXiv:2606.30049v1 Announce Type: new Abstract: Visibility distance is critical to maritime navigational safety because it determines the effective observation range of shipborne and shore-based monit

model-releasesarxiv-cs-cv
30 Jun 2026
Safety

HJ-SafeDMP: Hamilton-Jacobi Reachability-Guided Dynamic Movement Primitives for Provably Safe Robot Motion

DGX agent

arXiv:2606.28995v1 Announce Type: new Abstract: Robots deployed in safety-critical environments must execute motions that are simultaneously robust to disturbances and provably safe from collisions. D

safetyarxiv-cs-ro
30 Jun 2026
Safety

Beyond Feedforward Networks: Reentry Neural Systems as the Fundamental Basis of Subjecthood and Intrinsic Safety of Next-Generation AGI

DGX agent

arXiv:2606.26406v1 Announce Type: cross Abstract: We propose a complete architectural blueprint for safe artificial general intelligence based on a closed reentry loop (D I cycle). In contrast to feed

safetyarxiv-cs-ai
26 Jun 2026
Model Releases

Distributed Model Predictive Control with Adaptive Safety Zones for Multi-Fleet Drone Operations

DGX agent

arXiv:2606.20651v1 Announce Type: cross Abstract: Autonomous drone swarms in space-constrained environments such as warehouses, inspection corridors, and urban delivery routes must share limited airsp

model-releasesarxiv-cs-ro
23 Jun 2026
← Previous
1…1011121314…255
Next →