AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Model Releases

Moral Safety in LLMs: Exposing Performative Compliance with Puzzled Cues

DGX agent

arXiv:2606.31644v1 Announce Type: new Abstract: As large language models take on morally consequential roles in healthcare, legal, and hiring contexts, we need to examine whether their ethical behavio

model-releasesarxiv-cs-cl
1 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

DGX agent

arXiv:2606.28843v1 Announce Type: cross Abstract: Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work has shown th

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation

DGX agent

arXiv:2606.25782v1 Announce Type: new Abstract: With the widespread adoption of large language models (LLMs) in chatbots and everyday applications, companies increasingly need guardrails that are effe

model-releasesarxiv-cs-cl
25 Jun 2026
Model Releases

SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety

DGX agent

arXiv:2606.18936v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly embedded in AI for Science (AI4Science) workflows, from scientific question answering and literature a

model-releasesarxiv-cs-ai
25 Jun 2026
Model Releases

Online Shift Detection and Conformal Adaptation for Deployed Safety Classifiers

DGX agent

arXiv:2606.11949v1 Announce Type: new Abstract: We present an online monitoring system for distributional shift in deployed safety classifiers, using calibrated sequential statistics to detect when a

model-releasesarxiv-cs-lg
11 Jun 2026
Model Releases

Your Model Already Knows: Attention-Guided Safety Filter for Vision-Language-Action Models

DGX agent

arXiv:2606.09749v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated impressive end-to-end performance across a variety of robotic manipulation tasks. However, these

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

Data Flow Control: Data Safety Policies for AI Agents

DGX agent

arXiv:2606.05679v1 Announce Type: cross Abstract: Agents increasingly generate SQL, orchestrate pipelines, and automate data analysis on behalf of users. While recent work improves query correctness,

model-releasesarxiv-cs-ai
6 Jun 2026
Local Ai

Edge-Aware and Content-Adaptive Infrared Gas Leak Detection for Industrial Safety Monitoring

DGX agent

arXiv:2512.23234v3 Announce Type: replace-cross Abstract: Infrared gas leak detection is important for industrial safety and environmental monitoring, but automatic detection remains challenging becau

local-aiarxiv-cs-ai
3 Jun 2026
Model Releases

EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents

DGX agent

arXiv:2606.03678v1 Announce Type: new Abstract: Generating safety-critical scenarios is essential for validating and improving autonomous driving systems, yet it inherently requires maximizing adversa

model-releasesarxiv-cs-ai
3 Jun 2026
Local Ai

From 3D Perception to Safety Reasoning: A Graph-Based Framework for Real-Time Underground Mine Monitoring

DGX agent

arXiv:2606.03460v1 Announce Type: new Abstract: Underground coal mining requires personnel and heavy equipment to operate within shared, confined, and poorly illuminated spaces where hazards such as e

local-aiarxiv-cs-cv
3 Jun 2026
Model Releases

StructBreak: Structural Cognitive Overload-Induced Safety Failures in MLLMs

DGX agent

arXiv:2605.25534v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) excel at structural reasoning yet suffer from a sharp logical brittleness in structural consistency. We term th

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

DGX agent

arXiv:2602.12316v2 Announce Type: replace Abstract: Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely ev

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Steins;Gate Drive: Semantic Safety Arbitration over Structured Futures for Latency-Decoupled LLM Planning

DGX agent

arXiv:2605.22456v1 Announce Type: new Abstract: Cloud-hosted LLM driver agents provide useful semantic judgments, but their inference latency exceeds stepwise vehicle-control windows. Learned world mo

model-releasesarxiv-cs-ro
22 May 2026
Model Releases

Adversarial Stress Testing of SPARK Humanoid Safety Filters

DGX agent

arXiv:2605.19009v1 Announce Type: new Abstract: Humanoid robots are difficult to deploy safely because they have high-dimensional bodies, many collision constraints, and must operate near people and o

model-releasesarxiv-cs-ro
20 May 2026
Local Ai

Conservative AI for Safety-Sensitive Medical Image Restoration: Residual-Bounded CT-CTA Enhancement for Intracranial Aneurysm-Relevant Signal Recovery

DGX agent

arXiv:2605.16458v1 Announce Type: cross Abstract: Image restoration models are increasingly applied to degraded medical scans, but in safety-sensitive settings they must improve image quality without

local-aiarxiv-cs-ai
19 May 2026
Model Releases

EvoQRE: Modeling Bounded Rationality in Safety-Critical Traffic Simulation via Evolutionary Quantal Response Equilibrium

DGX agent

arXiv:2601.05653v2 Announce Type: replace Abstract: Existing traffic simulation frameworks for autonomous vehicles typically rely on imitation learning or game-theoretic approaches that solve for Nash

model-releasesarxiv-cs-ro
19 May 2026
Model Releases

CommandSwarm: Safety-Aware Natural Language-to-Behavior-Tree Generation for Robotic Swarms

DGX agent

arXiv:2605.07764v1 Announce Type: new Abstract: Natural-language interfaces can make swarm robotics more accessible to non-expert operators, but they must translate ambiguous user intent into executab

model-releasesarxiv-cs-ro
11 May 2026
Model Releases

Self-Mined Hardness for Safety Fine-Tuning

DGX agent

arXiv:2605.03226v1 Announce Type: new Abstract: Safety fine-tuning of language models typically requires a curated adversarial dataset. We take a different approach: score each candidate prompt's diff

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

ARMOR 2025: A Military-Aligned Benchmark for Evaluating Large Language Model Safety Beyond Civilian Contexts

DGX agent

arXiv:2605.00245v1 Announce Type: new Abstract: Large language models (LLMs) are now being explored for defense applications that require reliable and legally compliant decision support. They also hol

model-releasesarxiv-cs-ai
5 May 2026
Model Releases

SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs

DGX agent

arXiv:2604.22134v1 Announce Type: new Abstract: Large Language Models (LLMs) have been widely explored in educational scenarios. We identify a critical vulnerability in current educational LLMs, pedag

model-releasesarxiv-cs-cl
27 Apr 2026
Model Releases

SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models

DGX agent

arXiv:2604.19638v1 Announce Type: new Abstract: Multimodal Large Language Models are increasingly adopted as autonomous agents in interactive environments, yet their ability to proactively address saf

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models

DGX agent

arXiv:2604.16980v1 Announce Type: new Abstract: Background: Large language models (LLMs) are increasingly proposed for diagnostic support, but few evaluations use real-world multimodal inpatient data,

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs

DGX agent

arXiv:2604.11120v1 Announce Type: new Abstract: Personality imbuing customizes LLM behavior, but safety evaluations almost always study prompt-based personas alone. We show this is incomplete: prompti

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

STARS: Skill-Triggered Audit for Request-Conditioned Invocation Safety in Agent Systems

DGX agent

arXiv:2604.10286v1 Announce Type: new Abstract: Autonomous language-model agents increasingly rely on installable skills and tools to complete user tasks. Static skill auditing can expose capability s

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents

DGX agent

arXiv:2608.09476v1 Announce Type: cross Abstract: Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavio

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses

DGX agent

arXiv:2608.06984v1 Announce Type: cross Abstract: Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts. However,

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

Risky Business: Measuring The Faithfulness-Safety Tension

DGX agent

arXiv:2608.03745v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring. However, monitoring relies on faithfulness, i.e., the model output str

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

DGX agent

arXiv:2409.07314v4 Announce Type: replace Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become satur

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Risk-Aware Motion Planning with Learned Trajectory Primitives and Probabilistic Safety Assessment

DGX agent

arXiv:2607.26802v1 Announce Type: new Abstract: This paper presents a radial basis function network (RBFN)-informed motion planning framework for safe and efficient urban autonomous driving. The propo

model-releasesarxiv-cs-ro
30 Jul 2026
Model Releases

AMPBench-MT: A Homology-Controlled Benchmark for Antimicrobial Peptide Potency, Spectrum, and Safety Prediction

DGX agent

arXiv:2607.25518v1 Announce Type: new Abstract: Computational AMP discovery is often evaluated through AMP/non-AMP recognition, yet follow-up decisions depend on assay-derived evidence such as target-

model-releasesarxiv-cs-lg
29 Jul 2026
Safety

Is Your Safe Controller Actually Safe? A Critical Review of CBF Tautologies and Hidden Assumptions

DGX agent

arXiv:2603.06954v2 Announce Type: replace Abstract: This tutorial provides a critical review of the practical application of Control Barrier Functions (CBFs) in robotic safety. While the theoretical f

safetyarxiv-cs-ro
24 Jul 2026
Model Releases

OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills

DGX agent

arXiv:2607.20121v1 Announce Type: new Abstract: LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra secur

model-releasesarxiv-cs-cl
23 Jul 2026
Safety

Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies

DGX agent

arXiv:2603.15136v2 Announce Type: replace-cross Abstract: Offline safe reinforcement learning (RL) seeks reward-maximizing policies from static datasets under strict safety constraints. Existing metho

safetyarxiv-cs-ai
10 Jul 2026
Model Releases

Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety

DGX agent

arXiv:2607.07695v1 Announce Type: new Abstract: We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

DGX agent

arXiv:2607.07229v1 Announce Type: new Abstract: Prior work has shown that chain-of-thought (CoT) reasoning is often unfaithful: a model's stated reasoning does not reliably reflect the process that pr

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives

DGX agent

arXiv:2604.16870v2 Announce Type: replace-cross Abstract: AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP). These tool calls are the age

model-releasesarxiv-cs-ai
7 Jul 2026
Safety

Addressing Over-Refusal in LLMs with Competing Rewards

DGX agent

arXiv:2606.31748v1 Announce Type: new Abstract: Safety training on language models often induces over-refusal: improved safety on harmful prompts at the cost of increased refusal on harmless ones. Tho

safetyarxiv-cs-lg
1 Jul 2026
Safety

ShardNet: Training Neural Controllers with Hard, Non-Convex Constraints

DGX agent

arXiv:2606.30935v1 Announce Type: cross Abstract: While neural network control policies are powerful, their deployment on safety critical systems depends on ensuring that they obey strict constraints.

safetyarxiv-cs-lg
1 Jul 2026
Safety

Safe Learning Control with Optimality and Stability Guarantees

DGX agent

arXiv:2501.15373v2 Announce Type: replace-cross Abstract: Merely pursuing performance may adversely affect safety, while a conservative policy for safe exploration will degrade the performance. How to

safetyarxiv-cs-lg
25 Jun 2026
Model Releases

Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming

DGX agent

arXiv:2606.05233v1 Announce Type: cross Abstract: Recent computer-using-agent (CUA) red-teaming papers report prompt-injection attack success rates (ASR) of 42-98%, but these headline numbers cluster

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety

DGX agent

arXiv:2606.04867v1 Announce Type: new Abstract: As AI companion platforms such as Replika and Character.AI rapidly grow, concerns about unsafe human-AI interactions have intensified. This study introd

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models

DGX agent

arXiv:2606.00027v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation

DGX agent

arXiv:2606.01481v1 Announce Type: new Abstract: With the rapid advancements in text-to-image diffusion models, generative video models (T2V models) like Sora can now produce short synthetic videos fro

model-releasesarxiv-cs-cv
2 Jun 2026
Safety

Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models

DGX agent

arXiv:2605.24550v1 Announce Type: new Abstract: Fine-tuning-as-a-Service (FaaS) enables personalization of large language models (LLMs), but it can weaken safety-alignment under harmful fine-tuning at

safetyarxiv-cs-ai
26 May 2026
Safety

Jailbreaking on Text-to-Video Models via Scene Splitting Strategy

DGX agent

arXiv:2509.22292v2 Announce Type: replace-cross Abstract: Along with the rapid advancement of numerous Text-to-Video (T2V) models, growing concerns have emerged regarding their safety risks. While rec

safetyarxiv-cs-ai
20 May 2026
Model Releases

Passive Construction Site Safety Monitoring via Persona-Scaffolded Adversarial Chain-of-Thought VLM Verification

DGX agent

arXiv:2605.19869v1 Announce Type: cross Abstract: Construction remains the deadliest industry sector in the United States, with 1,055 fatal worker injuries recorded in 2023, and the majority preventab

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Graph-Regularized Sparse Autoencoders for LLM Safety Steering

DGX agent

arXiv:2512.06655v3 Announce Type: replace-cross Abstract: Sparse autoencoders (SAEs) are increasingly used to extract activation directions for inference-time steering, but their standard sparsity obj

model-releasesarxiv-cs-ai
18 May 2026
Model Releases

Invisible Orchestrators Suppress Protective Behavior and Dissociate Power-Holders: Safety Risks in Multi-Agent LLM Systems

DGX agent

arXiv:2605.13851v1 Announce Type: new Abstract: Multi-agent orchestration -- in which a hidden coordinator manages specialized worker agents -- is becoming the default architecture for enterprise AI d

model-releasesarxiv-cs-ai
15 May 2026
← Previous
1…1213141516…255
Next →