AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Model Releases

Your Model Already Knows: Attention-Guided Safety Filter for Vision-Language-Action Models

DGX agent

arXiv:2606.09749v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated impressive end-to-end performance across a variety of robotic manipulation tasks. However, these

model-releasesarxiv-cs-lg
9 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Data Flow Control: Data Safety Policies for AI Agents

DGX agent

arXiv:2606.05679v1 Announce Type: cross Abstract: Agents increasingly generate SQL, orchestrate pipelines, and automate data analysis on behalf of users. While recent work improves query correctness,

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Safety by narrow control has shown to fail many times. Need more transparency on the absolute frontier, and openness close behind.

DGX agent

Safety by narrow control has shown to fail many times. Need more transparency on the absolute frontier, and openness close behind. I found another API that offers claude-oceanus-v1-p the pricing and t

model-releasesclem-delangue--x
4 Jun 2026
Local Ai

Edge-Aware and Content-Adaptive Infrared Gas Leak Detection for Industrial Safety Monitoring

DGX agent

arXiv:2512.23234v3 Announce Type: replace-cross Abstract: Infrared gas leak detection is important for industrial safety and environmental monitoring, but automatic detection remains challenging becau

local-aiarxiv-cs-ai
3 Jun 2026
Model Releases

EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents

DGX agent

arXiv:2606.03678v1 Announce Type: new Abstract: Generating safety-critical scenarios is essential for validating and improving autonomous driving systems, yet it inherently requires maximizing adversa

model-releasesarxiv-cs-ai
3 Jun 2026
Local Ai

From 3D Perception to Safety Reasoning: A Graph-Based Framework for Real-Time Underground Mine Monitoring

DGX agent

arXiv:2606.03460v1 Announce Type: new Abstract: Underground coal mining requires personnel and heavy equipment to operate within shared, confined, and poorly illuminated spaces where hazards such as e

local-aiarxiv-cs-cv
3 Jun 2026
Model Releases

glad to know Mythos' safety concerns have been addressed right as Anthropic also secured tens of billions in inference compute 👍

DGX agent

glad to know Mythos' safety concerns have been addressed right as Anthropic also secured tens of billions in inference compute 👍 JUST IN: Anthropic announces it will roll out Claude Mythos “in the com

model-releasesjeremy-howard--x
28 May 2026
Model Releases

StructBreak: Structural Cognitive Overload-Induced Safety Failures in MLLMs

DGX agent

arXiv:2605.25534v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) excel at structural reasoning yet suffer from a sharp logical brittleness in structural consistency. We term th

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

DGX agent

arXiv:2602.12316v2 Announce Type: replace Abstract: Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely ev

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Steins;Gate Drive: Semantic Safety Arbitration over Structured Futures for Latency-Decoupled LLM Planning

DGX agent

arXiv:2605.22456v1 Announce Type: new Abstract: Cloud-hosted LLM driver agents provide useful semantic judgments, but their inference latency exceeds stepwise vehicle-control windows. Learned world mo

model-releasesarxiv-cs-ro
22 May 2026
Model Releases

Adversarial Stress Testing of SPARK Humanoid Safety Filters

DGX agent

arXiv:2605.19009v1 Announce Type: new Abstract: Humanoid robots are difficult to deploy safely because they have high-dimensional bodies, many collision constraints, and must operate near people and o

model-releasesarxiv-cs-ro
20 May 2026
Local Ai

Conservative AI for Safety-Sensitive Medical Image Restoration: Residual-Bounded CT-CTA Enhancement for Intracranial Aneurysm-Relevant Signal Recovery

DGX agent

arXiv:2605.16458v1 Announce Type: cross Abstract: Image restoration models are increasingly applied to degraded medical scans, but in safety-sensitive settings they must improve image quality without

local-aiarxiv-cs-ai
19 May 2026
Model Releases

EvoQRE: Modeling Bounded Rationality in Safety-Critical Traffic Simulation via Evolutionary Quantal Response Equilibrium

DGX agent

arXiv:2601.05653v2 Announce Type: replace Abstract: Existing traffic simulation frameworks for autonomous vehicles typically rely on imitation learning or game-theoretic approaches that solve for Nash

model-releasesarxiv-cs-ro
19 May 2026
Model Releases

🎉 We published a new AI safety study: shopping agents fall for whimsical attacks and lose money. A whimsical attack is an absurd scenario a…

DGX agent

🎉 We published a new AI safety study: shopping agents fall for whimsical attacks and lose money. A whimsical attack is an absurd scenario a human would never try on another human. In one run, GPT-5.1

model-releasesemad-mostaque--x
13 May 2026
Model Releases

CommandSwarm: Safety-Aware Natural Language-to-Behavior-Tree Generation for Robotic Swarms

DGX agent

arXiv:2605.07764v1 Announce Type: new Abstract: Natural-language interfaces can make swarm robotics more accessible to non-expert operators, but they must translate ambiguous user intent into executab

model-releasesarxiv-cs-ro
11 May 2026
Model Releases

Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with…

DGX agent

Ever wished your agent could read PDFs, images, and Office documents as easily as plain text? Or combine the safety of a secure sandbox with the full power of Bash access? We built exactly that. Meet

model-releasesjerry-liu--x
11 May 2026
Model Releases

Anthropic details how it improved Claude's safety training after finding agentic misalignment in older models, such as Opus 4 blackmailing engineers (Anthropic)

DGX agent

Anthropic: Anthropic details how it improved Claude's safety training after finding agentic misalignment in older models, such as Opus 4 blackmailing engineers — Last year, we released a case study on

model-releasestechmeme
9 May 2026
Model Releases

Self-Mined Hardness for Safety Fine-Tuning

DGX agent

arXiv:2605.03226v1 Announce Type: new Abstract: Safety fine-tuning of language models typically requires a curated adversarial dataset. We take a different approach: score each candidate prompt's diff

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

ARMOR 2025: A Military-Aligned Benchmark for Evaluating Large Language Model Safety Beyond Civilian Contexts

DGX agent

arXiv:2605.00245v1 Announce Type: new Abstract: Large language models (LLMs) are now being explored for defense applications that require reliable and legally compliant decision support. They also hol

model-releasesarxiv-cs-ai
5 May 2026
Model Releases

SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs

DGX agent

arXiv:2604.22134v1 Announce Type: new Abstract: Large Language Models (LLMs) have been widely explored in educational scenarios. We identify a critical vulnerability in current educational LLMs, pedag

model-releasesarxiv-cs-cl
27 Apr 2026
Model Releases

SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models

DGX agent

arXiv:2604.19638v1 Announce Type: new Abstract: Multimodal Large Language Models are increasingly adopted as autonomous agents in interactive environments, yet their ability to proactively address saf

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models

DGX agent

arXiv:2604.16980v1 Announce Type: new Abstract: Background: Large language models (LLMs) are increasingly proposed for diagnostic support, but few evaluations use real-world multimodal inpatient data,

model-releasesarxiv-cs-lg
21 Apr 2026
Local Ai

I wrapped my Ollama agent with deterministic safety checks — here's the setup (catches bad JSON, prompt injection, and refusals before they hit your app) apache 2.0 [GitHub: https://github.com/qaysSE/AG-X]

DGX agent

This GitHub project (AG-X) implements deterministic safety checks for Ollama agents, designed to catch malformed JSON responses, detect prompt injection attempts, and handle model refusals before they

local-air-ollama
19 Apr 2026
Model Releases

Persona Non Grata: Single-Method Safety Evaluation Is Incomplete for Persona-Imbued LLMs

DGX agent

arXiv:2604.11120v1 Announce Type: new Abstract: Personality imbuing customizes LLM behavior, but safety evaluations almost always study prompt-based personas alone. We show this is incomplete: prompti

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

STARS: Skill-Triggered Audit for Request-Conditioned Invocation Safety in Agent Systems

DGX agent

arXiv:2604.10286v1 Announce Type: new Abstract: Autonomous language-model agents increasingly rely on installable skills and tools to complete user tasks. Static skill auditing can expose capability s

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents

DGX agent

arXiv:2608.09476v1 Announce Type: cross Abstract: Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavio

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses

DGX agent

arXiv:2608.06984v1 Announce Type: cross Abstract: Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts. However,

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

Risky Business: Measuring The Faithfulness-Safety Tension

DGX agent

arXiv:2608.03745v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring. However, monitoring relies on faithfulness, i.e., the model output str

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

DGX agent

arXiv:2409.07314v4 Announce Type: replace Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become satur

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Risk-Aware Motion Planning with Learned Trajectory Primitives and Probabilistic Safety Assessment

DGX agent

arXiv:2607.26802v1 Announce Type: new Abstract: This paper presents a radial basis function network (RBFN)-informed motion planning framework for safe and efficient urban autonomous driving. The propo

model-releasesarxiv-cs-ro
30 Jul 2026
Model Releases

AMPBench-MT: A Homology-Controlled Benchmark for Antimicrobial Peptide Potency, Spectrum, and Safety Prediction

DGX agent

arXiv:2607.25518v1 Announce Type: new Abstract: Computational AMP discovery is often evaluated through AMP/non-AMP recognition, yet follow-up decisions depend on assay-derived evidence such as target-

model-releasesarxiv-cs-lg
29 Jul 2026
Safety

Is Your Safe Controller Actually Safe? A Critical Review of CBF Tautologies and Hidden Assumptions

DGX agent

arXiv:2603.06954v2 Announce Type: replace Abstract: This tutorial provides a critical review of the practical application of Control Barrier Functions (CBFs) in robotic safety. While the theoretical f

safetyarxiv-cs-ro
24 Jul 2026
Model Releases

OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills

DGX agent

arXiv:2607.20121v1 Announce Type: new Abstract: LLM-based agents leverage third-party skills to extend their capabilities in open-world scenarios. However, third-party skills can introduce extra secur

model-releasesarxiv-cs-cl
23 Jul 2026
Safety

Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies

DGX agent

arXiv:2603.15136v2 Announce Type: replace-cross Abstract: Offline safe reinforcement learning (RL) seeks reward-maximizing policies from static datasets under strict safety constraints. Existing metho

safetyarxiv-cs-ai
10 Jul 2026
Model Releases

Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety

DGX agent

arXiv:2607.07695v1 Announce Type: new Abstract: We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

DGX agent

arXiv:2607.07229v1 Announce Type: new Abstract: Prior work has shown that chain-of-thought (CoT) reasoning is often unfaithful: a model's stated reasoning does not reliably reflect the process that pr

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives

DGX agent

arXiv:2604.16870v2 Announce Type: replace-cross Abstract: AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP). These tool calls are the age

model-releasesarxiv-cs-ai
7 Jul 2026
Safety

Addressing Over-Refusal in LLMs with Competing Rewards

DGX agent

arXiv:2606.31748v1 Announce Type: new Abstract: Safety training on language models often induces over-refusal: improved safety on harmful prompts at the cost of increased refusal on harmless ones. Tho

safetyarxiv-cs-lg
1 Jul 2026
Safety

ShardNet: Training Neural Controllers with Hard, Non-Convex Constraints

DGX agent

arXiv:2606.30935v1 Announce Type: cross Abstract: While neural network control policies are powerful, their deployment on safety critical systems depends on ensuring that they obey strict constraints.

safetyarxiv-cs-lg
1 Jul 2026
Safety

Safe Learning Control with Optimality and Stability Guarantees

DGX agent

arXiv:2501.15373v2 Announce Type: replace-cross Abstract: Merely pursuing performance may adversely affect safety, while a conservative policy for safe exploration will degrade the performance. How to

safetyarxiv-cs-lg
25 Jun 2026
Model Releases

Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming

DGX agent

arXiv:2606.05233v1 Announce Type: cross Abstract: Recent computer-using-agent (CUA) red-teaming papers report prompt-injection attack success rates (ASR) of 42-98%, but these headline numbers cluster

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety

DGX agent

arXiv:2606.04867v1 Announce Type: new Abstract: As AI companion platforms such as Replika and Character.AI rapidly grow, concerns about unsafe human-AI interactions have intensified. This study introd

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

In policy paper, OpenAI diverges from White House on AI safety

DGX agent

OpenAI Group PBC’s newly released proposal for how advanced artificial intelligence should be regulated differs slightly from the Trump administration’s executive order, also released this week. Relea

model-releasessiliconangle
4 Jun 2026
Model Releases

A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models

DGX agent

arXiv:2606.00027v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation

DGX agent

arXiv:2606.01481v1 Announce Type: new Abstract: With the rapid advancements in text-to-image diffusion models, generative video models (T2V models) like Sora can now produce short synthetic videos fro

model-releasesarxiv-cs-cv
2 Jun 2026
Safety

Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models

DGX agent

arXiv:2605.24550v1 Announce Type: new Abstract: Fine-tuning-as-a-Service (FaaS) enables personalization of large language models (LLMs), but it can weaken safety-alignment under harmful fine-tuning at

safetyarxiv-cs-ai
26 May 2026
Safety

Jailbreaking on Text-to-Video Models via Scene Splitting Strategy

DGX agent

arXiv:2509.22292v2 Announce Type: replace-cross Abstract: Along with the rapid advancement of numerous Text-to-Video (T2V) models, growing concerns have emerged regarding their safety risks. While rec

safetyarxiv-cs-ai
20 May 2026
Model Releases

Passive Construction Site Safety Monitoring via Persona-Scaffolded Adversarial Chain-of-Thought VLM Verification

DGX agent

arXiv:2605.19869v1 Announce Type: cross Abstract: Construction remains the deadliest industry sector in the United States, with 1,055 fatal worker injuries recorded in 2023, and the majority preventab

model-releasesarxiv-cs-ai
20 May 2026
← Previous
1…1415161718…297
Next →