AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlog
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,234 results
Safety

Prompt-Induced Score Variance in Zero-Shot Binary Vision-Language Safety Classification

DGX agent

arXiv:2605.00326v1 Announce Type: new Abstract: Single-prompt first-token probabilities from zero-shot vision-language model (VLM) safety classifiers are treated as decision scores, but we show they a

safetyarxiv-cs-cl
4 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning…

DGX agent

Cool paper from Meta FAIR. It's on self-improving LLMs but on the pretraining side. (bookmark it) Most LLM safety, factuality, and reasoning fixes get bolted on at post-training. By then, the patterns

safetydair-ai--x
1 May 2026
Safety

Focus Session: Autonomous Systems Dependability in the era of AI: Design Challenges in Safety, Security, Reliability and Certification

DGX agent

arXiv:2604.27807v1 Announce Type: new Abstract: The design of embedded safety-critical systems such as those used in next-generation automotive and autonomous platforms, is increasingly challenged by

safetyarxiv-cs-ai
1 May 2026
Safety

Towards Neuro-symbolic Causal Rule Synthesis, Verification, and Evaluation Grounded in Legal and Safety Principles

DGX agent

arXiv:2604.28087v1 Announce Type: cross Abstract: Rule-based systems remain central in safety-critical domains but often struggle with scalability, brittleness, and goal misspecification. These limita

safetyarxiv-cs-ai
1 May 2026
Safety

A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?

DGX agent

arXiv:2505.10924v4 Announce Type: replace-cross Abstract: Recently, AI-driven interactions with computing devices have advanced from basic prototype tools to sophisticated, LLM-based systems that emul

safetyarxiv-cs-ai
30 Apr 2026
Safety

Recipes for Calibration Checks in Safety-Critical Applications

DGX agent

arXiv:2604.26479v1 Announce Type: cross Abstract: Safety-critical prediction systems, such as autonomous vehicles, weather forecasters, and medical monitors, commonly rely on probabilistic forecasters

safetyarxiv-cs-lg
30 Apr 2026
Safety

Safety and innovation are not mutually exclusive: for many companies, especially in high-trust industries, AI’s risks are also a hindrance t…

DGX agent

Safety and innovation are not mutually exclusive: for many companies, especially in high-trust industries, AI’s risks are also a hindrance to adoption. In this op-ed in the @FT, I underscore that Euro

safetyyoshua-bengio--x
27 Apr 2026
Safety

US child safety group NCMEC received 1.5M reports of suspected CSAM with ties to AI in 2025, a significant surge compared to 67,000 in 2024 and 4,700 in 2023 (Bloomberg)

DGX agent

Bloomberg: US child safety group NCMEC received 1.5M reports of suspected CSAM with ties to AI in 2025, a significant surge compared to 67,000 in 2024 and 4,700 in 2023 — William Michael Haslach was a

safetytechmeme
23 Apr 2026
Safety

Beyond Linear Probes: Dynamic Safety Monitoring for Language Models

DGX agent

arXiv:2509.26238v4 Announce Type: replace Abstract: Monitoring large language models' (LLMs) activations is an effective way to detect harmful requests before they lead to unsafe outputs. However, tra

safetyarxiv-cs-lg
22 Apr 2026
Model Releases

How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study

DGX agent

arXiv:2505.15404v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have achieved remarkable success on reasoning-intensive tasks such as mathematics and programming. However, their enha

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models

DGX agent

arXiv:2604.17691v1 Announce Type: new Abstract: Safety alignment in large language models is remarkably shallow: it is concentrated in the first few output tokens and reversible by fine-tuning on as f

model-releasesarxiv-cs-lg
21 Apr 2026
Local Ai

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts

DGX agent

arXiv:2604.16542v1 Announce Type: cross Abstract: Safety guardrails have become an active area of research in AI safety, aimed at ensuring the appropriate behavior of large language models (LLMs). How

local-aiarxiv-cs-cl
21 Apr 2026
Safety

Trajectory Planning for Safe Dual Control with Active Exploration

DGX agent

arXiv:2604.15507v1 Announce Type: new Abstract: Planning safe trajectories under model uncertainty is a fundamental challenge. Robust planning ensures safety by considering worst-case realizations, ye

safetyarxiv-cs-ro
20 Apr 2026
Safety

Calibrate-Then-Delegate: Safety Monitoring with Risk and Budget Guarantees via Model Cascades

DGX agent

arXiv:2604.14251v1 Announce Type: new Abstract: Monitoring LLM safety at scale requires balancing cost and accuracy: a cheap latent-space probe can screen every input, but hard cases should be escalat

safetyarxiv-cs-lg
17 Apr 2026
Safety

I have found that asking for a sestina regularly triggers Opus 4.7's safety guardrails. The forbidden poetic form!

DGX agent

Ethan Mollick reported that requesting Claude Opus 4.7 to write sestinas—a complex poetic form with strict structural requirements—frequently triggers the model's safety guardrails, suggesting the AI

safetyethan-mollick--x
16 Apr 2026
Safety

ContextLens: Modeling Imperfect Privacy and Safety Context for Legal Compliance

DGX agent

arXiv:2604.12308v1 Announce Type: new Abstract: Individuals' concerns about data privacy and AI safety are highly contextualized and extend beyond sensitive patterns. Addressing these issues requires

safetyarxiv-cs-cl
15 Apr 2026
Model Releases

LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety

DGX agent

arXiv:2604.12710v1 Announce Type: cross Abstract: Large language models (LLMs) often demonstrate strong safety performance in high-resource languages, yet exhibit severe vulnerabilities when queried i

model-releasesarxiv-cs-ai
15 Apr 2026
Safety

LLM-based Realistic Safety-Critical Driving Video Generation

DGX agent

arXiv:2507.01264v2 Announce Type: replace-cross Abstract: Designing diverse and safety-critical driving scenarios is essential for evaluating autonomous driving systems. In this paper, we propose a no

safetyarxiv-cs-ai
14 Apr 2026
Safety

Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility

DGX agent

arXiv:2602.03402v3 Announce Type: replace Abstract: Vision language models (VLMs) extend the reasoning capabilities of large language models (LLMs) to cross-modal settings, yet remain highly vulnerabl

safetyarxiv-cs-ai
14 Apr 2026
Safety

Safety Guarantees in Zero-Shot Reinforcement Learning for Cascade Dynamical Systems

DGX agent

arXiv:2604.10429v1 Announce Type: new Abstract: This paper considers the problem of zero-shot safety guarantees for cascade dynamical systems. These are systems where a subset of the states (the inner

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

Precise Shield: Explaining and Aligning VLLM Safety via Neuron-Level Guidance

DGX agent

arXiv:2604.08881v1 Announce Type: new Abstract: In real-world deployments, Vision-Language Large Models (VLLMs) face critical challenges from multilingual and multimodal composite attacks: harmful ima

model-releasesarxiv-cs-cv
13 Apr 2026
Safety

Introducing the Child Safety Blueprint

DGX agent

OpenAI's Child Safety Blueprint, released in April 2026, is a policy framework aimed at combating the rise of AI-enabled child sexual exploitation by combining legal, operational, and technical app...

safetyopenai
8 Apr 2026
Safety

Safe and Nonconservative Contingency Planning for Autonomous Vehicles via Online Learning-Based Reachable Set Barriers

DGX agent

arXiv:2509.07464v2 Announce Type: replace Abstract: Autonomous vehicles must navigate dynamically uncertain environments while balancing safety and efficiency. This challenge is exacerbated by unpredi

safetyarxiv-cs-ro
16 Apr 2026
Safety

The Impact of Operational-Data Fidelity when Assessing Safety-Critical Autonomous-Vehicle Software

DGX agent

arXiv:2608.10025v1 Announce Type: new Abstract: For safety-critical software, data from the software's operational past (e.g. a sequence of success and failure events experienced by the software) can

safetyarxiv-cs-ro
12 Aug 2026
Model Releases

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

DGX agent

arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguist

model-releasesarxiv-cs-cl
11 Aug 2026
Safety

Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production

DGX agent

arXiv:2608.08471v1 Announce Type: new Abstract: Deployed LLM safety guardrails are predominantly static: trained once and frozen at release, while new jailbreak techniques and previously un-addressed

safetyarxiv-cs-ai
11 Aug 2026
Safety

Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics

DGX agent

arXiv:2608.05656v1 Announce Type: cross Abstract: Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating these risks fa

safetyarxiv-cs-ai
7 Aug 2026
Model Releases

Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs

DGX agent

arXiv:2511.00382v2 Announce Type: replace-cross Abstract: Organizations increasingly adapt Large Language Models (LLMs) from public repositories such as HuggingFace to downstream tasks. Prior work sho

model-releasesarxiv-cs-lg
4 Aug 2026
Safety

From Vessel Trajectories to Safety-Critical Encounter Scenarios: A Generative AI Framework for Autonomous Ship Digital Testing

DGX agent

arXiv:2603.28067v2 Announce Type: replace Abstract: Digital testing has emerged as a key paradigm for the development and verification of autonomous maritime navigation systems, yet the availability o

safetyarxiv-cs-lg
4 Aug 2026
Safety

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

DGX agent

arXiv:2607.29254v1 Announce Type: new Abstract: AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential

safetyarxiv-cs-ai
3 Aug 2026
Safety

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

DGX agent

arXiv:2607.27081v1 Announce Type: cross Abstract: Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers

safetyarxiv-cs-cl
30 Jul 2026
Safety

Retrieval-Augmented Generation in LLMs for Mental Health: Quantifying the Incremental Contribution of Retrieval Within a Layered Safety Architecture

DGX agent

arXiv:2607.24817v1 Announce Type: cross Abstract: Digital mental health interventions (DMHIs) offer scalable support, but ensuring they accurately detect users' intent during volatile situations can b

safetyarxiv-cs-ai
29 Jul 2026
Model Releases

Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B

DGX agent

arXiv:2607.22545v1 Announce Type: cross Abstract: Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt injection, re

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure

DGX agent

arXiv:2607.21151v1 Announce Type: new Abstract: As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. Counterintuiti

safetyarxiv-cs-ai
24 Jul 2026
Safety

Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents

DGX agent

arXiv:2607.19449v1 Announce Type: cross Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure

safetyarxiv-cs-ai
23 Jul 2026
Safety

it’s just for your safety

DGX agent

On July 23, 2026, a Polymarket tweet announced that Anthropic donated $20 million to a political nonprofit calling for stricter AI regulation ahead of the U.S. midterm elections. The post, posted at 4

safetyyann-lecun--x
23 Jul 2026
Safety

PC-Diffuser: Path-Consistent Capsule CBF Safety Filtering for Diffusion-Based Trajectory Planner

DGX agent

arXiv:2603.10330v2 Announce Type: replace-cross Abstract: Autonomous driving in complex traffic requires planners that generalize beyond hand-crafted rules, motivating data-driven approaches that lear

safetyarxiv-cs-ai
16 Jul 2026
Safety

Efficient Partitioning Method of Large-Scale Public Safety Spatio-Temporal Data based on Information Loss Constraints

DGX agent

arXiv:2306.12857v3 Announce Type: replace Abstract: The storage, management, and application of massive spatio-temporal data are widely used in practical scenarios, including public safety. However, d

safetyarxiv-cs-lg
10 Jul 2026
Safety

Understanding Annotator Safety Policy with Interpretability

DGX agent

Safety policies define what constitutes safe and unsafe AI outputs, guiding data annotation and model development. However, annotation disagreement is pervasive and can stem from multiple sources such

safetyapple-ml-research
6 Jul 2026
Safety

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment

DGX agent

arXiv:2607.00572v1 Announce Type: new Abstract: Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed a

safetyarxiv-cs-ai
2 Jul 2026
Safety

Managed Autonomy at Runtime: Gear-Based Safety and Governance for Single- and Multi-Agent Cyber-Physical Systems

DGX agent

arXiv:2607.00334v1 Announce Type: new Abstract: Autonomous agents, whether LLM-driven software agents or robotic physical agents, face a common class of failure modes when operating without continuous

safetyarxiv-cs-ai
2 Jul 2026
Model Releases

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios

DGX agent

arXiv:2603.29759v2 Announce Type: replace-cross Abstract: Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment. However, existing ben

model-releasesarxiv-cs-ai
1 Jul 2026
Safety

Agent Safety Is Action Alignment

DGX agent

arXiv:2606.28739v1 Announce Type: new Abstract: Large language models increasingly act as agents: they call tools, move money, delete records, and send messages on a user's behalf. To keep them safe,

safetyarxiv-cs-ai
30 Jun 2026
Model Releases

When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries

DGX agent

arXiv:2606.28332v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenarios remains p

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

RAS: Measuring LLM Safety Through Refusal Alignment

DGX agent

arXiv:2606.25750v1 Announce Type: cross Abstract: Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their

model-releasesarxiv-cs-cl
25 Jun 2026
Safety

AutoSpec: Safety Rule Evolution for LLM Agents via Inductive Logic Programming

DGX agent

arXiv:2606.24245v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly automate complex tasks by integrating language models with external tools and environments. However, th

safetyarxiv-cs-ai
24 Jun 2026
Safety

Nvidia introduces Halos for Robotics to bridge the physical AI safety gap

DGX agent

Nivida Corp. today announced Halos for Robotics, the industry’s first full framework for robotic safety systems that encompasses building, testing and managing complete artificial intelligence robotic

safetysiliconangle
22 Jun 2026
Safety

Canada introduces the Safe Social Media Act, a bill that would ban social media for children under 16 and establish safety standards for AI chatbots (Maria Cheng/Reuters)

DGX agent

Maria Cheng / Reuters: Canada introduces the Safe Social Media Act, a bill that would ban social media for children under 16 and establish safety standards for AI chatbots — The Canadian government in

safetytechmeme
10 Jun 2026
← Previous
1…56789…297
Next →