AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Safety

Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds

DGX agent

arXiv:2608.10056v1 Announce Type: cross Abstract: Following a target human in crowded environments involves an inherent conflict between staying close to the target and navigating safely among surroun

safetyarxiv-cs-ai
12 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Robust Safety Filtering for Input-Constrained Underactuated Linear Systems

DGX agent

arXiv:2608.10872v1 Announce Type: cross Abstract: We present a robust safety-filtering framework for input-constrained underactuated linear systems subject to unknown disturbances. A baseline H-infty

safetyarxiv-cs-ro
12 Aug 2026
Safety

Mood Matters: How Syntactic Sensitivity Undermines Safety Alignment

DGX agent

arXiv:2608.05409v1 Announce Type: new Abstract: Large language models typically undergo post-training to align them with safety policies but there exist many sophisticated jailbreaks that sidestep est

safetyarxiv-cs-cl
7 Aug 2026
Safety

NSF-HRPT: Neural Semantic Field meets Hierarchical Risk Perception Tree for Safety-Critical Scenario Assessment

DGX agent

arXiv:2608.04776v1 Announce Type: new Abstract: The ability to accurately assess and anticipate risks in safety-critical scenarios is crucial for autonomous driving systems. While existing research ha

safetyarxiv-cs-ai
6 Aug 2026
Safety

SCOPE: Field-of-View-Aware Path Planning in Unknown 3D Environments via Safety-Volume Certification

DGX agent

arXiv:2608.04420v1 Announce Type: new Abstract: Safe navigation with a body-mounted limited-field-of-view sensor requires the complete robot-inflated volume of an intended motion to be observed and ve

safetyarxiv-cs-ro
6 Aug 2026
Safety

ADMITBench: A Safety-Governed Reference Framework for Evaluating the Admissibility of Industrial LLM Advisories

DGX agent

arXiv:2608.03866v1 Announce Type: new Abstract: This white paper presents ADMITBench, a reference framework for evaluating industrial LLM advisories at the level of the proposed action. The framework

safetyarxiv-cs-ai
5 Aug 2026
Safety

Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos …

DGX agent

Yes, the AIs were given a cybersecurity challenge, with internet access enabled and safety filters disabled. But the extent to which Mythos 5 pursued its mission (fake identities, social engineering,

safetyethan-mollick--x
5 Aug 2026
Safety

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

DGX agent

arXiv:2508.05775v3 Announce Type: replace Abstract: Large Language Models (LLMs) have revolutionized content creation across digital platforms, offering unprecedented capabilities in natural language

safetyarxiv-cs-cl
29 Jul 2026
Safety

Forecasting the Emergence and Evolution of Crash Hotspots: A Unified Deep Learning Framework for Proactive Traffic Safety

DGX agent

arXiv:2607.24168v1 Announce Type: new Abstract: Road crashes remain among the gravest threats to public safety, and preventing them is a defining task of transportation systems worldwide. Much of that

safetyarxiv-cs-lg
28 Jul 2026
Safety

What Can Be Enforced? A Theory of Certified Runtime Safety for Tool-Using Agents

DGX agent

arXiv:2607.22868v1 Announce Type: new Abstract: Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge observes, and whe

safetyarxiv-cs-ai
28 Jul 2026
Safety

SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing

DGX agent

arXiv:2607.13594v1 Announce Type: new Abstract: LLM agents act on real-world environments through tool calls, and a single misjudged action can cause irreversible harm. The standard safeguard is a gua

safetyarxiv-cs-ai
16 Jul 2026
Model Releases

A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis

DGX agent

arXiv:2607.08038v1 Announce Type: new Abstract: Diagnostic error is a major threat to patient safety, yet current large language model (LLM) systems often treat diagnosis as a one-shot prediction task

model-releasesarxiv-cs-ai
10 Jul 2026
Safety

Efficient Safety Alignment of Language Models via Latent Personality Traits

DGX agent

arXiv:2607.07918v1 Announce Type: cross Abstract: Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Late

safetyarxiv-cs-ai
10 Jul 2026
Safety

At UMA, we build and own the full stack, from hardware to software. This allows us to bake safety in at all levels, rather than bolt it on a…

DGX agent

At UMA, we build and own the full stack, from hardware to software. This allows us to bake safety in at all levels, rather than bolt it on as an afterthought. Because trust is everything when robots s

safetyyann-lecun--x
7 Jul 2026
Safety

VLM-CASE: Vision-Language Model Enabled Context-Adaptive Safety Envelopes for Anticipatory Safe Autonomous Driving

DGX agent

arXiv:2607.05180v1 Announce Type: cross Abstract: Adverse driving conditions, such as bad weather, remain a principal barrier to autonomous driving because they degrade two things at once: what the ve

safetyarxiv-cs-cv
7 Jul 2026
Safety

Q&A with Agility Robotics CEO Peggy Johnson on why the startup is going public via SPAC, the physical layer as its proprietary advantage, safety, and more (Connie Loizos/TechCrunch)

DGX agent

Connie Loizos / TechCrunch: Q&A with Agility Robotics CEO Peggy Johnson on why the startup is going public via SPAC, the physical layer as its proprietary advantage, safety, and more — The humanoid ro

safetytechmeme
6 Jul 2026
Safety

Tesla rolls out its Robotaxi service without a safety monitor in Miami, its fifth city, as it aims to expand to a dozen US states by the end of 2026 (Grace Kay/The Information)

DGX agent

Grace Kay / The Information: Tesla rolls out its Robotaxi service without a safety monitor in Miami, its fifth city, as it aims to expand to a dozen US states by the end of 2026 — Tesla said it rolled

safetytechmeme
5 Jul 2026
Safety

Multi-modal Rail Crossing Safety Analysis

DGX agent

arXiv:2607.01365v1 Announce Type: cross Abstract: Given one or more images of a railway crossing, can we leverage visual cues that allow us to robustly estimate how safe it is? Can we improve our abil

safetyarxiv-cs-ai
3 Jul 2026
Safety

FastBridge: Closing the Model-Based Realization Gap in Safety Filters on 3D Gaussian Splatting for Fast Quadrotor Flight

DGX agent

arXiv:2607.01200v1 Announce Type: new Abstract: Fast quadrotor flight requires safe obstacle avoidance under tight onboard compute limits. While 3D Gaussian Splatting (3DGS) provides a continuous, geo

safetyarxiv-cs-ro
2 Jul 2026
Model Releases

Multimodal Benchmark for Safety Assessment in Industrial Inspection Scenarios

DGX agent

arXiv:2601.21173v2 Announce Type: replace-cross Abstract: With the rapid development of industrial intelligence and unmanned inspection, reliable perception and safety assessment for AI systems in com

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots

DGX agent

arXiv:2606.30256v1 Announce Type: new Abstract: Safety benchmarks often buy scalability by fixing the prompt, the language, and the turn structure. For emotional-support chatbots, that bargain hides p

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context

DGX agent

arXiv:2601.17642v2 Announce Type: replace Abstract: Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-refusal o

model-releasesarxiv-cs-ai
29 Jun 2026
Safety

Really impressive first few drives with FSD v14 lite! Big leap in capability and feature set from v12.6.4, it’s tuned for safety right now s…

DGX agent

Really impressive first few drives with FSD v14 lite! Big leap in capability and feature set from v12.6.4, it’s tuned for safety right now so it’s taking things slow and smoothly. Highway performance

safetyelon-musk--x
29 Jun 2026
Model Releases

Real-Time Safety Evaluation of Human Arm Operations Using a Wrist-Mounted IMU with PSM System

DGX agent

arXiv:2502.09241v2 Announce Type: replace Abstract: This paper presents a novel approach to real-time safety monitoring in human-robot collaborative manufacturing environments through a wrist-mounted

model-releasesarxiv-cs-ro
26 Jun 2026
Safety

It would be very useful to understand more about the government safety concerns associated with frontier AI releases so we could (a) know wh…

DGX agent

It would be very useful to understand more about the government safety concerns associated with frontier AI releases so we could (a) know what risks everyone will face if/when open source reaches Myth

safetyethan-mollick--x
25 Jun 2026
Safety

Are Safety Guarantees in Neural Networks Safe? How to Compute Trustworthy Robustness Certifications

DGX agent

arXiv:2606.23858v1 Announce Type: cross Abstract: A primary challenge in AI safety is the existence of adversarial examples -- slightly distorted inputs that cause a neural network (NN) to misclassify

safetyarxiv-cs-ai
24 Jun 2026
Model Releases

IndicGuard: A Multilingual Safety Guard Model and Dataset for Indic Languages

DGX agent

arXiv:2606.22841v1 Announce Type: cross Abstract: As Large Language Models (LLMs) achieve widespread integration across diverse linguistic landscapes, ensuring their safety and alignment with regional

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs

DGX agent

arXiv:2606.22686v1 Announce Type: cross Abstract: Modern Large Language Models (LLMs) rely on extensive safety alignment, yet the mechanistic basis of refusal remains opaque. In this work, we investig

model-releasesarxiv-cs-lg
23 Jun 2026
Safety

An interesting new paper by my recent PhD graduate on how AI agents' greed for visible incentives can lead them to abandon their safety alig…

DGX agent

An interesting new paper by my recent PhD graduate on how AI agents' greed for visible incentives can lead them to abandon their safety alignment. You can read it here: https://arxiv.org/abs/2606.1691

safetyyoshua-bengio--x
22 Jun 2026
Safety

Nvidia unveils Halos, a safety-focused OS developed from autonomous vehicle tech and designed to run on IGX Thor hardware for humanoid robots, and opens a lab (Ian King/Bloomberg)

DGX agent

Ian King / Bloomberg: Nvidia unveils Halos, a safety-focused OS developed from autonomous vehicle tech and designed to run on IGX Thor hardware for humanoid robots, and opens a lab — Nvidia Corp. is w

safetytechmeme
22 Jun 2026
Model Releases

Benchmarking Large Language Models for Safety Data Extraction

DGX agent

arXiv:2606.11204v1 Announce Type: new Abstract: Accurate extraction of structured information from Safety Data Sheets (SDS) remains challenging in industrial safety due to heterogeneous document forma

model-releasesarxiv-cs-cl
11 Jun 2026
Safety

Schutzen: Evaluating LLM Safety in Bulgarian and German Contexts

DGX agent

arXiv:2606.11316v1 Announce Type: new Abstract: Large language models are increasingly deployed across professional domains, bringing hard-to-predict risks, including the generation of harmful or disr

safetyarxiv-cs-cl
11 Jun 2026
Safety

An essay on policy responses to AI's exponential progress across regulation and public safety, macroeconomics and taxes, science, civil liberties, geopolitics (Dario Amodei)

DGX agent

Dario Amodei: An essay on policy responses to AI's exponential progress across regulation and public safety, macroeconomics and taxes, science, civil liberties, geopolitics — In one of the side plots

safetytechmeme
10 Jun 2026
Safety

CameraMatics, which uses AI to help fleet operators improve safety, reduce operational risk, and lower carbon emissions, raised €49M (Joe Brennan/The Irish Times)

DGX agent

Joe Brennan / The Irish Times: CameraMatics, which uses AI to help fleet operators improve safety, reduce operational risk, and lower carbon emissions, raised €49M — Tech focuses on accident preventio

safetytechmeme
10 Jun 2026
Safety

🔔 IF Anthropic is for real about safety, the should pause, for one month, and show leadership. If they won’t pause, even briefly, we have a…

DGX agent

🔔 IF Anthropic is for real about safety, the should pause, for one month, and show leadership. If they won’t pause, even briefly, we have a different kind of answer about what is real and what is mark

safetygary-marcus--x
10 Jun 2026
Safety

Two to Tango: Coupled Task-Reference Selection for Safe LLM Fine-tuning

DGX agent

arXiv:2606.09866v1 Announce Type: cross Abstract: Fine-tuning safety aligned large language models (LLMs) on downstream data improves adaptation but may erode learned safety behavior. Existing methods

safetyarxiv-cs-ai
10 Jun 2026
Safety

When you hear AI 'safety' you should hear 'censorship' and 'control' instead. All of us surveilled and spied by safeguards of loving grace. …

DGX agent

When you hear AI 'safety' you should hear 'censorship' and 'control' instead. All of us surveilled and spied by safeguards of loving grace. Today it's intelligent Terms of Service control. You can't d

safetyyann-lecun--x
10 Jun 2026
Safety

CARE: A Conformal Safety Layer for Medical Summarization

DGX agent

arXiv:2606.08969v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for medical summarization, but their outputs can omit medically important information and introduce

safetyarxiv-cs-ai
9 Jun 2026
Safety

fable’s safety guardrails broken within an hour. raise your hand if you are surprised.

DGX agent

fable’s safety guardrails broken within an hour. raise your hand if you are surprised. We tested Anthropic’s new @claudeai Fable 5. It did not fail like an ordinary jailbreak. It failed more quietly.

safetygary-marcus--x
9 Jun 2026
Model Releases

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs

DGX agent

arXiv:2606.09038v1 Announce Type: new Abstract: Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' preferences, contexts, and long-term histories. H

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Vision-Language Work Zone Intelligence for Safety-Critical Speed Regulation of Mixed-Autonomy Vehicles in Dynamic Environments

DGX agent

arXiv:2606.08860v1 Announce Type: new Abstract: Temporary work-zone speed limits are communicated through visually inconsistent signage and are often missing from digital maps, creating safety risks f

safetyarxiv-cs-cv
9 Jun 2026
Model Releases

Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs

DGX agent

arXiv:2511.20158v2 Announce Type: replace Abstract: While continual visual instruction tuning (CVIT) has shown promise in adapting multimodal large language models (MLLMs), existing studies predominan

model-releasesarxiv-cs-cv
5 Jun 2026
Safety

Not just foreseeable, but foreseen and called out. We ran a campaign to get them out of the first ever AI Safety Summit. We won that one. Ye…

DGX agent

Not just foreseeable, but foreseen and called out. We ran a campaign to get them out of the first ever AI Safety Summit. We won that one. Yet most of the field kept licking the boot of the companies a

safetyconnor-leahy--x
5 Jun 2026
Safety

Safety officials finally have a good idea of what a big rocket explosion can do

DGX agent

Safety officials gained detailed data on the impact of large rocket explosions when Blue Origin's New Glenn rocket experienced an anomaly during a hot-fire test at Cape Canaveral on May 28, causing a

safetyars-technica
5 Jun 2026
Safety

Sources and docs detail defense tech startup Shield AI's struggles to overcome years of technical hitches and safety concerns with its V-BAT autonomous drone (David Jeans/Reuters)

DGX agent

David Jeans / Reuters: Sources and docs detail defense tech startup Shield AI's struggles to overcome years of technical hitches and safety concerns with its V-BAT autonomous drone — A year ago, Ryan

safetytechmeme
5 Jun 2026
Safety

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

DGX agent

arXiv:2602.06911v2 Announce Type: replace-cross Abstract: As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications,

safetyarxiv-cs-ai
4 Jun 2026
Model Releases

Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight LLMs

DGX agent

arXiv:2606.04035v1 Announce Type: cross Abstract: We present a systematic study of domain-dependent safety behavior in open-weight LLMs: 7 standardized experiments across 7 ethical domains, testing 5

model-releasesarxiv-cs-ai
4 Jun 2026
Safety

Advancing youth safety and opportunity through global leadership

DGX agent

OpenAI outlines initiatives and commitments to enhance youth safety while creating educational and economic opportunities for young people globally. The effort emphasizes OpenAI's role in responsible

safetyopenai
2 Jun 2026
← Previous
1…7891011…297
Next →