AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Safety

Safety-oriented sidewalk and road segmentation for smartphone-based assistive navigation

DGX agent

arXiv:2607.21137v1 Announce Type: new Abstract: Independent sidewalk mobility is essential for blind and visually impaired pedestrians (BVIPs), yet smartphone-based assistive navigation requires perce

safetyarxiv-cs-cv
24 Jul 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Data Safety: Synthetic Data Quality Analysis Using CIFAKE Dataset

DGX agent

arXiv:2607.12165v1 Announce Type: new Abstract: Recently, the societal implementation of high-performance image classification models has expanded rapidly. While these models require vast amounts of t

safetyarxiv-cs-cv
15 Jul 2026
Model Releases

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

DGX agent

arXiv:2602.05088v4 Announce Type: replace Abstract: Millions of people now use generative AI chatbots for psychological support. Despite their promise, the most pressing question in AI for mental heal

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale

DGX agent

arXiv:2607.02714v1 Announce Type: cross Abstract: There is no doubt that safety alignment is an essential step in LLM training. However, conceptually it does not distinguish between various domains an

model-releasesarxiv-cs-ai
7 Jul 2026
Safety

Chemical accidents rise as Trump administration proposes weakening safety rules

DGX agent

Chemical accidents involving releases of dangerous chemicals rose 57 percent between 2021 and 2025, from 83 to 131 incidents , with injuries or deaths rising from 60 to 89 over the same period . The E

safetyars-technica
5 Jul 2026
Model Releases

Neuro-Symbolic Safety Guidance for Vision-Language-Action Models via Constrained Flow Matching

DGX agent

arXiv:2607.01378v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated promising generalization capabilities across robotic manipulation tasks, yet their real-world depl

model-releasesarxiv-cs-ro
3 Jul 2026
Safety

Online Safety Monitoring for LLMs

DGX agent

arXiv:2607.02510v1 Announce Type: new Abstract: Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safet

safetyarxiv-cs-ai
3 Jul 2026
Safety

Off the Rails: Hijacking the Scoring Head in Generative End-to-End Driving Planners with Safety-Violating Adversarial Perturbations

DGX agent

arXiv:2606.30807v1 Announce Type: cross Abstract: Generative models have recently seen rapid adoption in End-to-End (E2E) autonomous driving (AD), with diffusion-based denoising and vocabulary-based r

safetyarxiv-cs-cv
1 Jul 2026
Model Releases

OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation

DGX agent

arXiv:2606.31993v1 Announce Type: new Abstract: While robotic manipulation capabilities have advanced rapidly, physical safety remains a major barrier to deploying household robots: task success is in

model-releasesarxiv-cs-ro
1 Jul 2026
Safety

Revealing Safety-Critical Scenarios for UTM via Transformer

DGX agent

arXiv:2606.31114v1 Announce Type: new Abstract: Unmanned Traffic Management (UTM) systems are cloud-based platforms designed to manage and coordinate multiple aerial vehicles remotely. UTM systems are

safetyarxiv-cs-ai
1 Jul 2026
Model Releases

Necessary but Not Sufficient: Temperature Control and Reproducibility in LLM-as-Judge Safety Evaluations

DGX agent

arXiv:2606.26185v1 Announce Type: new Abstract: LLM-as-judge ('grader') components are now standard in evaluation harnesses, including safety evaluations where a pass/fail verdict may gate downstream

model-releasesarxiv-cs-lg
26 Jun 2026
Model Releases

Do Thinking Tokens Help with Safety?

DGX agent

arXiv:2606.25013v1 Announce Type: cross Abstract: Today's reasoning models use thinking tokens to attain stronger performance on benchmarks than their instruction-tuned counterparts. It is also genera

model-releasesarxiv-cs-cl
25 Jun 2026
Safety

Safe and Generalizable Hierarchical Multi-Agent RL via Constraint Manifold Control

DGX agent

arXiv:2606.24010v1 Announce Type: new Abstract: Multi-agent systems are widely used in safety-critical applications that require coordinated behavior under strict safety constraints. Existing approach

safetyarxiv-cs-ai
24 Jun 2026
Safety

Enhancing Road Safety: An IoT-Based Accident Detection and Prevention Mechanism

DGX agent

arXiv:2606.22381v1 Announce Type: cross Abstract: Road traffic accidents remain a critical global crisis, consistently serving as a primary driver of preventable mortality and severe injury. These inc

safetyarxiv-cs-cv
23 Jun 2026
Safety

SAFER-Nav: Enhancing Safety for Visual Robot Navigation via Segmentation-Aware Fine-Tuning

DGX agent

arXiv:2606.11636v1 Announce Type: new Abstract: Vision-based navigation models, particularly foundation models, generate viable trajectories from RGB observations alone. However, even state-of-the-art

safetyarxiv-cs-ro
11 Jun 2026
Model Releases

Anthropic secretly limiting Claude's usefulness for LLM development strengthens the argument that Anthropic is using AI safety to justify monopolistic behavior (Dean W. Ball/@deanwball)

DGX agent

Dean W. Ball / @deanwball: Anthropic secretly limiting Claude's usefulness for LLM development strengthens the argument that Anthropic is using AI safety to justify monopolistic behavior — My last obs

model-releasestechmeme
10 Jun 2026
Safety

Enhancing AI Interpretability and Safety through Localised Architectures

DGX agent

arXiv:2606.07998v1 Announce Type: cross Abstract: Recent advances in generative AI, especially powerful Large Language Models (LLMs) and Large Reasoning Models (LRMs), raise concerns over the interpre

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection

DGX agent

arXiv:2606.09204v1 Announce Type: new Abstract: We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injection Paradox -- in which prompt injections embedde

model-releasesarxiv-cs-lg
9 Jun 2026
Safety

Apple refines Liquid Glass design and expands child safety tools at WWDC

DGX agent

Apple Inc.’s WWDC 2026 keynote today focused on making its software platform feel more polished, more responsive and more tightly managed across the company’s device ecosystem with refinements to the

safetysiliconangle
8 Jun 2026
Model Releases

PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage

DGX agent

arXiv:2606.05463v1 Announce Type: new Abstract: Patient safety event triage, determining whether a clinical event is reportable under jurisdiction-specific policy, is a high-stakes task typically perf

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Low-Resource Safety Failures Are Action Failures, Not Representation Failures

DGX agent

arXiv:2606.01196v1 Announce Type: cross Abstract: Safety alignment learned in high-resource languages transfers poorly to low-resource languages. Models refuse harmful prompts in English but fail to r

model-releasesarxiv-cs-ai
2 Jun 2026
Safety

Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures

DGX agent

arXiv:2605.29629v1 Announce Type: new Abstract: Attack Success Rate (ASR) evaluates each jailbreak with a single yes/no label at the end of generation, telling us whether a failure happened but not ho

safetyarxiv-cs-ai
29 May 2026
Model Releases

Beyond English and Evasion: A Human-Annotated Multi-Domain Benchmark for High-Stakes LLM Safety Evaluation in Chinese

DGX agent

arXiv:2605.29667v1 Announce Type: new Abstract: When Large Language Models (LLMs) are deployed in Chinese-language settings, a troubling pattern emerges: safety systems that work well in English break

model-releasesarxiv-cs-cl
29 May 2026
Safety

ICAN-Deploy: Identity-Stable Canary Deployment for Safety-Critical Embodied Agents

DGX agent

arXiv:2605.28097v1 Announce Type: new Abstract: Canary deployment routes a fraction of traffic to a new software version, monitors metrics, and rolls back on regression. Mainstream controllers (Argo R

safetyarxiv-cs-ro
28 May 2026
Safety

VLM-Based Advanced Rider Assistance System for Motorcycle Safety

DGX agent

arXiv:2605.27948v1 Announce Type: new Abstract: Motorcycles face disproportionately high crash risks compared to cars due to limited protection and heightened sensitivity to surface hazards, yet Advan

safetyarxiv-cs-ro
28 May 2026
Safety

FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents

DGX agent

arXiv:2605.27333v1 Announce Type: new Abstract: Finance LLM agents must simultaneously block prompt-induced unauthorized actions and approve legitimate multi-step business workflows. However, boundary

safetyarxiv-cs-cl
27 May 2026
Safety

Learning to Balance Motor Thermal Safety and Quadrupedal Locomotion Performance with Residual Policy

DGX agent

arXiv:2605.27046v1 Announce Type: new Abstract: Motor thermal management is often overlooked in the context of electrically-actuated robots, particularly legged robots, but motor overheating is a key

safetyarxiv-cs-ro
27 May 2026
Safety

Safety-Oriented Routing Analysis of Mixtral MoE Under Benign and Harmful Prompts

DGX agent

arXiv:2605.24270v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) language models activate only a small subset of parameters for each token, making router behavior a central part of mode

safetyarxiv-cs-ai
26 May 2026
Safety

Towards trustworthy agentic AI: a comprehensive survey of safety, robustness, privacy, and system security

DGX agent

arXiv:2605.23989v1 Announce Type: new Abstract: Agentic AI systems -- Large Language Models (LLMs) augmented with planning, tool use, memory, and long-horizon interactions -- can execute complex tasks

safetyarxiv-cs-ai
26 May 2026
Safety

SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety

DGX agent

arXiv:2605.05704v2 Announce Type: replace-cross Abstract: Recent advances in foundation models have transformed LLMs from passive conversational systems into autonomous agents capable of reasoning and

safetyarxiv-cs-ai
25 May 2026
Model Releases

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

DGX agent

arXiv:2605.22643v1 Announce Type: new Abstract: Background. Traditional safety benchmarks for language models evaluate generated text: whether a model outputs toxic language, reproduces bias, or follo

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety

DGX agent

arXiv:2605.20203v1 Announce Type: cross Abstract: As older adults increasingly use LLM-based chatbots for companionship and assistance, a safety gap is emerging. Older adults may face vulnerabilities

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

HealthCraft: A Reinforcement Learning Safety Environment for Emergency Medicine

DGX agent

arXiv:2605.21496v1 Announce Type: cross Abstract: Frontier language models are being deployed into clinical workflows faster than the infrastructure to evaluate them safely. Static medical-QA benchmar

model-releasesarxiv-cs-cl
22 May 2026
Safety

Tech researchers are suing the Trump administration over the future of online safety

DGX agent

Since its earliest days back in office, the Trump administration has been going after researchers who study and try to counter hate speech, harassment, propaganda, and disinformation online. Now, some

safetymit-tech-review
21 May 2026
Safety

The Download: online safety’s future and climate tech’s big pivot

DGX agent

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Tech researchers are suing the Trump administration over the f

safetymit-tech-review
21 May 2026
Safety

Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework

DGX agent

arXiv:2603.11768v2 Announce Type: replace Abstract: Long-term memory has emerged as a foundational component of autonomous Large Language Model (LLM) agents, enabling continuous adaptation, lifelong m

safetyarxiv-cs-ai
20 May 2026
Model Releases

Measuring Safety Alignment Effects in Autonomous Security Agents

DGX agent

arXiv:2605.19722v1 Announce Type: cross Abstract: Do stock safety-aligned language models and their uncensored or abliterated derivatives behave differently when run as autonomous security agents? Sin

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

AgentWall: A Runtime Safety Layer for Local AI Agents

DGX agent

arXiv:2605.16265v1 Announce Type: new Abstract: The safety of autonomous AI agents is increasingly recognized as a critical open problem. As agents transition from passive text generators to active ac

model-releasesarxiv-cs-ai
19 May 2026
Safety

Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation

DGX agent

arXiv:2605.17268v1 Announce Type: new Abstract: We present the first systematic study of faithfulness in Vision-Language-Action (VLA) driving models, analyzing 300 Alpamayo-R1-10B inferences across 10

safetyarxiv-cs-ai
19 May 2026
Model Releases

When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents

DGX agent

arXiv:2601.17887v2 Announce Type: replace Abstract: Long-term memory enables large language model (LLM) agents to support personalized and sustained interactions. However, most work on personalized ag

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing

DGX agent

arXiv:2605.10146v1 Announce Type: new Abstract: Large language models (LLMs) increasingly rely on knowledge editing to support knowledge-intensive reasoning, but this flexibility also introduces criti

model-releasesarxiv-cs-ai
12 May 2026
Safety

SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints

DGX agent

arXiv:2512.23770v3 Announce Type: replace-cross Abstract: In safety-critical domains, reinforcement learning (RL) agents must often satisfy strict, zero-cost safety constraints while accomplishing tas

safetyarxiv-cs-ai
11 May 2026
Safety

Big AI Lobbyists: if you regulate us at all, we lose to China because they will never regulate ... Actual China: 'safety first, innovation second ... Development must be controllable and orderly.'

DGX agent

This post highlights a contradiction in AI industry arguments, contrasting Western tech company claims that regulation will disadvantage them competitively against China with evidence of China's own s

safetyr-chatgpt
10 May 2026
Safety

Google, Microsoft and xAI agree to allow government safety checks of their AI models prior to release

DGX agent

Google LLC, Microsoft Corp. and xAI have agreed to share unreleased versions of their artificial intelligence models with the U.S. Department of Commerce to ensure the technologies do not pose a threa

safetysiliconangle
6 May 2026
Model Releases

Safety and accuracy follow different scaling laws in clinical large language models

DGX agent

arXiv:2605.04039v1 Announce Type: new Abstract: Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation

model-releasesarxiv-cs-cl
6 May 2026
Safety

Set-Based Training of Neural Barrier Certificates for Safety Verification of Dynamical Systems

DGX agent

arXiv:2605.02526v1 Announce Type: cross Abstract: Barrier certificates are scalar functions over the state space of dynamical systems that separate all unsafe states from all reachable states. The exi

safetyarxiv-cs-ai
6 May 2026
Model Releases

A decoupled diffusion planner that adapts to changing cost limits by using cost-conditioned generation for safety and reward gradients for performance

DGX agent

arXiv:2605.02777v1 Announce Type: new Abstract: Offline safe reinforcement learning often requires policies to adapt at deployment time to safety budgets that vary across episodes or change within a s

model-releasesarxiv-cs-lg
5 May 2026
Safety

Value Functions for Temporal Logic: Optimal Policies and Safety Filters

DGX agent

arXiv:2605.01051v1 Announce Type: cross Abstract: While Bellman equations for basic reach, avoid, and reach-avoid problems are well studied, the relationship between value optimality and policy optima

safetyarxiv-cs-lg
5 May 2026
← Previous
1…1112131415…297
Next →