AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,234 results
Safety

CCFM: Collision-Constrained Flow Matching for Safety-Critical Scenario Generation

DGX agent

arXiv:2607.04451v1 Announce Type: new Abstract: Evaluation of autonomous vehicle (AV) planners in safety-critical closed-loop simulation is essential for real-world deployment. However, generating con

safetyarxiv-cs-cv
7 Jul 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Safety Targeted Embedding Exploit via Refinement

DGX agent

arXiv:2607.01859v1 Announce Type: new Abstract: Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-r

model-releasesarxiv-cs-ai
3 Jul 2026
Safety

Altman’s AI safety proposal: bail me out as we badly missed our revenue runway, or i will not be a multi-billionaire

DGX agent

Altman’s AI safety proposal: bail me out as we badly missed our revenue runway, or i will not be a multi-billionaire Altman’s AI safety proposal: let us win, or everybody loses https://ft.trib.al/UrDI

safetygary-marcus--x
2 Jul 2026
Safety

The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis

DGX agent

arXiv:2606.29581v1 Announce Type: cross Abstract: Modern LLM deployments routinely compress models and raise sampling temperature to reduce cost, latency, or repetition, yet safety evaluations usually

safetyarxiv-cs-ai
30 Jun 2026
Safety

Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety

DGX agent

arXiv:2510.16492v4 Announce Type: replace Abstract: As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical. While

safetyarxiv-cs-cl
29 Jun 2026
Model Releases

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models

DGX agent

arXiv:2606.25442v1 Announce Type: new Abstract: Safety alignment of large language models (LLMs) typically depends on high-quality supervision data, such as safe demonstrations or preference pairs. Ho

model-releasesarxiv-cs-cl
25 Jun 2026
Safety

SHIELD: Safety on Humanoids via CBFs In Expectation on Learned Dynamics

DGX agent

arXiv:2505.11494v3 Announce Type: replace Abstract: Robot learning has produced remarkably effective ``black-box'' controllers for complex tasks such as dynamic locomotion on humanoids. Yet ensuring d

safetyarxiv-cs-ro
23 Jun 2026
Safety

Ideagram 4 - Safety Filter?

DGX agent

Ideogram 4 includes runtime safety filters powered by Hive for prompt and output moderation , with NSFW prompts blocked by displaying 'Image blocked by safety filter' . Users have reported false posit

safetyr-stablediffusion
11 Jun 2026
Safety

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech

DGX agent

arXiv:2606.06037v2 Announce Type: cross Abstract: Large audio language models (LALMs) are increasingly deployed in real-world applications, yet their safety alignment is still primarily evaluated on m

safetyarxiv-cs-cl
10 Jun 2026
Safety

DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment

DGX agent

arXiv:2606.07678v1 Announce Type: cross Abstract: Safety alignment for large language models relies on preference data, but current pipelines often train on large, redundant datasets. Existing data se

safetyarxiv-cs-ai
9 Jun 2026
Safety

Shield-Loco: Shielding Locomotion Policies with Predictive Safety Filtering

DGX agent

arXiv:2606.07193v1 Announce Type: new Abstract: Reinforcement learning (RL) policies enable dynamic legged locomotion but lack mechanisms to avoid violations of safety constraints that are absent duri

safetyarxiv-cs-ro
8 Jun 2026
Model Releases

Uncertainty-Guided Label Rebalancing for CPS Safety Monitoring

DGX agent

arXiv:2603.25670v3 Announce Type: replace Abstract: Safety monitoring is essential for Cyber-Physical Systems (CPSs). However, unsafe events are rare in real-world CPS operations, creating an extreme

model-releasesarxiv-cs-lg
8 Jun 2026
Safety

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models

DGX agent

arXiv:2606.03793v1 Announce Type: new Abstract: Multimodal Large Language Models integrate visual perception into language reasoning, introducing a continuous attack surface susceptible to adversarial

safetyarxiv-cs-cl
3 Jun 2026
Model Releases

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

DGX agent

arXiv:2606.03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previ

model-releasesarxiv-cs-ai
3 Jun 2026
Safety

Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing

DGX agent

arXiv:2606.00686v1 Announce Type: new Abstract: The prevailing paradigm in large language model (LLM) alignment operates via erasure, filtering unsafe data or training models to strictly refuse harmfu

safetyarxiv-cs-lg
2 Jun 2026
Safety

Quantifying the Salience of Geo-Cultural Values for Pluralistic Safety Alignment

DGX agent

arXiv:2606.00369v1 Announce Type: cross Abstract: Safe global deployment of AI models requires alignment with human values that vary across cultures. Yet rater pools in safety evaluation datasets rema

safetyarxiv-cs-lg
2 Jun 2026
Safety

TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety

DGX agent

arXiv:2606.00611v1 Announce Type: new Abstract: Long-horizon LLM agents produce safety evidence across long trajectories, where sparse, delayed, and compositional risk signals often escape local moder

safetyarxiv-cs-ai
2 Jun 2026
Safety

LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories

DGX agent

arXiv:2605.31381v1 Announce Type: new Abstract: We evaluate the consistency of automated judges in conducting a multi-dimensional safety evaluation in a reference-free setup. Our results indicate that

safetyarxiv-cs-cl
1 Jun 2026
Model Releases

Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?

DGX agent

arXiv:2508.11011v2 Announce Type: replace Abstract: Construction safety inspections typically involve a human inspector identifying safety concerns on-site. With the rise of powerful Vision Language M

model-releasesarxiv-cs-cv
28 May 2026
Safety

SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models

DGX agent

arXiv:2605.28338v1 Announce Type: new Abstract: Large language models(LLMs) increasingly match expert performance on licensing examinations, yet routine clinical use remains limited because governance

safetyarxiv-cs-ai
28 May 2026
Safety

TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling

DGX agent

arXiv:2605.27690v1 Announce Type: new Abstract: LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate steps long be

safetyarxiv-cs-cl
28 May 2026
Safety

When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?

DGX agent

arXiv:2605.27932v1 Announce Type: cross Abstract: Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly underst

safetyarxiv-cs-ai
28 May 2026
Safety

KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models

DGX agent

arXiv:2605.26947v1 Announce Type: new Abstract: Kazakh is underrepresented in resources for evaluating the safety behavior of large language models. We present KZ-SafetyPrompts, a Kazakh prompt datase

safetyarxiv-cs-cl
27 May 2026
Safety

Is GPT-4o mini Blinded by its Own Safety Filters? Exposing the Multimodal-to-Unimodal Bottleneck in Hate Speech Detection

DGX agent

arXiv:2509.13608v2 Announce Type: replace Abstract: As Large Multimodal Models (LMMs) become integral to daily digital life, understanding their safety architectures is a critical problem for AI Align

safetyarxiv-cs-lg
26 May 2026
Safety

Towards Context-Invariant Safety Alignment for Large Language Models

DGX agent

arXiv:2605.20994v1 Announce Type: new Abstract: Preference-based post-training aligns LLMs with human intent, yet safety behavior often remains brittle. A model may refuse a harmful request in a stand

safetyarxiv-cs-cl
21 May 2026
Safety

Alignment and Safety of Diffusion Models via Reinforcement Learning and Reward Modeling: A Survey

DGX agent

arXiv:2505.17352v2 Announce Type: replace Abstract: Diffusion models have become a central paradigm for image and multimodal generation, yet their deployment raises persistent questions about alignmen

safetyarxiv-cs-cv
19 May 2026
Model Releases

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents

DGX agent

arXiv:2605.16282v1 Announce Type: cross Abstract: The rapid deployment of LLM-based autonomous agents has introduced safety risks that extend far beyond traditional LLM concerns, prompting a prolifera

model-releasesarxiv-cs-ai
19 May 2026
Safety

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection

DGX agent

arXiv:2602.07892v2 Announce Type: replace-cross Abstract: Safety post-training can improve the harmfulness and policy compliance of Large Language Models (LLMs), but it may also reduce general utility

safetyarxiv-cs-cl
13 May 2026
Safety

A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models

DGX agent

arXiv:2605.08513v1 Announce Type: cross Abstract: Safety alignment in language models operates through two mechanistically distinct systems: refusal neurons that gate whether harmful knowledge is expr

safetyarxiv-cs-ai
12 May 2026
Model Releases

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion

DGX agent

arXiv:2503.06223v5 Announce Type: replace Abstract: Large Vision-Language Models (VLMs) are increasingly deployed in open-ended environments, where ensuring reliable safety under multimodal inputs is

model-releasesarxiv-cs-cv
11 May 2026
Safety

ChatGPT’s ‘Trusted Contact’ will alert loved ones of safety concerns

DGX agent

OpenAI is launching an optional safety feature for ChatGPT that allows adult users to assign an emergency contact for mental health and safety concerns. Friends, family members, or caregivers designat

safetythe-verge-ai
7 May 2026
Safety

Conditional Flow-VAE for Safety-Critical Traffic Scenario Generation

DGX agent

arXiv:2605.04366v1 Announce Type: cross Abstract: Safety-critical scenarios are essential for the development of autonomous vehicles (AVs) but are rare in real-world driving data. While simulation off

safetyarxiv-cs-lg
7 May 2026
Safety

FORMULA: FORmation MPC with neUral barrier Learning for safety Assurance

DGX agent

arXiv:2604.04409v2 Announce Type: replace Abstract: Multi-robot systems (MRS) are essential for large-scale applications such as disaster response, material transport, and warehouse logistics, yet ens

safetyarxiv-cs-ro
6 May 2026
Model Releases

RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs

DGX agent

arXiv:2605.01913v1 Announce Type: cross Abstract: Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models vulnerable t

model-releasesarxiv-cs-cl
5 May 2026
Safety

Benchmarking the Safety of Large Language Models for Robotic Health Attendant Control

DGX agent

arXiv:2604.26577v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly considered for deployment as the control component of robotic health attendants, yet their safety in this

safetyarxiv-cs-ai
30 Apr 2026
Safety

The Alignment Flywheel: A Governance-Centric Hybrid MAS for Architecture-Agnostic Safety

DGX agent

arXiv:2603.02259v2 Announce Type: replace-cross Abstract: Multi-agent systems provide mature methodologies for role decomposition, coordination, and normative governance, capabilities that remain esse

safetyarxiv-cs-lg
30 Apr 2026
Safety

AI Safety Training Can be Clinically Harmful

DGX agent

arXiv:2604.23445v1 Announce Type: cross Abstract: Large language models are being deployed as mental health support agents at scale, yet only 16% of LLM-based chatbot interventions have undergone rigo

safetyarxiv-cs-ai
28 Apr 2026
Safety

Our commitment to community safety

DGX agent

OpenAI outlines its commitment to implementing safety measures and responsible practices in the development and deployment of AI systems to protect users and communities. The statement likely covers O

safetyopenai
28 Apr 2026
Safety

Learning Control Policies to Provably Satisfy Hard Affine Constraints for Black-Box Hybrid Dynamical Systems

DGX agent

arXiv:2604.22244v1 Announce Type: new Abstract: Ensuring safety for black-box hybrid dynamical systems presents significant challenges due to their instantaneous state jumps and unknown explicit nonli

safetyarxiv-cs-ro
27 Apr 2026
Safety

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem

DGX agent

arXiv:2506.17299v2 Announce Type: replace-cross Abstract: As large language models (LLMs) become increasingly deployed in safety-critical applications, the lack of systematic methods to assess their v

safetyarxiv-cs-ai
27 Apr 2026
Safety

Logic Jailbreak: Efficiently Unlocking LLM Safety Restrictions Through Formal Logical Expression

DGX agent

arXiv:2505.13527v3 Announce Type: replace-cross Abstract: Despite substantial advancements in aligning large language models (LLMs) with human values, current safety mechanisms remain susceptible to j

safetyarxiv-cs-ai
24 Apr 2026
Safety

A Hough transform approach to safety-aware scalar field mapping using Gaussian Processes

DGX agent

arXiv:2604.20799v1 Announce Type: new Abstract: This paper presents a framework for mapping unknown scalar fields using a sensor-equipped autonomous robot operating in unsafe environments. The unsafe

safetyarxiv-cs-ro
23 Apr 2026
Safety

LLM-Guided Safety Agent for Edge Robotics with an ISO-Compliant Perception-Compute-Control Architecture

DGX agent

arXiv:2604.20193v1 Announce Type: new Abstract: Ensuring functional safety in human-robot interaction is challenging because AI perception is inherently probabilistic, whereas industrial standards req

safetyarxiv-cs-ro
23 Apr 2026
Safety

The Alignment Waltz: Jointly Training Agents to Collaborate for Safety

DGX agent

arXiv:2510.08240v2 Announce Type: replace Abstract: Harnessing the power of LLMs requires a delicate dance between being helpful and harmless. This creates a fundamental tension between two competing

safetyarxiv-cs-cl
22 Apr 2026
Safety

APIs and limited releases for AI models are not a safety policy, they’re a business model (which is totally ok as long as you’re transparent…

DGX agent

APIs and limited releases for AI models are not a safety policy, they’re a business model (which is totally ok as long as you’re transparent about it). Especially on cyber-security, they give a false

safetyclem-delangue--x
21 Apr 2026
Safety

Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment

DGX agent

arXiv:2405.13068v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have revolutionized various applications, making robust safety alignment essential to prevent harmful outputs. Cu

safetyarxiv-cs-lg
21 Apr 2026
Safety

When Choices Become Risks: Safety Failures of Large Language Models under Multiple-Choice Constraints

DGX agent

arXiv:2604.16916v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is primarily evaluated under open-ended generation, where models can mitigate risk by refusing to respo

safetyarxiv-cs-cl
21 Apr 2026
Safety

Tesla expands its robotaxi service to Dallas and Houston after launching in Austin last year and starting to offer rides without safety drivers in January 2026 (Anthony Ha/TechCrunch)

DGX agent

Anthony Ha / TechCrunch: Tesla expands its robotaxi service to Dallas and Houston after launching in Austin last year and starting to offer rides without safety drivers in January 2026 — Tesla is expa

safetytechmeme
19 Apr 2026
← Previous
1…34567…297
Next →