AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
29 May 2026

GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models

Model ReleasesDGX agent

arXiv:2605.28848v1 Announce Type: cross Abstract: Deployed language models are evaluated in a non-stationary environment: model versions, retrieval layers, safety systems, and real-world inputs all ch

Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles

Model ReleasesDGX agent

arXiv:2605.29473v1 Announce Type: cross Abstract: Language models are increasingly being deployed for conversational support in informal caregiving contexts, where interactions often extend beyond inf

ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.29114v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between

28 May 2026

Adversarial Fine-tuning of Compressed Neural Networks for Joint Improvement of Robustness and Efficiency

Model ReleasesDGX agent

arXiv:2403.09441v2 Announce Type: replace Abstract: As deep learning (DL) models are increasingly being integrated into our everyday lives, ensuring their safety by making them robust against adversar

Bad cholesterol slashed 62% by single dose of gene-editing drug in small trial

IndustryDGX agent

An experimental gene-editing therapy called VERVE-102 aims to lower bad cholesterol long-term after a single infusion and showed promising results in a Phase I safety trial published in the New Englan

Benchmarking Ultrasound Foundation Models for Fetal Plane Classification

Model ReleasesDGX agent

arXiv:2605.27796v1 Announce Type: cross Abstract: Ultrasound is widely used in obstetric care due to its safety, accessibility, and real-time imaging. However, interpretation remains operator-dependen

https://mistral.ai/news/ai-now-summit-2026/

Model ReleasesDGX agent

Mistral AI announced its participation in or perspective on the AI Now Summit 2026, likely discussing developments in AI safety, ethics, or industry trends relevant to the conference. The announcement

MIRA: A Bilingual Benchmark for Medical Information Response Audit

Model ReleasesDGX agent

arXiv:2605.28025v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to provide public-facing health information, yet existing safety evaluations overlook whether respons

27 May 2026

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

Model ReleasesDGX agent

arXiv:2605.26438v1 Announce Type: cross Abstract: Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the

PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic

Model ReleasesDGX agent

arXiv:2511.20586v4 Announce Type: replace Abstract: Trustworthiness has become a key requirement for the deployment of artificial intelligence systems in safety-critical applications. Conventional eva

RCSP: Risk-Sensitive Conjectural Scenario Planning for Safe Dynamic Robot Navigation

Model ReleasesDGX agent

arXiv:2605.26348v1 Announce Type: new Abstract: Mobile robots can fail before they collide: a velocity that is safe now may commit the robot to a passage that moving obstacles will soon close. We stud

Towards Error-Free EHRs: Reasoning-Intensive Consistency Verification Between Clinical Notes and Structured Tables in Electronic Health Records

Model ReleasesDGX agent

arXiv:2605.26463v1 Announce Type: cross Abstract: Data consistency between unstructured clinical notes and structured tables in Electronic Health Records (EHRs) is essential for patient safety and cli

26 May 2026

An Effective-Rank Audit of Alignment-Induced Activation Shifts: Confound Control, Constructive Calibration, and Limits

Model ReleasesDGX agent

arXiv:2605.24583v1 Announce Type: cross Abstract: We audit alignment-induced shifts in residual-stream activations of three open-weight instruction-tuned LLMs (Llama-3.1-8B-Instruct, Gemma-2-9B-it, Qw

FoodMonitor: Benchmarking MLLMs for Explainable Compliance Analysis

Model ReleasesDGX agent

arXiv:2605.24503v1 Announce Type: cross Abstract: As AI-powered compliance monitoring becomes increasingly important in public governance and industrial safety, the ability to provide verifiable evide

Latent Q-Barrier Shielding for Safe In-Context Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.25267v1 Announce Type: cross Abstract: Safe in-context reinforcement learning (ICRL) adapts online from interaction history without test-time parameter updates while controlling episode cos

Retrying vs Resampling in AI Control

Model ReleasesDGX agent

arXiv:2605.26047v1 Announce Type: new Abstract: AI coding scaffolds like Claude Code and Codex use extit{retrying}: blocking actions flagged as risky and continuing the trajectory. We study retrying f

SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models

Model ReleasesDGX agent

arXiv:2605.25420v1 Announce Type: cross Abstract: Large language model safety evaluation remains heavily English-centered, leaving low-resource languages under-measured even when models are deployed g

25 May 2026

Decomposing and Measuring Evaluation Awareness

Model ReleasesDGX agent

arXiv:2605.23055v1 Announce Type: cross Abstract: Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results. Yet t

Evaluating Large Language Models in a Complex Hidden Role Game

Model ReleasesDGX agent

arXiv:2605.22826v1 Announce Type: cross Abstract: Quantifying the deceptive potential of Large Language Models (LLMs) is critical for AI safety, yet difficult to achieve in uncontrolled environments.

23 May 2026

Characterizing the Fault Response of the Intel Neural Compute Stick 2 Under Single-Pulse Electromagnetic Fault Injection

Model ReleasesDGX agent

arXiv:2605.22437v1 Announce Type: cross Abstract: Vision processing units and other commercial neural-network inference accelerators are increasingly deployed in safety-relevant edge applications, but

22 May 2026

A Task-Agnostic Algebraic Integrity Metric for Event-Camera Streams Toward SOTIF-Compliant Perception using Pearson Correlation Coefficient

Model ReleasesDGX agent

arXiv:2605.21500v1 Announce Type: cross Abstract: Event cameras have emerged as a high-bandwidth, low-latency sensing modality for safety-critical perception in automated driving systems (ADS), offeri

Heartbeat-Bound Hierarchical Credentials: Cryptographic Revocation for AI Agent Swarms

Local AiDGX agent

arXiv:2605.20704v1 Announce Type: cross Abstract: Autonomous AI agents that spawn sub-agent swarms create a safety gap: existing credential revocation mechanisms, OAuth~2.0 introspection, OCSP, and W3

VBFDD-Agent for Electric Vehicle Battery Fault Detection and Diagnosis: Descriptive Text Modeling of Battery Digital Signals

Local AiDGX agent

arXiv:2605.20742v1 Announce Type: new Abstract: With the rapid proliferation of electric vehicles, the safety and reliability of lithium-ion batteries have become critical concerns. Effective anomaly

21 May 2026

A Deployment Audit of Release-Side Risk in Conformal Triage under Prevalence Shift

Model ReleasesDGX agent

arXiv:2605.20956v1 Announce Type: new Abstract: Conformal triage converts predictive scores into deployment actions that either release a case, flag it for urgent attention, or defer it to human revie

Hyper-V2X: Hypernetworks for Estimating Epistemic and Aleatoric Uncertainty in Cooperative Bird's-Eye-View Semantic Segmentation

Model ReleasesDGX agent

arXiv:2605.21309v1 Announce Type: new Abstract: Cooperative perception enabled by Vehicle-to-Everything (V2X) communication enhances autonomous driving safety by creating a unified environmental repre

20 May 2026

Lens Privacy Sealing: A New Benchmark and Method for Physical Privacy-Preserving Action Recognition

Model ReleasesDGX agent

arXiv:2605.19578v1 Announce Type: cross Abstract: RGB camera-based surveillance systems enable human action recognition for public safety and healthcare, yet raise serious privacy concerns. Existing m

Minnesota prohibits prediction markets, promptly gets sued by Trump admin

IndustryDGX agent

Minnesota became the first state to ban prediction markets when Governor Tim Walz signed a public safety bill into law, making operation, promotion, or advertising of platforms like Kalshi and Polymar

Real-World On-Vehicle Evaluation of Embedding-Based Anomaly Detection

Model ReleasesDGX agent

arXiv:2605.19744v1 Announce Type: new Abstract: Detecting anomalies in traffic scenes is crucial for ensuring safety in autonomous driving, yet collecting representative anomalous data remains challen

The Evaluation Game: Beyond Static LLM Benchmarking

Model ReleasesDGX agent

arXiv:2605.19377v1 Announce Type: cross Abstract: As jailbreaks, adversarially crafted inputs that bypass safety constraints, continue to be discovered in Large Language Models, practitioners increasi

19 May 2026

Advancing content provenance for a safer, more transparent AI ecosystem

Model ReleasesDGX agent

OpenAI discusses methods for establishing content provenance—tracking the origin and history of digital content—to improve transparency and safety in AI systems. The work addresses how verifiable cont

Adversarial Agent Collaboration for Correctness Improvements of C to Safe Rust Translation

Model ReleasesDGX agent

arXiv:2510.03879v3 Announce Type: replace-cross Abstract: Translating C to memory-safe languages, like Rust, prevents critical memory safety vulnerabilities that are prevalent in legacy C software. Ev

How to safeguard AI workloads with Unity AI Gateway Guardrails

TutorialsDGX agent

Unity AI Gateway Guardrails is a Databricks feature designed to protect AI workloads by implementing safety measures and policy controls within the Unity Catalog framework. The tool enables organizati

M^2FedAQI: Multimodal Federated Learning for Air Quality Prediction on Heterogeneous Edge Devices

Model ReleasesDGX agent

arXiv:2605.16375v1 Announce Type: new Abstract: Accurate air quality prediction is essential for public health, environmental monitoring, and industrial safety. However, most existing approaches rely

SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening

Model ReleasesDGX agent

arXiv:2605.17610v1 Announce Type: cross Abstract: The rapid growth of online video platforms and AI-generated content has made reliable video guardrails a key challenge for safety and real-world deplo

SwordBench: Evaluating Orthogonality of Steering Image Representations

Model ReleasesDGX agent

arXiv:2605.16372v1 Announce Type: cross Abstract: Steering or intervening on model representations at inference time to correct predictions is essential for AI interpretability and safety, yet existin

18 May 2026

Ti-iLSTM: A TinyDL Approach for Logic-Level Anomaly Detection in Industrial Water Treatment Systems

Local AiDGX agent

arXiv:2605.15874v1 Announce Type: new Abstract: Industrial Water Treatment Systems (IWTS) are safety critical cyber-physical infrastructures and due to increased connectivity, these systems are expose

15 May 2026

MahaVar: OOD Detection via Class-wise Mahalanobis Distance Variance under Neural Collapse

Model ReleasesDGX agent

arXiv:2605.14413v1 Announce Type: cross Abstract: Out-of-distribution (OOD) detection is a critical component for ensuring the reliability of deep neural networks in safety-critical applications. In t

Safe Bayesian Optimization for Complex Control Systems via Additive Gaussian Processes

Model ReleasesDGX agent

arXiv:2408.16307v3 Announce Type: replace-cross Abstract: Automatic controller tuning is attractive for robotics and mechatronic systems whose dynamics are difficult to model accurately, but direct bl

14 May 2026

Inference-Time Machine Unlearning via Gated Activation Redirection

Model ReleasesDGX agent

arXiv:2605.12765v1 Announce Type: new Abstract: Large Language Models memorize vast amounts of training data, raising concerns regarding privacy, copyright infringement, and safety. Machine unlearning

Safe Bayesian Optimization for Uncertain Correlations Matrices in Linear Models of Co-Regionalization

Model ReleasesDGX agent

arXiv:2605.13302v1 Announce Type: new Abstract: This paper extends safety guarantees for multi-task Bayesian optimization with uncertain correlation matrices from intrinsic co-reginalization models to

VERA-MH Concept Paper

Model ReleasesDGX agent

arXiv:2510.15297v4 Announce Type: replace-cross Abstract: We introduce VERA-MH (Validation of Ethical and Responsible AI in Mental Health), an automated evaluation of the safety of AI chatbots used in

12 May 2026

Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values

Model ReleasesDGX agent

arXiv:2605.10365v1 Announce Type: new Abstract: Autonomous agents have rapidly matured as task executors and seen widespread deployment via harnesses such as OpenClaw. Safety concerns have rightly dra

Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers

Model ReleasesDGX agent

arXiv:2605.10901v1 Announce Type: new Abstract: Guardrail Classifiers defend production language models against harmful behavior, but although results seem promising in testing, they provide no formal

CALYREX: Cross-Attention LaYeR EXtended Transformers for System Prompt Anchoring

Model ReleasesDGX agent

arXiv:2605.09737v1 Announce Type: new Abstract: Modern large language models (LLMs) rely on system prompts to establish behavioral constraints and safety rules. Standard causal self-attention treats p

Smart Railway Obstruction Detection System using IoT and Computer Vision

Local AiDGX agent

arXiv:2605.08246v1 Announce Type: new Abstract: Railway track intrusions pose a critical safety challenge for Indian Railways, encompassing wildlife incursions and deliberate malicious obstructions. T

Uncensored LLM

Local AiDGX agent

Uncensored LLMs are architectures that have been modified or fine-tuned to remove standard safety alignment layers (guardrails) that limit a model's ability to discuss sensitive topics. Ollama offers

11 May 2026

CarCrashNet: A Large-Scale Dataset and Hierarchical Neural Solver for Data-Driven Structural Crash Simulation

Model ReleasesDGX agent

arXiv:2605.07098v1 Announce Type: new Abstract: Crash simulation is a cornerstone of modern vehicle development because it reduces the need for costly physical prototypes, accelerates safety-driven de

Effective and Memory-Efficient Alternatives to ECC for Reliable Large-Scale DNNs

Model ReleasesDGX agent

arXiv:2605.07417v1 Announce Type: cross Abstract: Modern Deep Learning (DL) workloads are increasingly deployed in safety-critical domains, such as automotive systems and hyperscale data centers, wher

Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs

Model ReleasesDGX agent

arXiv:2605.06669v1 Announce Type: cross Abstract: Educational LLM tutors face a core AI alignment challenge: they must follow user intent while preserving pedagogical constraints and safety policies.

Multi-Objective Constraint Inference using Inverse reinforcement learning

Model ReleasesDGX agent

arXiv:2605.06951v1 Announce Type: new Abstract: Constraint inference is widely considered essential to align reinforcement learning agents with safety boundaries and operational guidelines by observin

RuleSafe-VL: Evaluating Rule-Conditioned Decision Reasoning in Vision-Language Content Moderation

Model ReleasesDGX agent

arXiv:2605.07760v1 Announce Type: new Abstract: Platform content moderation applies explicit policy rules and context-dependent conditions to decide whether user content is allowed, restricted, or rem

8 May 2026

lots of very interesting items here. They talk about pretty much everything from how to accelerate capabilities to concerns about agent safe…

Model ReleasesDGX agent

lots of very interesting items here. They talk about pretty much everything from how to accelerate capabilities to concerns about agent safety. I'm devastated to inform doomers that 'full stack open s

Running Codex safely at OpenAI

AgentsDGX agent

This OpenAI article outlines safety practices and guidelines for responsibly deploying and using Codex, their code generation model. It likely covers potential risks associated with automated code gen

Trump reportedly plans to fire FDA Commissioner Marty Makary

IndustryDGX agent

President Trump plans to fire FDA Commissioner Marty Makary following months of chaos at the agency . Makary has faced criticism over his handling of flavored vapes and for slow-walking a safety study

7 May 2026

How BASF manages thousands of supply chain decisions with AlphaEvolve’s agentic algorithms

Model ReleasesDGX agent

The agricultural and crop protection supply chain is one of the most intricate networks in the world. It takes up to two years to turn active ingredients into the final products farmers need, and a si

6 May 2026

Training-Free Probabilistic Time-Series Forecasting with Conformal Seasonal Pools

Model ReleasesDGX agent

arXiv:2605.03789v1 Announce Type: cross Abstract: We propose Conformal Seasonal Pools (CSP), a training-free probabilistic time-series forecaster that mixes same-season empirical draws with signed res

5 May 2026

Bridging the Experimental Last Mile: Digitizing Laboratory Know-How for Safe AI-Assisted Support

Local AiDGX agent

arXiv:2604.16345v2 Announce Type: replace-cross Abstract: While advances in materials informatics have accelerated the development of Self-Driving Laboratories (SDLs), human-led experiments remain sta

ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming

Model ReleasesDGX agent

arXiv:2605.02647v1 Announce Type: new Abstract: Large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safety alignment and elicit harmful responses. A growing body of work sh

ENFORCE: Nonlinear Constrained Learning with Adaptive-depth Neural Projection

Model ReleasesDGX agent

arXiv:2502.06774v4 Announce Type: replace Abstract: Ensuring neural networks adhere to domain-specific constraints is crucial for addressing safety and trustworthiness while also enhancing inference a

Finite-Sample Analysis of Elimination in Active Hypothesis Testing

Model ReleasesDGX agent

arXiv:2605.01039v1 Announce Type: new Abstract: A fixed-confidence, finite-sample problem of active hypothesis testing arises in many safety-critical applications. Situated in the context of sequentia

← Previous
1…221222223224225…240
Next →