AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Local Ai

Ti-iLSTM: A TinyDL Approach for Logic-Level Anomaly Detection in Industrial Water Treatment Systems

DGX agent

arXiv:2605.15874v1 Announce Type: new Abstract: Industrial Water Treatment Systems (IWTS) are safety critical cyber-physical infrastructures and due to increased connectivity, these systems are expose

local-aiarxiv-cs-lg
18 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

MahaVar: OOD Detection via Class-wise Mahalanobis Distance Variance under Neural Collapse

DGX agent

arXiv:2605.14413v1 Announce Type: cross Abstract: Out-of-distribution (OOD) detection is a critical component for ensuring the reliability of deep neural networks in safety-critical applications. In t

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Safe Bayesian Optimization for Complex Control Systems via Additive Gaussian Processes

DGX agent

arXiv:2408.16307v3 Announce Type: replace-cross Abstract: Automatic controller tuning is attractive for robotics and mechatronic systems whose dynamics are difficult to model accurately, but direct bl

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Inference-Time Machine Unlearning via Gated Activation Redirection

DGX agent

arXiv:2605.12765v1 Announce Type: new Abstract: Large Language Models memorize vast amounts of training data, raising concerns regarding privacy, copyright infringement, and safety. Machine unlearning

model-releasesarxiv-cs-lg
14 May 2026
Model Releases

Safe Bayesian Optimization for Uncertain Correlations Matrices in Linear Models of Co-Regionalization

DGX agent

arXiv:2605.13302v1 Announce Type: new Abstract: This paper extends safety guarantees for multi-task Bayesian optimization with uncertain correlation matrices from intrinsic co-reginalization models to

model-releasesarxiv-cs-lg
14 May 2026
Model Releases

VERA-MH Concept Paper

DGX agent

arXiv:2510.15297v4 Announce Type: replace-cross Abstract: We introduce VERA-MH (Validation of Ethical and Responsible AI in Mental Health), an automated evaluation of the safety of AI chatbots used in

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values

DGX agent

arXiv:2605.10365v1 Announce Type: new Abstract: Autonomous agents have rapidly matured as task executors and seen widespread deployment via harnesses such as OpenClaw. Safety concerns have rightly dra

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers

DGX agent

arXiv:2605.10901v1 Announce Type: new Abstract: Guardrail Classifiers defend production language models against harmful behavior, but although results seem promising in testing, they provide no formal

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

CALYREX: Cross-Attention LaYeR EXtended Transformers for System Prompt Anchoring

DGX agent

arXiv:2605.09737v1 Announce Type: new Abstract: Modern large language models (LLMs) rely on system prompts to establish behavioral constraints and safety rules. Standard causal self-attention treats p

model-releasesarxiv-cs-lg
12 May 2026
Local Ai

Smart Railway Obstruction Detection System using IoT and Computer Vision

DGX agent

arXiv:2605.08246v1 Announce Type: new Abstract: Railway track intrusions pose a critical safety challenge for Indian Railways, encompassing wildlife incursions and deliberate malicious obstructions. T

local-aiarxiv-cs-cv
12 May 2026
Local Ai

Uncensored LLM

DGX agent

Uncensored LLMs are architectures that have been modified or fine-tuned to remove standard safety alignment layers (guardrails) that limit a model's ability to discuss sensitive topics. Ollama offers

local-air-ollama
12 May 2026
Model Releases

CarCrashNet: A Large-Scale Dataset and Hierarchical Neural Solver for Data-Driven Structural Crash Simulation

DGX agent

arXiv:2605.07098v1 Announce Type: new Abstract: Crash simulation is a cornerstone of modern vehicle development because it reduces the need for costly physical prototypes, accelerates safety-driven de

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Effective and Memory-Efficient Alternatives to ECC for Reliable Large-Scale DNNs

DGX agent

arXiv:2605.07417v1 Announce Type: cross Abstract: Modern Deep Learning (DL) workloads are increasingly deployed in safety-critical domains, such as automotive systems and hyperscale data centers, wher

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs

DGX agent

arXiv:2605.06669v1 Announce Type: cross Abstract: Educational LLM tutors face a core AI alignment challenge: they must follow user intent while preserving pedagogical constraints and safety policies.

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Multi-Objective Constraint Inference using Inverse reinforcement learning

DGX agent

arXiv:2605.06951v1 Announce Type: new Abstract: Constraint inference is widely considered essential to align reinforcement learning agents with safety boundaries and operational guidelines by observin

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

RuleSafe-VL: Evaluating Rule-Conditioned Decision Reasoning in Vision-Language Content Moderation

DGX agent

arXiv:2605.07760v1 Announce Type: new Abstract: Platform content moderation applies explicit policy rules and context-dependent conditions to decide whether user content is allowed, restricted, or rem

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

lots of very interesting items here. They talk about pretty much everything from how to accelerate capabilities to concerns about agent safe…

DGX agent

lots of very interesting items here. They talk about pretty much everything from how to accelerate capabilities to concerns about agent safety. I'm devastated to inform doomers that 'full stack open s

model-releasesclem-delangue--x
8 May 2026
Agents

Running Codex safely at OpenAI

DGX agent

This OpenAI article outlines safety practices and guidelines for responsibly deploying and using Codex, their code generation model. It likely covers potential risks associated with automated code gen

agentsopenai
8 May 2026
Industry

Trump reportedly plans to fire FDA Commissioner Marty Makary

DGX agent

President Trump plans to fire FDA Commissioner Marty Makary following months of chaos at the agency . Makary has faced criticism over his handling of flavored vapes and for slow-walking a safety study

industryars-technica
8 May 2026
Model Releases

How BASF manages thousands of supply chain decisions with AlphaEvolve’s agentic algorithms

DGX agent

The agricultural and crop protection supply chain is one of the most intricate networks in the world. It takes up to two years to turn active ingredients into the final products farmers need, and a si

model-releasesgoogle-cloud-ai
7 May 2026
Model Releases

Training-Free Probabilistic Time-Series Forecasting with Conformal Seasonal Pools

DGX agent

arXiv:2605.03789v1 Announce Type: cross Abstract: We propose Conformal Seasonal Pools (CSP), a training-free probabilistic time-series forecaster that mixes same-season empirical draws with signed res

model-releasesarxiv-cs-lg
6 May 2026
Local Ai

Bridging the Experimental Last Mile: Digitizing Laboratory Know-How for Safe AI-Assisted Support

DGX agent

arXiv:2604.16345v2 Announce Type: replace-cross Abstract: While advances in materials informatics have accelerated the development of Self-Driving Laboratories (SDLs), human-led experiments remain sta

local-aiarxiv-cs-ai
5 May 2026
Model Releases

ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming

DGX agent

arXiv:2605.02647v1 Announce Type: new Abstract: Large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safety alignment and elicit harmful responses. A growing body of work sh

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

ENFORCE: Nonlinear Constrained Learning with Adaptive-depth Neural Projection

DGX agent

arXiv:2502.06774v4 Announce Type: replace Abstract: Ensuring neural networks adhere to domain-specific constraints is crucial for addressing safety and trustworthiness while also enhancing inference a

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

Finite-Sample Analysis of Elimination in Active Hypothesis Testing

DGX agent

arXiv:2605.01039v1 Announce Type: new Abstract: A fixed-confidence, finite-sample problem of active hypothesis testing arises in many safety-critical applications. Situated in the context of sequentia

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models

DGX agent

arXiv:2605.00123v1 Announce Type: new Abstract: Safety trained large language models (LLMs) can often be induced to answer harmful requests through jailbreak prompts. Because we lack a robust understa

model-releasesarxiv-cs-ai
5 May 2026
Model Releases

Understanding Emergent Misalignment via Feature Superposition Geometry

DGX agent

arXiv:2605.00842v1 Announce Type: cross Abstract: Emergent misalignment, where fine-tuning on narrow, non-harmful tasks induces harmful behaviors, poses a key challenge for AI safety in LLMs. Despite

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

WILD SAM: A Simulated-and-Real Data Augmentation for Autonomous Driving Perception under Challenging Weather

DGX agent

arXiv:2605.01081v1 Announce Type: new Abstract: The performance of state-of-the-art object detectors degrades significantly under adverse weather, causing a safety-critical domain shift problem for au

model-releasesarxiv-cs-cv
5 May 2026
Local Ai

OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving

DGX agent

arXiv:2512.14044v3 Announce Type: replace-cross Abstract: The deployment of Vision-Language Models (VLMs) in safety-critical domains like autonomous driving (AD) is critically hindered by reliability

local-aiarxiv-cs-ai
1 May 2026
Model Releases

Lyapunov-Guided Self-Alignment: Test-Time Adaptation for Offline Safe Reinforcement Learning

DGX agent

arXiv:2604.26516v1 Announce Type: cross Abstract: Offline reinforcement learning (RL) agents often fail when deployed, as the gap between training datasets and real environments leads to unsafe behavi

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

Below-Chance Blindness: Prompted Underperformance in Small LLMs Produces Positional Bias Rather than Answer Avoidance

DGX agent

arXiv:2604.25249v1 Announce Type: new Abstract: Detecting sandbagging--the deliberate underperformance on capability evaluations--is an open problem in AI safety. We tested whether symptom validity te

model-releasesarxiv-cs-cl
29 Apr 2026
Model Releases

CAN-QA: A Question-Answering Benchmark for Reasoning over In-Vehicle CAN Traffic

DGX agent

arXiv:2604.24935v1 Announce Type: cross Abstract: The Controller Area Network (CAN) is a safety-critical in-vehicle communication protocol that lacks built-in security mechanisms, making intrusion det

model-releasesarxiv-cs-lg
29 Apr 2026
Model Releases

Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver

DGX agent

arXiv:2604.25067v1 Announce Type: cross Abstract: Forecasting when AI systems will become capable of meaningfully accelerating AI research is a central challenge for AI safety. Existing benchmarks mea

model-releasesarxiv-cs-lg
29 Apr 2026
Industry

School-shooting lawsuits accuse OpenAI of hiding violent ChatGPT users

DGX agent

Seven California-based lawsuits allege that OpenAI could have prevented Canada’s deadliest school shooting by not reporting a flagged ChatGPT account to law‑enforcement. Safety‑team experts had identi

industryars-technica
29 Apr 2026
Model Releases

GCP: Guarded Collaborative Perception with Spatial-Temporal Aware Malicious Agent Detection

DGX agent

arXiv:2501.02450v2 Announce Type: replace Abstract: Collaborative perception significantly enhances autonomous driving safety by extending each vehicle's perception range through message sharing among

model-releasesarxiv-cs-cv
28 Apr 2026
Model Releases

Learning Under Low Illumination: A Dataset and Algorithm for Traffic Sign Recognition

DGX agent

arXiv:2511.17183v2 Announce Type: replace Abstract: Traffic signboards are vital for road safety and intelligent transportation systems, enabling navigation and autonomous driving. Yet, recognizing tr

model-releasesarxiv-cs-cv
28 Apr 2026
Model Releases

LOCAL AI MODELS ARE CATCHING UP TO FRONTIER MODELS WAY FASTER THAN ANYONE EXPECTED this guy ran qwen 3.6 27B locally on a base macbook pro M…

DGX agent

LOCAL AI MODELS ARE CATCHING UP TO FRONTIER MODELS WAY FASTER THAN ANYONE EXPECTED this guy ran qwen 3.6 27B locally on a base macbook pro M4 with 24GB of memory quantized and stripped of safety guard

model-releasesclem-delangue--x
28 Apr 2026
Model Releases

Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings

DGX agent

arXiv:2604.23130v1 Announce Type: cross Abstract: Large language models (LLMs) can still be jailbroken into producing harmful outputs despite safety alignment. Existing attacks show this vulnerability

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications

DGX agent

arXiv:2604.24668v1 Announce Type: new Abstract: Given the increased use of LLMs in financial systems today, it becomes important to evaluate the safety and robustness of such systems. One failure mode

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

TRACE: Topology-aware Reconstruction of Accidents in CARLA for AV Evaluation

DGX agent

arXiv:2604.22068v1 Announce Type: cross Abstract: Validating Autonomous Vehicles (AVs) requires exposure to rare, safety-critical scenarios, infrequent in routine driving data. Existing benchmarks add

model-releasesarxiv-cs-ro
27 Apr 2026
Model Releases

Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers

DGX agent

arXiv:2604.21700v1 Announce Type: cross Abstract: The growing application of large language models (LLMs) in safety-critical domains has raised urgent concerns about their security. Many recent studie

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models

DGX agent

arXiv:2604.21860v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly integrated into sensitive workflows, raising the stakes for adversarial robustness and safety. This pape

model-releasesarxiv-cs-ai
24 Apr 2026
Local Ai

Ufil: A Unified Framework for Infrastructure-based Localization

DGX agent

arXiv:2604.21471v1 Announce Type: new Abstract: Infrastructure-based localization enhances road safety and traffic management by providing state estimates of road users. Development is hindered by fra

local-aiarxiv-cs-ro
24 Apr 2026
Model Releases

When to Trust the Answer: Question-Aligned Semantic Nearest Neighbor Entropy for Safer Surgical VQA

DGX agent

arXiv:2511.01458v2 Announce Type: replace-cross Abstract: Safety and reliability are critical for deploying visual question answering (VQA) systems in surgery, where incorrect or ambiguous responses c

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

1. We believe in iterative deployment; although GPT-5.5 is already a smart model, we expect rapid improvements. Iterative deployment is a bi…

DGX agent

1. We believe in iterative deployment; although GPT-5.5 is already a smart model, we expect rapid improvements. Iterative deployment is a big part of our safety strategy; we believe the world will be

model-releasessam-altman--x
23 Apr 2026
Model Releases

Automated Detection of Dosing Errors in Clinical Trial Narratives: A Multi-Modal Feature Engineering Approach with LightGBM

DGX agent

arXiv:2604.19759v1 Announce Type: new Abstract: Clinical trials require strict adherence to medication protocols, yet dosing errors remain a persistent challenge affecting patient safety and trial int

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control

DGX agent

arXiv:2601.02896v2 Announce Type: replace Abstract: Controlling emergent behavioral personas (e.g., sycophancy, hallucination) in Large Language Models (LLMs) is critical for AI safety, yet remains a

model-releasesarxiv-cs-lg
23 Apr 2026
Model Releases

CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs

DGX agent

arXiv:2604.20460v1 Announce Type: new Abstract: Safety-critical traffic reasoning requires contrastive consistency: models must detect true hazards when an accident occurs, and reliably reject plausib

model-releasesarxiv-cs-cv
23 Apr 2026
← Previous
1…275276277278279…297
Next →