AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Model Releases

Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles

DGX agent

arXiv:2605.29473v1 Announce Type: cross Abstract: Language models are increasingly being deployed for conversational support in informal caregiving contexts, where interactions often extend beyond inf

model-releasesarxiv-cs-ai
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

DGX agent

arXiv:2605.29114v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Adversarial Fine-tuning of Compressed Neural Networks for Joint Improvement of Robustness and Efficiency

DGX agent

arXiv:2403.09441v2 Announce Type: replace Abstract: As deep learning (DL) models are increasingly being integrated into our everyday lives, ensuring their safety by making them robust against adversar

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Benchmarking Ultrasound Foundation Models for Fetal Plane Classification

DGX agent

arXiv:2605.27796v1 Announce Type: cross Abstract: Ultrasound is widely used in obstetric care due to its safety, accessibility, and real-time imaging. However, interpretation remains operator-dependen

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

MIRA: A Bilingual Benchmark for Medical Information Response Audit

DGX agent

arXiv:2605.28025v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to provide public-facing health information, yet existing safety evaluations overlook whether respons

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

DGX agent

arXiv:2605.26438v1 Announce Type: cross Abstract: Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic

DGX agent

arXiv:2511.20586v4 Announce Type: replace Abstract: Trustworthiness has become a key requirement for the deployment of artificial intelligence systems in safety-critical applications. Conventional eva

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

RCSP: Risk-Sensitive Conjectural Scenario Planning for Safe Dynamic Robot Navigation

DGX agent

arXiv:2605.26348v1 Announce Type: new Abstract: Mobile robots can fail before they collide: a velocity that is safe now may commit the robot to a passage that moving obstacles will soon close. We stud

model-releasesarxiv-cs-ro
27 May 2026
Model Releases

Towards Error-Free EHRs: Reasoning-Intensive Consistency Verification Between Clinical Notes and Structured Tables in Electronic Health Records

DGX agent

arXiv:2605.26463v1 Announce Type: cross Abstract: Data consistency between unstructured clinical notes and structured tables in Electronic Health Records (EHRs) is essential for patient safety and cli

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

An Effective-Rank Audit of Alignment-Induced Activation Shifts: Confound Control, Constructive Calibration, and Limits

DGX agent

arXiv:2605.24583v1 Announce Type: cross Abstract: We audit alignment-induced shifts in residual-stream activations of three open-weight instruction-tuned LLMs (Llama-3.1-8B-Instruct, Gemma-2-9B-it, Qw

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

FoodMonitor: Benchmarking MLLMs for Explainable Compliance Analysis

DGX agent

arXiv:2605.24503v1 Announce Type: cross Abstract: As AI-powered compliance monitoring becomes increasingly important in public governance and industrial safety, the ability to provide verifiable evide

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Latent Q-Barrier Shielding for Safe In-Context Reinforcement Learning

DGX agent

arXiv:2605.25267v1 Announce Type: cross Abstract: Safe in-context reinforcement learning (ICRL) adapts online from interaction history without test-time parameter updates while controlling episode cos

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Retrying vs Resampling in AI Control

DGX agent

arXiv:2605.26047v1 Announce Type: new Abstract: AI coding scaffolds like Claude Code and Codex use extit{retrying}: blocking actions flagged as risky and continuing the trajectory. We study retrying f

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models

DGX agent

arXiv:2605.25420v1 Announce Type: cross Abstract: Large language model safety evaluation remains heavily English-centered, leaving low-resource languages under-measured even when models are deployed g

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Decomposing and Measuring Evaluation Awareness

DGX agent

arXiv:2605.23055v1 Announce Type: cross Abstract: Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results. Yet t

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Evaluating Large Language Models in a Complex Hidden Role Game

DGX agent

arXiv:2605.22826v1 Announce Type: cross Abstract: Quantifying the deceptive potential of Large Language Models (LLMs) is critical for AI safety, yet difficult to achieve in uncontrolled environments.

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Characterizing the Fault Response of the Intel Neural Compute Stick 2 Under Single-Pulse Electromagnetic Fault Injection

DGX agent

arXiv:2605.22437v1 Announce Type: cross Abstract: Vision processing units and other commercial neural-network inference accelerators are increasingly deployed in safety-relevant edge applications, but

model-releasesarxiv-cs-lg
23 May 2026
Model Releases

A Task-Agnostic Algebraic Integrity Metric for Event-Camera Streams Toward SOTIF-Compliant Perception using Pearson Correlation Coefficient

DGX agent

arXiv:2605.21500v1 Announce Type: cross Abstract: Event cameras have emerged as a high-bandwidth, low-latency sensing modality for safety-critical perception in automated driving systems (ADS), offeri

model-releasesarxiv-cs-cv
22 May 2026
Local Ai

Heartbeat-Bound Hierarchical Credentials: Cryptographic Revocation for AI Agent Swarms

DGX agent

arXiv:2605.20704v1 Announce Type: cross Abstract: Autonomous AI agents that spawn sub-agent swarms create a safety gap: existing credential revocation mechanisms, OAuth~2.0 introspection, OCSP, and W3

local-aiarxiv-cs-ai
22 May 2026
Local Ai

VBFDD-Agent for Electric Vehicle Battery Fault Detection and Diagnosis: Descriptive Text Modeling of Battery Digital Signals

DGX agent

arXiv:2605.20742v1 Announce Type: new Abstract: With the rapid proliferation of electric vehicles, the safety and reliability of lithium-ion batteries have become critical concerns. Effective anomaly

local-aiarxiv-cs-ai
22 May 2026
Model Releases

A Deployment Audit of Release-Side Risk in Conformal Triage under Prevalence Shift

DGX agent

arXiv:2605.20956v1 Announce Type: new Abstract: Conformal triage converts predictive scores into deployment actions that either release a case, flag it for urgent attention, or defer it to human revie

model-releasesarxiv-cs-lg
21 May 2026
Model Releases

Hyper-V2X: Hypernetworks for Estimating Epistemic and Aleatoric Uncertainty in Cooperative Bird's-Eye-View Semantic Segmentation

DGX agent

arXiv:2605.21309v1 Announce Type: new Abstract: Cooperative perception enabled by Vehicle-to-Everything (V2X) communication enhances autonomous driving safety by creating a unified environmental repre

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

Lens Privacy Sealing: A New Benchmark and Method for Physical Privacy-Preserving Action Recognition

DGX agent

arXiv:2605.19578v1 Announce Type: cross Abstract: RGB camera-based surveillance systems enable human action recognition for public safety and healthcare, yet raise serious privacy concerns. Existing m

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Real-World On-Vehicle Evaluation of Embedding-Based Anomaly Detection

DGX agent

arXiv:2605.19744v1 Announce Type: new Abstract: Detecting anomalies in traffic scenes is crucial for ensuring safety in autonomous driving, yet collecting representative anomalous data remains challen

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

The Evaluation Game: Beyond Static LLM Benchmarking

DGX agent

arXiv:2605.19377v1 Announce Type: cross Abstract: As jailbreaks, adversarially crafted inputs that bypass safety constraints, continue to be discovered in Large Language Models, practitioners increasi

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Adversarial Agent Collaboration for Correctness Improvements of C to Safe Rust Translation

DGX agent

arXiv:2510.03879v3 Announce Type: replace-cross Abstract: Translating C to memory-safe languages, like Rust, prevents critical memory safety vulnerabilities that are prevalent in legacy C software. Ev

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

M^2FedAQI: Multimodal Federated Learning for Air Quality Prediction on Heterogeneous Edge Devices

DGX agent

arXiv:2605.16375v1 Announce Type: new Abstract: Accurate air quality prediction is essential for public health, environmental monitoring, and industrial safety. However, most existing approaches rely

model-releasesarxiv-cs-lg
19 May 2026
Model Releases

SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening

DGX agent

arXiv:2605.17610v1 Announce Type: cross Abstract: The rapid growth of online video platforms and AI-generated content has made reliable video guardrails a key challenge for safety and real-world deplo

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

SwordBench: Evaluating Orthogonality of Steering Image Representations

DGX agent

arXiv:2605.16372v1 Announce Type: cross Abstract: Steering or intervening on model representations at inference time to correct predictions is essential for AI interpretability and safety, yet existin

model-releasesarxiv-cs-ai
19 May 2026
Local Ai

Ti-iLSTM: A TinyDL Approach for Logic-Level Anomaly Detection in Industrial Water Treatment Systems

DGX agent

arXiv:2605.15874v1 Announce Type: new Abstract: Industrial Water Treatment Systems (IWTS) are safety critical cyber-physical infrastructures and due to increased connectivity, these systems are expose

local-aiarxiv-cs-lg
18 May 2026
Model Releases

MahaVar: OOD Detection via Class-wise Mahalanobis Distance Variance under Neural Collapse

DGX agent

arXiv:2605.14413v1 Announce Type: cross Abstract: Out-of-distribution (OOD) detection is a critical component for ensuring the reliability of deep neural networks in safety-critical applications. In t

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Safe Bayesian Optimization for Complex Control Systems via Additive Gaussian Processes

DGX agent

arXiv:2408.16307v3 Announce Type: replace-cross Abstract: Automatic controller tuning is attractive for robotics and mechatronic systems whose dynamics are difficult to model accurately, but direct bl

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Inference-Time Machine Unlearning via Gated Activation Redirection

DGX agent

arXiv:2605.12765v1 Announce Type: new Abstract: Large Language Models memorize vast amounts of training data, raising concerns regarding privacy, copyright infringement, and safety. Machine unlearning

model-releasesarxiv-cs-lg
14 May 2026
Model Releases

Safe Bayesian Optimization for Uncertain Correlations Matrices in Linear Models of Co-Regionalization

DGX agent

arXiv:2605.13302v1 Announce Type: new Abstract: This paper extends safety guarantees for multi-task Bayesian optimization with uncertain correlation matrices from intrinsic co-reginalization models to

model-releasesarxiv-cs-lg
14 May 2026
Model Releases

VERA-MH Concept Paper

DGX agent

arXiv:2510.15297v4 Announce Type: replace-cross Abstract: We introduce VERA-MH (Validation of Ethical and Responsible AI in Mental Health), an automated evaluation of the safety of AI chatbots used in

model-releasesarxiv-cs-ai
14 May 2026
Model Releases

Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values

DGX agent

arXiv:2605.10365v1 Announce Type: new Abstract: Autonomous agents have rapidly matured as task executors and seen widespread deployment via harnesses such as OpenClaw. Safety concerns have rightly dra

model-releasesarxiv-cs-ai
12 May 2026
Model Releases

Beyond Red-Teaming: Formal Guarantees of LLM Guardrail Classifiers

DGX agent

arXiv:2605.10901v1 Announce Type: new Abstract: Guardrail Classifiers defend production language models against harmful behavior, but although results seem promising in testing, they provide no formal

model-releasesarxiv-cs-lg
12 May 2026
Model Releases

CALYREX: Cross-Attention LaYeR EXtended Transformers for System Prompt Anchoring

DGX agent

arXiv:2605.09737v1 Announce Type: new Abstract: Modern large language models (LLMs) rely on system prompts to establish behavioral constraints and safety rules. Standard causal self-attention treats p

model-releasesarxiv-cs-lg
12 May 2026
Local Ai

Smart Railway Obstruction Detection System using IoT and Computer Vision

DGX agent

arXiv:2605.08246v1 Announce Type: new Abstract: Railway track intrusions pose a critical safety challenge for Indian Railways, encompassing wildlife incursions and deliberate malicious obstructions. T

local-aiarxiv-cs-cv
12 May 2026
Model Releases

CarCrashNet: A Large-Scale Dataset and Hierarchical Neural Solver for Data-Driven Structural Crash Simulation

DGX agent

arXiv:2605.07098v1 Announce Type: new Abstract: Crash simulation is a cornerstone of modern vehicle development because it reduces the need for costly physical prototypes, accelerates safety-driven de

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Effective and Memory-Efficient Alternatives to ECC for Reliable Large-Scale DNNs

DGX agent

arXiv:2605.07417v1 Announce Type: cross Abstract: Modern Deep Learning (DL) workloads are increasingly deployed in safety-critical domains, such as automotive systems and hyperscale data centers, wher

model-releasesarxiv-cs-lg
11 May 2026
Model Releases

Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs

DGX agent

arXiv:2605.06669v1 Announce Type: cross Abstract: Educational LLM tutors face a core AI alignment challenge: they must follow user intent while preserving pedagogical constraints and safety policies.

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Multi-Objective Constraint Inference using Inverse reinforcement learning

DGX agent

arXiv:2605.06951v1 Announce Type: new Abstract: Constraint inference is widely considered essential to align reinforcement learning agents with safety boundaries and operational guidelines by observin

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

RuleSafe-VL: Evaluating Rule-Conditioned Decision Reasoning in Vision-Language Content Moderation

DGX agent

arXiv:2605.07760v1 Announce Type: new Abstract: Platform content moderation applies explicit policy rules and context-dependent conditions to decide whether user content is allowed, restricted, or rem

model-releasesarxiv-cs-ai
11 May 2026
Model Releases

Training-Free Probabilistic Time-Series Forecasting with Conformal Seasonal Pools

DGX agent

arXiv:2605.03789v1 Announce Type: cross Abstract: We propose Conformal Seasonal Pools (CSP), a training-free probabilistic time-series forecaster that mixes same-season empirical draws with signed res

model-releasesarxiv-cs-lg
6 May 2026
Local Ai

Bridging the Experimental Last Mile: Digitizing Laboratory Know-How for Safe AI-Assisted Support

DGX agent

arXiv:2604.16345v2 Announce Type: replace-cross Abstract: While advances in materials informatics have accelerated the development of Self-Driving Laboratories (SDLs), human-led experiments remain sta

local-aiarxiv-cs-ai
5 May 2026
Model Releases

ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming

DGX agent

arXiv:2605.02647v1 Announce Type: new Abstract: Large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safety alignment and elicit harmful responses. A growing body of work sh

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

ENFORCE: Nonlinear Constrained Learning with Adaptive-depth Neural Projection

DGX agent

arXiv:2502.06774v4 Announce Type: replace Abstract: Ensuring neural networks adhere to domain-specific constraints is crucial for addressing safety and trustworthiness while also enhancing inference a

model-releasesarxiv-cs-lg
5 May 2026
← Previous
1…236237238239240…255
Next →