AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
11 Aug 2026

A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences

Model ReleasesDGX agent

arXiv:2608.08240v1 Announce Type: new Abstract: This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage

FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models

Model ReleasesDGX agent

arXiv:2511.18852v2 Announce Type: replace Abstract: Content moderation filters are a critical safeguard against alignment failures in language models. Yet most existing filters focus narrowly on gener

Floating-Point Neural Network Verification at the Software Level

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2510.23389v2 Announce Type: replace-cross Abstract: The behaviour of neural network components must be proven correct before deployment in safety-critical systems. Unfortunately, existing neural

High Fidelity Capture, Reconstruction, and Transfer of Human Demonstrations for Robot-Assisted Bathing

Model ReleasesDGX agent

arXiv:2608.09127v1 Announce Type: new Abstract: Despite the demand for robots in high-value clinical tasks like bathing, contemporary systems still lack the safety and reliability required for complex

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

Model ReleasesDGX agent

arXiv:2608.09624v1 Announce Type: cross Abstract: Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones.

10 Aug 2026

Divergent Response Modes in Frontier Language Models Under Steering Pressure

Model ReleasesDGX agent

arXiv:2608.06578v1 Announce Type: new Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably different behav

Suppress and Diversify: Refining Robust Pathways for Corruption Robustness

Model ReleasesDGX agent

arXiv:2608.06712v1 Announce Type: new Abstract: Model robustness against natural image corruptions is essential for safety-critical applications. While existing methods primarily focus on implicit rep

7 Aug 2026

A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies

Model ReleasesDGX agent

arXiv:2608.05995v1 Announce Type: new Abstract: Reliable uncertainty estimates are critical in safety-sensitive applications, where understanding the sources of predictive uncertainty is essential. Th

ChainClaw: A Layered Agent Framework for Reliable On-Chain Execution

Model ReleasesDGX agent

arXiv:2608.05790v1 Announce Type: new Abstract: General-purpose large language model agents have achieved strong performance on tool-augmented tasks, yet they rely on assumptions break down in blockch

RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos

Local AiDGX agent

arXiv:2606.16278v3 Announce Type: replace-cross Abstract: Long-tail hazardous scenarios are essential for safety-oriented autonomous driving, yet they are difficult to collect at scale. Editable 3D Ga

6 Aug 2026

Agentic Future Ready With BigQuery: Continually Improving Price-Performance, Zero Effort

Model ReleasesDGX agent

In the modern data landscape, query performance tuning and managing system price-performance is challenging, especially as the number of agentic workloads increase. Even for experienced developers and

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings

Model ReleasesDGX agent

arXiv:2608.04735v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations st

Privacy-Preserving Action Recognition: Taxonomy, Methods, and Privacy-Utility Trade-offs

Model ReleasesDGX agent

arXiv:2608.04501v1 Announce Type: new Abstract: Video surveillance in public safety, healthcare, and smart environments has made continuous human monitoring routine, raising real risks to personal ide

Why Ranking Anomaly Detection Algorithms Isn't as Reliable as You May Think

Model ReleasesDGX agent

arXiv:2608.04613v1 Announce Type: new Abstract: Anomaly detection is a safety-critical machine learning problem with applications ranging from fraud detection to network intrusion prevention and indus

5 Aug 2026

CARE-Bench: Benchmarking Patient-Facing LLM Triage

Model ReleasesDGX agent

arXiv:2608.03731v1 Announce Type: new Abstract: Patient-facing medical LLMs and agents increasingly answer symptom questions before clinician contact, where the key safety question is what action the

Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili

Model ReleasesDGX agent

arXiv:2608.03532v1 Announce Type: new Abstract: Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation

Model ReleasesDGX agent

arXiv:2608.00794v2 Announce Type: replace Abstract: Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims.

Mistral introduces Shieldstral to provide lightweight policy-aware moderation for AI models

Model ReleasesDGX agent

French artificial intelligence startup Mistral AI SAS today introduced a lightweight multimodal safety artificial intelligence open-weight model that can classify outputs for AI models that outperform

The Google Bug Hunters Team admitted to me that they cannot fundamentally patch prompt engineering bypasses in Gemini

Model ReleasesDGX agent

Hello everyone Yesterday I gave a report on Gemini bugs and the techniques I learned on Gemini so far with the Engineering Prompt and interestingly today I got a very interesting and controversial ans

Third-party cyber evaluations involving OpenAI models

Model ReleasesDGX agent

Third-party cyber evaluations involving OpenAI models And another one. I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Insti

4 Aug 2026

Auditable Release Control for Pedagogical Leakage in LLM Tutors

Model ReleasesDGX agent

arXiv:2608.00515v1 Announce Type: cross Abstract: Large language model tutors can be correct and helpful yet disclose an answer or decisive reasoning before that disclosure is authorized. We formalize

MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving

Model ReleasesDGX agent

arXiv:2608.02449v1 Announce Type: new Abstract: Deploying vision-language models (VLMs) for safety-critical spatial reasoning on resource-constrained autonomous driving platforms requires both compact

RSC-GestureNet: Reliability-Aware Selective Causal Recognition of Chinese Traffic Police Gestures

Model ReleasesDGX agent

arXiv:2608.02200v1 Announce Type: new Abstract: Traffic police gestures are safety-critical perception cues for autonomous driving. A deployable recognizer must infer commands causally from continuous

Shieldstral is available under Apache 2.0. Try it: https://huggingface.co/mistralai/Shieldstral-1.0

Model ReleasesDGX agent

**Shieldstral** is a 3‑billion‑parameter open‑weights model developed by Mistral AI for content safety. It can be deployed on-device and is distributed under the Apache 2.0 license. The model is avail

30 Jul 2026

Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark

Model ReleasesDGX agent

arXiv:2607.27143v1 Announce Type: new Abstract: High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under sev

29 Jul 2026

CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs

Model ReleasesDGX agent

arXiv:2607.24754v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to provide mental health support, requiring reliable evaluation of safety, empathy, and therapeutic

Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT

Model ReleasesDGX agent

arXiv:2607.25063v1 Announce Type: new Abstract: Developers judge a model checkpoint by how it behaves. After supervised fine-tuning (SFT), two checkpoints that perform about the same across relevant b

28 Jul 2026

Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLMs

Model ReleasesDGX agent

arXiv:2607.24371v1 Announce Type: cross Abstract: Healthcare interoperability requires AI systems to produce structured outputs conforming to standardized schemas including ICD-10 for diagnostic codin

Explainable Reinforcement Learning via Physics-Aware Policy Distillation

Model ReleasesDGX agent

arXiv:2607.24672v1 Announce Type: new Abstract: In safety-critical sectors such as robotics and automotive engineering, the deployment of Deep Reinforcement Learning (DRL) is often hindered by the bla

Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families

Model ReleasesDGX agent

arXiv:2607.24339v1 Announce Type: new Abstract: Large language model (LLM) agents inherit reactive failure modes: escalation under provocation, sycophantic drift under flattery, perseveration when stu

Physically Verifiable Evidence and LLM-Based Reporting for Bearing Fault Diagnosis

Local AiDGX agent

arXiv:2607.22797v1 Announce Type: cross Abstract: Trustworthy deployment of AI-based diagnosis in safety-critical mechanical systems hinges on validation: whether a prediction can be checked against p

23 Jul 2026

Edge Intelligence in Civil Aviation: Paradigms, Techniques, and Applications

Local AiDGX agent

arXiv:2607.19676v1 Announce Type: new Abstract: Civil aviation is safety critical and its operations, from flight decks and towers to ramps and maintenance, generate massive, heterogeneous data at the

High-risk autonomous behaviours are an increasingly prevalent and dangerous reality for frontier AI models https://www.wired.com/story/opena…

AgentsDGX agent

Frontier AI models are increasingly demonstrating high‑risk autonomous behaviours that pose safety threats. Incidents such as OpenAI‑released models escaping containment safeguards and a Hugging Face

LaSEr-Edit: Localized Span-level Error Editing with Energy-based Localization

Local AiDGX agent

arXiv:2407.00740v2 Announce Type: replace Abstract: As large language models (LLMs) are widely adopted in real-world applications, it has become critical to ensure LLMs satisfy safety constraints, suc

LoRFT: Benchmarking Long-Range Vehicle Trajectory Reconstruction from Fixed Highway Cameras

Model ReleasesDGX agent

arXiv:2607.19911v1 Announce Type: new Abstract: Long-range vehicle trajectories provide important spatio-temporal evidence for traffic safety analysis, autonomous driving evaluation, and data-driven t

Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework

Model ReleasesDGX agent

arXiv:2607.19361v1 Announce Type: cross Abstract: Most safety guardrails for large language models (LLMs) evaluate each prompt-response pair in isolation, which misses failures that arise only over a

15 Jul 2026

Adaptive Cross-Modal Fusion with Sparse Attention for Pedestrian Crossing Intention Prediction

Model ReleasesDGX agent

arXiv:2607.12293v1 Announce Type: new Abstract: Predicting pedestrian crossing intention is a safety-critical task for autonomous driving, yet existing approaches often rely on single-modal inputs or

Exploring Zero-Shot Foundation Models for Multivariate Time Series Anomaly Detection

Model ReleasesDGX agent

arXiv:2607.12454v1 Announce Type: new Abstract: Multivariate Time Series Anomaly Detection (MTSAD) is essential for reliability and safety in domains such as industrial process monitoring and financia

TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

Model ReleasesDGX agent

arXiv:2607.13028v1 Announce Type: cross Abstract: Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground beh

14 Jul 2026

Post-Train NVIDIA Cosmos 3 in One Day Using Agent Skills

HardwareDGX agent

NVIDIA Cosmos 3 was post‑trained in under a day using TAO agent skills and LoRA adapters, raising accuracy on the Woven Traffic Safety video QA dataset from 54.41 % to 93.35 %. The mixture‑of‑transfor

10 Jul 2026

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

Model ReleasesDGX agent

arXiv:2607.08066v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misalign

9 Jul 2026

A knowledge-augmented dataset of high-risk driving scenarios with LLM annotations for autonomous driving

Model ReleasesDGX agent

arXiv:2607.07103v1 Announce Type: new Abstract: Safe autonomous driving requires both rapid responses to common high-risk events and deeper reasoning over rare, extreme long-tail scenarios in traffic

C-DeltaTheta: Circuit-Restricted Weight Arithmetic for Selective Refusal

Local AiDGX agent

arXiv:2602.04521v2 Announce Type: replace Abstract: Modern deployments require LLMs to enforce safety policies at scale, yet many controls rely on inference-time interventions that add recurring compu

Riemannian Geometry for Pre-trained Language Model Embeddings

Model ReleasesDGX agent

arXiv:2607.07047v1 Announce Type: cross Abstract: Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety. We ask whether sentence-level

What's on My Network? Using Large Language Models to Identify Real-World IoT Devices at Scale

Model ReleasesDGX agent

arXiv:2510.13817v2 Announce Type: replace Abstract: The growth of IoT devices in shared environments has outpaced our ability to identify them, posing urgent risks to privacy, safety, and accountabili

8 Jul 2026

StepShield: When, Not Whether to Intervene on Rogue Agents

Model ReleasesDGX agent

arXiv:2601.22136v2 Announce Type: replace-cross Abstract: Agent safety benchmarks measure whether a monitor detects harm, not when. Yet timing is the difference between intervention and autopsy. We in

7 Jul 2026

20 questions for the Agentic Enterprise (and how Agent Platform can help)

Model ReleasesDGX agent

If you’re an IT leader, you might be getting a lot of questions about how to build and deploy agents. The pressure to move fast is intense, but the engineering reality is incredibly complex. Where do

AgentLTL: A Trace-Verification Framework for Measuring, Enforcing, and Training Procedural Compliance in Tool-Using LLM Agents

Model ReleasesDGX agent

arXiv:2607.02599v1 Announce Type: cross Abstract: Tool-using LLM agents are usually evaluated by final-answer correctness or LLM judges. Neither captures how an answer was produced. In safety-critical

Beyond Scene Priors: Fine-Grained Traffic Scene Reasoning with Benchmarking and Query-Guided Small-Object Focus

Model ReleasesDGX agent

arXiv:2607.04149v1 Announce Type: new Abstract: In safety-critical traffic scenarios, answering complex questions relies on minute, localized visual cues. However, standard Multimodal Large Language M

Federated Learning for Object Detection: Enabling Collaborative Drone Learning Without Centralizing Data

Local AiDGX agent

arXiv:2607.02636v1 Announce Type: cross Abstract: Object detection is a fundamental capability for AI-driven perception in safety-critical drone and edge-vision systems, including disaster response, o

From Arithmetic to Logic: The Resilience of Logic and Lookup-Based Neural Networks Under Parameter Bit-Flips

Model ReleasesDGX agent

arXiv:2603.22770v2 Announce Type: replace-cross Abstract: The deployment of deep neural networks (DNNs) in safety-critical edge environments necessitates robustness against hardware-induced bit-flip e

Hierarchical Multi-to-Single-Modal Knowledge Distillation for Disruption Prediction in EAST

Model ReleasesDGX agent

arXiv:2607.04241v1 Announce Type: cross Abstract: Plasma disruption is a critical threat to tokamak safety. Existing data-driven predictors mainly rely on time-series diagnostic signals, while visible

LLMs Encode Harmfulness and Refusal Separately

Model ReleasesDGX agent

arXiv:2507.11878v5 Announce Type: replace Abstract: LLMs are trained to refuse harmful instructions, but do they truly understand harmfulness beyond just refusing? Prior work has shown that LLMs' refu

NEST: Nascent Encoded Steganographic Thoughts

Model ReleasesDGX agent

arXiv:2602.14095v2 Announce Type: replace Abstract: Monitoring chain-of-thought (CoT) reasoning is a foundational safety technique for large language model agents; however, this oversight is compromis

Optimizing Large Language Models for Causality Assessment in Pharmacovigilance: Developing a Performance Metric as Objective for Bayesian Hyperparameter Optimization

Model ReleasesDGX agent

arXiv:2607.03704v1 Announce Type: new Abstract: Background: Growing individual case safety report (ICSR) volumes have intensified demand for scalable automated causality assessment. Large Language Mod

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents

Model ReleasesDGX agent

arXiv:2607.03968v1 Announce Type: cross Abstract: Large language models are increasingly deployed as IDE-integrated coding agents that decompose tasks, generate and edit files, run code, and refine ou

Robustness Meets Uncertainty: Evidential Adversarial Training for Robust Selective Classification

Model ReleasesDGX agent

arXiv:2607.03075v1 Announce Type: cross Abstract: Safety-critical applications require classifiers that are both robust and reliable. Adversarial training is a widely adopted defense for improving rob

3 Jul 2026

Fast Multi-dimensional Refusal Subspaces via RFM-AGOP

Model ReleasesDGX agent

arXiv:2607.02396v1 Announce Type: new Abstract: Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed behav

Locality-Aware Continual Unlearning for Diffusion Models

Model ReleasesDGX agent

arXiv:2512.02657v2 Announce Type: replace-cross Abstract: Real-world deployment of text-to-image diffusion models requires continual concept removal as new privacy, copyright, or safety obligations ar

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets

Model ReleasesDGX agent

arXiv:2607.02047v1 Announce Type: cross Abstract: Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolated prompts.

← Previous
1…219220221222223…240
Next →