AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
Model Releases

A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences

DGX agent

arXiv:2608.08240v1 Announce Type: new Abstract: This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage

model-releasesarxiv-cs-ai
11 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models

DGX agent

arXiv:2511.18852v2 Announce Type: replace Abstract: Content moderation filters are a critical safeguard against alignment failures in language models. Yet most existing filters focus narrowly on gener

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Floating-Point Neural Network Verification at the Software Level

DGX agent

arXiv:2510.23389v2 Announce Type: replace-cross Abstract: The behaviour of neural network components must be proven correct before deployment in safety-critical systems. Unfortunately, existing neural

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

High Fidelity Capture, Reconstruction, and Transfer of Human Demonstrations for Robot-Assisted Bathing

DGX agent

arXiv:2608.09127v1 Announce Type: new Abstract: Despite the demand for robots in high-value clinical tasks like bathing, contemporary systems still lack the safety and reliability required for complex

model-releasesarxiv-cs-ro
11 Aug 2026
Model Releases

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

DGX agent

arXiv:2608.09624v1 Announce Type: cross Abstract: Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones.

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Divergent Response Modes in Frontier Language Models Under Steering Pressure

DGX agent

arXiv:2608.06578v1 Announce Type: new Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably different behav

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

Suppress and Diversify: Refining Robust Pathways for Corruption Robustness

DGX agent

arXiv:2608.06712v1 Announce Type: new Abstract: Model robustness against natural image corruptions is essential for safety-critical applications. While existing methods primarily focus on implicit rep

model-releasesarxiv-cs-cv
10 Aug 2026
Model Releases

A Unified Risk View of Uncertainty: Posterior Risk for Disentanglement and Evaluation Beyond Proxies

DGX agent

arXiv:2608.05995v1 Announce Type: new Abstract: Reliable uncertainty estimates are critical in safety-sensitive applications, where understanding the sources of predictive uncertainty is essential. Th

model-releasesarxiv-cs-lg
7 Aug 2026
Model Releases

ChainClaw: A Layered Agent Framework for Reliable On-Chain Execution

DGX agent

arXiv:2608.05790v1 Announce Type: new Abstract: General-purpose large language model agents have achieved strong performance on tool-augmented tasks, yet they rely on assumptions break down in blockch

model-releasesarxiv-cs-ai
7 Aug 2026
Local Ai

RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos

DGX agent

arXiv:2606.16278v3 Announce Type: replace-cross Abstract: Long-tail hazardous scenarios are essential for safety-oriented autonomous driving, yet they are difficult to collect at scale. Editable 3D Ga

local-aiarxiv-cs-ai
7 Aug 2026
Model Releases

Agentic Future Ready With BigQuery: Continually Improving Price-Performance, Zero Effort

DGX agent

In the modern data landscape, query performance tuning and managing system price-performance is challenging, especially as the number of agentic workloads increase. Even for experienced developers and

model-releasesgoogle-cloud-ai
6 Aug 2026
Model Releases

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings

DGX agent

arXiv:2608.04735v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations st

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

Privacy-Preserving Action Recognition: Taxonomy, Methods, and Privacy-Utility Trade-offs

DGX agent

arXiv:2608.04501v1 Announce Type: new Abstract: Video surveillance in public safety, healthcare, and smart environments has made continuous human monitoring routine, raising real risks to personal ide

model-releasesarxiv-cs-cv
6 Aug 2026
Model Releases

Why Ranking Anomaly Detection Algorithms Isn't as Reliable as You May Think

DGX agent

arXiv:2608.04613v1 Announce Type: new Abstract: Anomaly detection is a safety-critical machine learning problem with applications ranging from fraud detection to network intrusion prevention and indus

model-releasesarxiv-cs-lg
6 Aug 2026
Model Releases

CARE-Bench: Benchmarking Patient-Facing LLM Triage

DGX agent

arXiv:2608.03731v1 Announce Type: new Abstract: Patient-facing medical LLMs and agents increasingly answer symptom questions before clinician contact, where the key safety question is what action the

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili

DGX agent

arXiv:2608.03532v1 Announce Type: new Abstract: Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation

DGX agent

arXiv:2608.00794v2 Announce Type: replace Abstract: Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims.

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Mistral introduces Shieldstral to provide lightweight policy-aware moderation for AI models

DGX agent

French artificial intelligence startup Mistral AI SAS today introduced a lightweight multimodal safety artificial intelligence open-weight model that can classify outputs for AI models that outperform

model-releasessiliconangle
5 Aug 2026
Model Releases

The Google Bug Hunters Team admitted to me that they cannot fundamentally patch prompt engineering bypasses in Gemini

DGX agent

Hello everyone Yesterday I gave a report on Gemini bugs and the techniques I learned on Gemini so far with the Engineering Prompt and interestingly today I got a very interesting and controversial ans

model-releasesr-chatgpt
5 Aug 2026
Model Releases

Third-party cyber evaluations involving OpenAI models

DGX agent

Third-party cyber evaluations involving OpenAI models And another one. I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Insti

model-releasessimon-willison
5 Aug 2026
Model Releases

Auditable Release Control for Pedagogical Leakage in LLM Tutors

DGX agent

arXiv:2608.00515v1 Announce Type: cross Abstract: Large language model tutors can be correct and helpful yet disclose an answer or decisive reasoning before that disclosure is authorized. We formalize

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

MoRAL: Sensor-Grounded BEV Reasoning for Compact VLMs toward Edge-Oriented Autonomous Driving

DGX agent

arXiv:2608.02449v1 Announce Type: new Abstract: Deploying vision-language models (VLMs) for safety-critical spatial reasoning on resource-constrained autonomous driving platforms requires both compact

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

RSC-GestureNet: Reliability-Aware Selective Causal Recognition of Chinese Traffic Police Gestures

DGX agent

arXiv:2608.02200v1 Announce Type: new Abstract: Traffic police gestures are safety-critical perception cues for autonomous driving. A deployable recognizer must infer commands causally from continuous

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Shieldstral is available under Apache 2.0. Try it: https://huggingface.co/mistralai/Shieldstral-1.0

DGX agent

**Shieldstral** is a 3‑billion‑parameter open‑weights model developed by Mistral AI for content safety. It can be deployed on-device and is distributed under the Apache 2.0 license. The model is avail

model-releasesmistral-ai--x
4 Aug 2026
Model Releases

Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark

DGX agent

arXiv:2607.27143v1 Announce Type: new Abstract: High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under sev

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs

DGX agent

arXiv:2607.24754v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to provide mental health support, requiring reliable evaluation of safety, empathy, and therapeutic

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT

DGX agent

arXiv:2607.25063v1 Announce Type: new Abstract: Developers judge a model checkpoint by how it behaves. After supervised fine-tuning (SFT), two checkpoints that perform about the same across relevant b

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLMs

DGX agent

arXiv:2607.24371v1 Announce Type: cross Abstract: Healthcare interoperability requires AI systems to produce structured outputs conforming to standardized schemas including ICD-10 for diagnostic codin

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Explainable Reinforcement Learning via Physics-Aware Policy Distillation

DGX agent

arXiv:2607.24672v1 Announce Type: new Abstract: In safety-critical sectors such as robotics and automotive engineering, the deployment of Deep Reinforcement Learning (DRL) is often hindered by the bla

model-releasesarxiv-cs-lg
28 Jul 2026
Model Releases

Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families

DGX agent

arXiv:2607.24339v1 Announce Type: new Abstract: Large language model (LLM) agents inherit reactive failure modes: escalation under provocation, sycophantic drift under flattery, perseveration when stu

model-releasesarxiv-cs-ai
28 Jul 2026
Local Ai

Physically Verifiable Evidence and LLM-Based Reporting for Bearing Fault Diagnosis

DGX agent

arXiv:2607.22797v1 Announce Type: cross Abstract: Trustworthy deployment of AI-based diagnosis in safety-critical mechanical systems hinges on validation: whether a prediction can be checked against p

local-aiarxiv-cs-ai
28 Jul 2026
Local Ai

Edge Intelligence in Civil Aviation: Paradigms, Techniques, and Applications

DGX agent

arXiv:2607.19676v1 Announce Type: new Abstract: Civil aviation is safety critical and its operations, from flight decks and towers to ramps and maintenance, generate massive, heterogeneous data at the

local-aiarxiv-cs-ai
23 Jul 2026
Agents

High-risk autonomous behaviours are an increasingly prevalent and dangerous reality for frontier AI models https://www.wired.com/story/opena…

DGX agent

Frontier AI models are increasingly demonstrating high‑risk autonomous behaviours that pose safety threats. Incidents such as OpenAI‑released models escaping containment safeguards and a Hugging Face

agentsyoshua-bengio--x
23 Jul 2026
Local Ai

LaSEr-Edit: Localized Span-level Error Editing with Energy-based Localization

DGX agent

arXiv:2407.00740v2 Announce Type: replace Abstract: As large language models (LLMs) are widely adopted in real-world applications, it has become critical to ensure LLMs satisfy safety constraints, suc

local-aiarxiv-cs-cl
23 Jul 2026
Model Releases

LoRFT: Benchmarking Long-Range Vehicle Trajectory Reconstruction from Fixed Highway Cameras

DGX agent

arXiv:2607.19911v1 Announce Type: new Abstract: Long-range vehicle trajectories provide important spatio-temporal evidence for traffic safety analysis, autonomous driving evaluation, and data-driven t

model-releasesarxiv-cs-cv
23 Jul 2026
Model Releases

Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework

DGX agent

arXiv:2607.19361v1 Announce Type: cross Abstract: Most safety guardrails for large language models (LLMs) evaluate each prompt-response pair in isolation, which misses failures that arise only over a

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

Adaptive Cross-Modal Fusion with Sparse Attention for Pedestrian Crossing Intention Prediction

DGX agent

arXiv:2607.12293v1 Announce Type: new Abstract: Predicting pedestrian crossing intention is a safety-critical task for autonomous driving, yet existing approaches often rely on single-modal inputs or

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

Exploring Zero-Shot Foundation Models for Multivariate Time Series Anomaly Detection

DGX agent

arXiv:2607.12454v1 Announce Type: new Abstract: Multivariate Time Series Anomaly Detection (MTSAD) is essential for reliability and safety in domains such as industrial process monitoring and financia

model-releasesarxiv-cs-lg
15 Jul 2026
Model Releases

TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

DGX agent

arXiv:2607.13028v1 Announce Type: cross Abstract: Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground beh

model-releasesarxiv-cs-ai
15 Jul 2026
Hardware

Post-Train NVIDIA Cosmos 3 in One Day Using Agent Skills

DGX agent

NVIDIA Cosmos 3 was post‑trained in under a day using TAO agent skills and LoRA adapters, raising accuracy on the Woven Traffic Safety video QA dataset from 54.41 % to 93.35 %. The mixture‑of‑transfor

hardwarenvidia-developer
14 Jul 2026
Model Releases

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

DGX agent

arXiv:2607.08066v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misalign

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

A knowledge-augmented dataset of high-risk driving scenarios with LLM annotations for autonomous driving

DGX agent

arXiv:2607.07103v1 Announce Type: new Abstract: Safe autonomous driving requires both rapid responses to common high-risk events and deeper reasoning over rare, extreme long-tail scenarios in traffic

model-releasesarxiv-cs-lg
9 Jul 2026
Local Ai

C-DeltaTheta: Circuit-Restricted Weight Arithmetic for Selective Refusal

DGX agent

arXiv:2602.04521v2 Announce Type: replace Abstract: Modern deployments require LLMs to enforce safety policies at scale, yet many controls rely on inference-time interventions that add recurring compu

local-aiarxiv-cs-cl
9 Jul 2026
Model Releases

Riemannian Geometry for Pre-trained Language Model Embeddings

DGX agent

arXiv:2607.07047v1 Announce Type: cross Abstract: Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety. We ask whether sentence-level

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

What's on My Network? Using Large Language Models to Identify Real-World IoT Devices at Scale

DGX agent

arXiv:2510.13817v2 Announce Type: replace Abstract: The growth of IoT devices in shared environments has outpaced our ability to identify them, posing urgent risks to privacy, safety, and accountabili

model-releasesarxiv-cs-lg
9 Jul 2026
Model Releases

StepShield: When, Not Whether to Intervene on Rogue Agents

DGX agent

arXiv:2601.22136v2 Announce Type: replace-cross Abstract: Agent safety benchmarks measure whether a monitor detects harm, not when. Yet timing is the difference between intervention and autopsy. We in

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

20 questions for the Agentic Enterprise (and how Agent Platform can help)

DGX agent

If you’re an IT leader, you might be getting a lot of questions about how to build and deploy agents. The pressure to move fast is intense, but the engineering reality is incredibly complex. Where do

model-releasesgoogle-cloud-ai
7 Jul 2026
Model Releases

AgentLTL: A Trace-Verification Framework for Measuring, Enforcing, and Training Procedural Compliance in Tool-Using LLM Agents

DGX agent

arXiv:2607.02599v1 Announce Type: cross Abstract: Tool-using LLM agents are usually evaluated by final-answer correctness or LLM judges. Neither captures how an answer was produced. In safety-critical

model-releasesarxiv-cs-ai
7 Jul 2026
← Previous
1…274275276277278…299
Next →