AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,202 results
Model Releases

Finite-Sample Analysis of Elimination in Active Hypothesis Testing

DGX agent

arXiv:2605.01039v1 Announce Type: new Abstract: A fixed-confidence, finite-sample problem of active hypothesis testing arises in many safety-critical applications. Situated in the context of sequentia

model-releasesarxiv-cs-lg
5 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models

DGX agent

arXiv:2605.00123v1 Announce Type: new Abstract: Safety trained large language models (LLMs) can often be induced to answer harmful requests through jailbreak prompts. Because we lack a robust understa

model-releasesarxiv-cs-ai
5 May 2026
Model Releases

Understanding Emergent Misalignment via Feature Superposition Geometry

DGX agent

arXiv:2605.00842v1 Announce Type: cross Abstract: Emergent misalignment, where fine-tuning on narrow, non-harmful tasks induces harmful behaviors, poses a key challenge for AI safety in LLMs. Despite

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

WILD SAM: A Simulated-and-Real Data Augmentation for Autonomous Driving Perception under Challenging Weather

DGX agent

arXiv:2605.01081v1 Announce Type: new Abstract: The performance of state-of-the-art object detectors degrades significantly under adverse weather, causing a safety-critical domain shift problem for au

model-releasesarxiv-cs-cv
5 May 2026
Local Ai

OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving

DGX agent

arXiv:2512.14044v3 Announce Type: replace-cross Abstract: The deployment of Vision-Language Models (VLMs) in safety-critical domains like autonomous driving (AD) is critically hindered by reliability

local-aiarxiv-cs-ai
1 May 2026
Model Releases

Lyapunov-Guided Self-Alignment: Test-Time Adaptation for Offline Safe Reinforcement Learning

DGX agent

arXiv:2604.26516v1 Announce Type: cross Abstract: Offline reinforcement learning (RL) agents often fail when deployed, as the gap between training datasets and real environments leads to unsafe behavi

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

Below-Chance Blindness: Prompted Underperformance in Small LLMs Produces Positional Bias Rather than Answer Avoidance

DGX agent

arXiv:2604.25249v1 Announce Type: new Abstract: Detecting sandbagging--the deliberate underperformance on capability evaluations--is an open problem in AI safety. We tested whether symptom validity te

model-releasesarxiv-cs-cl
29 Apr 2026
Model Releases

CAN-QA: A Question-Answering Benchmark for Reasoning over In-Vehicle CAN Traffic

DGX agent

arXiv:2604.24935v1 Announce Type: cross Abstract: The Controller Area Network (CAN) is a safety-critical in-vehicle communication protocol that lacks built-in security mechanisms, making intrusion det

model-releasesarxiv-cs-lg
29 Apr 2026
Model Releases

Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver

DGX agent

arXiv:2604.25067v1 Announce Type: cross Abstract: Forecasting when AI systems will become capable of meaningfully accelerating AI research is a central challenge for AI safety. Existing benchmarks mea

model-releasesarxiv-cs-lg
29 Apr 2026
Model Releases

GCP: Guarded Collaborative Perception with Spatial-Temporal Aware Malicious Agent Detection

DGX agent

arXiv:2501.02450v2 Announce Type: replace Abstract: Collaborative perception significantly enhances autonomous driving safety by extending each vehicle's perception range through message sharing among

model-releasesarxiv-cs-cv
28 Apr 2026
Model Releases

Learning Under Low Illumination: A Dataset and Algorithm for Traffic Sign Recognition

DGX agent

arXiv:2511.17183v2 Announce Type: replace Abstract: Traffic signboards are vital for road safety and intelligent transportation systems, enabling navigation and autonomous driving. Yet, recognizing tr

model-releasesarxiv-cs-cv
28 Apr 2026
Model Releases

Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings

DGX agent

arXiv:2604.23130v1 Announce Type: cross Abstract: Large language models (LLMs) can still be jailbroken into producing harmful outputs despite safety alignment. Existing attacks show this vulnerability

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications

DGX agent

arXiv:2604.24668v1 Announce Type: new Abstract: Given the increased use of LLMs in financial systems today, it becomes important to evaluate the safety and robustness of such systems. One failure mode

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

TRACE: Topology-aware Reconstruction of Accidents in CARLA for AV Evaluation

DGX agent

arXiv:2604.22068v1 Announce Type: cross Abstract: Validating Autonomous Vehicles (AVs) requires exposure to rare, safety-critical scenarios, infrequent in routine driving data. Existing benchmarks add

model-releasesarxiv-cs-ro
27 Apr 2026
Model Releases

Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers

DGX agent

arXiv:2604.21700v1 Announce Type: cross Abstract: The growing application of large language models (LLMs) in safety-critical domains has raised urgent concerns about their security. Many recent studie

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models

DGX agent

arXiv:2604.21860v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly integrated into sensitive workflows, raising the stakes for adversarial robustness and safety. This pape

model-releasesarxiv-cs-ai
24 Apr 2026
Local Ai

Ufil: A Unified Framework for Infrastructure-based Localization

DGX agent

arXiv:2604.21471v1 Announce Type: new Abstract: Infrastructure-based localization enhances road safety and traffic management by providing state estimates of road users. Development is hindered by fra

local-aiarxiv-cs-ro
24 Apr 2026
Model Releases

When to Trust the Answer: Question-Aligned Semantic Nearest Neighbor Entropy for Safer Surgical VQA

DGX agent

arXiv:2511.01458v2 Announce Type: replace-cross Abstract: Safety and reliability are critical for deploying visual question answering (VQA) systems in surgery, where incorrect or ambiguous responses c

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

Automated Detection of Dosing Errors in Clinical Trial Narratives: A Multi-Modal Feature Engineering Approach with LightGBM

DGX agent

arXiv:2604.19759v1 Announce Type: new Abstract: Clinical trials require strict adherence to medication protocols, yet dosing errors remain a persistent challenge affecting patient safety and trial int

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control

DGX agent

arXiv:2601.02896v2 Announce Type: replace Abstract: Controlling emergent behavioral personas (e.g., sycophancy, hallucination) in Large Language Models (LLMs) is critical for AI safety, yet remains a

model-releasesarxiv-cs-lg
23 Apr 2026
Model Releases

CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs

DGX agent

arXiv:2604.20460v1 Announce Type: new Abstract: Safety-critical traffic reasoning requires contrastive consistency: models must detect true hazards when an accident occurs, and reliably reject plausib

model-releasesarxiv-cs-cv
23 Apr 2026
Model Releases

Evaluating the Quality of the Quantified Uncertainty for (Re)Calibration of Data-Driven Regression Models

DGX agent

arXiv:2508.17761v3 Announce Type: replace Abstract: In safety-critical applications data-driven models must not only be accurate but also provide reliable uncertainty estimates. This property, commonl

model-releasesarxiv-cs-lg
23 Apr 2026
Model Releases

Model Capability Assessment and Safeguards for Biological Weaponization

DGX agent

arXiv:2604.19811v1 Announce Type: cross Abstract: AI leaders and safety reports increasingly warn that advances in model reasoning may enable biological misuse, including by low-expertise users, while

model-releasesarxiv-cs-ai
23 Apr 2026
Model Releases

PipeMFL-240K: A Large-scale Dataset and Benchmark for Object Detection in Pipeline Magnetic Flux Leakage Imaging

DGX agent

arXiv:2602.07044v2 Announce Type: replace-cross Abstract: Pipeline integrity is critical to industrial safety and environmental protection, with Magnetic Flux Leakage (MFL) detection being a primary n

model-releasesarxiv-cs-ai
23 Apr 2026
Local Ai

A Heterogeneous Long-Micro Scale Cascading Architecture for General Aviation Health Management

DGX agent

arXiv:2603.22885v4 Announce Type: replace Abstract: BACKGROUND: General aviation fleet expansion demands intelligent health monitoring under computational constraints. Real-world aircraft health diagn

local-aiarxiv-cs-lg
22 Apr 2026
Local Ai

HardNet++: Nonlinear Constraint Enforcement in Neural Networks

DGX agent

arXiv:2604.19669v1 Announce Type: new Abstract: Enforcing constraint satisfaction in neural network outputs is critical for safety, reliability, and physical fidelity in many control and decision-maki

local-aiarxiv-cs-lg
22 Apr 2026
Model Releases

Harness as an Asset: Enforcing Determinism via the Convergent AI Agent Framework (CAAF)

DGX agent

arXiv:2604.17025v1 Announce Type: cross Abstract: Large Language Models (LLMs) produce a controllability gap in safety-critical engineering: even low rates of undetected constraint violations render a

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Learning to Retrieve User History and Generate User Profiles for Personalized Persuasiveness Prediction

DGX agent

arXiv:2601.05654v3 Announce Type: replace Abstract: Estimating the persuasiveness of messages is critical in various applications, from recommender systems to safety assessment of LLMs. While it is im

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Screen Before You Interpret: A Portable Validity Protocol for Benchmark-Based LLM Confidence Signals

DGX agent

arXiv:2604.17714v1 Announce Type: new Abstract: LLM confidence signals are used for abstention, routing, and safety-critical decisions. No standard practice exists for checking whether a confidence si

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Automating Crash Diagram Generation Using Vision-Language Models: A Case Study on Multi-Lane Roundabouts

DGX agent

arXiv:2604.15332v1 Announce Type: cross Abstract: Crash diagrams are essential tools in transportation safety analysis, yet their manual preparation remains time-consuming and prone to human variabili

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

LinuxArena: A Control Setting for AI Agents in Live Production Software Environments

DGX agent

arXiv:2604.15384v1 Announce Type: cross Abstract: We introduce LinuxArena, a control setting in which agents operate directly on live, multi-service production environments. LinuxArena contains 20 env

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs

DGX agent

arXiv:2604.15780v1 Announce Type: cross Abstract: Machine learning models are increasingly deployed in real-world applications, but even aligned models such as Mistral and LLaVA still exhibit unsafe b

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing

DGX agent

arXiv:2604.15725v1 Announce Type: cross Abstract: Large Reasoning Models (LRMs) have demonstrated strong capabilities in generating step-by-step reasoning chains alongside final answers, enabling thei

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models

DGX agent

arXiv:2601.03699v2 Announce Type: replace Abstract: As large language models (LLMs) become integral to safety-critical applications, ensuring their robustness against adversarial prompts is paramount.

model-releasesarxiv-cs-cl
20 Apr 2026
Local Ai

MS-SSE-Net: A Multi-Scale Spatial Squeeze-and-Excitation Network for Structural Damage Detection in Civil and Geotechnical Engineering

DGX agent

arXiv:2604.14711v1 Announce Type: new Abstract: Structural damage detection is essential for maintaining the safety and reliability of civil infrastructure. However, accurately identifying different t

local-aiarxiv-cs-cv
17 Apr 2026
Model Releases

DeEscalWild: A Real-World Benchmark for Automated De-Escalation Training with SLMs

DGX agent

arXiv:2604.13075v1 Announce Type: new Abstract: Effective de-escalation is critical for law enforcement safety and community trust, yet traditional training methods lack scalability and realism. While

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

Dental-TriageBench: Benchmarking Multimodal Reasoning for Hierarchical Dental Triage

DGX agent

arXiv:2604.13060v1 Announce Type: new Abstract: Dental triage is a safety-critical clinical routing task that requires integrating multimodal clinical information (e.g., patient complaints and radiogr

model-releasesarxiv-cs-cl
16 Apr 2026
Model Releases

FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content Moderation

DGX agent

arXiv:2602.23636v3 Announce Type: replace Abstract: Ensuring the safety of LLM-generated content is essential for real-world deployment. Most existing guardrail models formulate moderation as a fixed

model-releasesarxiv-cs-lg
16 Apr 2026
Model Releases

GF-Score: Certified Class-Conditional Robustness Evaluation with Fairness Guarantees

DGX agent

arXiv:2604.12757v1 Announce Type: cross Abstract: Adversarial robustness is essential for deploying neural networks in safety-critical applications, yet standard evaluation methods either require expe

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

Monte Carlo Stochastic Depth for Uncertainty Estimation in Deep Learning

DGX agent

arXiv:2604.12719v1 Announce Type: new Abstract: The deployment of deep neural networks in safety-critical systems necessitates reliable and efficient uncertainty quantification (UQ). A practical and w

model-releasesarxiv-cs-lg
15 Apr 2026
Model Releases

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models

DGX agent

arXiv:2604.12371v1 Announce Type: new Abstract: We study typographic prompt injection attacks on vision-language models (VLMs), where adversarial text is rendered as images to bypass safety mechanisms

model-releasesarxiv-cs-cv
15 Apr 2026
Model Releases

Doc-PP: Document Policy Preservation Benchmark for Large Vision-Language Models

DGX agent

arXiv:2601.03926v2 Announce Type: replace Abstract: The deployment of Large Vision-Language Models (LVLMs) for real-world document question answering is often constrained by dynamic, user-defined poli

model-releasesarxiv-cs-cl
14 Apr 2026
Model Releases

FedKLPR: KL-Guided Pruning-Aware Federated Learning for Person Re-Identification

DGX agent

arXiv:2508.17431v3 Announce Type: replace-cross Abstract: Person re-identification (re-ID) is a fundamental task in intelligent surveillance and public safety. Federated learning (FL) provides a priva

model-releasesarxiv-cs-ai
14 Apr 2026
Local Ai

Global monitoring of methane point sources using deep learning on hyperspectral radiance measurements from EMIT

DGX agent

arXiv:2604.10094v1 Announce Type: new Abstract: Anthropogenic methane (CH4) point sources drive near-term climate forcing, safety hazards, and system inefficiencies. Space-based imaging spectroscopy i

local-aiarxiv-cs-cv
14 Apr 2026
Model Releases

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models

DGX agent

arXiv:2604.10866v1 Announce Type: new Abstract: AI agents are expected to perform professional work across hundreds of occupational domains (from emergency department triage to nuclear reactor safety

model-releasesarxiv-cs-cl
14 Apr 2026
Model Releases

Quantitative Introspection in Language Models: Tracking Emotive States Across Conversation

DGX agent

arXiv:2603.18893v2 Announce Type: replace Abstract: Tracking the internal states of large language models across conversations is important for safety, interpretability, and model welfare, yet current

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

SimScale: Learning to Drive via Real-World Simulation at Scale

DGX agent

arXiv:2511.23369v3 Announce Type: replace Abstract: Achieving fully autonomous driving systems requires learning rational decisions in a wide span of scenarios, including safety-critical and out-of-di

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

VOLTA: The Surprising Ineffectiveness of Auxiliary Losses for Calibrated Deep Learning

DGX agent

arXiv:2604.08639v1 Announce Type: cross Abstract: Uncertainty quantification (UQ) is essential for deploying deep learning models in safety critical applications, yet no consensus exists on which UQ m

model-releasesarxiv-cs-ai
13 Apr 2026
← Previous
1…237238239240241…255
Next →