AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
14 Apr 2026

AWARE: Adaptive Whole-body Active Rotating Control for Enhanced LiDAR-Inertial Odometry under Human-in-the-Loop Interaction

SafetyDGX agent

arXiv:2604.10598v1 Announce Type: new Abstract: Human-in-the-loop (HITL) UAV operation is essential in complex and safety-critical aerial surveying environments, where human operators provide navigati

Prompt Injection as Role Confusion

SafetyDGX agent

arXiv:2603.12277v3 Announce Type: replace-cross Abstract: Language models remain vulnerable to prompt injection attacks despite extensive safety training. We trace this failure to role confusion: mode

QShield: Securing Neural Networks Against Adversarial Attacks using Quantum Circuits

SafetyDGX agent

arXiv:2604.10933v1 Announce Type: cross Abstract: Deep neural networks remain highly vulnerable to adversarial perturbations, limiting their reliability in security- and safety-critical applications.

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering

SafetyDGX agent

arXiv:2508.11290v3 Announce Type: replace Abstract: LLMs increasingly exhibit over-refusal behavior, where safety mechanisms cause models to reject benign instructions that seemingly resemble harmful

13 Apr 2026

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs

SafetyDGX agent

arXiv:2604.08846v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have been shown to be vulnerable to malicious queries that can elicit unsafe responses. Recent work uses prom

Large Reasoning Models Learn Better Alignment from Flawed Thinking

SafetyDGX agent

arXiv:2510.00938v2 Announce Type: replace Abstract: Large reasoning models (LRMs) 'think' by generating structured chain-of-thought (CoT) before producing a final answer, yet they still lack the abili

10 Apr 2026

AEROS: A Single-Agent Operating Architecture with Embodied Capability Modules

SafetyDGX agent

arXiv:2604.07039v1 Announce Type: cross Abstract: Robotic systems lack a principled abstraction for organizing intelligence, capabilities, and execution in a unified manner. Existing approaches either

CMP: Robust Whole-Body Tracking for Loco-Manipulation via Competence Manifold Projection

SafetyDGX agent

arXiv:2604.07457v1 Announce Type: new Abstract: While decoupled control schemes for legged mobile manipulators have shown robustness, learning holistic whole-body control policies for tracking global

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs

SafetyDGX agent

arXiv:2604.07655v1 Announce Type: cross Abstract: Hard-gated safety checkers often over-refuse and misalign with a vendor's model spec; prevailing taxonomies also neglect robustness and honesty, yield

LLM-based Schema-Guided Extraction and Validation of Missing-Person Intelligence from Heterogeneous Data Sources

SafetyDGX agent

arXiv:2604.06571v1 Announce Type: cross Abstract: Missing-person and child-safety investigations rely on heterogeneous case documents, including structured forms, bulletin-style posters, and narrative

VLMShield: Efficient and Robust Defense of Vision-Language Models against Malicious Prompts

SafetyDGX agent

arXiv:2604.06502v1 Announce Type: new Abstract: Vision-Language Models (VLMs) face significant safety vulnerabilities from malicious prompt attacks due to weakened alignment during visual integration.

9 Apr 2026

Guardrails at the gateway: Securing AI inference on GKE with Model Armor

SafetyDGX agent

Enterprises are rapidly moving AI workloads from experimentation to production on Google Kubernetes Engine (GKE), using its scalability to serve powerful inference endpoints. However, as these models

8 Apr 2026

Chilling. The only thing I got wrong here in @politico was the year. This is exactly where we are now.

SafetyDGX agent

* **What it likely covers:** This entry likely discusses contemporary concerns regarding the safety and restriction of freedom of speech or civil liberties, drawing a parallel between past discus...

13 Aug 2026

A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization

SafetyDGX agent

arXiv:2608.11483v1 Announce Type: new Abstract: Hit-to-lead optimization requires iterative design of hit analogs across competing potency, selectivity, physicochemical, pharmacokinetic, safety, and s

BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model

SafetyDGX agent

arXiv:2608.11244v1 Announce Type: new Abstract: Construction standards are critical for building safety and sustainability. Existing standard application workflows rely on keyword-based document retri

Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers

SafetyDGX agent

arXiv:2608.12166v1 Announce Type: cross Abstract: Algorithm registers have been championed as a means of providing transparency on the use of algorithms in public services. Yet potential publics diffe

Energy-Aware Wind-Resilient Routing for Truck-Assisted Multi-UAV Delivery under Wind Uncertainty

SafetyDGX agent

arXiv:2608.11641v1 Announce Type: cross Abstract: Energy feasibility under wind uncertainty is a critical safety issue for low-altitude air-ground delivery. In truck-UAV systems, UAVs complete assigne

Inverse-dynamics observer design for a linear single-track vehicle model with distributed tire dynamics

SafetyDGX agent

arXiv:2603.07499v3 Announce Type: replace-cross Abstract: Accurate estimation of the vehicle's sideslip angle and tire forces is essential for enhancing safety and handling performances in unknown dri

Video2Track: From Real-World Interaction Videos to Steerable Adversarial Closed-Track Testing for Automated Driving Systems

SafetyDGX agent

arXiv:2608.11592v1 Announce Type: new Abstract: Closed-track testing plays a fundamental role in the verification and validation of automated driving systems (ADS), particularly for safety-critical sc

12 Aug 2026

CRHT: A Continuous Regression Hybrid Transformer for Vessel Trajectory Prediction with Online Cluster Sampling

SafetyDGX agent

arXiv:2608.10256v1 Announce Type: new Abstract: Accurate vessel trajectory prediction is critical for maritime safety and anomaly detection, yet existing models often struggle with geographic bias and

Introspective Attention Modulation for Safe Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2607.14945v2 Announce Type: replace Abstract: State-of-the-art flow based text-to-image (T2I) models exhibit remarkable generative abilities but remain vulnerable to producing unsafe content. Pr

SeFaR: Semantic Feature-aware Robustness Testing of Deep Neural Networks

SafetyDGX agent

arXiv:2608.10289v1 Announce Type: new Abstract: Deep neural networks are increasingly deployed in safety-critical domains as perception modules, where failures are often caused due to rare and under-r

11 Aug 2026

City Sentinel: A Unified AI-Based Smart Surveillance Framework for Real-Time Multi-Threat Detection Using Deep Learning

SafetyDGX agent

arXiv:2608.08887v1 Announce Type: new Abstract: Rapid urbanization has increased the need for surveillance systems that can monitor multiple public safety risks at the same time. Traditional systems o

FlowErase-OPD: Multi-Concept Erasure via Anchored On-Policy Distillation in Flow Matching Models

SafetyDGX agent

arXiv:2608.07620v1 Announce Type: new Abstract: Recent advances in flow matching models have substantially improved the quality of text-to-image generation, but have also raised increasing safety conc

Machine-Learning-Based Diagnostic Framework for Passive Ultrasonic Detection of Railway Wheel Defects

SafetyDGX agent

arXiv:2608.08301v1 Announce Type: new Abstract: Reliable identification of railway wheel defects is important for safety and maintenance. This study develops a machine-learning-based diagnostic framew

PAM: Training Policy-Aligned Moderation Filters at Scale

SafetyDGX agent

arXiv:2505.19766v4 Announce Type: replace Abstract: Large language models (LLMs) remain vulnerable to misalignment and jailbreaks, making external safeguards like moderation filters essential, yet exi

Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models

SafetyDGX agent

arXiv:2508.16406v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attacks, which attempt to elicit harmful responses from LLMs. The evolving nature

10 Aug 2026

KnifeHunter: Structured Local Representation Learning for Fine-Grained Knife Image Retrieval in Law Enforcement

SafetyDGX agent

arXiv:2608.07057v1 Announce Type: new Abstract: Knife-enabled violence presents a major public safety challenge, and law enforcement agencies require scalable tools for catalogue-level knife identific

Online Conformal Prediction Beyond Feedback

SafetyDGX agent

arXiv:2608.07139v1 Announce Type: new Abstract: Uncertainty quantification is essential when deploying machine learning models in safety-critical applications. Online conformal prediction (OCP) provid

7 Aug 2026

Deterministic World Models for Closed-loop Reachability Analysis of End-to-End Vision-based Control

SafetyDGX agent

arXiv:2512.08991v3 Announce Type: replace Abstract: End-to-end image controllers that map raw camera frames directly to control actions are increasingly deployed in safety-critical systems. However, f

Failing Gracefully: Mitigating Impact of Inevitable Robot Failures

SafetyDGX agent

arXiv:2608.05313v1 Announce Type: cross Abstract: Service robots operate in household environments shared with humans, pets, and everyday objects, where they are highly susceptible to failures such as

Symbol Grounding in Neuro-Symbolic AI: A Gentle Introduction to Reasoning Shortcuts

SafetyDGX agent

arXiv:2510.14538v3 Announce Type: replace Abstract: Neuro-symbolic (NeSy) AI aims to develop deep neural networks whose predictions comply with prior knowledge encoding, e.g. safety or structural cons

Worst-Case Distance-Aware Error Bounds for Neural Networks

SafetyDGX agent

arXiv:2510.22021v3 Announce Type: replace Abstract: Safety-critical applications of machine learning require uncertainty estimates that support reliable worst-case analysis. Neural networks (NNs) prov

6 Aug 2026

An Inline Control Architecture for Language Models in Intelligent Transportation Systems

SafetyDGX agent

arXiv:2608.04065v1 Announce Type: cross Abstract: Vehicle-to-everything (V2X) systems increasingly incorporate large language models (LLMs) for semantic tasks such as message summarization, operator a

Approximate Multi-Objective Search Under Rulebooks

SafetyDGX agent

arXiv:2608.04398v1 Announce Type: cross Abstract: Robotic planning often involves multiple objectives with complex priority relationships, such as safety, efficiency, and regulatory compliance. Rulebo

Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification

SafetyDGX agent

arXiv:2607.16868v2 Announce Type: replace Abstract: Large Language Models often produce confidently stated yet unreliable outputs, posing critical challenges for deployment in safety-sensitive applica

CheckOne: Lightweight Fault Detection and Mitigation for Vision Transformers

SafetyDGX agent

arXiv:2608.04035v1 Announce Type: cross Abstract: The wide adoption of Vision Transformers (ViTs) in safety-critical applications raises reliability concerns related to hardware faults. Algorithm-Base

muSync-GS: Physics-Synchronized Driving Video Synthesis for Weather and Geometric Road Hazards

SafetyDGX agent

arXiv:2608.04412v1 Announce Type: new Abstract: High-quality driving data are essential for autonomous-driving systems and generative world models. However, rare and safety-critical scenarios involvin

5 Aug 2026

Double Down on Defense: Strengthening Deep Perceptual Hashes against Evasion Attacks without Retraining

SafetyDGX agent

arXiv:2608.03101v1 Announce Type: new Abstract: Near-duplicate image matching is crucial for trust and safety, provenance verification, copyright enforcement, and large-scale visual search. Modern pla

Evading Chain-of-Thought Monitoring Through Model Poisoning

SafetyDGX agent

arXiv:2608.02820v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring is an increasingly important component of AI safety stacks but relies on the assumption that a model's reasoning tra

Just had to create an 'accidental-cyberattacks' tag on my blog We're up to four now: the original OpenAI+Hugging Face one, Anthropic's me-to…

SafetyDGX agent

Just had to create an 'accidental-cyberattacks' tag on my blog We're up to four now: the original OpenAI+Hugging Face one, Anthropic's me-too attacks, then two new ones from the UK AI Safety Institute

Optimising for Flourishing: Flourishing Metrics and Return on Flourishing as Success Criteria for Artificial Intelligence and Post-AGI Economic Systems

SafetyDGX agent

arXiv:2608.00151v2 Announce Type: replace-cross Abstract: Current evaluation frameworks for artificial intelligence focus mainly on capability, safety, and proxies such as adoption, engagement, effici

Rethinking Uncertainty Quantification and Entanglement in Image Segmentation

SafetyDGX agent

arXiv:2603.18792v2 Announce Type: replace Abstract: Uncertainty quantification (UQ) is crucial in safety-critical applications such as medical image segmentation. Total uncertainty is typically decomp

Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI…

SafetyDGX agent

Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI model framework. I'm not surprised at all, and it's actuall

Staying on Spec: Real-Time Monitoring under Uncertainty with a Maritime Case Study

SafetyDGX agent

arXiv:2608.02811v1 Announce Type: new Abstract: Robotic systems must operate under uncertainty while satisfying complex task and safety specifications. Monitoring such specifications under uncertainty

4 Aug 2026

Announcing Cloudflare Wallets: the programmable wallet for the agentic Internet

SafetyDGX agent

Cloudflare Wallets will provide AI agents with native payments and verifiable identity on the web. Using the x402 protocol, agents can autonomously purchase APIs and content within clear safety guardr

CoLI: A Reproducible Platform for Continuum Robot Learning via Monolithic 3D Printing and Isomorphic Teleoperation

SafetyDGX agent

arXiv:2606.20389v2 Announce Type: replace Abstract: Continuum robots offer strong potential for manipulation tasks due to their high degrees of freedom, compliant structures, and operational safety. H

EnvShip: A Unified Framework for Context-Aware and Cross-Region Vessel Trajectory Forecasting

SafetyDGX agent

arXiv:2606.15240v2 Announce Type: replace Abstract: Accurate vessel trajectory forecasting is essential for maritime situational awareness, navigation safety, traffic management, and autonomous naviga

Mind the Gap: Zero-Query Jailbreaks via Filter-Generator Discrepancy in Text-to-Image Systems

SafetyDGX agent

arXiv:2608.00973v1 Announce Type: new Abstract: Text-to-image (T2I) systems typically have prompt-level safety filters before the generator to block unsafe requests, yet such systems remain vulnerable

On the Limits of Support-Preserving Alignment and Bounded Filtering

SafetyDGX agent

arXiv:2607.18295v2 Announce Type: replace Abstract: We study whether alignment schemes that reshape a base model's output distribution, combined with bounded safety filters, can drive the probability

Proteus: A Truncation-Robust Entropy Model for Progressive LiDAR Compression

SafetyDGX agent

arXiv:2608.00687v1 Announce Type: new Abstract: LiDAR point clouds provide explicit, deterministic physical boundaries critical for collaborative safety-critical perception. However, wireless channels

Provably Safe Generative Sampling with Constricting Barrier Functions

SafetyDGX agent

arXiv:2602.21429v3 Announce Type: replace Abstract: Flow-based generative models, such as diffusion models and flow matching models, have achieved remarkable success in learning complex data distribut

SPIRIT: Spatio-temporal Pairwise Relational Modeling of Instrument-Tissue Interactions for Surgical Action Triplet Recognition

SafetyDGX agent

arXiv:2608.02188v1 Announce Type: new Abstract: Fine-grained understanding of surgical activity is essential for context-aware assistance in the operating room, including safety monitoring, adverse ev

Today, we’re launching Alpamayo 2 Super, our frontier open reasoning model for autonomous vehicles. Beyond seeing, Alpamayo understands and …

SafetyDGX agent

Today, we’re launching Alpamayo 2 Super, our frontier open reasoning model for autonomous vehicles. Beyond seeing, Alpamayo understands and reasons through the complex world - thinks before it acts. I

3 Aug 2026

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

SafetyDGX agent

arXiv:2607.28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the

Safe Vision Language Action Models via Barrier Enhanced Flow Matching

Model ReleasesDGX agent

arXiv:2607.29569v1 Announce Type: new Abstract: This article presents a modular inference framework that integrates Flow Matching generative models with formal Control Barrier Function (CBF) safety gu

Sensitivity Analysis of GRU, LSTM and Transformer Encoder in Classification of Automated Driving Systems

SafetyDGX agent

arXiv:2607.28665v1 Announce Type: cross Abstract: Automated driving systems (ADSs) are becoming ubiquitous. Future Software Defined Vehicles (SDVs) may be able to run multiple ADSs, both native and af

TransGraspNet: Physically and Geometrically Consistent Manipulation of Transparent Labware

SafetyDGX agent

arXiv:2607.29567v1 Announce Type: new Abstract: Manipulating transparent laboratory glassware that contains liquid is inherently safety-critical: even small geometric errors can cause unstable grasps

31 Jul 2026

Context-Informed Ship Trajectory Prediction via Conditional Attention

SafetyDGX agent

arXiv:2607.27418v1 Announce Type: new Abstract: Long-term ship trajectory prediction is a fundamental capability for maritime safety and autonomous navigation. While recent Transformer-based architect

Hierarchical Multilevel Monte Carlo for Order-Optimal Neural Actor-Critic in Average-Reward CMDPs

SafetyDGX agent

arXiv:2607.28390v1 Announce Type: new Abstract: Constrained Markov Decision Processes (CMDPs) provide a natural framework for reinforcement learning in safety-critical applications, where agents maxim

← Previous
1…1819202122…240
Next →