AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
20 Jul 2026

SSI should open-source an opus-tier model. 1. it nearly closes the us/china gap on the os frontier 2. they don't want to waste compute alloc…

SafetyDGX agent

SSI should open-source an opus-tier model. 1. it nearly closes the us/china gap on the os frontier 2. they don't want to waste compute allocation on inference, releasing os will enable continued full

16 Jul 2026

A Comparative Evaluation of Large Vision-Language Models for 2D Object Detection under SOTIF Conditions

Model ReleasesDGX agent

arXiv:2601.22830v2 Announce Type: replace Abstract: Reliable environmental perception remains one of the main obstacles for safe operation of automated vehicles. Safety of the Intended Functionality (

Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2607.13036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolving. When the

Distributionally Robust and Safe Imitation Learning

SafetyDGX agent

arXiv:2607.13436v1 Announce Type: new Abstract: Imitation learning (IL) has achieved remarkable success in complex decision-making tasks. However, its performance is highly sensitive to distribution s

Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake Education

SafetyDGX agent

arXiv:2607.14046v1 Announce Type: new Abstract: This paper presents Earthquaker-AI, a hybrid educational framework building upon a previously implemented educational robotics project by integrating a

Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants

Model ReleasesDGX agent

arXiv:2607.13039v1 Announce Type: cross Abstract: Safety evaluations for dual-use biology assistants often measure base-model capability, refusal behavior, or jailbreak success. These metrics miss a d

15 Jul 2026

A Neurosymbolic Approach to Natural Language Formalization and Verification

SafetyDGX agent

arXiv:2511.09008v2 Announce Type: replace-cross Abstract: Large Language Models perform well at natural language interpretation and reasoning, but their lack of formal correctness guarantees limits th

Can Induced Emotion Bias LLM Behaviors in Sequential Decision Making?

SafetyDGX agent

arXiv:2607.12631v1 Announce Type: cross Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents in high-stakes domains, understanding contextual factors that may modul

Expert Knowledge-driven Reinforcement Learning for Autonomous Racing via Trajectory Guidance and Dynamics Constraints

SafetyDGX agent

arXiv:2603.05842v2 Announce Type: replace Abstract: Reinforcement learning has demonstrated significant potential in the field of autonomous driving. However, it suffers from defects such as training

From Sentiment to Actionable Insights: Public Sentiment Analysis of Advanced Air Mobility

SafetyDGX agent

arXiv:2606.20751v2 Announce Type: replace Abstract: Advanced Air Mobility (AAM) is an emerging low-altitude transportation system whose successful deployment depends on both technological progress and

StratMamba: Strategic and Reactive Stream Partitioning for Path-Efficient LiDAR-Based Obstacle Avoidance

SafetyDGX agent

arXiv:2607.12370v1 Announce Type: new Abstract: This paper proposes StratMamba, a dual-stream Mamba-based temporal modeling architecture, to more efficiently capture long-horizon temporal dependencies

10 Jul 2026

Alignment Plausibility: A New Standard for Assuring AI in Healthcare

SafetyDGX agent

arXiv:2607.07766v1 Announce Type: new Abstract: Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy whose operatio

Post-Training in End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2607.08072v1 Announce Type: new Abstract: End-to-end models that map multimodal inputs directly to future trajectories/maneuvers have emerged as an increasingly prominent research paradigm in au

Securing Autonomous Vehicle Systems via Twin-Aware Federated Reinforcement Learning

SafetyDGX agent

arXiv:2607.08137v1 Announce Type: cross Abstract: Federated reinforcement learning (FRL) is crucial for enabling collaborative learning across multiple agents without sharing raw data, thereby enhanci

9 Jul 2026

Online Data Selection Is Implicit Alignment

SafetyDGX agent

arXiv:2607.07023v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) is often treated as a capability-adaptation step, while alignment is attributed to later preference optimization or reinfor

Progressive Crystallization: Turning Agent Exploration into Deterministic, Lower-Cost Workflows in Production

SafetyDGX agent

arXiv:2607.07052v1 Announce Type: cross Abstract: AI agents deployed for IT operations are typically permanent cost centers because every execution requires full LLM inference, even for previously sol

8 Jul 2026

Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique

SafetyDGX agent

arXiv:2602.13213v2 Announce Type: replace Abstract: Commercial insurance underwriting is a labor-intensive process that requires manual review of extensive documentation to assess risk and determine p

7 Jul 2026

Closed-loop vs. Open-loop Kalman Filter Architectures in Airborne Aided Inertial Navigation

SafetyDGX agent

arXiv:2607.03338v1 Announce Type: new Abstract: Closed-loop (or feedback) error-state Kalman filters with their relatives and offspring are the state-of-the-art in modern aided inertial navigation res

Directional Curvature from Armijo Backtracking: A Low-Cost Sharpness Probe and a Calibration-Free Learning-Rate Safeguard for Adam

SafetyDGX agent

arXiv:2607.03998v1 Announce Type: new Abstract: The local sharpness of the loss, the top Hessian eigenvalue lambda_1, determines the largest stable gradient step, but measuring it normally requires La

Drone delivery is about to be cheaper than cars. But the drone is only 15% of the work to get there. A decade after starting in Rwanda flyin…

SafetyDGX agent

Drone delivery is about to be cheaper than cars. But the drone is only 15% of the work to get there. A decade after starting in Rwanda flying blood to one hospital, @zipline has flown 140M+ autonomous

Integrating Physics-Informed Neural Networks for Safe Reinforcement Learning in a 1-DoF Helicopter System

SafetyDGX agent

arXiv:2607.03125v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) offers powerful control for industrial cyber-physical systems (ICPSs), but its 'black-box' exploration risks violating

Interaction Dynamics for Dexterous Manipulation

SafetyDGX agent

arXiv:2606.14606v2 Announce Type: replace Abstract: Dexterous manipulation is fundamentally a problem of interaction dynamics: the hand must track precise finger trajectories, regulate the contact for

NavEYE: Vision-Centered Multi-Sensor Fusion-Based Situational Awareness System for Intelligent Surface Vehicles

SafetyDGX agent

arXiv:2607.03915v1 Announce Type: new Abstract: With the rapid development of sensor and artificial intelligence (AI) technologies, intelligent surface vehicles (ISVs) have gained increasing attention

Pretraining Curricula Enable Selective Fine-tuning

SafetyDGX agent

arXiv:2607.04846v1 Announce Type: cross Abstract: Transformers follow implicit curricula whereby some tasks are learned before others. However, how explicit pretraining curricula influence learning, g

Safe Inference-Time Alignment via Lagrangian Reward Augmentation

SafetyDGX agent

arXiv:2607.02781v1 Announce Type: cross Abstract: Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates.

Toward Trustworthy Large Language Model Agents in Healthcare

Model ReleasesDGX agent

arXiv:2607.05055v1 Announce Type: new Abstract: Healthcare appointment scheduling remains a persistent operational bottleneck, driven by manual coordination, fragmented legacy systems, and high admini

2 Jul 2026

Distributed Multi Robot Lunar Cargo Transportation via Phase Decomposed Reinforcement Learning

SafetyDGX agent

arXiv:2607.00160v1 Announce Type: new Abstract: Modular reconfigurable robotic systems provide a scalable solution for cooperative surface operations in future lunar missions. However, cooperative car

NeHMO: Neural Hamilton-Jacobi Reachability Learning for Decentralized Safe Multi-Arm Motion Planning

SafetyDGX agent

arXiv:2607.00326v1 Announce Type: new Abstract: Safe multi-arm motion planning is a challenging problem in robotics due to its high dimensionality, coupled configuration space, and complex collision c

1 Jul 2026

AI systems that are accessible or even better open-source are safer for a simple reason: more people can inspect them, test them, stress the…

SafetyDGX agent

AI systems that are accessible or even better open-source are safer for a simple reason: more people can inspect them, test them, stress them, and report what breaks or harms to keep the builders acco

Automating Cause-Effect Specification with Knowledge Graphs and Large Language Models

SafetyDGX agent

arXiv:2606.31614v1 Announce Type: cross Abstract: Engineering specifications such as interlocks, alarm rationalization tables, and cause-and-effect (C&E) matrices remain central to process control and

Verification-Gated Agentic Mission-State Governance for Intelligent Industrial Multi-Robot Systems

SafetyDGX agent

arXiv:2606.31339v1 Announce Type: new Abstract: Agentic artificial intelligence is increasingly used to decompose industrial tasks, propose robot actions, and adapt execution plans in dynamic cyber-ph

30 Jun 2026

Being at the launch of Cybercab two years ago was magic. A decade from now there will be millions of these things rolling around the world. …

SafetyDGX agent

Being at the launch of Cybercab two years ago was magic. A decade from now there will be millions of these things rolling around the world. The war over autonomous vehicles is on, as everyone in San F

Carolina Guide: A Multi-Agent RAG System with Institutional Guardrails for Academic Policy Assistance

SafetyDGX agent

arXiv:2606.28360v1 Announce Type: cross Abstract: University students often struggle to navigate complex academic policies, leading to advising bottlenecks and delayed access to critical information.

OGM-CBF: Occupancy Grid Map-based Control Barrier Function for Safe Mobile Robot Control with Memory of out of View Obstacles

SafetyDGX agent

arXiv:2405.10703v4 Announce Type: replace Abstract: Safe control in unknown environments is a key challenge in mobile robotics. Control Barrier Functions (CBFs) provide a principled framework for guar

SPACE: Swarm Pheromone Fields for Adaptive Collision-Aware Exploration

SafetyDGX agent

arXiv:2606.29372v1 Announce Type: new Abstract: Massive robot swarms can explore unknown environments quickly, but adding robots eventually stops helping. Doorways and dense traffic create congestion,

29 Jun 2026

Low-Agreeableness Persona Conditioning for Safe LLM Fine-Tuning

SafetyDGX agent

arXiv:2606.27709v1 Announce Type: cross Abstract: Recent work has shown that fine-tuning large language models (LLMs) for social warmth degrades factual reliability and increases sycophancy. We invest

26 Jun 2026

Human-AI Complementarity: A Goal for Amplified Oversight

SafetyDGX agent

arXiv:2510.26518v2 Announce Type: replace Abstract: Human feedback is critical for aligning AI systems to human values. As AI capabilities improve and AI is used to tackle more challenging tasks, veri

Proposal-Conditioned Latent Diffusion for Closed-Loop Traffic Scenario Generation

SafetyDGX agent

arXiv:2606.27123v1 Announce Type: cross Abstract: Closed-loop traffic simulation remains challenging because it must generate interactive multi-agent behaviors that are scene-consistent and controllab

Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks

SafetyDGX agent

arXiv:2606.27147v1 Announce Type: cross Abstract: Unlike diffusion-based models that operate in continuous latent spaces, autoregressive unified multimodal models produce images by sequentially predic

25 Jun 2026

G2DP: Diffusion Planning with Spatio-Temporal Grid Guidance

SafetyDGX agent

arXiv:2606.26017v1 Announce Type: new Abstract: In autonomous driving, diffusion-based planners have emerged as a promising paradigm for robust motion planning in dense and interactive traffic, as the

MedGuards: Multi-Agent System for Reliable Medical Error Detection and Correction

SafetyDGX agent

arXiv:2606.25651v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed in healthcare settings, accurate error detection and correction in generated or existing text

24 Jun 2026

A Geometry-Informed Computer Vision Method for Detecting and Examining Overtaking Vehicles From A Bicycle

SafetyDGX agent

arXiv:2606.23699v1 Announce Type: new Abstract: Instrumented bicycle studies have produced direct field evidence on vehicle passing behavior, but extracting overtaking events from continuous rear-faci

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning

SafetyDGX agent

arXiv:2606.24231v1 Announce Type: new Abstract: Multimodal driving planning faces a long-standing tension between two paradigms: scoring-based methods benefit from dense reward supervision but are con

HiPath: Hierarchical Vision-Language Alignment for Structured Pathology Report Prediction

SafetyDGX agent

arXiv:2603.19957v2 Announce Type: replace-cross Abstract: Pathology reports are structured, multi-granular documents encoding diagnostic conclusions, histological grades, and ancillary test results ac

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

SafetyDGX agent

arXiv:2606.24388v1 Announce Type: new Abstract: We introduce a large-scale, open-source dataset of pre-generated adversarial attacks for vision-language models (VLMs). The dataset is designed to be di

23 Jun 2026

BiliVLA: Scene-Aware Vision-Language-Action Model with Reinforcement Learning for Autonomous Biliary Endoscopic Navigation

SafetyDGX agent

arXiv:2606.23531v1 Announce Type: new Abstract: Endoscopic retrograde cholangiopancreatography (ERCP) demands precise endoscopic navigation and stable biliary cannulation within a narrow monocular fie

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness

SafetyDGX agent

arXiv:2606.15980v2 Announce Type: replace Abstract: Activation monitors -- lightweight probes trained on a language model's internal representations -- are an increasingly common layer in deployment s

Expert Consensus on Criteria for the Automated Assessment of Laparoscopic Camera Navigation

SafetyDGX agent

arXiv:2606.23131v1 Announce Type: new Abstract: Background: Laparoscopic camera navigation (LCN) is a critical skill, yet its current assessment typically relies on manual rating systems which are tim

Intent-Handover: Grounding Language in Human-Usage Regions for Trustworthy Robot-to-Human Handovers

SafetyDGX agent

arXiv:2503.03579v2 Announce Type: replace-cross Abstract: Spoken instructions in robot-to-human handovers may specify either an object ('the cup') or an intended use ('pour water'); in both cases, suc

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception

SafetyDGX agent

arXiv:2606.20764v1 Announce Type: new Abstract: Reliable spatial decision automation, such as autonomous driving and maritime surveillance, critically depends on robust visual perception. However, rea

SkillHarness: Harnessing Safe Skills for Computer-Use Agents

SafetyDGX agent

arXiv:2606.20636v1 Announce Type: cross Abstract: Computer-Use Agents (CUAs) are increasingly deployed in dynamic interactive environments, creating a growing need for continual skill learning during

11 Jun 2026

AerialClaw: An Open-Source Framework for LLM-Driven Autonomous Aerial Agents

SafetyDGX agent

arXiv:2606.12142v1 Announce Type: cross Abstract: Unmanned aerial vehicles (UAVs) are increasingly used in inspection, search and rescue, environmental monitoring, and emergency response. However, mos

Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code

SafetyDGX agent

arXiv:2606.11817v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for code generation, raising concerns that they may be misused to produce malicious code. Meanwhile

UGV-Conditioned Multi-UAV Informative Planning on a Shared Exposure Belief

SafetyDGX agent

arXiv:2606.12306v1 Announce Type: new Abstract: Safe ground navigation in large, threat-augmented environments requires aerial support that actively reduces the risks that a ground vehicle faces along

10 Jun 2026

EM-Fall: Embodied mmWave Sensing for Day-and-Night Fall Detection on Humanoid Robots

SafetyDGX agent

arXiv:2606.11109v1 Announce Type: new Abstract: Falls are one of the leading causes of injury and hospitalization among elderly individuals, making reliable fall awareness an essential capability for

Why are AI research restrictions treated differently from every other safeguard? @theemozilla and @karan4d, co-founders of Nous Research: 'O…

SafetyDGX agent

Why are AI research restrictions treated differently from every other safeguard? @theemozilla and @karan4d, co-founders of Nous Research: 'On the bio stuff... just saying no to the user and being hone

9 Jun 2026

AI Assurance in UK Defence: Challenges in Operationalising JSP 936

SafetyDGX agent

arXiv:2606.09414v1 Announce Type: cross Abstract: This report examines practical challenges in operationalising JSP 936 Part 1 for AI assurance in UK Defence. Using a structured interpretive review of

Autonomous Incident Resolution at Hyperscale: An Agentic AI Architecture for Network Operations

SafetyDGX agent

arXiv:2606.09122v1 Announce Type: cross Abstract: Cloud network infrastructure at hyperscale presents unique operational challenges where traditional human-driven incident response cannot keep pace wi

Can the Environment Speak for Itself? T^{2}-GRPO: A Turn-Trajectory Group Relative Policy Optimization for Caregiver Agents

SafetyDGX agent

arXiv:2606.08875v1 Announce Type: new Abstract: Optimizing large language models (LLMs) for long-horizon caregiver agents requires balancing delayed task objectives with immediate environment dynamics

Fast LLM-Based Semantic Filtering: From a Unified Framework to an Adaptive Two-Phase Method

SafetyDGX agent

arXiv:2606.08090v1 Announce Type: cross Abstract: Evaluating a natural-language yes/no predicate over a document corpus under an accuracy target - the semantic filter - is a cornerstone of LLM-based d

← Previous
1…2728293031…240
Next →