AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
16 Jul 2026

A VAE-Driven Multi-Task Satellite-Aided Semantic Communication Framework for 6G-Enabled Connected Autonomous Vehicles

SafetyDGX agent

arXiv:2607.13494v1 Announce Type: new Abstract: The development of smart transportation systems and the introduction of 6G wireless communication technologies have significantly changed vehicle networ

Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management

SafetyDGX agent

arXiv:2607.13239v1 Announce Type: new Abstract: Foundation models, including large language models (LLMs) and vision-language models (VLMs), are increasingly used for transportation management center

Layered Risk Mapping for Autonomous Patient Transport in Expeditionary Medical Facilities

SafetyDGX agent

arXiv:2607.13497v1 Announce Type: new Abstract: In expeditionary medical facilities, routine patient transport imposes a compounding burden of personal protective equipment consumption, staff diversio

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities

SafetyDGX agent

arXiv:2607.13596v1 Announce Type: cross Abstract: When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by acknowledgi

STITCHER: Constrained Trajectory Planning in Complex Environments with Real-Time Motion Primitive Search

SafetyDGX agent

arXiv:2510.14893v4 Announce Type: replace Abstract: Autonomous high-speed navigation through large, complex environments requires real-time generation of agile trajectories that are dynamically feasib

Temperature Scaling Is Not Enough: Calibration Gaps Under Human Label Distributions

SafetyDGX agent

arXiv:2607.13423v1 Announce Type: new Abstract: Temperature scaling is the dominant post-hoc calibration method in modern deep learning. Its theoretical justification rests on an assumption that is ra

Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems

SafetyDGX agent

arXiv:2607.13048v1 Announce Type: cross Abstract: Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) that provide rich semantic understanding at

15 Jul 2026

Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs

SafetyDGX agent

arXiv:2607.12273v1 Announce Type: cross Abstract: As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-world risks, w

G-SHARE: A Guideline-Based Structured Reasoning Framework for Human-Factor Event Diagnosis

SafetyDGX agent

arXiv:2607.11892v1 Announce Type: cross Abstract: Human-factor event diagnosis is essential for learning from operational events in nuclear power plants, yet its quality depends strongly on expert int

Git-Assistant: Planning-Based Support for Updating Git Repositories

SafetyDGX agent

arXiv:2607.09224v2 Announce Type: replace-cross Abstract: Version control systems are essential for collaborative software development, yet tools like git remain challenging for many practitioners. Re

Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels

Model ReleasesDGX agent

arXiv:2607.12792v1 Announce Type: cross Abstract: Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach. Such evaluations, however, are se

Streamlining stereo differentiable rendering for marker-free real-time tracking of surgical robots

SafetyDGX agent

arXiv:2607.12604v1 Announce Type: new Abstract: Purpose: Marker-based tracking of surgical robots is occlusion-prone in cluttered operating rooms. We evaluate stereo differentiable rendering for marke

Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap

SafetyDGX agent

arXiv:2607.12113v1 Announce Type: cross Abstract: One year ago, the AISLE roadmap argued that autonomous laboratories operated as isolated islands and proposed a grassroots network organized around fi

Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents

SafetyDGX agent

arXiv:2607.12790v1 Announce Type: new Abstract: Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption: a reliable e

11 Jul 2026

The Anglo-Scottish Enlightenment – the real antidote to Rousseau and Voltaire The French Enlightenment and the Anglo-Scottish Enlightenment …

SafetyDGX agent

The Anglo-Scottish Enlightenment – the real antidote to Rousseau and Voltaire The French Enlightenment and the Anglo-Scottish Enlightenment happened simultaneously, in the same century, reading the sa

10 Jul 2026

Detecting Ladder Logic Bombs in IEC 61131-3 PLC Programs using ESBMC-PLC+: A Formal Verification Approach with Trigger Synthesis

SafetyDGX agent

arXiv:2607.08417v1 Announce Type: new Abstract: A Ladder Logic Bomb (LLB) is malicious control logic in a Programmable Logic Controller (PLC) program that lies dormant until a trigger activates a payl

Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning

SafetyDGX agent

arXiv:2607.07859v1 Announce Type: new Abstract: Reinforcement learning (RL) research has increasingly shifted focus towards alignment, ensuring agents learn behaviors adhering to human values. While h

From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents

SafetyDGX agent

arXiv:2607.08028v1 Announce Type: new Abstract: Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context. Productization

In preliminary findings, the EU Commission said Facebook's and Instagram's 'addictive design' violates the DSA, telling Meta to make changes or risk hefty fines (Adam Satariano/New York Times)

SafetyDGX agent

Adam Satariano / New York Times: In preliminary findings, the EU Commission said Facebook's and Instagram's “addictive design” violates the DSA, telling Meta to make changes or risk hefty fines — Euro

In vivo feasibility study of humanoid robots in surgery

SafetyDGX agent

arXiv:2607.07972v1 Announce Type: new Abstract: Recent advances in actuation, control and learning have rapidly pushed humanoid robots from a distant vision towards near-term real-world deployment. He

It Takes a MAESTRO To Prune Bad Experts

SafetyDGX agent

arXiv:2607.08601v1 Announce Type: new Abstract: Sparsely-activated Mixture-of-Experts (MoE) language models achieve remarkable inference efficiency by activating only a small fraction of parameters pe

Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs

SafetyDGX agent

arXiv:2607.07903v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing appro

9 Jul 2026

Agentic Data Environments

SafetyDGX agent

arXiv:2607.07397v1 Announce Type: new Abstract: Autonomous agents promise substantial gains in speed, scale, and labor efficiency, but their failures can impose abrupt and often irreversible costs. Th

Behavior Foundations for Quadruped Robots: ABot-C0 Technical Report

SafetyDGX agent

arXiv:2607.07370v1 Announce Type: cross Abstract: In embodied intelligence systems, the motion controller serves as the critical bridge between semantic reasoning and physical execution. Humanoid cont

Manual, Joystick, or Haptic Control? An In Vitro Comparison of Navigation Strategies for Robotic Interventional Neuroradiology Procedures

SafetyDGX agent

arXiv:2607.07253v1 Announce Type: new Abstract: Objective: To evaluate robotic controller interfaces for interventional neuroradiology procedures in-vitro incorporating a force-sensing platform to ass

Mechanistic Interpretability for Neural Networks: Circuits, Sparse Features and Symbolic Reasoning

SafetyDGX agent

arXiv:2607.07316v1 Announce Type: new Abstract: This article offers a comprehensive overview of mechanistic interpretability, an emerging field that seeks to reverse-engineer the internal algorithms o

Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors

SafetyDGX agent

arXiv:2607.07368v1 Announce Type: cross Abstract: AI control is a family of techniques to prevent an AI with malicious goals from subverting its operator's intent. AI Control usually studies a single

NonTextual Target Attack

SafetyDGX agent

arXiv:2510.02999v5 Announce Type: replace-cross Abstract: Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with

R^3: Advertisement Compliance Rectification via Group-Relative Experience Extractor and Curriculum Reinforcement

SafetyDGX agent

arXiv:2607.07318v1 Announce Type: new Abstract: Rigorous content moderation is crucial for online advertising but leads to millions of daily rejections. This scale renders manual rectification infeasi

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

SafetyDGX agent

arXiv:2607.07663v1 Announce Type: new Abstract: AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses during deployment, training on data t

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents

SafetyDGX agent

arXiv:2607.07436v1 Announce Type: new Abstract: A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the stru

The Power of Backdoor Absorption in Community Training

SafetyDGX agent

arXiv:2607.06643v1 Announce Type: cross Abstract: Backdoor attacks severely threaten large-scale AI models. When model owners delegate training to external compute providers within a decentralized tra

Zero-Human Demonstration End-to-end Autonomous Driving with Trajectory Scorer

SafetyDGX agent

arXiv:2510.24108v2 Announce Type: replace-cross Abstract: Human demonstrations are widely considered the cornerstone of end-to-end (E2E) autonomous driving despite human demonstration's scarcity for l

8 Jul 2026

AEGIS: A Mechanism-Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models

Local AiDGX agent

arXiv:2607.06120v1 Announce Type: new Abstract: Text-to-image diffusion models have achieved high visual fidelity and broad adoption, but remain vulnerable to safety violations when adversaries exploi

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability

Model ReleasesDGX agent

arXiv:2607.06196v1 Announce Type: new Abstract: Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional law

Property-Driven Synthetic Data Engineering for Data-Scarce Software Systems: Reflections from the Breast Cancer Domain

SafetyDGX agent

arXiv:2607.06133v1 Announce Type: cross Abstract: Modern software systems increasingly depend on data for analysis, prediction, testing, and decision-making. Yet many important domains, including medi

The Large Cancer Assistant (LCA): A Model-Agnostic Orchestration Framework for Scalable Clinical Decision Support in Oncology

SafetyDGX agent

arXiv:2607.06531v1 Announce Type: new Abstract: - Objective: Multimodal deep learning models in oncology are currently limited by monolithic designs that rigidly couple data ingestion, clinical routin

TILDE: TILt-based Distributional Erasure for Concept Unlearning

SafetyDGX agent

arXiv:2607.06432v1 Announce Type: cross Abstract: Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright disputes,

Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations

SafetyDGX agent

arXiv:2504.05294v3 Announce Type: replace Abstract: Chain-of-thought explanations are widely used to inspect the decision process of large language models (LLMs) and to evaluate the trustworthiness of

7 Jul 2026

A Precedent-Guided Co-Scientist for Side-Effect-Aware Drug Redesign

SafetyDGX agent

arXiv:2607.02944v1 Announce Type: cross Abstract: We propose PRECEDE, a precedent-guided co-scientist for side-effect-aware drug redesign that revises a parent compound to mitigate a specified side ef

A User-driven Design Framework for Robotaxi

SafetyDGX agent

arXiv:2602.19107v3 Announce Type: replace Abstract: Robotaxis are emerging as a promising form of urban mobility, but removing human drivers fundamentally reshapes passenger-vehicle interaction and ra

Agentic Artificial Intelligence for Multistage Physics Experiments at a Large-Scale User Facility Particle Accelerator

SafetyDGX agent

arXiv:2509.17255v2 Announce Type: replace-cross Abstract: We present the first language-model-driven agentic artificial intelligence (AI) system to autonomously execute multi-stage physics experiments

Best-of-Better-N: Generating Pre-Aligned Responses with In-Context Learning

SafetyDGX agent

arXiv:2607.03453v1 Announce Type: cross Abstract: Inference-time alignment methods, such as Best-of-N, offer a flexible alternative to training-based alignment by using reward models to select high-qu

BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations

SafetyDGX agent

arXiv:2603.06576v2 Announce Type: replace-cross Abstract: The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic

CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI

SafetyDGX agent

arXiv:2607.03510v1 Announce Type: cross Abstract: Enterprise artificial intelligence is moving from experimentation into operational workflows. Early programs focused on model access and retrieval-aug

CN-CBF: Composite Neural Control Barrier Function for Robot Navigation in Dynamic Environments

SafetyDGX agent

arXiv:2603.06921v2 Announce Type: replace-cross Abstract: Safe navigation of autonomous robots remains one of the core challenges in the field, especially in dynamic and uncertain environments. One pr

Continuous-Time Gaussian Belief Trees for Motion Planning

Model ReleasesDGX agent

arXiv:2607.02884v1 Announce Type: new Abstract: We address sampling-based motion planning for continuous-time stochastic systems under process and measurement uncertainty, with probabilistic guarantee

CONTRA: Red-Teaming Configurations of Personalizable Agents

SafetyDGX agent

arXiv:2607.03220v1 Announce Type: cross Abstract: Recent tools such as OpenClaw have extended the capabilities of LLM-based agents from simple dialog-based systems to fully autonomous agents. These sy

CRRL: A Causality-Based Reinforcement Learning Framework for Autonomous System Recovery

SafetyDGX agent

arXiv:2607.03177v1 Announce Type: cross Abstract: Traditional reinforcement learning (RL) for recovery in autonomous systems lacks causal understanding and generalizes poorly to novel failure scenario

Defending from GeoLocalization through Adversarial Road Trips

SafetyDGX agent

arXiv:2607.03277v1 Announce Type: new Abstract: Retrieval-based image geolocalization has emerged as a powerful technique for determining the location of a query image by matching it against a large,

Designing Touch for Trauma-Informed Social Robots: A Design Space for Direct and Indirect Actuation

SafetyDGX agent

arXiv:2607.04981v1 Announce Type: new Abstract: Touch is a fundamental communication modality in human-robot interaction and may support grounding, emotional regulation, and stress reduction in therap

Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG

SafetyDGX agent

arXiv:2607.02966v1 Announce Type: new Abstract: Cross-lingual retrieval-augmented generation (RAG) is often deployed in an English-evidence regime, where users query in diverse languages but retrieved

Faithfulness to Refusal: A Causal Audit of Neuron Selectors

SafetyDGX agent

arXiv:2607.05355v1 Announce Type: new Abstract: Attribution scores increasingly identify which neuron rows of a language model matter for applications such as pruning, interpretability, and editing fo

FLOAT Drone for Physical Interaction: Lateral Airflow Reduction, Wrench Modeling, and Adaptive Control

SafetyDGX agent

arXiv:2607.04260v1 Announce Type: new Abstract: Aerial physical interaction represents a promising direction for next-generation unmanned aerial vehicles (UAVs), but it requires an aerial platform tha

How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krugel, and Uhl (2025)

SafetyDGX agent

arXiv:2603.22730v2 Announce Type: replace Abstract: Pfeffer, Krugel, and Uhl (2025) report that OpenAI's reasoning model o1-mini produces more utilitarian responses to the trolley problem and footbrid

Integrated Graph Search and Model Predictive Control for Smooth and Efficient Path Planning in Autonomous Vehicles

SafetyDGX agent

arXiv:2607.04259v1 Announce Type: new Abstract: Path planning is a fundamental component of autonomous vehicles, where achieving safe, comfortable, and dynamically feasible paths while ensuring comput

NeSy-CSA: A Neuro-Symbolic Framework for Open-Ended Critical Scenario Attribution

SafetyDGX agent

arXiv:2607.03847v1 Announce Type: new Abstract: Understanding why discovered scenarios become critical in scenario-based testing is essential for effectively leveraging them in decision-making systems

Open-Attribute Person Retrieval: Finding People Through Distinctive and Novel Attributes

SafetyDGX agent

arXiv:2508.01389v3 Announce Type: replace Abstract: Person retrieval in surveillance videos often depends on attributes described by witnesses or operators. However, the most useful cues in practice a

Position: Use Sparse Autoencoders to Discover Unknowns

SafetyDGX agent

arXiv:2506.23845v2 Announce Type: replace-cross Abstract: While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their usef

Road-Aware Anomaly Segmentation with Query-Guided Polygons and CLIP in Autonomous Driving

SafetyDGX agent

arXiv:2607.04304v1 Announce Type: new Abstract: Traditional semantic segmentation models operate under a closed-set assumption and struggle to recognize unknown or unexpected objects-an essential capa

← Previous
1…3637383940…240
Next →