AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
Safety

Intent-Handover: Grounding Language in Human-Usage Regions for Trustworthy Robot-to-Human Handovers

DGX agent

arXiv:2503.03579v2 Announce Type: replace-cross Abstract: Spoken instructions in robot-to-human handovers may specify either an object ('the cup') or an intended use ('pour water'); in both cases, suc

safetyarxiv-cs-lg
23 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception

DGX agent

arXiv:2606.20764v1 Announce Type: new Abstract: Reliable spatial decision automation, such as autonomous driving and maritime surveillance, critically depends on robust visual perception. However, rea

safetyarxiv-cs-cv
23 Jun 2026
Safety

SkillHarness: Harnessing Safe Skills for Computer-Use Agents

DGX agent

arXiv:2606.20636v1 Announce Type: cross Abstract: Computer-Use Agents (CUAs) are increasingly deployed in dynamic interactive environments, creating a growing need for continual skill learning during

safetyarxiv-cs-lg
23 Jun 2026
Safety

AerialClaw: An Open-Source Framework for LLM-Driven Autonomous Aerial Agents

DGX agent

arXiv:2606.12142v1 Announce Type: cross Abstract: Unmanned aerial vehicles (UAVs) are increasingly used in inspection, search and rescue, environmental monitoring, and emergency response. However, mos

safetyarxiv-cs-cv
11 Jun 2026
Safety

Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code

DGX agent

arXiv:2606.11817v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for code generation, raising concerns that they may be misused to produce malicious code. Meanwhile

safetyarxiv-cs-ai
11 Jun 2026
Safety

UGV-Conditioned Multi-UAV Informative Planning on a Shared Exposure Belief

DGX agent

arXiv:2606.12306v1 Announce Type: new Abstract: Safe ground navigation in large, threat-augmented environments requires aerial support that actively reduces the risks that a ground vehicle faces along

safetyarxiv-cs-ro
11 Jun 2026
Safety

EM-Fall: Embodied mmWave Sensing for Day-and-Night Fall Detection on Humanoid Robots

DGX agent

arXiv:2606.11109v1 Announce Type: new Abstract: Falls are one of the leading causes of injury and hospitalization among elderly individuals, making reliable fall awareness an essential capability for

safetyarxiv-cs-ro
10 Jun 2026
Safety

Why are AI research restrictions treated differently from every other safeguard? @theemozilla and @karan4d, co-founders of Nous Research: 'O…

DGX agent

Why are AI research restrictions treated differently from every other safeguard? @theemozilla and @karan4d, co-founders of Nous Research: 'On the bio stuff... just saying no to the user and being hone

safetynous-research--x
10 Jun 2026
Safety

AI Assurance in UK Defence: Challenges in Operationalising JSP 936

DGX agent

arXiv:2606.09414v1 Announce Type: cross Abstract: This report examines practical challenges in operationalising JSP 936 Part 1 for AI assurance in UK Defence. Using a structured interpretive review of

safetyarxiv-cs-ai
9 Jun 2026
Safety

Autonomous Incident Resolution at Hyperscale: An Agentic AI Architecture for Network Operations

DGX agent

arXiv:2606.09122v1 Announce Type: cross Abstract: Cloud network infrastructure at hyperscale presents unique operational challenges where traditional human-driven incident response cannot keep pace wi

safetyarxiv-cs-ai
9 Jun 2026
Safety

Can the Environment Speak for Itself? T^{2}-GRPO: A Turn-Trajectory Group Relative Policy Optimization for Caregiver Agents

DGX agent

arXiv:2606.08875v1 Announce Type: new Abstract: Optimizing large language models (LLMs) for long-horizon caregiver agents requires balancing delayed task objectives with immediate environment dynamics

safetyarxiv-cs-ai
9 Jun 2026
Safety

Fast LLM-Based Semantic Filtering: From a Unified Framework to an Adaptive Two-Phase Method

DGX agent

arXiv:2606.08090v1 Announce Type: cross Abstract: Evaluating a natural-language yes/no predicate over a document corpus under an accuracy target - the semantic filter - is a cornerstone of LLM-based d

safetyarxiv-cs-ai
9 Jun 2026
Safety

MC-CPO: Mastery-Conditioned Constrained Policy Optimization for Pedagogically Safe Intelligent Tutoring Systems

DGX agent

arXiv:2604.04251v2 Announce Type: replace Abstract: Intelligent tutoring systems increasingly rely on reinforcement learning to personalise instruction, yet optimising for observable engagement signal

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Safe-RULE: Safe Reinforcement UnLEarning

DGX agent

arXiv:2606.09559v1 Announce Type: cross Abstract: Offline safe reinforcement learning (Safe RL) enables policy learning without online interactions, making it suitable for safety-critical systems such

model-releasesarxiv-cs-ai
9 Jun 2026
Safety

Semantic Quorum Assurance: Collective Certification for Non-Deterministic AI Infrastructure

DGX agent

arXiv:2606.08021v1 Announce Type: cross Abstract: As large language model (LLM) agents are integrated into autonomous cloud operations, distributed systems face a semantic reliability problem: propose

safetyarxiv-cs-ai
9 Jun 2026
Safety

Workflow-to-Skill: Skill Creation via Routing-Workflow-Semantics-Attachments Decomposition

DGX agent

arXiv:2606.06893v1 Announce Type: new Abstract: Large language model agents increasingly rely on Skills to encode procedural knowledge, yet high-quality Skills remain costly to hand-write. This paper

safetyarxiv-cs-ai
8 Jun 2026
Safety

UNIVID: Unified Vision-Language Model for Video Moderation

DGX agent

arXiv:2606.05748v1 Announce Type: cross Abstract: Global-scale video moderation faces a dual challenge: the need for fine-grained multi-modal reasoning and the demand for interpretable outputs to supp

safetyarxiv-cs-cl
5 Jun 2026
Safety

Distribution-Free Risk-Aware Planning and Control Under Uncertainty Using Conformal Spectral Risk Control

DGX agent

arXiv:2606.04185v1 Announce Type: new Abstract: Safe navigation in dynamic and uncertain environments often relies on accurate estimation of, or assumptions about, the true underlying uncertainty. How

safetyarxiv-cs-ro
4 Jun 2026
Safety

From Agent Traces to Trust: Evidence Tracing and Execution Provenance in LLM Agents

DGX agent

arXiv:2606.04990v1 Announce Type: cross Abstract: Large language model (LLM)-based agents increasingly solve complex tasks by interacting with external tools, retrieval systems, memory modules, enviro

safetyarxiv-cs-ai
4 Jun 2026
Model Releases

Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use

DGX agent

arXiv:2603.03205v2 Announce Type: replace Abstract: Agentic language models operate in a fundamentally different safety regime than chat models: they must plan, call tools, and execute long-horizon ac

model-releasesarxiv-cs-cl
4 Jun 2026
Model Releases

Reasoning over Boundaries: Enhancing Specification Alignment via Test-time Deliberation

DGX agent

arXiv:2509.14760v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied in diverse real-world scenarios, each governed by bespoke behavioral and safety specifications

model-releasesarxiv-cs-cl
4 Jun 2026
Safety

RSC: Decentralized Rigid Formation Flocking for Large-Scale Swarms via Hybrid Predictive Control and Online Reconfiguration

DGX agent

arXiv:2606.04248v1 Announce Type: new Abstract: Decentralized rigid formation flocking requires a swarm of autonomous agents to maintain a predetermined geometric configuration while moving, relying s

safetyarxiv-cs-ro
4 Jun 2026
Safety

Bridging Predictive Uncertainty and Safe Action: Sample-Conditioned Differentiable Planning for Autonomous Driving

DGX agent

arXiv:2606.03296v1 Announce Type: new Abstract: Complex, dynamic, and interactive driving environments pose significant challenges for autonomous driving, primarily due to the pervasive uncertainty of

safetyarxiv-cs-ro
3 Jun 2026
Local Ai

DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair

DGX agent

arXiv:2606.03601v1 Announce Type: cross Abstract: While safety alignment and guardrails help large language models (LLMs) avoid harmful outputs, they can also induce overrefusal, i.e., unwarranted rej

local-aiarxiv-cs-ai
3 Jun 2026
Safety

Latent Activation Editing: Inference-Time Refinement of Learned Policies for Safer Multirobot Navigation

DGX agent

arXiv:2509.20623v2 Announce Type: replace Abstract: Reinforcement learning has enabled significant progress in complex domains such as coordinating and navigating multiple quadrotors. However, even we

safetyarxiv-cs-ro
3 Jun 2026
Safety

OpenAI public policy agenda

DGX agent

OpenAI's public policy agenda outlines the organization's priorities and positions regarding AI regulation and governance. The document likely details OpenAI's recommendations for government policies,

safetyopenai
3 Jun 2026
Safety

Towards a Science of AI Agent Reliability

DGX agent

arXiv:2602.16666v3 Announce Type: replace Abstract: AI agents are increasingly deployed to execute important tasks. While rising accuracy scores on standard benchmarks suggest rapid progress, many age

safetyarxiv-cs-ai
3 Jun 2026
Safety

When Models Refuse: Political Steerability and Feature Richness as Measures of Ideological Depth

DGX agent

arXiv:2508.21448v3 Announce Type: replace Abstract: Large language models (LLMs) sometimes refuse to follow benign instructions, such as declining to argue a political position or adopt a stated perso

safetyarxiv-cs-cl
3 Jun 2026
Safety

Infeasible optimization problems and the hierarchical augmented Lagrangian method in imitation learning

DGX agent

arXiv:2606.00730v1 Announce Type: new Abstract: Imitation learning (IL) is an effective approach to train complex robotics policies. Recent works have introduced hard constraints into imitation-learni

safetyarxiv-cs-ro
2 Jun 2026
Safety

Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing

DGX agent

arXiv:2606.00033v1 Announce Type: cross Abstract: While mechanistic interpretability (MI) has produced important insights into neural network internals, the field has yet to establish a standardized s

safetyarxiv-cs-ai
2 Jun 2026
Safety

Safe2Drive: Evaluating Safe Driving Behaviors of E2E Autonomous Driving Models

DGX agent

arXiv:2606.00191v1 Announce Type: cross Abstract: Recent end-to-end (E2E) autonomous driving policies achieve high driving scores in closed-loop simulations. Yet it remains unclear whether these polic

safetyarxiv-cs-cv
2 Jun 2026
Safety

Tether-Aware Dynamic Collision Avoidance for USV-HROV Systems

DGX agent

arXiv:2606.01112v1 Announce Type: new Abstract: Heterogeneous marine robotic systems composed of an unmanned surface vehicle (USV) and a hybrid remotely operated vehicle (HROV) have shown great potent

safetyarxiv-cs-ro
2 Jun 2026
Safety

Stateful Online Monitoring Catches Distributed Agent Attacks

DGX agent

arXiv:2605.31593v1 Announce Type: cross Abstract: Language models can find thousands of severe software vulnerabilities, and agents are increasingly being misused for cyberattacks. To avoid detection,

safetyarxiv-cs-ai
1 Jun 2026
Safety

TARIC: Memory-Augmented Traversability-Aware Outdoor VLN under Interrupted Semantic Cues

DGX agent

arXiv:2605.31121v1 Announce Type: cross Abstract: Outdoor vision-language navigation (VLN) in long-range, open-world environments is frequently disrupted by semantic-cue interruptions, where informati

safetyarxiv-cs-ai
1 Jun 2026
Safety

Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency

DGX agent

arXiv:2605.30208v1 Announce Type: cross Abstract: AI-assisted coding tools have altered software production. At Meta, significant lines of code per human-landed diff grew by 105.9% year over year and

safetyarxiv-cs-ai
29 May 2026
Safety

LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedback

DGX agent

arXiv:2605.30273v1 Announce Type: cross Abstract: Large language models (LLMs) show promise in generating supportive responses for mental health queries, but improving their usefulness, empathy, and s

safetyarxiv-cs-ai
29 May 2026
Local Ai

A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving

DGX agent

arXiv:2605.27763v1 Announce Type: new Abstract: Safety evaluations of language models often treat serving configuration as fixed background infrastructure, but batch condition is an untested treatment

local-aiarxiv-cs-lg
28 May 2026
Safety

Chance-Constrained MPPI under State and Dynamic Object Prediction Uncertainty and the Evaluation of Collision Risk Calibration

DGX agent

arXiv:2605.28330v1 Announce Type: new Abstract: Chance-constrained Model Predictive Path Integral (MPPI) control is increasingly adopted for navigation in dynamic environments to explicitly bound coll

safetyarxiv-cs-ro
28 May 2026
Safety

CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders

DGX agent

arXiv:2604.01604v2 Announce Type: replace Abstract: While modern LLMs are aligned to refuse harmful requests, it is essential to understand the underlying mechanistic basis of this refusal behavior fo

safetyarxiv-cs-ai
28 May 2026
Safety

Grounded Cache Routing for Retrieval-Augmented Generation: When Is It Safe to Reuse an Answer?

DGX agent

arXiv:2605.27494v1 Announce Type: cross Abstract: Modern retrieval-augmented generation(RAG) deployments increasingly rely on caching to reduce token cost and time-to-first-token(TTFT). Prefix-level K

safetyarxiv-cs-ai
28 May 2026
Safety

LACUNA: Safe Agents as Recursive Program Holes

DGX agent

arXiv:2605.28617v1 Announce Type: new Abstract: LLM agents increasingly act by writing code, yet a split persists between the runtime that drives the agent and the code the model writes. The runtime o

safetyarxiv-cs-ai
28 May 2026
Model Releases

Models That Know How Evaluations Are Designed Score Safer

DGX agent

arXiv:2605.28591v1 Announce Type: cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified tes

model-releasesarxiv-cs-ai
28 May 2026
Safety

Trump loses more control over AI regulation as Illinois passes landmark law

DGX agent

Illinois' House of Representatives passed SB 315, a landmark bill requiring frontier AI companies like OpenAI and Anthropic to create, publish and annually update plans addressing severe or catastroph

safetyars-technica
28 May 2026
Safety

Beyond Fixed Benchmarks and Worst-Case Attacks: Dynamic Boundary Evaluation for Language Models

DGX agent

arXiv:2605.06213v2 Announce Type: replace Abstract: Evaluating large language models (LLMs) today rests on fixed benchmarks that apply the same set of items to any model, producing ceiling and floor e

safetyarxiv-cs-ai
27 May 2026
Safety

Bridging Control with Neural Network Verifier alpha-beta-CROWN: A Tutorial

DGX agent

arXiv:2605.26577v1 Announce Type: cross Abstract: Learning-based methods for synthesizing controllers have gained popularity due to their high expressiveness and strong empirical performance. However,

safetyarxiv-cs-ai
27 May 2026
Safety

ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation

DGX agent

arXiv:2605.26542v1 Announce Type: cross Abstract: Tool-using agents increasingly operate in open-ended deployment environments, where they compose file systems, web APIs, code interpreters, and enterp

safetyarxiv-cs-ai
27 May 2026
Safety

Modernising Reinforcement Learning-Based Navigation for Embodied Semantic Scene Graph Generation

DGX agent

arXiv:2603.25415v2 Announce Type: replace Abstract: Semantic world models enable embodied agents to reason about objects, relations, and spatial context beyond purely geometric representations. In Org

safetyarxiv-cs-ai
27 May 2026
Safety

Provably Safe Motion Planning Under Unknown Disturbances

DGX agent

arXiv:2605.26625v1 Announce Type: new Abstract: We present a provably safe sampling-based motion planning algorithm for robotic systems affected by random disturbances of unknown distribution. We cons

safetyarxiv-cs-ro
27 May 2026
← Previous
1…3536373839…299
Next →