AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors

DGX agent

arXiv:2607.07368v1 Announce Type: cross Abstract: AI control is a family of techniques to prevent an AI with malicious goals from subverting its operator's intent. AI Control usually studies a single

safetyarxiv-cs-ai
9 Jul 2026
Safety

NonTextual Target Attack

DGX agent
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2510.02999v5 Announce Type: replace-cross Abstract: Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with

safetyarxiv-cs-ai
9 Jul 2026
Safety

R^3: Advertisement Compliance Rectification via Group-Relative Experience Extractor and Curriculum Reinforcement

DGX agent

arXiv:2607.07318v1 Announce Type: new Abstract: Rigorous content moderation is crucial for online advertising but leads to millions of daily rejections. This scale renders manual rectification infeasi

safetyarxiv-cs-cl
9 Jul 2026
Safety

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

DGX agent

arXiv:2607.07663v1 Announce Type: new Abstract: AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses during deployment, training on data t

safetyarxiv-cs-ai
9 Jul 2026
Safety

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents

DGX agent

arXiv:2607.07436v1 Announce Type: new Abstract: A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the stru

safetyarxiv-cs-ai
9 Jul 2026
Safety

The Power of Backdoor Absorption in Community Training

DGX agent

arXiv:2607.06643v1 Announce Type: cross Abstract: Backdoor attacks severely threaten large-scale AI models. When model owners delegate training to external compute providers within a decentralized tra

safetyarxiv-cs-lg
9 Jul 2026
Safety

Zero-Human Demonstration End-to-end Autonomous Driving with Trajectory Scorer

DGX agent

arXiv:2510.24108v2 Announce Type: replace-cross Abstract: Human demonstrations are widely considered the cornerstone of end-to-end (E2E) autonomous driving despite human demonstration's scarcity for l

safetyarxiv-cs-cv
9 Jul 2026
Local Ai

AEGIS: A Mechanism-Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models

DGX agent

arXiv:2607.06120v1 Announce Type: new Abstract: Text-to-image diffusion models have achieved high visual fidelity and broad adoption, but remain vulnerable to safety violations when adversaries exploi

local-aiarxiv-cs-cv
8 Jul 2026
Model Releases

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability

DGX agent

arXiv:2607.06196v1 Announce Type: new Abstract: Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional law

model-releasesarxiv-cs-cl
8 Jul 2026
Safety

Property-Driven Synthetic Data Engineering for Data-Scarce Software Systems: Reflections from the Breast Cancer Domain

DGX agent

arXiv:2607.06133v1 Announce Type: cross Abstract: Modern software systems increasingly depend on data for analysis, prediction, testing, and decision-making. Yet many important domains, including medi

safetyarxiv-cs-ai
8 Jul 2026
Safety

The Large Cancer Assistant (LCA): A Model-Agnostic Orchestration Framework for Scalable Clinical Decision Support in Oncology

DGX agent

arXiv:2607.06531v1 Announce Type: new Abstract: - Objective: Multimodal deep learning models in oncology are currently limited by monolithic designs that rigidly couple data ingestion, clinical routin

safetyarxiv-cs-ai
8 Jul 2026
Safety

TILDE: TILt-based Distributional Erasure for Concept Unlearning

DGX agent

arXiv:2607.06432v1 Announce Type: cross Abstract: Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright disputes,

safetyarxiv-cs-ai
8 Jul 2026
Safety

Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations

DGX agent

arXiv:2504.05294v3 Announce Type: replace Abstract: Chain-of-thought explanations are widely used to inspect the decision process of large language models (LLMs) and to evaluate the trustworthiness of

safetyarxiv-cs-cl
8 Jul 2026
Safety

A Precedent-Guided Co-Scientist for Side-Effect-Aware Drug Redesign

DGX agent

arXiv:2607.02944v1 Announce Type: cross Abstract: We propose PRECEDE, a precedent-guided co-scientist for side-effect-aware drug redesign that revises a parent compound to mitigate a specified side ef

safetyarxiv-cs-ai
7 Jul 2026
Safety

A User-driven Design Framework for Robotaxi

DGX agent

arXiv:2602.19107v3 Announce Type: replace Abstract: Robotaxis are emerging as a promising form of urban mobility, but removing human drivers fundamentally reshapes passenger-vehicle interaction and ra

safetyarxiv-cs-ro
7 Jul 2026
Safety

Agentic Artificial Intelligence for Multistage Physics Experiments at a Large-Scale User Facility Particle Accelerator

DGX agent

arXiv:2509.17255v2 Announce Type: replace-cross Abstract: We present the first language-model-driven agentic artificial intelligence (AI) system to autonomously execute multi-stage physics experiments

safetyarxiv-cs-ai
7 Jul 2026
Safety

Best-of-Better-N: Generating Pre-Aligned Responses with In-Context Learning

DGX agent

arXiv:2607.03453v1 Announce Type: cross Abstract: Inference-time alignment methods, such as Best-of-N, offer a flexible alternative to training-based alignment by using reward models to select high-qu

safetyarxiv-cs-ai
7 Jul 2026
Safety

BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations

DGX agent

arXiv:2603.06576v2 Announce Type: replace-cross Abstract: The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and semantic

safetyarxiv-cs-ai
7 Jul 2026
Safety

CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI

DGX agent

arXiv:2607.03510v1 Announce Type: cross Abstract: Enterprise artificial intelligence is moving from experimentation into operational workflows. Early programs focused on model access and retrieval-aug

safetyarxiv-cs-ai
7 Jul 2026
Safety

CN-CBF: Composite Neural Control Barrier Function for Robot Navigation in Dynamic Environments

DGX agent

arXiv:2603.06921v2 Announce Type: replace-cross Abstract: Safe navigation of autonomous robots remains one of the core challenges in the field, especially in dynamic and uncertain environments. One pr

safetyarxiv-cs-lg
7 Jul 2026
Model Releases

Continuous-Time Gaussian Belief Trees for Motion Planning

DGX agent

arXiv:2607.02884v1 Announce Type: new Abstract: We address sampling-based motion planning for continuous-time stochastic systems under process and measurement uncertainty, with probabilistic guarantee

model-releasesarxiv-cs-ro
7 Jul 2026
Safety

CONTRA: Red-Teaming Configurations of Personalizable Agents

DGX agent

arXiv:2607.03220v1 Announce Type: cross Abstract: Recent tools such as OpenClaw have extended the capabilities of LLM-based agents from simple dialog-based systems to fully autonomous agents. These sy

safetyarxiv-cs-ai
7 Jul 2026
Safety

CRRL: A Causality-Based Reinforcement Learning Framework for Autonomous System Recovery

DGX agent

arXiv:2607.03177v1 Announce Type: cross Abstract: Traditional reinforcement learning (RL) for recovery in autonomous systems lacks causal understanding and generalizes poorly to novel failure scenario

safetyarxiv-cs-ai
7 Jul 2026
Safety

Defending from GeoLocalization through Adversarial Road Trips

DGX agent

arXiv:2607.03277v1 Announce Type: new Abstract: Retrieval-based image geolocalization has emerged as a powerful technique for determining the location of a query image by matching it against a large,

safetyarxiv-cs-cv
7 Jul 2026
Safety

Designing Touch for Trauma-Informed Social Robots: A Design Space for Direct and Indirect Actuation

DGX agent

arXiv:2607.04981v1 Announce Type: new Abstract: Touch is a fundamental communication modality in human-robot interaction and may support grounding, emotional regulation, and stress reduction in therap

safetyarxiv-cs-ro
7 Jul 2026
Safety

Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG

DGX agent

arXiv:2607.02966v1 Announce Type: new Abstract: Cross-lingual retrieval-augmented generation (RAG) is often deployed in an English-evidence regime, where users query in diverse languages but retrieved

safetyarxiv-cs-cl
7 Jul 2026
Safety

Faithfulness to Refusal: A Causal Audit of Neuron Selectors

DGX agent

arXiv:2607.05355v1 Announce Type: new Abstract: Attribution scores increasingly identify which neuron rows of a language model matter for applications such as pruning, interpretability, and editing fo

safetyarxiv-cs-cl
7 Jul 2026
Safety

FLOAT Drone for Physical Interaction: Lateral Airflow Reduction, Wrench Modeling, and Adaptive Control

DGX agent

arXiv:2607.04260v1 Announce Type: new Abstract: Aerial physical interaction represents a promising direction for next-generation unmanned aerial vehicles (UAVs), but it requires an aerial platform tha

safetyarxiv-cs-ro
7 Jul 2026
Safety

How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krugel, and Uhl (2025)

DGX agent

arXiv:2603.22730v2 Announce Type: replace Abstract: Pfeffer, Krugel, and Uhl (2025) report that OpenAI's reasoning model o1-mini produces more utilitarian responses to the trolley problem and footbrid

safetyarxiv-cs-cl
7 Jul 2026
Safety

Integrated Graph Search and Model Predictive Control for Smooth and Efficient Path Planning in Autonomous Vehicles

DGX agent

arXiv:2607.04259v1 Announce Type: new Abstract: Path planning is a fundamental component of autonomous vehicles, where achieving safe, comfortable, and dynamically feasible paths while ensuring comput

safetyarxiv-cs-ro
7 Jul 2026
Safety

NeSy-CSA: A Neuro-Symbolic Framework for Open-Ended Critical Scenario Attribution

DGX agent

arXiv:2607.03847v1 Announce Type: new Abstract: Understanding why discovered scenarios become critical in scenario-based testing is essential for effectively leveraging them in decision-making systems

safetyarxiv-cs-lg
7 Jul 2026
Safety

Open-Attribute Person Retrieval: Finding People Through Distinctive and Novel Attributes

DGX agent

arXiv:2508.01389v3 Announce Type: replace Abstract: Person retrieval in surveillance videos often depends on attributes described by witnesses or operators. However, the most useful cues in practice a

safetyarxiv-cs-cv
7 Jul 2026
Safety

Position: Use Sparse Autoencoders to Discover Unknowns

DGX agent

arXiv:2506.23845v2 Announce Type: replace-cross Abstract: While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their usef

safetyarxiv-cs-ai
7 Jul 2026
Safety

Road-Aware Anomaly Segmentation with Query-Guided Polygons and CLIP in Autonomous Driving

DGX agent

arXiv:2607.04304v1 Announce Type: new Abstract: Traditional semantic segmentation models operate under a closed-set assumption and struggle to recognize unknown or unexpected objects-an essential capa

safetyarxiv-cs-cv
7 Jul 2026
Safety

Scalable Dexterous Robot Learning with AR-based Remote Human-Robot Interactions

DGX agent

arXiv:2602.07341v2 Announce Type: replace Abstract: This paper focuses on the scalable robot learning for manipulation in the dexterous robot arm-hand systems, where the remote human-robot interaction

safetyarxiv-cs-lg
7 Jul 2026
Safety

Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement

DGX agent

arXiv:2607.04277v1 Announce Type: cross Abstract: The pursuit of self-evolving AI raises a critical question: when is autonomous self-improvement sustainable rather than degenerative? Drawing an analo

safetyarxiv-cs-ai
7 Jul 2026
Safety

Short-Horizon Position Accuracy of Single-Track Models: Implications for Motion Planning of Autonomous Vehicles

DGX agent

arXiv:2606.14216v2 Announce Type: replace Abstract: Accurate and computationally efficient vehicle models are essential for motion planning of autonomous vehicles, where positional accuracy directly a

safetyarxiv-cs-ro
7 Jul 2026
Safety

Silicon Sampling via Cross-Survey Transfer

DGX agent

arXiv:2607.03091v1 Announce Type: new Abstract: Silicon sampling-using large language models (LLMs) to simulate human survey respondents-has emerged as a promising approach for augmenting traditional

safetyarxiv-cs-ai
7 Jul 2026
Safety

Training Verifiably Robust Agents Using Set-Based Reinforcement Learning

DGX agent

arXiv:2408.09112v2 Announce Type: replace Abstract: Reinforcement learning policies parametrized by deep neural networks have achieved strong performance for continuous control, yet even small input p

safetyarxiv-cs-lg
7 Jul 2026
Safety

Adaptive Companionship for Group-Following Robots: Handling Dynamically Changing Group Formations

DGX agent

arXiv:2607.01287v1 Announce Type: cross Abstract: Accompanying a group of humans is an essential aspect of developing human-like social cognition in robots. However, human groups typically do not foll

safetyarxiv-cs-ai
3 Jul 2026
Safety

Copewell: A Multi-Agent Swarm Architecture for Equitable Mental Wellness Support

DGX agent

arXiv:2607.02245v1 Announce Type: new Abstract: Mental health disorders affect nearly one billion people globally, yet 75% of individuals in low- and middle-income countries receive no treatment due t

safetyarxiv-cs-ai
3 Jul 2026
Safety

DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving

DGX agent

arXiv:2603.18315v2 Announce Type: replace-cross Abstract: Traditional reinforcement learning (RL) methods rely on manually engineered rewards or sparse collision signals, which fail to capture the ric

safetyarxiv-cs-ai
3 Jul 2026
Safety

Epistemic Goggles: A Pretrained Module that Induces an Epistemic Frame via Gradient Editing

DGX agent

arXiv:2607.01690v1 Announce Type: new Abstract: Finetuning a language model on documents that are explicitly annotated as fictional results in a model that still actually believes the documents' core

safetyarxiv-cs-ai
3 Jul 2026
Safety

ESC: Emotional Self-Correction for Reliable Vision-Language Models

DGX agent

arXiv:2607.02089v1 Announce Type: cross Abstract: Vision-language models (VLMs) have achieved strong performance across diverse multimodal tasks, yet they remain vulnerable to unreliable reasoning. Ex

safetyarxiv-cs-ai
3 Jul 2026
Safety

Fast and Accurate Anomaly Detection in Time Series

DGX agent

arXiv:2607.02046v1 Announce Type: new Abstract: Anomaly detection is a critical and evolving field in Machine Learning, with applications targeting different domains such as cybersecurity, finance, he

safetyarxiv-cs-lg
3 Jul 2026
Safety

kNNGuard: Turning LLM Hidden Activations into a Training-Free Configurable Guardrail

DGX agent

arXiv:2607.02072v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in domains requiring guardrails to detect unsafe, off-topic, or adversarial prompts. Existing g

safetyarxiv-cs-ai
3 Jul 2026
Safety

One Demonstration Is Enough for Real-World Robotic Reinforcement Learning

DGX agent

arXiv:2607.01651v1 Announce Type: new Abstract: Learning effective robot control policies on physical hardware is challenging due to costly data collection and the difficulty of reward specification.

safetyarxiv-cs-ro
3 Jul 2026
Safety

Overthink-Triggered Slowdown Attacks on LVLM-Based Robotic Systems

DGX agent

arXiv:2607.01518v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have been increasingly integrated into robotic systems. However, these models may exhibit overthinking behaviors,

safetyarxiv-cs-ro
3 Jul 2026
← Previous
1…4142434445…257
Next →