AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System

DGX agent

arXiv:2604.18789v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is central to aligning Large Language Models (LLMs), yet it introduces a critical vulnerability: an im

safetyarxiv-cs-ai
22 Apr 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models

DGX agent

arXiv:2510.26782v3 Announce Type: replace-cross Abstract: A world model is an internal model that simulates how the world evolves. Given past observations and actions, it predicts the future physical

safetyarxiv-cs-ai
22 Apr 2026
Safety

Integrating Anomaly Detection into Agentic AI for Proactive Risk Management in Human Activity

DGX agent

arXiv:2604.19538v1 Announce Type: new Abstract: Agentic AI, with goal-directed, proactive, and autonomous decision-making capabilities, offers a compelling opportunity to address movement-related risk

safetyarxiv-cs-ai
22 Apr 2026
Safety

Large Language Models Exhibit Normative Conformity

DGX agent

arXiv:2604.19301v1 Announce Type: new Abstract: The conformity bias exhibited by large language models (LLMs) can pose a significant challenge to decision-making in LLM-based multi-agent systems (LLM-

safetyarxiv-cs-ai
22 Apr 2026
Safety

Jailbreaking Large Language Models with Morality Attacks

DGX agent

arXiv:2604.17053v1 Announce Type: new Abstract: Pluralism alignment with AI has the sophisticated and necessary goal of creating AI that can coexist with and serve morally multifaceted humanity. Resea

safetyarxiv-cs-cl
21 Apr 2026
Safety

SaFeR-Steer: Evolving Multi-Turn MLLMs via Synthetic Bootstrapping and Feedback Dynamics

DGX agent

arXiv:2604.16358v1 Announce Type: cross Abstract: MLLMs are increasingly deployed in multi-turn settings, where attackers can escalate unsafe intent through the evolving visual-text history and exploi

safetyarxiv-cs-cl
21 Apr 2026
Safety

STL-Based Motion Planning and Uncertainty-Aware Risk Analysis for Human-Robot Collaboration with a Multi-Rotor Aerial Vehicle

DGX agent

arXiv:2509.10692v3 Announce Type: replace Abstract: This paper presents a motion planning and risk analysis framework for enhancing human-robot collaboration with a Multi-Rotor Aerial Vehicle. The pro

safetyarxiv-cs-ro
21 Apr 2026
Safety

Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning

DGX agent

arXiv:2604.15705v1 Announce Type: new Abstract: Reinforcement Fine-Tuning (RFT) has established itself as a critical paradigm for the alignment of Multi-modal Large Language Models (MLLMs) with comple

safetyarxiv-cs-lg
20 Apr 2026
Safety

SeaAlert: Critical Information Extraction From Maritime Distress Communications with Large Language Models

DGX agent

arXiv:2604.14163v1 Announce Type: new Abstract: Maritime distress communications transmitted over very high frequency (VHF) radio are safety-critical voice messages used to report emergencies at sea.

safetyarxiv-cs-cl
17 Apr 2026
Safety

AWARE: Adaptive Whole-body Active Rotating Control for Enhanced LiDAR-Inertial Odometry under Human-in-the-Loop Interaction

DGX agent

arXiv:2604.10598v1 Announce Type: new Abstract: Human-in-the-loop (HITL) UAV operation is essential in complex and safety-critical aerial surveying environments, where human operators provide navigati

safetyarxiv-cs-ro
14 Apr 2026
Safety

Prompt Injection as Role Confusion

DGX agent

arXiv:2603.12277v3 Announce Type: replace-cross Abstract: Language models remain vulnerable to prompt injection attacks despite extensive safety training. We trace this failure to role confusion: mode

safetyarxiv-cs-ai
14 Apr 2026
Safety

QShield: Securing Neural Networks Against Adversarial Attacks using Quantum Circuits

DGX agent

arXiv:2604.10933v1 Announce Type: cross Abstract: Deep neural networks remain highly vulnerable to adversarial perturbations, limiting their reliability in security- and safety-critical applications.

safetyarxiv-cs-ai
14 Apr 2026
Safety

SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering

DGX agent

arXiv:2508.11290v3 Announce Type: replace Abstract: LLMs increasingly exhibit over-refusal behavior, where safety mechanisms cause models to reject benign instructions that seemingly resemble harmful

safetyarxiv-cs-cl
14 Apr 2026
Safety

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs

DGX agent

arXiv:2604.08846v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have been shown to be vulnerable to malicious queries that can elicit unsafe responses. Recent work uses prom

safetyarxiv-cs-ai
13 Apr 2026
Safety

Large Reasoning Models Learn Better Alignment from Flawed Thinking

DGX agent

arXiv:2510.00938v2 Announce Type: replace Abstract: Large reasoning models (LRMs) 'think' by generating structured chain-of-thought (CoT) before producing a final answer, yet they still lack the abili

safetyarxiv-cs-lg
13 Apr 2026
Safety

AEROS: A Single-Agent Operating Architecture with Embodied Capability Modules

DGX agent

arXiv:2604.07039v1 Announce Type: cross Abstract: Robotic systems lack a principled abstraction for organizing intelligence, capabilities, and execution in a unified manner. Existing approaches either

safetyarxiv-cs-ai
10 Apr 2026
Safety

CMP: Robust Whole-Body Tracking for Loco-Manipulation via Competence Manifold Projection

DGX agent

arXiv:2604.07457v1 Announce Type: new Abstract: While decoupled control schemes for legged mobile manipulators have shown robustness, learning holistic whole-body control policies for tracking global

safetyarxiv-cs-ro
10 Apr 2026
Safety

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs

DGX agent

arXiv:2604.07655v1 Announce Type: cross Abstract: Hard-gated safety checkers often over-refuse and misalign with a vendor's model spec; prevailing taxonomies also neglect robustness and honesty, yield

safetyarxiv-cs-cl
10 Apr 2026
Safety

LLM-based Schema-Guided Extraction and Validation of Missing-Person Intelligence from Heterogeneous Data Sources

DGX agent

arXiv:2604.06571v1 Announce Type: cross Abstract: Missing-person and child-safety investigations rely on heterogeneous case documents, including structured forms, bulletin-style posters, and narrative

safetyarxiv-cs-ai
10 Apr 2026
Safety

VLMShield: Efficient and Robust Defense of Vision-Language Models against Malicious Prompts

DGX agent

arXiv:2604.06502v1 Announce Type: new Abstract: Vision-Language Models (VLMs) face significant safety vulnerabilities from malicious prompt attacks due to weakened alignment during visual integration.

safetyarxiv-cs-lg
10 Apr 2026
Safety

A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization

DGX agent

arXiv:2608.11483v1 Announce Type: new Abstract: Hit-to-lead optimization requires iterative design of hit analogs across competing potency, selectivity, physicochemical, pharmacokinetic, safety, and s

safetyarxiv-cs-ai
13 Aug 2026
Safety

BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model

DGX agent

arXiv:2608.11244v1 Announce Type: new Abstract: Construction standards are critical for building safety and sustainability. Existing standard application workflows rely on keyword-based document retri

safetyarxiv-cs-ai
13 Aug 2026
Safety

Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers

DGX agent

arXiv:2608.12166v1 Announce Type: cross Abstract: Algorithm registers have been championed as a means of providing transparency on the use of algorithms in public services. Yet potential publics diffe

safetyarxiv-cs-ai
13 Aug 2026
Safety

Energy-Aware Wind-Resilient Routing for Truck-Assisted Multi-UAV Delivery under Wind Uncertainty

DGX agent

arXiv:2608.11641v1 Announce Type: cross Abstract: Energy feasibility under wind uncertainty is a critical safety issue for low-altitude air-ground delivery. In truck-UAV systems, UAVs complete assigne

safetyarxiv-cs-ro
13 Aug 2026
Safety

Inverse-dynamics observer design for a linear single-track vehicle model with distributed tire dynamics

DGX agent

arXiv:2603.07499v3 Announce Type: replace-cross Abstract: Accurate estimation of the vehicle's sideslip angle and tire forces is essential for enhancing safety and handling performances in unknown dri

safetyarxiv-cs-ro
13 Aug 2026
Safety

Video2Track: From Real-World Interaction Videos to Steerable Adversarial Closed-Track Testing for Automated Driving Systems

DGX agent

arXiv:2608.11592v1 Announce Type: new Abstract: Closed-track testing plays a fundamental role in the verification and validation of automated driving systems (ADS), particularly for safety-critical sc

safetyarxiv-cs-ro
13 Aug 2026
Safety

CRHT: A Continuous Regression Hybrid Transformer for Vessel Trajectory Prediction with Online Cluster Sampling

DGX agent

arXiv:2608.10256v1 Announce Type: new Abstract: Accurate vessel trajectory prediction is critical for maritime safety and anomaly detection, yet existing models often struggle with geographic bias and

safetyarxiv-cs-lg
12 Aug 2026
Model Releases

Introspective Attention Modulation for Safe Text-to-Image Generation

DGX agent

arXiv:2607.14945v2 Announce Type: replace Abstract: State-of-the-art flow based text-to-image (T2I) models exhibit remarkable generative abilities but remain vulnerable to producing unsafe content. Pr

model-releasesarxiv-cs-cv
12 Aug 2026
Safety

SeFaR: Semantic Feature-aware Robustness Testing of Deep Neural Networks

DGX agent

arXiv:2608.10289v1 Announce Type: new Abstract: Deep neural networks are increasingly deployed in safety-critical domains as perception modules, where failures are often caused due to rare and under-r

safetyarxiv-cs-cv
12 Aug 2026
Safety

City Sentinel: A Unified AI-Based Smart Surveillance Framework for Real-Time Multi-Threat Detection Using Deep Learning

DGX agent

arXiv:2608.08887v1 Announce Type: new Abstract: Rapid urbanization has increased the need for surveillance systems that can monitor multiple public safety risks at the same time. Traditional systems o

safetyarxiv-cs-cv
11 Aug 2026
Safety

FlowErase-OPD: Multi-Concept Erasure via Anchored On-Policy Distillation in Flow Matching Models

DGX agent

arXiv:2608.07620v1 Announce Type: new Abstract: Recent advances in flow matching models have substantially improved the quality of text-to-image generation, but have also raised increasing safety conc

safetyarxiv-cs-cv
11 Aug 2026
Safety

Machine-Learning-Based Diagnostic Framework for Passive Ultrasonic Detection of Railway Wheel Defects

DGX agent

arXiv:2608.08301v1 Announce Type: new Abstract: Reliable identification of railway wheel defects is important for safety and maintenance. This study develops a machine-learning-based diagnostic framew

safetyarxiv-cs-lg
11 Aug 2026
Safety

PAM: Training Policy-Aligned Moderation Filters at Scale

DGX agent

arXiv:2505.19766v4 Announce Type: replace Abstract: Large language models (LLMs) remain vulnerable to misalignment and jailbreaks, making external safeguards like moderation filters essential, yet exi

safetyarxiv-cs-cl
11 Aug 2026
Safety

Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models

DGX agent

arXiv:2508.16406v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attacks, which attempt to elicit harmful responses from LLMs. The evolving nature

safetyarxiv-cs-cl
11 Aug 2026
Safety

KnifeHunter: Structured Local Representation Learning for Fine-Grained Knife Image Retrieval in Law Enforcement

DGX agent

arXiv:2608.07057v1 Announce Type: new Abstract: Knife-enabled violence presents a major public safety challenge, and law enforcement agencies require scalable tools for catalogue-level knife identific

safetyarxiv-cs-cv
10 Aug 2026
Safety

Online Conformal Prediction Beyond Feedback

DGX agent

arXiv:2608.07139v1 Announce Type: new Abstract: Uncertainty quantification is essential when deploying machine learning models in safety-critical applications. Online conformal prediction (OCP) provid

safetyarxiv-cs-lg
10 Aug 2026
Safety

Deterministic World Models for Closed-loop Reachability Analysis of End-to-End Vision-based Control

DGX agent

arXiv:2512.08991v3 Announce Type: replace Abstract: End-to-end image controllers that map raw camera frames directly to control actions are increasingly deployed in safety-critical systems. However, f

safetyarxiv-cs-cv
7 Aug 2026
Safety

Failing Gracefully: Mitigating Impact of Inevitable Robot Failures

DGX agent

arXiv:2608.05313v1 Announce Type: cross Abstract: Service robots operate in household environments shared with humans, pets, and everyday objects, where they are highly susceptible to failures such as

safetyarxiv-cs-ai
7 Aug 2026
Safety

Symbol Grounding in Neuro-Symbolic AI: A Gentle Introduction to Reasoning Shortcuts

DGX agent

arXiv:2510.14538v3 Announce Type: replace Abstract: Neuro-symbolic (NeSy) AI aims to develop deep neural networks whose predictions comply with prior knowledge encoding, e.g. safety or structural cons

safetyarxiv-cs-ai
7 Aug 2026
Safety

Worst-Case Distance-Aware Error Bounds for Neural Networks

DGX agent

arXiv:2510.22021v3 Announce Type: replace Abstract: Safety-critical applications of machine learning require uncertainty estimates that support reliable worst-case analysis. Neural networks (NNs) prov

safetyarxiv-cs-lg
7 Aug 2026
Safety

An Inline Control Architecture for Language Models in Intelligent Transportation Systems

DGX agent

arXiv:2608.04065v1 Announce Type: cross Abstract: Vehicle-to-everything (V2X) systems increasingly incorporate large language models (LLMs) for semantic tasks such as message summarization, operator a

safetyarxiv-cs-ai
6 Aug 2026
Safety

Approximate Multi-Objective Search Under Rulebooks

DGX agent

arXiv:2608.04398v1 Announce Type: cross Abstract: Robotic planning often involves multiple objectives with complex priority relationships, such as safety, efficiency, and regulatory compliance. Rulebo

safetyarxiv-cs-ai
6 Aug 2026
Safety

Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification

DGX agent

arXiv:2607.16868v2 Announce Type: replace Abstract: Large Language Models often produce confidently stated yet unreliable outputs, posing critical challenges for deployment in safety-sensitive applica

safetyarxiv-cs-ai
6 Aug 2026
Safety

CheckOne: Lightweight Fault Detection and Mitigation for Vision Transformers

DGX agent

arXiv:2608.04035v1 Announce Type: cross Abstract: The wide adoption of Vision Transformers (ViTs) in safety-critical applications raises reliability concerns related to hardware faults. Algorithm-Base

safetyarxiv-cs-ai
6 Aug 2026
Safety

muSync-GS: Physics-Synchronized Driving Video Synthesis for Weather and Geometric Road Hazards

DGX agent

arXiv:2608.04412v1 Announce Type: new Abstract: High-quality driving data are essential for autonomous-driving systems and generative world models. However, rare and safety-critical scenarios involvin

safetyarxiv-cs-cv
6 Aug 2026
Safety

Double Down on Defense: Strengthening Deep Perceptual Hashes against Evasion Attacks without Retraining

DGX agent

arXiv:2608.03101v1 Announce Type: new Abstract: Near-duplicate image matching is crucial for trust and safety, provenance verification, copyright enforcement, and large-scale visual search. Modern pla

safetyarxiv-cs-cv
5 Aug 2026
Safety

Evading Chain-of-Thought Monitoring Through Model Poisoning

DGX agent

arXiv:2608.02820v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring is an increasingly important component of AI safety stacks but relies on the assumption that a model's reasoning tra

safetyarxiv-cs-ai
5 Aug 2026
Safety

Optimising for Flourishing: Flourishing Metrics and Return on Flourishing as Success Criteria for Artificial Intelligence and Post-AGI Economic Systems

DGX agent

arXiv:2608.00151v2 Announce Type: replace-cross Abstract: Current evaluation frameworks for artificial intelligence focus mainly on capability, safety, and proxies such as adoption, engagement, effici

safetyarxiv-cs-ai
5 Aug 2026
← Previous
1…2021222324…257
Next →