AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
Safety

Mismatch-Aware Adaptive Constraint Tightening for Bicycle-Model Trajectory Optimization

DGX agent

arXiv:2605.09376v1 Announce Type: new Abstract: Trajectory optimization for autonomous vehicles usually relies on the kinematic bicycle model because of its computational simplicity. However, when the

safetyarxiv-cs-ro
12 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models

DGX agent

arXiv:2501.03544v5 Announce Type: replace-cross Abstract: Recent text-to-image (T2I) models have exhibited remarkable performance in generating high-quality images from text descriptions. However, the

safetyarxiv-cs-ai
12 May 2026
Safety

Sam Altman swearing to tell the whole truth, and then failing to do so. May 2023.

DGX agent

Sam Altman made statements under oath in May 2023 regarding AI safety and OpenAI's practices, but Gary Marcus critiqued these statements as incomplete or misleading, suggesting Altman failed to fully

safetygary-marcus--x
12 May 2026
Safety

Activation Differences Reveal Backdoors: A Comparison of SAE Architectures

DGX agent

arXiv:2605.07324v1 Announce Type: cross Abstract: Backdoor attacks on language models pose a significant threat to AI safety, where models behave normally on most inputs but exhibit harmful behavior w

safetyarxiv-cs-ai
11 May 2026
Safety

Good summary of today, @katiemiller, but then again it is getting hard to track of the total number of ex-board members who have called Altm…

DGX agent

Good summary of today, @katiemiller, but then again it is getting hard to track of the total number of ex-board members who have called Altman a liar 🤷‍♂️ Also hard to keep track how many OpenAI safet

safetygary-marcus--x
8 May 2026
Safety

A Vision-Based Shared-Control Teleoperation Scheme for Controlling the Robotic Arm of a Four-Legged Robot

DGX agent

arXiv:2508.14994v3 Announce Type: replace-cross Abstract: In hazardous and remote environments, robotic systems perform critical tasks demanding improved safety and efficiency. Among these, quadruped

safetyarxiv-cs-cv
6 May 2026
Safety

EvoJail: Evolutionary Diverse Jailbreak Prompt Generation for Large Language Models

DGX agent

arXiv:2605.02921v1 Announce Type: cross Abstract: As LLMs continue to shape real-world applications, automated jailbreak generation becomes essential to reveal safety weaknesses and guide model improv

safetyarxiv-cs-lg
6 May 2026
Safety

False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models

DGX agent

arXiv:2601.07885v2 Announce Type: replace-cross Abstract: Emoticons are widely used in digital communication to convey affective intent, yet their safety implications for Large Language Models (LLMs)

safetyarxiv-cs-ai
6 May 2026
Safety

Autonomous Reliability Qualification of Ga_2O_3-based Hydrogen and Temperature Sensors via Safe Active Learning

DGX agent

arXiv:2605.00868v1 Announce Type: cross Abstract: We present a Safe Active Learning (SAL) framework for autonomous reliability characterization of rectifying Ga_2O_3-based devices under coupled therma

safetyarxiv-cs-lg
5 May 2026
Safety

Beyond Crash: Hijacking Your Autonomous Vehicle for Fun and Profit

DGX agent

arXiv:2602.07249v2 Announce Type: replace-cross Abstract: Autonomous Vehicles (AVs), especially vision-based AVs, are rapidly being deployed without human operators. As AVs operate in safety-critical

safetyarxiv-cs-lg
5 May 2026
Safety

Disciplined Diffusion: Text-to-Image Diffusion Model against NSFW Generation

DGX agent

arXiv:2605.01113v1 Announce Type: new Abstract: Text-to-image (T2I) diffusion models have the ability to build high-quality pictures from text prompts, but they pose safety concerns because they can g

safetyarxiv-cs-cv
5 May 2026
Safety

Evidence-Based Landing Site Selection and Vison-Based Landing for UAVs in Unstructured Environments

DGX agent

arXiv:2605.01432v1 Announce Type: new Abstract: Autonomous landing in cluttered or unstructured environments remains a safety-critical challenge for unmanned aerial vehicles (UAVs), particularly under

safetyarxiv-cs-ro
5 May 2026
Safety

A Deep Learning-Based CCTV System for Automatic Smoking Detection in Fire Exit Zones

DGX agent

arXiv:2508.11696v3 Announce Type: replace Abstract: A deep learning real-time smoking detection system for CCTV surveillance of fire exit areas is proposed due to critical safety requirements. The dat

safetyarxiv-cs-cv
4 May 2026
Safety

Beyond Suffixes: Token Position in GCG Adversarial Attacks on Large Language Models

DGX agent

arXiv:2602.03265v2 Announce Type: replace Abstract: Large Language Models (LLMs) have seen widespread adoption across multiple domains, creating an urgent need for robust safety alignment mechanisms.

safetyarxiv-cs-lg
4 May 2026
Safety

OpenAI o1 System Card

DGX agent

arXiv:2412.16720v2 Announce Type: replace Abstract: The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provi

safetyarxiv-cs-ai
1 May 2026
Safety

No Pedestrian Left Behind: Real-Time Detection and Tracking of Vulnerable Road Users for Adaptive Traffic Signal Control

DGX agent

arXiv:2604.25887v1 Announce Type: new Abstract: Current pedestrian crossing signals operate on fixed timing without adjustment to pedestrian behavior, which can leave vulnerable road users (VRUs) such

safetyarxiv-cs-cv
29 Apr 2026
Safety

A Self-Supervised Framework for Space Object Behaviour Characterisation

DGX agent

arXiv:2504.06176v3 Announce Type: replace-cross Abstract: Foundation Models, which leverage large neural networks pre-trained on unlabelled data before fine-tuning for specific tasks, are increasingly

safetyarxiv-cs-ai
28 Apr 2026
Safety

FastOMOP: A Foundational Architecture for Reliable Agentic Real-World Evidence Generation on OMOP CDM data

DGX agent

arXiv:2604.24572v1 Announce Type: new Abstract: The Observational Medical Outcomes Partnership Common Data Model (OMOP CDM), maintained by the Observational Health Data Sciences and Informatics (OHDSI

safetyarxiv-cs-ai
28 Apr 2026
Safety

Pedestrians play chicken with an autonomous vehicle

DGX agent

arXiv:2604.24384v1 Announce Type: new Abstract: Automated vehicles (AVs) are commonly programmed to yield unconditionally to pedestrians in the interest of safety. However, this design choice can give

safetyarxiv-cs-ro
28 Apr 2026
Safety

Risk-Aware Robust Learning: Reducing Clinical Risk under Label Noise in Medical Image Classification

DGX agent

arXiv:2604.23875v1 Announce Type: cross Abstract: Noisy labels are a pervasive challenge in medical image classification, where annotation errors arise from inter-observer variability and diagnostic a

safetyarxiv-cs-ai
28 Apr 2026
Safety

How Many Visual Levers Drive Urban Perception? Interventional Counterfactuals via Multiple Localised Edits

DGX agent

arXiv:2604.22103v1 Announce Type: cross Abstract: Street-view perception models predict subjective attributes such as safety at scale, but remain correlational: they do not identify which localized vi

safetyarxiv-cs-cv
27 Apr 2026
Safety

If extremely violent criminals are not imprisoned, eventually they will murder innocent people

DGX agent

If extremely violent criminals are not imprisoned, eventually they will murder innocent people This bodega owner told ABC a year ago that he fears for his safety in NY Last night, he was kiIIed by a s

safetyelon-musk--x
27 Apr 2026
Safety

Are LLMs really more important than fire or electricity? “Honestly, a ton of what we’ve developed in my lifetime amounts to scaling up the d…

DGX agent

Are LLMs really more important than fire or electricity? “Honestly, a ton of what we’ve developed in my lifetime amounts to scaling up the delivery of information and entertainment and the frictionles

safetygary-marcus--x
24 Apr 2026
Safety

Survey on Evaluation of LLM-based Agents

DGX agent

arXiv:2503.16416v2 Announce Type: replace Abstract: LLM-based agents represent a paradigm shift in AI, enabling autonomous systems to plan, reason, and use tools while interacting with dynamic environ

safetyarxiv-cs-ai
24 Apr 2026
Safety

Aligning Human-AI-Interaction Trust for Mental Health Support: Survey and Position for Multi-Stakeholders

DGX agent

arXiv:2604.20166v1 Announce Type: new Abstract: Building trustworthy AI systems for mental health support is a shared priority across stakeholders from multiple disciplines. However, 'trustworthy' rem

safetyarxiv-cs-cl
23 Apr 2026
Safety

Interval POMDP Shielding for Imperfect-Perception Agents

DGX agent

arXiv:2604.20728v1 Announce Type: new Abstract: Autonomous systems that rely on learned perception can make unsafe decisions when sensor readings are misclassified. We study shielding for this setting

safetyarxiv-cs-ai
23 Apr 2026
Safety

ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System

DGX agent

arXiv:2604.18789v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is central to aligning Large Language Models (LLMs), yet it introduces a critical vulnerability: an im

safetyarxiv-cs-ai
22 Apr 2026
Safety

Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models

DGX agent

arXiv:2510.26782v3 Announce Type: replace-cross Abstract: A world model is an internal model that simulates how the world evolves. Given past observations and actions, it predicts the future physical

safetyarxiv-cs-ai
22 Apr 2026
Safety

Integrating Anomaly Detection into Agentic AI for Proactive Risk Management in Human Activity

DGX agent

arXiv:2604.19538v1 Announce Type: new Abstract: Agentic AI, with goal-directed, proactive, and autonomous decision-making capabilities, offers a compelling opportunity to address movement-related risk

safetyarxiv-cs-ai
22 Apr 2026
Safety

Large Language Models Exhibit Normative Conformity

DGX agent

arXiv:2604.19301v1 Announce Type: new Abstract: The conformity bias exhibited by large language models (LLMs) can pose a significant challenge to decision-making in LLM-based multi-agent systems (LLM-

safetyarxiv-cs-ai
22 Apr 2026
Safety

Jailbreaking Large Language Models with Morality Attacks

DGX agent

arXiv:2604.17053v1 Announce Type: new Abstract: Pluralism alignment with AI has the sophisticated and necessary goal of creating AI that can coexist with and serve morally multifaceted humanity. Resea

safetyarxiv-cs-cl
21 Apr 2026
Safety

SaFeR-Steer: Evolving Multi-Turn MLLMs via Synthetic Bootstrapping and Feedback Dynamics

DGX agent

arXiv:2604.16358v1 Announce Type: cross Abstract: MLLMs are increasingly deployed in multi-turn settings, where attackers can escalate unsafe intent through the evolving visual-text history and exploi

safetyarxiv-cs-cl
21 Apr 2026
Safety

STL-Based Motion Planning and Uncertainty-Aware Risk Analysis for Human-Robot Collaboration with a Multi-Rotor Aerial Vehicle

DGX agent

arXiv:2509.10692v3 Announce Type: replace Abstract: This paper presents a motion planning and risk analysis framework for enhancing human-robot collaboration with a Multi-Rotor Aerial Vehicle. The pro

safetyarxiv-cs-ro
21 Apr 2026
Safety

Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning

DGX agent

arXiv:2604.15705v1 Announce Type: new Abstract: Reinforcement Fine-Tuning (RFT) has established itself as a critical paradigm for the alignment of Multi-modal Large Language Models (MLLMs) with comple

safetyarxiv-cs-lg
20 Apr 2026
Safety

SeaAlert: Critical Information Extraction From Maritime Distress Communications with Large Language Models

DGX agent

arXiv:2604.14163v1 Announce Type: new Abstract: Maritime distress communications transmitted over very high frequency (VHF) radio are safety-critical voice messages used to report emergencies at sea.

safetyarxiv-cs-cl
17 Apr 2026
Safety

NHTSA autonomous vehicle crash data has been updated through March 15, 2026, for AVs including Tesla Robotaxi. This includes unsupervised Te…

DGX agent

NHTSA autonomous vehicle crash data has been updated through March 15, 2026, for AVs including Tesla Robotaxi. This includes unsupervised Tesla-driven robotaxis. • Waymo: 58 incidents • Zoox: 3 incide

safetyelon-musk--x
16 Apr 2026
Safety

AWARE: Adaptive Whole-body Active Rotating Control for Enhanced LiDAR-Inertial Odometry under Human-in-the-Loop Interaction

DGX agent

arXiv:2604.10598v1 Announce Type: new Abstract: Human-in-the-loop (HITL) UAV operation is essential in complex and safety-critical aerial surveying environments, where human operators provide navigati

safetyarxiv-cs-ro
14 Apr 2026
Safety

Prompt Injection as Role Confusion

DGX agent

arXiv:2603.12277v3 Announce Type: replace-cross Abstract: Language models remain vulnerable to prompt injection attacks despite extensive safety training. We trace this failure to role confusion: mode

safetyarxiv-cs-ai
14 Apr 2026
Safety

QShield: Securing Neural Networks Against Adversarial Attacks using Quantum Circuits

DGX agent

arXiv:2604.10933v1 Announce Type: cross Abstract: Deep neural networks remain highly vulnerable to adversarial perturbations, limiting their reliability in security- and safety-critical applications.

safetyarxiv-cs-ai
14 Apr 2026
Safety

SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering

DGX agent

arXiv:2508.11290v3 Announce Type: replace Abstract: LLMs increasingly exhibit over-refusal behavior, where safety mechanisms cause models to reject benign instructions that seemingly resemble harmful

safetyarxiv-cs-cl
14 Apr 2026
Safety

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs

DGX agent

arXiv:2604.08846v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have been shown to be vulnerable to malicious queries that can elicit unsafe responses. Recent work uses prom

safetyarxiv-cs-ai
13 Apr 2026
Safety

Large Reasoning Models Learn Better Alignment from Flawed Thinking

DGX agent

arXiv:2510.00938v2 Announce Type: replace Abstract: Large reasoning models (LRMs) 'think' by generating structured chain-of-thought (CoT) before producing a final answer, yet they still lack the abili

safetyarxiv-cs-lg
13 Apr 2026
Safety

AEROS: A Single-Agent Operating Architecture with Embodied Capability Modules

DGX agent

arXiv:2604.07039v1 Announce Type: cross Abstract: Robotic systems lack a principled abstraction for organizing intelligence, capabilities, and execution in a unified manner. Existing approaches either

safetyarxiv-cs-ai
10 Apr 2026
Safety

CMP: Robust Whole-Body Tracking for Loco-Manipulation via Competence Manifold Projection

DGX agent

arXiv:2604.07457v1 Announce Type: new Abstract: While decoupled control schemes for legged mobile manipulators have shown robustness, learning holistic whole-body control policies for tracking global

safetyarxiv-cs-ro
10 Apr 2026
Safety

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs

DGX agent

arXiv:2604.07655v1 Announce Type: cross Abstract: Hard-gated safety checkers often over-refuse and misalign with a vendor's model spec; prevailing taxonomies also neglect robustness and honesty, yield

safetyarxiv-cs-cl
10 Apr 2026
Safety

LLM-based Schema-Guided Extraction and Validation of Missing-Person Intelligence from Heterogeneous Data Sources

DGX agent

arXiv:2604.06571v1 Announce Type: cross Abstract: Missing-person and child-safety investigations rely on heterogeneous case documents, including structured forms, bulletin-style posters, and narrative

safetyarxiv-cs-ai
10 Apr 2026
Safety

VLMShield: Efficient and Robust Defense of Vision-Language Models against Malicious Prompts

DGX agent

arXiv:2604.06502v1 Announce Type: new Abstract: Vision-Language Models (VLMs) face significant safety vulnerabilities from malicious prompt attacks due to weakened alignment during visual integration.

safetyarxiv-cs-lg
10 Apr 2026
Safety

Guardrails at the gateway: Securing AI inference on GKE with Model Armor

DGX agent

Enterprises are rapidly moving AI workloads from experimentation to production on Google Kubernetes Engine (GKE), using its scalability to serve powerful inference endpoints. However, as these models

safetygoogle-cloud-ai
9 Apr 2026
← Previous
1…2223242526…299
Next →