AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

Steered LLM Activations are Non-Surjective

DGX agent

arXiv:2604.09839v1 Announce Type: new Abstract: Activation steering is a popular white-box control technique that modifies model activations to elicit an abstract change in output behavior. It has als

safetyarxiv-cs-ai
14 Apr 2026
Safety

Towards Adaptive Open-Set Object Detection via Category-Level Collaboration Knowledge Mining

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2604.11195v1 Announce Type: cross Abstract: Existing object detectors often struggle to generalize across domains while adapting to emerging novel categories. Adaptive open-set object detection

safetyarxiv-cs-ai
14 Apr 2026
Safety

Trajectory-based actuator identification via differentiable simulation

DGX agent

arXiv:2604.10351v1 Announce Type: new Abstract: Accurate actuation models are critical for bridging the gap between simulation and real robot behavior, yet obtaining high-fidelity actuator dynamics ty

safetyarxiv-cs-ro
14 Apr 2026
Safety

Weird Generalization is Weirdly Brittle

DGX agent

arXiv:2604.10022v1 Announce Type: new Abstract: Weird generalization is a phenomenon in which models fine-tuned on data from a narrow domain (e.g. insecure code) develop surprising traits that manifes

safetyarxiv-cs-cl
14 Apr 2026
Safety

Better Eyes, Better Thoughts: Why Vision Chain-of-Thought Fails in Medicine

DGX agent

arXiv:2603.06665v2 Announce Type: replace-cross Abstract: Large vision-language models (VLMs) often benefit from chain-of-thought (CoT) prompting in general domains, yet its efficacy in medical vision

safetyarxiv-cs-ai
13 Apr 2026
Safety

CausalVAD: De-confounding End-to-End Autonomous Driving via Causal Intervention

DGX agent

arXiv:2603.18561v2 Announce Type: replace Abstract: Planning-oriented end-to-end driving models show great promise, yet they fundamentally learn statistical correlations instead of true causal relatio

safetyarxiv-cs-cv
13 Apr 2026
Safety

ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion

DGX agent

arXiv:2604.09450v1 Announce Type: cross Abstract: Chest X-ray report generation (CXR-RG) has the potential to substantially alleviate radiologists' workload. However, conventional autoregressive visio

safetyarxiv-cs-ai
13 Apr 2026
Safety

Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism

DGX agent

arXiv:2604.09544v1 Announce Type: cross Abstract: Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks routinely

safetyarxiv-cs-ai
13 Apr 2026
Safety

On the Spectral Geometry of Cross-Modal Representations: A Functional Map Diagnostic for Multimodal Alignment

DGX agent

arXiv:2604.08579v1 Announce Type: cross Abstract: We study cross-modal alignment between independently pretrained vision (DINOv2) and language (all-MiniLM-L6-v2) encoders using the functional map fram

safetyarxiv-cs-ai
13 Apr 2026
Safety

Semantic Intent Fragmentation: A Single-Shot Compositional Attack on Multi-Agent AI Pipelines

DGX agent

arXiv:2604.08608v1 Announce Type: cross Abstract: We introduce Semantic Intent Fragmentation (SIF), an attack class against LLM orchestration systems where a single, legitimately phrased request cause

safetyarxiv-cs-ai
13 Apr 2026
Safety

Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection

DGX agent

arXiv:2604.07831v1 Announce Type: cross Abstract: Existing red-teaming studies on GUI agents have important limitations. Adversarial perturbations typically require white-box access, which is unavaila

safetyarxiv-cs-cl
10 Apr 2026
Safety

Learning Without Losing Identity: Capability Evolution for Embodied Agents

DGX agent

arXiv:2604.07799v1 Announce Type: new Abstract: Embodied agents are expected to operate persistently in dynamic physical environments, continuously acquiring new capabilities over time. Existing appro

safetyarxiv-cs-ro
10 Apr 2026
Safety

Machine Unlearning in the Era of Quantum Machine Learning: An Empirical Study

DGX agent

arXiv:2512.19253v4 Announce Type: replace-cross Abstract: We present the first empirical study of machine unlearning (MU) in hybrid quantum-classical neural networks. While MU has been extensively exp

safetyarxiv-cs-ai
10 Apr 2026
Safety

SymptomWise: A Deterministic Reasoning Layer for Reliable and Efficient AI Systems

DGX agent

arXiv:2604.06375v1 Announce Type: new Abstract: AI-driven symptom analysis systems face persistent challenges in reliability, interpretability, and hallucination. End-to-end generative approaches ofte

safetyarxiv-cs-ai
10 Apr 2026
Safety

AI Guardrail Survival under Single-Cycle Agentic Self-Summarization

DGX agent

arXiv:2608.11392v1 Announce Type: cross Abstract: Long-running agents periodically compact their context, replacing the transcript with a model-generated summary.Recent work shows that dropping a stan

safetyarxiv-cs-ai
13 Aug 2026
Safety

Clustered Randomized Smoothing for Stochastic Prediction Functions

DGX agent

arXiv:2608.12037v1 Announce Type: new Abstract: Modern stochastic predictors can model rich, multi-modal outcome distributions. However, this expressive power comes with challenges in ensuring robust

safetyarxiv-cs-lg
13 Aug 2026
Safety

Confidence Calibration of Deep Learning Systems

DGX agent

arXiv:2608.12100v1 Announce Type: cross Abstract: In high-stakes applications, reliable confidence estimates are as important as the predictions themselves. Confidence calibration ensures that predict

safetyarxiv-cs-ai
13 Aug 2026
Safety

Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models

DGX agent

arXiv:2608.12304v1 Announce Type: new Abstract: Dynamic Master Logic (DML) provides a hierarchical framework for representing system behavior by linking functional objectives to underlying structural

safetyarxiv-cs-ai
13 Aug 2026
Safety

Contextual Quality-Diversity Evolutionary Reinforcement Learning for HVAC Control in Tropical Commercial Buildings

DGX agent

arXiv:2608.11324v1 Announce Type: cross Abstract: This paper proposes a contextual quality-diversity evolutionary reinforcement-learning controller, CQD-ERL, for the supervisory control of a tropical,

safetyarxiv-cs-ai
13 Aug 2026
Safety

Do Not Forget the Obvious - RISC: A Risk-Informed Slice-Coverage Protocol for Safe Autonomous Driving

DGX agent

arXiv:2608.12051v1 Announce Type: new Abstract: Aggregate metrics may not fully reflect performance in insufficiently examined high-risk driving conditions. We propose RISC (Risk-Informed Slice Covera

safetyarxiv-cs-cv
13 Aug 2026
Safety

Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion

DGX agent

arXiv:2608.12083v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) achieve strong predictive performance on graph-structured data across domains such as chemistry, biology, and network ana

safetyarxiv-cs-ai
13 Aug 2026
Safety

Forecasting Side Effects of Activation Steering

DGX agent

arXiv:2608.11227v1 Announce Type: new Abstract: Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes without retr

safetyarxiv-cs-ai
13 Aug 2026
Safety

GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

DGX agent

arXiv:2608.11787v1 Announce Type: cross Abstract: Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound judgment,

safetyarxiv-cs-ai
13 Aug 2026
Safety

Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning

DGX agent

arXiv:2608.11658v1 Announce Type: cross Abstract: Many reinforcement learning systems, from fleet management to traffic signal control, must serve an objective that changes dynamically after deploymen

safetyarxiv-cs-ai
13 Aug 2026
Safety

Multi-Agent Embodied Autonomous Driving: From V2X Information Exchange to Shared World Models

DGX agent

arXiv:2606.13840v2 Announce Type: replace-cross Abstract: Autonomous driving is shifting from isolated vehicle intelligence toward multi-agent embodied systems that share perception, infer intent, and

safetyarxiv-cs-cv
13 Aug 2026
Safety

On the Definition of Intelligence

DGX agent

arXiv:2507.22423v3 Announce Type: replace Abstract: To engineer AGI, we should first capture the essence of intelligence in a species-agnostic form that can be evaluated, while being sufficiently gene

safetyarxiv-cs-ai
13 Aug 2026
Safety

TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement

DGX agent

arXiv:2608.11951v1 Announce Type: cross Abstract: Extreme events in air transport, such as severe arrival delays and abnormal air times, cause cascading network disruptions with substantial operationa

safetyarxiv-cs-ai
13 Aug 2026
Safety

Topology-Aware Query Selection for Surgical Instrument Instance Segmentation

DGX agent

arXiv:2608.11607v1 Announce Type: new Abstract: Accurate foreground masks can still form an incorrect surgical-instrument instance set: duplicate, fragmented, merged, missed, or empty-frame prediction

safetyarxiv-cs-cv
13 Aug 2026
Safety

Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits

DGX agent

arXiv:2608.11410v1 Announce Type: new Abstract: Offline reinforcement learning (RL) offers considerable promise for optimizing ICU treatment decisions, yet standard evaluation metrics Mean Squared Err

safetyarxiv-cs-lg
13 Aug 2026
Safety

A Convolutional Layer Activation Dimensionality Reduction for Out-of-Distribution and Adversarial Attack Detection Methods

DGX agent

arXiv:2608.10203v1 Announce Type: new Abstract: Despite the success of convolutional neural networks in image classification tasks and their general application in multi-modal models, their susceptibi

safetyarxiv-cs-cv
12 Aug 2026
Safety

A Neural Network Based Teleoperation for Remote Controlled Vehicles

DGX agent

arXiv:2608.10367v1 Announce Type: new Abstract: Direct teleoperation of vehicles faces critical technical bottlenecks: communication latency and the operator's inability to physically perceive unmodel

safetyarxiv-cs-ro
12 Aug 2026
Safety

ConfTriage: A Calibration-Aware LLM Triage Framework for Pulmonary Nodule Malignancy with Selective Specialist Deferral

DGX agent

arXiv:2608.10885v1 Announce Type: new Abstract: Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imagin

safetyarxiv-cs-cv
12 Aug 2026
Safety

Data Attribution of Emergent Misalignment with Persona Features

DGX agent

arXiv:2608.11025v1 Announce Type: new Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leadi

safetyarxiv-cs-cl
12 Aug 2026
Safety

Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving

DGX agent

arXiv:2608.10386v1 Announce Type: new Abstract: Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias. While world mod

safetyarxiv-cs-lg
12 Aug 2026
Local Ai

Easy3D-Labels: Supervising Semantic Occupancy Estimation with 3D Pseudo-Labels for Automotive Perception

DGX agent

arXiv:2509.26087v5 Announce Type: replace Abstract: In perception for automated vehicles, safety is critical not only for the driver but also for other agents in the scene, particularly vulnerable roa

local-aiarxiv-cs-cv
12 Aug 2026
Safety

Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement

DGX agent

arXiv:2608.10339v1 Announce Type: cross Abstract: Hospital quality improvement (QI) programs routinely face multiple candidate interventions to optimize hospital flow, but existing methods struggle to

safetyarxiv-cs-ai
12 Aug 2026
Safety

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

DGX agent

arXiv:2608.11171v1 Announce Type: cross Abstract: The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings pap

safetyarxiv-cs-ai
12 Aug 2026
Safety

LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations

DGX agent

arXiv:2602.09924v4 Announce Type: replace-cross Abstract: Running LLMs with extended reasoning on every problem is expensive, but determining which inputs actually require additional compute remains c

safetyarxiv-cs-ai
12 Aug 2026
Model Releases

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

DGX agent

arXiv:2608.10669v1 Announce Type: new Abstract: Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interact

model-releasesarxiv-cs-ai
12 Aug 2026
Safety

Similarity Gates Approve Reversals: A Validity Audit of Embedding-Cosine Thresholds in Agent Systems

DGX agent

arXiv:2608.10216v1 Announce Type: cross Abstract: Agent frameworks ship quality gates that compare text blocks by embedding-cosine similarity and decide at a fixed cutoff. Deduplication filters, seman

safetyarxiv-cs-ai
12 Aug 2026
Safety

Smart Enough to Go Extinct? An Evolutionary Challenge to the Value of General Intelligence and Its Ethical Implications for AGI

DGX agent

arXiv:2608.10730v1 Announce Type: cross Abstract: The pursuit of artificial general intelligence (AGI) rests on a seemingly self-evident premise: that general intelligence, the kind of flexible, domai

safetyarxiv-cs-ai
12 Aug 2026
Safety

Toward a Theory of Value in AI Alignment

DGX agent

arXiv:2608.10327v1 Announce Type: new Abstract: Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms

safetyarxiv-cs-ai
12 Aug 2026
Safety

UPAIR: Diagnosing Reasoning States via Uncertainty-Progress Alignment for Selective Intervention

DGX agent

arXiv:2607.17188v2 Announce Type: replace Abstract: While test-time scaling improves the problem-solving ability of large reasoning models (LRMs) through additional inference-time computation, it can

safetyarxiv-cs-ai
12 Aug 2026
Safety

VIDS-Seg: Towards Reliable Uncertainty Quantification in Pediatric Cardiac Ultrasound Segmentation

DGX agent

arXiv:2608.10903v1 Announce Type: new Abstract: Reliable clinical deployment of machine learning requires models that know when they are likely to fail, particularly for subgroups underrepresented in

safetyarxiv-cs-cv
12 Aug 2026
Safety

A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning

DGX agent

arXiv:2608.08158v1 Announce Type: new Abstract: Sparse, delayed, and weakly informative rewards remain central obstacles to efficient reinforcement learning. Reward shaping addresses these limitations

safetyarxiv-cs-ai
11 Aug 2026
Safety

Adaptive Sequential Test Planning for Multi-Mechanism Reliability Qualification via Bayesian Monte Carlo Tree Search

DGX agent

arXiv:2608.09622v1 Announce Type: new Abstract: Reliability qualification of advanced semiconductor devices requires sequential stress decisions that balance characterization objectives against multip

safetyarxiv-cs-ai
11 Aug 2026
Safety

Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy

DGX agent

arXiv:2608.09857v1 Announce Type: cross Abstract: Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on

safetyarxiv-cs-ai
11 Aug 2026
Safety

An Explainable GNN Framework for Component-Level Anomaly Diagnosis

DGX agent

arXiv:2608.09246v1 Announce Type: new Abstract: Industrial processes are complex systems composed of multiple interacting sensors that generate multivariate time series (MTS). Detecting anomalies in s

safetyarxiv-cs-ai
11 Aug 2026
← Previous
1…3637383940…257
Next →