AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs

DGX agent

arXiv:2604.12616v1 Announce Type: new Abstract: The rapid evolution of Vision-Language Models (VLMs) has catalyzed unprecedented capabilities in artificial intelligence; however, this continuous modal

safetyarxiv-cs-ai
15 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

How Transformers Learn to Plan via Multi-Token Prediction

DGX agent

arXiv:2604.11912v1 Announce Type: cross Abstract: While next-token prediction (NTP) has been the standard objective for training language models, it often struggles to capture global structure in reas

safetyarxiv-cs-ai
15 Apr 2026
Safety

Physics-Grounded Monocular Vehicle Distance Estimation Using Standardized License Plate Typography

DGX agent

arXiv:2604.12239v1 Announce Type: new Abstract: Accurate inter-vehicle distance estimation is a cornerstone of Advanced Driver Assistance Systems (ADAS) and autonomous driving. While LiDAR and radar p

safetyarxiv-cs-cv
15 Apr 2026
Safety

Probabilistic Feature Imputation and Uncertainty-Aware Multimodal Federated Aggregation

DGX agent

arXiv:2604.12970v1 Announce Type: cross Abstract: Multimodal federated learning enables privacy-preserving collaborative model training across healthcare institutions. However, a fundamental challenge

safetyarxiv-cs-cv
15 Apr 2026
Safety

Reasoning about Intent for Ambiguous Requests

DGX agent

arXiv:2511.10453v3 Announce Type: replace-cross Abstract: Large language models often respond to ambiguous requests by implicitly committing to one interpretation, frustrating users and creating safet

safetyarxiv-cs-ai
15 Apr 2026
Safety

Reliability-Guided Depth Fusion for Glare-Resilient Navigation Costmaps

DGX agent

arXiv:2604.12753v1 Announce Type: new Abstract: Specular glare on reflective floors and glass surfaces frequently corrupts RGB-D depth measurements, producing holes and spikes that accumulate as persi

safetyarxiv-cs-ro
15 Apr 2026
Safety

Ternary Logic Encodings of Temporal Behavior Trees with Application to Control Synthesis

DGX agent

arXiv:2604.12092v1 Announce Type: new Abstract: Behavior Trees (BTs) provide designers an intuitive graphical interface to construct long-horizon plans for autonomous systems. To ensure their correctn

safetyarxiv-cs-ro
15 Apr 2026
Safety

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment

DGX agent

arXiv:2604.12116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as tool-augmented agents capable of executing system-level operations. While existing benchmarks

safetyarxiv-cs-ai
15 Apr 2026
Safety

Uncertainty-Aware Image Classification In Biomedical Imaging Using Spectral-normalized Neural Gaussian Processes

DGX agent

arXiv:2602.02370v2 Announce Type: replace Abstract: Accurate histopathologic interpretation is key for clinical decision-making; however, current deep learning models for digital pathology are often o

safetyarxiv-cs-cv
15 Apr 2026
Safety

AI Integrity: A New Paradigm for Verifiable AI Governance

DGX agent

arXiv:2604.11065v1 Announce Type: new Abstract: AI systems increasingly shape high-stakes decisions in healthcare, law, defense, and education, yet existing governance paradigms -- AI Ethics, AI Safet

safetyarxiv-cs-ai
14 Apr 2026
Safety

Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward

DGX agent

arXiv:2604.09748v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is an emerging paradigm that significantly boosts a Large Language Model's (LLM's) reasoning abi

safetyarxiv-cs-ai
14 Apr 2026
Safety

Belief-Aware VLM Model for Human-like Reasoning

DGX agent

arXiv:2604.09686v1 Announce Type: new Abstract: Traditional neural network models for intent inference rely heavily on observable states and struggle to generalize across diverse tasks and dynamic env

safetyarxiv-cs-ai
14 Apr 2026
Safety

CAGenMol: Condition-Aware Diffusion Language Model for Goal-Directed Molecular Generation

DGX agent

arXiv:2604.11483v1 Announce Type: new Abstract: Goal-directed molecular generation requires satisfying heterogeneous constraints such as protein--ligand compatibility and multi-objective drug-like pro

safetyarxiv-cs-lg
14 Apr 2026
Safety

CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation

DGX agent

arXiv:2604.09746v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed as autonomous agents, understanding how strategic behavior emerges in multi-agent environmen

safetyarxiv-cs-ai
14 Apr 2026
Safety

Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails

DGX agent

arXiv:2603.18280v2 Announce Type: replace-cross Abstract: Current alignment evaluation mostly measures whether models encode dangerous concepts and whether they refuse harmful requests. Both miss the

safetyarxiv-cs-ai
14 Apr 2026
Safety

Distributionally Robust PAC-Bayesian Control

DGX agent

arXiv:2604.10588v1 Announce Type: new Abstract: We present a distributionally robust PAC-Bayesian framework for certifying the performance of learning-based finite-horizon controllers. While existing

safetyarxiv-cs-lg
14 Apr 2026
Safety

Enabling and Inhibitory Pathways of Students' AI Use Concealment Intention in Higher Education: Evidence from SEM and fsQCA

DGX agent

arXiv:2604.10978v1 Announce Type: cross Abstract: This study investigates students' AI use concealment intention in higher education by integrating the cognition-affect-conation (CAC) framework with a

safetyarxiv-cs-ai
14 Apr 2026
Safety

Explainability and Certification of AI-Generated Educational Assessments

DGX agent

arXiv:2604.09622v1 Announce Type: cross Abstract: The rapid adoption of generative artificial intelligence (AI) in educational assessment has created new opportunities for scalable item creation, pers

safetyarxiv-cs-ai
14 Apr 2026
Safety

Explainable Planning for Hybrid Systems

DGX agent

arXiv:2604.09578v1 Announce Type: new Abstract: The recent advancement in artificial intelligence (AI) technologies facilitates a paradigm shift toward automation. Autonomous systems are fully or part

safetyarxiv-cs-ai
14 Apr 2026
Safety

FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning

DGX agent

arXiv:2604.10693v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting has improved LLM reasoning, but models often generate explanations that appear coherent while containing unfaithful int

safetyarxiv-cs-ai
14 Apr 2026
Safety

Fake-HR1: Rethinking Reasoning of Vision Language Model for Synthetic Image Detection

DGX agent

arXiv:2602.10042v3 Announce Type: replace-cross Abstract: Recent studies have demonstrated that incorporating Chain-of-Thought (CoT) reasoning into the detection process can enhance a model's ability

safetyarxiv-cs-ai
14 Apr 2026
Safety

From Answers to Arguments: Toward Trustworthy Clinical Diagnostic Reasoning with Toulmin-Guided Curriculum Goal-Conditioned Learning

DGX agent

arXiv:2604.11137v1 Announce Type: new Abstract: The integration of Large Language Models (LLMs) into clinical decision support is critically obstructed by their opaque and often unreliable reasoning.

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

GAMBIT: A Gamified Jailbreak Framework for Multimodal Large Language Models

DGX agent

arXiv:2601.03416v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have become widely deployed, yet their safety alignment remains fragile under adversarial inputs. Previous

model-releasesarxiv-cs-cv
14 Apr 2026
Safety

Hubble: An LLM-Driven Agentic Framework for Safe and Automated Alpha Factor Discovery

DGX agent

arXiv:2604.09601v1 Announce Type: new Abstract: Discovering predictive alpha factors in quantitative finance remains a formidable challenge due to the vast combinatorial search space and inherently lo

safetyarxiv-cs-ai
14 Apr 2026
Safety

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs

DGX agent

arXiv:2604.10403v1 Announce Type: new Abstract: We address jailbreaks, backdoors, and unlearning for large language models (LLMs). Unlike prior work, which trains LLMs based on their actions when give

safetyarxiv-cs-lg
14 Apr 2026
Safety

LLM-as-Judge on a Budget

DGX agent

arXiv:2602.15481v2 Announce Type: replace Abstract: LLM-as-a-judge has emerged as a cornerstone technique for evaluating large language models by leveraging LLM reasoning to score prompt-response pair

safetyarxiv-cs-lg
14 Apr 2026
Safety

Minimal Embodiment Enables Efficient Learning of Number Concepts in Robot

DGX agent

arXiv:2604.11373v1 Announce Type: cross Abstract: Robots are increasingly entering human-interactive scenarios that require understanding of quantity. How intelligent systems acquire abstract numerica

safetyarxiv-cs-ai
14 Apr 2026
Safety

MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

DGX agent

arXiv:2308.12067v3 Announce Type: replace-cross Abstract: Multimodal large language models are typically trained in two stages: first pre-training on image-text pairs, and then fine-tuning using super

safetyarxiv-cs-ai
14 Apr 2026
Safety

MoRI: Mixture of RL and IL Experts for Long-Horizon Manipulation Tasks

DGX agent

arXiv:2604.10165v1 Announce Type: new Abstract: Reinforcement Learning (RL) and Imitation Learning (IL) are the standard frameworks for policy acquisition in manipulation. While IL offers efficient po

safetyarxiv-cs-ro
14 Apr 2026
Safety

Principles Do Not Apply Themselves: A Hermeneutic Perspective on AI Alignment

DGX agent

arXiv:2604.10673v1 Announce Type: new Abstract: AI alignment is often framed as the task of ensuring that an AI system follows a set of stated principles or human preferences, but general principles r

safetyarxiv-cs-ai
14 Apr 2026
Safety

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging

DGX agent

arXiv:2604.11399v1 Announce Type: cross Abstract: Multimodal adaptation equips large language models (LLMs) with perceptual capabilities, but often weakens the reasoning ability inherited from languag

safetyarxiv-cs-cl
14 Apr 2026
Safety

Resilient Write: A Six-Layer Durable Write Surface for LLM Coding Agents

DGX agent

arXiv:2604.10842v1 Announce Type: cross Abstract: LLM-powered coding agents increasingly rely on tool-use protocols such as the Model Context Protocol~(MCP) to read and write files on a developer's wo

safetyarxiv-cs-ai
14 Apr 2026
Safety

Spatiotemporal-Aware Bit-Flip Injection on DNN-based Advanced Driver Assistance Systems (extended version)

DGX agent

arXiv:2604.03753v2 Announce Type: replace-cross Abstract: Modern advanced driver assistance systems (ADAS) rely on deep neural networks (DNNs) for perception and planning. Since DNNs' parameters resid

safetyarxiv-cs-lg
14 Apr 2026
Safety

Speaking to No One: Ontological Dissonance and the Double Bind of Conversational AI

DGX agent

arXiv:2604.10833v1 Announce Type: cross Abstract: Recent reports indicate that sustained interaction with conversational artificial intelligence (AI) systems can, in a small subset of users, contribut

safetyarxiv-cs-ai
14 Apr 2026
Safety

VLMaterial: Vision-Language Model-Based Camera-Radar Fusion for Physics-Grounded Material Identification

DGX agent

arXiv:2604.11671v1 Announce Type: cross Abstract: Accurate material recognition is a fundamental capability for intelligent perception systems to interact safely and effectively with the physical worl

safetyarxiv-cs-ro
14 Apr 2026
Safety

What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data

DGX agent

arXiv:2510.26202v2 Announce Type: replace-cross Abstract: Human feedback can alter language models in unpredictable and undesirable ways, as practitioners lack a clear understanding of what feedback d

safetyarxiv-cs-ai
14 Apr 2026
Safety

EGLOCE: Training-Free Energy-Guided Latent Optimization for Concept Erasure

DGX agent

arXiv:2604.09405v1 Announce Type: new Abstract: As text-to-image diffusion models grow increasingly prevalent, the ability to remove specific concepts-mostly explicit content and many copyrighted char

safetyarxiv-cs-cv
13 Apr 2026
Safety

EthicMind: A Risk-Aware Framework for Ethical-Emotional Alignment in Multi-Turn Dialogue

DGX agent

arXiv:2604.09265v1 Announce Type: new Abstract: Intelligent dialogue systems are increasingly deployed in emotionally and ethically sensitive settings, where failures in either emotional attunement or

safetyarxiv-cs-cl
13 Apr 2026
Safety

Hierarchical Alignment: Enforcing Hierarchical Instruction-Following in LLMs through Logical Consistency

DGX agent

arXiv:2604.09075v1 Announce Type: new Abstract: Large language models increasingly operate under multiple instructions from heterogeneous sources with different authority levels, including system poli

safetyarxiv-cs-cl
13 Apr 2026
Safety

Learning Vision-Language-Action World Models for Autonomous Driving

DGX agent

arXiv:2604.09059v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have recently achieved notable progress in end-to-end autonomous driving by integrating perception, reasoning, and

safetyarxiv-cs-ai
13 Apr 2026
Safety

Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection

DGX agent

arXiv:2604.09024v1 Announce Type: cross Abstract: Multi-modal large language models (MLLMs) have emerged as powerful tools for analyzing Internet-scale image data, offering significant benefits but al

safetyarxiv-cs-ai
13 Apr 2026
Safety

LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving

DGX agent

arXiv:2604.08719v1 Announce Type: cross Abstract: Recent years have seen remarkable progress in autonomous driving, yet generalization to long-tail and open-world scenarios remains a major bottleneck

safetyarxiv-cs-ai
13 Apr 2026
Safety

Many Preferences, Few Policies: Towards Scalable Language Model Personalization

DGX agent

arXiv:2604.04144v2 Announce Type: replace-cross Abstract: The holy grail of LLM personalization is a single LLM for each user, perfectly aligned with that user's preferences. However, maintaining a se

safetyarxiv-cs-ai
13 Apr 2026
Safety

Mosaic: Multimodal Jailbreak against Closed-Source VLMs via Multi-View Ensemble Optimization

DGX agent

arXiv:2604.09253v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are powerful but remain vulnerable to multimodal jailbreak attacks. Existing attacks mainly rely on either explicit visu

safetyarxiv-cs-ai
13 Apr 2026
Safety

Scheming in the wild: detecting real-world AI scheming incidents with open-source intelligence

DGX agent

arXiv:2604.09104v1 Announce Type: cross Abstract: Scheming, the covert pursuit of misaligned goals by AI systems, represents a potentially catastrophic risk, yet scheming research suffers from signifi

safetyarxiv-cs-ai
13 Apr 2026
Safety

Verbalizing LLMs' assumptions to explain and control sycophancy

DGX agent

arXiv:2604.03058v2 Announce Type: replace-cross Abstract: LLMs can be socially sycophantic, affirming users when they ask questions like 'am I in the wrong?' rather than providing genuine assessment.

safetyarxiv-cs-ai
13 Apr 2026
Safety

A Unified Multi-Layer Framework for Skill Acquisition from Imperfect Human Demonstrations

DGX agent

arXiv:2604.08341v1 Announce Type: new Abstract: Current Human-Robot Interaction (HRI) systems for skill teaching are fragmented, and existing approaches in the literature do not offer a cohesive frame

safetyarxiv-cs-ro
10 Apr 2026
Safety

Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge

DGX agent

arXiv:2510.18196v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are commonly used as evaluators in various applications, but the reliability of the outcomes remains a challenge.

safetyarxiv-cs-ai
10 Apr 2026
← Previous
1…5556575859…257
Next →