AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,357 results
15 Apr 2026

How Transformers Learn to Plan via Multi-Token Prediction

SafetyDGX agent

arXiv:2604.11912v1 Announce Type: cross Abstract: While next-token prediction (NTP) has been the standard objective for training language models, it often struggles to capture global structure in reas

Physics-Grounded Monocular Vehicle Distance Estimation Using Standardized License Plate Typography

SafetyDGX agent

arXiv:2604.12239v1 Announce Type: new Abstract: Accurate inter-vehicle distance estimation is a cornerstone of Advanced Driver Assistance Systems (ADAS) and autonomous driving. While LiDAR and radar p

🇧🇪 Positive news for FSD Supervised in Belgium! I just received an official response from the cabinet of @MDiependaele , Minister-Presiden…

SafetyDGX agent

🇧🇪 Positive news for FSD Supervised in Belgium! I just received an official response from the cabinet of @MDiependaele , Minister-President of the Flemish Government. A few days ago, the Dutch vehicle

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Probabilistic Feature Imputation and Uncertainty-Aware Multimodal Federated Aggregation

SafetyDGX agent

arXiv:2604.12970v1 Announce Type: cross Abstract: Multimodal federated learning enables privacy-preserving collaborative model training across healthcare institutions. However, a fundamental challenge

Reasoning about Intent for Ambiguous Requests

SafetyDGX agent

arXiv:2511.10453v3 Announce Type: replace-cross Abstract: Large language models often respond to ambiguous requests by implicitly committing to one interpretation, frustrating users and creating safet

Reliability-Guided Depth Fusion for Glare-Resilient Navigation Costmaps

SafetyDGX agent

arXiv:2604.12753v1 Announce Type: new Abstract: Specular glare on reflective floors and glass surfaces frequently corrupts RGB-D depth measurements, producing holes and spikes that accumulate as persi

Ternary Logic Encodings of Temporal Behavior Trees with Application to Control Synthesis

SafetyDGX agent

arXiv:2604.12092v1 Announce Type: new Abstract: Behavior Trees (BTs) provide designers an intuitive graphical interface to construct long-horizon plans for autonomous systems. To ensure their correctn

The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment

SafetyDGX agent

arXiv:2604.12116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed as tool-augmented agents capable of executing system-level operations. While existing benchmarks

This is what a black politician in South Africa said… “We will kill white women, we will kill white children, and we will even kill your pet…

SafetyDGX agent

This is what a black politician in South Africa said… “We will kill white women, we will kill white children, and we will even kill your pets' Violent language entering politics is a serious warning s

Uncertainty-Aware Image Classification In Biomedical Imaging Using Spectral-normalized Neural Gaussian Processes

SafetyDGX agent

arXiv:2602.02370v2 Announce Type: replace Abstract: Accurate histopathologic interpretation is key for clinical decision-making; however, current deep learning models for digital pathology are often o

14 Apr 2026

AI Integrity: A New Paradigm for Verifiable AI Governance

SafetyDGX agent

arXiv:2604.11065v1 Announce Type: new Abstract: AI systems increasingly shape high-stakes decisions in healthcare, law, defense, and education, yet existing governance paradigms -- AI Ethics, AI Safet

Awesome work by @jiaxinwen22, @liangqiu_1994, Joe Benton, and @janhkirchner! For more details, check out the blog post 👇 https://anthropic.…

SafetyDGX agent

Jan Leike praised collaborative work by researchers Jiaxin Wen, Liang Qiu, Joe Benton, and Jan Kirchner, directing followers to an Anthropic blog post for further details. The post appears to highligh

Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward

SafetyDGX agent

arXiv:2604.09748v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is an emerging paradigm that significantly boosts a Large Language Model's (LLM's) reasoning abi

Belief-Aware VLM Model for Human-like Reasoning

SafetyDGX agent

arXiv:2604.09686v1 Announce Type: new Abstract: Traditional neural network models for intent inference rely heavily on observable states and struggle to generalize across diverse tasks and dynamic env

CAGenMol: Condition-Aware Diffusion Language Model for Goal-Directed Molecular Generation

SafetyDGX agent

arXiv:2604.11483v1 Announce Type: new Abstract: Goal-directed molecular generation requires satisfying heterogeneous constraints such as protein--ligand compatibility and multi-objective drug-like pro

CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation

SafetyDGX agent

arXiv:2604.09746v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed as autonomous agents, understanding how strategic behavior emerges in multi-agent environmen

Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails

SafetyDGX agent

arXiv:2603.18280v2 Announce Type: replace-cross Abstract: Current alignment evaluation mostly measures whether models encode dangerous concepts and whether they refuse harmful requests. Both miss the

Distributionally Robust PAC-Bayesian Control

SafetyDGX agent

arXiv:2604.10588v1 Announce Type: new Abstract: We present a distributionally robust PAC-Bayesian framework for certifying the performance of learning-based finite-horizon controllers. While existing

Enabling and Inhibitory Pathways of Students' AI Use Concealment Intention in Higher Education: Evidence from SEM and fsQCA

SafetyDGX agent

arXiv:2604.10978v1 Announce Type: cross Abstract: This study investigates students' AI use concealment intention in higher education by integrating the cognition-affect-conation (CAC) framework with a

Explainability and Certification of AI-Generated Educational Assessments

SafetyDGX agent

arXiv:2604.09622v1 Announce Type: cross Abstract: The rapid adoption of generative artificial intelligence (AI) in educational assessment has created new opportunities for scalable item creation, pers

Explainable Planning for Hybrid Systems

SafetyDGX agent

arXiv:2604.09578v1 Announce Type: new Abstract: The recent advancement in artificial intelligence (AI) technologies facilitates a paradigm shift toward automation. Autonomous systems are fully or part

FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning

SafetyDGX agent

arXiv:2604.10693v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting has improved LLM reasoning, but models often generate explanations that appear coherent while containing unfaithful int

Fake-HR1: Rethinking Reasoning of Vision Language Model for Synthetic Image Detection

SafetyDGX agent

arXiv:2602.10042v3 Announce Type: replace-cross Abstract: Recent studies have demonstrated that incorporating Chain-of-Thought (CoT) reasoning into the detection process can enhance a model's ability

From Answers to Arguments: Toward Trustworthy Clinical Diagnostic Reasoning with Toulmin-Guided Curriculum Goal-Conditioned Learning

SafetyDGX agent

arXiv:2604.11137v1 Announce Type: new Abstract: The integration of Large Language Models (LLMs) into clinical decision support is critically obstructed by their opaque and often unreliable reasoning.

GAMBIT: A Gamified Jailbreak Framework for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2601.03416v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have become widely deployed, yet their safety alignment remains fragile under adversarial inputs. Previous

Here's my first drive on Tesla FSD V14.3.1. It feels polished vs 14.3 and it my opinion is ready for wide release. • With 14.3.1, you can no…

SafetyDGX agent

Here's my first drive on Tesla FSD V14.3.1. It feels polished vs 14.3 and it my opinion is ready for wide release. • With 14.3.1, you can now tap on the new 'P' parking icon and bring up your differen

Hubble: An LLM-Driven Agentic Framework for Safe and Automated Alpha Factor Discovery

SafetyDGX agent

arXiv:2604.09601v1 Announce Type: new Abstract: Discovering predictive alpha factors in quantitative finance remains a formidable challenge due to the vast combinatorial search space and inherently lo

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs

SafetyDGX agent

arXiv:2604.10403v1 Announce Type: new Abstract: We address jailbreaks, backdoors, and unlearning for large language models (LLMs). Unlike prior work, which trains LLMs based on their actions when give

LLM-as-Judge on a Budget

SafetyDGX agent

arXiv:2602.15481v2 Announce Type: replace Abstract: LLM-as-a-judge has emerged as a cornerstone technique for evaluating large language models by leveraging LLM reasoning to score prompt-response pair

Minimal Embodiment Enables Efficient Learning of Number Concepts in Robot

SafetyDGX agent

arXiv:2604.11373v1 Announce Type: cross Abstract: Robots are increasingly entering human-interactive scenarios that require understanding of quantity. How intelligent systems acquire abstract numerica

MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

SafetyDGX agent

arXiv:2308.12067v3 Announce Type: replace-cross Abstract: Multimodal large language models are typically trained in two stages: first pre-training on image-text pairs, and then fine-tuning using super

MoRI: Mixture of RL and IL Experts for Long-Horizon Manipulation Tasks

SafetyDGX agent

arXiv:2604.10165v1 Announce Type: new Abstract: Reinforcement Learning (RL) and Imitation Learning (IL) are the standard frameworks for policy acquisition in manipulation. While IL offers efficient po

🦔OpenAI is backing an Illinois state bill that would shield AI labs from liability in cases where their models cause mass casualties or lar…

SafetyDGX agent

🦔OpenAI is backing an Illinois state bill that would shield AI labs from liability in cases where their models cause mass casualties or large-scale financial disasters, defined as death or serious inj

Principles Do Not Apply Themselves: A Hermeneutic Perspective on AI Alignment

SafetyDGX agent

arXiv:2604.10673v1 Announce Type: new Abstract: AI alignment is often framed as the task of ensuring that an AI system follows a set of stated principles or human preferences, but general principles r

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging

SafetyDGX agent

arXiv:2604.11399v1 Announce Type: cross Abstract: Multimodal adaptation equips large language models (LLMs) with perceptual capabilities, but often weakens the reasoning ability inherited from languag

Resilient Write: A Six-Layer Durable Write Surface for LLM Coding Agents

SafetyDGX agent

arXiv:2604.10842v1 Announce Type: cross Abstract: LLM-powered coding agents increasingly rely on tool-use protocols such as the Model Context Protocol~(MCP) to read and write files on a developer's wo

Spatiotemporal-Aware Bit-Flip Injection on DNN-based Advanced Driver Assistance Systems (extended version)

SafetyDGX agent

arXiv:2604.03753v2 Announce Type: replace-cross Abstract: Modern advanced driver assistance systems (ADAS) rely on deep neural networks (DNNs) for perception and planning. Since DNNs' parameters resid

Speaking to No One: Ontological Dissonance and the Double Bind of Conversational AI

SafetyDGX agent

arXiv:2604.10833v1 Announce Type: cross Abstract: Recent reports indicate that sustained interaction with conversational artificial intelligence (AI) systems can, in a small subset of users, contribut

VLMaterial: Vision-Language Model-Based Camera-Radar Fusion for Physics-Grounded Material Identification

SafetyDGX agent

arXiv:2604.11671v1 Announce Type: cross Abstract: Accurate material recognition is a fundamental capability for intelligent perception systems to interact safely and effectively with the physical worl

What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data

SafetyDGX agent

arXiv:2510.26202v2 Announce Type: replace-cross Abstract: Human feedback can alter language models in unpredictable and undesirable ways, as practitioners lack a clear understanding of what feedback d

13 Apr 2026

A new standard for research: How UC Riverside is securing the path to federal grants with Google Public Sector

SafetyDGX agent

At the University of California, Riverside (UCR), scientific breakthroughs depend on quickly moving from a hypothesis to a finished study. Yet for many researchers, the path to federal grants is often

Anthropic says its $20M donation to Public First Action can't be 'used to influence federal elections' and is to educate the public on AI policy (Veronica Irwin/Transformer)

SafetyDGX agent

Veronica Irwin / Transformer: Anthropic says its $20M donation to Public First Action can't be “used to influence federal elections” and is to educate the public on AI policy — The company's money isn

EGLOCE: Training-Free Energy-Guided Latent Optimization for Concept Erasure

SafetyDGX agent

arXiv:2604.09405v1 Announce Type: new Abstract: As text-to-image diffusion models grow increasingly prevalent, the ability to remove specific concepts-mostly explicit content and many copyrighted char

@ESYudkowsky My horrendous nightmare of a political lifecycle, ladies and gentlemen and others.

SafetyDGX agent

Connor Leahy shared a post on X (formerly Twitter) quoting or referencing Eliezer Yudkowsky's account (@ESYudkowsky), describing what he characterizes as a 'horrendous nightmare of a political lifecyc

EthicMind: A Risk-Aware Framework for Ethical-Emotional Alignment in Multi-Turn Dialogue

SafetyDGX agent

arXiv:2604.09265v1 Announce Type: new Abstract: Intelligent dialogue systems are increasingly deployed in emotionally and ethically sensitive settings, where failures in either emotional attunement or

Hierarchical Alignment: Enforcing Hierarchical Instruction-Following in LLMs through Logical Consistency

SafetyDGX agent

arXiv:2604.09075v1 Announce Type: new Abstract: Large language models increasingly operate under multiple instructions from heterogeneous sources with different authority levels, including system poli

Import AI 453: Breaking AI agents; MirrorCode; and ten views on gradual disempowerment

SafetyDGX agent

Import AI issue 453 covers research and developments around vulnerabilities in AI agent systems, including methods for breaking or adversarially manipulating AI agents. The issue also features MirrorC

Learning Vision-Language-Action World Models for Autonomous Driving

SafetyDGX agent

arXiv:2604.09059v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have recently achieved notable progress in end-to-end autonomous driving by integrating perception, reasoning, and

Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection

SafetyDGX agent

arXiv:2604.09024v1 Announce Type: cross Abstract: Multi-modal large language models (MLLMs) have emerged as powerful tools for analyzing Internet-scale image data, offering significant benefits but al

LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving

SafetyDGX agent

arXiv:2604.08719v1 Announce Type: cross Abstract: Recent years have seen remarkable progress in autonomous driving, yet generalization to long-tail and open-world scenarios remains a major bottleneck

Many Preferences, Few Policies: Towards Scalable Language Model Personalization

SafetyDGX agent

arXiv:2604.04144v2 Announce Type: replace-cross Abstract: The holy grail of LLM personalization is a single LLM for each user, perfectly aligned with that user's preferences. However, maintaining a se

Mosaic: Multimodal Jailbreak against Closed-Source VLMs via Multi-View Ensemble Optimization

SafetyDGX agent

arXiv:2604.09253v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are powerful but remain vulnerable to multimodal jailbreak attacks. Existing attacks mainly rely on either explicit visu

Scheming in the wild: detecting real-world AI scheming incidents with open-source intelligence

SafetyDGX agent

arXiv:2604.09104v1 Announce Type: cross Abstract: Scheming, the covert pursuit of misaligned goals by AI systems, represents a potentially catastrophic risk, yet scheming research suffers from signifi

Verbalizing LLMs' assumptions to explain and control sycophancy

SafetyDGX agent

arXiv:2604.03058v2 Announce Type: replace-cross Abstract: LLMs can be socially sycophantic, affirming users when they ask questions like 'am I in the wrong?' rather than providing genuine assessment.

⚡️ with @staysaasy, our second anonymous pod ever: https://www.youtube.com/watch?v=5KnCKadxSPY A conversation with the Stay Sassy duo on how…

SafetyDGX agent

⚡️ with @staysaasy, our second anonymous pod ever: https://www.youtube.com/watch?v=5KnCKadxSPY A conversation with the Stay Sassy duo on how AI is changing software teams, management, internal tooling

12 Apr 2026

Each workflow in Thoth is a full LangGraph agent which can call subagents which themselves are LangGraph agents. Just describe what you want…

SafetyDGX agent

Each workflow in Thoth is a full LangGraph agent which can call subagents which themselves are LangGraph agents. Just describe what you want in plain English and it builds a full multi-step pipeline.

🇧🇪Good news for Belgian Tesla owners! The Netherlands’ RDW just issued the first European type approval for Tesla FSD Supervised. Attached…

SafetyDGX agent

🇧🇪Good news for Belgian Tesla owners! The Netherlands’ RDW just issued the first European type approval for Tesla FSD Supervised. Attached letter from the Belgian federal administration (Minister Jean

Restore Britain would allow free public speech without imprisonment

SafetyDGX agent

Restore Britain would allow free public speech without imprisonment Platforms hosting lawful content must be shielded from government pressure to censor. We would require transparency in content moder

The Biden administration actively flew illegals into America with no vetting of their violent criminal past into America and paid for their …

SafetyDGX agent

The Biden administration actively flew illegals into America with no vetting of their violent criminal past into America and paid for their flights via NGOs. This is a war crime. Mayorkas and his budd

Under South Africa's Employment Equity policy, every employer with more than 50 staff must comply with RACIAL quota targets set by the gover…

SafetyDGX agent

Under South Africa's Employment Equity policy, every employer with more than 50 staff must comply with RACIAL quota targets set by the government. Under these rules, in roles such as 'skilled technici

← Previous
1…4950515253…240
Next →