AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,356 results
9 Jun 2026

Position: Anthropomorphic Misalignment Research Needs Stronger Evidence

SafetyDGX agent

arXiv:2606.07612v1 Announce Type: cross Abstract: We argue that many Anthropomorphic Misalignment Research (AMR) studies need stronger evidence to ensure that they can provide a robust foundation for

Revisiting the shutdown problem

SafetyDGX agent

arXiv:2606.08296v1 Announce Type: new Abstract: A key premise in leading arguments for existential risk from artificial intelligence is that malfunctioning artificial agents could not be easily shut d

RPO-PDT: Demonstrating Role-Play-Based Knowledge Adaptation for Student Support Dialogue (Demonstration System)

SafetyDGX agent

arXiv:2606.09255v1 Announce Type: new Abstract: We present RPO-PDT: a retrieval-grounded, role-play-based dialogue system for adaptive student support in higher education. RPO-PDT is: (1) able to prov

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SAD-Flower: Flow Matching for Safe, Admissible, and Dynamically Consistent Planning

SafetyDGX agent

arXiv:2511.05355v3 Announce Type: replace Abstract: Flow matching (FM) has shown promising results in data-driven planning. However, it inherently lacks formal guarantees for ensuring state and action

SafeRun: Enabling Determinism in LLM Planning for Running

Model ReleasesDGX agent

arXiv:2606.09027v1 Announce Type: cross Abstract: Large Language Models enable flexible natural-language planning but remain unreliable in determinism-critical domains due to their probabilistic natur

Stain-Aware Wavelet Regularization for Instant Adversarial Purification in Histopathology

SafetyDGX agent

arXiv:2606.08745v1 Announce Type: new Abstract: Deep learning has become prevalent in computational pathology pipelines that support tasks such as cancer screening and digital pathology analysis. Howe

Toward autocorrection of chemical process flowsheets using large language models

SafetyDGX agent

arXiv:2312.02873v2 Announce Type: replace-cross Abstract: The process engineering domain widely uses Process Flow Diagrams (PFDs) and Process and Instrumentation Diagrams (P&IDs) to represent process

TRACER: Token ReAssignment for Concept ERasure in Generative Recommendation

SafetyDGX agent

arXiv:2606.07688v1 Announce Type: cross Abstract: Generative recommendation formulates next-item prediction as autoregressive generation over semantic ID (SID) sequences derived from users' historical

Transforming Police-Car Swerving for Mitigating Isolated Stop-and-Go Traffic Waves: A Practice-Oriented Jam-Absorption Driving Strategy

SafetyDGX agent

arXiv:2602.10234v3 Announce Type: replace-cross Abstract: Stop-and-go traffic waves, a major form of freeway congestion, impose severe and persistent adverse impacts, including reduced traffic efficie

Video2Sim2Real: Full-Stack Autonomous Dexterous Skill Acquisition from a Single Human Video

SafetyDGX agent

arXiv:2606.08828v1 Announce Type: new Abstract: Human manipulation videos are a convenient and intuitive source for robot learning. However, directly transferring human dexterity to robots remains cha

8 Jun 2026

An Abstract Architecture for Explainable Autonomy in Hazardous Environments

SafetyDGX agent

arXiv:2606.07211v1 Announce Type: cross Abstract: Autonomous robotic systems are being proposed for use in hazardous environments, often to reduce the risks to human workers. In the immediate future,

Bounded-Abstention Pairwise Learning to Rank

SafetyDGX agent

arXiv:2505.23437v2 Announce Type: replace-cross Abstract: Ranking systems influence decision-making in high-stakes domains like health, education, and employment, where they can have substantial econo

CARVE-Q: Quantum-Proposed, Classically Certified Interactive Driving Repair

SafetyDGX agent

arXiv:2606.06531v1 Announce Type: new Abstract: The critical question after a correct driving veto is not only whether a maneuver is unsafe, but whether the blocked interaction admits a lawful, audita

check out Marcus on AI

SafetyDGX agent

Gary Marcus, a prominent AI researcher and critic, shared a post on X directing followers to learn more about his perspectives on artificial intelligence. The post likely promotes his work, writings,

Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing

SafetyDGX agent

This Import AI newsletter issue covers three main topics: societal implications of reward hacking (optimizing for measurable metrics at the expense of intended goals), new reinforcement learning data

Position: Don't Just 'Fix it in Post': A Science of AI Must Study Training Dynamics

SafetyDGX agent

arXiv:2606.06533v1 Announce Type: new Abstract: What would it mean to have a scientific understanding of AI? Models are not static objects: they are snapshots of time-evolving processes shaped by data

UK PM Keir Starmer says tech companies must introduce 'device controls' that stop kids from sending and receiving nude images or face laws forcing them to do so (Reuters)

SafetyDGX agent

Reuters: UK PM Keir Starmer says tech companies must introduce “device controls” that stop kids from sending and receiving nude images or face laws forcing them to do so — Big tech firms operating in

Which title is better? Seven [Lies/Myths/Bitter Truths] About AI

SafetyDGX agent

Gary Marcus critiques common misconceptions about AI, presenting seven corrective perspectives on prevailing beliefs about artificial intelligence's capabilities and limitations. The post likely addre

6 Jun 2026

Consistency Training Along the Transformer Stack

SafetyDGX agent

arXiv:2606.05817v1 Announce Type: cross Abstract: Consistency training encourages models to behave similarly across different contexts, and has shown promise for reducing misalignment. We broaden the

If leading AI companies are indeed approaching the point of recursive self-improvement, a coordinated, verifiable, and universally applied p…

SafetyDGX agent

If leading AI companies are indeed approaching the point of recursive self-improvement, a coordinated, verifiable, and universally applied pause is probably the only responsible solution to mitigate s

Risk Assessment of Autonomous Driving: Integrating Technical Failures, Ethical Dilemmas, and Policy Frameworks

SafetyDGX agent

arXiv:2606.06396v1 Announce Type: new Abstract: Autonomous driving technology has the potential to reduce the large number of road traffic accidents caused by human error each year, but it also brings

Towards World Models in Biomedical Research

SafetyDGX agent

arXiv:2606.05925v1 Announce Type: new Abstract: A central goal of biomedicine is to understand, predict and ultimately control the dynamic mechanisms by which biological systems respond to perturbatio

Unsupervised Pattern Analysis in Japanese Veterinary Toxicology: A Regulatory-Compliant Framework for Cross-Species Risk Assessment

SafetyDGX agent

arXiv:2606.06207v1 Announce Type: new Abstract: Veterinary pharmacovigilance systems are essential for monitoring adverse drug events (ADEs), yet existing approaches often fail to capture region-speci

Willing but Unable: Separating Refusal from Capability in Code LLMs via Abliteration

SafetyDGX agent

arXiv:2606.05396v1 Announce Type: cross Abstract: Producing a labeled vulnerable code at scale is a recurring obstacle for learning-based vulnerability detection: mined corpora carry substantial label

5 Jun 2026

Alignment Risks from Capability-Seeking RL Training

SafetyDGX agent

arXiv:2602.12124v2 Announce Type: replace-cross Abstract: While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capab

Forgive or forget: Understanding the context of hate in audio retrieval systems

SafetyDGX agent

arXiv:2606.05857v1 Announce Type: new Abstract: Handling toxic retrieval in text-to-audio systems is challenging due to contextual dependencies. Existing strategies (e.g., rephrasing, summarization) r

HANDOFF: Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers

SafetyDGX agent

arXiv:2606.06493v1 Announce Type: new Abstract: For a humanoid robot to be deployed in the real world, the choice of command space (i.e., the interface between task planning and whole-body control) is

Literally the only reason to bailout OpenAI.

SafetyDGX agent

Gary Marcus argues for a specific rationale supporting a potential OpenAI bailout, though the exact reasoning is not detailed in the available information. Given Marcus's background as an AI researche

What's in a Name? Morphological Shortcuts by LLMs in Pharmacology

SafetyDGX agent

arXiv:2606.05616v1 Announce Type: new Abstract: The morphological form of a word can often give cues to its meaning, but purely relying on these mappings can lead to overgeneralization in high-stakes

4 Jun 2026

A Pathology Foundation Model for Gastric Cancer with Real-World Validation

SafetyDGX agent

arXiv:2606.04792v1 Announce Type: new Abstract: Gastric cancer remains a major cause of cancer mortality, yet its histological and molecular heterogeneity complicates diagnosis and risk stratification

Activation Steering of Video Generation Models via Reduced-Order Linear Optimal Control

SafetyDGX agent

arXiv:2606.04775v1 Announce Type: cross Abstract: Text-to-video (T2V) models trained on large-scale web data can generate undesired content, motivating interventions that reduce harmful outputs withou

Certified Neural Approximations of Nonlinear Dynamics

SafetyDGX agent

arXiv:2505.15497v3 Announce Type: replace Abstract: Neural networks hold great potential to act as approximate models of nonlinear dynamical systems, with the resulting neural approximations enabling

Formal Semantics for Agentic Tool Protocols: A Process Calculus Approach

SafetyDGX agent

arXiv:2603.24747v2 Announce Type: replace Abstract: The emergence of large language model agents capable of invoking external tools has created urgent need for formal verification of agent protocols.

Learning Empirically Admissible Neural Heuristics for Combinatorial Search

SafetyDGX agent

arXiv:2606.04860v1 Announce Type: cross Abstract: Finding optimal solution paths for combinatorial puzzles like the Rubik's Cube, sliding tile puzzles, and Lights Out remains a classical challenge in

MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models

SafetyDGX agent

arXiv:2606.04027v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safe

Off-Distribution Voices: Fanfiction Subgenres as Universal Vernacular Jailbreaks for Aligned LLMs

SafetyDGX agent

arXiv:2606.04483v1 Announce Type: new Abstract: Existing jailbreaks against aligned LLMs are discrete artifacts whose surface forms are easy to fingerprint and patch. We argue that the real failure mo

PerceptTwin: Semantic Scene Reconstruction for Iterative LLM Planning and Verification

SafetyDGX agent

arXiv:2606.04226v1 Announce Type: cross Abstract: Simulation environments are useful for both robot policy learning and planning verification and validation. Traditionally, the process of creating a s

Plan First, Judge Later, Run Better: A DMAIC-Inspired Agentic System for Industrial Anomaly Detection

SafetyDGX agent

arXiv:2606.04599v1 Announce Type: new Abstract: Large language model (LLM) agents have shown promise in automating complex data-analysis workflows, but their reliable deployment remains challenging in

Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models

SafetyDGX agent

arXiv:2502.01576v2 Announce Type: replace Abstract: Multi-modal Large Language Models (MLLMs) excel in vision-language tasks but remain vulnerable to visual adversarial perturbations that can induce h

Two House lawmakers unveil bipartisan AI legislation that would override some state AI laws and require top AI developers to implement risk-management plans (Politico)

SafetyDGX agent

Politico: Two House lawmakers unveil bipartisan AI legislation that would override some state AI laws and require top AI developers to implement risk-management plans — But it's the proposal to preemp

veriFIRE: an Industrial Case Study in Verifying Consistency Properties for a DNN-Based Wildfire Detection System

SafetyDGX agent

arXiv:2606.04121v1 Announce Type: cross Abstract: We present our ongoing work on the veriFIRE project: a collaboration between industry and academia, aimed at applying verification to increase the rel

What Can Eye Gaze Teach Us About Real-World Cycling? Insights From the Oxford RobotCycle Project

SafetyDGX agent

arXiv:2606.04989v1 Announce Type: cross Abstract: Although much is known about the physical danger of cycling situations, less is understood about the perceived danger of cycling. Furthermore, percept

3 Jun 2026

AI Agents Enable Adaptive Computer Worms

SafetyDGX agent

arXiv:2606.03811v1 Announce Type: cross Abstract: A computer worm is malware that spreads on a network by replicating itself from one machine to another. Traditional worms, like WannaCry, exploited pr

Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs

SafetyDGX agent

arXiv:2606.03785v1 Announce Type: new Abstract: Backdoor attacks in Large Language Models (LLMs) are a growing security concern, where models can generate adversary-chosen content. Existing defenses t

Every action that I partake is animated by two ideals: Truth and freedom. Seeing the endless attacks on both ideals throughout the West is s…

SafetyDGX agent

Every action that I partake is animated by two ideals: Truth and freedom. Seeing the endless attacks on both ideals throughout the West is soul-crushing. We did not lose a war of aggression. We decide

Extreme Motion Generation via Hybrid Null-Space Control for Straight-Line Path Following

SafetyDGX agent

arXiv:2606.03390v1 Announce Type: new Abstract: This work studies ``extreme motion generation'', which aims to maximize the Cartesian path length along a pre-defined trajectory within the manipulator'

How Kodiak trains the brain behind 28 driverless trucks

SafetyDGX agent

Twenty-eight trucks, and no humans in the cab. As of March 31, 2026, Kodiak's autonomous driving system, the Kodiak Driver, runs commercial freight on public roads across long-haul trucking, and indus

How OpenAI, Anthropic, and AI startups are pursuing 'recursive self-improvement', in a bid to build AI that can improve itself with little to no human input (Financial Times)

SafetyDGX agent

Financial Times: How OpenAI, Anthropic, and AI startups are pursuing “recursive self-improvement”, in a bid to build AI that can improve itself with little to no human input — Industry chiefs say tech

LC-SAC: Lyapunov-Constrained Soft Actor-Critic via Koopman Operator Theory for Trajectory Tracking and Stabilization

SafetyDGX agent

arXiv:2602.04132v4 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) has achieved remarkable success in solving complex sequential decision-making problems. However, its application t

Localized, High-resolution Geographic Representations with Slepian Functions

SafetyDGX agent

arXiv:2602.00392v2 Announce Type: replace Abstract: Geographic data is fundamentally local. Disease outbreaks cluster in population centers, ecological patterns emerge along coastlines, and economic a

Measuring Weak-to-Strong Legibility of Reasoning Models

SafetyDGX agent

arXiv:2603.20508v2 Announce Type: replace-cross Abstract: Reasoning language models (RLMs) and the intermediate chains of thought they emit play an increasingly central role in multi-agent setups such

Rethinking Neural Width for Alternating Current Optimal Power Flow Proxies

SafetyDGX agent

arXiv:2606.03125v1 Announce Type: new Abstract: Deep learning proxies for Alternating Current Optimal Power Flow (ACOPF) lack systematic methods for determining architectural size. This paper conducts

Validation-Gated Multi-Agent Governance for Online Adaptation of Thermal-Hydraulic Surrogate Models under Operating-Regime Shift

SafetyDGX agent

arXiv:2606.03321v1 Announce Type: new Abstract: Artificial-intelligence surrogates can support second-by-second thermal-hydraulic forecasting, but models selected and frozen offline may become conditi

2 Jun 2026

A Predictive Control Strategy to Offset-Point Tracking for Agricultural Mobile Robots

SafetyDGX agent

arXiv:2603.28439v2 Announce Type: replace Abstract: Robots are increasingly being deployed in agriculture to support sustainable practices and improve productivity. They offer strong potential to enab

Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults

SafetyDGX agent

arXiv:2606.00914v1 Announce Type: new Abstract: LLM agents increasingly act after consuming ranked external information streams such as social feeds, search results, retrieval contexts, and email queu

Adversarially Robust Control of Conditional Value-at-Risk via Rockafellar-Uryasev Conformal Inference

SafetyDGX agent

arXiv:2606.00320v1 Announce Type: new Abstract: We present an online, distribution-free framework for controlling the Conditional Value-at-Risk (CVaR), extending conformal tail risk control to non-sta

Agent Operating Systems (AOS): Integrating Agentic Control Planes into, and Beyond, Traditional Operating Systems

SafetyDGX agent

arXiv:2606.01508v1 Announce Type: cross Abstract: Traditional operating systems were designed around deterministic programs, explicit control flow, and human initiated workflows. Their core abstractio

Balancing Accuracy and Efficiency: Adaptive Dynamics Orchestration for Model Predictive Control

SafetyDGX agent

arXiv:2606.00085v1 Announce Type: new Abstract: Model Predictive Control (MPC) for autonomous navigation faces a fundamental trade-off between model accuracy and real-time efficiency. High-fidelity dy

Bridging the Last Mile of Time Series Forecasting with LLM Agents

SafetyDGX agent

arXiv:2606.02497v1 Announce Type: new Abstract: Time series forecasting has advanced rapidly, especially with the emergence of foundation models that show strong zero-shot performance on numerical ext

Collaborative Space Object Detection with Multi-Satellite Viewpoints in LEO Constellations

SafetyDGX agent

arXiv:2606.01895v1 Announce Type: cross Abstract: With the growing number of satellites in low Earth orbit (LEO) constellations, the near-Earth space environment has become increasingly congested, mak

← Previous
1…4041424344…240
Next →