AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,813 results
3 Jun 2026

Adaptive Causal Alignment for High-Confidence Adversarial Training

SafetyDGX agent

arXiv:2606.03925v1 Announce Type: new Abstract: Inverse adversarial training leverages high-confidence predictions to stabilize robust learning, yet we uncover a critical paradox: high confidence ofte

AI Agents Enable Adaptive Computer Worms

SafetyDGX agent

arXiv:2606.03811v1 Announce Type: cross Abstract: A computer worm is malware that spreads on a network by replicating itself from one machine to another. Traditional worms, like WannaCry, exploited pr

AirDreamer: Generalist Drone Navigation with World Models

SafetyDGX agent

arXiv:2606.03252v1 Announce Type: cross Abstract: Navigating a drone in unseen and cluttered environments requires reliable generalization to unseen scene layouts and understanding of environmental st


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Aletheia: What Makes RLVR For Code Verifiers Tick?

SafetyDGX agent

arXiv:2601.12186v3 Announce Type: replace-cross Abstract: Multi-domain thinking verifiers trained via Reinforcement Learning with Verifiable Rewards (RLVR) are a cornerstone of modern post-training. H

Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model Enhancement

SafetyDGX agent

arXiv:2412.01282v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) bring powerful understanding and reasoning capabilities to multimodal tasks. Meanwhile, the great need for capab

Aligning Data-Driven Predictors with Allocation: A Decision-Focused Approach to Survival Analysis

SafetyDGX agent

arXiv:2606.02671v1 Announce Type: cross Abstract: Machine learning predictors have become essential tools for guiding automated decision making. However, a major misalignment persists: predictive mode

Alignment-Aware Decoding

SafetyDGX agent

arXiv:2509.26169v2 Announce Type: replace Abstract: Alignment of large language models remains a central challenge in natural language processing. Preference optimization has emerged as a popular and

an astonishing statistic, consistent with the claim I made the other day that Elon’s best days may be behind him:

SafetyDGX agent

Gary Marcus shared a statistic on X that he claims supports his earlier assertion that Elon Musk's most successful period may be in the past. The post likely presents data related to Musk's business p

Are we really tilting? The mechanics of reward guidance in flow and diffusion models

SafetyDGX agent

arXiv:2606.02884v1 Announce Type: cross Abstract: Reward guidance algorithms steer a learned generative process toward the reward-tilted measure at inference time. While empirically powerful, these me

ASAP: Exploiting the Satisficing Generalization Edge in Neural Combinatorial Optimization

SafetyDGX agent

arXiv:2501.17377v4 Announce Type: replace-cross Abstract: Deep Reinforcement Learning (DRL) has emerged as a promising approach for solving Combinatorial Optimization (CO) problems, such as the 3D Bin

Assessing Region-Level EEG Contributions to Cognitive Workload Prediction

SafetyDGX agent

arXiv:2606.02598v1 Announce Type: new Abstract: Accurate and generalizable estimation of cognitive workload from electroencephalography (EEG) is critical for human-centered and safety-critical systems

ASymPO: Asymmetric-Scale Policy Optimization for Asynchronous LLM Post-Training Without Behavior Information

SafetyDGX agent

arXiv:2606.03070v1 Announce Type: cross Abstract: Asynchronous reinforcement learning can improve language-model post-training throughput by decoupling response generation from policy optimization, bu

Attention Calibration for Position-Fair Dense Information Retrieval

SafetyDGX agent

arXiv:2606.02737v1 Announce Type: cross Abstract: Dense retrieval models exhibit positional bias: retrieval effectiveness degrades when relevant information appears later in a passage (Zeng et al., 20

Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs

SafetyDGX agent

arXiv:2606.03785v1 Announce Type: new Abstract: Backdoor attacks in Large Language Models (LLMs) are a growing security concern, where models can generate adversary-chosen content. Existing defenses t

Bayesian Tensor Decomposition with Diffusion Model Prior

SafetyDGX agent

arXiv:2606.03212v1 Announce Type: new Abstract: Low-rank tensor decomposition (TD) is usually effective on clean, fully observed data, but it often degrades under severe missingness or noise. Low-rank

Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching

SafetyDGX agent

arXiv:2602.12221v2 Announce Type: replace Abstract: We propose UniDFlow, a unified discrete flow-matching framework for multimodal understanding, generation, and editing. It decouples understanding an

Bionic Human-Motion Style Transfer for Physically Executable Whole-Body Control of Humanoid Robots

SafetyDGX agent

arXiv:2606.03536v1 Announce Type: new Abstract: Expressive whole-body motion is important for humanoid robots operating in human environments, where robots are expected to move stably while presenting

Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs

SafetyDGX agent

arXiv:2606.03647v1 Announce Type: cross Abstract: Accurately evaluating adversarial robustness is a longstanding challenge. A flawed attack design can inflate robustness estimates, making deployment r

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL

SafetyDGX agent

arXiv:2510.08977v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) efficiently scales the reasoning ability of large language models (LLMs) but is bottlene

Bridging Predictive Uncertainty and Safe Action: Sample-Conditioned Differentiable Planning for Autonomous Driving

SafetyDGX agent

arXiv:2606.03296v1 Announce Type: new Abstract: Complex, dynamic, and interactive driving environments pose significant challenges for autonomous driving, primarily due to the pervasive uncertainty of

Brief Announcement: Generative Markov Model for Distributed Computing Systems

SafetyDGX agent

arXiv:2606.03061v1 Announce Type: cross Abstract: Emerging distributed computing paradigms, such as the computing continuum, are inherently heterogeneous, stochastic, and complex. Efficiently and effe

Building Better Activation Oracles

SafetyDGX agent

arXiv:2606.02609v1 Announce Type: cross Abstract: Activation Oracles (AOs) are promising methods for interpreting residual stream activations. However, current AOs face important issues, such as hallu

Coherence Maximization Improves Pluralistic Alignment

SafetyDGX agent

arXiv:2606.03110v1 Announce Type: new Abstract: Aligning AI systems with diverse human values requires value specifications grounded in concrete examples, but generating such examples without extensiv

Collab-REC: An LLM-based Agentic Framework for Balancing Recommendations in Tourism

SafetyDGX agent

arXiv:2508.15030v5 Announce Type: replace Abstract: We propose COLLAB-REC, a multi-agent framework designed to counteract popularity bias and improve diversity in tourism recommendations. In our setup

Consistency Training Can Entrench Misalignment

SafetyDGX agent

arXiv:2606.03810v1 Announce Type: cross Abstract: Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures. Such methods are simple, scalable, an

Constitutional On-Policy Safe Distillation

SafetyDGX agent

arXiv:2606.03089v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged information to prov

ConTrack: Constrained Hand Motion Tracking with Adaptive Trade-off Control

SafetyDGX agent

arXiv:2606.03177v1 Announce Type: new Abstract: Human demonstrations provide strong priors for robot manipulation, yet it is non-trivial to transfer them to execute on real robots due to the kinematic

ConTraIRL: Factorized Contrastive Abstractions for Transferable IRL

SafetyDGX agent

arXiv:2606.03017v1 Announce Type: cross Abstract: Reward transfer in Inverse Reinforcement Learning (IRL) is unreliable when policies must generalize to unseen combinations of environment dynamics and

Contrastive Neural Algorithmic Reasoning for Graph Coloring

SafetyDGX agent

arXiv:2606.03923v1 Announce Type: new Abstract: Graph coloring seeks to assigns colors to a graph's nodes so that adjacent nodes receive different colors, using as few colors as possible. Here, we stu

Correcting Neural Operator Spectral Bias via Diffusion Posterior Sampling with Sparse Observations

SafetyDGX agent

arXiv:2606.03936v1 Announce Type: new Abstract: Neural operator surrogates (NO) approximate PDE solutions orders of magnitude faster than numerical solvers, but suffer from spectral bias: high-frequen

CP-Agent: Context-Aware Multimodal Reasoning for Cellular Morphological Profiling under Chemical Perturbations

SafetyDGX agent

arXiv:2606.03435v1 Announce Type: new Abstract: Cell Painting combines multiplexed fluorescent staining, high-content imaging, and quantitative analysis to generate high-dimensional phenotypic readout

Curriculum-Adapted Robust Reinforcement Learning for UAV Deconfliction in Adversarial Environments

SafetyDGX agent

arXiv:2506.21129v2 Announce Type: replace-cross Abstract: Autonomous unmanned aerial vehicles (UAVs) increasingly rely on reinforcement learning (RL) for navigation. However, global navigation satelli

D-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output Rewriting

SafetyDGX agent

arXiv:2606.02640v1 Announce Type: cross Abstract: Multi-turn jailbreak attacks pose a growing threat to large language model (LLM) safety because they exploit feedback from auxiliary judge models to i

Data- and Variance-dependent Regret Bounds for Online Tabular MDPs

SafetyDGX agent

arXiv:2602.01903v2 Announce Type: replace Abstract: This work studies online episodic tabular Markov decision processes (MDPs) with known transitions and develops best-of-both-worlds algorithms that a

Denoise First, Orthogonalize Later: Understanding Momentum in Muon via Spectral Filtering

SafetyDGX agent

arXiv:2606.03899v1 Announce Type: new Abstract: Muon has recently demonstrated strong empirical performance in large language model training, but the theoretical role of momentum in Muon remains uncle

Denoising Tells When to Replan: Denoising-Variance Adaptive Chunking for Flow-Based Robot Policies

SafetyDGX agent

arXiv:2606.03847v1 Announce Type: new Abstract: Action chunking has become a common inference strategy for flow-based robot policies, improving action coherence by modeling multi-step temporal depende

Discovering autonomous quantum error correction via deep reinforcement learning

SafetyDGX agent

arXiv:2511.12482v2 Announce Type: replace-cross Abstract: Quantum error correction is essential for fault-tolerant quantum computing. However, standard methods relying on active measurements may intro

Do Explanations Increase the Risk of Decision Logic Leakage? Explanation-Guided Stealing of Graph Models

SafetyDGX agent

arXiv:2506.03087v2 Announce Type: replace-cross Abstract: Graph Neural Networks (GNNs) have become essential tools for analyzing graph-structured data in domains such as drug discovery and financial a

Do Neural Retrievers Prefer Certain Documents? Evidence of Learned Relevance Priors

SafetyDGX agent

arXiv:2606.02814v1 Announce Type: cross Abstract: Neural retrievers are trained to estimate query-document relevance from annotated query-document pairs. Yet annotation protocols may not purely reflec

DriftSched: Adaptive QoS-Aware Scheduling under Runtime Token Drift for Multi-Tenant GPU Inference

SafetyDGX agent

arXiv:2606.02982v1 Announce Type: cross Abstract: The rapid growth of large language model (LLM) inference services has increased the demand for efficient multi-tenant GPU scheduling. While modern inf

Dynamic Short Convolutions Improve Transformers

SafetyDGX agent

arXiv:2606.03825v1 Announce Type: cross Abstract: Transformers have become the dominant architecture for large language models, largely due to the scalability and flexibility of attention, feed-forwar

Easy-to-Use Shielding for Reinforcement Learning

SafetyDGX agent

arXiv:2606.03804v1 Announce Type: new Abstract: Safe exploration is a key challenge in Reinforcement Learning (RL) that aims to prevent agents from making harmful decisions while exploring their envir

Effect of Demographic Bias on Skin Lesion Classification

SafetyDGX agent

arXiv:2606.03214v1 Announce Type: new Abstract: In this study, we evaluate the performance of skin lesion classification using ResNet-based convolutional models, focusing on the impact of demographic

Enhancing Operational Safety via Agentic Dialogue Hazard Identification Analysis

SafetyDGX agent

arXiv:2606.03812v1 Announce Type: new Abstract: Operational safety in high-stakes domains such as industrial process control, autonomous, and safety-critical systems, demand reliable hazard identifica

Entropy Is Not Enough: Unlocking Effective Reinforcement Learning for Visual Reasoning via Vision-Anchored Token Selection

SafetyDGX agent

arXiv:2606.03937v1 Announce Type: new Abstract: While token-level entropy is commonly recognized as effective for credit assignment in text-only reinforcement learning with verifiable rewards (RLVR),

Estimating Bidirectional Causal Effects with Large Scale Online Kernel Learning

SafetyDGX agent

arXiv:2511.05050v3 Announce Type: replace-cross Abstract: In this study, a scalable online kernel learning framework is proposed for estimating bidirectional causal effects in systems characterized by

Evaluating Transformer and LSTM Frameworks for Prediction in Ungauged Basins

SafetyDGX agent

arXiv:2606.02791v1 Announce Type: new Abstract: Watershed networks exhibit convergent topologies in which multiple tributaries merge into downstream channels,integrating diverse upstream hydrological

Every action that I partake is animated by two ideals: Truth and freedom. Seeing the endless attacks on both ideals throughout the West is s…

SafetyDGX agent

Every action that I partake is animated by two ideals: Truth and freedom. Seeing the endless attacks on both ideals throughout the West is soul-crushing. We did not lose a war of aggression. We decide

EvoMemNav: Efficient Self-Evolving Fine-Grained Memory for Zero-Shot Embodied Navigation

SafetyDGX agent

arXiv:2606.03509v1 Announce Type: new Abstract: Building memory is essential for long-horizon planning in zero-shot embodied navigation. Detector-centric scene graphs often compress observations into

Explainable Forecasting of Scientific Breakthroughs from Concept Network Dynamics

SafetyDGX agent

arXiv:2606.03864v1 Announce Type: cross Abstract: We introduce an explainable machine-learning approach that forecasts the structural precursors of scientific breakthroughs -- the emergence and intens

Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models

SafetyDGX agent

arXiv:2606.03793v1 Announce Type: new Abstract: Multimodal Large Language Models integrate visual perception into language reasoning, introducing a continuous attack surface susceptible to adversarial

Extreme Motion Generation via Hybrid Null-Space Control for Straight-Line Path Following

SafetyDGX agent

arXiv:2606.03390v1 Announce Type: new Abstract: This work studies ``extreme motion generation'', which aims to maximize the Cartesian path length along a pre-defined trajectory within the manipulator'

FAF-CD: Frequency-Aware Fusion for Change Detection under Imperfect Multimodal Remote Sensing

SafetyDGX agent

arXiv:2606.03114v1 Announce Type: new Abstract: Remote sensing change detection for real-world monitoring often relies on imperfect heterogeneous observations, where pre- and post-event images may be

Fairness Definitions and Metrics in Deep Reinforcement Learning for Drug Discovery in Healthcare: A Rapid Evidence Review

SafetyDGX agent

arXiv:2606.02902v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) is increasingly applied to de novo molecular design, but choices in data, rewards, and evaluation can yield uneven p

Fast Organic Crystal Structure Prediction with Unit Cell Flow Matching

SafetyDGX agent

arXiv:2606.03199v1 Announce Type: new Abstract: Organic crystal structure prediction (CSP) is a requirement for computational modelling of organic solids, but traditionally costs several CPU-years per

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data

SafetyDGX agent

arXiv:2606.03094v1 Announce Type: new Abstract: Recent advances in language models have established reinforcement learning as the primary paradigm for eliciting self-correction and long-chain reasonin

Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation

SafetyDGX agent

arXiv:2606.02684v1 Announce Type: cross Abstract: On-Policy distillation (OPD) in large language models is shifting from full-trace KL supervision toward more selective training paradigms. Recent OPD

First it was MIT and McKinsey. Now Bain finds that returns to corporate AI investments are disappointing.

SafetyDGX agent

A major consulting firm (Bain) has found that corporate returns on AI investments are underwhelming, following similar findings from MIT and McKinsey. This suggests that despite significant spending a

Flow Learners for PDEs: Toward a Physics-to-Physics Paradigm for Scientific Computing

SafetyDGX agent

arXiv:2604.07366v2 Announce Type: replace Abstract: Partial differential equations (PDEs) govern nearly every physical process in science and engineering, but solving them at scale remains prohibitive

Follow-Your-Preference++: Rethinking Preference Alignment for Image Inpainting

SafetyDGX agent

arXiv:2606.03216v1 Announce Type: new Abstract: We study preference alignment for image inpainting. Rather than proposing yet another method, we revisit the problem from first principles and reassess

← Previous
1…9293949596…214
Next →