AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlog
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,485 results
Safety

Diffusion Model's Generalization Can Be Characterized by Inductive Biases toward a Data-Dependent Ridge Manifold

DGX agent

arXiv:2602.06021v2 Announce Type: replace-cross Abstract: We study a data-dependent notion of diffusion-model generalization: when a model does not memorize the training set, where do its generated sa

safetyarxiv-cs-lg
14 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Do Fair Models Reason Fairly? Counterfactual Explanation Consistency for Procedural Fairness in Credit Decisions

DGX agent

arXiv:2605.12701v1 Announce Type: cross Abstract: Machine learning algorithms in socially sensitive domains (e.g., credit decisions) often focus on equalizing predictive outcomes. However, satisfying

safetyarxiv-cs-ai
14 May 2026
Safety

DP-Muon: Differentially Private Optimization via Matrix-Orthogonalized Momentum

DGX agent

arXiv:2605.12994v1 Announce Type: new Abstract: We study differentially private (DP) training with Muon, a matrix-valued optimizer that updates hidden-layer weights using momentum followed by Newton--

safetyarxiv-cs-lg
14 May 2026
Safety

Driving Intents Amplify Planning-Oriented Reinforcement Learning

DGX agent

arXiv:2605.12625v1 Announce Type: cross Abstract: Continuous-action policies trained on a single demonstrated trajectory per scene suffer from mode collapse: samples cluster around the demonstrated ma

safetyarxiv-cs-cv
14 May 2026
Safety

Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer

DGX agent

arXiv:2605.12798v1 Announce Type: cross Abstract: Fine-tuning LLMs on narrow harmful datasets can induce Emergent Misalignment (EM), where models exhibit misaligned behavior far beyond the fine-tuning

safetyarxiv-cs-ai
14 May 2026
Safety

Ergodic Trajectory Design by Learned Pushforward Maps: Provable Coverage via Conditional Flow Matching

DGX agent

arXiv:2605.13063v1 Announce Type: new Abstract: Designing continuous trajectories whose time-averaged occupancy provably matches a prescribed spatial density (the ergodic coverage problem) is central

safetyarxiv-cs-lg
14 May 2026
Safety

ERPPO: Entropy Regularization-based Proximal Policy Optimization

DGX agent

arXiv:2605.13131v1 Announce Type: new Abstract: Multi-Agent Proximal Policy Optimization (MAPPO) is a variant of the Proximal Policy Optimization (PPO) algorithm, specifically tailored for multi-agent

safetyarxiv-cs-lg
14 May 2026
Safety

F-GRPO: Factorized Group-Relative Policy Optimization for Unified Candidate Generation and Ranking

DGX agent

arXiv:2605.12995v1 Announce Type: new Abstract: Traditional retrieval pipelines optimize utility through stages of candidate retrieval and reranking, where ranking operates over a predefined candidate

safetyarxiv-cs-lg
14 May 2026
Safety

Flow Matching for Offline Reinforcement Learning with Discrete Actions

DGX agent

arXiv:2602.06138v2 Announce Type: replace Abstract: Generative policies based on diffusion models and flow matching have shown strong promise for offline reinforcement learning (RL), but their applica

safetyarxiv-cs-lg
14 May 2026
Model Releases

Follow-Bench: A Unified Motion Planning Benchmark for Socially-Aware Robot Person Following

DGX agent

arXiv:2509.10796v4 Announce Type: replace Abstract: Robot person following (RPF) -- mobile robots that follow and assist a specific person -- has emerging applications in personal assistance, security

model-releasesarxiv-cs-ro
14 May 2026
Safety

FrameSkip: Learning from Fewer but More Informative Frames in VLA Training

DGX agent

arXiv:2605.13757v1 Announce Type: new Abstract: Vision-Language-Action (VLA) policies are commonly trained from dense robot demonstration trajectories, often collected through teleoperation, by sampli

safetyarxiv-cs-ro
14 May 2026
Safety

Frequency Bias and OOD Generalization in Neural Operators under a Variable-Coefficient Wave Equation

DGX agent

arXiv:2605.12997v1 Announce Type: new Abstract: Neural operators learn to map initial conditions to the terminal solution of partial differential equations (PDEs), providing a surrogate for the full o

safetyarxiv-cs-lg
14 May 2026
Safety

GAGPO: Generalized Advantage Grouped Policy Optimization

DGX agent

arXiv:2605.13217v1 Announce Type: cross Abstract: Reinforcement learning has become a powerful paradigm for post-training large language model agents, yet credit assignment in multi-turn environments

safetyarxiv-cs-lg
14 May 2026
Safety

GRACE: Gradient-aligned Reasoning Data Curation for Efficient Post-training

DGX agent

arXiv:2605.13130v1 Announce Type: new Abstract: Existing reasoning data curation pipelines score whole samples, treating every intermediate step as equally valuable. In reality, steps within a trace c

safetyarxiv-cs-ai
14 May 2026
Safety

Graph-Based Financial Fraud Detection with Calibrated Risk Scoring and Structural Regularization

DGX agent

arXiv:2605.12782v1 Announce Type: new Abstract: Financial transaction fraud prevention faces challenges such as complex relationship structures, concealed behavioral patterns, and dynamically changing

safetyarxiv-cs-lg
14 May 2026
Safety

Helping ChatGPT better recognize context in sensitive conversations

DGX agent

OpenAI implemented improvements to help ChatGPT better understand and respond appropriately to context in sensitive conversations, such as those involving mental health, abuse, or other delicate topic

safetyopenai
14 May 2026
Safety

HIR-ALIGN: Enhancing Hyperspectral Image Restoration via Diffusion-Based Data Generation

DGX agent

arXiv:2605.13581v1 Announce Type: new Abstract: Hyperspectral image (HSI) restoration is crucial for reliable analysis, as real HSIs suffer from degradations like noise, blur, and resolution loss. How

safetyarxiv-cs-cv
14 May 2026
Safety

Improving Classifier-Free Guidance of Flow Matching via Manifold Projection

DGX agent

arXiv:2601.21892v2 Announce Type: replace-cross Abstract: Classifier-free guidance (CFG) is a widely used technique for controllable generation in diffusion and flow-based models. Despite its empirica

safetyarxiv-cs-ai
14 May 2026
Safety

Improving Code Translation with Syntax-Guided and Semantic-aware Preference Optimization

DGX agent

arXiv:2605.13229v1 Announce Type: new Abstract: LLMs have shown immense potential for code translation, yet they often struggle to ensure both syntactic correctness and semantic consistency. While pre

safetyarxiv-cs-ai
14 May 2026
Safety

Improving Diffusion Posterior Samplers with Lagged Temporal Corrections for Image Restoration

DGX agent

arXiv:2605.12573v1 Announce Type: cross Abstract: Diffusion-based posterior sampling (PS) is a leading framework for imaging inverse problems, combining learned priors with measurement constraints. Ye

safetyarxiv-cs-ai
14 May 2026
Safety

In a policy paper, Anthropic urges the US and allies to enforce export controls, curb distillation attacks, and export US AI to hold the lead over China by 2028 (Anthropic)

DGX agent

Anthropic: In a policy paper, Anthropic urges the US and allies to enforce export controls, curb distillation attacks, and export US AI to hold the lead over China by 2028 — We're releasing a new pape

safetytechmeme
14 May 2026
Safety

In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores

DGX agent

arXiv:2605.12530v1 Announce Type: cross Abstract: LLM fairness should be evaluated through in-situ conversational behavior rather than standardized-test Q&A benchmarks. We show that the standardized-t

safetyarxiv-cs-ai
14 May 2026
Safety

In the @nytimes, Media Lab Prof. @kesvelt and other scientists call for stronger oversight and regulation of AI technologies, including chat…

DGX agent

In the @nytimes, Media Lab Prof. @kesvelt and other scientists call for stronger oversight and regulation of AI technologies, including chatbots that can provide information on producing lethal biolog

safetygary-marcus--x
14 May 2026
Safety

interwhen: A Generalizable Framework for Steering Reasoning Models with Test-time Verification

DGX agent

arXiv:2602.11202v3 Announce Type: replace-cross Abstract: Reasoning models produce long traces of intermediate decisions and tool calls, making test-time verification important for ensuring correctnes

safetyarxiv-cs-ai
14 May 2026
Safety

Is Video Anomaly Detection Misframed? Evidence from LLM-Based and Multi-Scene Models

DGX agent

arXiv:2605.12725v1 Announce Type: new Abstract: Recent video anomaly detection research has expanded rapidly with an emphasis on general models of normality intended to work across many different scen

safetyarxiv-cs-cv
14 May 2026
Model Releases

Learning Responsibility-Attributed Adversarial Scenarios for Testing Autonomous Vehicles

DGX agent

arXiv:2605.13751v1 Announce Type: new Abstract: Establishing trustworthy safety assurance for autonomous driving systems (ADSs) requires evidence that failures arise from avoidable system deficiencies

model-releasesarxiv-cs-ro
14 May 2026
Safety

Learning to Decide with AI Assistance under Human-Alignment

DGX agent

arXiv:2605.12646v1 Announce Type: cross Abstract: It is widely agreed that when AI models assist decision-makers in high-stakes domains by predicting an outcome of interest, they should communicate th

safetyarxiv-cs-ai
14 May 2026
Safety

Learning Transferable Latent User Preferences for Human-Aligned Decision Making

DGX agent

arXiv:2605.12682v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as reasoning modules in many applications. While they are efficient in certain tasks, LLMs often stru

safetyarxiv-cs-ai
14 May 2026
Safety

Learning with Rare Success but Rich Feedback via Reflection-Enhanced Self-Distillation

DGX agent

arXiv:2605.12741v1 Announce Type: new Abstract: Enabling Large Language Models (LLMs) to continuously improve from environmental interactions is a central challenge in post-training. While on-policy s

safetyarxiv-cs-lg
14 May 2026
Safety

Macro-Action Based Multi-Agent Instruction Following through Value Cancellation

DGX agent

arXiv:2605.12655v1 Announce Type: new Abstract: Multi-agent reinforcement learning (MARL) in real-world use cases may need to adapt to external natural language instructions that interrupt ongoing beh

safetyarxiv-cs-ai
14 May 2026
Safety

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs

DGX agent

arXiv:2506.12876v2 Announce Type: replace Abstract: The rapid scaling of large language models~(LLMs) has made inference efficiency a primary bottleneck in the practical deployment. To address this, s

safetyarxiv-cs-lg
14 May 2026
Safety

MinT: Managed Infrastructure for Training and Serving Millions of LLMs

DGX agent

arXiv:2605.13779v1 Announce Type: cross Abstract: We present MindLab Toolkit (MinT), a managed infrastructure system for Low-Rank Adaptation (LoRA) post-training and online serving. MinT targets a set

safetyarxiv-cs-ai
14 May 2026
Safety

MUJICA: Multi-skill Unified Joint Integration of Control Architecture for Wheeled-Legged Robots

DGX agent

arXiv:2605.13058v1 Announce Type: new Abstract: Wheeled-legged robots hold promise for traversing complex terrains and offer superior mobility compared to legged robots. However, wheeled-legged robots

safetyarxiv-cs-ro
14 May 2026
Safety

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization

DGX agent

arXiv:2605.13641v1 Announce Type: new Abstract: Complex reinforcement learning environments frequently employ multi-task and mixed-reward formulations. In these settings, heterogeneous reward distribu

safetyarxiv-cs-lg
14 May 2026
Safety

Multi-Rollout On-Policy Distillation via Peer Successes and Failures

DGX agent

arXiv:2605.12652v1 Announce Type: cross Abstract: Large language models are often post-trained with sparse verifier rewards, which indicate whether a sampled trajectory succeeds but provide limited gu

safetyarxiv-cs-ai
14 May 2026
Safety

Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy

DGX agent

arXiv:2605.12991v1 Announce Type: cross Abstract: LLM-based multi-agent pipelines flip from correct to incorrect answers under simulated peer disagreement at rates we term yield, a vulnerability widel

safetyarxiv-cs-ai
14 May 2026
Safety

ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization

DGX agent

arXiv:2605.12667v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) utilizes Reinforcement Learning from AI Feedback (RLAIF) for non-verifiable domains such as long-form qu

safetyarxiv-cs-ai
14 May 2026
Safety

On the Generalization of Knowledge Distillation: An Information-Theoretic View

DGX agent

arXiv:2605.13143v1 Announce Type: cross Abstract: Knowledge distillation is widely used to improve generalization in practice, yet its theoretical understanding remains elusive. In the standard distil

safetyarxiv-cs-lg
14 May 2026
Safety

On the Sample Complexity of Differentially Private Policy Optimization

DGX agent

arXiv:2510.21060v3 Announce Type: replace-cross Abstract: Policy optimization (PO) is a cornerstone of modern reinforcement learning (RL), with diverse applications spanning robotics, healthcare, and

safetyarxiv-cs-ai
14 May 2026
Safety

Oof. One of the few things Americans of all parties appear to agree on.

DGX agent

This post likely discusses a topic of broad bipartisan agreement among Americans, with Gary Marcus commenting on its significance on social media. The specific topic of agreement is not determinable f

safetygary-marcus--x
14 May 2026
Safety

OptMap: Geometric Map Distillation via Submodular Maximization

DGX agent

arXiv:2512.07775v2 Announce Type: replace Abstract: Autonomous robots rely on geometric maps to inform a diverse set of perception and decision-making algorithms. As autonomy requires reasoning and pl

safetyarxiv-cs-ro
14 May 2026
Safety

Pareto-Guided Optimal Transport for Multi-Reward Alignment

DGX agent

arXiv:2605.13155v1 Announce Type: new Abstract: Text-to-image generation models have achieved remarkable progress in preference optimization, yet achieving robust alignment across diverse reward model

safetyarxiv-cs-cv
14 May 2026
Safety

PG-LRF: Physiology-Guided Latent Rectified Flow for Electro-Hemodynamic PPG-to-ECG Generation

DGX agent

arXiv:2605.12541v1 Announce Type: cross Abstract: Electrocardiography (ECG) is the clinical standard for cardiac assessment but requires dedicated hardware that does not scale to daily-life monitoring

safetyarxiv-cs-ai
14 May 2026
Safety

Position: Assistive Agents Need Accessibility Alignment

DGX agent

arXiv:2605.13579v1 Announce Type: new Abstract: Assistive agents for Blind and Visually Impaired (BVI) users require accessibility alignment as a first-class design objective. Despite rapid progress i

safetyarxiv-cs-ai
14 May 2026
Safety

PRA-PoE: Robust Alzheimer's Diagnosis with Arbitrary Missing Modalities

DGX agent

arXiv:2605.13081v1 Announce Type: new Abstract: Missing modalities are prevalent in real-world Alzheimer's disease (AD) assessment and pose a significant challenge to multimodal learning, particularly

safetyarxiv-cs-cv
14 May 2026
Safety

Precautionary Governance of Autonomous AI: Legal Personhood as Functional Instrument

DGX agent

arXiv:2605.12505v1 Announce Type: cross Abstract: Autonomous AI systems generate responsibility gaps: consequential actions that cannot be satisfactorily attributed to developers, operators, or users

safetyarxiv-cs-ai
14 May 2026
Safety

Pretraining Language Models with Subword Regularization: An Empirical Study of BPE Dropout in Low-Resource NLP

DGX agent

arXiv:2605.13436v1 Announce Type: cross Abstract: Subword regularization methods such as BPE dropout are typically applied only during fine-tuning, while pretraining is usually done with deterministic

safetyarxiv-cs-lg
14 May 2026
Safety

Protocol-Driven Development: Governing Generated Software Through Invariants and Evidence

DGX agent

arXiv:2605.12981v1 Announce Type: cross Abstract: Automated program synthesis has reduced the cost of producing candidate implementations, but it introduces a harder governance problem: determining wh

safetyarxiv-cs-ai
14 May 2026
← Previous
1…210211212213214…302
Next →