AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
All
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,812 results
Safety

History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions

DGX agent

arXiv:2605.13825v1 Announce Type: new Abstract: Frontier LLMs are increasingly deployed as agents that pick the next action after a long log of prior tool calls produced by the same or a different mod

safetyarxiv-cs-ai
14 May 2026
Safety
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Humanwashing -- It Should Leave You Feeling Dirty

DGX agent

arXiv:2605.13723v1 Announce Type: cross Abstract: The phrase 'human in the loop' is increasingly used to imply a sense of safety in relation to AI decision systems. It shouldn't. There are contexts wh

safetyarxiv-cs-ai
14 May 2026
Safety

Improving Classifier-Free Guidance of Flow Matching via Manifold Projection

DGX agent

arXiv:2601.21892v2 Announce Type: replace-cross Abstract: Classifier-free guidance (CFG) is a widely used technique for controllable generation in diffusion and flow-based models. Despite its empirica

safetyarxiv-cs-ai
14 May 2026
Safety

Improving Code Translation with Syntax-Guided and Semantic-aware Preference Optimization

DGX agent

arXiv:2605.13229v1 Announce Type: new Abstract: LLMs have shown immense potential for code translation, yet they often struggle to ensure both syntactic correctness and semantic consistency. While pre

safetyarxiv-cs-ai
14 May 2026
Safety

Improving Diffusion Posterior Samplers with Lagged Temporal Corrections for Image Restoration

DGX agent

arXiv:2605.12573v1 Announce Type: cross Abstract: Diffusion-based posterior sampling (PS) is a leading framework for imaging inverse problems, combining learned priors with measurement constraints. Ye

safetyarxiv-cs-ai
14 May 2026
Safety

Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling

DGX agent

arXiv:2605.13801v1 Announce Type: cross Abstract: As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of th

safetyarxiv-cs-ai
14 May 2026
Safety

In a policy paper, Anthropic urges the US and allies to enforce export controls, curb distillation attacks, and export US AI to hold the lead over China by 2028 (Anthropic)

DGX agent

Anthropic: In a policy paper, Anthropic urges the US and allies to enforce export controls, curb distillation attacks, and export US AI to hold the lead over China by 2028 — We're releasing a new pape

safetytechmeme
14 May 2026
Safety

In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores

DGX agent

arXiv:2605.12530v1 Announce Type: cross Abstract: LLM fairness should be evaluated through in-situ conversational behavior rather than standardized-test Q&A benchmarks. We show that the standardized-t

safetyarxiv-cs-ai
14 May 2026
Safety

In the @nytimes, Media Lab Prof. @kesvelt and other scientists call for stronger oversight and regulation of AI technologies, including chat…

DGX agent

In the @nytimes, Media Lab Prof. @kesvelt and other scientists call for stronger oversight and regulation of AI technologies, including chatbots that can provide information on producing lethal biolog

safetygary-marcus--x
14 May 2026
Safety

Integration of an Agent Model into an Open Simulation Architecture for Scenario-Based Testing of Automated Vehicles

DGX agent

arXiv:2605.13539v1 Announce Type: new Abstract: Simulative and scenario-based testing are crucial methods in the safety assurance for automated driving systems. To ensure that simulation results are r

safetyarxiv-cs-ro
14 May 2026
Safety

Interesting position paper on agentic AI as a foreseeable pathway to AGI. (bookmark it) There has been strong debate on whether a larger sin…

DGX agent

Interesting position paper on agentic AI as a foreseeable pathway to AGI. (bookmark it) There has been strong debate on whether a larger single model get us there or a multi-agent system. The authors

safetydair-ai--x
14 May 2026
Safety

interwhen: A Generalizable Framework for Steering Reasoning Models with Test-time Verification

DGX agent

arXiv:2602.11202v3 Announce Type: replace-cross Abstract: Reasoning models produce long traces of intermediate decisions and tool calls, making test-time verification important for ensuring correctnes

safetyarxiv-cs-ai
14 May 2026
Safety

Is Video Anomaly Detection Misframed? Evidence from LLM-Based and Multi-Scene Models

DGX agent

arXiv:2605.12725v1 Announce Type: new Abstract: Recent video anomaly detection research has expanded rapidly with an emphasis on general models of normality intended to work across many different scen

safetyarxiv-cs-cv
14 May 2026
Safety

Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

DGX agent

arXiv:2605.12729v1 Announce Type: cross Abstract: Large language models are increasingly being used to support network operations (NetOps) and artificial intelligence for IT operations (AIOps), includ

safetyarxiv-cs-ai
14 May 2026
Safety

Learning to Decide with AI Assistance under Human-Alignment

DGX agent

arXiv:2605.12646v1 Announce Type: cross Abstract: It is widely agreed that when AI models assist decision-makers in high-stakes domains by predicting an outcome of interest, they should communicate th

safetyarxiv-cs-ai
14 May 2026
Safety

Learning Transferable Latent User Preferences for Human-Aligned Decision Making

DGX agent

arXiv:2605.12682v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as reasoning modules in many applications. While they are efficient in certain tasks, LLMs often stru

safetyarxiv-cs-ai
14 May 2026
Safety

Learning When to Act: Communication-Efficient Reinforcement Learning via Run-Time Assurance

DGX agent

arXiv:2605.12561v1 Announce Type: new Abstract: Safe reinforcement learning (RL) typically asks extit{what} an agent should do. We ask extit{when} it needs to act, and show that a single policy can jo

safetyarxiv-cs-lg
14 May 2026
Safety

Learning with Rare Success but Rich Feedback via Reflection-Enhanced Self-Distillation

DGX agent

arXiv:2605.12741v1 Announce Type: new Abstract: Enabling Large Language Models (LLMs) to continuously improve from environmental interactions is a central challenge in post-training. While on-policy s

safetyarxiv-cs-lg
14 May 2026
Safety

Loiter UAV Reinsertion Guidance for Fixed-wing UAV Corridors

DGX agent

arXiv:2605.13822v1 Announce Type: new Abstract: This paper considers fixed-wing unmanned aerial vehicle (UAV) corridors comprising a main lane, a circular loiter lane for managing traffic congestion,

safetyarxiv-cs-ro
14 May 2026
Safety

Macro-Action Based Multi-Agent Instruction Following through Value Cancellation

DGX agent

arXiv:2605.12655v1 Announce Type: new Abstract: Multi-agent reinforcement learning (MARL) in real-world use cases may need to adapt to external natural language instructions that interrupt ongoing beh

safetyarxiv-cs-ai
14 May 2026
Safety

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs

DGX agent

arXiv:2506.12876v2 Announce Type: replace Abstract: The rapid scaling of large language models~(LLMs) has made inference efficiency a primary bottleneck in the practical deployment. To address this, s

safetyarxiv-cs-lg
14 May 2026
Safety

MinT: Managed Infrastructure for Training and Serving Millions of LLMs

DGX agent

arXiv:2605.13779v1 Announce Type: cross Abstract: We present MindLab Toolkit (MinT), a managed infrastructure system for Low-Rank Adaptation (LoRA) post-training and online serving. MinT targets a set

safetyarxiv-cs-ai
14 May 2026
Safety

MoCCA: A Movable Circle Probability of Collision Approximation

DGX agent

arXiv:2605.13125v1 Announce Type: new Abstract: In automated driving, crash mitigation is crucial to ensure passenger safety. Accurate avoidance requires precise knowledge of the object's position and

safetyarxiv-cs-ro
14 May 2026
Safety

MUJICA: Multi-skill Unified Joint Integration of Control Architecture for Wheeled-Legged Robots

DGX agent

arXiv:2605.13058v1 Announce Type: new Abstract: Wheeled-legged robots hold promise for traversing complex terrains and offer superior mobility compared to legged robots. However, wheeled-legged robots

safetyarxiv-cs-ro
14 May 2026
Safety

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization

DGX agent

arXiv:2605.13641v1 Announce Type: new Abstract: Complex reinforcement learning environments frequently employ multi-task and mixed-reward formulations. In these settings, heterogeneous reward distribu

safetyarxiv-cs-lg
14 May 2026
Safety

Multi-Rollout On-Policy Distillation via Peer Successes and Failures

DGX agent

arXiv:2605.12652v1 Announce Type: cross Abstract: Large language models are often post-trained with sparse verifier rewards, which indicate whether a sampled trajectory succeeds but provide limited gu

safetyarxiv-cs-ai
14 May 2026
Safety

NeuroRisk: Physics-Informed Neural Optimization for Risk-Aware Traffic Engineering

DGX agent

arXiv:2605.12862v1 Announce Type: cross Abstract: In production Wide-Area Networks (WANs), correlated failures dominate availability losses, forcing operators to reserve large safety margins that leav

safetyarxiv-cs-lg
14 May 2026
Safety

No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills

DGX agent

arXiv:2605.13044v1 Announce Type: cross Abstract: LLM-powered agents can silently delete documents, leak credentials, or transfer funds on a routine user request, not because the agent was attacked, b

safetyarxiv-cs-ai
14 May 2026
Safety

Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy

DGX agent

arXiv:2605.12991v1 Announce Type: cross Abstract: LLM-based multi-agent pipelines flip from correct to incorrect answers under simulated peer disagreement at rates we term yield, a vulnerability widel

safetyarxiv-cs-ai
14 May 2026
Safety

ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization

DGX agent

arXiv:2605.12667v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) utilizes Reinforcement Learning from AI Feedback (RLAIF) for non-verifiable domains such as long-form qu

safetyarxiv-cs-ai
14 May 2026
Safety

On the Generalization of Knowledge Distillation: An Information-Theoretic View

DGX agent

arXiv:2605.13143v1 Announce Type: cross Abstract: Knowledge distillation is widely used to improve generalization in practice, yet its theoretical understanding remains elusive. In the standard distil

safetyarxiv-cs-lg
14 May 2026
Safety

On the Sample Complexity of Differentially Private Policy Optimization

DGX agent

arXiv:2510.21060v3 Announce Type: replace-cross Abstract: Policy optimization (PO) is a cornerstone of modern reinforcement learning (RL), with diverse applications spanning robotics, healthcare, and

safetyarxiv-cs-ai
14 May 2026
Safety

Oof. One of the few things Americans of all parties appear to agree on.

DGX agent

This post likely discusses a topic of broad bipartisan agreement among Americans, with Gary Marcus commenting on its significance on social media. The specific topic of agreement is not determinable f

safetygary-marcus--x
14 May 2026
Safety

OptMap: Geometric Map Distillation via Submodular Maximization

DGX agent

arXiv:2512.07775v2 Announce Type: replace Abstract: Autonomous robots rely on geometric maps to inform a diverse set of perception and decision-making algorithms. As autonomy requires reasoning and pl

safetyarxiv-cs-ro
14 May 2026
Safety

Pareto-Guided Optimal Transport for Multi-Reward Alignment

DGX agent

arXiv:2605.13155v1 Announce Type: new Abstract: Text-to-image generation models have achieved remarkable progress in preference optimization, yet achieving robust alignment across diverse reward model

safetyarxiv-cs-cv
14 May 2026
Safety

Perception with Guarantees: Certified Pose Estimation via Reachability Analysis

DGX agent

arXiv:2602.10032v2 Announce Type: replace Abstract: Agents in cyber-physical systems are increasingly entrusted with safety-critical tasks. Ensuring safety of these agents often requires localizing th

safetyarxiv-cs-cv
14 May 2026
Safety

PG-LRF: Physiology-Guided Latent Rectified Flow for Electro-Hemodynamic PPG-to-ECG Generation

DGX agent

arXiv:2605.12541v1 Announce Type: cross Abstract: Electrocardiography (ECG) is the clinical standard for cardiac assessment but requires dedicated hardware that does not scale to daily-life monitoring

safetyarxiv-cs-ai
14 May 2026
Safety

Position: Assistive Agents Need Accessibility Alignment

DGX agent

arXiv:2605.13579v1 Announce Type: new Abstract: Assistive agents for Blind and Visually Impaired (BVI) users require accessibility alignment as a first-class design objective. Despite rapid progress i

safetyarxiv-cs-ai
14 May 2026
Safety

PRA-PoE: Robust Alzheimer's Diagnosis with Arbitrary Missing Modalities

DGX agent

arXiv:2605.13081v1 Announce Type: new Abstract: Missing modalities are prevalent in real-world Alzheimer's disease (AD) assessment and pose a significant challenge to multimodal learning, particularly

safetyarxiv-cs-cv
14 May 2026
Safety

Precautionary Governance of Autonomous AI: Legal Personhood as Functional Instrument

DGX agent

arXiv:2605.12505v1 Announce Type: cross Abstract: Autonomous AI systems generate responsibility gaps: consequential actions that cannot be satisfactorily attributed to developers, operators, or users

safetyarxiv-cs-ai
14 May 2026
Safety

Pretraining Language Models with Subword Regularization: An Empirical Study of BPE Dropout in Low-Resource NLP

DGX agent

arXiv:2605.13436v1 Announce Type: cross Abstract: Subword regularization methods such as BPE dropout are typically applied only during fine-tuning, while pretraining is usually done with deterministic

safetyarxiv-cs-lg
14 May 2026
Safety

Protocol-Driven Development: Governing Generated Software Through Invariants and Evidence

DGX agent

arXiv:2605.12981v1 Announce Type: cross Abstract: Automated program synthesis has reduced the cost of producing candidate implementations, but it introduces a harder governance problem: determining wh

safetyarxiv-cs-ai
14 May 2026
Safety

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution

DGX agent

arXiv:2602.06239v2 Announce Type: replace Abstract: We introduce PEPO (Pessimistic Ensemble based Preference Optimization), a single-step Direct Preference Optimization (DPO)-like algorithm to mitigat

safetyarxiv-cs-lg
14 May 2026
Safety

Proximal-Based Generative Modeling for Bayesian Inverse Problems

DGX agent

arXiv:2605.13278v1 Announce Type: cross Abstract: Score-based diffusion models demonstrate superior performance in generative tasks but encounter fundamental bottlenecks in inverse problems due to the

safetyarxiv-cs-lg
14 May 2026
Safety

Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation

DGX agent

arXiv:2605.13111v1 Announce Type: new Abstract: Autoregressive video generation enables streaming and open-ended long video synthesis, but still suffers from long-term degradation caused by accumulate

safetyarxiv-cs-cv
14 May 2026
Safety

Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy

DGX agent

arXiv:2605.13435v1 Announce Type: cross Abstract: There is growing interest in utilizing flow-based models as decision-making policies in reinforcement learning due to their high expressive capacity.

safetyarxiv-cs-ai
14 May 2026
Safety

Quantifying LLM Safety Degradation Under Repeated Attacks Using Survival Analysis

DGX agent

arXiv:2605.12869v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in a wide range of applications, yet remain vulnerable to adversarial jailbreak attacks that ci

safetyarxiv-cs-ai
14 May 2026
Safety

Quantifying Sensitivity for Tree Ensembles: A symbolic and compositional approach

DGX agent

arXiv:2605.13830v1 Announce Type: new Abstract: Decision tree ensembles (DTE) are a popular model for a wide range of AI classification tasks, used in multiple safety critical domains, and hence verif

safetyarxiv-cs-ai
14 May 2026
← Previous
1…179180181182183…267
Next →