AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,806 results
14 May 2026

In a policy paper, Anthropic urges the US and allies to enforce export controls, curb distillation attacks, and export US AI to hold the lead over China by 2028 (Anthropic)

SafetyDGX agent

Anthropic: In a policy paper, Anthropic urges the US and allies to enforce export controls, curb distillation attacks, and export US AI to hold the lead over China by 2028 — We're releasing a new pape

In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores

SafetyDGX agent

arXiv:2605.12530v1 Announce Type: cross Abstract: LLM fairness should be evaluated through in-situ conversational behavior rather than standardized-test Q&A benchmarks. We show that the standardized-t

In the @nytimes, Media Lab Prof. @kesvelt and other scientists call for stronger oversight and regulation of AI technologies, including chat…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety
DGX agent

In the @nytimes, Media Lab Prof. @kesvelt and other scientists call for stronger oversight and regulation of AI technologies, including chatbots that can provide information on producing lethal biolog

Integration of an Agent Model into an Open Simulation Architecture for Scenario-Based Testing of Automated Vehicles

SafetyDGX agent

arXiv:2605.13539v1 Announce Type: new Abstract: Simulative and scenario-based testing are crucial methods in the safety assurance for automated driving systems. To ensure that simulation results are r

Interesting position paper on agentic AI as a foreseeable pathway to AGI. (bookmark it) There has been strong debate on whether a larger sin…

SafetyDGX agent

Interesting position paper on agentic AI as a foreseeable pathway to AGI. (bookmark it) There has been strong debate on whether a larger single model get us there or a multi-agent system. The authors

interwhen: A Generalizable Framework for Steering Reasoning Models with Test-time Verification

SafetyDGX agent

arXiv:2602.11202v3 Announce Type: replace-cross Abstract: Reasoning models produce long traces of intermediate decisions and tool calls, making test-time verification important for ensuring correctnes

Is Video Anomaly Detection Misframed? Evidence from LLM-Based and Multi-Scene Models

SafetyDGX agent

arXiv:2605.12725v1 Announce Type: new Abstract: Recent video anomaly detection research has expanded rapidly with an emphasis on general models of normality intended to work across many different scen

Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety

SafetyDGX agent

arXiv:2605.12729v1 Announce Type: cross Abstract: Large language models are increasingly being used to support network operations (NetOps) and artificial intelligence for IT operations (AIOps), includ

Learning to Decide with AI Assistance under Human-Alignment

SafetyDGX agent

arXiv:2605.12646v1 Announce Type: cross Abstract: It is widely agreed that when AI models assist decision-makers in high-stakes domains by predicting an outcome of interest, they should communicate th

Learning Transferable Latent User Preferences for Human-Aligned Decision Making

SafetyDGX agent

arXiv:2605.12682v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as reasoning modules in many applications. While they are efficient in certain tasks, LLMs often stru

Learning When to Act: Communication-Efficient Reinforcement Learning via Run-Time Assurance

SafetyDGX agent

arXiv:2605.12561v1 Announce Type: new Abstract: Safe reinforcement learning (RL) typically asks extit{what} an agent should do. We ask extit{when} it needs to act, and show that a single policy can jo

Learning with Rare Success but Rich Feedback via Reflection-Enhanced Self-Distillation

SafetyDGX agent

arXiv:2605.12741v1 Announce Type: new Abstract: Enabling Large Language Models (LLMs) to continuously improve from environmental interactions is a central challenge in post-training. While on-policy s

Loiter UAV Reinsertion Guidance for Fixed-wing UAV Corridors

SafetyDGX agent

arXiv:2605.13822v1 Announce Type: new Abstract: This paper considers fixed-wing unmanned aerial vehicle (UAV) corridors comprising a main lane, a circular loiter lane for managing traffic congestion,

Macro-Action Based Multi-Agent Instruction Following through Value Cancellation

SafetyDGX agent

arXiv:2605.12655v1 Announce Type: new Abstract: Multi-agent reinforcement learning (MARL) in real-world use cases may need to adapt to external natural language instructions that interrupt ongoing beh

MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs

SafetyDGX agent

arXiv:2506.12876v2 Announce Type: replace Abstract: The rapid scaling of large language models~(LLMs) has made inference efficiency a primary bottleneck in the practical deployment. To address this, s

MinT: Managed Infrastructure for Training and Serving Millions of LLMs

SafetyDGX agent

arXiv:2605.13779v1 Announce Type: cross Abstract: We present MindLab Toolkit (MinT), a managed infrastructure system for Low-Rank Adaptation (LoRA) post-training and online serving. MinT targets a set

MoCCA: A Movable Circle Probability of Collision Approximation

SafetyDGX agent

arXiv:2605.13125v1 Announce Type: new Abstract: In automated driving, crash mitigation is crucial to ensure passenger safety. Accurate avoidance requires precise knowledge of the object's position and

MUJICA: Multi-skill Unified Joint Integration of Control Architecture for Wheeled-Legged Robots

SafetyDGX agent

arXiv:2605.13058v1 Announce Type: new Abstract: Wheeled-legged robots hold promise for traversing complex terrains and offer superior mobility compared to legged robots. However, wheeled-legged robots

Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization

SafetyDGX agent

arXiv:2605.13641v1 Announce Type: new Abstract: Complex reinforcement learning environments frequently employ multi-task and mixed-reward formulations. In these settings, heterogeneous reward distribu

Multi-Rollout On-Policy Distillation via Peer Successes and Failures

SafetyDGX agent

arXiv:2605.12652v1 Announce Type: cross Abstract: Large language models are often post-trained with sparse verifier rewards, which indicate whether a sampled trajectory succeeds but provide limited gu

NeuroRisk: Physics-Informed Neural Optimization for Risk-Aware Traffic Engineering

SafetyDGX agent

arXiv:2605.12862v1 Announce Type: cross Abstract: In production Wide-Area Networks (WANs), correlated failures dominate availability losses, forcing operators to reserve large safety margins that leav

No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills

SafetyDGX agent

arXiv:2605.13044v1 Announce Type: cross Abstract: LLM-powered agents can silently delete documents, leak credentials, or transfer funds on a routine user request, not because the agent was attacked, b

Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy

SafetyDGX agent

arXiv:2605.12991v1 Announce Type: cross Abstract: LLM-based multi-agent pipelines flip from correct to incorrect answers under simulated peer disagreement at rates we term yield, a vulnerability widel

ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization

SafetyDGX agent

arXiv:2605.12667v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) utilizes Reinforcement Learning from AI Feedback (RLAIF) for non-verifiable domains such as long-form qu

On the Generalization of Knowledge Distillation: An Information-Theoretic View

SafetyDGX agent

arXiv:2605.13143v1 Announce Type: cross Abstract: Knowledge distillation is widely used to improve generalization in practice, yet its theoretical understanding remains elusive. In the standard distil

On the Sample Complexity of Differentially Private Policy Optimization

SafetyDGX agent

arXiv:2510.21060v3 Announce Type: replace-cross Abstract: Policy optimization (PO) is a cornerstone of modern reinforcement learning (RL), with diverse applications spanning robotics, healthcare, and

Oof. One of the few things Americans of all parties appear to agree on.

SafetyDGX agent

This post likely discusses a topic of broad bipartisan agreement among Americans, with Gary Marcus commenting on its significance on social media. The specific topic of agreement is not determinable f

OptMap: Geometric Map Distillation via Submodular Maximization

SafetyDGX agent

arXiv:2512.07775v2 Announce Type: replace Abstract: Autonomous robots rely on geometric maps to inform a diverse set of perception and decision-making algorithms. As autonomy requires reasoning and pl

Pareto-Guided Optimal Transport for Multi-Reward Alignment

SafetyDGX agent

arXiv:2605.13155v1 Announce Type: new Abstract: Text-to-image generation models have achieved remarkable progress in preference optimization, yet achieving robust alignment across diverse reward model

Perception with Guarantees: Certified Pose Estimation via Reachability Analysis

SafetyDGX agent

arXiv:2602.10032v2 Announce Type: replace Abstract: Agents in cyber-physical systems are increasingly entrusted with safety-critical tasks. Ensuring safety of these agents often requires localizing th

PG-LRF: Physiology-Guided Latent Rectified Flow for Electro-Hemodynamic PPG-to-ECG Generation

SafetyDGX agent

arXiv:2605.12541v1 Announce Type: cross Abstract: Electrocardiography (ECG) is the clinical standard for cardiac assessment but requires dedicated hardware that does not scale to daily-life monitoring

Position: Assistive Agents Need Accessibility Alignment

SafetyDGX agent

arXiv:2605.13579v1 Announce Type: new Abstract: Assistive agents for Blind and Visually Impaired (BVI) users require accessibility alignment as a first-class design objective. Despite rapid progress i

PRA-PoE: Robust Alzheimer's Diagnosis with Arbitrary Missing Modalities

SafetyDGX agent

arXiv:2605.13081v1 Announce Type: new Abstract: Missing modalities are prevalent in real-world Alzheimer's disease (AD) assessment and pose a significant challenge to multimodal learning, particularly

Precautionary Governance of Autonomous AI: Legal Personhood as Functional Instrument

SafetyDGX agent

arXiv:2605.12505v1 Announce Type: cross Abstract: Autonomous AI systems generate responsibility gaps: consequential actions that cannot be satisfactorily attributed to developers, operators, or users

Pretraining Language Models with Subword Regularization: An Empirical Study of BPE Dropout in Low-Resource NLP

SafetyDGX agent

arXiv:2605.13436v1 Announce Type: cross Abstract: Subword regularization methods such as BPE dropout are typically applied only during fine-tuning, while pretraining is usually done with deterministic

Protocol-Driven Development: Governing Generated Software Through Invariants and Evidence

SafetyDGX agent

arXiv:2605.12981v1 Announce Type: cross Abstract: Automated program synthesis has reduced the cost of producing candidate implementations, but it introduces a harder governance problem: determining wh

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution

SafetyDGX agent

arXiv:2602.06239v2 Announce Type: replace Abstract: We introduce PEPO (Pessimistic Ensemble based Preference Optimization), a single-step Direct Preference Optimization (DPO)-like algorithm to mitigat

Proximal-Based Generative Modeling for Bayesian Inverse Problems

SafetyDGX agent

arXiv:2605.13278v1 Announce Type: cross Abstract: Score-based diffusion models demonstrate superior performance in generative tasks but encounter fundamental bottlenecks in inverse problems due to the

Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation

SafetyDGX agent

arXiv:2605.13111v1 Announce Type: new Abstract: Autoregressive video generation enables streaming and open-ended long video synthesis, but still suffers from long-term degradation caused by accumulate

Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy

SafetyDGX agent

arXiv:2605.13435v1 Announce Type: cross Abstract: There is growing interest in utilizing flow-based models as decision-making policies in reinforcement learning due to their high expressive capacity.

Quantifying LLM Safety Degradation Under Repeated Attacks Using Survival Analysis

SafetyDGX agent

arXiv:2605.12869v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in a wide range of applications, yet remain vulnerable to adversarial jailbreak attacks that ci

Quantifying Sensitivity for Tree Ensembles: A symbolic and compositional approach

SafetyDGX agent

arXiv:2605.13830v1 Announce Type: new Abstract: Decision tree ensembles (DTE) are a popular model for a wide range of AI classification tasks, used in multiple safety critical domains, and hence verif

Quantitative Certification of Agentic Tool Selection

SafetyDGX agent

arXiv:2510.03992v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly deployed in agentic systems, where a fundamental task is mapping user intents to relevant extern

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow

SafetyDGX agent

arXiv:2605.13838v1 Announce Type: new Abstract: Video-guided 3D animation holds immense potential for content creation, offering intuitive and precise control over dynamic assets. However, practical d

RAG-GNN: Integrating Retrieved Knowledge with Graph Neural Networks for Precision Medicine

SafetyDGX agent

arXiv:2602.00586v2 Announce Type: replace-cross Abstract: Network topology excels at structural predictions but fails to capture functional semantics encoded in biomedical literature. We present RAG-G

Real2Sim: A Physics-driven and Editable Gaussian Splatting Framework for Autonomous Driving Scenes

SafetyDGX agent

arXiv:2605.13591v1 Announce Type: new Abstract: Reliable autonomous driving relies on large-scale, well-labeled data and robust models. However, manual data collection is resource-intensive, and tradi

Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning

SafetyDGX agent

arXiv:2605.13255v1 Announce Type: new Abstract: On-policy self-distillation trains a reasoning model on its own rollouts while a teacher, often the same model conditioned on privileged context, provid

Revealing Interpretable Failure Modes of VLMs

SafetyDGX agent

arXiv:2605.12674v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are increasingly used in safety-critical applications because of their broad reasoning capabilities and ability to general

Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency

SafetyDGX agent

arXiv:2605.13047v1 Announce Type: cross Abstract: Evaluating whether large vision-language models (VLMs) align with human perception for high-level semantic scene comprehension remains a challenge. Tr

Revisiting DAgger in the Era of LLM-Agents

SafetyDGX agent

arXiv:2605.12913v1 Announce Type: new Abstract: Long-horizon LM agents learn from multi-turn interaction, where a single early mistake can alter the subsequent state distribution and derail the whole

Reward-Weighted On-Policy Distillation with an Open Property-Equivalence Verifier for NL-to-SVA Generation

SafetyDGX agent

arXiv:2605.13501v1 Announce Type: cross Abstract: LLM-based generation of SystemVerilog Assertions (SVA) is often reported as nearing saturation, with the strongest specialized model reaching {sim}76%

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data

SafetyDGX agent

arXiv:2605.13775v1 Announce Type: cross Abstract: The scalability of robotic manipulation is fundamentally bottlenecked by the scarcity of task-aligned physical interaction data. While vision-language

Robot Squid Game: Quadrupedal Locomotion for Traversing Narrow Tunnels

SafetyDGX agent

arXiv:2605.13665v1 Announce Type: new Abstract: Quadruped robots demonstrate exceptional potential for navigating complex terrain in critical applications such as search and rescue missions and infras

Robust and Safe Multi-Agent Reinforcement Learning with Communication for Autonomous Vehicles: From Simulation to Hardware

SafetyDGX agent

arXiv:2506.00982v3 Announce Type: replace Abstract: Deep multi-agent reinforcement learning (MARL) has been demonstrated effectively in simulations for multi-robot problems. For autonomous vehicles, t

Runtime Monitoring of Perception-Based Autonomous Systems via Embedding Temporal Logic

SafetyDGX agent

arXiv:2605.12651v1 Announce Type: new Abstract: Runtime monitoring of autonomous systems traditionally relies on mapping continuous sensor observations to discrete logical propositions defined over lo

ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic Profiles

SafetyDGX agent

arXiv:2605.13725v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent simulation offers a powerful testbed for studying social opinion dynamics. Yet current approaches often ado

SECOND-Grasp: Semantic Contact-guided Dexterous Grasping

SafetyDGX agent

arXiv:2605.13117v1 Announce Type: cross Abstract: Achieving reliable robotic manipulation, such as dexterous grasping, requires a synergy between physically stable interactions and semantic task guida

Selective Off-Policy Reference Tuning with Plan Guidance

SafetyDGX agent

arXiv:2605.11505v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT a

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation

SafetyDGX agent

arXiv:2605.13554v1 Announce Type: cross Abstract: Contrastive reinforcement learning (CRL) learns goal-conditioned Q-values through a contrastive objective over state-action and goal representations,

Sharpness-Guided Group Relative Policy Optimization via Probability Shaping

SafetyDGX agent

arXiv:2511.00066v4 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a practical route to improve large language model reasoning, and Group Relative Pol

← Previous
1…143144145146147…214
Next →