AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlog
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,485 results
Safety

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution

DGX agent

arXiv:2602.06239v2 Announce Type: replace Abstract: We introduce PEPO (Pessimistic Ensemble based Preference Optimization), a single-step Direct Preference Optimization (DPO)-like algorithm to mitigat

safetyarxiv-cs-lg
14 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Proximal-Based Generative Modeling for Bayesian Inverse Problems

DGX agent

arXiv:2605.13278v1 Announce Type: cross Abstract: Score-based diffusion models demonstrate superior performance in generative tasks but encounter fundamental bottlenecks in inverse problems due to the

safetyarxiv-cs-lg
14 May 2026
Safety

Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation

DGX agent

arXiv:2605.13111v1 Announce Type: new Abstract: Autoregressive video generation enables streaming and open-ended long video synthesis, but still suffers from long-term degradation caused by accumulate

safetyarxiv-cs-cv
14 May 2026
Safety

Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy

DGX agent

arXiv:2605.13435v1 Announce Type: cross Abstract: There is growing interest in utilizing flow-based models as decision-making policies in reinforcement learning due to their high expressive capacity.

safetyarxiv-cs-ai
14 May 2026
Safety

R-DMesh: Video-Guided 3D Animation via Rectified Dynamic Mesh Flow

DGX agent

arXiv:2605.13838v1 Announce Type: new Abstract: Video-guided 3D animation holds immense potential for content creation, offering intuitive and precise control over dynamic assets. However, practical d

safetyarxiv-cs-cv
14 May 2026
Safety

RAG-GNN: Integrating Retrieved Knowledge with Graph Neural Networks for Precision Medicine

DGX agent

arXiv:2602.00586v2 Announce Type: replace-cross Abstract: Network topology excels at structural predictions but fails to capture functional semantics encoded in biomedical literature. We present RAG-G

safetyarxiv-cs-ai
14 May 2026
Safety

Real2Sim: A Physics-driven and Editable Gaussian Splatting Framework for Autonomous Driving Scenes

DGX agent

arXiv:2605.13591v1 Announce Type: new Abstract: Reliable autonomous driving relies on large-scale, well-labeled data and robust models. However, manual data collection is resource-intensive, and tradi

safetyarxiv-cs-cv
14 May 2026
Safety

Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning

DGX agent

arXiv:2605.13255v1 Announce Type: new Abstract: On-policy self-distillation trains a reasoning model on its own rollouts while a teacher, often the same model conditioned on privileged context, provid

safetyarxiv-cs-ai
14 May 2026
Safety

Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency

DGX agent

arXiv:2605.13047v1 Announce Type: cross Abstract: Evaluating whether large vision-language models (VLMs) align with human perception for high-level semantic scene comprehension remains a challenge. Tr

safetyarxiv-cs-ai
14 May 2026
Safety

Revisiting DAgger in the Era of LLM-Agents

DGX agent

arXiv:2605.12913v1 Announce Type: new Abstract: Long-horizon LM agents learn from multi-turn interaction, where a single early mistake can alter the subsequent state distribution and derail the whole

safetyarxiv-cs-lg
14 May 2026
Safety

Reward-Weighted On-Policy Distillation with an Open Property-Equivalence Verifier for NL-to-SVA Generation

DGX agent

arXiv:2605.13501v1 Announce Type: cross Abstract: LLM-based generation of SystemVerilog Assertions (SVA) is often reported as nearing saturation, with the strongest specialized model reaching {sim}76%

safetyarxiv-cs-lg
14 May 2026
Safety

RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data

DGX agent

arXiv:2605.13775v1 Announce Type: cross Abstract: The scalability of robotic manipulation is fundamentally bottlenecked by the scarcity of task-aligned physical interaction data. While vision-language

safetyarxiv-cs-cv
14 May 2026
Safety

Robot Squid Game: Quadrupedal Locomotion for Traversing Narrow Tunnels

DGX agent

arXiv:2605.13665v1 Announce Type: new Abstract: Quadruped robots demonstrate exceptional potential for navigating complex terrain in critical applications such as search and rescue missions and infras

safetyarxiv-cs-ro
14 May 2026
Safety

ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic Profiles

DGX agent

arXiv:2605.13725v1 Announce Type: new Abstract: Large language model (LLM)-based multi-agent simulation offers a powerful testbed for studying social opinion dynamics. Yet current approaches often ado

safetyarxiv-cs-ai
14 May 2026
Safety

SECOND-Grasp: Semantic Contact-guided Dexterous Grasping

DGX agent

arXiv:2605.13117v1 Announce Type: cross Abstract: Achieving reliable robotic manipulation, such as dexterous grasping, requires a synergy between physically stable interactions and semantic task guida

safetyarxiv-cs-ai
14 May 2026
Safety

Selective Off-Policy Reference Tuning with Plan Guidance

DGX agent

arXiv:2605.11505v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT a

safetyarxiv-cs-ai
14 May 2026
Safety

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation

DGX agent

arXiv:2605.13554v1 Announce Type: cross Abstract: Contrastive reinforcement learning (CRL) learns goal-conditioned Q-values through a contrastive objective over state-action and goal representations,

safetyarxiv-cs-ai
14 May 2026
Safety

Sharpness-Guided Group Relative Policy Optimization via Probability Shaping

DGX agent

arXiv:2511.00066v4 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a practical route to improve large language model reasoning, and Group Relative Pol

safetyarxiv-cs-lg
14 May 2026
Safety

SID: Sliding into Distribution for Robust Few-Demonstration Manipulation

DGX agent

arXiv:2605.13428v1 Announce Type: new Abstract: Generalizing robotic manipulation across object poses, viewpoints, and dynamic disturbances is difficult, especially with only a few demonstrations. End

safetyarxiv-cs-ro
14 May 2026
Safety

SP-GCRL: Influence Maximization on Incomplete Social Graphs

DGX agent

arXiv:2605.12513v1 Announce Type: cross Abstract: Influence maximization (IM) in real platforms is challenged by incomplete, noisy social graphs and non-stationary diffusion dynamics. We propose SP-GC

safetyarxiv-cs-ai
14 May 2026
Safety

Spatiotemporal downscaling and nowcasting of urban land surface temperatures with deep neural networks

DGX agent

arXiv:2605.13566v1 Announce Type: new Abstract: Land Surface Temperature (LST) is a key variable for various applications, such as urban climate and ecology studies. Yet, existing satellite-derived LS

safetyarxiv-cs-lg
14 May 2026
Safety

Spectral Energy Centroid: a Metric for Improving Performance and Analyzing Spectral Bias in Implicit Neural Representations

DGX agent

arXiv:2605.12709v1 Announce Type: new Abstract: Implicit Neural Representations (INRs) model continuous signals using multilayer perceptrons (MLPs), enabling compact, differentiable, and high-fidelity

safetyarxiv-cs-lg
14 May 2026
Safety

STAR: Semantic-Temporal Adaptive Representation Learning for Few-Shot Action Recognition

DGX agent

arXiv:2605.13202v1 Announce Type: cross Abstract: Few-shot action recognition (FSAR) requires models to generalize to novel action categories from only a handful of annotated samples. Despite progress

safetyarxiv-cs-ai
14 May 2026
Safety

Structural Diversity Drives Disruptive Scientific Innovation

DGX agent

arXiv:2605.12514v1 Announce Type: cross Abstract: Scientific innovation increasingly depends on collaboration, yet the organizational structure that fosters breakthrough ideas remains poorly understoo

safetyarxiv-cs-cv
14 May 2026
Safety

Switching Successor Measures for Hierarchical Zero-shot Reinforcement Learning

DGX agent

arXiv:2605.13207v1 Announce Type: new Abstract: Hierarchical reinforcement learning can improve generalization by decomposing long-horizon decision-making into simpler subproblems. However, existing a

safetyarxiv-cs-lg
14 May 2026
Safety

Teacher-Guided Policy Optimization for LLM Distillation

DGX agent

arXiv:2605.13230v1 Announce Type: cross Abstract: The convergence of reinforcement learning and imitation learning has positioned Reverse KL (RKL) as a promising paradigm for on-policy LLM distillatio

safetyarxiv-cs-ai
14 May 2026
Safety

TeleGate: Whole-Body Humanoid Teleoperation via Gated Expert Selection with Motion Prior

DGX agent

arXiv:2602.09628v2 Announce Type: replace Abstract: Real-time whole-body teleoperation is a critical method for humanoid robots to perform complex tasks in unstructured environments. However, developi

safetyarxiv-cs-ro
14 May 2026
Safety

Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment

DGX agent

arXiv:2605.13537v1 Announce Type: cross Abstract: Inference-time alignment techniques offer a lightweight alternative or complement to costly reinforcement learning, while enabling continual adaptatio

safetyarxiv-cs-ai
14 May 2026
Safety

Test-time Offline Reinforcement Learning on Goal-related Experience

DGX agent

arXiv:2507.18809v2 Announce Type: replace Abstract: Foundation models compress a large amount of information in a single, large neural network, which can then be queried for individual tasks. There ar

safetyarxiv-cs-lg
14 May 2026
Safety

Test-time Sparsity for Extreme Fast Action Diffusion

DGX agent

arXiv:2605.13316v1 Announce Type: new Abstract: Action diffusion excels at high-fidelity action generation but incurs heavy computational costs owing to its iterative denoising nature. Despite current

safetyarxiv-cs-cv
14 May 2026
Safety

The End Justifies the Mean: A Linear Ranking Rule for Proportional Sequential Decisions

DGX agent

arXiv:2605.12717v1 Announce Type: cross Abstract: AI alignment and participatory design motivate a new democratic design problem: how to collectively choose a decision rule to use repeatedly. We study

safetyarxiv-cs-ai
14 May 2026
Safety

The Horizon Threshold in Cooperative Multi-Agent Reward-Free Exploration

DGX agent

arXiv:2602.01453v3 Announce Type: replace Abstract: We study cooperative multi-agent reinforcement learning in the setting of reward-free exploration, where multiple agents jointly explore an unknown

safetyarxiv-cs-lg
14 May 2026
Safety

Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents

DGX agent

arXiv:2605.12620v1 Announce Type: new Abstract: Building generalist embodied agents capable of solving complex real-world tasks remains a fundamental challenge in AI. Multimodal Large Language Models

safetyarxiv-cs-ai
14 May 2026
Safety

Tight Sample Complexity Bounds for Entropic Best Policy Identification

DGX agent

arXiv:2605.13717v1 Announce Type: new Abstract: We study best-policy identification for finite-horizon risk-sensitive reinforcement learning under the entropic risk measure. Recent work established a

safetyarxiv-cs-lg
14 May 2026
Safety

Topology-Preserving Neural Operator Learning via Hodge Decomposition

DGX agent

arXiv:2605.13834v1 Announce Type: cross Abstract: In this paper, we study solution operators of physical field equations on geometric meshes from a function-space perspective. We reveal that Hodge ort

safetyarxiv-cs-ai
14 May 2026
Safety

Towards a holistic understanding of Selection Bias for Causal Effect Identification

DGX agent

arXiv:2605.13430v1 Announce Type: cross Abstract: Selection bias is pervasive in observational studies. For example, large scale biobanks data can exhibit ``healthy volunteer bias'' when respondents a

safetyarxiv-cs-ai
14 May 2026
Safety

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning

DGX agent

arXiv:2602.06475v2 Announce Type: replace Abstract: Large language models (LLMs) excel at complex tasks with advances in reasoning capabilities. However, existing reward mechanisms remain tightly coup

safetyarxiv-cs-lg
14 May 2026
Safety

TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking

DGX agent

arXiv:2605.12587v1 Announce Type: new Abstract: Dense 3D tracking from monocular video is fundamental to dynamic scene understanding. While recent 3D foundation models provide reliable per-frame geome

safetyarxiv-cs-cv
14 May 2026
Safety

Trajectory-Level Data Augmentation for Offline Reinforcement Learning

DGX agent

arXiv:2605.13401v1 Announce Type: new Abstract: We propose a data augmentation method for offline reinforcement learning, motivated by active positioning problems. Particularly, our approach enables t

safetyarxiv-cs-lg
14 May 2026
Safety

Uncertainty-aware Spatial-Frequency Registration and Fusion for Infrared and Visible Images

DGX agent

arXiv:2605.13049v1 Announce Type: new Abstract: Infrared and Visible Image Fusion (IVIF) has shown promise in visual tasks under challenging environments, but fusion under unregistered conditions face

safetyarxiv-cs-cv
14 May 2026
Safety

Unifying Entropy Regularization in Optimal Control: From and Back to Classical Objectives via Iterated Soft Policies and Path Integral Solutions

DGX agent

arXiv:2512.06109v3 Announce Type: replace-cross Abstract: This paper develops a unified perspective on several optimal control formulations through the lens of Kullback-Leibler (KL) regularization. We

safetyarxiv-cs-lg
14 May 2026
Safety

UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning

DGX agent

arXiv:2510.10642v3 Announce Type: replace-cross Abstract: Building generalist robot policies that can handle diverse tasks in open-ended environments is a central challenge in robotics. To leverage kn

safetyarxiv-cs-ai
14 May 2026
Safety

Unweighted ranking for value-based decision making with uncertainty

DGX agent

arXiv:2605.13601v1 Announce Type: new Abstract: As intelligent systems are increasingly implemented in our society to make autonomous decisions, their commitment to human values raises serious concern

safetyarxiv-cs-ai
14 May 2026
Safety

VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority

DGX agent

arXiv:2605.12571v1 Announce Type: cross Abstract: Long video question answering requires locating sparse, time-scattered visual evidence within highly redundant content. Although current MLLMs perform

safetyarxiv-cs-ai
14 May 2026
Safety

wading through bots and LLM-written replies here is getting more tedious by the day. retweet if you agree.

DGX agent

Gary Marcus expresses frustration about the increasing prevalence of bot-generated and LLM-written responses on social media platforms, noting that filtering through such content has become increasing

safetygary-marcus--x
14 May 2026
Safety

WD-FQDet: Multispectral Detection Transformer via Wavelet Decomposition and Frequency-aware Query Learning

DGX agent

arXiv:2605.13621v1 Announce Type: new Abstract: Infrared-visible object detection improves detection performance by combining complementary features from multispectral images. Existing backbone-specif

safetyarxiv-cs-cv
14 May 2026
Safety

“we are working harder to manage our tools than we are to solve the actual problems they were meant to fix.”

DGX agent

“we are working harder to manage our tools than we are to solve the actual problems they were meant to fix.” Harvard Business Review research reveals that excessive interaction with AI is causing a sp

safetygary-marcus--x
14 May 2026
Safety

What properties of reasoning supervision are associated with improved downstream model quality?

DGX agent

arXiv:2605.13290v1 Announce Type: new Abstract: Validating training data for reasoning models typically requires expensive trial-and-error fine-tuning cycles. In this work, we investigate whether the

safetyarxiv-cs-ai
14 May 2026
← Previous
1…211212213214215…302
Next →