AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA

DGX agent

arXiv:2604.13731v1 Announce Type: new Abstract: Multi-page Document Visual Question Answering requires reasoning over semantics, layouts, and visual elements in long, visually dense documents. Existin

safetyarxiv-cs-cl
16 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models

DGX agent

arXiv:2602.20981v3 Announce Type: replace Abstract: Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and

safetyarxiv-cs-cv
16 Apr 2026
Safety

Enhanced Text-to-Image Generation by Fine-grained Multimodal Reasoning

DGX agent

arXiv:2604.13491v1 Announce Type: new Abstract: With the rapid progress of Multimodal Large Language Models (MLLMs), unified MLLMs that jointly perform image understanding and generation have advanced

safetyarxiv-cs-cv
16 Apr 2026
Safety

Enhancing Reinforcement Learning for Radiology Report Generation with Evidence-aware Rewards and Self-correcting Preference Learning

DGX agent

arXiv:2604.13598v1 Announce Type: new Abstract: Recent reinforcement learning (RL) approaches have advanced radiology report generation (RRG), yet two core limitations persist: (1) report-level reward

safetyarxiv-cs-lg
16 Apr 2026
Safety

Estimating Continuous Treatment Effects with Two-Stage Kernel Ridge Regression

DGX agent

arXiv:2604.13410v1 Announce Type: cross Abstract: We study the problem of estimating the effect function for a continuous treatment, which maps each treatment value to a population-averaged outcome. A

safetyarxiv-cs-lg
16 Apr 2026
Safety

Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization

DGX agent

arXiv:2604.13533v1 Announce Type: cross Abstract: Achieving general-purpose robotics requires empowering robots to adapt and evolve based on their environment and feedback. Traditional methods face li

safetyarxiv-cs-cv
16 Apr 2026
Safety

Failure Identification in Imitation Learning Via Statistical and Semantic Filtering

DGX agent

arXiv:2604.13788v1 Announce Type: cross Abstract: Imitation learning (IL) policies in robotics deliver strong performance in controlled settings but remain brittle in real-world deployments: rare even

safetyarxiv-cs-cv
16 Apr 2026
Safety

FCBV-Net: Category-Level Robotic Garment Smoothing via Feature-Conditioned Bimanual Value Prediction

DGX agent

arXiv:2508.05153v2 Announce Type: replace Abstract: Category-level generalization for robotic garment manipulation, such as bimanual smoothing, remains a significant hurdle due to high dimensionality,

safetyarxiv-cs-ro
16 Apr 2026
Safety

First-See-Then-Design: A Multi-Stakeholder View for Optimal Performance-Fairness Trade-Offs

DGX agent

arXiv:2604.14035v1 Announce Type: new Abstract: Fairness in algorithmic decision-making is often defined in the predictive space, where predictive performance - used as a proxy for decision-maker (DM)

safetyarxiv-cs-lg
16 Apr 2026
Safety

Foresight Optimization for Strategic Reasoning in Large Language Models

DGX agent

arXiv:2604.13592v1 Announce Type: new Abstract: Reasoning capabilities in large language models (LLMs) have generally advanced significantly. However, it is still challenging for existing reasoning-ba

safetyarxiv-cs-cl
16 Apr 2026
Safety

From Alignment to Prediction: A Study of Self-Supervised Learning and Predictive Representation Learning

DGX agent

arXiv:2604.13518v1 Announce Type: new Abstract: Self-supervised learning has emerged as a major technique for the task of learning from unlabeled data, where the current methods mostly revolve around

safetyarxiv-cs-lg
16 Apr 2026
Safety

From Instruction to Event: Sound-Triggered Mobile Manipulation

DGX agent

arXiv:2601.21667v2 Announce Type: replace-cross Abstract: Current mobile manipulation research predominantly follows an instruction-driven paradigm, where agents rely on predefined textual commands to

safetyarxiv-cs-cv
16 Apr 2026
Safety

From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space

DGX agent

arXiv:2604.14142v1 Announce Type: cross Abstract: While reinforcement learning with verifiable rewards (RLVR) significantly enhances LLM reasoning by optimizing the conditional distribution P(y|x), it

safetyarxiv-cs-cl
16 Apr 2026
Safety

From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction

DGX agent

arXiv:2604.13067v1 Announce Type: cross Abstract: SpeechLLMs process spoken language directly from audio, but accent and vocal identity cues can lead to biased behaviour. Current bias evaluations ofte

safetyarxiv-cs-cl
16 Apr 2026
Safety

GRITS: A Spillage-Aware Guided Diffusion Policy for Robot Food Scooping Tasks

DGX agent

arXiv:2510.00573v2 Announce Type: replace Abstract: Robotic food scooping is a critical manipulation skill for food preparation and service robots. However, existing robot learning algorithms, especia

safetyarxiv-cs-ro
16 Apr 2026
Safety

HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy

DGX agent

arXiv:2510.00695v3 Announce Type: replace-cross Abstract: Inherently, robotic manipulation tasks are history-dependent: leveraging past context could be beneficial. However, most existing Vision-Langu

safetyarxiv-cs-cv
16 Apr 2026
Safety

Heavy-Tailed Class-Conditional Priors for Long-Tailed Generative Modeling

DGX agent

arXiv:2509.02154v2 Announce Type: replace-cross Abstract: Variational Autoencoders (VAEs) with global priors trained under an imbalanced empirical class distribution can lead to underrepresentation of

safetyarxiv-cs-cv
16 Apr 2026
Model Releases

HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

DGX agent

arXiv:2604.13954v1 Announce Type: new Abstract: Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions.

model-releasesarxiv-cs-lg
16 Apr 2026
Safety

IGen: Scalable Data Generation for Robot Learning from Open-World Images

DGX agent

arXiv:2512.01773v2 Announce Type: replace Abstract: The rise of generalist robotic policies has created an exponential demand for large-scale training data. However, on-robot data collection is labor-

safetyarxiv-cs-ro
16 Apr 2026
Safety

Irregularly Sampled Time Series Interpolation for Binary Evolution Simulations Using Dynamic Time Warping

DGX agent

arXiv:2604.13604v1 Announce Type: cross Abstract: Binary stellar evolution simulations are computationally expensive. Stellar population synthesis relies on these detailed evolution models at a fundam

safetyarxiv-cs-lg
16 Apr 2026
Safety

Jump-Start Reinforcement Learning with Vision-Language-Action Regularization

DGX agent

arXiv:2604.13733v1 Announce Type: new Abstract: Reinforcement learning (RL) enables high-frequency, closed-loop control for robotic manipulation, but scaling to long-horizon tasks with sparse or imper

safetyarxiv-cs-lg
16 Apr 2026
Safety

Learning Class Difficulty in Imbalanced Histopathology Segmentation via Dynamic Focal Attention

DGX agent

arXiv:2604.13479v1 Announce Type: cross Abstract: Semantic segmentation of histopathology images under class imbalance is typically addressed through frequency-based loss reweighting, which implicitly

safetyarxiv-cs-cv
16 Apr 2026
Safety

Med-CAM: Minimal Evidence for Explaining Medical Decision Making

DGX agent

arXiv:2604.13695v1 Announce Type: new Abstract: Reliable and interpretable decision-making is essential in medical imaging, where diagnostic outcomes directly influence patient care. Despite advances

safetyarxiv-cs-cv
16 Apr 2026
Safety

Neural Mean-Field Games: Extending Mean-Field Game Theory with Neural Stochastic Differential Equations

DGX agent

arXiv:2504.13228v4 Announce Type: replace Abstract: Mean-field game theory relies on approximating games that are intractable to model due to a very large to infinite population of players. While thes

safetyarxiv-cs-lg
16 Apr 2026
Safety

Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning

DGX agent

arXiv:2506.08125v3 Announce Type: replace-cross Abstract: Large language models (LLMs) show strong reasoning abilities but often produce unnecessarily long explanations that reduce efficiency. Althoug

safetyarxiv-cs-cl
16 Apr 2026
Safety

On the Fundamental Limitations of Dual Static CVaR Decompositions in Markov Decision Processes

DGX agent

arXiv:2507.14005v2 Announce Type: replace Abstract: It was recently shown that dynamic programming (DP) methods for finding static CVaR-optimal policies in Markov Decision Processes (MDPs) can fail wh

safetyarxiv-cs-lg
16 Apr 2026
Safety

Pareto-Optimal Offline Reinforcement Learning via Smooth Tchebysheff Scalarization

DGX agent

arXiv:2604.13175v1 Announce Type: new Abstract: Large language models can be aligned with human preferences through offline reinforcement learning (RL) on small labeled datasets. While single-objectiv

safetyarxiv-cs-lg
16 Apr 2026
Safety

Reducing Hallucinations in LLMs via Factuality-Aware Preference Learning

DGX agent

arXiv:2601.03027v3 Announce Type: replace Abstract: Preference alignment methods such as RLHF and Direct Preference Optimization (DPO) improve instruction following, but they can also reinforce halluc

safetyarxiv-cs-cl
16 Apr 2026
Safety

Representation over Routing: Overcoming Surrogate Hacking in Multi-Timescale PPO

DGX agent

arXiv:2604.13517v1 Announce Type: new Abstract: Temporal credit assignment in reinforcement learning has long been a central challenge. Inspired by the multi-timescale encoding of the dopamine system

safetyarxiv-cs-lg
16 Apr 2026
Safety

Restless Bandits with Individual Penalty Constraints: A New Near-Optimal Index Policy and How to Learn It

DGX agent

arXiv:2604.04101v2 Announce Type: replace Abstract: This paper investigates the Restless Multi-Armed Bandit (RMAB) framework under individual penalty constraints to address resource allocation challen

safetyarxiv-cs-lg
16 Apr 2026
Safety

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

DGX agent

arXiv:2508.00222v5 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Reward (RLVR) has significantly advanced the complex reasoning abilities of Large Language Models (LLMs

safetyarxiv-cs-cl
16 Apr 2026
Safety

RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph

DGX agent

arXiv:2511.07717v2 Announce Type: replace-cross Abstract: Estimating robot pose from a monocular RGB image is a challenge in robotics and computer vision. Existing methods typically build networks on

safetyarxiv-cs-cv
16 Apr 2026
Safety

SFT-GRPO Data Overlap as a Post-Training Hyperparameter for Autoformalization

DGX agent

arXiv:2604.13515v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) followed by Group Relative Policy Optimization (GRPO) is a common post-training recipe. We conduct a controlled ablation ov

safetyarxiv-cs-lg
16 Apr 2026
Safety

Soft Q(lambda): A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces

DGX agent

arXiv:2604.13780v1 Announce Type: new Abstract: Soft Q-learning has emerged as a versatile model-free method for entropy-regularised reinforcement learning, optimising for returns augmented with a pen

safetyarxiv-cs-lg
16 Apr 2026
Safety

Spectral methods: crucial for machine learning, natural for quantum computers?

DGX agent

arXiv:2603.24654v2 Announce Type: replace-cross Abstract: This article presents an argument for why quantum computers could unlock new methods for machine learning. We argue that spectral methods, in

safetyarxiv-cs-lg
16 Apr 2026
Safety

SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models

DGX agent

arXiv:2510.09541v3 Announce Type: replace Abstract: Diffusion large language models (dLLMs) are emerging as an efficient alternative to autoregressive models due to their ability to decode multiple to

safetyarxiv-cs-cl
16 Apr 2026
Safety

Temporally Consistent Long-Term Memory for 3D Single Object Tracking

DGX agent

arXiv:2604.13789v1 Announce Type: new Abstract: 3D Single Object Tracking (3D-SOT) aims to localize a target object across a sequence of LiDAR point clouds, given its 3D bounding box in the first fram

safetyarxiv-cs-cv
16 Apr 2026
Safety

The Long Delay to Arithmetic Generalization: When Learned Representations Outrun Behavior

DGX agent

arXiv:2604.13082v1 Announce Type: new Abstract: Grokking in transformers trained on algorithmic tasks is characterized by a long delay between training-set fit and abrupt generalization, but the sourc

safetyarxiv-cs-lg
16 Apr 2026
Safety

Think Outside the Policy: In-Context Steered Policy Optimization

DGX agent

arXiv:2510.26519v3 Announce Type: replace Abstract: Existing Reinforcement Learning from Verifiable Rewards (RLVR) methods, such as Group Relative Policy Optimization (GRPO), have achieved remarkable

safetyarxiv-cs-lg
16 Apr 2026
Safety

TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks

DGX agent

arXiv:2601.10245v2 Announce Type: replace-cross Abstract: Multi-step reasoning tasks like mathematical problem solving are vulnerable to cascading failures, where a single incorrect step leads to comp

safetyarxiv-cs-cl
16 Apr 2026
Safety

UMI-3D: Extending Universal Manipulation Interface from Vision-Limited to 3D Spatial Perception

DGX agent

arXiv:2604.14089v1 Announce Type: new Abstract: We present UMI-3D, a multimodal extension of the Universal Manipulation Interface (UMI) for robust and scalable data collection in embodied manipulation

safetyarxiv-cs-ro
16 Apr 2026
Safety

UNBOX: Unveiling Black-box visual models with Natural-language

DGX agent

arXiv:2603.08639v2 Announce Type: replace Abstract: Ensuring trustworthiness in open-world visual recognition requires models that are interpretable, fair, and robust to distribution shifts. Yet moder

safetyarxiv-cs-cv
16 Apr 2026
Safety

What Are We Really Measuring? Rethinking Dataset Bias in Web-Scale Natural Image Collections via Unsupervised Semantic Clustering

DGX agent

arXiv:2604.13610v1 Announce Type: new Abstract: In computer vision, a prevailing method for quantifying dataset bias is to train a model to distinguish between datasets. High classification accuracy i

safetyarxiv-cs-cv
16 Apr 2026
Safety

Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking

DGX agent

arXiv:2604.13776v1 Announce Type: cross Abstract: Watermarking is becoming the default mechanism for AI content authentication, with governance policies and frameworks referencing it as infrastructure

safetyarxiv-cs-cl
16 Apr 2026
Safety

Why MLLMs Struggle to Determine Object Orientations

DGX agent

arXiv:2604.13321v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) struggle with tasks that require reasoning about 2D object orientation in images, as documented in prior work.

safetyarxiv-cs-cv
16 Apr 2026
Safety

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks

DGX agent

arXiv:2604.13403v1 Announce Type: new Abstract: In-context learning (ICL) enables models to adapt to new tasks via inference-time demonstrations. Despite its success in large language models, the exte

safetyarxiv-cs-cv
16 Apr 2026
Safety

A Dataset and Evaluation for Complex 4D Markerless Human Motion Capture

DGX agent

arXiv:2604.12765v1 Announce Type: new Abstract: Marker-based motion capture (MoCap) systems have long been the gold standard for accurate 4D human modeling, yet their reliance on specialized hardware

safetyarxiv-cs-cv
15 Apr 2026
Safety

A document is worth a structured record: Principled inductive bias design for document recognition

DGX agent

arXiv:2507.08458v2 Announce Type: replace-cross Abstract: Many document types use intrinsic, convention-driven structures that serve to encode precise and structured information, such as the conventio

safetyarxiv-cs-ai
15 Apr 2026
← Previous
1…224225226227228…257
Next →