AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

Binary Rewards and Reinforcement Learning: Fundamental Challenges

DGX agent

arXiv:2605.02375v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard approach for improving reasoning in language models, yet models trained with

safetyarxiv-cs-lg
5 May 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Bolek: A Multimodal Language Model for Molecular Reasoning

DGX agent

arXiv:2605.02745v1 Announce Type: new Abstract: Molecular property models increasingly support high-stakes drug-discovery decisions, but their outputs are often difficult to audit: classical predictor

safetyarxiv-cs-lg
5 May 2026
Safety

Breaking the Computational Barrier: Provably Efficient Actor-Critic for Low-Rank MDPs

DGX agent

arXiv:2605.01242v1 Announce Type: new Abstract: Reinforcement learning (RL) is a fundamental framework for sequential decision-making, in which an agent learns an optimal policy through interactions w

safetyarxiv-cs-lg
5 May 2026
Safety

Bridging the Gap Between Average and Discounted TD Learning

DGX agent

arXiv:2605.02103v1 Announce Type: new Abstract: The analysis of Temporal Difference (TD) learning in the average-reward setting faces notable theoretical difficulties because the Bellman operator is n

safetyarxiv-cs-lg
5 May 2026
Safety

Bringing Order to Asynchronous SGD: Towards Optimality under Data-Dependent Delays with Momentum

DGX agent

arXiv:2605.02043v1 Announce Type: new Abstract: Asynchronous stochastic gradient descent (SGD) enables scalable distributed training but suffers from gradient staleness. Existing mitigation strategies

safetyarxiv-cs-lg
5 May 2026
Safety

Colinearity Decay: Training Quantization-Friendly ViTs with Outlier Decay

DGX agent

arXiv:2605.01330v1 Announce Type: new Abstract: Low-bit quantization is a practical route for efficiently deploying vision Transformers, yet activation outliers complicate fully quantized deployment.

safetyarxiv-cs-cv
5 May 2026
Safety

Combining Trained Models in Reinforcement Learning

DGX agent

arXiv:2605.02159v1 Announce Type: new Abstract: Deep reinforcement learning (DRL) has delivered strong results in domains such as Atari and Go, but it still suffers from high sample cost and weak tran

safetyarxiv-cs-lg
5 May 2026
Safety

Compared to What? Baselines and Metrics for Counterfactual Prompting

DGX agent

arXiv:2605.01048v1 Announce Type: new Abstract: Counterfactual prompting (i.e., perturbing a single factor and measuring output change) is widely used to evaluate things like LLM bias and CoT faithful

safetyarxiv-cs-cl
5 May 2026
Safety

Compliance-Aware Agentic Payments on Stablecoin Rails

DGX agent

arXiv:2605.00071v1 Announce Type: cross Abstract: Agentic payment systems extend delegated action to financial transfers, but scaling them on stablecoin rails in regulated settings requires safeguards

safetyarxiv-cs-ai
5 May 2026
Safety

Contrastive Residual Energy Test-time Adaptation

DGX agent

arXiv:2505.19607v2 Announce Type: replace Abstract: Test-time adaptation (TTA) enhances model robustness by enabling adaptation to target distributions that differ from training distributions, improvi

safetyarxiv-cs-lg
5 May 2026
Safety

CUE: Concept-Aware Multi-Label Expansion to Mitigate Concept Confusion in Long-Tailed Learning

DGX agent

arXiv:2605.01309v1 Announce Type: new Abstract: Long-tailed distributions are common in real-world recognition tasks, where a few head classes have many samples while most tail classes have very few.

safetyarxiv-cs-cv
5 May 2026
Safety

CycleRL: Sim-to-Real Deep Reinforcement Learning for Robust Autonomous Bicycle Control

DGX agent

arXiv:2603.15013v2 Announce Type: replace Abstract: Autonomous bicycles offer a promising agile solution for urban mobility and last-mile logistics. However, conventional control strategies often stru

safetyarxiv-cs-ro
5 May 2026
Safety

Decision Boundary-aware Generation for Long-tailed Learning

DGX agent

arXiv:2605.01468v1 Announce Type: new Abstract: Long-tailed data bias decision boundaries toward head classes and degrade tail class accuracy. Diffusion-based generative augmentation address this prob

safetyarxiv-cs-cv
5 May 2026
Safety

DeepStage: Learning Autonomous Defense Policies Against Multi-Stage APT Campaigns

DGX agent

arXiv:2603.16969v2 Announce Type: replace-cross Abstract: This paper presents DeepStage, a deep reinforcement learning (DRL) framework for adaptive and stage-aware defense against Advanced Persistent

safetyarxiv-cs-lg
5 May 2026
Safety

Delayed homomorphic reinforcement learning for environments with delayed feedback

DGX agent

arXiv:2604.03641v2 Announce Type: replace Abstract: Reinforcement learning in real-world systems often involves delayed feedback, which breaks the Markov assumption and impedes both learning and contr

safetyarxiv-cs-lg
5 May 2026
Safety

Differential Parity: Relative Fairness Between Two Sets of Decisions

DGX agent

arXiv:2112.11279v4 Announce Type: replace Abstract: With AI systems widely applied to assist humans in decision-making processes such as talent hiring, school admission, and loan approval; there is an

safetyarxiv-cs-lg
5 May 2026
Safety

Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models

DGX agent

arXiv:2605.01896v1 Announce Type: new Abstract: Emerging multi-modal world models attempt to jointly generate videos across diverse modalities (e.g., RGB, depth, and mask), yet they fail to fully expl

safetyarxiv-cs-cv
5 May 2026
Safety

Do Large Language Models Plan Answer Positions? Position Bias in Multiple-Choice Question Generation

DGX agent

arXiv:2605.01846v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate multiple-choice questions (MCQs), where correct answers should ideally be uniformly distr

safetyarxiv-cs-cl
5 May 2026
Safety

DR-SNE: Density-Regularized Stochastic Neighbor Embedding

DGX agent

arXiv:2605.02060v1 Announce Type: new Abstract: Dimensionality reduction methods such as t-SNE are designed to preserve local neighborhood structure but do not explicitly account for how probability m

safetyarxiv-cs-lg
5 May 2026
Safety

Dynamics Aware Quadrupedal Locomotion via Intrinsic Dynamics Head

DGX agent

arXiv:2605.01227v1 Announce Type: new Abstract: Quadrupedal locomotion plays a critical role in enabling agile, versatile movement across complex terrains. Understanding and estimating the underlying

safetyarxiv-cs-ro
5 May 2026
Safety

Dynamics Distillation for Efficient and Transferable Control Learning

DGX agent

arXiv:2605.01516v1 Announce Type: new Abstract: Robust control policy learning for autonomous driving requires training environments to be both physically realistic and computationally scalable, prope

safetyarxiv-cs-ro
5 May 2026
Safety

Exploring Data-Free LoRA Transferability for Video Diffusion Models

DGX agent

arXiv:2605.01929v1 Announce Type: new Abstract: Video diffusion models leveraging step distillation or causal distillation have achieved remarkable performance. However, adapting existing LoRAs to the

safetyarxiv-cs-cv
5 May 2026
Safety

Exploring Entropy-based Active Learning for Fair Brain Segmentation

DGX agent

arXiv:2605.01706v1 Announce Type: new Abstract: Active learning (AL) has emerged as a crucial strategy for reducing the prohibitive costs associated with medical image segmentation. However, standard

safetyarxiv-cs-cv
5 May 2026
Safety

Exploring Prompt Alignment with Clinical Factors in Zero-Shot Segmentation VLMs for NSCLC Tumor Segmentation

DGX agent

arXiv:2605.01266v1 Announce Type: new Abstract: Zero-shot vision-language models (VLMs) offer a promptable alternative to task-specific training for gross tumor volume (GTV) delineation in non-small-c

safetyarxiv-cs-cv
5 May 2026
Safety

ExpoCM: Exposure-Aware One-Step Generative Single-Image HDR Reconstruction

DGX agent

arXiv:2605.02464v1 Announce Type: new Abstract: Single-image HDR reconstruction aims to recover high dynamic range radiance from a single low dynamic range (LDR) input, but remains highly ill-posed du

safetyarxiv-cs-cv
5 May 2026
Safety

FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control

DGX agent

arXiv:2603.12612v2 Announce Type: replace Abstract: Scaling Maximum Entropy Reinforcement Learning (RL) to high-dimensional humanoid control remains a fundamental challenge, as the ''curse of dimensio

safetyarxiv-cs-lg
5 May 2026
Safety

Fine-Grained Class-Conditional Distribution Balancing for Debiased Learning

DGX agent

arXiv:2505.06831v2 Announce Type: replace Abstract: Achieving group-robust generalization in the presence of spurious correlations remains a significant challenge, particularly when bias annotations a

safetyarxiv-cs-cv
5 May 2026
Safety

FLoRA: Fusion-Latent for Optical Reconstruction and Flood Area Segmentation via Cross-Modal Multi-Task Distillation Network

DGX agent

arXiv:2605.02137v1 Announce Type: new Abstract: Accurate flood water mapping is critical for disaster management, yet current methods struggle to fully exploit the potential of spaceborne imagery. Opt

safetyarxiv-cs-cv
5 May 2026
Safety

Foundation Models to Unlock Real-World Evidence from Nationwide Medical Claims

DGX agent

arXiv:2605.02740v1 Announce Type: cross Abstract: Evidence derived from large-scale real-world data (RWD) is increasingly informing regulatory evaluation and healthcare decision-making. Administrative

safetyarxiv-cs-cl
5 May 2026
Safety

From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Release for Offline-to-Online Reinforcement Learning

DGX agent

arXiv:2511.03828v2 Announce Type: replace Abstract: Offline-to-online reinforcement learning (O2O RL) faces a central challenge between retaining offline conservatism and adapting to online feedback u

safetyarxiv-cs-lg
5 May 2026
Safety

General Frameworks for Conditional Two-Sample Testing

DGX agent

arXiv:2410.16636v2 Announce Type: replace-cross Abstract: We study the problem of conditional two-sample testing, which aims to determine whether two populations have the same distribution after accou

safetyarxiv-cs-lg
5 May 2026
Safety

Generalized Distributional Alignment Games for Unbiased Answer-Level Fine-Tuning

DGX agent

arXiv:2605.02435v1 Announce Type: new Abstract: The Distributional Alignment Game framework provides a powerful variational perspective on Answer-Level Fine-Tuning (ALFT). However, standard algorithms

safetyarxiv-cs-lg
5 May 2026
Safety

Geometric and Spectral Alignment for Deep Neural Network I

DGX agent

arXiv:2605.02108v1 Announce Type: new Abstract: Deep residual architectures are modeled as products of near-identity Jacobians. This paper proves deterministic quotient-geometric estimates for singula

safetyarxiv-cs-lg
5 May 2026
Safety

Geometric and Spectral Alignment for Deep Neural Network II

DGX agent

arXiv:2605.02111v1 Announce Type: new Abstract: This paper develops the angular and static-channel component of Geometric and Spectral Alignment for residual Jacobian chains. Starting from Cartan-coor

safetyarxiv-cs-lg
5 May 2026
Safety

GETA-3DGS: Automatic Joint Structured Pruning and Quantization for 3D Gaussian Splatting

DGX agent

arXiv:2605.02086v1 Announce Type: new Abstract: 3D Gaussian splatting (3DGS) is a state-of-the-art representation for real-time photorealistic novel-view synthesis, yet a single high-fidelity scene ty

safetyarxiv-cs-lg
5 May 2026
Safety

Good in Bad (GiB): Sifting Through End-user Demonstrations for Learning a Better Policy

DGX agent

arXiv:2605.01529v1 Announce Type: new Abstract: Imitation learning offers a promising framework for enabling robots to acquire diverse skills from human users. However, most imitation learning algorit

safetyarxiv-cs-ro
5 May 2026
Safety

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models

DGX agent

arXiv:2605.02626v1 Announce Type: new Abstract: Preference optimization has become a central paradigm for aligning large language models with human feedback. Direct Preference Optimization (DPO) simpl

safetyarxiv-cs-lg
5 May 2026
Safety

Green Energy Management for Sustainable Data Centers Using Deep Reinforcement Learning

DGX agent

arXiv:2507.21153v2 Announce Type: replace Abstract: The exponential growth of digital services has positioned data centers among the most energy-intensive infrastructures in the modern economy, raisin

safetyarxiv-cs-lg
5 May 2026
Safety

HandelBot: Real-World Piano Playing via Fast Adaptation of Dexterous Robot Policies

DGX agent

arXiv:2603.12243v3 Announce Type: replace Abstract: Mastering dexterous manipulation with multi-fingered hands has been a grand challenge in robotics for decades. Despite its potential, the difficulty

safetyarxiv-cs-ro
5 May 2026
Safety

HeteroRAG: A Heterogeneous Retrieval-Augmented Generation Framework for Medical Vision Language Tasks

DGX agent

arXiv:2508.12778v2 Announce Type: replace Abstract: Medical large vision-language Models (Med-LVLMs) have shown promise in clinical applications but suffer from factual inaccuracies and unreliable out

safetyarxiv-cs-cl
5 May 2026
Safety

High entropy leads to symmetry equivariant policies in Dec-POMDPs

DGX agent

arXiv:2511.22581v3 Announce Type: replace Abstract: We prove that in any Dec-POMDP, sufficiently high entropy regularization ensures that the policy gradient flow with tabular softmax parametrization

safetyarxiv-cs-lg
5 May 2026
Safety

How Can One Choose the Best CAM-Based Explainability Method for a CNN Model?

DGX agent

arXiv:2605.02007v1 Announce Type: cross Abstract: In recent years, several advances have been observed in Deep Learning with surprising results. Models in this area have been increasingly used in nume

safetyarxiv-cs-cv
5 May 2026
Safety

HumanSplatHMR: Closing the Loop Between Human Mesh Recovery and Gaussian Splatting Avatar

DGX agent

arXiv:2605.02784v1 Announce Type: new Abstract: Accurately recovering human pose and appearance from video is an essential component of scene reconstruction, with applications to motion capture, motio

safetyarxiv-cs-cv
5 May 2026
Safety

Hybrid Quantum Reinforcement Learning with QAOA for Improved Vehicle Routing Optimization

DGX agent

arXiv:2605.01574v1 Announce Type: new Abstract: Vehicle Routing Problem (VRP) is one of the most complex NP-hard combinatorial optimization problem in transportation and logistics that requires a dyna

safetyarxiv-cs-lg
5 May 2026
Safety

Hydra-DP3: Frequency-Aware Right-Sizing of 3D Diffusion Policies for Visuomotor Control

DGX agent

arXiv:2605.01581v1 Announce Type: new Abstract: Diffusion-based visuomotor policies perform well in robotic manipulation, yet current methods still inherit image-generation-style decoders and multi-st

safetyarxiv-cs-ro
5 May 2026
Safety

Implicature in Interaction: Understanding Implicature Improves Alignment in Human-LLM Interaction

DGX agent

arXiv:2510.25426v2 Announce Type: replace Abstract: The rapid advancement of Large Language Models (LLMs) is positioning language at the core of human-computer interaction (HCI). We argue that advanci

safetyarxiv-cs-cl
5 May 2026
Safety

Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression

DGX agent

arXiv:2605.01402v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) struggle with numerical regression under long-tailed target distributions. Token-level supervised fine-tuning (

safetyarxiv-cs-cl
5 May 2026
Safety

Investigating Anthropometric Fidelity in SAM 3D Body

DGX agent

arXiv:2601.06035v2 Announce Type: replace-cross Abstract: The release of SAM 3D Body is a recent development in human mesh recovery, demonstrating improved performance in producing clean, topologicall

safetyarxiv-cs-cv
5 May 2026
← Previous
1…200201202203204…257
Next →