AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs

DGX agent

arXiv:2408.11513v2 Announce Type: replace Abstract: This paper focuses on learning a Constrained Markov Decision Process (CMDP) via general parameterized policies. We propose a Primal-Dual based Regul

safetyarxiv-cs-lg
4 May 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding

DGX agent

arXiv:2605.00642v1 Announce Type: cross Abstract: Graphical User Interface (GUI) grounding maps natural language instructions to the visual coordinates of target elements and serves as a core capabili

safetyarxiv-cs-cv
4 May 2026
Safety

Learning Coarse-to-Fine Osteoarthritis Representations under Noisy Hierarchical Labels

DGX agent

arXiv:2605.00718v1 Announce Type: new Abstract: Knee osteoarthritis (OA) assessment involves a natural but often underused label hierarchy: a coarse binary OA decision and a fine-grained Kellgren--Law

safetyarxiv-cs-cv
4 May 2026
Safety

Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory

DGX agent

arXiv:2605.00702v1 Announce Type: new Abstract: Large language model (LLM) agents require long-term user memory for consistent personalization, but limited context windows hinder tracking evolving pre

safetyarxiv-cs-cl
4 May 2026
Safety

Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

DGX agent

arXiv:2605.00416v1 Announce Type: new Abstract: Generalist robot policies increasingly benefit from large-scale pretraining, but offline data alone is insufficient for robust real-world deployment. De

safetyarxiv-cs-ro
4 May 2026
Safety

MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents

DGX agent

arXiv:2605.00356v1 Announce Type: new Abstract: Long-term conversational agents must decide which turns to store in external memory, yet recent systems rely on autoregressive LLM generation at every t

safetyarxiv-cs-cl
4 May 2026
Safety

Meritocratic Fairness in Budgeted Combinatorial Multi-armed Bandits via Shapley Values

DGX agent

arXiv:2605.00762v1 Announce Type: new Abstract: We propose a new framework for meritocratic fairness in budgeted combinatorial multi-armed bandits with full-bandit feedback (BCMAB-FBF). Unlike semi-ba

safetyarxiv-cs-lg
4 May 2026
Safety

Mesh Field Theory: Port-Hamiltonian Formulation of Mesh-Based Physics

DGX agent

arXiv:2605.00394v1 Announce Type: new Abstract: We present Mesh Field Theory (MeshFT) and its neural realization, MeshFT-Net: a structure-preserving framework for mesh-based continuum physics that cle

safetyarxiv-cs-lg
4 May 2026
Safety

Model-Based Reinforcement Learning with Double Oracle Efficiency in Policy Optimization and Offline Estimation

DGX agent

arXiv:2605.00393v1 Announce Type: new Abstract: Reinforcement learning (RL) in large environments often suffers from severe computational bottlenecks, as conventional regret minimization algorithms re

safetyarxiv-cs-lg
4 May 2026
Safety

Online Self-Calibration Against Hallucination in Vision-Language Models

DGX agent

arXiv:2605.00323v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) often suffer from hallucinations, generating descriptions that include visual details absent from the input image.

safetyarxiv-cs-cv
4 May 2026
Safety

Optimal Spatio-Temporal Decoupling for Bayesian Conformal Prediction

DGX agent

arXiv:2605.00432v1 Announce Type: new Abstract: Online Conformal Prediction (CP) struggles to balance temporal adaptability and structural stability. Feedback-driven methods (e.g., Adaptive Conformal

safetyarxiv-cs-lg
4 May 2026
Safety

Optimizing Resource-Constrained Non-Pharmaceutical Interventions for Multi-Cluster Outbreak Control Using Hierarchical Reinforcement Learning

DGX agent

arXiv:2603.19397v2 Announce Type: replace Abstract: Non-pharmaceutical interventions (NPIs), such as diagnostic testing and quarantine, are crucial for controlling infectious disease outbreaks but are

safetyarxiv-cs-lg
4 May 2026
Safety

PORTool: Importance-Aware Policy Optimization with Rewarded Tree for Multi-Tool-Integrated Reasoning

DGX agent

arXiv:2510.26020v2 Announce Type: replace Abstract: Multi-tool-integrated reasoning enables LLM-empowered tool-use agents to solve complex tasks by interleaving natural-language reasoning with calls t

safetyarxiv-cs-cl
4 May 2026
Safety

Pose-Aware Diffusion for 3D Generation

DGX agent

arXiv:2605.00345v1 Announce Type: new Abstract: Generating pose-aligned 3D objects is challenging due to the spatial mismatches and transformation ambiguities inherent in decoupled canonical-then-rota

safetyarxiv-cs-cv
4 May 2026
Safety

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

DGX agent

arXiv:2411.02327v4 Announce Type: replace Abstract: In the past year, video-based large language models (Video LLMs) have achieved impressive progress, particularly in their ability to process long vi

safetyarxiv-cs-cv
4 May 2026
Safety

PrefMoE: Robust Preference Modeling with Mixture-of-Experts Reward Learning

DGX agent

arXiv:2605.00384v1 Announce Type: new Abstract: Preference-based reinforcement learning offers a scalable alternative to manual reward engineering by learning reward structures from comparative feedba

safetyarxiv-cs-ro
4 May 2026
Safety

Provable and scalable quantum Gaussian processes for quantum learning

DGX agent

arXiv:2605.00099v1 Announce Type: cross Abstract: Despite rapid recent advances in quantum machine learning, the field is in many ways stuck. Existing approaches can exhibit serious limitations, and w

safetyarxiv-cs-lg
4 May 2026
Safety

Recovering Hidden Reward in Diffusion-Based Policies

DGX agent

arXiv:2605.00623v1 Announce Type: new Abstract: This paper introduces EnergyFlow, a framework that unifies generative action modeling with inverse reinforcement learning by parameterizing a scalar ene

safetyarxiv-cs-ro
4 May 2026
Safety

Reinforcement Learning for LLM Post-Training: A Survey

DGX agent

arXiv:2407.16216v3 Announce Type: replace Abstract: Large language models (LLMs) trained via pretraining and supervised fine-tuning (SFT) can still produce harmful and misaligned outputs, or struggle

safetyarxiv-cs-cl
4 May 2026
Safety

Reinforcement Learning with LLM-Guided Action Spaces for Synthesizable Lead Optimization

DGX agent

arXiv:2604.07669v2 Announce Type: replace Abstract: Lead optimization in drug discovery requires improving therapeutic properties while ensuring that molecular modifications correspond to feasible syn

safetyarxiv-cs-lg
4 May 2026
Safety

Reinforcement Learning with Markov Risk Measures and Multipattern Risk Approximation

DGX agent

arXiv:2605.00654v1 Announce Type: new Abstract: For a risk-averse finite-horizon Markov Decision Problem, we introduce a special class of Markov coherent risk measures, called mini-batch measures. We

safetyarxiv-cs-lg
4 May 2026
Safety

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning

DGX agent

arXiv:2605.00380v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) enhances reasoning of Large Language Models (LLMs) but usually exhibits limited generation diver

safetyarxiv-cs-cl
4 May 2026
Safety

Resting Neurons, Active Insights: Robustify Activation Sparsity for Large Language Models

DGX agent

arXiv:2512.12744v3 Announce Type: replace Abstract: Activation sparsity offers a compelling route to accelerate large language model (LLM) inference by selectively suppressing hidden activations, yet

safetyarxiv-cs-lg
4 May 2026
Safety

SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters

DGX agent

arXiv:2605.00528v1 Announce Type: cross Abstract: AI agents execute tens to hundreds of chained LLM calls per task, yet GPU schedulers treat each call as independent, discarding gigabytes of intermedi

safetyarxiv-cs-lg
4 May 2026
Safety

SAVGO: Learning State-Action Value Geometry with Cosine Similarity for Continuous Control

DGX agent

arXiv:2605.00787v1 Announce Type: new Abstract: While representation and similarity learning have improved the sample efficiency of Reinforcement Learning (RL), they are rarely used to shape policy up

safetyarxiv-cs-lg
4 May 2026
Safety

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration

DGX agent

arXiv:2605.00444v1 Announce Type: new Abstract: Multi-modal large language models (MLLMs) advance vision language understanding but face inherent limitations in long-video tasks due to bounded percept

safetyarxiv-cs-cv
4 May 2026
Safety

SIMON: Saliency-aware Integrative Multi-view Object-centric Neural Decoding

DGX agent

arXiv:2605.00401v1 Announce Type: new Abstract: Recent EEG-to-image retrieval methods leverage pretrained vision encoders and foveation-inspired priors, but typically assume a fixed, center-focused vi

safetyarxiv-cs-cv
4 May 2026
Safety

Soft Graph Diffusion Transformer for MIMO Detection

DGX agent

arXiv:2605.00449v1 Announce Type: cross Abstract: Learning-based MIMO detection has shown strong empirical performance, yet existing methods typically rely on fixed-depth architectures without explici

safetyarxiv-cs-lg
4 May 2026
Safety

Soft-MSM: Differentiable Context-Aware Elastic Alignment for Time Series

DGX agent

arXiv:2605.00069v1 Announce Type: new Abstract: Elastic distances like dynamic time warping (DTW) are central to time series machine learning because they compare sequences under local temporal misali

safetyarxiv-cs-lg
4 May 2026
Safety

Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium

DGX agent

arXiv:2503.10990v2 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with diverse human preferences is critical for ensuring fairness and informed outcomes when deploying th

safetyarxiv-cs-lg
4 May 2026
Safety

The Determinism of Randomness: Latent Space Degeneracy in Diffusion Model

DGX agent

arXiv:2511.07756v4 Announce Type: replace Abstract: Diffusion models initialize generation from an isotropic Gaussian latent, yet changing only the random seed can substantially alter prompt faithfuln

safetyarxiv-cs-cv
4 May 2026
Safety

Towards A Generative Protein Evolution Machine with DPLM-Evo

DGX agent

arXiv:2605.00182v1 Announce Type: new Abstract: Proteins are shaped by gradual evolution under biophysical and functional constraints. Protein language models learn rich evolutionary constraints from

safetyarxiv-cs-lg
4 May 2026
Safety

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity

DGX agent

arXiv:2605.00365v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has achieved substantial gains in single-attempt accuracy (Pass@1) on reasoning tasks, yet often

safetyarxiv-cs-cl
4 May 2026
Safety

UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors

DGX agent

arXiv:2605.00658v1 Announce Type: new Abstract: Recent progress has shown that video diffusion models (VDMs) can be repurposed for diverse multimodal graphics tasks. However, existing methods often tr

safetyarxiv-cs-cv
4 May 2026
Safety

Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards

DGX agent

arXiv:2510.00072v2 Announce Type: replace Abstract: Training robust reasoning vision-language models (VLMs) in rare domains (such as geospatial) is fundamentally constrained by supervision scarcity. W

safetyarxiv-cs-cv
4 May 2026
Safety

Unpaired Image Deraining Using Reward-Guided Self-Reinforcement Strategy

DGX agent

arXiv:2605.00719v1 Announce Type: new Abstract: Unsupervised deraining has attracted attention for its ability to learn the real-world distribution of rain without paired supervision. However, the lac

safetyarxiv-cs-cv
4 May 2026
Safety

VGR: Visual Grounded Reasoning

DGX agent

arXiv:2506.11991v3 Announce Type: replace-cross Abstract: In the field of multimodal chain-of-thought (CoT) reasoning, existing approaches predominantly rely on reasoning on pure language space, which

safetyarxiv-cs-cl
4 May 2026
Safety

VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation

DGX agent

arXiv:2509.21723v4 Announce Type: replace Abstract: Achieving generalizable bimanual manipulation requires systems that can learn efficiently from minimal human input while adapting to real-world unce

safetyarxiv-cs-ro
4 May 2026
Safety

Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback

DGX agent

arXiv:2605.00155v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has become a core post-training step for aligning large language models, yet the reward signal used

safetyarxiv-cs-cl
4 May 2026
Safety

What Physics do Data-Driven MoCap-to-Radar Models Learn?

DGX agent

arXiv:2605.00018v1 Announce Type: new Abstract: Data-driven MoCap-to-radar models generate plausible micro-Doppler spectrograms, but do they actually learn the underlying physics? We introduce a physi

safetyarxiv-cs-lg
4 May 2026
Safety

World Model for Robot Learning: A Comprehensive Survey

DGX agent

arXiv:2605.00080v1 Announce Type: cross Abstract: World models, which are predictive representations of how environments evolve under actions, have become a central component of robot learning. They s

safetyarxiv-cs-cv
4 May 2026
Safety

A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations

DGX agent

arXiv:2604.27162v1 Announce Type: cross Abstract: Reinforcement Learning (RL) algorithms exhibit high sample complexity, particularly when applied to Decentralized Partially Observable Markov Decision

safetyarxiv-cs-lg
1 May 2026
Safety

A Woman with a Knife or A Knife with a Woman? Measuring Directional Bias Amplification in Image Captions

DGX agent

arXiv:2503.07878v5 Announce Type: replace-cross Abstract: When we train models on biased datasets, they not only reproduce data biases, but can worsen them at test time - a phenomenon called bias ampl

safetyarxiv-cs-ai
1 May 2026
Safety

Accelerating Policy Synthesis in Large-Scale MDPs via Hierarchical Adaptive Refinement

DGX agent

arXiv:2506.17792v2 Announce Type: replace Abstract: Software-intensive systems, such as software product lines and robotics, utilise Markov decision processes (MDPs) to capture uncertainty and analyse

safetyarxiv-cs-ai
1 May 2026
Safety

Addressing the Reality Gap: A Three-Tension Framework for Agentic AI Adoption

DGX agent

arXiv:2604.27245v1 Announce Type: cross Abstract: Generative AI has rapidly entered education through free consumer tools, outpacing the ability of schools and universities to respond. Now a new wave

safetyarxiv-cs-ai
1 May 2026
Safety

AEGIS: Authentic Edge Growth In Sparsity for Link Prediction in Edge-Sparse Bipartite Knowledge Graphs

DGX agent

arXiv:2509.22017v4 Announce Type: replace Abstract: Bipartite knowledge graphs in niche domains are typically data-poor and edge-sparse, which hinders link prediction. We introduce AEGIS (Authentic Ed

safetyarxiv-cs-lg
1 May 2026
Safety

Agent-Agnostic Evaluation of SQL Accuracy in Production Text-to-SQL Systems

DGX agent

arXiv:2604.28049v1 Announce Type: new Abstract: Text-to-SQL (T2SQL) evaluation in production environments poses fundamental challenges that existing benchmarks do not address. Current evaluation metho

safetyarxiv-cs-ai
1 May 2026
Safety

Agent Name Service (ANS): A Proof-of-Concept Trust Layer for Secure AI Agent Discovery, Identity, and Governance in Kubernetes

DGX agent

arXiv:2604.26997v1 Announce Type: cross Abstract: Autonomous AI agent ecosystems require stronger mechanisms for secure discovery, identity verification, capability attestation, and policy governance.

safetyarxiv-cs-ai
1 May 2026
← Previous
1…203204205206207…257
Next →