AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,435 results
Safety

Few-Shot Neural Differentiable Simulator: Real-to-Sim Rigid-Contact Modeling

DGX agent

arXiv:2603.06218v2 Announce Type: replace Abstract: Accurate physics simulation is essential for robotic learning and control, yet analytical simulators often fail to capture complex contact dynamics,

safetyarxiv-cs-ro
26 May 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Flat Minima and Generalization: Insights from Stochastic Convex Optimization

DGX agent

arXiv:2511.03548v2 Announce Type: replace Abstract: Understanding the generalization behavior of learning algorithms is a central goal of learning theory. A recently emerging explanation is that learn

safetyarxiv-cs-lg
26 May 2026
Safety

From Reasoning to Code: GRPO Optimization for Underrepresented Languages

DGX agent

arXiv:2506.11027v3 Announce Type: replace-cross Abstract: Generating accurate and executable code using Large Language Models (LLMs) remains a significant challenge for underrepresented programming la

safetyarxiv-cs-ai
26 May 2026
Safety

From Simulation to Enaction: Post-trained language models recognize and react to their own generations

DGX agent

arXiv:2605.25459v1 Announce Type: cross Abstract: Language models are pretrained as passive predictors with no incentive to model the consequences of their own outputs. Post-training changes this: a m

safetyarxiv-cs-ai
26 May 2026
Safety

FusionCore: A 23-State Unscented Kalman Filter for IMU, Wheel Encoder, GPS, and Visual SLAM Fusion in ROS 2

DGX agent

arXiv:2605.25239v1 Announce Type: new Abstract: We present FusionCore, an open-source ROS 2 sensor fusion package that fuses IMU, wheel encoder odometry, GPS, and Visual SLAM pose into a single 100 Hz

safetyarxiv-cs-ro
26 May 2026
Safety

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning

DGX agent

arXiv:2603.10250v2 Announce Type: replace Abstract: A commonly used family of RL algorithms for diffusion policies conducts softmax reweighting over samples from the behavior policy, which often induc

safetyarxiv-cs-lg
26 May 2026
Safety

Generative OOD-regularized Model-based Policy Optimization

DGX agent

arXiv:2605.24405v1 Announce Type: cross Abstract: We study sequential decision-making with offline reinforcement learning (RL). Traditional offline RL policies may result in out-of-distribution (OOD)

safetyarxiv-cs-ai
26 May 2026
Safety

Generative Visual Code Mobile World Models

DGX agent

arXiv:2602.01576v2 Announce Type: replace-cross Abstract: Mobile Graphical User Interface (GUI) World Models (WMs) offer a promising path for improving mobile GUI agent performance at train- and infer

safetyarxiv-cs-ai
26 May 2026
Safety

GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation

DGX agent

arXiv:2605.25447v1 Announce Type: new Abstract: Generating structured, editable diagrams remains a significant challenge for contemporary large language models, despite their proficiency in general-pu

safetyarxiv-cs-cl
26 May 2026
Safety

GIBLy: Improving 3D Semantic Segmentation through an Architecture-Agnostic Lightweight Geometric Inductive Bias Layer

DGX agent

arXiv:2605.24243v1 Announce Type: cross Abstract: In 3D scene understanding, deep learning models rely on large models and extensive training to capture basic geometric structures that are present in

safetyarxiv-cs-ai
26 May 2026
Safety

Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning

DGX agent

arXiv:2605.26078v1 Announce Type: new Abstract: Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action

safetyarxiv-cs-lg
26 May 2026
Safety

Global linear convergence of entropy-regularized softmax policy gradient beyond tabular MDPs

DGX agent

arXiv:2605.24939v1 Announce Type: new Abstract: We study the global convergence of policy gradient for infinite-horizon entropy-regularized Markov decision processes (MDPs) with continuous state and a

safetyarxiv-cs-lg
26 May 2026
Model Releases

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

DGX agent

arXiv:2605.24636v1 Announce Type: new Abstract: While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios re

model-releasesarxiv-cs-ai
26 May 2026
Safety

Grouter: Decoupling Routing from Representation for Accelerated MoE Training

DGX agent

arXiv:2603.06626v2 Announce Type: replace-cross Abstract: Traditional Mixture-of-Experts (MoE) training typically proceeds without any structural priors, effectively requiring the model to simultaneou

safetyarxiv-cs-ai
26 May 2026
Safety

Grow-Prune-Freeze Networks: Adaptive & Continual Learning Technique for Olfactory Navigation

DGX agent

arXiv:2605.25170v1 Announce Type: cross Abstract: Training data for olfaction is scattered through disparate, non-standardized datasets that limit the ability to build representative world models. Olf

safetyarxiv-cs-ai
26 May 2026
Safety

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models

DGX agent

arXiv:2605.25443v1 Announce Type: new Abstract: Post-training has significantly enhanced the reasoning capability of Large Reasoning Models (LRMs), especially with Reinforcement Learning (RL) like Gro

safetyarxiv-cs-cl
26 May 2026
Safety

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio

DGX agent

arXiv:2605.25967v1 Announce Type: new Abstract: As policy catches up with the capabilities of generative AI, watermarking is central to content provenance efforts. Inference-time watermarks for autore

safetyarxiv-cs-lg
26 May 2026
Safety

Hide-and-Shill: A Reinforcement Learning Framework for Market Manipulation Detection in Symphony-a Decentralized Multi-Agent System

DGX agent

arXiv:2507.09179v3 Announce Type: replace Abstract: Decentralized finance (DeFi) has introduced a new era of permissionless financial innovation but also led to unprecedented market manipulation. With

safetyarxiv-cs-ai
26 May 2026
Safety

Hide to Guide: Learning via Semantic Masking

DGX agent

arXiv:2605.25198v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a powerful paradigm for improving language models on reasoning-intensive tasks, but i

safetyarxiv-cs-ai
26 May 2026
Safety

How Much Structure Do LLMs Need? Evaluating LLMs for Bibliometric Cluster Description

DGX agent

arXiv:2605.24351v1 Announce Type: new Abstract: Large language models (LLMs) can support scientific literature synthesis, but remain prone to hallucinated references, uneven coverage, and weakly groun

safetyarxiv-cs-cl
26 May 2026
Safety

How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis

DGX agent

arXiv:2605.24749v1 Announce Type: cross Abstract: Reward modeling is not only a prediction problem: in KL-regularized policy optimization, the learned reward is exponentiated to define the deployed po

safetyarxiv-cs-lg
26 May 2026
Safety

How to Mitigate the Distribution Shift Problem in Robotics Control: A Robust and Adaptive Approach Based on Offline to Online Imitation Learning

DGX agent

arXiv:2605.25414v1 Announce Type: new Abstract: Distribution shift in imitation learning refers to the problem that the agent cannot plan proper actions for a state that has not been visited during th

safetyarxiv-cs-ro
26 May 2026
Safety

HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos

DGX agent

arXiv:2605.24934v1 Announce Type: cross Abstract: Human egocentric video captures rich manipulation demonstrations without any robot hardware, yet transferring these skills to robots remains challengi

safetyarxiv-cs-ai
26 May 2026
Safety

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks

DGX agent

arXiv:2605.24217v1 Announce Type: new Abstract: As Large Language Models (LLMs) transition from research environments to production deployments, evaluating their performance against strict Service Lev

safetyarxiv-cs-ai
26 May 2026
Safety

Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression

DGX agent

arXiv:2603.05691v2 Announce Type: replace Abstract: It is increasingly common in machine learning to use learned models to label data and then employ such data to train more capable models. The phenom

safetyarxiv-cs-lg
26 May 2026
Safety

Improving Ensemble CAPE Forecasts with a Diffusion Model Incorporating Aerosol Information

DGX agent

arXiv:2605.24009v1 Announce Type: cross Abstract: Convective available potential energy (CAPE) is an important variable for forecasting severe weather and understanding deep convection and precipitati

safetyarxiv-cs-lg
26 May 2026
Safety

Improving the Completeness and Comparability of Segment Disclosures: A Large Language Model Approach

DGX agent

arXiv:2605.23924v1 Announce Type: new Abstract: Segment-level disclosures are a central component of financial reporting, providing insight into firms' internal organization and the allocation of econ

safetyarxiv-cs-cl
26 May 2026
Safety

Inference-Time Alignment of Diffusion Models via Trust-Region Iterative Twisted Sequential Monte Carlo

DGX agent

arXiv:2605.25123v1 Announce Type: cross Abstract: We study inference-time alignment for diffusion-based generative models, aiming to steer a base model toward high-reward outputs without updating its

safetyarxiv-cs-ai
26 May 2026
Safety

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning

DGX agent

arXiv:2605.05226v2 Announce Type: replace-cross Abstract: The central challenge of reinforcement learning for reasoning lies not only in the sparsity of outcome-level supervision, but more fundamental

safetyarxiv-cs-ai
26 May 2026
Safety

Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization

DGX agent

arXiv:2605.24960v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) faithfulness, i.e., whether CoTs genuinely reflect large language models' (LLM) underlying behavior, is typically evaluated und

safetyarxiv-cs-ai
26 May 2026
Safety

Is Decentralized AI Governable? From Regulative Policy to Constitutive Protocol

DGX agent

arXiv:2605.24538v1 Announce Type: cross Abstract: Every major framework for governing artificial intelligence presupposes an identifiable entity -- a developer, deployer, or operator -- who can be hel

safetyarxiv-cs-ai
26 May 2026
Safety

IsaacIPC: Coupling High-Fidelity Simulation and Realistic Rendering for Contact-Rich Robotic Systems

DGX agent

arXiv:2605.24339v1 Announce Type: new Abstract: We present IsaacIPC, a robotic simulation framework that couples GPU accelerated incremental potential contact (IPC) with IsaacSim/Lab. IsaacIPC maps si

safetyarxiv-cs-ro
26 May 2026
Safety

Iterative Feature Space Optimization through Incremental Adaptive Evaluation

DGX agent

arXiv:2501.14889v2 Announce Type: replace Abstract: Iterative feature space optimization involves systematically evaluating and adjusting the feature space to improve downstream task performance. Howe

safetyarxiv-cs-lg
26 May 2026
Safety

Iterative Refinement Neural Operators are Learned Fixed-Point Solvers: A Principled Approach to Spectral Bias Mitigation

DGX agent

arXiv:2605.24041v1 Announce Type: cross Abstract: Neural operators serve as fast, data-driven surrogates for scientific modeling but typically rely on a monolithic, single-pass inference procedure tha

safetyarxiv-cs-ai
26 May 2026
Safety

IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning

DGX agent

arXiv:2605.23997v1 Announce Type: cross Abstract: Multimodal large language models via reinforcement learning (RL) have demonstrated remarkable capabilities in complex visual reasoning tasks, yet they

safetyarxiv-cs-ai
26 May 2026
Safety

Joint Optimization of Training and Inference in Federated Edge Learning via Constrained Multi-Objective Deep Reinforcement Learning

DGX agent

arXiv:2605.25916v1 Announce Type: new Abstract: Federated edge learning (FEEL) has recently emerged as a promising paradigm for achieving edge intelligence (EI) via enabling collaborative model traini

safetyarxiv-cs-lg
26 May 2026
Safety

KYA: A Framework-Agnostic Trust Layer for Autonomous Systems with Verifiable Provenance and Hierarchical Policy Composition

DGX agent

arXiv:2605.25376v1 Announce Type: cross Abstract: Observability tells operators when an agent is slow. KYA tells operators when an agent is wrong, drifting, leaking, or quietly going rogue. We present

safetyarxiv-cs-ai
26 May 2026
Safety

Label-NTK Alignments and A Tighter Convergence Bound in the NTK Regime

DGX agent

arXiv:2605.25275v1 Announce Type: new Abstract: The Neural Tangent Kernel (NTK) framework explains optimization in over-parameterized neural networks via approximately linearized dynamics, yielding ex

safetyarxiv-cs-lg
26 May 2026
Safety

Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation

DGX agent

arXiv:2605.25036v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) extend large language models with visual understanding, but remain vulnerable to hallucination, where outputs are

safetyarxiv-cs-ai
26 May 2026
Safety

LAPLEX: The FFT of Learnable Laplace Kernels

DGX agent

arXiv:2605.24584v1 Announce Type: cross Abstract: Fast linear algebra in deep learning usually comes with a choice: fixed geometry and exact computation, as in the Fourier transform, or adaptive geome

safetyarxiv-cs-ai
26 May 2026
Safety

Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning

DGX agent

arXiv:2605.25740v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (GCRL) provides a practical framework for obtaining goal-reaching policies from fixed datasets. However,

safetyarxiv-cs-lg
26 May 2026
Safety

LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition

DGX agent

arXiv:2605.24005v1 Announce Type: new Abstract: The evolution of Large Language Model (LLM) reasoning is bottlenecked by the scarcity of high-quality process data. While self-alignment via endogenous

safetyarxiv-cs-ai
26 May 2026
Safety

Learning Concepts, Not Tokens: Self-Supervised Semantic Alignment for Language Models

DGX agent

arXiv:2603.29123v2 Announce Type: replace Abstract: The next-token prediction (NTP) objective trains language models to predict a single token at each step, even though many continuations can express

safetyarxiv-cs-cl
26 May 2026
Safety

Learning High-Frequency Continuous Action Chunks in Latent Space

DGX agent

arXiv:2605.24931v1 Announce Type: new Abstract: Modern robotic policies increasingly rely on action chunking to execute complex tasks in the physical world. While action chunking improves temporal con

safetyarxiv-cs-ro
26 May 2026
Safety

Learning in Low-Dimensional Subspaces: Orthogonal Bottlenecks for Reinforcement Learning

DGX agent

arXiv:2605.26012v1 Announce Type: cross Abstract: Deep reinforcement learning (RL) agents commonly rely on high-dimensional neural representations, despite growing evidence that task-relevant value an

safetyarxiv-cs-ai
26 May 2026
Safety

Learning to Route Languages for Multilingual Policy Optimization

DGX agent

arXiv:2605.25360v1 Announce Type: new Abstract: Large language models~(LLMs) are trained on heterogeneous multilingual corpora, yet existing policy optimization methods often implicitly restrict each

safetyarxiv-cs-cl
26 May 2026
Safety

Locality Matters for Training-Free Audio Token Compression in Audio-Language Models

DGX agent

arXiv:2605.25179v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used for audio captioning, question answering, and open-ended audio understanding, but their inference cos

safetyarxiv-cs-cl
26 May 2026
Safety

Machine Psychometrics: A Mathematical Psychology of Artificial Intelligence

DGX agent

arXiv:2605.23952v1 Announce Type: new Abstract: Artificial agents now generate behavior rich enough to invite trust, surprise, and concern, yet our evaluation tools still privilege capability scores o

safetyarxiv-cs-ai
26 May 2026
← Previous
1…160161162163164…260
Next →