AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,487 results
26 May 2026

FairJudge: Abstention-Aware Multimodal Judges for Fairness and Alignment Evaluation in Text-to-Image Models

SafetyDGX agent

arXiv:2510.22827v3 Announce Type: replace-cross Abstract: Evaluating text-to-image (T2I) systems requires judging not only whether an image matches a prompt, but also whether socially salient attribut

Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges

SafetyDGX agent

arXiv:2605.23970v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as automatic judges for summarization and dialogue evaluation. Prior work has documented biases such

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning

SafetyDGX agent

arXiv:2605.24286v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning is useful for monitoring language models only when the reasoning trace faithfully reflects the computation that produ

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Few-Shot Neural Differentiable Simulator: Real-to-Sim Rigid-Contact Modeling

SafetyDGX agent

arXiv:2603.06218v2 Announce Type: replace Abstract: Accurate physics simulation is essential for robotic learning and control, yet analytical simulators often fail to capture complex contact dynamics,

Flat Minima and Generalization: Insights from Stochastic Convex Optimization

SafetyDGX agent

arXiv:2511.03548v2 Announce Type: replace Abstract: Understanding the generalization behavior of learning algorithms is a central goal of learning theory. A recently emerging explanation is that learn

From Reasoning to Code: GRPO Optimization for Underrepresented Languages

SafetyDGX agent

arXiv:2506.11027v3 Announce Type: replace-cross Abstract: Generating accurate and executable code using Large Language Models (LLMs) remains a significant challenge for underrepresented programming la

From Simulation to Enaction: Post-trained language models recognize and react to their own generations

SafetyDGX agent

arXiv:2605.25459v1 Announce Type: cross Abstract: Language models are pretrained as passive predictors with no incentive to model the consequences of their own outputs. Post-training changes this: a m

FusionCore: A 23-State Unscented Kalman Filter for IMU, Wheel Encoder, GPS, and Visual SLAM Fusion in ROS 2

SafetyDGX agent

arXiv:2605.25239v1 Announce Type: new Abstract: We present FusionCore, an open-source ROS 2 sensor fusion package that fuses IMU, wheel encoder odometry, GPS, and Visual SLAM pose into a single 100 Hz

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning

SafetyDGX agent

arXiv:2603.10250v2 Announce Type: replace Abstract: A commonly used family of RL algorithms for diffusion policies conducts softmax reweighting over samples from the behavior policy, which often induc

Generative OOD-regularized Model-based Policy Optimization

SafetyDGX agent

arXiv:2605.24405v1 Announce Type: cross Abstract: We study sequential decision-making with offline reinforcement learning (RL). Traditional offline RL policies may result in out-of-distribution (OOD)

Generative Visual Code Mobile World Models

SafetyDGX agent

arXiv:2602.01576v2 Announce Type: replace-cross Abstract: Mobile Graphical User Interface (GUI) World Models (WMs) offer a promising path for improving mobile GUI agent performance at train- and infer

GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation

SafetyDGX agent

arXiv:2605.25447v1 Announce Type: new Abstract: Generating structured, editable diagrams remains a significant challenge for contemporary large language models, despite their proficiency in general-pu

GIBLy: Improving 3D Semantic Segmentation through an Architecture-Agnostic Lightweight Geometric Inductive Bias Layer

SafetyDGX agent

arXiv:2605.24243v1 Announce Type: cross Abstract: In 3D scene understanding, deep learning models rely on large models and extensive training to capture basic geometric structures that are present in

Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning

SafetyDGX agent

arXiv:2605.26078v1 Announce Type: new Abstract: Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action

Global linear convergence of entropy-regularized softmax policy gradient beyond tabular MDPs

SafetyDGX agent

arXiv:2605.24939v1 Announce Type: new Abstract: We study the global convergence of policy gradient for infinite-horizon entropy-regularized Markov decision processes (MDPs) with continuous state and a

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

Model ReleasesDGX agent

arXiv:2605.24636v1 Announce Type: new Abstract: While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios re

Grouter: Decoupling Routing from Representation for Accelerated MoE Training

SafetyDGX agent

arXiv:2603.06626v2 Announce Type: replace-cross Abstract: Traditional Mixture-of-Experts (MoE) training typically proceeds without any structural priors, effectively requiring the model to simultaneou

Grow-Prune-Freeze Networks: Adaptive & Continual Learning Technique for Olfactory Navigation

SafetyDGX agent

arXiv:2605.25170v1 Announce Type: cross Abstract: Training data for olfaction is scattered through disparate, non-standardized datasets that limit the ability to build representative world models. Olf

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models

SafetyDGX agent

arXiv:2605.25443v1 Announce Type: new Abstract: Post-training has significantly enhanced the reasoning capability of Large Reasoning Models (LRMs), especially with Reinforcement Learning (RL) like Gro

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio

SafetyDGX agent

arXiv:2605.25967v1 Announce Type: new Abstract: As policy catches up with the capabilities of generative AI, watermarking is central to content provenance efforts. Inference-time watermarks for autore

Hide-and-Shill: A Reinforcement Learning Framework for Market Manipulation Detection in Symphony-a Decentralized Multi-Agent System

SafetyDGX agent

arXiv:2507.09179v3 Announce Type: replace Abstract: Decentralized finance (DeFi) has introduced a new era of permissionless financial innovation but also led to unprecedented market manipulation. With

Hide to Guide: Learning via Semantic Masking

SafetyDGX agent

arXiv:2605.25198v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a powerful paradigm for improving language models on reasoning-intensive tasks, but i

How Much Structure Do LLMs Need? Evaluating LLMs for Bibliometric Cluster Description

SafetyDGX agent

arXiv:2605.24351v1 Announce Type: new Abstract: Large language models (LLMs) can support scientific literature synthesis, but remain prone to hallucinated references, uneven coverage, and weakly groun

How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis

SafetyDGX agent

arXiv:2605.24749v1 Announce Type: cross Abstract: Reward modeling is not only a prediction problem: in KL-regularized policy optimization, the learned reward is exponentiated to define the deployed po

How to Mitigate the Distribution Shift Problem in Robotics Control: A Robust and Adaptive Approach Based on Offline to Online Imitation Learning

SafetyDGX agent

arXiv:2605.25414v1 Announce Type: new Abstract: Distribution shift in imitation learning refers to the problem that the agent cannot plan proper actions for a state that has not been visited during th

HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos

SafetyDGX agent

arXiv:2605.24934v1 Announce Type: cross Abstract: Human egocentric video captures rich manipulation demonstrations without any robot hardware, yet transferring these skills to robots remains challengi

I feel like I'm eating crazy pills when I read the countless bad takes around how the Vatican would have virtually anointed Anthropic. When …

SafetyDGX agent

I feel like I'm eating crazy pills when I read the countless bad takes around how the Vatican would have virtually anointed Anthropic. When if you read the Pope's encyclical it's actually a COMPLETE r

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks

SafetyDGX agent

arXiv:2605.24217v1 Announce Type: new Abstract: As Large Language Models (LLMs) transition from research environments to production deployments, evaluating their performance against strict Service Lev

⚠️⚠️⚠️if you believe this you and think you are safe, you don’t understand how S&P is about to change the rules and what that means. run don…

SafetyDGX agent

⚠️⚠️⚠️if you believe this you and think you are safe, you don’t understand how S&P is about to change the rules and what that means. run don’t walk to read my essay “This one weird trick might cost yo

if you don’t follow or think you know better i urge you to read my essay “This one weird trick might cost your retirement fund billions”

SafetyDGX agent

Gary Marcus discusses a potentially overlooked financial risk that could have massive implications for retirement savings, presenting his analysis in an essay format on X. The post appears to emphasiz

Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression

SafetyDGX agent

arXiv:2603.05691v2 Announce Type: replace Abstract: It is increasingly common in machine learning to use learned models to label data and then employ such data to train more capable models. The phenom

Improving Ensemble CAPE Forecasts with a Diffusion Model Incorporating Aerosol Information

SafetyDGX agent

arXiv:2605.24009v1 Announce Type: cross Abstract: Convective available potential energy (CAPE) is an important variable for forecasting severe weather and understanding deep convection and precipitati

Improving the Completeness and Comparability of Segment Disclosures: A Large Language Model Approach

SafetyDGX agent

arXiv:2605.23924v1 Announce Type: new Abstract: Segment-level disclosures are a central component of financial reporting, providing insight into firms' internal organization and the allocation of econ

In our new paper, EDGE-OPD (EviDence GuidEd On-Policy Distillation), we introduce two key ideas for improving On-Policy Distillation (OPD). …

SafetyDGX agent

In our new paper, EDGE-OPD (EviDence GuidEd On-Policy Distillation), we introduce two key ideas for improving On-Policy Distillation (OPD). First, we use guided rollouts that inject privileged context

Inference-Time Alignment of Diffusion Models via Trust-Region Iterative Twisted Sequential Monte Carlo

SafetyDGX agent

arXiv:2605.25123v1 Announce Type: cross Abstract: We study inference-time alignment for diffusion-based generative models, aiming to steer a base model toward high-reward outputs without updating its

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning

SafetyDGX agent

arXiv:2605.05226v2 Announce Type: replace-cross Abstract: The central challenge of reinforcement learning for reasoning lies not only in the sparsity of outcome-level supervision, but more fundamental

Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization

SafetyDGX agent

arXiv:2605.24960v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) faithfulness, i.e., whether CoTs genuinely reflect large language models' (LLM) underlying behavior, is typically evaluated und

Is Decentralized AI Governable? From Regulative Policy to Constitutive Protocol

SafetyDGX agent

arXiv:2605.24538v1 Announce Type: cross Abstract: Every major framework for governing artificial intelligence presupposes an identifiable entity -- a developer, deployer, or operator -- who can be hel

IsaacIPC: Coupling High-Fidelity Simulation and Realistic Rendering for Contact-Rich Robotic Systems

SafetyDGX agent

arXiv:2605.24339v1 Announce Type: new Abstract: We present IsaacIPC, a robotic simulation framework that couples GPU accelerated incremental potential contact (IPC) with IsaacSim/Lab. IsaacIPC maps si

Iterative Feature Space Optimization through Incremental Adaptive Evaluation

SafetyDGX agent

arXiv:2501.14889v2 Announce Type: replace Abstract: Iterative feature space optimization involves systematically evaluating and adjusting the feature space to improve downstream task performance. Howe

Iterative Refinement Neural Operators are Learned Fixed-Point Solvers: A Principled Approach to Spectral Bias Mitigation

SafetyDGX agent

arXiv:2605.24041v1 Announce Type: cross Abstract: Neural operators serve as fast, data-driven surrogates for scientific modeling but typically rely on a monolithic, single-pass inference procedure tha

It's a shame the Pope didn't ask Chris Olah what happened to his plan to give up to 10% of Anthropic to the authors of the work they train o…

SafetyDGX agent

It's a shame the Pope didn't ask Chris Olah what happened to his plan to give up to 10% of Anthropic to the authors of the work they train on. (Spoiler: it never happened, and the authors on whose wor

IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning

SafetyDGX agent

arXiv:2605.23997v1 Announce Type: cross Abstract: Multimodal large language models via reinforcement learning (RL) have demonstrated remarkable capabilities in complex visual reasoning tasks, yet they

Joint Optimization of Training and Inference in Federated Edge Learning via Constrained Multi-Objective Deep Reinforcement Learning

SafetyDGX agent

arXiv:2605.25916v1 Announce Type: new Abstract: Federated edge learning (FEEL) has recently emerged as a promising paradigm for achieving edge intelligence (EI) via enabling collaborative model traini

KYA: A Framework-Agnostic Trust Layer for Autonomous Systems with Verifiable Provenance and Hierarchical Policy Composition

SafetyDGX agent

arXiv:2605.25376v1 Announce Type: cross Abstract: Observability tells operators when an agent is slow. KYA tells operators when an agent is wrong, drifting, leaking, or quietly going rogue. We present

Label-NTK Alignments and A Tighter Convergence Bound in the NTK Regime

SafetyDGX agent

arXiv:2605.25275v1 Announce Type: new Abstract: The Neural Tangent Kernel (NTK) framework explains optimization in over-parameterized neural networks via approximately linearized dynamics, yielding ex

Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective Mitigation

SafetyDGX agent

arXiv:2605.25036v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) extend large language models with visual understanding, but remain vulnerable to hallucination, where outputs are

LAPLEX: The FFT of Learnable Laplace Kernels

SafetyDGX agent

arXiv:2605.24584v1 Announce Type: cross Abstract: Fast linear algebra in deep learning usually comes with a choice: fixed geometry and exact computation, as in the Fourier transform, or adaptive geome

Latent Representation Alignment for Offline Goal-Conditioned Reinforcement Learning

SafetyDGX agent

arXiv:2605.25740v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (GCRL) provides a practical framework for obtaining goal-reaching policies from fixed datasets. However,

LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition

SafetyDGX agent

arXiv:2605.24005v1 Announce Type: new Abstract: The evolution of Large Language Model (LLM) reasoning is bottlenecked by the scarcity of high-quality process data. While self-alignment via endogenous

Learning Concepts, Not Tokens: Self-Supervised Semantic Alignment for Language Models

SafetyDGX agent

arXiv:2603.29123v2 Announce Type: replace Abstract: The next-token prediction (NTP) objective trains language models to predict a single token at each step, even though many continuations can express

Learning High-Frequency Continuous Action Chunks in Latent Space

SafetyDGX agent

arXiv:2605.24931v1 Announce Type: new Abstract: Modern robotic policies increasingly rely on action chunking to execute complex tasks in the physical world. While action chunking improves temporal con

Learning in Low-Dimensional Subspaces: Orthogonal Bottlenecks for Reinforcement Learning

SafetyDGX agent

arXiv:2605.26012v1 Announce Type: cross Abstract: Deep reinforcement learning (RL) agents commonly rely on high-dimensional neural representations, despite growing evidence that task-relevant value an

Learning to Route Languages for Multilingual Policy Optimization

SafetyDGX agent

arXiv:2605.25360v1 Announce Type: new Abstract: Large language models~(LLMs) are trained on heterogeneous multilingual corpora, yet existing policy optimization methods often implicitly restrict each

Locality Matters for Training-Free Audio Token Compression in Audio-Language Models

SafetyDGX agent

arXiv:2605.25179v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used for audio captioning, question answering, and open-ended audio understanding, but their inference cos

lol. OpenAI as the WeWork of AI. Literally called it, in those exact words, @CNBC w @carlquintanilla, 2024. Now even SoftBank is worried.

SafetyDGX agent

lol. OpenAI as the WeWork of AI. Literally called it, in those exact words, @CNBC w @carlquintanilla, 2024. Now even SoftBank is worried. SoftBank's own executives think Sam Altman is scamming their C

Machine Psychometrics: A Mathematical Psychology of Artificial Intelligence

SafetyDGX agent

arXiv:2605.23952v1 Announce Type: new Abstract: Artificial agents now generate behavior rich enough to invite trust, surprise, and concern, yet our evaluation tools still privilege capability scores o

MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models

SafetyDGX agent

arXiv:2605.26004v1 Announce Type: cross Abstract: Instruction tuning of large vision-language models (LVLMs) increasingly depends on massive multimodal corpora, yet these datasets contain samples with

MAPLE: Multi-State Aggregated Policy Evaluation for AlphaZero in Imperfect-Information Games

SafetyDGX agent

arXiv:2605.24139v1 Announce Type: new Abstract: Imperfect-information games (IIGs) are challenging, as players must make decisions without fully observing the true game state. While AlphaZero has achi

MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling

SafetyDGX agent

arXiv:2602.17658v2 Announce Type: replace-cross Abstract: Reward modeling is central to alignment pipelines such as RLHF, RLAIF, and PPO-based policy optimization, yet its reliability is constrained b

← Previous
1…146147148149150…242
Next →