AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
16 Apr 2026

GRITS: A Spillage-Aware Guided Diffusion Policy for Robot Food Scooping Tasks

SafetyDGX agent

arXiv:2510.00573v2 Announce Type: replace Abstract: Robotic food scooping is a critical manipulation skill for food preparation and service robots. However, existing robot learning algorithms, especia

HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy

SafetyDGX agent

arXiv:2510.00695v3 Announce Type: replace-cross Abstract: Inherently, robotic manipulation tasks are history-dependent: leveraging past context could be beneficial. However, most existing Vision-Langu

Heavy-Tailed Class-Conditional Priors for Long-Tailed Generative Modeling

SafetyDGX agent

arXiv:2509.02154v2 Announce Type: replace-cross Abstract: Variational Autoencoders (VAEs) with global priors trained under an imbalanced empirical class distribution can lead to underrepresentation of

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

Model ReleasesDGX agent

arXiv:2604.13954v1 Announce Type: new Abstract: Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions.

' If a superintelligence is built, humanity will lose control over its future.' Before Canadian Senators, ControlAI's US Director Connor Lea…

SafetyDGX agent

' If a superintelligence is built, humanity will lose control over its future.' Before Canadian Senators, ControlAI's US Director Connor Leahy (@NPCollapse) explains why superintelligence could lead t

IGen: Scalable Data Generation for Robot Learning from Open-World Images

SafetyDGX agent

arXiv:2512.01773v2 Announce Type: replace Abstract: The rise of generalist robotic policies has created an exponential demand for large-scale training data. However, on-robot data collection is labor-

Irregularly Sampled Time Series Interpolation for Binary Evolution Simulations Using Dynamic Time Warping

SafetyDGX agent

arXiv:2604.13604v1 Announce Type: cross Abstract: Binary stellar evolution simulations are computationally expensive. Stellar population synthesis relies on these detailed evolution models at a fundam

Jump-Start Reinforcement Learning with Vision-Language-Action Regularization

SafetyDGX agent

arXiv:2604.13733v1 Announce Type: new Abstract: Reinforcement learning (RL) enables high-frequency, closed-loop control for robotic manipulation, but scaling to long-horizon tasks with sparse or imper

Learning Class Difficulty in Imbalanced Histopathology Segmentation via Dynamic Focal Attention

SafetyDGX agent

arXiv:2604.13479v1 Announce Type: cross Abstract: Semantic segmentation of histopathology images under class imbalance is typically addressed through frequency-based loss reweighting, which implicitly

Med-CAM: Minimal Evidence for Explaining Medical Decision Making

SafetyDGX agent

arXiv:2604.13695v1 Announce Type: new Abstract: Reliable and interpretable decision-making is essential in medical imaging, where diagnostic outcomes directly influence patient care. Despite advances

Neural Mean-Field Games: Extending Mean-Field Game Theory with Neural Stochastic Differential Equations

SafetyDGX agent

arXiv:2504.13228v4 Announce Type: replace Abstract: Mean-field game theory relies on approximating games that are intractable to model due to a very large to infinite population of players. While thes

Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning

SafetyDGX agent

arXiv:2506.08125v3 Announce Type: replace-cross Abstract: Large language models (LLMs) show strong reasoning abilities but often produce unnecessarily long explanations that reduce efficiency. Althoug

On the Fundamental Limitations of Dual Static CVaR Decompositions in Markov Decision Processes

SafetyDGX agent

arXiv:2507.14005v2 Announce Type: replace Abstract: It was recently shown that dynamic programming (DP) methods for finding static CVaR-optimal policies in Markov Decision Processes (MDPs) can fail wh

OpenAI's global policy chief, Chris Lehane, calls out AI 'doomers' and says 'when you put some of those thoughts and ideas out there, they do have consequences' (Caroline O'Donovan/The San Francisco ...)

SafetyDGX agent

Caroline O'Donovan / The San Francisco Standard: OpenAI's global policy chief, Chris Lehane, calls out AI “doomers” and says “when you put some of those thoughts and ideas out there, they do have cons

Pareto-Optimal Offline Reinforcement Learning via Smooth Tchebysheff Scalarization

SafetyDGX agent

arXiv:2604.13175v1 Announce Type: new Abstract: Large language models can be aligned with human preferences through offline reinforcement learning (RL) on small labeled datasets. While single-objectiv

Reducing Hallucinations in LLMs via Factuality-Aware Preference Learning

SafetyDGX agent

arXiv:2601.03027v3 Announce Type: replace Abstract: Preference alignment methods such as RLHF and Direct Preference Optimization (DPO) improve instruction following, but they can also reinforce halluc

Representation over Routing: Overcoming Surrogate Hacking in Multi-Timescale PPO

SafetyDGX agent

arXiv:2604.13517v1 Announce Type: new Abstract: Temporal credit assignment in reinforcement learning has long been a central challenge. Inspired by the multi-timescale encoding of the dopamine system

Restless Bandits with Individual Penalty Constraints: A New Near-Optimal Index Policy and How to Learn It

SafetyDGX agent

arXiv:2604.04101v2 Announce Type: replace Abstract: This paper investigates the Restless Multi-Armed Bandit (RMAB) framework under individual penalty constraints to address resource allocation challen

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

SafetyDGX agent

arXiv:2508.00222v5 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Reward (RLVR) has significantly advanced the complex reasoning abilities of Large Language Models (LLMs

RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph

SafetyDGX agent

arXiv:2511.07717v2 Announce Type: replace-cross Abstract: Estimating robot pose from a monocular RGB image is a challenge in robotics and computer vision. Existing methods typically build networks on

SFT-GRPO Data Overlap as a Post-Training Hyperparameter for Autoformalization

SafetyDGX agent

arXiv:2604.13515v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) followed by Group Relative Policy Optimization (GRPO) is a common post-training recipe. We conduct a controlled ablation ov

Soft Q(lambda): A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces

SafetyDGX agent

arXiv:2604.13780v1 Announce Type: new Abstract: Soft Q-learning has emerged as a versatile model-free method for entropy-regularised reinforcement learning, optimising for returns augmented with a pen

Spectral methods: crucial for machine learning, natural for quantum computers?

SafetyDGX agent

arXiv:2603.24654v2 Announce Type: replace-cross Abstract: This article presents an argument for why quantum computers could unlock new methods for machine learning. We argue that spectral methods, in

SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models

SafetyDGX agent

arXiv:2510.09541v3 Announce Type: replace Abstract: Diffusion large language models (dLLMs) are emerging as an efficient alternative to autoregressive models due to their ability to decode multiple to

Supreme Court Justice Clarence Thomas on Wednesday delivered a televised broadside against progressivism, a political philosophy he describe…

SafetyDGX agent

Supreme Court Justice Clarence Thomas on Wednesday delivered a televised broadside against progressivism, a political philosophy he described as an existential threat to America and the principles tha

Temporally Consistent Long-Term Memory for 3D Single Object Tracking

SafetyDGX agent

arXiv:2604.13789v1 Announce Type: new Abstract: 3D Single Object Tracking (3D-SOT) aims to localize a target object across a sequence of LiDAR point clouds, given its 3D bounding box in the first fram

The Long Delay to Arithmetic Generalization: When Learned Representations Outrun Behavior

SafetyDGX agent

arXiv:2604.13082v1 Announce Type: new Abstract: Grokking in transformers trained on algorithmic tasks is characterized by a long delay between training-set fit and abrupt generalization, but the sourc

Think Outside the Policy: In-Context Steered Policy Optimization

SafetyDGX agent

arXiv:2510.26519v3 Announce Type: replace Abstract: Existing Reinforcement Learning from Verifiable Rewards (RLVR) methods, such as Group Relative Policy Optimization (GRPO), have achieved remarkable

TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks

SafetyDGX agent

arXiv:2601.10245v2 Announce Type: replace-cross Abstract: Multi-step reasoning tasks like mathematical problem solving are vulnerable to cascading failures, where a single incorrect step leads to comp

UMI-3D: Extending Universal Manipulation Interface from Vision-Limited to 3D Spatial Perception

SafetyDGX agent

arXiv:2604.14089v1 Announce Type: new Abstract: We present UMI-3D, a multimodal extension of the Universal Manipulation Interface (UMI) for robust and scalable data collection in embodied manipulation

UNBOX: Unveiling Black-box visual models with Natural-language

SafetyDGX agent

arXiv:2603.08639v2 Announce Type: replace Abstract: Ensuring trustworthiness in open-world visual recognition requires models that are interpretable, fair, and robust to distribution shifts. Yet moder

What Are We Really Measuring? Rethinking Dataset Bias in Web-Scale Natural Image Collections via Unsupervised Semantic Clustering

SafetyDGX agent

arXiv:2604.13610v1 Announce Type: new Abstract: In computer vision, a prevailing method for quantifying dataset bias is to train a model to distinguish between datasets. High classification accuracy i

Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking

SafetyDGX agent

arXiv:2604.13776v1 Announce Type: cross Abstract: Watermarking is becoming the default mechanism for AI content authentication, with governance policies and frameworks referencing it as infrastructure

Why dynamically routing multi-timescale advantages in PPO causes policy collapse (and a simple decoupled fix) [R]

SafetyDGX agent

This Reddit post discusses a known instability in PPO when advantage estimates operating across different temporal scales (e.g., short-horizon and long-horizon returns) are dynamically routed or mixed

Why MLLMs Struggle to Determine Object Orientations

SafetyDGX agent

arXiv:2604.13321v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) struggle with tasks that require reasoning about 2D object orientation in images, as documented in prior work.

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks

SafetyDGX agent

arXiv:2604.13403v1 Announce Type: new Abstract: In-context learning (ICL) enables models to adapt to new tasks via inference-time demonstrations. Despite its success in large language models, the exte

15 Apr 2026

5/5 The takeaway: If your agent relies on an LLM judge for selection accuracy, measuring code quality isn’t enough; you need a measure of th…

SafetyDGX agent

5/5 The takeaway: If your agent relies on an LLM judge for selection accuracy, measuring code quality isn’t enough; you need a measure of the model's inductive bias toward the 'fingerprint' of a gold

A Dataset and Evaluation for Complex 4D Markerless Human Motion Capture

SafetyDGX agent

arXiv:2604.12765v1 Announce Type: new Abstract: Marker-based motion capture (MoCap) systems have long been the gold standard for accurate 4D human modeling, yet their reliance on specialized hardware

A document is worth a structured record: Principled inductive bias design for document recognition

SafetyDGX agent

arXiv:2507.08458v2 Announce Type: replace-cross Abstract: Many document types use intrinsic, convention-driven structures that serve to encode precise and structured information, such as the conventio

A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning

SafetyDGX agent

arXiv:2603.08291v3 Announce Type: replace Abstract: Multimodal Mathematical Reasoning (MMR) has recently attracted increasing attention for its capability to solve mathematical problems involving both

A TTP analysis finds dozens of nudify apps in Apple and Google app stores via search, despite company policy prohibiting them; the apps earned $122M+ in revenue (Bloomberg)

SafetyDGX agent

Bloomberg: A TTP analysis finds dozens of nudify apps in Apple and Google app stores via search, despite company policy prohibiting them; the apps earned $122M+ in revenue — Apple Inc. and Google have

A Two-Timescale Primal-Dual Framework for Reinforcement Learning via Online Dual Variable Guidance

SafetyDGX agent

arXiv:2505.04494v3 Announce Type: replace-cross Abstract: We study reinforcement learning by combining recent advances in regularized linear programming formulations with the classical theory of stoch

AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage Margin

SafetyDGX agent

arXiv:2505.14264v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as an effective approach for enhancing the reasoning capabilities of large language models (LLMs), esp

Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning

SafetyDGX agent

arXiv:2505.17086v4 Announce Type: replace Abstract: Large Language Models (LLMs) equipped with modern Retrieval-Augmented Generation (RAG) systems often employ multi-turn interaction pipelines to inte

ART-VITON: Measurement-Guided Latent Diffusion for Artifact-Free Virtual Try-On

SafetyDGX agent

arXiv:2509.25749v2 Announce Type: cross Abstract: Virtual try-on (VITON) aims to generate realistic images of a person wearing a target garment, requiring precise garment alignment in try-on regions a

BayMOTH: Bayesian optiMizatiOn with meTa-lookahead -- a simple approacH

SafetyDGX agent

arXiv:2604.12005v1 Announce Type: cross Abstract: Bayesian optimization (BO) has for sequential optimization of expensive black-box functions demonstrated practicality and effectiveness in many real-w

Beyond Factual Grounding: The Case for Opinion-Aware Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2604.12138v1 Announce Type: new Abstract: RAG systems have transformed how LLMs access external knowledge, but we find that current implementations exhibit a bias toward factual, objective conte

Beyond Transcription: Unified Audio Schema for Perception-Aware AudioLLMs

SafetyDGX agent

arXiv:2604.12506v1 Announce Type: new Abstract: Recent Audio Large Language Models (AudioLLMs) exhibit a striking performance inversion: while excelling at complex reasoning tasks, they consistently u

Black-Box Optimization From Small Offline Datasets via Meta Learning with Synthetic Tasks

SafetyDGX agent

arXiv:2604.12325v1 Announce Type: cross Abstract: We consider the problem of offline black-box optimization, where the goal is to discover optimal designs (e.g., molecules or materials) from past expe

BRAIN: Bias-Mitigation Continual Learning Approach to Vision-Brain Understanding

SafetyDGX agent

arXiv:2508.18187v2 Announce Type: replace-cross Abstract: Memory decay makes it harder for the human brain to recognize visual objects and retain details. Consequently, recorded brain signals become w

Brain-DiT: A Universal Multi-state fMRI Foundation Model with Metadata-Conditioned Pretraining

SafetyDGX agent

arXiv:2604.12683v1 Announce Type: new Abstract: Current fMRI foundation models primarily rely on a limited range of brain states and mismatched pretraining tasks, restricting their ability to learn ge

Bridging the Micro--Macro Gap: Frequency-Aware Semantic Alignment for Image Manipulation Localization

SafetyDGX agent

arXiv:2604.12341v1 Announce Type: new Abstract: As generative image editing advances, image manipulation localization (IML) must handle both traditional manipulations with conspicuous forensic artifac

Calibration-Aware Policy Optimization for Reasoning LLMs

SafetyDGX agent

arXiv:2604.12632v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) enhances LLM reasoning but often induces overconfidence, where incorrect responses yield lower perplexity th

CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning

SafetyDGX agent

arXiv:2602.00181v3 Announce Type: replace-cross Abstract: Understanding camera dynamics is a fundamental pillar of video spatial intelligence. However, existing multimodal models predominantly treat t

Causal Diffusion Models for Counterfactual Outcome Distributions in Longitudinal Data

SafetyDGX agent

arXiv:2604.12992v1 Announce Type: cross Abstract: Predicting counterfactual outcomes in longitudinal data, where sequential treatment decisions heavily depend on evolving patient states, is critical y

Challenging Vision-Language Models with Physically Deployable Multimodal Semantic Lighting Attacks

SafetyDGX agent

arXiv:2604.12833v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown remarkable performance, yet their security remains insufficiently understood. Existing adversarial studies focu

CIA: Inferring the Communication Topology from LLM-based Multi-Agent Systems

SafetyDGX agent

arXiv:2604.12461v1 Announce Type: new Abstract: LLM-based Multi-Agent Systems (MAS) have demonstrated remarkable capabilities in solving complex tasks. Central to MAS is the communication topology whi

CLEAR: Cross-Lingual Enhancement in Alignment via Reverse-training

SafetyDGX agent

arXiv:2604.05821v2 Announce Type: replace Abstract: Existing multilingual embedding models often encounter challenges in cross-lingual scenarios due to imbalanced linguistic resources and less conside

Combating Pattern and Content Bias: Adversarial Feature Learning for Generalized AI-Generated Image Detection

SafetyDGX agent

arXiv:2604.12353v1 Announce Type: new Abstract: In recent years, the rapid development of generative artificial intelligence technology has significantly lowered the barrier to creating high-quality f

Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring

SafetyDGX agent

arXiv:2604.12645v1 Announce Type: cross Abstract: Although autonomous underwater vehicles promise the capability of marine ecosystem monitoring, their deployment is fundamentally limited by the diffic

← Previous
1…208209210211212…240
Next →