AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,351 results
Safety

Estimating Continuous Treatment Effects with Two-Stage Kernel Ridge Regression

DGX agent

arXiv:2604.13410v1 Announce Type: cross Abstract: We study the problem of estimating the effect function for a continuous treatment, which maps each treatment value to a population-averaged outcome. A

safetyarxiv-cs-lg
16 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization

DGX agent

arXiv:2604.13533v1 Announce Type: cross Abstract: Achieving general-purpose robotics requires empowering robots to adapt and evolve based on their environment and feedback. Traditional methods face li

safetyarxiv-cs-cv
16 Apr 2026
Safety

Failure Identification in Imitation Learning Via Statistical and Semantic Filtering

DGX agent

arXiv:2604.13788v1 Announce Type: cross Abstract: Imitation learning (IL) policies in robotics deliver strong performance in controlled settings but remain brittle in real-world deployments: rare even

safetyarxiv-cs-cv
16 Apr 2026
Safety

FCBV-Net: Category-Level Robotic Garment Smoothing via Feature-Conditioned Bimanual Value Prediction

DGX agent

arXiv:2508.05153v2 Announce Type: replace Abstract: Category-level generalization for robotic garment manipulation, such as bimanual smoothing, remains a significant hurdle due to high dimensionality,

safetyarxiv-cs-ro
16 Apr 2026
Safety

First-See-Then-Design: A Multi-Stakeholder View for Optimal Performance-Fairness Trade-Offs

DGX agent

arXiv:2604.14035v1 Announce Type: new Abstract: Fairness in algorithmic decision-making is often defined in the predictive space, where predictive performance - used as a proxy for decision-maker (DM)

safetyarxiv-cs-lg
16 Apr 2026
Safety

Foresight Optimization for Strategic Reasoning in Large Language Models

DGX agent

arXiv:2604.13592v1 Announce Type: new Abstract: Reasoning capabilities in large language models (LLMs) have generally advanced significantly. However, it is still challenging for existing reasoning-ba

safetyarxiv-cs-cl
16 Apr 2026
Safety

From Alignment to Prediction: A Study of Self-Supervised Learning and Predictive Representation Learning

DGX agent

arXiv:2604.13518v1 Announce Type: new Abstract: Self-supervised learning has emerged as a major technique for the task of learning from unlabeled data, where the current methods mostly revolve around

safetyarxiv-cs-lg
16 Apr 2026
Safety

From Instruction to Event: Sound-Triggered Mobile Manipulation

DGX agent

arXiv:2601.21667v2 Announce Type: replace-cross Abstract: Current mobile manipulation research predominantly follows an instruction-driven paradigm, where agents rely on predefined textual commands to

safetyarxiv-cs-cv
16 Apr 2026
Safety

From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space

DGX agent

arXiv:2604.14142v1 Announce Type: cross Abstract: While reinforcement learning with verifiable rewards (RLVR) significantly enhances LLM reasoning by optimizing the conditional distribution P(y|x), it

safetyarxiv-cs-cl
16 Apr 2026
Safety

From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction

DGX agent

arXiv:2604.13067v1 Announce Type: cross Abstract: SpeechLLMs process spoken language directly from audio, but accent and vocal identity cues can lead to biased behaviour. Current bias evaluations ofte

safetyarxiv-cs-cl
16 Apr 2026
Safety

@GaryMarcus Essentially your prediction that LLMs to scale wouldn’t solve key fundamental problems, reasoning, factual reliability, composit…

DGX agent

@GaryMarcus Essentially your prediction that LLMs to scale wouldn’t solve key fundamental problems, reasoning, factual reliability, compositionality, grounded understanding etc are now elephantine chi

safetygary-marcus--x
16 Apr 2026
Safety

@GaryMarcus @sama We need to rediscover the power of Satyagraha (principled, ardent, nonviolent resistence)

DGX agent

Gary Marcus advocates for applying Satyagraha—a principle of nonviolent resistance emphasizing moral conviction—as a response to contemporary challenges, likely in the context of AI development and go

safetygary-marcus--x
16 Apr 2026
Safety

GRITS: A Spillage-Aware Guided Diffusion Policy for Robot Food Scooping Tasks

DGX agent

arXiv:2510.00573v2 Announce Type: replace Abstract: Robotic food scooping is a critical manipulation skill for food preparation and service robots. However, existing robot learning algorithms, especia

safetyarxiv-cs-ro
16 Apr 2026
Safety

HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy

DGX agent

arXiv:2510.00695v3 Announce Type: replace-cross Abstract: Inherently, robotic manipulation tasks are history-dependent: leveraging past context could be beneficial. However, most existing Vision-Langu

safetyarxiv-cs-cv
16 Apr 2026
Safety

Heavy-Tailed Class-Conditional Priors for Long-Tailed Generative Modeling

DGX agent

arXiv:2509.02154v2 Announce Type: replace-cross Abstract: Variational Autoencoders (VAEs) with global priors trained under an imbalanced empirical class distribution can lead to underrepresentation of

safetyarxiv-cs-cv
16 Apr 2026
Model Releases

HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

DGX agent

arXiv:2604.13954v1 Announce Type: new Abstract: Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions.

model-releasesarxiv-cs-lg
16 Apr 2026
Safety

' If a superintelligence is built, humanity will lose control over its future.' Before Canadian Senators, ControlAI's US Director Connor Lea…

DGX agent

' If a superintelligence is built, humanity will lose control over its future.' Before Canadian Senators, ControlAI's US Director Connor Leahy (@NPCollapse) explains why superintelligence could lead t

safetyconnor-leahy--x
16 Apr 2026
Safety

IGen: Scalable Data Generation for Robot Learning from Open-World Images

DGX agent

arXiv:2512.01773v2 Announce Type: replace Abstract: The rise of generalist robotic policies has created an exponential demand for large-scale training data. However, on-robot data collection is labor-

safetyarxiv-cs-ro
16 Apr 2026
Safety

Irregularly Sampled Time Series Interpolation for Binary Evolution Simulations Using Dynamic Time Warping

DGX agent

arXiv:2604.13604v1 Announce Type: cross Abstract: Binary stellar evolution simulations are computationally expensive. Stellar population synthesis relies on these detailed evolution models at a fundam

safetyarxiv-cs-lg
16 Apr 2026
Safety

Jump-Start Reinforcement Learning with Vision-Language-Action Regularization

DGX agent

arXiv:2604.13733v1 Announce Type: new Abstract: Reinforcement learning (RL) enables high-frequency, closed-loop control for robotic manipulation, but scaling to long-horizon tasks with sparse or imper

safetyarxiv-cs-lg
16 Apr 2026
Safety

Learning Class Difficulty in Imbalanced Histopathology Segmentation via Dynamic Focal Attention

DGX agent

arXiv:2604.13479v1 Announce Type: cross Abstract: Semantic segmentation of histopathology images under class imbalance is typically addressed through frequency-based loss reweighting, which implicitly

safetyarxiv-cs-cv
16 Apr 2026
Safety

Med-CAM: Minimal Evidence for Explaining Medical Decision Making

DGX agent

arXiv:2604.13695v1 Announce Type: new Abstract: Reliable and interpretable decision-making is essential in medical imaging, where diagnostic outcomes directly influence patient care. Despite advances

safetyarxiv-cs-cv
16 Apr 2026
Safety

Neural Mean-Field Games: Extending Mean-Field Game Theory with Neural Stochastic Differential Equations

DGX agent

arXiv:2504.13228v4 Announce Type: replace Abstract: Mean-field game theory relies on approximating games that are intractable to model due to a very large to infinite population of players. While thes

safetyarxiv-cs-lg
16 Apr 2026
Safety

Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning

DGX agent

arXiv:2506.08125v3 Announce Type: replace-cross Abstract: Large language models (LLMs) show strong reasoning abilities but often produce unnecessarily long explanations that reduce efficiency. Althoug

safetyarxiv-cs-cl
16 Apr 2026
Safety

On the Fundamental Limitations of Dual Static CVaR Decompositions in Markov Decision Processes

DGX agent

arXiv:2507.14005v2 Announce Type: replace Abstract: It was recently shown that dynamic programming (DP) methods for finding static CVaR-optimal policies in Markov Decision Processes (MDPs) can fail wh

safetyarxiv-cs-lg
16 Apr 2026
Safety

OpenAI's global policy chief, Chris Lehane, calls out AI 'doomers' and says 'when you put some of those thoughts and ideas out there, they do have consequences' (Caroline O'Donovan/The San Francisco ...)

DGX agent

Caroline O'Donovan / The San Francisco Standard: OpenAI's global policy chief, Chris Lehane, calls out AI “doomers” and says “when you put some of those thoughts and ideas out there, they do have cons

safetytechmeme
16 Apr 2026
Safety

Pareto-Optimal Offline Reinforcement Learning via Smooth Tchebysheff Scalarization

DGX agent

arXiv:2604.13175v1 Announce Type: new Abstract: Large language models can be aligned with human preferences through offline reinforcement learning (RL) on small labeled datasets. While single-objectiv

safetyarxiv-cs-lg
16 Apr 2026
Safety

Reducing Hallucinations in LLMs via Factuality-Aware Preference Learning

DGX agent

arXiv:2601.03027v3 Announce Type: replace Abstract: Preference alignment methods such as RLHF and Direct Preference Optimization (DPO) improve instruction following, but they can also reinforce halluc

safetyarxiv-cs-cl
16 Apr 2026
Safety

Representation over Routing: Overcoming Surrogate Hacking in Multi-Timescale PPO

DGX agent

arXiv:2604.13517v1 Announce Type: new Abstract: Temporal credit assignment in reinforcement learning has long been a central challenge. Inspired by the multi-timescale encoding of the dopamine system

safetyarxiv-cs-lg
16 Apr 2026
Safety

Restless Bandits with Individual Penalty Constraints: A New Near-Optimal Index Policy and How to Learn It

DGX agent

arXiv:2604.04101v2 Announce Type: replace Abstract: This paper investigates the Restless Multi-Armed Bandit (RMAB) framework under individual penalty constraints to address resource allocation challen

safetyarxiv-cs-lg
16 Apr 2026
Safety

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

DGX agent

arXiv:2508.00222v5 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Reward (RLVR) has significantly advanced the complex reasoning abilities of Large Language Models (LLMs

safetyarxiv-cs-cl
16 Apr 2026
Safety

RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph

DGX agent

arXiv:2511.07717v2 Announce Type: replace-cross Abstract: Estimating robot pose from a monocular RGB image is a challenge in robotics and computer vision. Existing methods typically build networks on

safetyarxiv-cs-cv
16 Apr 2026
Safety

SFT-GRPO Data Overlap as a Post-Training Hyperparameter for Autoformalization

DGX agent

arXiv:2604.13515v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) followed by Group Relative Policy Optimization (GRPO) is a common post-training recipe. We conduct a controlled ablation ov

safetyarxiv-cs-lg
16 Apr 2026
Safety

Soft Q(lambda): A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces

DGX agent

arXiv:2604.13780v1 Announce Type: new Abstract: Soft Q-learning has emerged as a versatile model-free method for entropy-regularised reinforcement learning, optimising for returns augmented with a pen

safetyarxiv-cs-lg
16 Apr 2026
Safety

Spectral methods: crucial for machine learning, natural for quantum computers?

DGX agent

arXiv:2603.24654v2 Announce Type: replace-cross Abstract: This article presents an argument for why quantum computers could unlock new methods for machine learning. We argue that spectral methods, in

safetyarxiv-cs-lg
16 Apr 2026
Safety

SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models

DGX agent

arXiv:2510.09541v3 Announce Type: replace Abstract: Diffusion large language models (dLLMs) are emerging as an efficient alternative to autoregressive models due to their ability to decode multiple to

safetyarxiv-cs-cl
16 Apr 2026
Safety

Supreme Court Justice Clarence Thomas on Wednesday delivered a televised broadside against progressivism, a political philosophy he describe…

DGX agent

Supreme Court Justice Clarence Thomas on Wednesday delivered a televised broadside against progressivism, a political philosophy he described as an existential threat to America and the principles tha

safetyelon-musk--x
16 Apr 2026
Safety

Temporally Consistent Long-Term Memory for 3D Single Object Tracking

DGX agent

arXiv:2604.13789v1 Announce Type: new Abstract: 3D Single Object Tracking (3D-SOT) aims to localize a target object across a sequence of LiDAR point clouds, given its 3D bounding box in the first fram

safetyarxiv-cs-cv
16 Apr 2026
Safety

The Long Delay to Arithmetic Generalization: When Learned Representations Outrun Behavior

DGX agent

arXiv:2604.13082v1 Announce Type: new Abstract: Grokking in transformers trained on algorithmic tasks is characterized by a long delay between training-set fit and abrupt generalization, but the sourc

safetyarxiv-cs-lg
16 Apr 2026
Safety

Think Outside the Policy: In-Context Steered Policy Optimization

DGX agent

arXiv:2510.26519v3 Announce Type: replace Abstract: Existing Reinforcement Learning from Verifiable Rewards (RLVR) methods, such as Group Relative Policy Optimization (GRPO), have achieved remarkable

safetyarxiv-cs-lg
16 Apr 2026
Safety

TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks

DGX agent

arXiv:2601.10245v2 Announce Type: replace-cross Abstract: Multi-step reasoning tasks like mathematical problem solving are vulnerable to cascading failures, where a single incorrect step leads to comp

safetyarxiv-cs-cl
16 Apr 2026
Safety

UMI-3D: Extending Universal Manipulation Interface from Vision-Limited to 3D Spatial Perception

DGX agent

arXiv:2604.14089v1 Announce Type: new Abstract: We present UMI-3D, a multimodal extension of the Universal Manipulation Interface (UMI) for robust and scalable data collection in embodied manipulation

safetyarxiv-cs-ro
16 Apr 2026
Safety

UNBOX: Unveiling Black-box visual models with Natural-language

DGX agent

arXiv:2603.08639v2 Announce Type: replace Abstract: Ensuring trustworthiness in open-world visual recognition requires models that are interpretable, fair, and robust to distribution shifts. Yet moder

safetyarxiv-cs-cv
16 Apr 2026
Safety

What Are We Really Measuring? Rethinking Dataset Bias in Web-Scale Natural Image Collections via Unsupervised Semantic Clustering

DGX agent

arXiv:2604.13610v1 Announce Type: new Abstract: In computer vision, a prevailing method for quantifying dataset bias is to train a model to distinguish between datasets. High classification accuracy i

safetyarxiv-cs-cv
16 Apr 2026
Safety

Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking

DGX agent

arXiv:2604.13776v1 Announce Type: cross Abstract: Watermarking is becoming the default mechanism for AI content authentication, with governance policies and frameworks referencing it as infrastructure

safetyarxiv-cs-cl
16 Apr 2026
Safety

Why dynamically routing multi-timescale advantages in PPO causes policy collapse (and a simple decoupled fix) [R]

DGX agent

This Reddit post discusses a known instability in PPO when advantage estimates operating across different temporal scales (e.g., short-horizon and long-horizon returns) are dynamically routed or mixed

safetyr-machinelearning
16 Apr 2026
Safety

Why MLLMs Struggle to Determine Object Orientations

DGX agent

arXiv:2604.13321v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) struggle with tasks that require reasoning about 2D object orientation in images, as documented in prior work.

safetyarxiv-cs-cv
16 Apr 2026
Safety

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks

DGX agent

arXiv:2604.13403v1 Announce Type: new Abstract: In-context learning (ICL) enables models to adapt to new tasks via inference-time demonstrations. Despite its success in large language models, the exte

safetyarxiv-cs-cv
16 Apr 2026
← Previous
1…260261262263264…299
Next →