AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,600 results
16 Apr 2026

Heavy-Tailed Class-Conditional Priors for Long-Tailed Generative Modeling

SafetyDGX agent

arXiv:2509.02154v2 Announce Type: replace-cross Abstract: Variational Autoencoders (VAEs) with global priors trained under an imbalanced empirical class distribution can lead to underrepresentation of

I have found that asking for a sestina regularly triggers Opus 4.7's safety guardrails. The forbidden poetic form!

SafetyDGX agent

Ethan Mollick reported that requesting Claude Opus 4.7 to write sestinas—a complex poetic form with strict structural requirements—frequently triggers the model's safety guardrails, suggesting the AI

' If a superintelligence is built, humanity will lose control over its future.' Before Canadian Senators, ControlAI's US Director Connor Lea…

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

' If a superintelligence is built, humanity will lose control over its future.' Before Canadian Senators, ControlAI's US Director Connor Leahy (@NPCollapse) explains why superintelligence could lead t

IGen: Scalable Data Generation for Robot Learning from Open-World Images

SafetyDGX agent

arXiv:2512.01773v2 Announce Type: replace Abstract: The rise of generalist robotic policies has created an exponential demand for large-scale training data. However, on-robot data collection is labor-

Irregularly Sampled Time Series Interpolation for Binary Evolution Simulations Using Dynamic Time Warping

SafetyDGX agent

arXiv:2604.13604v1 Announce Type: cross Abstract: Binary stellar evolution simulations are computationally expensive. Stellar population synthesis relies on these detailed evolution models at a fundam

Jump-Start Reinforcement Learning with Vision-Language-Action Regularization

SafetyDGX agent

arXiv:2604.13733v1 Announce Type: new Abstract: Reinforcement learning (RL) enables high-frequency, closed-loop control for robotic manipulation, but scaling to long-horizon tasks with sparse or imper

Learning Class Difficulty in Imbalanced Histopathology Segmentation via Dynamic Focal Attention

SafetyDGX agent

arXiv:2604.13479v1 Announce Type: cross Abstract: Semantic segmentation of histopathology images under class imbalance is typically addressed through frequency-based loss reweighting, which implicitly

Learning Probabilistic Responsibility Allocations for Multi-Agent Interactions

SafetyDGX agent

arXiv:2604.13128v1 Announce Type: cross Abstract: Human behavior in interactive settings is shaped not only by individual objectives but also by shared constraints with others, such as safety. Underst

Maybe because of this paper? https://x.com/emollick/status/1991624198855561508?s=20

SafetyDGX agent

Maybe because of this paper? https://x.com/emollick/status/1991624198855561508?s=20 Tell all the truth but tell it slant— Success in Circuit lies Too bright for our infirm Delight The Truth's superb s

Med-CAM: Minimal Evidence for Explaining Medical Decision Making

SafetyDGX agent

arXiv:2604.13695v1 Announce Type: new Abstract: Reliable and interpretable decision-making is essential in medical imaging, where diagnostic outcomes directly influence patient care. Despite advances

Multi-Dimensional Knowledge Profiling with Large-Scale Literature Database and Hierarchical Retrieval

SafetyDGX agent

arXiv:2601.15170v2 Announce Type: replace Abstract: The rapid expansion of research across machine learning, vision, and language has produced a volume of publications that is increasingly difficult t

Neural Mean-Field Games: Extending Mean-Field Game Theory with Neural Stochastic Differential Equations

SafetyDGX agent

arXiv:2504.13228v4 Announce Type: replace Abstract: Mean-field game theory relies on approximating games that are intractable to model due to a very large to infinite population of players. While thes

NHTSA autonomous vehicle crash data has been updated through March 15, 2026, for AVs including Tesla Robotaxi. This includes unsupervised Te…

SafetyDGX agent

NHTSA autonomous vehicle crash data has been updated through March 15, 2026, for AVs including Tesla Robotaxi. This includes unsupervised Tesla-driven robotaxis. • Waymo: 58 incidents • Zoox: 3 incide

Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning

SafetyDGX agent

arXiv:2506.08125v3 Announce Type: replace-cross Abstract: Large language models (LLMs) show strong reasoning abilities but often produce unnecessarily long explanations that reduce efficiency. Althoug

On the Fundamental Limitations of Dual Static CVaR Decompositions in Markov Decision Processes

SafetyDGX agent

arXiv:2507.14005v2 Announce Type: replace Abstract: It was recently shown that dynamic programming (DP) methods for finding static CVaR-optimal policies in Markov Decision Processes (MDPs) can fail wh

OpenAI's global policy chief, Chris Lehane, calls out AI 'doomers' and says 'when you put some of those thoughts and ideas out there, they do have consequences' (Caroline O'Donovan/The San Francisco ...)

SafetyDGX agent

Caroline O'Donovan / The San Francisco Standard: OpenAI's global policy chief, Chris Lehane, calls out AI “doomers” and says “when you put some of those thoughts and ideas out there, they do have cons

Pareto-Optimal Offline Reinforcement Learning via Smooth Tchebysheff Scalarization

SafetyDGX agent

arXiv:2604.13175v1 Announce Type: new Abstract: Large language models can be aligned with human preferences through offline reinforcement learning (RL) on small labeled datasets. While single-objectiv

Reducing Hallucinations in LLMs via Factuality-Aware Preference Learning

SafetyDGX agent

arXiv:2601.03027v3 Announce Type: replace Abstract: Preference alignment methods such as RLHF and Direct Preference Optimization (DPO) improve instruction following, but they can also reinforce halluc

Representation over Routing: Overcoming Surrogate Hacking in Multi-Timescale PPO

SafetyDGX agent

arXiv:2604.13517v1 Announce Type: new Abstract: Temporal credit assignment in reinforcement learning has long been a central challenge. Inspired by the multi-timescale encoding of the dopamine system

Restless Bandits with Individual Penalty Constraints: A New Near-Optimal Index Policy and How to Learn It

SafetyDGX agent

arXiv:2604.04101v2 Announce Type: replace Abstract: This paper investigates the Restless Multi-Armed Bandit (RMAB) framework under individual penalty constraints to address resource allocation challen

Rethinking Uncertainty in Segmentation: From Estimation to Decision

SafetyDGX agent

arXiv:2604.13262v1 Announce Type: new Abstract: In medical image segmentation, uncertainty estimates are often reported but rarely used to guide decisions. We study the missing step: how uncertainty m

RFK Jr. forces FDA to reconsider 12 unproven peptides after 2023 ban

SafetyDGX agent

In 2023, the FDA removed 19 peptides from the list of drugs that compounding pharmacies could produce, and in 2026 the FDA announced it will review whether to add back 7 of these peptides following pr

RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization

SafetyDGX agent

arXiv:2508.00222v5 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Reward (RLVR) has significantly advanced the complex reasoning abilities of Large Language Models (LLMs

RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph

SafetyDGX agent

arXiv:2511.07717v2 Announce Type: replace-cross Abstract: Estimating robot pose from a monocular RGB image is a challenge in robotics and computer vision. Existing methods typically build networks on

Robust Energy-Aware Routing for Air-Ground Cooperative Multi-UAV Delivery in Wind-Uncertain Environments

SafetyDGX agent

arXiv:2604.13441v1 Announce Type: new Abstract: Ensuring energy feasibility under wind uncertainty is critical for the safety and reliability of UAV delivery missions. In realistic truck-drone logisti

Robust Verification of Controllers under State Uncertainty via Hamilton-Jacobi Reachability Analysis

SafetyDGX agent

arXiv:2511.14755v2 Announce Type: replace-cross Abstract: As perception-based controllers for autonomous systems become increasingly popular in the real world, it is important that we can formally ver

Safe and Nonconservative Contingency Planning for Autonomous Vehicles via Online Learning-Based Reachable Set Barriers

SafetyDGX agent

arXiv:2509.07464v2 Announce Type: replace Abstract: Autonomous vehicles must navigate dynamically uncertain environments while balancing safety and efficiency. This challenge is exacerbated by unpredi

See&Say: Vision Language Guided Safe Zone Detection for Autonomous Package Delivery Drones

SafetyDGX agent

arXiv:2604.13292v1 Announce Type: new Abstract: Autonomous drone delivery systems are rapidly advancing, but ensuring safe and reliable package drop-offs remains highly challenging in cluttered urban

Self-adaptive Multi-Access Edge Architectures: A Robotics Case

SafetyDGX agent

arXiv:2604.13542v1 Announce Type: new Abstract: The growth of compute-intensive AI tasks highlights the need to mitigate the processing costs and improve performance and energy efficiency. This necess

SFT-GRPO Data Overlap as a Post-Training Hyperparameter for Autoformalization

SafetyDGX agent

arXiv:2604.13515v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) followed by Group Relative Policy Optimization (GRPO) is a common post-training recipe. We conduct a controlled ablation ov

Soft Q(lambda): A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces

SafetyDGX agent

arXiv:2604.13780v1 Announce Type: new Abstract: Soft Q-learning has emerged as a versatile model-free method for entropy-regularised reinforcement learning, optimising for returns augmented with a pen

Spectral methods: crucial for machine learning, natural for quantum computers?

SafetyDGX agent

arXiv:2603.24654v2 Announce Type: replace-cross Abstract: This article presents an argument for why quantum computers could unlock new methods for machine learning. We argue that spectral methods, in

SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models

SafetyDGX agent

arXiv:2510.09541v3 Announce Type: replace Abstract: Diffusion large language models (dLLMs) are emerging as an efficient alternative to autoregressive models due to their ability to decode multiple to

Supreme Court Justice Clarence Thomas on Wednesday delivered a televised broadside against progressivism, a political philosophy he describe…

SafetyDGX agent

Supreme Court Justice Clarence Thomas on Wednesday delivered a televised broadside against progressivism, a political philosophy he described as an existential threat to America and the principles tha

Temporally Consistent Long-Term Memory for 3D Single Object Tracking

SafetyDGX agent

arXiv:2604.13789v1 Announce Type: new Abstract: 3D Single Object Tracking (3D-SOT) aims to localize a target object across a sequence of LiDAR point clouds, given its 3D bounding box in the first fram

The Long Delay to Arithmetic Generalization: When Learned Representations Outrun Behavior

SafetyDGX agent

arXiv:2604.13082v1 Announce Type: new Abstract: Grokking in transformers trained on algorithmic tasks is characterized by a long delay between training-set fit and abrupt generalization, but the sourc

Think Outside the Policy: In-Context Steered Policy Optimization

SafetyDGX agent

arXiv:2510.26519v3 Announce Type: replace Abstract: Existing Reinforcement Learning from Verifiable Rewards (RLVR) methods, such as Group Relative Policy Optimization (GRPO), have achieved remarkable

TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks

SafetyDGX agent

arXiv:2601.10245v2 Announce Type: replace-cross Abstract: Multi-step reasoning tasks like mathematical problem solving are vulnerable to cascading failures, where a single incorrect step leads to comp

UMI-3D: Extending Universal Manipulation Interface from Vision-Limited to 3D Spatial Perception

SafetyDGX agent

arXiv:2604.14089v1 Announce Type: new Abstract: We present UMI-3D, a multimodal extension of the Universal Manipulation Interface (UMI) for robust and scalable data collection in embodied manipulation

UNBOX: Unveiling Black-box visual models with Natural-language

SafetyDGX agent

arXiv:2603.08639v2 Announce Type: replace Abstract: Ensuring trustworthiness in open-world visual recognition requires models that are interpretable, fair, and robust to distribution shifts. Yet moder

Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images

SafetyDGX agent

arXiv:2603.08486v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) face safety misalignment, where visual inputs enable harmful outputs. To address this, existing methods req

VRAG-DFD: Verifiable Retrieval-Augmentation for MLLM-based Deepfake Detection

SafetyDGX agent

arXiv:2604.13660v1 Announce Type: new Abstract: In Deepfake Detection (DFD) tasks, researchers proposed two types of MLLM-based methods: complementary combination with small DFD detectors, or static f

What Are We Really Measuring? Rethinking Dataset Bias in Web-Scale Natural Image Collections via Unsupervised Semantic Clustering

SafetyDGX agent

arXiv:2604.13610v1 Announce Type: new Abstract: In computer vision, a prevailing method for quantifying dataset bias is to train a model to distinguish between datasets. High classification accuracy i

Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking

SafetyDGX agent

arXiv:2604.13776v1 Announce Type: cross Abstract: Watermarking is becoming the default mechanism for AI content authentication, with governance policies and frameworks referencing it as infrastructure

Why dynamically routing multi-timescale advantages in PPO causes policy collapse (and a simple decoupled fix) [R]

SafetyDGX agent

This Reddit post discusses a known instability in PPO when advantage estimates operating across different temporal scales (e.g., short-horizon and long-horizon returns) are dynamically routed or mixed

Why MLLMs Struggle to Determine Object Orientations

SafetyDGX agent

arXiv:2604.13321v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) struggle with tasks that require reasoning about 2D object orientation in images, as documented in prior work.

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks

SafetyDGX agent

arXiv:2604.13403v1 Announce Type: new Abstract: In-context learning (ICL) enables models to adapt to new tasks via inference-time demonstrations. Despite its success in large language models, the exte

ZK-APEX: Zero-Knowledge Approximate Personalized Unlearning with Executable Proofs

SafetyDGX agent

arXiv:2512.09953v2 Announce Type: replace-cross Abstract: Machine unlearning aims to remove the influence of specific data points from a trained model to satisfy privacy, copyright, and safety require

15 Apr 2026

4/5 We upgraded our original 3-line “be correct” prompt → a much more detailed prompt that enforced a hierarchy of constraints for correctne…

SafetyDGX agent

4/5 We upgraded our original 3-line “be correct” prompt → a much more detailed prompt that enforced a hierarchy of constraints for correctness, regression safety, and minimality. Basically, get the LL

5/5 The takeaway: If your agent relies on an LLM judge for selection accuracy, measuring code quality isn’t enough; you need a measure of th…

SafetyDGX agent

5/5 The takeaway: If your agent relies on an LLM judge for selection accuracy, measuring code quality isn’t enough; you need a measure of the model's inductive bias toward the 'fingerprint' of a gold

A Comparison of Reinforcement Learning and Optimal Control Methods for Path Planning

SafetyDGX agent

arXiv:2604.12628v1 Announce Type: cross Abstract: Path-planning for autonomous vehicles in threat-laden environments is a fundamental challenge. While traditional optimal control methods can find idea

A Dataset and Evaluation for Complex 4D Markerless Human Motion Capture

SafetyDGX agent

arXiv:2604.12765v1 Announce Type: new Abstract: Marker-based motion capture (MoCap) systems have long been the gold standard for accurate 4D human modeling, yet their reliance on specialized hardware

A document is worth a structured record: Principled inductive bias design for document recognition

SafetyDGX agent

arXiv:2507.08458v2 Announce Type: replace-cross Abstract: Many document types use intrinsic, convention-driven structures that serve to encode precise and structured information, such as the conventio

A longitudinal health agent framework

SafetyDGX agent

arXiv:2604.12019v1 Announce Type: new Abstract: Although artificial intelligence (AI) agents are increasingly proposed to support potentially longitudinal health tasks, such as symptom management, beh

A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning

SafetyDGX agent

arXiv:2603.08291v3 Announce Type: replace Abstract: Multimodal Mathematical Reasoning (MMR) has recently attracted increasing attention for its capability to solve mathematical problems involving both

A TTP analysis finds dozens of nudify apps in Apple and Google app stores via search, despite company policy prohibiting them; the apps earned $122M+ in revenue (Bloomberg)

SafetyDGX agent

Bloomberg: A TTP analysis finds dozens of nudify apps in Apple and Google app stores via search, despite company policy prohibiting them; the apps earned $122M+ in revenue — Apple Inc. and Google have

A Two-Timescale Primal-Dual Framework for Reinforcement Learning via Online Dual Variable Guidance

SafetyDGX agent

arXiv:2505.04494v3 Announce Type: replace-cross Abstract: We study reinforcement learning by combining recent advances in regularized linear programming formulations with the classical theory of stoch

AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage Margin

SafetyDGX agent

arXiv:2505.14264v3 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as an effective approach for enhancing the reasoning capabilities of large language models (LLMs), esp

Active Imitation Learning for Thermal- and Kernel-Aware LFM Inference on 3D S-NUCA Many-Cores

SafetyDGX agent

arXiv:2604.11948v1 Announce Type: new Abstract: Large Foundation Model (LFM) inference is both memory- and compute-intensive, traditionally relying on GPUs. However, the limited availability and high

Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning

SafetyDGX agent

arXiv:2505.17086v4 Announce Type: replace Abstract: Large Language Models (LLMs) equipped with modern Retrieval-Augmented Generation (RAG) systems often employ multi-turn interaction pipelines to inte

← Previous
1…194195196197198…210
Next →