AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,357 results
7 May 2026

OracleProto: A Reproducible Framework for Benchmarking LLM Native Forecasting via Knowledge Cutoff and Temporal Masking

SafetyDGX agent

arXiv:2605.03762v1 Announce Type: new Abstract: Large language models are moving from static text generators toward real-world decision-support systems, where forecasting is a composite capability tha

Order Matters: Improving Domain Adaptation by Reordering Data

SafetyDGX agent

arXiv:2605.05084v1 Announce Type: new Abstract: Domain shift remains a key challenge in deploying machine learning models to the real world. Unsupervised domain adaptation (UDA) aims to address this b

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage

SafetyDGX agent

arXiv:2506.07548v2 Announce Type: replace-cross Abstract: Multi-agent reinforcement learning (MARL) has reached competitive performance on cooperative tasks against scripted adversaries, yet most meth

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Overcoming reward signal challenges: Verifiable rewards-based reinforcement learning with GRPO on SageMaker AI

SafetyDGX agent

In this post, you will learn how to implement reinforcement learning with verifiable rewards (RLVR) to introduce verification and transparency into reward signals to improve training performance. This

PhySe-RPO: Physics and Semantics Guided Relative Policy Optimization for Diffusion-Based Surgical Smoke Removal

SafetyDGX agent

arXiv:2603.22844v4 Announce Type: replace Abstract: Surgical smoke severely degrades intraoperative video quality, obscuring anatomical structures and limiting surgical perception. Existing learning-b

POMA-3D: The Point Map Way to 3D Scene Understanding

SafetyDGX agent

arXiv:2511.16567v3 Announce Type: replace Abstract: In this paper, we introduce POMA-3D, the first self-supervised 3D representation model learned from point maps. Point maps encode explicit 3D coordi

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

SafetyDGX agent

arXiv:2605.05040v1 Announce Type: new Abstract: On-policy distillation is an efficient alternative to reinforcement learning, offering dense token-level training signals. However, its reliance on a st

ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection

SafetyDGX agent

arXiv:2601.09195v3 Announce Type: replace Abstract: Supervised fine-tuning (SFT) is a fundamental post-training strategy to align Large Language Models (LLMs) with human intent. However, traditional S

Provable imitation learning for control of instability in partially-observed Vlasov--Poisson equations

SafetyDGX agent

arXiv:2605.05081v1 Announce Type: new Abstract: We consider the stabilization of Vlasov--Poisson plasma dynamics, a central control problem in nuclear fusion. Our focus is the gap between what an idea

Purdah and Patriarchy: Evaluating and Mitigating South Asian Biases in Open-Ended Multilingual LLM Generations

SafetyDGX agent

arXiv:2505.18466v2 Announce Type: replace Abstract: Evaluations of Large Language Models (LLMs) often overlook intersectional and culturally specific biases, particularly in underrepresented multiling

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents

SafetyDGX agent

arXiv:2604.03976v2 Announce Type: replace Abstract: Prior work on trustworthy AI emphasizes model-internal properties such as bias mitigation, adversarial robustness, and interpretability. As AI syste

Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization

SafetyDGX agent

arXiv:2605.04920v1 Announce Type: cross Abstract: Compositional generalization refers to correctly interpret novel combinations of known primitives, which remains a major challenge. Existing approache

Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime

SafetyDGX agent

arXiv:2605.05112v1 Announce Type: new Abstract: SWE-bench-style agentic reinforcement learning relies on expensive stateful trajectories, yet substantial compute is wasted on sampled rollout groups wi

S1-MMAlign: A Large-Scale, Multi-Disciplinary Dataset for Scientific Figure-Text Understanding

SafetyDGX agent

arXiv:2601.00264v2 Announce Type: replace Abstract: Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap be

Scalable inference of spatial regions and temporal signatures from time series

SafetyDGX agent

arXiv:2605.05008v1 Announce Type: cross Abstract: Regionalization aims to partition a spatial domain into contiguous regions that share similar characteristics, enabling more effective spatial analysi

Scalable Multi Agent Diffusion Policies for Coverage Control

SafetyDGX agent

arXiv:2509.17244v2 Announce Type: replace Abstract: We propose MADP, a novel diffusion-model-based approach for collaboration in decentralized robot swarms. MADP leverages diffusion models to generate

Scalable Policy Maximization Under Network Interference

SafetyDGX agent

arXiv:2505.18118v2 Announce Type: replace-cross Abstract: Many interventions, such as vaccines in clinical trials or coupons in online marketplaces, must be assigned sequentially without full knowledg

Sequential Strategic Classification with Multi-Stage Selective Classifiers

SafetyDGX agent

arXiv:2605.04202v1 Announce Type: new Abstract: Strategic classification studies the problem where self-interested individuals or agents manipulate their response to obtain favorable decision outcomes

Structural Equivalence and Learning Dynamics in Delayed MARL

SafetyDGX agent

arXiv:2605.04345v1 Announce Type: new Abstract: We formally establish the equivalence between Observation Delay (OD) and Action Delay (AD) in cooperative partially observable multi-agent systems using

Temporal Structure Matters for Efficient Test-Time Adaptation in Wearable Human Activity Recognition

SafetyDGX agent

arXiv:2605.04617v1 Announce Type: new Abstract: Wearable human activity recognition (WHAR) models often suffer from performance degradation under real-world cross-user distribution shifts. Test-time a

The balcony solar boom is coming to the US

SafetyDGX agent

Dozens of US states are considering legislation to allow people to install plug-in solar systems, often called balcony solar. These small arrays require little to no setup and could help cut emissions

The illustration on this story is quite funny, but the study itself has quite big implications I think. The point is that there might be bet…

SafetyDGX agent

The illustration on this story is quite funny, but the study itself has quite big implications I think. The point is that there might be better ways to design AI systems so that they don’t simply do e

Theories are proven by accurate predictions. I first read @GaryMarcus in the late ’90s, and his ideas helped shape my deterministic ML frame…

SafetyDGX agent

Theories are proven by accurate predictions. I first read @GaryMarcus in the late ’90s, and his ideas helped shape my deterministic ML frameworks, GAIuS & KATO. He’s been proven right repeatedly, yet

“they hadn’t figured out how OpenAI would pay for it” may turn out to be the epitaph for an entire era. 🪦 scoop from @anissagardizy8 @thein…

SafetyDGX agent

“they hadn’t figured out how OpenAI would pay for it” may turn out to be the epitaph for an entire era. 🪦 scoop from @anissagardizy8 @theinformation, and credit her with the great line “OpenAI has mad

Threshold-Guided Optimization for Visual Generative Models

SafetyDGX agent

arXiv:2605.04653v1 Announce Type: new Abstract: Aligning large visual generative models with human feedback is often performed through pairwise preference optimization. While such approaches are conce

Time series causal discovery with variable lags

SafetyDGX agent

arXiv:2605.04081v1 Announce Type: new Abstract: Causal Bayesian Networks (CBNs) are a powerful tool for reasoning under uncertainty about complex real-world problems. Such problems evolve over time, r

To Fuse or to Drop? Dual-Path Learning for Resolving Modality Conflicts in Multimodal Emotion Recognition

SafetyDGX agent

arXiv:2605.04877v1 Announce Type: cross Abstract: Multimodal emotion recognition (MER) benefits from combining text, audio, and vision, yet standard fusion often fails when modalities conflict. Crucia

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium

SafetyDGX agent

arXiv:2605.04494v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has been popular for aligning text-to-image (T2I) diffusion models with human preferences. As a main

Training-Time Batch Normalization Reshapes Local Partition Geometry in Piecewise-Affine Networks

SafetyDGX agent

arXiv:2605.04946v1 Announce Type: new Abstract: Batch normalization (BN) is central to modern deep networks, but its effect on the realized function during training remains less understood than its op

Two years ago. I stand by both predictions.

SafetyDGX agent

Gary Marcus reflects on predictions he made two years prior and reaffirms his confidence in their accuracy. The post suggests Marcus is reviewing his track record on forecasts, likely related to artif

UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning

SafetyDGX agent

arXiv:2508.11196v2 Announce Type: replace Abstract: Recent advances in vision-language models (VLMs) have demonstrated strong generalization in natural image tasks. However, their performance often de

UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization

SafetyDGX agent

arXiv:2511.08195v3 Announce Type: replace Abstract: UI-to-code aims to translate UI screenshots into executable front-end code. Despite progress with vision-language models (VLMs), most existing metho

ULF-Loc: Unbiased Landmark Feature for Robust Visual Localization with 3D Gaussian Splatting

SafetyDGX agent

arXiv:2605.04730v1 Announce Type: new Abstract: Visual localization is a core technology for augmented reality and autonomous navigation. Recent methods combine the efficient rendering of 3D Gaussian

Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models

SafetyDGX agent

arXiv:2605.04874v1 Announce Type: cross Abstract: Direct Preference Optimization (DPO) has proven to be an effective solution for mitigating hallucination in Multimodal Large Language Models (MLLMs) b

Unifying Dynamical Systems and Graph Theory to Mechanistically Understand Computation in Neural Networks

SafetyDGX agent

arXiv:2605.03598v2 Announce Type: cross Abstract: Understanding how biological and artificial neural networks implement computation from connectivity is a central problem in neuroscience and machine l

UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings

SafetyDGX agent

arXiv:2505.11815v2 Announce Type: replace Abstract: Current vision-language models have been explored for multi-modal embedding tasks like information retrieval. However, they face significant challen

Using Common Random Numbers for Simulation-based Planning with Rollouts

SafetyDGX agent

arXiv:2605.04732v1 Announce Type: new Abstract: Simulation-based planning with rollouts is a widely-deployed technique for decision making in stochastic environments. The primary instrument of simulat

Variance Matters: Improving Domain Adaptation via Stratified Sampling

SafetyDGX agent

arXiv:2512.05226v2 Announce Type: replace Abstract: Domain shift remains a key challenge in deploying machine learning models to the real world. Unsupervised domain adaptation (UDA) aims to address th

What happens when AIs become smarter than us? Why would they keep humans around if given the choice? Our new paper argues that only trying t…

SafetyDGX agent

What happens when AIs become smarter than us? Why would they keep humans around if given the choice? Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutu

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning

SafetyDGX agent

arXiv:2605.05172v1 Announce Type: new Abstract: Behavior Cloning (BC) has emerged as a highly effective paradigm for robot learning. However, BC lacks a self-guided mechanism for online improvement af

Why Expert Alignment Is Hard: Evidence from Subjective Evaluation

SafetyDGX agent

arXiv:2605.04972v1 Announce Type: new Abstract: Aligning large language models with expert judgment is especially difficult in subjective evaluation tasks, where experts may disagree, rely on tacit cr

Wow this paper has been “just published” many times for nearly a year. I have called out at least two other “influencers” for the same thing…

SafetyDGX agent

Wow this paper has been “just published” many times for nearly a year. I have called out at least two other “influencers” for the same thing on the same paper, including another one earlier this week.

yep, really embarassing

SafetyDGX agent

yep, really embarassing One of the richest VCs, Marc Andreessen, has shown again that he thinks generative AI works perfectly well if you use the right prompts, thus ignoring the statistical nature of

6 May 2026

100% stood the test of time: “What Ilya saw” was Sam’s bad behavior, not AGI.

SafetyDGX agent

100% stood the test of time: “What Ilya saw” was Sam’s bad behavior, not AGI. When Altman got fired “What did Ilya see?” became a wildly popular conspiracy meme. I think we can safely say now that wha

2nd episode of The Roman Forum is an interview with AI Safety/Governance expert Connor Leahy @NPCollapse. Connor is a great speaker and is l…

SafetyDGX agent

2nd episode of The Roman Forum is an interview with AI Safety/Governance expert Connor Leahy @NPCollapse. Connor is a great speaker and is lobbying to get government to ban Superintelligence. My first

A Robust Unsupervised Domain Adaptation Framework for Medical Image Classification Using RKHS-MMD

SafetyDGX agent

arXiv:2605.03787v1 Announce Type: new Abstract: Labeling medical images is a major bottleneck in the field of medical imaging, as it requires domain-specific expertise, and it gets further complicated

A Universal Reproducing Kernel Hilbert Space from Polynomial Alignment and IMQ Distance

SafetyDGX agent

arXiv:2605.03262v1 Announce Type: new Abstract: We introduce the Yat kernel $k_{b,arepsilon}(mathbf{w},mathbf{x})=frac{(mathbf{w}^opmathbf{x}+b)^2}{|mathbf{x}-mathbf{w}|^2+arepsilon},qquad bge 0, arep

A US appeals court strikes down a 2023 FCC rule banning broadband access discrimination based on income, race, and more; Chair Brendan Carr welcomes the ruling (Jon Brodkin/Ars Technica)

SafetyDGX agent

Jon Brodkin / Ars Technica: A US appeals court strikes down a 2023 FCC rule banning broadband access discrimination based on income, race, and more; Chair Brendan Carr welcomes the ruling — An appeals

Adaptive 3D-RoPE: Physics-Aligned Rotary Positional Encoding for Wireless Foundation Models

SafetyDGX agent

arXiv:2605.00968v1 Announce Type: cross Abstract: Positional encoding plays a pivotal role in determin?ing the extrapolation and generalization performance of wireless foundation models for channel st

ADAPTS: Agentic Decomposition for Automated Protocol-agnostic Tracking of Symptoms

SafetyDGX agent

arXiv:2605.03212v1 Announce Type: cross Abstract: Modeling latent clinical constructs from unconstrained clinical interactions is a unique challenge in affective computing. We present ADAPTS (Agentic

Agentic AI-Based Joint Computing and Networking via Mixture of Experts and Large Language Models

SafetyDGX agent

arXiv:2605.02911v1 Announce Type: new Abstract: Future sixth-generation (6G) mobile networks are envisioned to be equipped with a diverse set of powerful, yet highly specialized, optimization experts.

Aligning Inductive Bias for Data-Efficient Generalization in State Space Models

SafetyDGX agent

arXiv:2509.20789v4 Announce Type: replace Abstract: The remarkable success of modern AI has been closely tied to scaling laws, yet the finite supply of high-quality data makes data efficiency--learnin

🚨Alphabet $GOOGL trades at 133x free cash flow. For context: its pre-COVID multiple was ~20x. And free cash flow hasn't grown since 2021. G…

SafetyDGX agent

🚨Alphabet GOOGL trades at 133x free cash flow. For context: its pre-COVID multiple was ~20x. And free cash flow hasn't grown since 2021. GQG Partners — one of the world's top institutional investors —

Am I right that hyperscaling compute is the biggest bet in history? Any counter examples? It’s way more expensive than the Manhattan Project…

SafetyDGX agent

Am I right that hyperscaling compute is the biggest bet in history? Any counter examples? It’s way more expensive than the Manhattan Project, the Apollo project, and railways across the US. If it does

Amazing the shit X gave me in November 2023 for saying Sam wasn’t always candid. He wasn’t. Period.

SafetyDGX agent

Amazing the shit X gave me in November 2023 for saying Sam wasn’t always candid. He wasn’t. Period. Murati testified that, by fall 2023, Sam Altman was “not always” candid with her, and undermined her

And how many times have I told you that Sam is no longer the right CEO for OpenAI?

SafetyDGX agent

And how many times have I told you that Sam is no longer the right CEO for OpenAI? Murati: My issues with Sam were very much around management and providing direction to the organization and decision

Anthropic researchers detail 'model spec midtraining', which adds a stage between pretraining and fine-tuning to improve generalization from alignment training (Anthropic)

SafetyDGX agent

Anthropic: Anthropic researchers detail “model spec midtraining”, which adds a stage between pretraining and fine-tuning to improve generalization from alignment training — Sara Price2, Samuel Marks2,

Beyond Activation Alignment: The Geometry of Neural Sensitivity

SafetyDGX agent

arXiv:2605.03222v1 Announce Type: new Abstract: Activation-alignment measures such as Representational Similarity Analysis (RSA), Canonical Correlation Analysis (CCA), and Centered Kernel Alignment (C

BifrostUMI: Bridging Robot-Free Demonstrations and Humanoid Whole-Body Manipulation

SafetyDGX agent

arXiv:2605.03452v1 Announce Type: new Abstract: High-quality data collection is a fundamental cornerstone for training humanoid whole-body visuomotor policies. Current data acquisition paradigms predo

Bootstrapped Mixed Rewards for RL Post-Training: Injecting Canonical Action Order

SafetyDGX agent

arXiv:2512.04277v3 Announce Type: replace Abstract: Post-training with reinforcement learning (RL) typically optimizes a single scalar objective and ignores structure in how solutions are produced. We

← Previous
1…182183184185186…240
Next →