AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,707 results
7 May 2026

New Bigtable in-memory tier for sub-millisecond read latency

SafetyDGX agent

In the high-stakes world of digital infrastructure, speed isn't just a metric — it’s currency. At Google Cloud Next ‘26 we announced the Bigtable in-memory tier, a breakthrough for our fully managed c

On-line Learning in Tree MDPs by Treating Policies as Bandit Arms

SafetyDGX agent

arXiv:2605.04979v1 Announce Type: cross Abstract: A Tree Markov Decision Problem (T-MDP) is a finite-horizon MDP with a starting state s_{1}, in which every state is reachable from s_{1} through exact

On the Hardness of Junking LLMs

SafetyDGX agent

arXiv:2605.05116v1 Announce Type: new Abstract: Large language models (LLMs) are known to be vulnerable to jailbreak attacks, which typically rely on carefully designed prompts containing explicit sem


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

One of the things that made the Mythos release hard to interpret is that Anthropic held back details on most vulns they found, to give defen…

SafetyDGX agent

One of the things that made the Mythos release hard to interpret is that Anthropic held back details on most vulns they found, to give defenders time to patch. 1 month later, info from orgs with acces

One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender Serving

SafetyDGX agent

arXiv:2605.04450v1 Announce Type: cross Abstract: Generative Recommender (GR) inference places embedding hot caches (EMB) and KV caches in direct competition for limited GPU HBM: allocating more memor

OracleProto: A Reproducible Framework for Benchmarking LLM Native Forecasting via Knowledge Cutoff and Temporal Masking

SafetyDGX agent

arXiv:2605.03762v1 Announce Type: new Abstract: Large language models are moving from static text generators toward real-world decision-support systems, where forecasting is a composite capability tha

Order Matters: Improving Domain Adaptation by Reordering Data

SafetyDGX agent

arXiv:2605.05084v1 Announce Type: new Abstract: Domain shift remains a key challenge in deploying machine learning models to the real world. Unsupervised domain adaptation (UDA) aims to address this b

Overcoming Environmental Meta-Stationarity in MARL via Adaptive Curriculum and Counterfactual Group Advantage

SafetyDGX agent

arXiv:2506.07548v2 Announce Type: replace-cross Abstract: Multi-agent reinforcement learning (MARL) has reached competitive performance on cooperative tasks against scripted adversaries, yet most meth

Overcoming reward signal challenges: Verifiable rewards-based reinforcement learning with GRPO on SageMaker AI

SafetyDGX agent

In this post, you will learn how to implement reinforcement learning with verifiable rewards (RLVR) to introduce verification and transparency into reward signals to improve training performance. This

PhySe-RPO: Physics and Semantics Guided Relative Policy Optimization for Diffusion-Based Surgical Smoke Removal

SafetyDGX agent

arXiv:2603.22844v4 Announce Type: replace Abstract: Surgical smoke severely degrades intraoperative video quality, obscuring anatomical structures and limiting surgical perception. Existing learning-b

POMA-3D: The Point Map Way to 3D Scene Understanding

SafetyDGX agent

arXiv:2511.16567v3 Announce Type: replace Abstract: In this paper, we introduce POMA-3D, the first self-supervised 3D representation model learned from point maps. Point maps encode explicit 3D coordi

Practical validation of synthetic pre-crash scenarios

SafetyDGX agent

arXiv:2605.04564v1 Announce Type: new Abstract: The representativeness of synthetic pre-crash scenarios is crucial for assessing the safety impact of Driving Automation Systems through virtual simulat

Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs

SafetyDGX agent

arXiv:2605.04215v1 Announce Type: new Abstract: Diffusion-based Large Language Models (D-LLMs) represent a promising frontier in generative AI, offering fully parallel token generation that can lead t

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

SafetyDGX agent

arXiv:2605.05040v1 Announce Type: new Abstract: On-policy distillation is an efficient alternative to reinforcement learning, offering dense token-level training signals. However, its reliance on a st

ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection

SafetyDGX agent

arXiv:2601.09195v3 Announce Type: replace Abstract: Supervised fine-tuning (SFT) is a fundamental post-training strategy to align Large Language Models (LLMs) with human intent. However, traditional S

Provable imitation learning for control of instability in partially-observed Vlasov--Poisson equations

SafetyDGX agent

arXiv:2605.05081v1 Announce Type: new Abstract: We consider the stabilization of Vlasov--Poisson plasma dynamics, a central control problem in nuclear fusion. Our focus is the gap between what an idea

Purdah and Patriarchy: Evaluating and Mitigating South Asian Biases in Open-Ended Multilingual LLM Generations

SafetyDGX agent

arXiv:2505.18466v2 Announce Type: replace Abstract: Evaluations of Large Language Models (LLMs) often overlook intersectional and culturally specific biases, particularly in underrepresented multiling

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents

SafetyDGX agent

arXiv:2604.03976v2 Announce Type: replace Abstract: Prior work on trustworthy AI emphasizes model-internal properties such as bias mitigation, adversarial robustness, and interpretability. As AI syste

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments

SafetyDGX agent

arXiv:2508.04204v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have demonstrated impressive performance in reasoning-intensive tasks, but they remain vulnerable to harmful content g

Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization

SafetyDGX agent

arXiv:2605.04920v1 Announce Type: cross Abstract: Compositional generalization refers to correctly interpret novel combinations of known primitives, which remains a major challenge. Existing approache

Road Risk Monitor: A Deployable U.S. Road Incident Forecasting System with Live Weather and Road-Level Tiles

SafetyDGX agent

arXiv:2605.04242v1 Announce Type: new Abstract: Nationwide road-incident forecasting is a systems problem before it is a modeling problem. A usable service must connect historical incident archives, h

Robust Agent Compensation (RAC): Teaching AI Agents to Compensate

SafetyDGX agent

arXiv:2605.03409v1 Announce Type: new Abstract: We present Robust Agent Compensation (RAC), a log-based recovery paradigm (providing a safety net) implemented through an architectural extension that c

Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime

SafetyDGX agent

arXiv:2605.05112v1 Announce Type: new Abstract: SWE-bench-style agentic reinforcement learning relies on expensive stateful trajectories, yet substantial compute is wasted on sampled rollout groups wi

S1-MMAlign: A Large-Scale, Multi-Disciplinary Dataset for Scientific Figure-Text Understanding

SafetyDGX agent

arXiv:2601.00264v2 Announce Type: replace Abstract: Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap be

SafeRedir: Prompt Embedding Redirection for Robust Unlearning in Image Generation Models

SafetyDGX agent

arXiv:2601.08623v2 Announce Type: replace Abstract: Image generation models (IGMs), while capable of producing impressive and creative content, often memorize a wide range of undesirable concepts from

Safety by Invariance, Liveness through Refinement: Heterogeneous Contract Framework for Co-Design of Layered Control

SafetyDGX agent

arXiv:2605.04222v1 Announce Type: cross Abstract: Real-world control systems must achieve long-horizon objectives (liveness) while respecting continuous-time safety constraints, a combination that mot

Safety Must Precede the Deployment of Open-Ended AI

SafetyDGX agent

arXiv:2502.04512v3 Announce Type: replace Abstract: AI advancements have been significantly driven by a combination of foundation models and curiosity-driven learning aimed at increasing capability an

Scalable inference of spatial regions and temporal signatures from time series

SafetyDGX agent

arXiv:2605.05008v1 Announce Type: cross Abstract: Regionalization aims to partition a spatial domain into contiguous regions that share similar characteristics, enabling more effective spatial analysi

Scalable Multi Agent Diffusion Policies for Coverage Control

SafetyDGX agent

arXiv:2509.17244v2 Announce Type: replace Abstract: We propose MADP, a novel diffusion-model-based approach for collaboration in decentralized robot swarms. MADP leverages diffusion models to generate

Scalable Policy Maximization Under Network Interference

SafetyDGX agent

arXiv:2505.18118v2 Announce Type: replace-cross Abstract: Many interventions, such as vaccines in clinical trials or coupons in online marketplaces, must be assigned sequentially without full knowledg

Sequential Strategic Classification with Multi-Stage Selective Classifiers

SafetyDGX agent

arXiv:2605.04202v1 Announce Type: new Abstract: Strategic classification studies the problem where self-interested individuals or agents manipulate their response to obtain favorable decision outcomes

Software Engineering for Self-Adaptive Robotics: A Research Agenda

SafetyDGX agent

arXiv:2505.19629v3 Announce Type: replace-cross Abstract: Self-adaptive robotic systems operate autonomously in dynamic and uncertain environments, requiring robust real-time monitoring and adaptive b

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization

SafetyDGX agent

arXiv:2605.04700v1 Announce Type: cross Abstract: Jailbreak attacks on audio language models (ALMs) optimize audio perturbations to elicit unsafe generations, and they typically update the entire wave

Structural Equivalence and Learning Dynamics in Delayed MARL

SafetyDGX agent

arXiv:2605.04345v1 Announce Type: new Abstract: We formally establish the equivalence between Observation Delay (OD) and Action Delay (AD) in cooperative partially observable multi-agent systems using

Temporal Structure Matters for Efficient Test-Time Adaptation in Wearable Human Activity Recognition

SafetyDGX agent

arXiv:2605.04617v1 Announce Type: new Abstract: Wearable human activity recognition (WHAR) models often suffer from performance degradation under real-world cross-user distribution shifts. Test-time a

Terence Tao recognized that plausibility and veracity are not the same, and that current tools are better at the former than the latter. [ed…

SafetyDGX agent

Terence Tao recognized that plausibility and veracity are not the same, and that current tools are better at the former than the latter. [edit: the video is from 2024, and affirms what i said in 2019

The balcony solar boom is coming to the US

SafetyDGX agent

Dozens of US states are considering legislation to allow people to install plug-in solar systems, often called balcony solar. These small arrays require little to no setup and could help cut emissions

The illustration on this story is quite funny, but the study itself has quite big implications I think. The point is that there might be bet…

SafetyDGX agent

The illustration on this story is quite funny, but the study itself has quite big implications I think. The point is that there might be better ways to design AI systems so that they don’t simply do e

Theories are proven by accurate predictions. I first read @GaryMarcus in the late ’90s, and his ideas helped shape my deterministic ML frame…

SafetyDGX agent

Theories are proven by accurate predictions. I first read @GaryMarcus in the late ’90s, and his ideas helped shape my deterministic ML frameworks, GAIuS & KATO. He’s been proven right repeatedly, yet

“they hadn’t figured out how OpenAI would pay for it” may turn out to be the epitaph for an entire era. 🪦 scoop from @anissagardizy8 @thein…

SafetyDGX agent

“they hadn’t figured out how OpenAI would pay for it” may turn out to be the epitaph for an entire era. 🪦 scoop from @anissagardizy8 @theinformation, and credit her with the great line “OpenAI has mad

Threshold-Guided Optimization for Visual Generative Models

SafetyDGX agent

arXiv:2605.04653v1 Announce Type: new Abstract: Aligning large visual generative models with human feedback is often performed through pairwise preference optimization. While such approaches are conce

Time series causal discovery with variable lags

SafetyDGX agent

arXiv:2605.04081v1 Announce Type: new Abstract: Causal Bayesian Networks (CBNs) are a powerful tool for reasoning under uncertainty about complex real-world problems. Such problems evolve over time, r

To Fuse or to Drop? Dual-Path Learning for Resolving Modality Conflicts in Multimodal Emotion Recognition

SafetyDGX agent

arXiv:2605.04877v1 Announce Type: cross Abstract: Multimodal emotion recognition (MER) benefits from combining text, audio, and vision, yet standard fusion often fails when modalities conflict. Crucia

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium

SafetyDGX agent

arXiv:2605.04494v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has been popular for aligning text-to-image (T2I) diffusion models with human preferences. As a main

Training-Time Batch Normalization Reshapes Local Partition Geometry in Piecewise-Affine Networks

SafetyDGX agent

arXiv:2605.04946v1 Announce Type: new Abstract: Batch normalization (BN) is central to modern deep networks, but its effect on the realized function during training remains less understood than its op

Two years ago. I stand by both predictions.

SafetyDGX agent

Gary Marcus reflects on predictions he made two years prior and reaffirms his confidence in their accuracy. The post suggests Marcus is reviewing his track record on forecasts, likely related to artif

UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning

SafetyDGX agent

arXiv:2508.11196v2 Announce Type: replace Abstract: Recent advances in vision-language models (VLMs) have demonstrated strong generalization in natural image tasks. However, their performance often de

UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization

SafetyDGX agent

arXiv:2511.08195v3 Announce Type: replace Abstract: UI-to-code aims to translate UI screenshots into executable front-end code. Despite progress with vision-language models (VLMs), most existing metho

ULF-Loc: Unbiased Landmark Feature for Robust Visual Localization with 3D Gaussian Splatting

SafetyDGX agent

arXiv:2605.04730v1 Announce Type: new Abstract: Visual localization is a core technology for augmented reality and autonomous navigation. Recent methods combine the efficient rendering of 3D Gaussian

Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models

SafetyDGX agent

arXiv:2605.04874v1 Announce Type: cross Abstract: Direct Preference Optimization (DPO) has proven to be an effective solution for mitigating hallucination in Multimodal Large Language Models (MLLMs) b

Unifying Dynamical Systems and Graph Theory to Mechanistically Understand Computation in Neural Networks

SafetyDGX agent

arXiv:2605.03598v2 Announce Type: cross Abstract: Understanding how biological and artificial neural networks implement computation from connectivity is a central problem in neuroscience and machine l

UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings

SafetyDGX agent

arXiv:2505.11815v2 Announce Type: replace Abstract: Current vision-language models have been explored for multi-modal embedding tasks like information retrieval. However, they face significant challen

Using Common Random Numbers for Simulation-based Planning with Rollouts

SafetyDGX agent

arXiv:2605.04732v1 Announce Type: new Abstract: Simulation-based planning with rollouts is a widely-deployed technique for decision making in stochastic environments. The primary instrument of simulat

Variance Matters: Improving Domain Adaptation via Stratified Sampling

SafetyDGX agent

arXiv:2512.05226v2 Announce Type: replace Abstract: Domain shift remains a key challenge in deploying machine learning models to the real world. Unsupervised domain adaptation (UDA) aims to address th

What happens when AIs become smarter than us? Why would they keep humans around if given the choice? Our new paper argues that only trying t…

SafetyDGX agent

What happens when AIs become smarter than us? Why would they keep humans around if given the choice? Our new paper argues that only trying to control AIs is a limited strategy, and that a stable, mutu

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning

SafetyDGX agent

arXiv:2605.05172v1 Announce Type: new Abstract: Behavior Cloning (BC) has emerged as a highly effective paradigm for robot learning. However, BC lacks a self-guided mechanism for online improvement af

Why Expert Alignment Is Hard: Evidence from Subjective Evaluation

SafetyDGX agent

arXiv:2605.04972v1 Announce Type: new Abstract: Aligning large language models with expert judgment is especially difficult in subjective evaluation tasks, where experts may disagree, rely on tacit cr

Worth remembering when @janleike quit over OpenAI safety concerns.

SafetyDGX agent

Jan Leike, OpenAI's safety and alignment lead, resigned in May 2024, citing concerns about the company's commitment to safety practices and the deprioritization of safety work relative to product deve

Wow this paper has been “just published” many times for nearly a year. I have called out at least two other “influencers” for the same thing…

SafetyDGX agent

Wow this paper has been “just published” many times for nearly a year. I have called out at least two other “influencers” for the same thing on the same paper, including another one earlier this week.

yep, really embarassing

SafetyDGX agent

yep, really embarassing One of the richest VCs, Marc Andreessen, has shown again that he thinks generative AI works perfectly well if you use the right prompts, thus ignoring the statistical nature of

← Previous
1…160161162163164…212
Next →