AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,812 results
21 May 2026

Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving

SafetyDGX agent

arXiv:2605.21139v1 Announce Type: new Abstract: Current end-to-end autonomous driving models are fundamentally constrained by the behavioral cloning ceiling of imitation learning. While reinforcement

Distributed Direct Preference Optimization

SafetyDGX agent

arXiv:2605.20696v1 Announce Type: new Abstract: Preference-based reinforcement learning (RL) is a key paradigm for aligning policies with human judgments, yet its theoretical behavior in distributed s

Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression

SafetyDGX agent

arXiv:2605.20740v1 Announce Type: cross Abstract: Large language models can predict real-valued quantities from heterogeneous inputs such as text, code, and molecular strings, but most training object


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning

SafetyDGX agent

arXiv:2605.20730v1 Announce Type: new Abstract: In-context learning (ICL) allows large language models (LLMs) to adapt to new tasks through demonstrations, yet it suffers from escalating inference cos

Divide-Prompt-Refine: a Training-Free, Structure-Aware Framework for Biomedical Abstract Generation

SafetyDGX agent

arXiv:2605.20628v1 Announce Type: new Abstract: Biomedical abstracts play a critical role in downstream NLP applications, such as information retrieval, biocuration, and biomedical knowledge discovery

Do agents listen to you… or themselves? While evaling subagent behavior in deep agent systems, we noticed an interesting quirk in our agents…

SafetyDGX agent

Do agents listen to you… or themselves? While evaling subagent behavior in deep agent systems, we noticed an interesting quirk in our agents' alignment with hand-written system prompts vs. the instruc

Do No Harm? Hallucination and Actor-Level Abuse in Web-Deployed Medical Large Language Models

SafetyDGX agent

arXiv:2605.20591v1 Announce Type: new Abstract: Medical large language models (LLMs), including custom medical GPTs (MedGPTs) and open-source models, are increasingly deployed on web platforms to prov

Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards

SafetyDGX agent

arXiv:2605.21180v1 Announce Type: new Abstract: Large language models show strong potential for automated code generation, but lack guarantees for correctness, quality, safety, and domain-specific con

Enhancing Speech Large Language Models through Reinforced Behavior Alignment

SafetyDGX agent

arXiv:2509.03526v2 Announce Type: replace Abstract: The recent advancements of Large Language Models (LLMs) have spurred considerable research interest in extending their linguistic capabilities beyon

EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models

SafetyDGX agent

arXiv:2510.05942v3 Announce Type: replace Abstract: We present EvalMORAAL, a transparent chain-of-thought (CoT) framework that uses two scoring methods (log-probabilities and direct ratings) plus a mo

extit{Stochastic} MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent

SafetyDGX agent

arXiv:2605.21282v1 Announce Type: new Abstract: Online off-policy reinforcement learning (RL) is shaped by two coupled choices: the policy class and the update rule. Gaussian policies are fast and hav

FEAT: A Linear-Complexity Foundation Model for Extremely Large Structured Data

SafetyDGX agent

arXiv:2603.16513v3 Announce Type: replace Abstract: Structured data is widely used in domains such as healthcare, finance, and scientific data management. Recent studies on structured data foundation

FineVision: Open Data Is All You Need

SafetyDGX agent

arXiv:2510.17269v2 Announce Type: replace Abstract: The advancement of vision-language models (VLMs) is hampered by a fragmented landscape of inconsistent and contaminated public datasets. We introduc

FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching

SafetyDGX agent

arXiv:2605.20910v1 Announce Type: new Abstract: Extending the generation horizon of video diffusion models to long sequences remains a long-standing and important challenge. Existing training-free app

GAMR: Geometric-Aware Manifold Regularization with Virtual Outlier Synthesis for Learning with Noisy Labels

SafetyDGX agent

arXiv:2605.20727v1 Announce Type: new Abstract: Deep neural networks (DNNs) experience significant performance degradation when processing noisy labels, primarily due to overfitting on mislabeled data

Graph Transductive Sharpening: Leveraging Unlabeled Predictions in Node Classification

SafetyDGX agent

arXiv:2605.20248v1 Announce Type: new Abstract: In the transductive setting, where the full graph is observed but node labels are only partially available, progress in semi-supervised node classificat

GraphDiffMed: Knowledge-Constrained Differential Attention with Pharmacological Graph Priors for Medication Recommendation

SafetyDGX agent

arXiv:2605.20188v1 Announce Type: new Abstract: Recommending safe and effective medication combinations from electronic health records (EHRs) is a core clinical AI problem, yet it remains difficult be

GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents

SafetyDGX agent

arXiv:2605.20246v1 Announce Type: new Abstract: Recently, vision-language model (VLM) agents have shown promising progress in open-world tasks, where successful task completion often requires multiple

Here's how the corruption works: Thursday: RJ Reynolds donates $5M to Trump Saturday: Trump invites RJR execs to Mar a Lago; execs ask to lo…

SafetyDGX agent

Here's how the corruption works: Thursday: RJ Reynolds donates $5M to Trump Saturday: Trump invites RJR execs to Mar a Lago; execs ask to loosen regs on flavored vapes; Trump calls up RFK Jr. and tell

HORST: Composing Optimizer Geometries for Sparse Transformer Training

SafetyDGX agent

arXiv:2605.21104v1 Announce Type: new Abstract: Sparsifying transformers remains a fundamental challenge, as standard optimizers fail to simultaneously encourage sparsity and maintain training stabili

I don't think anyone has a good intuitive sense about what this means, and that failure of imagination is a generally bad thing for planning…

SafetyDGX agent

I don't think anyone has a good intuitive sense about what this means, and that failure of imagination is a generally bad thing for planning, investment, and policy. I also don't have an easy solution

I was happy to join the @ScienceBoard_UN podcast to discuss deception among frontier AI models and the need to manage its potential global i…

SafetyDGX agent

I was happy to join the @ScienceBoard_UN podcast to discuss deception among frontier AI models and the need to manage its potential global impacts and risks. 🎙️What happens when AI learns to lie? Scie

if you own a tesla, you are the product

SafetyDGX agent

if you own a tesla, you are the product Astonishing potential, especially if one takes all the galaxies into account and factors them into the calculation. 😃 And this comes from a man who removed Auto

Improved convergence rate of kNN graph Laplacians: differentiable self-tuned affinity

SafetyDGX agent

arXiv:2410.23212v2 Announce Type: replace-cross Abstract: In graph-based data analysis, k-nearest neighbor (kNN) graphs are widely used due to their adaptivity to local data densities. Allowing weight

Inference Time Policy Optimization for Offline RL with Differentiable World Models

SafetyDGX agent

arXiv:2603.22430v2 Announce Type: replace Abstract: Offline Reinforcement Learning (RL) learns optimal policies from fixed datasets, training a policy once and deploying it at inference time without f

It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs

SafetyDGX agent

arXiv:2605.20258v1 Announce Type: new Abstract: Contextual Integrity (CI) defines privacy not merely as keeping information hidden, but as governing information flows according to the norms of a given

It’s crazy that OpenAI and xAI were built partly as ways to mitigate @demishassabis’s power and how both are more reckless and less focused …

SafetyDGX agent

It’s crazy that OpenAI and xAI were built partly as ways to mitigate @demishassabis’s power and how both are more reckless and less focused on the good part, viz AI for science. And then Anthropic was

LASH: Adaptive Semantic Hybridization for Black-Box Jailbreaking of Large Language Models

SafetyDGX agent

arXiv:2605.21362v1 Announce Type: new Abstract: Jailbreak attacks expose a persistent gap between the intended safety behavior of aligned large language models and their behavior under adversarial pro

Latent Geometry as a Structural Monitor: Eigenspace Alignment for Anomaly Detection in Anonymity Networks

SafetyDGX agent

arXiv:2605.20391v1 Announce Type: cross Abstract: Traditional anomaly detection marks events when measured signals cross predefined thresholds. This captures the moment of transition but not the struc

Learning Robust Dexterous In-Hand Manipulation from Joint Sensors with Proprioceptive Transformer

SafetyDGX agent

arXiv:2605.21330v1 Announce Type: new Abstract: In-hand object manipulation is a fundamental yet challenging capability for dexterous robots. Despite significant progress in dexterous manipulation, ex

Learning to Think in Physics: Breaking Shortcut Learning in Scientific Diffusion via Representation Alignment

SafetyDGX agent

arXiv:2605.20780v1 Announce Type: cross Abstract: Physics-informed diffusion models typically enforce PDE constraints only on final outputs, leaving intermediate representations unconstrained and pron

Letting Trajectories Spread: Quality-Preserving Control for Diverse Flow Matching

SafetyDGX agent

arXiv:2510.09060v2 Announce Type: replace-cross Abstract: Flow-based text-to-image models follow deterministic trajectories, making it costly to explore diverse modes under limited sampling budgets. E

Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models

SafetyDGX agent

arXiv:2605.21123v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) is successful for alignment in LLMs but still faces challenges in text-to-image generation. Existing studies are co

Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex

SafetyDGX agent

arXiv:2605.06139v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a standard approach for large language models (LLMs) post-training to incentivize r

LLM Pretraining Shapes a Generalizable Manifold: Insights into Cross-Modal Transfer to Time Series

SafetyDGX agent

arXiv:2605.20449v1 Announce Type: new Abstract: Can language-pretrained transformers become effective time-series forecasters, and why? In this paper, we show that cross-modal transfer arises because

longer comment on yesterday’s stories

SafetyDGX agent

longer comment on yesterday’s stories the nuance between yesterday’s stories that X influencers won’t tell you: https://open.substack.com/pub/garymarcus/p/checking-the-math-behind-openai-and?r=8tdk6&u

Lost in Fog: Sensor Perturbations Expose Reasoning Fragility in Driving VLAs

SafetyDGX agent

arXiv:2605.21446v1 Announce Type: new Abstract: Interpretable autonomous driving planners depend not only on generating explanations, but also on those explanations remaining reliable under real-world

Mahjax: A GPU-Accelerated Mahjong Simulator for Reinforcement Learning in JAX

SafetyDGX agent

arXiv:2605.20577v1 Announce Type: cross Abstract: Riichi Mahjong is a multi-player, imperfect-information game characterized by stochasticity and high-dimensional state spaces. These attributes presen

Mechanistic Interpretability for Learning Assurance of a Vision-Based Landing System

SafetyDGX agent

arXiv:2605.20607v1 Announce Type: cross Abstract: EASA's learning-assurance guidance requires data-driven aviation systems to build and monitor their own situation representation, yet for neural netwo

Mind the Sim-to-Real Gap & Think Like a Scientist

SafetyDGX agent

arXiv:2605.21458v1 Announce Type: cross Abstract: Suppose a planner has a pre-trained simulator of a sequential decision problem and the option to run real experiments in the field. The simulator is c

Mitigating Label Bias with Interpretable Rubric Embeddings

SafetyDGX agent

arXiv:2605.21455v1 Announce Type: new Abstract: Statistical decision algorithms are increasingly deployed in domains where ground-truth labels are hard to obtain, such as hiring, university admissions

Mobile UMI: Cross-View Diffusion Policy with Decoupled Kinematics for Mobile Manipulation

SafetyDGX agent

arXiv:2605.20894v1 Announce Type: new Abstract: Mobile imitation learning on portable demonstration interfaces faces two coupled bottlenecks: locomotion-contaminated action labels and inference-induce

Multi-Agent Reinforcement Learning for Safe Autonomous Driving Under Pedestrian Behavioral Uncertainty

SafetyDGX agent

arXiv:2605.20255v1 Announce Type: new Abstract: Simulation-based testing of self-driving cars (SDCs) typically relies on scripted or simplified pedestrian models that do not capture the heterogeneity

Multi-Head Attention as Ensemble Nadaraya-Watson Estimation: Variance Reduction, Decorrelation, and Optimal Head Diversity

SafetyDGX agent

arXiv:2605.20271v1 Announce Type: cross Abstract: We develop a rigorous statistical theory of multi-head attention (MHA) as an ensemble of Nadaraya-Watson (NW) kernel regression estimators. Building o

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards

SafetyDGX agent

arXiv:2605.20865v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) plays a pivotal role in improving the reasoning ability of large language models. However, widely

Multimodal LLMs under Pairwise Modalities

SafetyDGX agent

arXiv:2605.21059v1 Announce Type: new Abstract: Despite the impressive results achieved by multimodal large language models (MLLMs), their training typically relies on jointly curated multimodal data,

NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control

SafetyDGX agent

arXiv:2605.20209v1 Announce Type: cross Abstract: Achieving precise, versatile whole-body character control in physics-based animation remains challenging. Recent diffusion-based policies generate ric

Neural Collapse by Design: Learning Class Prototypes on the Hypersphere

SafetyDGX agent

arXiv:2605.20302v1 Announce Type: cross Abstract: Supervised classification has a theoretical optimum, Neural Collapse (NC), yet neither of its two dominant paradigms reaches it in practice. Cross ent

nice measure of how insane SpaceX’s claims are

SafetyDGX agent

nice measure of how insane SpaceX’s claims are I don't offer investment advice, but when I see that SpaceX claims that their total addressable market is the size of the *complete* US economy (with eve

nor do we know how the (new) model works nor how it does on anything else nor how it was trained. scientists wait for facts; cheerleaders (o…

SafetyDGX agent

nor do we know how the (new) model works nor how it does on anything else nor how it was trained. scientists wait for facts; cheerleaders (over and over) rush to judgments that have often been wrong.

note that @jeffBezos is not betting on chatbots, but rather focusing on domain-specific tools for engineering, which is also something I wou…

SafetyDGX agent

note that @jeffBezos is not betting on chatbots, but rather focusing on domain-specific tools for engineering, which is also something I would encourage. the financial story for what he is doing may b

OcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generation

SafetyDGX agent

arXiv:2605.21343v1 Announce Type: new Abstract: Recent layout-to-image models have achieved remarkable progress in spatial controllability. However, they still struggle with inter-object occlusion. Wh

👀OpenAI Q1 operating margin was -122% even when excluding SBC and other items.

SafetyDGX agent

OpenAI reported a negative 122% operating margin in Q1, indicating the company spent significantly more on operations than it generated in revenue, even after excluding stock-based compensation and ot

our company, which lost money last quarter, is going to take a big slice of $28.5 trillion market that we project will exist, primarily in a…

SafetyDGX agent

our company, which lost money last quarter, is going to take a big slice of 28.5 trillion market that we project will exist, primarily in a field in which we have almost no current presence, which is

Parallel LLM Reasoning for Bias-Resilient, Robust Conceptual Abstraction

SafetyDGX agent

arXiv:2605.20194v1 Announce Type: new Abstract: Large language models (LLMs) have been increasingly used to analyze text. However, they are often plagued with contextual reasoning limitations when ana

Pareto-Enhanced Portrait Generation: Vision-Aligned Text Supervision for Alignment, Realism, and Aesthetics

SafetyDGX agent

arXiv:2605.20640v1 Announce Type: new Abstract: Text-to-image diffusion models often face a severe trilemma in human portrait generation: text-image alignment, photorealism, and human-perceived aesthe

PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction

SafetyDGX agent

arXiv:2605.21414v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation by leveraging large pretrained vision-languag

Preference-aware Influence-function-based Data Selection Method for Efficient Fine-Tuning

SafetyDGX agent

arXiv:2605.21422v1 Announce Type: new Abstract: As LLMs continue to scale, improving training efficiency increasingly depends on using data more effectively. Data selection addresses this problem by a

PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment

SafetyDGX agent

arXiv:2605.21225v1 Announce Type: new Abstract: We address the problem of making a pre-trained reinforcement learning (RL) policy safety-aware by incorporating cost constraints without retraining it f

Pretty profound for an ipo filing: “For the entirety of its existence, human civilization has lived on a single celestial body: Earth. The c…

SafetyDGX agent

Pretty profound for an ipo filing: “For the entirety of its existence, human civilization has lived on a single celestial body: Earth. The current paradigm, in which human civilization is confined to

← Previous
1…126127128129130…214
Next →