AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,809 results
20 May 2026

RoHIL: Robust Human-in-the-Loop Robotic Reinforcement Learning Against Illumination Variations

SafetyDGX agent

arXiv:2605.19924v1 Announce Type: new Abstract: Human-in-the-loop reinforcement learning systems achieve near-perfect success on the workstation where they are trained, but collapse when the same robo

RoVLA: Multi-Consistency Constraints for Robust Vision-Language-Action Models

SafetyDGX agent

arXiv:2605.19678v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong performance on embodied manipulation, yet they remain brittle under visual observation changes, pa

Safe Continual Reinforcement Learning under Nonstationarity via Adaptive Safety Constraints

SafetyDGX agent

arXiv:2605.18842v1 Announce Type: new Abstract: Safe reinforcement learning in nonstationary environments requires safety mechanisms that adapt as environmental conditions change. Standard safe reinfo


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SafeAlign-VLA: A Negative-Enhanced Safe Alignment Framework for Risk-Aware Autonomous Driving

SafetyDGX agent

arXiv:2605.19524v1 Announce Type: cross Abstract: End-to-end autonomous driving systems excel in common scenarios but struggle with safety-critical long-tail cases. Vision-Language-Action (VLA) models

SAGE: Scalable Automatic Gating Ensemble for Confident Negative Harvesting in Fraud Detection

SafetyDGX agent

arXiv:2605.20157v1 Announce Type: new Abstract: Music streaming fraud, where bad actors artificially inflate stream counts to manipulate chart rankings and royalty payments, poses a significant threat

SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs

SafetyDGX agent

arXiv:2605.18864v1 Announce Type: cross Abstract: Recent studies observe that reinforcement learning with verifiable rewards (RLVR) reliably improves pass@1 on reasoning tasks, yet often fails to yiel

Sampling-Based Safe Reinforcement Learning

SafetyDGX agent

arXiv:2605.19469v1 Announce Type: cross Abstract: Safe exploration remains a fundamental challenge in reinforcement learning (RL), limiting the deployment of RL agents in the real world. We propose Sa

Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling

SafetyDGX agent

arXiv:2503.06310v4 Announce Type: replace Abstract: Generating coherent long-form video sequences from discrete text prompts remains challenging due to difficulties in maintaining temporal coherence,

SceneCode: Executable World Programs for Editable Indoor Scenes with Articulated Objects

SafetyDGX agent

arXiv:2605.19587v1 Announce Type: new Abstract: Indoor scene synthesis underpins embodied AI, robotic manipulation, and simulation-based policy evaluation, where a useful scene must specify not only w

Self-Creative Text-to-Object Generation using Semantic-Aware Spatial Weighting

SafetyDGX agent

arXiv:2605.19554v1 Announce Type: new Abstract: Instilling creativity in text-to-image (T2I) generation presents a significant challenge, as it requires synthesized images to exhibit not only visual n

Set-Valued Policy Learning

SafetyDGX agent

arXiv:2605.19830v1 Announce Type: new Abstract: Conventional treatment policies map patient covariates to a single recommended intervention in order to maximize expected clinical outcomes. Although a

SimGym: A Framework for A/B Test Simulation in E-Commerce with Traffic-Grounded VLM Agents

SafetyDGX agent

arXiv:2605.19219v1 Announce Type: new Abstract: A/B testing remains the gold standard for evaluating modifications to e-commerce storefronts, yet it diverts traffic, requires weeks to reach statistica

Smooth Partial Lotteries for Stable Randomized Selection

SafetyDGX agent

arXiv:2605.20069v1 Announce Type: new Abstract: Competitive selection processes, from scientific funding to admissions and hiring, use evaluations to score candidates, and eventually choose a subset o

Spatially Prompted Visual Trajectory Prediction for Egocentric Manipulation

SafetyDGX agent

arXiv:2605.20085v1 Announce Type: new Abstract: Robotic manipulation is often specified through language instructions or task identifiers, yet cluttered environments with similar objects are better ha

Stitched Value Model for Diffusion Alignment

SafetyDGX agent

arXiv:2605.19804v1 Announce Type: cross Abstract: For practical use, diffusion- or flow-based generative models must be aligned with task-specific rewards, such as prompt fidelity or aesthetic prefere

Structural Energy Guidance for View-Consistent Text-to-3D Generation

SafetyDGX agent

arXiv:2605.19876v1 Announce Type: new Abstract: Text-to-3D generation based on diffusion models often suffers from the Janus problem, leading to inconsistent geometry across viewpoints. This work iden

StruMPL: Multi-task Dense Regression under Disjoint Partial Supervision and MNAR Labels

SafetyDGX agent

arXiv:2605.19931v1 Announce Type: cross Abstract: Estimating forest aboveground biomass (AGB) from Earth observation combines two structurally incompatible label sources: spaceborne lidar provides can

Surviving the Unseen: Predictive Defense for Novel Multi-Turn Multimodal Attacks

SafetyDGX agent

arXiv:2605.18988v1 Announce Type: cross Abstract: The expansion of Multimodal Large Language Models (MLLMs) and their integration into autonomous agentic workflows has introduced a non-stationary atta

Swimming with Whales: Analysis of Power Imbalances in Stake-Weighted Governance

SafetyDGX agent

arXiv:2605.19264v1 Announce Type: new Abstract: Voting methods weighted by stakes are the fundamental governance paradigm in Proof-of-Stake (PoS) blockchains. Such a paradigm is known to be prone to p

Symmetry in the Wild: The Role of Equivariance in Neural Fluid Surrogates

SafetyDGX agent

arXiv:2605.18816v1 Announce Type: cross Abstract: Neural surrogates enable orders-of-magnitude acceleration of computational fluid dynamics (CFD) simulations, with the potential to transform engineeri

TEA-Time: Transporting Effects Across Time

SafetyDGX agent

arXiv:2603.07018v2 Announce Type: replace-cross Abstract: Treatment effects estimated from a randomized controlled trial are local not only to the study population but also to the time at which the tr

TEMPO: Temporal Enforcement via Mode-Separated Policy Optimization for Trustworthy LLM Backtesting

SafetyDGX agent

arXiv:2605.18843v1 Announce Type: new Abstract: Backtesting large language models on historical events requires reasoning exclusively from information available before a specified cutoff date. Yet mod

Text-to-SPARQL Generation with Reinforcement Learning: A GRPO-based Approach on DBLP

SafetyDGX agent

arXiv:2605.20066v1 Announce Type: new Abstract: Knowledge graph question answering seeks to translate natural language questions into executable queries over knowledge graphs, but existing approaches

The Accessibility Capability Boundary: Operational Limits and Expansion Potential of AI-Generated Browser-Native Accessibility Systems

SafetyDGX agent

arXiv:2605.19638v1 Announce Type: cross Abstract: As large language models (LLMs) demonstrate increasing competence in synthesizing functional user interfaces, a fundamental question emerges in access

The consensus of analysts is that AI capital investments will rise by 20% a year for five years while revenues are expected to grow 15% resu…

SafetyDGX agent

The consensus of analysts is that AI capital investments will rise by 20% a year for five years while revenues are expected to grow 15% resulting in negative returns. 'The AI boom will become a story

the crazy part is that people are “clowning” me without knowing anything about the training or whether anything else other than scaled chang…

SafetyDGX agent

the crazy part is that people are “clowning” me without knowing anything about the training or whether anything else other than scaled changed or how the model does on anything else. (or what it costs

The Uncomfortable Truth About AI “Reasoning” | World Science Festival https://youtu.be/iFYF_e1GSGI?si=pFNjvhp48TMKp-PE via @YouTube @GaryMar…

SafetyDGX agent

This video from the World Science Festival, shared by cognitive scientist Gary Marcus, examines the gap between AI systems' claimed reasoning capabilities and their actual mechanisms, likely arguing t

think you know @garymarcus because read a few of his tweets? try watching this, to see the real deal, and why, for example, the US Senate in…

SafetyDGX agent

think you know @garymarcus because read a few of his tweets? try watching this, to see the real deal, and why, for example, the US Senate invited him to testify. Superb conversation between @bgreene a

ThoughtTrace: Understanding User Thoughts in Real-World LLM Interactions

SafetyDGX agent

arXiv:2605.20087v1 Announce Type: cross Abstract: Conversational AI has now reached billions of users, yet existing datasets capture only what people say, not what they think. We introduce ThoughtTrac

Toward an AI-Powered Computational Testbed for Workforce Policy

SafetyDGX agent

arXiv:2605.19064v1 Announce Type: cross Abstract: Workforce transformations are difficult to forecast and costly to mismanage. In particular, the integration of artificial intelligence into knowledge

Towards Distillation Guarantees under Algorithmic Alignment for Combinatorial Optimization

SafetyDGX agent

arXiv:2605.20074v1 Announce Type: new Abstract: Distillation transfers knowledge from a large model trained on broad data to a smaller, more efficient model suitable for deployment. In structured pred

Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On

SafetyDGX agent

arXiv:2605.19035v1 Announce Type: new Abstract: The rapid advancement of Large Language Models has given rise to autonomous LLM-based agents capable of complex reasoning and execution. As these agents

TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agents

SafetyDGX agent

arXiv:2602.11767v3 Announce Type: replace Abstract: Advances in large language models (LLMs) are driving a shift toward using reinforcement learning (RL) to train agents from iterative, multi-turn int

two trillion dollars to build “pathologically dishonest” AI

SafetyDGX agent

two trillion dollars to build “pathologically dishonest” AI I think @RyanPGreenblatt's recent post summarizes the vibe of this behavior well. Not the sort of thing that would be acceptable for a human

Universal Skeleton Understanding via Differentiable Rendering and MLLMs

SafetyDGX agent

arXiv:2603.18003v4 Announce Type: replace Abstract: Multimodal large language models (MLLMs) exhibit strong visual-language reasoning, yet cannot process structured, non-visual data such as human skel

When Critics Disagree: Adaptive Reward Poisoning Attacks in RIS-Aided Wireless Control System

SafetyDGX agent

arXiv:2605.20037v1 Announce Type: cross Abstract: Reward-poisoning attacks present a significant risk to learning-based wireless control systems. Given this, we propose a Disagreement-Guided Reward Po

When Preference Labels Fall Short: Aligning Diffusion Models from Real Data

SafetyDGX agent

arXiv:2605.19839v1 Announce Type: new Abstract: Preference alignment aims to guide generative models by learning from comparisons between preferred and non-preferred samples. In practice, most existin

When Tabular Foundation Models Meet Strategic Tabular Data: A Prior Alignment Approach

SafetyDGX agent

arXiv:2605.19662v1 Announce Type: new Abstract: Tabular foundation models based on pretrained prior-data fitted networks~(PFNs) have shown strong generalization on diverse tabular tasks, but they are

When to Stop Reusing: Dynamic Gradient Gating for Sample-Efficient RLVR

SafetyDGX agent

arXiv:2605.19425v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become the dominant paradigm for advanced reasoning in Large Language Models (LLMs), but rol

Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making

SafetyDGX agent

arXiv:2602.07008v2 Announce Type: replace Abstract: Reliable models should not only predict correctly, but also justify decisions with acceptable evidence. Yet conventional supervised learning typical

Worst-Group Equalized Odds Regularization for Multi-Attribute Fair Medical Image Classification

SafetyDGX agent

arXiv:2605.19214v1 Announce Type: cross Abstract: Diagnostic performance in medical AI varies systematically across demographic groups, yet subgroup AUC can mask clinically important disparities. At a

19 May 2026

A Fourier perspective on the learning dynamics of neural networks: from sample complexities to mechanistic insights

SafetyDGX agent

arXiv:2605.16913v1 Announce Type: cross Abstract: Neural networks trained with gradient-based methods exhibit a strong simplicity bias: they learn simpler statistical features of their data before mov

A Simplex Witness Certificate for Constant Collapse in Variational Autoencoders

SafetyDGX agent

arXiv:2605.18224v1 Announce Type: cross Abstract: This note studies exact constant collapse in variational autoencoders, where the encoder mean becomes independent of the input. The goal is to make th

A study done on 361,645 job applications in almost 30 countries over the last 40 years discovered the hiring bias in society is actually aga…

SafetyDGX agent

A large-scale meta-analysis examining over 361,000 job applications across nearly 30 countries spanning 40 years found that hiring discrimination based on protected characteristics (such as race, gend

A Visual Reinforcement Learning-Based Separate Primitive Policy for Peg-in-Hole Tasks

SafetyDGX agent

arXiv:2504.14820v2 Announce Type: replace Abstract: For peg-in-hole tasks, humans rely on binocular visual perception to locate the peg above the hole surface and then proceed with insertion. This pap

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications

SafetyDGX agent

arXiv:2605.17413v1 Announce Type: cross Abstract: Safety-aligned language models often refuse cybersecurity requests whose wording resembles misuse, even when the task is authorized and defensive. Thi

Actionable World Representation

SafetyDGX agent

arXiv:2605.18743v1 Announce Type: new Abstract: Inspired by the emergent behaviors in large language models that generalized human intelligence, the research community is pursuing similar emergent cap

Activation Steering with a Feedback Controller

SafetyDGX agent

arXiv:2510.04309v3 Announce Type: replace Abstract: Controlling the behaviors of large language models (LLM) is fundamental to their safety alignment and reliable deployment. However, existing steerin

Adaptive Control in Autonomous Driving via Real-Time Recurrent RL

SafetyDGX agent

arXiv:2602.02236v4 Announce Type: replace-cross Abstract: We study online fine-tuning of pretrained control policies for autonomous driving using Real-Time Recurrent Reinforcement Learning (RTRRL), a

Adaptive Experimentation for Censored Survival Outcomes

SafetyDGX agent

arXiv:2605.18459v1 Announce Type: new Abstract: Adaptive experimentation enables efficient estimation of causal effects, but existing methods are not designed for survival data with censoring, where e

Adaptive Generate-Rank-Verify: Inference-Time Search with Costly Verification

SafetyDGX agent

arXiv:2605.17609v1 Announce Type: new Abstract: Many inference-time language-model pipelines combine a cheap reward signal with an expensive verifier, such as exact answer checking in mathematical rea

Adversarial Fragility and Language Vulnerability in Clinical AI: A Systematic Audit of Diagnostic Collapse Under Imperceptible Perturbations and Cross-Lingual Drift in Low-Resource Healthcare Settings

SafetyDGX agent

arXiv:2605.16993v1 Announce Type: cross Abstract: Current clinical artificial intelligence (AI) systems are evaluated almost exclusively on clean, standardised, English-language inputs, conditions tha

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment

SafetyDGX agent

arXiv:2605.17517v1 Announce Type: new Abstract: Recent advances in Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation. However, the visual represe

Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces

SafetyDGX agent

arXiv:2605.17698v1 Announce Type: new Abstract: The deployment of Large Language Models (LLMs) as autonomous economic agents introduces systemic risks that extend beyond individual capability failures

AI Agents May Always Fall for Prompt Injections

SafetyDGX agent

arXiv:2605.17634v1 Announce Type: cross Abstract: Prompt injection is the most critical vulnerability in deployed AI agents. Despite recent progress, we show that the prevailing defense paradigm (data

AI Alignment Breaks at the Edge

SafetyDGX agent

arXiv:2602.20042v2 Announce Type: replace Abstract: General Alignment has improved average-case helpfulness and safety, but current alignment practice still rewards confident, single-turn responses. T

AI of the People, by the People, for the People: A Social Choice Approach to Collective Control of Artificial Intelligence

SafetyDGX agent

arXiv:2605.16291v1 Announce Type: cross Abstract: With the growing adoption of AI systems, reasoning about how society can exert control over AI becomes an increasingly urgent problem. Existing work o

AIM: Adversarial Information Masking for Faithfulness Evaluation of Saliency Maps

SafetyDGX agent

arXiv:2605.16905v1 Announce Type: cross Abstract: Post-hoc saliency methods are widely used to interpret deep neural networks, but their faithfulness is difficult to evaluate reliably. Existing evalua

Algorithmic Cultivation: How Social Media Feeds Shape User Language

SafetyDGX agent

arXiv:2605.17010v1 Announce Type: cross Abstract: Algorithmic feeds have become primary environments for encountering information online, yet while they shape what people see, less is known about how

Alignment and Safety of Diffusion Models via Reinforcement Learning and Reward Modeling: A Survey

SafetyDGX agent

arXiv:2505.17352v2 Announce Type: replace Abstract: Diffusion models have become a central paradigm for image and multimodal generation, yet their deployment raises persistent questions about alignmen

← Previous
1…130131132133134…214
Next →