AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,813 results
26 May 2026

ECHO: Terminal Agents Learn World Models for Free

SafetyDGX agent

arXiv:2605.24517v1 Announce Type: cross Abstract: CLI agents are the closest thing language models have to an embodied setting: the model emits commands, the terminal executes them, and the returned s

ECo-MoE: Embodiment-Conditioned Mixture of Experts Increases the Evolvability of Robots

SafetyDGX agent

arXiv:2605.24225v1 Announce Type: new Abstract: In this paper, we introduce a model of evolution and learning in robots that co-optimizes a distribution of latent design vectors (genotypes) and a mixt

EMA: Effort Metric Attention for Anatomical Effort-Guided Human Motion Diffusion

SafetyDGX agent

arXiv:2605.24566v1 Announce Type: cross Abstract: Human motion diffusion models can synthesize action sequences from text, but controlling motion intensity remains challenging. Existing approaches rel


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Emergent Analogical Reasoning in Transformers

SafetyDGX agent

arXiv:2602.01992v4 Announce Type: replace Abstract: Analogy is a central faculty of human intelligence, enabling abstract patterns discovered in one domain to be applied to another. Despite its centra

Energy Shields for Fairness

SafetyDGX agent

arXiv:2605.24926v1 Announce Type: new Abstract: Runtime fairness is not a one-time constraint but a dynamic property evaluated over a sequence of decisions. To ensure fairness at runtime, it is necess

EPPC-OASIS: Ontology-Aware Adaptation and Structured Inference Refinement for Electronic Patient-Provider Communication Mining in Secure Messages

SafetyDGX agent

arXiv:2605.24172v1 Announce Type: new Abstract: Secure patient-provider messages contain clinically important communication behaviors that are difficult to characterize manually at scale. The Electron

Eureka: Intelligent Feature Engineering for Enterprise AI Cloud Resource Demand Prediction

SafetyDGX agent

arXiv:2605.25297v1 Announce Type: cross Abstract: Effective features are crucial for predictive model performance, but creating them often requires domain expertise, limiting scalability across applic

Europe’s sovereign AI moment arrives in Heilbronn next week. At TECH by Handelsblatt 2026, Cohere CEO and Co-founder, @aidangomez, will join…

SafetyDGX agent

Europe’s sovereign AI moment arrives in Heilbronn next week. At TECH by Handelsblatt 2026, Cohere CEO and Co-founder, @aidangomez, will join leaders across business, policy, and industry to discuss ho

even by the standards of the last few years, we are in some truly insane territory here.

SafetyDGX agent

even by the standards of the last few years, we are in some truly insane territory here. “Unserious, empty, hallucinatory, and borderline dishonest” - the prospectus for a company that the S&P 500 is

Evidence-Linked Radiology Reporting: A Human-Supervised Reference Architecture for Structured Imaging Intelligence

SafetyDGX agent

arXiv:2605.25120v1 Announce Type: cross Abstract: Radiology reports remain the primary mechanism by which imaging findings are communicated to clinical teams. However, much of the structured informati

Evolutionary Enhanced Multi-Agent Reinforcement Learning for Cooperative Air Combat

SafetyDGX agent

arXiv:2605.25091v1 Announce Type: new Abstract: As modern air combat evolves toward beyond-visual-range (BVR) multi-aircraft cooperative engagements, autonomous decision-making for unmanned combat aer

Evolving Robustness--Exploration Trade-off in Online Reinforcement Learning via Quantile Bayesian Risk MDPs

SafetyDGX agent

arXiv:2605.24345v1 Announce Type: new Abstract: In online reinforcement learning, data scarcity creates epistemic uncertainty that makes robustness important early in learning, whereas sufficient expl

Extracting Training Data from Diffusion Language Models via Infilling

SafetyDGX agent

arXiv:2605.24173v1 Announce Type: cross Abstract: Memorization in large language models has been studied almost exclusively through prefix-conditioned extraction, a natural choice for autoregressive m

Extreme Region Policy Distillation

SafetyDGX agent

arXiv:2605.25582v1 Announce Type: cross Abstract: Reinforcement learning for large language models faces a fundamental trade-off between sample efficiency and asymptotic performance: strictly on-polic

Factored Latent Action World Models

SafetyDGX agent

arXiv:2602.16229v2 Announce Type: replace Abstract: Learning latent actions from action-free video has emerged as a powerful paradigm for scaling up controllable world model learning. Latent actions p

FairJudge: Abstention-Aware Multimodal Judges for Fairness and Alignment Evaluation in Text-to-Image Models

SafetyDGX agent

arXiv:2510.22827v3 Announce Type: replace-cross Abstract: Evaluating text-to-image (T2I) systems requires judging not only whether an image matches a prompt, but also whether socially salient attribut

Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges

SafetyDGX agent

arXiv:2605.23970v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as automatic judges for summarization and dialogue evaluation. Prior work has documented biases such

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning

SafetyDGX agent

arXiv:2605.24286v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning is useful for monitoring language models only when the reasoning trace faithfully reflects the computation that produ

Few-Shot Neural Differentiable Simulator: Real-to-Sim Rigid-Contact Modeling

SafetyDGX agent

arXiv:2603.06218v2 Announce Type: replace Abstract: Accurate physics simulation is essential for robotic learning and control, yet analytical simulators often fail to capture complex contact dynamics,

First, do no harm: Breaking suicidogenic echo chambers in media recommendation

SafetyDGX agent

arXiv:2605.25258v1 Announce Type: cross Abstract: Recommender systems generally optimises user engagement, but this approach is dangerous in mental health contexts. When vulnerable users show signs of

Flat Minima and Generalization: Insights from Stochastic Convex Optimization

SafetyDGX agent

arXiv:2511.03548v2 Announce Type: replace Abstract: Understanding the generalization behavior of learning algorithms is a central goal of learning theory. A recently emerging explanation is that learn

From Automation to Collaboration: Human-in-the-Loop Methods for Safe and Trustworthy NLP

SafetyDGX agent

arXiv:2605.25226v1 Announce Type: new Abstract: Large language models are widely deployed in high-stakes NLP tasks, yet risks such as bias, hallucination, adversarial vulnerability and unreliable gene

From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlas

SafetyDGX agent

arXiv:2602.00491v2 Announce Type: replace Abstract: Public health reasoning requires population level inference grounded in scientific evidence, expert consensus, and safety constraints. However, it r

From Multi-Agent Systems and the Semantic Web to Agentic AI: A Unified Narrative of the Web of Agents

SafetyDGX agent

arXiv:2507.10644v4 Announce Type: replace Abstract: The Web of Agents (WoA) transforms the document-centric Web into an environment of autonomous agents acting on users' behalf, a vision newly tractab

From Reasoning to Code: GRPO Optimization for Underrepresented Languages

SafetyDGX agent

arXiv:2506.11027v3 Announce Type: replace-cross Abstract: Generating accurate and executable code using Large Language Models (LLMs) remains a significant challenge for underrepresented programming la

From Simulation to Enaction: Post-trained language models recognize and react to their own generations

SafetyDGX agent

arXiv:2605.25459v1 Announce Type: cross Abstract: Language models are pretrained as passive predictors with no incentive to model the consequences of their own outputs. Post-training changes this: a m

FusionCore: A 23-State Unscented Kalman Filter for IMU, Wheel Encoder, GPS, and Visual SLAM Fusion in ROS 2

SafetyDGX agent

arXiv:2605.25239v1 Announce Type: new Abstract: We present FusionCore, an open-source ROS 2 sensor fusion package that fuses IMU, wheel encoder odometry, GPS, and Visual SLAM pose into a single 100 Hz

GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning

SafetyDGX agent

arXiv:2603.10250v2 Announce Type: replace Abstract: A commonly used family of RL algorithms for diffusion policies conducts softmax reweighting over samples from the behavior policy, which often induc

Generative OOD-regularized Model-based Policy Optimization

SafetyDGX agent

arXiv:2605.24405v1 Announce Type: cross Abstract: We study sequential decision-making with offline reinforcement learning (RL). Traditional offline RL policies may result in out-of-distribution (OOD)

Generative Visual Code Mobile World Models

SafetyDGX agent

arXiv:2602.01576v2 Announce Type: replace-cross Abstract: Mobile Graphical User Interface (GUI) World Models (WMs) offer a promising path for improving mobile GUI agent performance at train- and infer

GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation

SafetyDGX agent

arXiv:2605.25447v1 Announce Type: new Abstract: Generating structured, editable diagrams remains a significant challenge for contemporary large language models, despite their proficiency in general-pu

GIBLy: Improving 3D Semantic Segmentation through an Architecture-Agnostic Lightweight Geometric Inductive Bias Layer

SafetyDGX agent

arXiv:2605.24243v1 Announce Type: cross Abstract: In 3D scene understanding, deep learning models rely on large models and extensive training to capture basic geometric structures that are present in

Global Convergence of Wasserstein Policy Gradient for Entropy-Regularized Reinforcement Learning

SafetyDGX agent

arXiv:2605.26078v1 Announce Type: new Abstract: Wasserstein policy gradient (WPG) is a policy optimization method for reinforcement learning (RL) that exploits the optimal-transport geometry of action

Global linear convergence of entropy-regularized softmax policy gradient beyond tabular MDPs

SafetyDGX agent

arXiv:2605.24939v1 Announce Type: new Abstract: We study the global convergence of policy gradient for infinite-horizon entropy-regularized Markov decision processes (MDPs) with continuous state and a

Grouter: Decoupling Routing from Representation for Accelerated MoE Training

SafetyDGX agent

arXiv:2603.06626v2 Announce Type: replace-cross Abstract: Traditional Mixture-of-Experts (MoE) training typically proceeds without any structural priors, effectively requiring the model to simultaneou

Grow-Prune-Freeze Networks: Adaptive & Continual Learning Technique for Olfactory Navigation

SafetyDGX agent

arXiv:2605.25170v1 Announce Type: cross Abstract: Training data for olfaction is scattered through disparate, non-standardized datasets that limit the ability to build representative world models. Olf

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models

SafetyDGX agent

arXiv:2605.25443v1 Announce Type: new Abstract: Post-training has significantly enhanced the reasoning capability of Large Reasoning Models (LRMs), especially with Reinforcement Learning (RL) like Gro

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio

SafetyDGX agent

arXiv:2605.25967v1 Announce Type: new Abstract: As policy catches up with the capabilities of generative AI, watermarking is central to content provenance efforts. Inference-time watermarks for autore

Hidden-State Privacy Has an Empty Middle

SafetyDGX agent

arXiv:2605.24042v1 Announce Type: cross Abstract: Of 1{,}536 Gaussian release covariances we tested for single-layer hidden-state privacy, zero achieve both moderate utility and moderate privacy again

Hide-and-Shill: A Reinforcement Learning Framework for Market Manipulation Detection in Symphony-a Decentralized Multi-Agent System

SafetyDGX agent

arXiv:2507.09179v3 Announce Type: replace Abstract: Decentralized finance (DeFi) has introduced a new era of permissionless financial innovation but also led to unprecedented market manipulation. With

Hide to Guide: Learning via Semantic Masking

SafetyDGX agent

arXiv:2605.25198v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a powerful paradigm for improving language models on reasoning-intensive tasks, but i

HoLoArm: Deformable Arms for Collision-Tolerant Quadrotor Flight

SafetyDGX agent

arXiv:2605.25790v1 Announce Type: new Abstract: The increasing use of drones in human-centric applications highlights the need for designs that can survive collisions and recover rapidly, minimizing r

How Much Structure Do LLMs Need? Evaluating LLMs for Bibliometric Cluster Description

SafetyDGX agent

arXiv:2605.24351v1 Announce Type: new Abstract: Large language models (LLMs) can support scientific literature synthesis, but remain prone to hallucinated references, uneven coverage, and weakly groun

How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis

SafetyDGX agent

arXiv:2605.24749v1 Announce Type: cross Abstract: Reward modeling is not only a prediction problem: in KL-regularized policy optimization, the learned reward is exponentiated to define the deployed po

How to Mitigate the Distribution Shift Problem in Robotics Control: A Robust and Adaptive Approach Based on Offline to Online Imitation Learning

SafetyDGX agent

arXiv:2605.25414v1 Announce Type: new Abstract: Distribution shift in imitation learning refers to the problem that the agent cannot plan proper actions for a state that has not been visited during th

HumanEgo: Zero-Shot Robot Learning from Minutes of Human Egocentric Videos

SafetyDGX agent

arXiv:2605.24934v1 Announce Type: cross Abstract: Human egocentric video captures rich manipulation demonstrations without any robot hardware, yet transferring these skills to robots remains challengi

HumanFlow -- Diffusion-Driven MAV Navigation Among Humans via Tightly-Coupled Motion Tracking, Forecasting, and Control

SafetyDGX agent

arXiv:2605.25685v1 Announce Type: new Abstract: Robust and accurate perception of humans in their 3D scene context is essential for integrating robots into everyday environments. Existing approaches,

I feel like I'm eating crazy pills when I read the countless bad takes around how the Vatican would have virtually anointed Anthropic. When …

SafetyDGX agent

I feel like I'm eating crazy pills when I read the countless bad takes around how the Vatican would have virtually anointed Anthropic. When if you read the Pope's encyclical it's actually a COMPLETE r

Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks

SafetyDGX agent

arXiv:2605.24217v1 Announce Type: new Abstract: As Large Language Models (LLMs) transition from research environments to production deployments, evaluating their performance against strict Service Lev

⚠️⚠️⚠️if you believe this you and think you are safe, you don’t understand how S&P is about to change the rules and what that means. run don…

SafetyDGX agent

⚠️⚠️⚠️if you believe this you and think you are safe, you don’t understand how S&P is about to change the rules and what that means. run don’t walk to read my essay “This one weird trick might cost yo

if you don’t follow or think you know better i urge you to read my essay “This one weird trick might cost your retirement fund billions”

SafetyDGX agent

Gary Marcus discusses a potentially overlooked financial risk that could have massive implications for retirement savings, presenting his analysis in an essay format on X. The post appears to emphasiz

Import AI 458: Reckoning with the future; and a singularity story

SafetyDGX agent

Import AI 458 discusses perspectives on AI's future trajectory and potential long-term scenarios, likely including analysis of singularity concepts and their implications. The newsletter entry examine

Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression

SafetyDGX agent

arXiv:2603.05691v2 Announce Type: replace Abstract: It is increasingly common in machine learning to use learned models to label data and then employ such data to train more capable models. The phenom

Improving Ensemble CAPE Forecasts with a Diffusion Model Incorporating Aerosol Information

SafetyDGX agent

arXiv:2605.24009v1 Announce Type: cross Abstract: Convective available potential energy (CAPE) is an important variable for forecasting severe weather and understanding deep convection and precipitati

Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation

SafetyDGX agent

arXiv:2605.24247v1 Announce Type: cross Abstract: Many automated labeling pipelines classify inputs into categories defined by a written specification, content moderation being a prominent use case. S

Improving the Completeness and Comparability of Segment Disclosures: A Large Language Model Approach

SafetyDGX agent

arXiv:2605.23924v1 Announce Type: new Abstract: Segment-level disclosures are a central component of financial reporting, providing insight into firms' internal organization and the allocation of econ

In our new paper, EDGE-OPD (EviDence GuidEd On-Policy Distillation), we introduce two key ideas for improving On-Policy Distillation (OPD). …

SafetyDGX agent

In our new paper, EDGE-OPD (EviDence GuidEd On-Policy Distillation), we introduce two key ideas for improving On-Policy Distillation (OPD). First, we use guided rollouts that inject privileged context

Inference-Time Alignment of Diffusion Models via Trust-Region Iterative Twisted Sequential Monte Carlo

SafetyDGX agent

arXiv:2605.25123v1 Announce Type: cross Abstract: We study inference-time alignment for diffusion-based generative models, aiming to steer a base model toward high-reward outputs without updating its

Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcement Learning for Reasoning

SafetyDGX agent

arXiv:2605.05226v2 Announce Type: replace-cross Abstract: The central challenge of reinforcement learning for reasoning lies not only in the sparsity of outcome-level supervision, but more fundamental

Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications

SafetyDGX agent

arXiv:2605.24883v1 Announce Type: new Abstract: The widespread integration of Large Language Models (LLMs) necessitates rigorous and systematic safety evaluation. Existing paradigms either rely on con

← Previous
1…116117118119120…214
Next →