AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,490 results
4 Jun 2026

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success

SafetyDGX agent

arXiv:2601.18175v2 Announce Type: replace Abstract: A widely used technique for improving policies is success conditioning, in which one collects trajectories, identifies those that achieve a desired

Test-time reward-guided alignment of language models by importance sampling on pre-logit space

SafetyDGX agent

arXiv:2510.26219v3 Announce Type: replace-cross Abstract: Test-time alignment of large language models (LLMs) attracts attention because fine-tuning of LLMs requires high computational costs. In this

The Accountability Horizon: An Impossibility Theorem for Governing Human-Agent Collectives

SafetyDGX agent

arXiv:2604.07778v2 Announce Type: replace Abstract: Existing accountability frameworks for AI systems, legal, ethical, and regulatory, rest on a shared assumption: for any consequential outcome, at le

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

The Digital Apprentice: A Framework for Human-Directed Agentic AI Development

SafetyDGX agent

arXiv:2606.04321v1 Announce Type: new Abstract: Agentic AI deployments face a recurring design tension: heavy human oversight limits scale, while broad autonomy outruns accountability. Neither posture

The Invisible Lottery: How Subtle Cues Steer Algorithm Choice in LLM Code Generation

SafetyDGX agent

arXiv:2606.04057v1 Announce Type: cross Abstract: Large language models (LLMs) now generate substantial production code, often for tasks with multiple valid algorithmic solutions. Incidental prompt cu

The Loss Is Not Enough: Sampling Conditions and Inductive Bias in Contrastive Representation Learning

SafetyDGX agent

arXiv:2606.04280v1 Announce Type: cross Abstract: Contrastive learning has become a leading paradigm for self-supervised representation learning, yet the conditions under which it recovers meaningful

The Right Measure for Physics-Constrained Generation: A Co-Area Correction for Posterior-Consistent PDE Inverse Problems

SafetyDGX agent

arXiv:2606.04804v1 Announce Type: new Abstract: Generative models -- diffusion and flow matching -- are increasingly used to solve partial differential equation (PDE) inverse problems, enforcing the g

Think Fast and Far: Long-Horizon Online POMDP Planning via Rapid State Sampling

SafetyDGX agent

arXiv:2606.04355v1 Announce Type: new Abstract: Partially Observable Markov Decision Processes (POMDPs) are a general and principled framework for motion planning under uncertainty. Despite tremendous

THIS. I fully agree with @andrewyang, especially since GenAi leverages IP from a huge range of humans that were not adequately compensated.

SafetyDGX agent

THIS. I fully agree with @andrewyang, especially since GenAi leverages IP from a huge range of humans that were not adequately compensated. We should tax AI. https://www.cnbc.com/video/2026/06/04/andr

Three Predictions: 1. Some form of AI, probably neurosymbolic in nature, will come that is far more economical and data- and energy-efficien…

SafetyDGX agent

Three Predictions: 1. Some form of AI, probably neurosymbolic in nature, will come that is far more economical and data- and energy-efficient than LLMs, and it will make an absolute fortune. 2. LLMs,

Toward Multi-Domain and Long-Tailed Quantization via Feature Alignment and Scaling

SafetyDGX agent

arXiv:2606.04920v1 Announce Type: cross Abstract: Quantizing deep neural networks is essential for efficient inference on resource-constrained devices. However, most existing methods are designed for

Towards Pretraining Text Encoders for TabPFN

SafetyDGX agent

arXiv:2606.04876v1 Announce Type: new Abstract: Tabular foundation models, such as TabPFN, achieve strong performance on tabular datasets with numerical and categorical data, but do not natively handl

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning

SafetyDGX agent

arXiv:2606.04735v1 Announce Type: cross Abstract: Temporal credit assignment is central to both biological and artificial intelligence, yet its interaction with non-linear function approximation is po

Transferable Multi-Bit Watermarking Across Frozen Diffusion Models via Latent Consistency Bridges

SafetyDGX agent

arXiv:2603.20304v2 Announce Type: replace Abstract: As generative AI advances, global governance frameworks increasingly mandate verifiable content provenance. However, existing watermarking technique

TransTac: Visuo-Tactile Modality Transition via Ultraviolet-Encoded Transparent Elastomers

SafetyDGX agent

arXiv:2606.04477v1 Announce Type: new Abstract: Vision-based tactile sensors (VBTS) recover high-resolution contact geometry but typically rely on opaque elastomer layers that prevent visual transpare

Trump’s budget director Russ Vought is the most dangerous person you’ve never heard of. And he just proposed turning every federal grant int…

SafetyDGX agent

Trump’s budget director Russ Vought is the most dangerous person you’ve never heard of. And he just proposed turning every federal grant into a loyalty test. His plan would make funding for cancer res

U-Net-Accelerated Quality-Diversity Optimization for Climate-Adaptive Urban Layouts

SafetyDGX agent

arXiv:2606.04658v1 Announce Type: cross Abstract: Optimizing urban layouts for climate adaptation requires balancing building density with cold-air ventilation. Because physics-based climate simulatio

UniFair: A unified fair clustering approach based on separation and compactness

SafetyDGX agent

arXiv:2606.04777v1 Announce Type: new Abstract: Clustering is increasingly used to support high-impact decisions, yet standard objectives such as k-means can produce clusterings that treat demographic

Unlocking Proactivity in Task-Oriented Dialogue

SafetyDGX agent

arXiv:2605.22240v2 Announce Type: replace Abstract: Proactive task-oriented dialogue (TOD), such as outbound sales, demands a persuasive agent that actively probes the user's concerns and steers the c

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models

SafetyDGX agent

arXiv:2602.19101v2 Announce Type: replace-cross Abstract: Value alignment of Large Language Models (LLMs) requires us to empirically measure these models' actual, acquired representation of value. Amo

VentAgent: When LLMs Learn to Breathe -- Multi-Objective Arbitration for ARDS Ventilation

SafetyDGX agent

arXiv:2606.04632v1 Announce Type: cross Abstract: Mechanical ventilation for Acute Respiratory Distress Syndrome (ARDS) requires balancing competing physiological goals, including oxygenation, lung pr

vibe coding meme of the day, via @emanuelmaiberg @404mediaco

SafetyDGX agent

This post likely references a humorous or critical meme about 'vibe coding'—a colloquial term for writing code based on intuition rather than rigorous testing or established best practices. The post a

VT-3DAD: Cross-Category 3D Anomaly Detection via Visual-Text Normal Space Alignment

SafetyDGX agent

arXiv:2606.04369v1 Announce Type: new Abstract: Few-shot cross-category 3D anomaly detection aims to determine whether an unknown point cloud belongs to a target normal category using only a few norma

WAM-Nav: Asymmetric Latent World-Action Modeling for Unified Visual Navigation

SafetyDGX agent

arXiv:2606.04907v1 Announce Type: new Abstract: Visual navigation requires generating smooth and collision-free trajectories under complex geometric and physical constraints. Existing reactive policie

Watch the goal post shift unfold in real time: AGI used to be doing anything a person, including an expert, could do, and by the end of the …

SafetyDGX agent

Watch the goal post shift unfold in real time: AGI used to be doing anything a person, including an expert, could do, and by the end of the interview it’s not. I call this the AI bait and switch. As I

We are saying no to a data centre in Île-des-Chênes because there are big threats to the environment and not much benefit to the economy. Ou…

SafetyDGX agent

We are saying no to a data centre in Île-des-Chênes because there are big threats to the environment and not much benefit to the economy. Our message to any tech company out there… if you want to buil

What Type of Inference is Active Inference?

SafetyDGX agent

arXiv:2606.04935v1 Announce Type: new Abstract: Active inference casts decision-making as inference, with the Expected Free Energy (EFE) unifying goal-directed and information-seeking behavior. Recent

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks

SafetyDGX agent

arXiv:2606.04476v1 Announce Type: new Abstract: In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target functio

X4Val: Learning Neural Surrogates for Variance-Reduced Policy Evaluation

SafetyDGX agent

arXiv:2606.05159v1 Announce Type: new Abstract: Rigorous evaluation of learning-based robotic systems is an essential prerequisite for deployment. However, real-world test data is expensive to gather;

You know this is about Sam from line 1, even before you get to the picture.

SafetyDGX agent

You know this is about Sam from line 1, even before you get to the picture. TL;DR: He’s going to keep bullshitting his way to IPO. This entire industry is based on an illusion. It deliberately mistake

ZeroWBC: Learning Natural Whole-Body Humanoid Interaction from Human Egocentric Data

SafetyDGX agent

arXiv:2603.09170v2 Announce Type: replace-cross Abstract: Achieving versatile and natural whole-body humanoid interaction control remains challenging due to the high cost of whole-body teleoperation d

3 Jun 2026

A blueprint for democratic governance of frontier AI

SafetyDGX agent

This OpenAI publication outlines a proposed framework for establishing democratic oversight and governance structures for advanced AI systems, addressing how frontier AI development should be regulate

A Cartesian-3j Framework for Machine Learning Interatomic Potentials

SafetyDGX agent

arXiv:2512.16882v2 Announce Type: replace-cross Abstract: Machine learning interatomic potentials (MLIPs) have brought substantial gains in the extrapolation capability in computational chemistry. How

A Negative Result on Cross-Model Activation Transfer in a Pythia Multi-Hop Setting

SafetyDGX agent

arXiv:2606.03280v1 Announce Type: new Abstract: Recent work shows that language models can transmit behavioural traits through hidden signals in generated data during training. We ask whether a more d

A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026

SafetyDGX agent

arXiv:2606.03948v1 Announce Type: new Abstract: We implement simultaneous translation capability with the offline direct speech-to-text translation model Canary, using the state-of-the-art policy Alig

“a train wreck’ - @edels0n Not worth half of what you will be paying – Morningstar At most, worth half of what you will be paying – Martin P…

SafetyDGX agent

“a train wreck’ - @edels0n Not worth half of what you will be paying – Morningstar At most, worth half of what you will be paying – Martin Peers, The Information “financial terrorism” - @egrefen Get y

Adaptive Causal Alignment for High-Confidence Adversarial Training

SafetyDGX agent

arXiv:2606.03925v1 Announce Type: new Abstract: Inverse adversarial training leverages high-confidence predictions to stabilize robust learning, yet we uncover a critical paradox: high confidence ofte

AirDreamer: Generalist Drone Navigation with World Models

SafetyDGX agent

arXiv:2606.03252v1 Announce Type: cross Abstract: Navigating a drone in unseen and cluttered environments requires reliable generalization to unseen scene layouts and understanding of environmental st

Aletheia: What Makes RLVR For Code Verifiers Tick?

SafetyDGX agent

arXiv:2601.12186v3 Announce Type: replace-cross Abstract: Multi-domain thinking verifiers trained via Reinforcement Learning with Verifiable Rewards (RLVR) are a cornerstone of modern post-training. H

Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model Enhancement

SafetyDGX agent

arXiv:2412.01282v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) bring powerful understanding and reasoning capabilities to multimodal tasks. Meanwhile, the great need for capab

Aligning Data-Driven Predictors with Allocation: A Decision-Focused Approach to Survival Analysis

SafetyDGX agent

arXiv:2606.02671v1 Announce Type: cross Abstract: Machine learning predictors have become essential tools for guiding automated decision making. However, a major misalignment persists: predictive mode

Alignment-Aware Decoding

SafetyDGX agent

arXiv:2509.26169v2 Announce Type: replace Abstract: Alignment of large language models remains a central challenge in natural language processing. Preference optimization has emerged as a popular and

an astonishing statistic, consistent with the claim I made the other day that Elon’s best days may be behind him:

SafetyDGX agent

Gary Marcus shared a statistic on X that he claims supports his earlier assertion that Elon Musk's most successful period may be in the past. The post likely presents data related to Musk's business p

Are we really tilting? The mechanics of reward guidance in flow and diffusion models

SafetyDGX agent

arXiv:2606.02884v1 Announce Type: cross Abstract: Reward guidance algorithms steer a learned generative process toward the reward-tilted measure at inference time. While empirically powerful, these me

ASAP: Exploiting the Satisficing Generalization Edge in Neural Combinatorial Optimization

SafetyDGX agent

arXiv:2501.17377v4 Announce Type: replace-cross Abstract: Deep Reinforcement Learning (DRL) has emerged as a promising approach for solving Combinatorial Optimization (CO) problems, such as the 3D Bin

ASymPO: Asymmetric-Scale Policy Optimization for Asynchronous LLM Post-Training Without Behavior Information

SafetyDGX agent

arXiv:2606.03070v1 Announce Type: cross Abstract: Asynchronous reinforcement learning can improve language-model post-training throughput by decoupling response generation from policy optimization, bu

Attention Calibration for Position-Fair Dense Information Retrieval

SafetyDGX agent

arXiv:2606.02737v1 Announce Type: cross Abstract: Dense retrieval models exhibit positional bias: retrieval effectiveness degrades when relevant information appears later in a passage (Zeng et al., 20

Bayesian Tensor Decomposition with Diffusion Model Prior

SafetyDGX agent

arXiv:2606.03212v1 Announce Type: new Abstract: Low-rank tensor decomposition (TD) is usually effective on clean, fully observed data, but it often degrades under severe missingness or noise. Low-rank

Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching

SafetyDGX agent

arXiv:2602.12221v2 Announce Type: replace Abstract: We propose UniDFlow, a unified discrete flow-matching framework for multimodal understanding, generation, and editing. It decouples understanding an

Bionic Human-Motion Style Transfer for Physically Executable Whole-Body Control of Humanoid Robots

SafetyDGX agent

arXiv:2606.03536v1 Announce Type: new Abstract: Expressive whole-body motion is important for humanoid robots operating in human environments, where robots are expected to move stably while presenting

Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs

SafetyDGX agent

arXiv:2606.03647v1 Announce Type: cross Abstract: Accurately evaluating adversarial robustness is a longstanding challenge. A flawed attack design can inflate robustness estimates, making deployment r

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL

SafetyDGX agent

arXiv:2510.08977v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) efficiently scales the reasoning ability of large language models (LLMs) but is bottlene

Brief Announcement: Generative Markov Model for Distributed Computing Systems

SafetyDGX agent

arXiv:2606.03061v1 Announce Type: cross Abstract: Emerging distributed computing paradigms, such as the computing continuum, are inherently heterogeneous, stochastic, and complex. Efficiently and effe

Building Better Activation Oracles

SafetyDGX agent

arXiv:2606.02609v1 Announce Type: cross Abstract: Activation Oracles (AOs) are promising methods for interpreting residual stream activations. However, current AOs face important issues, such as hallu

Coherence Maximization Improves Pluralistic Alignment

SafetyDGX agent

arXiv:2606.03110v1 Announce Type: new Abstract: Aligning AI systems with diverse human values requires value specifications grounded in concrete examples, but generating such examples without extensiv

Collab-REC: An LLM-based Agentic Framework for Balancing Recommendations in Tourism

SafetyDGX agent

arXiv:2508.15030v5 Announce Type: replace Abstract: We propose COLLAB-REC, a multi-agent framework designed to counteract popularity bias and improve diversity in tourism recommendations. In our setup

Consistency Training Can Entrench Misalignment

SafetyDGX agent

arXiv:2606.03810v1 Announce Type: cross Abstract: Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures. Such methods are simple, scalable, an

ConTrack: Constrained Hand Motion Tracking with Adaptive Trade-off Control

SafetyDGX agent

arXiv:2606.03177v1 Announce Type: new Abstract: Human demonstrations provide strong priors for robot manipulation, yet it is non-trivial to transfer them to execute on real robots due to the kinematic

ConTraIRL: Factorized Contrastive Abstractions for Transferable IRL

SafetyDGX agent

arXiv:2606.03017v1 Announce Type: cross Abstract: Reward transfer in Inverse Reinforcement Learning (IRL) is unreliable when policies must generalize to unseen combinations of environment dynamics and

Contrastive Neural Algorithmic Reasoning for Graph Coloring

SafetyDGX agent

arXiv:2606.03923v1 Announce Type: new Abstract: Graph coloring seeks to assigns colors to a graph's nodes so that adjacent nodes receive different colors, using as few colors as possible. Here, we stu

← Previous
1…125126127128129…242
Next →