AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,813 results
4 Jun 2026

Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models

SafetyDGX agent

arXiv:2502.01576v2 Announce Type: replace Abstract: Multi-modal Large Language Models (MLLMs) excel in vision-language tasks but remain vulnerable to visual adversarial perturbations that can induce h

RSC: Decentralized Rigid Formation Flocking for Large-Scale Swarms via Hybrid Predictive Control and Online Reconfiguration

SafetyDGX agent

arXiv:2606.04248v1 Announce Type: new Abstract: Decentralized rigid formation flocking requires a swarm of autonomous agents to maintain a predetermined geometric configuration while moving, relying s

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety

SafetyDGX agent

arXiv:2606.04051v1 Announce Type: cross Abstract: The evolution of LLMs into tool-enabled agents creates a new class of safety challenges associated with real-world execution rather than simple text g


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Scaling Self-Evolving Agents via Parametric Memory

SafetyDGX agent

arXiv:2606.04536v1 Announce Type: new Abstract: Existing memory-augmented LLM agents store past experience exclusively in prompt space, as textual summaries or retrieved passages, while keeping model

Scenario Generation for Risk-Aware Reinforcement Learning with Probably Approximately Safe Guarantees

SafetyDGX agent

arXiv:2606.04812v1 Announce Type: cross Abstract: Guaranteeing safety is critical to the deployment of reinforcement learning (RL) agents in the real-world, especially as policies learned using deep R

See also: https://x.com/emollick/status/1799046218917822682?s=20

SafetyDGX agent

See also: https://x.com/emollick/status/1799046218917822682?s=20 Alignment, of a sort: this paper conducts what they call a “moral Turing Test,” asking people to compare GPT-4o to humans on ethical qu

Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing

SafetyDGX agent

arXiv:2512.08094v2 Announce Type: replace Abstract: The goal of this work is to develop a universal approach for aligning subtitles (i.e., spoken language text with corresponding timestamps) to contin

Selective Coupling of Decoupled Informative Regions: Masked Attention Alignment for Data-Free Quantization of Vision Transformers

SafetyDGX agent

arXiv:2606.04373v1 Announce Type: cross Abstract: Data-Free Quantization (DFQ) addresses data security concerns by synthesizing samples, without accessing real data. It has garnered increasing attenti

Self-Distilled Policy Gradient

SafetyDGX agent

arXiv:2606.04036v1 Announce Type: new Abstract: On-policy self-distillation, where a language model conditions on privileged context to supervise its own generations, is a promising source of dense su

Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model

SafetyDGX agent

arXiv:2512.21917v3 Announce Type: replace-cross Abstract: Policy alignment to preference data typically assumes a known link function between observed preferences and latent rewards (e.g., Bradley-Ter

shit is getting real

SafetyDGX agent

shit is getting real Sam Altman, Dario Amodei, Demis Hassabis and many others have signed a letter urging Congress to increase security on orders of synthetic nucleic acids - and the equipment needed

Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents

SafetyDGX agent

arXiv:2510.13704v2 Announce Type: replace-cross Abstract: Recent works have proposed accelerating the wall-clock training time of actor-critic methods via the use of large-scale environment paralleliz

Smart Transportation Without Neurons -- Fair Metro Network Expansion with Tabular Reinforcement Learning

SafetyDGX agent

arXiv:2606.04167v1 Announce Type: cross Abstract: We tackle the Metro Network Expansion Problem (MNEP), a subset of the Transport Network Design Problem (TNDP), which focuses on expanding metro system

SocialCoach: Personalized Social Skill Learning with RL-based Agentic Tutoring and Practice

SafetyDGX agent

arXiv:2606.04155v1 Announce Type: cross Abstract: Social skills such as negotiation and leadership are crucial for personal and professional success in today's interconnected world. However, scalable

SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization

SafetyDGX agent

arXiv:2505.11166v3 Announce Type: replace-cross Abstract: Despite advances in pretraining with extended context sizes, large language models (LLMs) still face challenges in effectively utilizing real-

SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space

SafetyDGX agent

arXiv:2511.20102v3 Announce Type: replace Abstract: Sparse attention reduces the quadratic complexity of full self-attention but faces two challenges: (1) an attention gap, where applying sparse atten

Starmer is right about this; would be ironic if the net of it was to keep Starmer in power for longer.

SafetyDGX agent

Starmer is right about this; would be ironic if the net of it was to keep Starmer in power for longer. British Prime Minister Keir Starmer has accused Elon Musk of 'whipping up division' and interferi

Stumbling Into AI Emotional Dependence: How Routine AI Interactions Reshape Human Connection

SafetyDGX agent

arXiv:2606.04150v1 Announce Type: new Abstract: Public discourse and emerging policy typically assume that AI emotional support is a deliberate act: a lonely user consciously seeking comfort from a de

Success Conditioning as Policy Improvement: The Optimization Problem Solved by Imitating Success

SafetyDGX agent

arXiv:2601.18175v2 Announce Type: replace Abstract: A widely used technique for improving policies is success conditioning, in which one collects trajectories, identifies those that achieve a desired

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

SafetyDGX agent

arXiv:2602.06911v2 Announce Type: replace-cross Abstract: As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications,

Test-time reward-guided alignment of language models by importance sampling on pre-logit space

SafetyDGX agent

arXiv:2510.26219v3 Announce Type: replace-cross Abstract: Test-time alignment of large language models (LLMs) attracts attention because fine-tuning of LLMs requires high computational costs. In this

Testing Neural Networks via Bayesian-Guided Exploration of Decision Landscapes

SafetyDGX agent

arXiv:2606.04314v1 Announce Type: new Abstract: As neural networks are increasingly deployed in safety-critical domains, testing is essential to evaluate and improve their reliability. Existing testin

The Accountability Horizon: An Impossibility Theorem for Governing Human-Agent Collectives

SafetyDGX agent

arXiv:2604.07778v2 Announce Type: replace Abstract: Existing accountability frameworks for AI systems, legal, ethical, and regulatory, rest on a shared assumption: for any consequential outcome, at le

The Digital Apprentice: A Framework for Human-Directed Agentic AI Development

SafetyDGX agent

arXiv:2606.04321v1 Announce Type: new Abstract: Agentic AI deployments face a recurring design tension: heavy human oversight limits scale, while broad autonomy outruns accountability. Neither posture

The Invisible Lottery: How Subtle Cues Steer Algorithm Choice in LLM Code Generation

SafetyDGX agent

arXiv:2606.04057v1 Announce Type: cross Abstract: Large language models (LLMs) now generate substantial production code, often for tasks with multiple valid algorithmic solutions. Incidental prompt cu

The Loss Is Not Enough: Sampling Conditions and Inductive Bias in Contrastive Representation Learning

SafetyDGX agent

arXiv:2606.04280v1 Announce Type: cross Abstract: Contrastive learning has become a leading paradigm for self-supervised representation learning, yet the conditions under which it recovers meaningful

The Right Measure for Physics-Constrained Generation: A Co-Area Correction for Posterior-Consistent PDE Inverse Problems

SafetyDGX agent

arXiv:2606.04804v1 Announce Type: new Abstract: Generative models -- diffusion and flow matching -- are increasingly used to solve partial differential equation (PDE) inverse problems, enforcing the g

Think Fast and Far: Long-Horizon Online POMDP Planning via Rapid State Sampling

SafetyDGX agent

arXiv:2606.04355v1 Announce Type: new Abstract: Partially Observable Markov Decision Processes (POMDPs) are a general and principled framework for motion planning under uncertainty. Despite tremendous

THIS. I fully agree with @andrewyang, especially since GenAi leverages IP from a huge range of humans that were not adequately compensated.

SafetyDGX agent

THIS. I fully agree with @andrewyang, especially since GenAi leverages IP from a huge range of humans that were not adequately compensated. We should tax AI. https://www.cnbc.com/video/2026/06/04/andr

Three Predictions: 1. Some form of AI, probably neurosymbolic in nature, will come that is far more economical and data- and energy-efficien…

SafetyDGX agent

Three Predictions: 1. Some form of AI, probably neurosymbolic in nature, will come that is far more economical and data- and energy-efficient than LLMs, and it will make an absolute fortune. 2. LLMs,

Toward Multi-Domain and Long-Tailed Quantization via Feature Alignment and Scaling

SafetyDGX agent

arXiv:2606.04920v1 Announce Type: cross Abstract: Quantizing deep neural networks is essential for efficient inference on resource-constrained devices. However, most existing methods are designed for

Towards Pretraining Text Encoders for TabPFN

SafetyDGX agent

arXiv:2606.04876v1 Announce Type: new Abstract: Tabular foundation models, such as TabPFN, achieve strong performance on tabular datasets with numerical and categorical data, but do not natively handl

Trace-Mediated Peak Bias: Bridging Temporal Credit Assignment and Cognitive Heuristics in Deep Reinforcement Learning

SafetyDGX agent

arXiv:2606.04735v1 Announce Type: cross Abstract: Temporal credit assignment is central to both biological and artificial intelligence, yet its interaction with non-linear function approximation is po

Transferable Multi-Bit Watermarking Across Frozen Diffusion Models via Latent Consistency Bridges

SafetyDGX agent

arXiv:2603.20304v2 Announce Type: replace Abstract: As generative AI advances, global governance frameworks increasingly mandate verifiable content provenance. However, existing watermarking technique

TransTac: Visuo-Tactile Modality Transition via Ultraviolet-Encoded Transparent Elastomers

SafetyDGX agent

arXiv:2606.04477v1 Announce Type: new Abstract: Vision-based tactile sensors (VBTS) recover high-resolution contact geometry but typically rely on opaque elastomer layers that prevent visual transpare

Trump’s budget director Russ Vought is the most dangerous person you’ve never heard of. And he just proposed turning every federal grant int…

SafetyDGX agent

Trump’s budget director Russ Vought is the most dangerous person you’ve never heard of. And he just proposed turning every federal grant into a loyalty test. His plan would make funding for cancer res

Two House lawmakers unveil bipartisan AI legislation that would override some state AI laws and require top AI developers to implement risk-management plans (Politico)

SafetyDGX agent

Politico: Two House lawmakers unveil bipartisan AI legislation that would override some state AI laws and require top AI developers to implement risk-management plans — But it's the proposal to preemp

U-Net-Accelerated Quality-Diversity Optimization for Climate-Adaptive Urban Layouts

SafetyDGX agent

arXiv:2606.04658v1 Announce Type: cross Abstract: Optimizing urban layouts for climate adaptation requires balancing building density with cold-air ventilation. Because physics-based climate simulatio

UniFair: A unified fair clustering approach based on separation and compactness

SafetyDGX agent

arXiv:2606.04777v1 Announce Type: new Abstract: Clustering is increasingly used to support high-impact decisions, yet standard objectives such as k-means can produce clusterings that treat demographic

Unlocking Proactivity in Task-Oriented Dialogue

SafetyDGX agent

arXiv:2605.22240v2 Announce Type: replace Abstract: Proactive task-oriented dialogue (TOD), such as outbound sales, demands a persuasive agent that actively probes the user's concerns and steers the c

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models

SafetyDGX agent

arXiv:2602.19101v2 Announce Type: replace-cross Abstract: Value alignment of Large Language Models (LLMs) requires us to empirically measure these models' actual, acquired representation of value. Amo

VentAgent: When LLMs Learn to Breathe -- Multi-Objective Arbitration for ARDS Ventilation

SafetyDGX agent

arXiv:2606.04632v1 Announce Type: cross Abstract: Mechanical ventilation for Acute Respiratory Distress Syndrome (ARDS) requires balancing competing physiological goals, including oxygenation, lung pr

veriFIRE: an Industrial Case Study in Verifying Consistency Properties for a DNN-Based Wildfire Detection System

SafetyDGX agent

arXiv:2606.04121v1 Announce Type: cross Abstract: We present our ongoing work on the veriFIRE project: a collaboration between industry and academia, aimed at applying verification to increase the rel

vibe coding meme of the day, via @emanuelmaiberg @404mediaco

SafetyDGX agent

This post likely references a humorous or critical meme about 'vibe coding'—a colloquial term for writing code based on intuition rather than rigorous testing or established best practices. The post a

VT-3DAD: Cross-Category 3D Anomaly Detection via Visual-Text Normal Space Alignment

SafetyDGX agent

arXiv:2606.04369v1 Announce Type: new Abstract: Few-shot cross-category 3D anomaly detection aims to determine whether an unknown point cloud belongs to a target normal category using only a few norma

WAM-Nav: Asymmetric Latent World-Action Modeling for Unified Visual Navigation

SafetyDGX agent

arXiv:2606.04907v1 Announce Type: new Abstract: Visual navigation requires generating smooth and collision-free trajectories under complex geometric and physical constraints. Existing reactive policie

Watch the goal post shift unfold in real time: AGI used to be doing anything a person, including an expert, could do, and by the end of the …

SafetyDGX agent

Watch the goal post shift unfold in real time: AGI used to be doing anything a person, including an expert, could do, and by the end of the interview it’s not. I call this the AI bait and switch. As I

We are saying no to a data centre in Île-des-Chênes because there are big threats to the environment and not much benefit to the economy. Ou…

SafetyDGX agent

We are saying no to a data centre in Île-des-Chênes because there are big threats to the environment and not much benefit to the economy. Our message to any tech company out there… if you want to buil

What Can Eye Gaze Teach Us About Real-World Cycling? Insights From the Oxford RobotCycle Project

SafetyDGX agent

arXiv:2606.04989v1 Announce Type: cross Abstract: Although much is known about the physical danger of cycling situations, less is understood about the perceived danger of cycling. Furthermore, percept

What Type of Inference is Active Inference?

SafetyDGX agent

arXiv:2606.04935v1 Announce Type: new Abstract: Active inference casts decision-making as inference, with the Expected Free Energy (EFE) unifying goal-directed and information-seeking behavior. Recent

When Autoregressive Consistency Hurts Safety Alignment

SafetyDGX agent

arXiv:2606.04168v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is fragile in part because it is often shallow: fine-tuning mainly reshapes the model's behavior near t

When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks

SafetyDGX agent

arXiv:2606.04476v1 Announce Type: new Abstract: In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target functio

X4Val: Learning Neural Surrogates for Variance-Reduced Policy Evaluation

SafetyDGX agent

arXiv:2606.05159v1 Announce Type: new Abstract: Rigorous evaluation of learning-based robotic systems is an essential prerequisite for deployment. However, real-world test data is expensive to gather;

You know this is about Sam from line 1, even before you get to the picture.

SafetyDGX agent

You know this is about Sam from line 1, even before you get to the picture. TL;DR: He’s going to keep bullshitting his way to IPO. This entire industry is based on an illusion. It deliberately mistake

ZeroWBC: Learning Natural Whole-Body Humanoid Interaction from Human Egocentric Data

SafetyDGX agent

arXiv:2603.09170v2 Announce Type: replace-cross Abstract: Achieving versatile and natural whole-body humanoid interaction control remains challenging due to the high cost of whole-body teleoperation d

3 Jun 2026

A blueprint for democratic governance of frontier AI

SafetyDGX agent

This OpenAI publication outlines a proposed framework for establishing democratic oversight and governance structures for advanced AI systems, addressing how frontier AI development should be regulate

A Cartesian-3j Framework for Machine Learning Interatomic Potentials

SafetyDGX agent

arXiv:2512.16882v2 Announce Type: replace-cross Abstract: Machine learning interatomic potentials (MLIPs) have brought substantial gains in the extrapolation capability in computational chemistry. How

A Negative Result on Cross-Model Activation Transfer in a Pythia Multi-Hop Setting

SafetyDGX agent

arXiv:2606.03280v1 Announce Type: new Abstract: Recent work shows that language models can transmit behavioural traits through hidden signals in generated data during training. We ask whether a more d

A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026

SafetyDGX agent

arXiv:2606.03948v1 Announce Type: new Abstract: We implement simultaneous translation capability with the offline direct speech-to-text translation model Canary, using the state-of-the-art policy Alig

“a train wreck’ - @edels0n Not worth half of what you will be paying – Morningstar At most, worth half of what you will be paying – Martin P…

SafetyDGX agent

“a train wreck’ - @edels0n Not worth half of what you will be paying – Morningstar At most, worth half of what you will be paying – Martin Peers, The Information “financial terrorism” - @egrefen Get y

← Previous
1…9192939495…214
Next →