AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,813 results
27 May 2026

Self-Improvement Imitation with Biologically Guided Search for Protein Design Under Oracle Budgets

SafetyDGX agent

arXiv:2605.26690v1 Announce Type: cross Abstract: Protein sequence optimization under tight oracle budgets requires methods that explore vast combinatorial spaces while making each evaluation informat

Semantic Robustness Probing via Inpainting: An Interactive Tool for Safety-Critical Object Detection

SafetyDGX agent

arXiv:2605.27155v1 Announce Type: cross Abstract: Testing object detectors in safety-critical domains requires semantically meaningful probes beyond pixel-level corruptions. We present SemProbe, a too

Signal-to-Noise Ratio and Sample Size Govern Representational Alignment in Neural Networks

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.26973v1 Announce Type: cross Abstract: Neural networks are known to develop latent representations that are aligned, namely structurally similar across networks trained with different archi

SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing

SafetyDGX agent

arXiv:2512.14140v2 Announce Type: replace Abstract: Sketch editing requires jointly handling high-level semantic changes and precise local redrawing, a combination that is particularly challenging for

SL-BiLEM: Structured Learnable Behavior-in-the-Loop Epidemic Modeling for Forecasting and Policy Evaluation

SafetyDGX agent

arXiv:2605.26704v1 Announce Type: cross Abstract: Epidemic forecasting faces a fundamental challenge: human behavior dynamically responds to disease spread, creating feedback loops that induce distrib

some data i shared yesterday on anthropic revenue possibly slowing down aren’t as a compelling as i thought; i have deleted my posts and awa…

SafetyDGX agent

some data i shared yesterday on anthropic revenue possibly slowing down aren’t as a compelling as i thought; i have deleted my posts and await better data before drawing conclusions. h/t @GergelyOrosz

Spectral Principal Paths: A Spectral Perspective on Linear Representation Formation in LLMs

SafetyDGX agent

arXiv:2506.08543v3 Announce Type: replace Abstract: High-level representations have become a central focus in enhancing AI transparency and control, shifting attention from individual neurons or circu

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training

SafetyDGX agent

arXiv:2605.26606v1 Announce Type: cross Abstract: Reinforcement learning (RL) is the dominant paradigm for post-training large language models. However, in the online, on-policy setting, rollout gener

SQARL: A Size-Agnostic Reinforcement Learning approach for Circuit Allocation in Distributed Quantum Architectures

SafetyDGX agent

arXiv:2605.27027v1 Announce Type: new Abstract: The scaling of quantum processors is currently limited by technical challenges such as decoherence and cross-talk. As the number of qubits grows, interf

Starbucks learned the hard way: you literally can’t even trust (current) AI to count.

SafetyDGX agent

Starbucks learned the hard way: you literally can’t even trust (current) AI to count. For some reason OpenAI doesn’t seem to be talking about how Starbucks spent years creating and testing an AI inven

StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

SafetyDGX agent

arXiv:2605.27140v1 Announce Type: new Abstract: Reinforcement learning for multi-turn agents suffers from a credit-assignment mismatch: rewards are sparse and trajectory-level, while success often hin

Stochastic Decision Horizons for Constrained Reinforcement Learning

SafetyDGX agent

arXiv:2602.04599v2 Announce Type: replace Abstract: We propose stochastic decision horizons (SDH), a theoretically grounded framework for solving constrained RL problems with every-step constraint sat

TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment

SafetyDGX agent

arXiv:2601.05729v2 Announce Type: replace Abstract: Recent studies have demonstrated the efficacy of integrating Group Relative Policy Optimization (GRPO) into flow matching models, particularly for t

The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance

SafetyDGX agent

arXiv:2601.07085v2 Announce Type: replace-cross Abstract: Large language model (LLM)-based conversational AI systems present a challenge to human cognition that current frameworks for understanding mi

The Coverage Illusion: From Pre-retrieval Routing Failure to Post-retrieval Cascades in a Production RAG System

SafetyDGX agent

arXiv:2605.27220v1 Announce Type: new Abstract: In modern RAG pipelines, query augmentation methods such as HyDE and query expansion are applied to every query, resulting in substantial LLM inference

The Labyrinth and the Thread: Rethinking Regularizations in Sequential Knowledge Editing for Large Language Models

SafetyDGX agent

arXiv:2605.26670v1 Announce Type: cross Abstract: Sequential editing of structured knowledge in large language models allows targeted factual updates without retraining, yet existing methods often rel

The Rescue Effect: Spatio-Semantic Early Exit Bypasses Quantization Collapse in CLIP

SafetyDGX agent

arXiv:2605.26415v1 Announce Type: cross Abstract: Deploying Vision-Language Models on resource-constrained hardware typically requires INT8 quantization, but in joint-embedding architectures such as C

The Role of Causal Features in Strategic Classification for Robustness and Alignment

SafetyDGX agent

arXiv:2605.27163v1 Announce Type: new Abstract: In strategic classification, an institution (e.g., a bank) anticipates adaptation from users who change their features to increase utility in a classifi

Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction

SafetyDGX agent

arXiv:2505.11063v3 Announce Type: replace Abstract: LLM-based agents solve complex tasks through iterative reasoning, tool use, and environment interaction, where each intermediate thought directly sh

This is fearmongering, @davidsacks, afaik. I see no candidate regulation being taken seriously that is an actual threat to the trillion doll…

SafetyDGX agent

This is fearmongering, @davidsacks, afaik. I see no candidate regulation being taken seriously that is an actual threat to the trillion dollar AI companies (which can easily afford whatever compliance

To model human linguistic prediction, make LLMs less superhuman

SafetyDGX agent

arXiv:2510.05141v2 Announce Type: replace Abstract: When we read, we make predictions about upcoming words; these predictions influence our reading behavior. The success of large language models (LLMs

“Tokens got burned for millions of dollars without any real significant ROI to show for it.” hearing this over and over again

SafetyDGX agent

“Tokens got burned for millions of dollars without any real significant ROI to show for it.” hearing this over and over again The same conversation is happening across tech right now and many of us sa

TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving

SafetyDGX agent

arXiv:2605.27038v1 Announce Type: new Abstract: Vision-Language Models (VLMs) provide a promising foundation for autonomous driving planning, yet bridging semantic reasoning and precise 3D spatial for

Triadic Dynamics Aware Diffusion Posterior Sampling for Inverse Problems: Optimizing Guidance and Stochasticity Schedules

SafetyDGX agent

arXiv:2605.26470v1 Announce Type: new Abstract: Generative posterior sampling using diffusion models has emerged as a dominant paradigm for solving inverse problems in imaging, which usually consists

Trust, Geometry, and Rules: A Credibility-Aware Reinforcement Learning Framework for Safe USV Navigation under Uncertainty

SafetyDGX agent

arXiv:2605.26974v1 Announce Type: new Abstract: Autonomous navigation of Unmanned Surface Vehicles (USVs) that is safe and compliant with the International Regulations for Preventing Collisions at Sea

Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges

SafetyDGX agent

arXiv:2605.26156v1 Announce Type: cross Abstract: The known stylistic biases in LLM judges, such as a preference for verbosity or specific sentence structures, present an underexplored security vulner

UCPO: Uncertainty-Aware Policy Optimization

SafetyDGX agent

arXiv:2601.22648v2 Announce Type: replace Abstract: The key to building trustworthy large language models (LLMs) lies in endowing them with inherent uncertainty expression capabilities, thereby mitiga

Uniboost: Global Coordination with Value Alignment for Fair and Efficient Traffic Allocation

SafetyDGX agent

arXiv:2605.26424v1 Announce Type: cross Abstract: With the rapid evolution of internet services, recommendation systems have become indispensable. In particular, the blending (re-ranking) stage plays

Unique Lives, Shared World: Learning from Single-Life Videos

SafetyDGX agent

arXiv:2512.04085v2 Announce Type: replace Abstract: We introduce the 'single-life' learning paradigm, where we train a distinct vision model exclusively on egocentric videos captured by one individual

V2V3D: View-to-View Denoised 3D Reconstruction for Light-Field Microscopy

SafetyDGX agent

arXiv:2504.07853v2 Announce Type: replace Abstract: Light field microscopy (LFM) has gained significant attention due to its ability to capture snapshot-based, large-scale 3D fluorescence images. Howe

VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models

SafetyDGX agent

arXiv:2510.17759v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) extend large language models with visual reasoning, but their multimodal design also introduces new, underexplor

VR-DAgger: Immersive VR for Dexterous Data Collection and Uncertainty-Guided On-Policy Correction

SafetyDGX agent

arXiv:2605.27114v1 Announce Type: new Abstract: Learning from demonstrations is effective for robotic manipulation, but collecting sufficient task-specific data remains a major bottleneck. Under distr

we will have the most science and math submissions ever in the next few years. but most ≠ best. to get to best, quality will have to beat sl…

SafetyDGX agent

we will have the most science and math submissions ever in the next few years. but most ≠ best. to get to best, quality will have to beat slop Terence Tao: AI is creating a “traffic jam” in math If AI

When Does LeJEPA Learn a World Model?

SafetyDGX agent

arXiv:2605.26379v1 Announce Type: cross Abstract: A representation that scrambles the true degrees of freedom of the world cannot support reliable planning or compositional generalization. We prove th

When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection

SafetyDGX agent

arXiv:2605.27348v1 Announce Type: cross Abstract: Recent generative models have largely closed the gap on low-level artifacts - pixel fingerprints, frequency anomalies, upsampling traces - particularl

Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning

SafetyDGX agent

arXiv:2605.26530v1 Announce Type: new Abstract: Legal reasoning requires distinguishing changes that matter from those that do not. Legal AI should remain stable under legally irrelevant perturbations

You reap what you sow.

SafetyDGX agent

You reap what you sow. OpenAI’s public image is becoming a bigger liability as political backlash against AI grows. The company has spoken with several communications executives but has yet to fill th

26 May 2026

A comparative study of accuracy and rollout stability of temporal surrogate models

SafetyDGX agent

arXiv:2605.24868v1 Announce Type: new Abstract: Temporal surrogate models are effective for predicting chaotic dynamical systems where computational cost can be prohibitive. Several deep neural networ

A Contractive Feedback Semantics for Reinforcement Learning

SafetyDGX agent

arXiv:2605.24759v1 Announce Type: new Abstract: Discounted reinforcement learning is usually presented through Bellman equations on closed Markov decision processes. This paper develops a compositiona

A Decentralized LiDAR-SLAM System with Certifiably Optimal Pose Graph Optimization

SafetyDGX agent

arXiv:2605.25051v1 Announce Type: new Abstract: Decentralized multi-robot LiDAR-SLAM is essential for collaborative missions but faces significant challenges in maintaining global consistency. Existin

A Formal gatekeeper Framework for Safe Dual Control with Active Exploration

SafetyDGX agent

arXiv:2510.06351v2 Announce Type: replace Abstract: Planning safe trajectories under model uncertainty is a fundamental challenge. Robust planning ensures safety by considering worst-case realizations

A governance horizon for ethical-use constraints in open-weight AI models

SafetyDGX agent

arXiv:2605.24383v1 Announce Type: new Abstract: Ethical constraints on open-weight AI models are both a reflection of societal concerns and a foundation for AI governance policy. They are expected to

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

SafetyDGX agent

arXiv:2605.26026v1 Announce Type: cross Abstract: Light sheet fluorescence microscopy (LSM) enables high-resolution, three-dimensional (3D) imaging of biological specimens, providing rich volumetric d

A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning

SafetyDGX agent

arXiv:2605.25540v1 Announce Type: cross Abstract: Alzheimer's disease (AD) is a progressive neurodegenerative disorder and the leading cause of dementia, affecting memory, reasoning, communication, an

A post about Pope Leo XIV's encyclical on AI. Why the Pope is right, but perhaps not right enough. Artificial intelligence is reshaping the …

SafetyDGX agent

A post about Pope Leo XIV's encyclical on AI. Why the Pope is right, but perhaps not right enough. Artificial intelligence is reshaping the world in front of our eyes: how we communicate, how we acces

a shout out to the paper: https://arxiv.org/html/2605.25376v1 'KYA: A Framework-Agnostic Trust Layer for Autonomous Systems with Verifiable …

SafetyDGX agent

KYA is a framework-agnostic trust layer designed for autonomous systems that provides verifiable guarantees, addressing the need for trustworthy and transparent operation of AI agents across different

A Sober Look at Agentic Misalignment in Automated Workflows

SafetyDGX agent

arXiv:2605.24197v1 Announce Type: new Abstract: We study a class of emergent misalignment in multi-agent systems (MAS), with a focus on automated workflows, which we refer to agentic misalignment. Alt

A Tertiary Review of Large Language Model-Based Code Generating Tasks: Trends, Challenges, and Future Directions

SafetyDGX agent

arXiv:2605.25536v1 Announce Type: cross Abstract: Context. Large language models (LLMs) are increasingly applied to code-generating tasks (CGTs) in software engineering. While reported results are pro

A Unified Python Framework for Direct PPO-based Control of AHUs with Economizer Logic and CO2-Constrained Ventilation

SafetyDGX agent

arXiv:2605.24406v1 Announce Type: new Abstract: Optimizing HVAC (Heating, Ventilation and Air Conditioning) can enhance a building's energy efficiency while providing comfort levels for its occupants.

Active Learning for Stochastic Contextual Linear Bandits

SafetyDGX agent

arXiv:2605.24803v1 Announce Type: new Abstract: A key goal in stochastic contextual linear bandits is to efficiently learn a near-optimal policy. Prior algorithms for this problem learn a policy by st

Adaptive Human-AI Coordination via Hierarchical Action Disentanglement

SafetyDGX agent

arXiv:2605.24343v1 Announce Type: new Abstract: Human-AI collaboration requires agents that can adapt to diverse partner behaviors and skill levels while remaining robust to unseen partners. Existing

Adaptive Preference Optimization with Uncertainty-aware Utility Anchor

SafetyDGX agent

arXiv:2509.10515v1 Announce Type: cross Abstract: Offline preference optimization methods are efficient for large language models (LLMs) alignment. Direct Preference optimization (DPO)-like learning,

AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models

SafetyDGX agent

arXiv:2605.26013v1 Announce Type: cross Abstract: We introduce AdvantageFlow, a forward-process reinforcement learning algorithm for rectified flow models. Unlike Flow-GRPO, which optimizes the revers

Adversarial Error Correction for Visual Autoregressive Generation

SafetyDGX agent

arXiv:2605.24843v1 Announce Type: cross Abstract: Visual Autoregressive (VAR) models have emerged as a powerful paradigm for image synthesis by performing hierarchical next-scale prediction. However,

Agent-Centric Social Trajectory Prediction: A Free Energy Principle Perspective

SafetyDGX agent

arXiv:2605.25748v1 Announce Type: new Abstract: Trajectory prediction methods have demonstrated remarkable capabilities in capturing complex motion patterns. However, existing methods rely on global s

Agent-Facing Information Design in LLM Tool Registries

SafetyDGX agent

arXiv:2605.23916v1 Announce Type: cross Abstract: LLM tool registries function as unregulated advertising platforms: providers write free-text descriptions that agents use for selection, yet no measur

Agent Learning via Early Experience

SafetyDGX agent

arXiv:2510.08558v3 Announce Type: replace Abstract: A long-term goal of language agents is to learn and improve through their own experience, ultimately outperforming humans in complex, real-world tas

Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning

SafetyDGX agent

arXiv:2605.24216v1 Announce Type: cross Abstract: Monitoring autonomous large language model (LLM) agents for covert malicious behavior is challenging due to delayed, context-dependent, and long-horiz

Agents and AI responsibility; nice clip from @thsottiaux and @siliconvalleymm

SafetyDGX agent

Gary Marcus shares a video clip discussing the intersection of AI agents and questions of responsibility, featuring contributors Thierry Souttiaux and Silicon Valley commentators. The post highlights

AI-Assisted Systematization for Evaluating GenAI Systems

SafetyDGX agent

arXiv:2605.26001v1 Announce Type: cross Abstract: Evaluating generative AI (GenAI) systems is challenging because many targets of evaluation are broad, contested concepts, such as 'reasoning,' 'fairne

← Previous
1…114115116117118…214
Next →