AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,816 results
Safety

Starbucks learned the hard way: you literally can’t even trust (current) AI to count.

DGX agent

Starbucks learned the hard way: you literally can’t even trust (current) AI to count. For some reason OpenAI doesn’t seem to be talking about how Starbucks spent years creating and testing an AI inven

safetygary-marcus--x
27 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

DGX agent

arXiv:2605.27140v1 Announce Type: new Abstract: Reinforcement learning for multi-turn agents suffers from a credit-assignment mismatch: rewards are sparse and trajectory-level, while success often hin

safetyarxiv-cs-ai
27 May 2026
Safety

Stochastic Decision Horizons for Constrained Reinforcement Learning

DGX agent

arXiv:2602.04599v2 Announce Type: replace Abstract: We propose stochastic decision horizons (SDH), a theoretically grounded framework for solving constrained RL problems with every-step constraint sat

safetyarxiv-cs-lg
27 May 2026
Safety

TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment

DGX agent

arXiv:2601.05729v2 Announce Type: replace Abstract: Recent studies have demonstrated the efficacy of integrating Group Relative Policy Optimization (GRPO) into flow matching models, particularly for t

safetyarxiv-cs-cv
27 May 2026
Safety

The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance

DGX agent

arXiv:2601.07085v2 Announce Type: replace-cross Abstract: Large language model (LLM)-based conversational AI systems present a challenge to human cognition that current frameworks for understanding mi

safetyarxiv-cs-ai
27 May 2026
Safety

The Coverage Illusion: From Pre-retrieval Routing Failure to Post-retrieval Cascades in a Production RAG System

DGX agent

arXiv:2605.27220v1 Announce Type: new Abstract: In modern RAG pipelines, query augmentation methods such as HyDE and query expansion are applied to every query, resulting in substantial LLM inference

safetyarxiv-cs-cl
27 May 2026
Safety

The Labyrinth and the Thread: Rethinking Regularizations in Sequential Knowledge Editing for Large Language Models

DGX agent

arXiv:2605.26670v1 Announce Type: cross Abstract: Sequential editing of structured knowledge in large language models allows targeted factual updates without retraining, yet existing methods often rel

safetyarxiv-cs-ai
27 May 2026
Safety

The Rescue Effect: Spatio-Semantic Early Exit Bypasses Quantization Collapse in CLIP

DGX agent

arXiv:2605.26415v1 Announce Type: cross Abstract: Deploying Vision-Language Models on resource-constrained hardware typically requires INT8 quantization, but in joint-embedding architectures such as C

safetyarxiv-cs-ai
27 May 2026
Safety

The Role of Causal Features in Strategic Classification for Robustness and Alignment

DGX agent

arXiv:2605.27163v1 Announce Type: new Abstract: In strategic classification, an institution (e.g., a bank) anticipates adaptation from users who change their features to increase utility in a classifi

safetyarxiv-cs-lg
27 May 2026
Safety

Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction

DGX agent

arXiv:2505.11063v3 Announce Type: replace Abstract: LLM-based agents solve complex tasks through iterative reasoning, tool use, and environment interaction, where each intermediate thought directly sh

safetyarxiv-cs-ai
27 May 2026
Safety

This is fearmongering, @davidsacks, afaik. I see no candidate regulation being taken seriously that is an actual threat to the trillion doll…

DGX agent

This is fearmongering, @davidsacks, afaik. I see no candidate regulation being taken seriously that is an actual threat to the trillion dollar AI companies (which can easily afford whatever compliance

safetygary-marcus--x
27 May 2026
Safety

To model human linguistic prediction, make LLMs less superhuman

DGX agent

arXiv:2510.05141v2 Announce Type: replace Abstract: When we read, we make predictions about upcoming words; these predictions influence our reading behavior. The success of large language models (LLMs

safetyarxiv-cs-cl
27 May 2026
Safety

“Tokens got burned for millions of dollars without any real significant ROI to show for it.” hearing this over and over again

DGX agent

“Tokens got burned for millions of dollars without any real significant ROI to show for it.” hearing this over and over again The same conversation is happening across tech right now and many of us sa

safetygary-marcus--x
27 May 2026
Safety

TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving

DGX agent

arXiv:2605.27038v1 Announce Type: new Abstract: Vision-Language Models (VLMs) provide a promising foundation for autonomous driving planning, yet bridging semantic reasoning and precise 3D spatial for

safetyarxiv-cs-ro
27 May 2026
Safety

Triadic Dynamics Aware Diffusion Posterior Sampling for Inverse Problems: Optimizing Guidance and Stochasticity Schedules

DGX agent

arXiv:2605.26470v1 Announce Type: new Abstract: Generative posterior sampling using diffusion models has emerged as a dominant paradigm for solving inverse problems in imaging, which usually consists

safetyarxiv-cs-cv
27 May 2026
Safety

Trust, Geometry, and Rules: A Credibility-Aware Reinforcement Learning Framework for Safe USV Navigation under Uncertainty

DGX agent

arXiv:2605.26974v1 Announce Type: new Abstract: Autonomous navigation of Unmanned Surface Vehicles (USVs) that is safe and compliant with the International Regulations for Preventing Collisions at Sea

safetyarxiv-cs-ro
27 May 2026
Safety

Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges

DGX agent

arXiv:2605.26156v1 Announce Type: cross Abstract: The known stylistic biases in LLM judges, such as a preference for verbosity or specific sentence structures, present an underexplored security vulner

safetyarxiv-cs-ai
27 May 2026
Safety

UCPO: Uncertainty-Aware Policy Optimization

DGX agent

arXiv:2601.22648v2 Announce Type: replace Abstract: The key to building trustworthy large language models (LLMs) lies in endowing them with inherent uncertainty expression capabilities, thereby mitiga

safetyarxiv-cs-ai
27 May 2026
Safety

Uniboost: Global Coordination with Value Alignment for Fair and Efficient Traffic Allocation

DGX agent

arXiv:2605.26424v1 Announce Type: cross Abstract: With the rapid evolution of internet services, recommendation systems have become indispensable. In particular, the blending (re-ranking) stage plays

safetyarxiv-cs-ai
27 May 2026
Safety

Unique Lives, Shared World: Learning from Single-Life Videos

DGX agent

arXiv:2512.04085v2 Announce Type: replace Abstract: We introduce the 'single-life' learning paradigm, where we train a distinct vision model exclusively on egocentric videos captured by one individual

safetyarxiv-cs-cv
27 May 2026
Safety

V2V3D: View-to-View Denoised 3D Reconstruction for Light-Field Microscopy

DGX agent

arXiv:2504.07853v2 Announce Type: replace Abstract: Light field microscopy (LFM) has gained significant attention due to its ability to capture snapshot-based, large-scale 3D fluorescence images. Howe

safetyarxiv-cs-cv
27 May 2026
Safety

VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models

DGX agent

arXiv:2510.17759v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) extend large language models with visual reasoning, but their multimodal design also introduces new, underexplor

safetyarxiv-cs-cl
27 May 2026
Safety

VR-DAgger: Immersive VR for Dexterous Data Collection and Uncertainty-Guided On-Policy Correction

DGX agent

arXiv:2605.27114v1 Announce Type: new Abstract: Learning from demonstrations is effective for robotic manipulation, but collecting sufficient task-specific data remains a major bottleneck. Under distr

safetyarxiv-cs-ro
27 May 2026
Safety

we will have the most science and math submissions ever in the next few years. but most ≠ best. to get to best, quality will have to beat sl…

DGX agent

we will have the most science and math submissions ever in the next few years. but most ≠ best. to get to best, quality will have to beat slop Terence Tao: AI is creating a “traffic jam” in math If AI

safetygary-marcus--x
27 May 2026
Safety

When Does LeJEPA Learn a World Model?

DGX agent

arXiv:2605.26379v1 Announce Type: cross Abstract: A representation that scrambles the true degrees of freedom of the world cannot support reliable planning or compositional generalization. We prove th

safetyarxiv-cs-lg
27 May 2026
Safety

When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection

DGX agent

arXiv:2605.27348v1 Announce Type: cross Abstract: Recent generative models have largely closed the gap on low-level artifacts - pixel fingerprints, frequency anomalies, upsampling traces - particularl

safetyarxiv-cs-ai
27 May 2026
Safety

Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning

DGX agent

arXiv:2605.26530v1 Announce Type: new Abstract: Legal reasoning requires distinguishing changes that matter from those that do not. Legal AI should remain stable under legally irrelevant perturbations

safetyarxiv-cs-ai
27 May 2026
Safety

You reap what you sow.

DGX agent

You reap what you sow. OpenAI’s public image is becoming a bigger liability as political backlash against AI grows. The company has spoken with several communications executives but has yet to fill th

safetygary-marcus--x
27 May 2026
Safety

A comparative study of accuracy and rollout stability of temporal surrogate models

DGX agent

arXiv:2605.24868v1 Announce Type: new Abstract: Temporal surrogate models are effective for predicting chaotic dynamical systems where computational cost can be prohibitive. Several deep neural networ

safetyarxiv-cs-lg
26 May 2026
Safety

A Contractive Feedback Semantics for Reinforcement Learning

DGX agent

arXiv:2605.24759v1 Announce Type: new Abstract: Discounted reinforcement learning is usually presented through Bellman equations on closed Markov decision processes. This paper develops a compositiona

safetyarxiv-cs-lg
26 May 2026
Safety

A Decentralized LiDAR-SLAM System with Certifiably Optimal Pose Graph Optimization

DGX agent

arXiv:2605.25051v1 Announce Type: new Abstract: Decentralized multi-robot LiDAR-SLAM is essential for collaborative missions but faces significant challenges in maintaining global consistency. Existin

safetyarxiv-cs-ro
26 May 2026
Safety

A Formal gatekeeper Framework for Safe Dual Control with Active Exploration

DGX agent

arXiv:2510.06351v2 Announce Type: replace Abstract: Planning safe trajectories under model uncertainty is a fundamental challenge. Robust planning ensures safety by considering worst-case realizations

safetyarxiv-cs-ro
26 May 2026
Safety

A governance horizon for ethical-use constraints in open-weight AI models

DGX agent

arXiv:2605.24383v1 Announce Type: new Abstract: Ethical constraints on open-weight AI models are both a reflection of societal concerns and a foundation for AI governance policy. They are expected to

safetyarxiv-cs-ai
26 May 2026
Safety

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

DGX agent

arXiv:2605.26026v1 Announce Type: cross Abstract: Light sheet fluorescence microscopy (LSM) enables high-resolution, three-dimensional (3D) imaging of biological specimens, providing rich volumetric d

safetyarxiv-cs-ai
26 May 2026
Safety

A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning

DGX agent

arXiv:2605.25540v1 Announce Type: cross Abstract: Alzheimer's disease (AD) is a progressive neurodegenerative disorder and the leading cause of dementia, affecting memory, reasoning, communication, an

safetyarxiv-cs-lg
26 May 2026
Safety

A post about Pope Leo XIV's encyclical on AI. Why the Pope is right, but perhaps not right enough. Artificial intelligence is reshaping the …

DGX agent

A post about Pope Leo XIV's encyclical on AI. Why the Pope is right, but perhaps not right enough. Artificial intelligence is reshaping the world in front of our eyes: how we communicate, how we acces

safetyyann-lecun--x
26 May 2026
Safety

a shout out to the paper: https://arxiv.org/html/2605.25376v1 'KYA: A Framework-Agnostic Trust Layer for Autonomous Systems with Verifiable …

DGX agent

KYA is a framework-agnostic trust layer designed for autonomous systems that provides verifiable guarantees, addressing the need for trustworthy and transparent operation of AI agents across different

safetyyohei-nakajima--x
26 May 2026
Safety

A Sober Look at Agentic Misalignment in Automated Workflows

DGX agent

arXiv:2605.24197v1 Announce Type: new Abstract: We study a class of emergent misalignment in multi-agent systems (MAS), with a focus on automated workflows, which we refer to agentic misalignment. Alt

safetyarxiv-cs-ai
26 May 2026
Safety

A Tertiary Review of Large Language Model-Based Code Generating Tasks: Trends, Challenges, and Future Directions

DGX agent

arXiv:2605.25536v1 Announce Type: cross Abstract: Context. Large language models (LLMs) are increasingly applied to code-generating tasks (CGTs) in software engineering. While reported results are pro

safetyarxiv-cs-ai
26 May 2026
Safety

A Unified Python Framework for Direct PPO-based Control of AHUs with Economizer Logic and CO2-Constrained Ventilation

DGX agent

arXiv:2605.24406v1 Announce Type: new Abstract: Optimizing HVAC (Heating, Ventilation and Air Conditioning) can enhance a building's energy efficiency while providing comfort levels for its occupants.

safetyarxiv-cs-lg
26 May 2026
Safety

Active Learning for Stochastic Contextual Linear Bandits

DGX agent

arXiv:2605.24803v1 Announce Type: new Abstract: A key goal in stochastic contextual linear bandits is to efficiently learn a near-optimal policy. Prior algorithms for this problem learn a policy by st

safetyarxiv-cs-lg
26 May 2026
Safety

Adaptive Human-AI Coordination via Hierarchical Action Disentanglement

DGX agent

arXiv:2605.24343v1 Announce Type: new Abstract: Human-AI collaboration requires agents that can adapt to diverse partner behaviors and skill levels while remaining robust to unseen partners. Existing

safetyarxiv-cs-ai
26 May 2026
Safety

Adaptive Preference Optimization with Uncertainty-aware Utility Anchor

DGX agent

arXiv:2509.10515v1 Announce Type: cross Abstract: Offline preference optimization methods are efficient for large language models (LLMs) alignment. Direct Preference optimization (DPO)-like learning,

safetyarxiv-cs-cl
26 May 2026
Safety

AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models

DGX agent

arXiv:2605.26013v1 Announce Type: cross Abstract: We introduce AdvantageFlow, a forward-process reinforcement learning algorithm for rectified flow models. Unlike Flow-GRPO, which optimizes the revers

safetyarxiv-cs-ai
26 May 2026
Safety

Adversarial Error Correction for Visual Autoregressive Generation

DGX agent

arXiv:2605.24843v1 Announce Type: cross Abstract: Visual Autoregressive (VAR) models have emerged as a powerful paradigm for image synthesis by performing hierarchical next-scale prediction. However,

safetyarxiv-cs-ai
26 May 2026
Safety

Agent-Centric Social Trajectory Prediction: A Free Energy Principle Perspective

DGX agent

arXiv:2605.25748v1 Announce Type: new Abstract: Trajectory prediction methods have demonstrated remarkable capabilities in capturing complex motion patterns. However, existing methods rely on global s

safetyarxiv-cs-ai
26 May 2026
Safety

Agent-Facing Information Design in LLM Tool Registries

DGX agent

arXiv:2605.23916v1 Announce Type: cross Abstract: LLM tool registries function as unregulated advertising platforms: providers write free-text descriptions that agents use for selection, yet no measur

safetyarxiv-cs-ai
26 May 2026
Safety

Agent Learning via Early Experience

DGX agent

arXiv:2510.08558v3 Announce Type: replace Abstract: A long-term goal of language agents is to learn and improve through their own experience, ultimately outperforming humans in complex, real-world tas

safetyarxiv-cs-ai
26 May 2026
← Previous
1…143144145146147…267
Next →