AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,488 results
Safety

Over-Alignment vs Over-Fitting: The Role of Feature Learning Strength in Generalization

DGX agent

arXiv:2602.00827v2 Announce Type: replace Abstract: Feature learning strength (FLS), i.e., the inverse of the effective output scaling of a model, plays a critical role in shaping the optimization dyn

safetyarxiv-cs-lg
27 May 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Pair-In, Pair-Out: Latent Multi-Token Prediction for Efficient LLMs

DGX agent

arXiv:2605.27255v1 Announce Type: cross Abstract: Long chain-of-thought reasoning has made autoregressive decoding the dominant inference cost of modern large language models. Existing methods target

safetyarxiv-cs-ai
27 May 2026
Safety

Palantir CEO Alex Karp goes after AI slop. The fight over AI “slop” is really a fight over whether software is performing or merely pretendi…

DGX agent

Palantir CEO Alex Karp goes after AI slop. The fight over AI “slop” is really a fight over whether software is performing or merely pretending. 'The appearance of software working is not software work

safetygary-marcus--x
27 May 2026
Safety

per comments from @GergelyOrosz below i don’t think these data are compelling after all, and am deleting the OP

DGX agent

Gary Marcus deleted an original post after Gergely Orosz provided comments questioning the compelling nature of the data presented. The post appears to have been withdrawn due to critical feedback tha

safetygary-marcus--x
27 May 2026
Safety

PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization

DGX agent

arXiv:2507.16679v3 Announce Type: replace-cross Abstract: In-Context Learning has shown great potential for aligning Large Language Models (LLMs) with human values, helping reduce harmful outputs and

safetyarxiv-cs-ai
27 May 2026
Safety

Position: Machine Learning for Heart Transplant Allocation Policy Optimization Should Account for Incentives

DGX agent

arXiv:2602.04990v3 Announce Type: replace Abstract: The allocation of scarce donor organs constitutes one of the most consequential algorithmic challenges in healthcare. While the field is rapidly tra

safetyarxiv-cs-lg
27 May 2026
Safety

PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation

DGX agent

arXiv:2508.02806v3 Announce Type: replace Abstract: Recently, a significant improvement in the accuracy of 3D human pose estimation has been achieved by combining convolutional neural networks (CNNs)

safetyarxiv-cs-cv
27 May 2026
Safety

Quantized Keys Steal Attention: Bias Correction for KV-Cache Compression in Video Diffusion

DGX agent

arXiv:2605.26266v1 Announce Type: cross Abstract: Chunk-wise autoregressive video diffusion models rely on a KV cache of previously generated chunks to avoid redundant computation, but this cache quic

safetyarxiv-cs-ai
27 May 2026
Safety

Real Images, Worse Judgments: Evaluating Vision-Language Models on Concreteness and Imagery

DGX agent

arXiv:2605.27315v1 Announce Type: new Abstract: Visual inputs are often assumed to improve language understanding in multimodal models. We examine this assumption by asking whether vision-language mod

safetyarxiv-cs-cl
27 May 2026
Safety

Rethinking the Trust Region in LLM Reinforcement Learning

DGX agent

arXiv:2602.04879v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become a cornerstone for fine-tuning Large Language Models (LLMs), with Proximal Policy Optimization (PPO) ser

safetyarxiv-cs-ai
27 May 2026
Safety

Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective

DGX agent

arXiv:2605.26441v1 Announce Type: cross Abstract: This paper addresses the challenging task of weakly-supervised video temporal grounding. Existing approaches are generally based on the moment proposa

safetyarxiv-cs-ai
27 May 2026
Safety

RICE-PO: Turning Retrieval Interactions into Credit Signals for Reasoning Agents

DGX agent

arXiv:2605.26352v1 Announce Type: new Abstract: Retrieval is increasingly moving from one-shot matching toward interactive reasoning, where language agents iteratively inspect evidence, reformulate qu

safetyarxiv-cs-cl
27 May 2026
Safety

Sample Complexity of Policy Gradient for Log-Growth Control

DGX agent

arXiv:2605.26640v1 Announce Type: cross Abstract: We study the sample complexity of policy gradient for log-growth control -- the problem of learning, from observed state transitions, a feedback gain

safetyarxiv-cs-lg
27 May 2026
Safety

Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization

DGX agent

arXiv:2605.26282v1 Announce Type: new Abstract: Model-based reinforcement learning (RL) can be effectively supported at scale through the use of world models. However, in practice, scaling such approa

safetyarxiv-cs-lg
27 May 2026
Safety

SCENT: Aligning Mass Spectra with Molecular Structure for Olfactory Perception

DGX agent

arXiv:2605.27009v1 Announce Type: new Abstract: Predicting human olfactory perception from molecular structure has seen remarkable progress, yet these approaches require explicit chemical structure at

safetyarxiv-cs-lg
27 May 2026
Safety

SCKAN: Structural Consensus-based KAN Prototype Learning for Semi-Supervised Pancreas Segmentation

DGX agent

arXiv:2605.27032v1 Announce Type: new Abstract: Accurate pancreas segmentation is critical for early cancer diagnosis, where annotation scarcity necessitates Semi-Supervised Learning (SSL). However, d

safetyarxiv-cs-cv
27 May 2026
Safety

Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation

DGX agent

arXiv:2510.19420v2 Announce Type: replace-cross Abstract: Multi-Agent Systems (MAS) have become a prevalent paradigm for Large Language Model (LLM) applications. However, the complex multi-agent desig

safetyarxiv-cs-ai
27 May 2026
Safety

Self-Improvement Imitation with Biologically Guided Search for Protein Design Under Oracle Budgets

DGX agent

arXiv:2605.26690v1 Announce Type: cross Abstract: Protein sequence optimization under tight oracle budgets requires methods that explore vast combinatorial spaces while making each evaluation informat

safetyarxiv-cs-ai
27 May 2026
Safety

Signal-to-Noise Ratio and Sample Size Govern Representational Alignment in Neural Networks

DGX agent

arXiv:2605.26973v1 Announce Type: cross Abstract: Neural networks are known to develop latent representations that are aligned, namely structurally similar across networks trained with different archi

safetyarxiv-cs-lg
27 May 2026
Safety

SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing

DGX agent

arXiv:2512.14140v2 Announce Type: replace Abstract: Sketch editing requires jointly handling high-level semantic changes and precise local redrawing, a combination that is particularly challenging for

safetyarxiv-cs-cv
27 May 2026
Safety

SL-BiLEM: Structured Learnable Behavior-in-the-Loop Epidemic Modeling for Forecasting and Policy Evaluation

DGX agent

arXiv:2605.26704v1 Announce Type: cross Abstract: Epidemic forecasting faces a fundamental challenge: human behavior dynamically responds to disease spread, creating feedback loops that induce distrib

safetyarxiv-cs-ai
27 May 2026
Safety

some data i shared yesterday on anthropic revenue possibly slowing down aren’t as a compelling as i thought; i have deleted my posts and awa…

DGX agent

some data i shared yesterday on anthropic revenue possibly slowing down aren’t as a compelling as i thought; i have deleted my posts and await better data before drawing conclusions. h/t @GergelyOrosz

safetygary-marcus--x
27 May 2026
Safety

Spectral Principal Paths: A Spectral Perspective on Linear Representation Formation in LLMs

DGX agent

arXiv:2506.08543v3 Announce Type: replace Abstract: High-level representations have become a central focus in enhancing AI transparency and control, shifting attention from individual neurons or circu

safetyarxiv-cs-cv
27 May 2026
Safety

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training

DGX agent

arXiv:2605.26606v1 Announce Type: cross Abstract: Reinforcement learning (RL) is the dominant paradigm for post-training large language models. However, in the online, on-policy setting, rollout gener

safetyarxiv-cs-ai
27 May 2026
Safety

SQARL: A Size-Agnostic Reinforcement Learning approach for Circuit Allocation in Distributed Quantum Architectures

DGX agent

arXiv:2605.27027v1 Announce Type: new Abstract: The scaling of quantum processors is currently limited by technical challenges such as decoherence and cross-talk. As the number of qubits grows, interf

safetyarxiv-cs-lg
27 May 2026
Safety

Starbucks learned the hard way: you literally can’t even trust (current) AI to count.

DGX agent

Starbucks learned the hard way: you literally can’t even trust (current) AI to count. For some reason OpenAI doesn’t seem to be talking about how Starbucks spent years creating and testing an AI inven

safetygary-marcus--x
27 May 2026
Safety

StepOPSD: Step-Aware Online Preference Distillation for Agent Reinforcement Learning

DGX agent

arXiv:2605.27140v1 Announce Type: new Abstract: Reinforcement learning for multi-turn agents suffers from a credit-assignment mismatch: rewards are sparse and trajectory-level, while success often hin

safetyarxiv-cs-ai
27 May 2026
Safety

TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment

DGX agent

arXiv:2601.05729v2 Announce Type: replace Abstract: Recent studies have demonstrated the efficacy of integrating Group Relative Policy Optimization (GRPO) into flow matching models, particularly for t

safetyarxiv-cs-cv
27 May 2026
Safety

The Coverage Illusion: From Pre-retrieval Routing Failure to Post-retrieval Cascades in a Production RAG System

DGX agent

arXiv:2605.27220v1 Announce Type: new Abstract: In modern RAG pipelines, query augmentation methods such as HyDE and query expansion are applied to every query, resulting in substantial LLM inference

safetyarxiv-cs-cl
27 May 2026
Safety

The Labyrinth and the Thread: Rethinking Regularizations in Sequential Knowledge Editing for Large Language Models

DGX agent

arXiv:2605.26670v1 Announce Type: cross Abstract: Sequential editing of structured knowledge in large language models allows targeted factual updates without retraining, yet existing methods often rel

safetyarxiv-cs-ai
27 May 2026
Safety

The Rescue Effect: Spatio-Semantic Early Exit Bypasses Quantization Collapse in CLIP

DGX agent

arXiv:2605.26415v1 Announce Type: cross Abstract: Deploying Vision-Language Models on resource-constrained hardware typically requires INT8 quantization, but in joint-embedding architectures such as C

safetyarxiv-cs-ai
27 May 2026
Safety

The Role of Causal Features in Strategic Classification for Robustness and Alignment

DGX agent

arXiv:2605.27163v1 Announce Type: new Abstract: In strategic classification, an institution (e.g., a bank) anticipates adaptation from users who change their features to increase utility in a classifi

safetyarxiv-cs-lg
27 May 2026
Safety

This is fearmongering, @davidsacks, afaik. I see no candidate regulation being taken seriously that is an actual threat to the trillion doll…

DGX agent

This is fearmongering, @davidsacks, afaik. I see no candidate regulation being taken seriously that is an actual threat to the trillion dollar AI companies (which can easily afford whatever compliance

safetygary-marcus--x
27 May 2026
Safety

To model human linguistic prediction, make LLMs less superhuman

DGX agent

arXiv:2510.05141v2 Announce Type: replace Abstract: When we read, we make predictions about upcoming words; these predictions influence our reading behavior. The success of large language models (LLMs

safetyarxiv-cs-cl
27 May 2026
Safety

“Tokens got burned for millions of dollars without any real significant ROI to show for it.” hearing this over and over again

DGX agent

“Tokens got burned for millions of dollars without any real significant ROI to show for it.” hearing this over and over again The same conversation is happening across tech right now and many of us sa

safetygary-marcus--x
27 May 2026
Safety

Triadic Dynamics Aware Diffusion Posterior Sampling for Inverse Problems: Optimizing Guidance and Stochasticity Schedules

DGX agent

arXiv:2605.26470v1 Announce Type: new Abstract: Generative posterior sampling using diffusion models has emerged as a dominant paradigm for solving inverse problems in imaging, which usually consists

safetyarxiv-cs-cv
27 May 2026
Safety

Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges

DGX agent

arXiv:2605.26156v1 Announce Type: cross Abstract: The known stylistic biases in LLM judges, such as a preference for verbosity or specific sentence structures, present an underexplored security vulner

safetyarxiv-cs-ai
27 May 2026
Safety

UCPO: Uncertainty-Aware Policy Optimization

DGX agent

arXiv:2601.22648v2 Announce Type: replace Abstract: The key to building trustworthy large language models (LLMs) lies in endowing them with inherent uncertainty expression capabilities, thereby mitiga

safetyarxiv-cs-ai
27 May 2026
Safety

Uniboost: Global Coordination with Value Alignment for Fair and Efficient Traffic Allocation

DGX agent

arXiv:2605.26424v1 Announce Type: cross Abstract: With the rapid evolution of internet services, recommendation systems have become indispensable. In particular, the blending (re-ranking) stage plays

safetyarxiv-cs-ai
27 May 2026
Safety

Unique Lives, Shared World: Learning from Single-Life Videos

DGX agent

arXiv:2512.04085v2 Announce Type: replace Abstract: We introduce the 'single-life' learning paradigm, where we train a distinct vision model exclusively on egocentric videos captured by one individual

safetyarxiv-cs-cv
27 May 2026
Safety

V2V3D: View-to-View Denoised 3D Reconstruction for Light-Field Microscopy

DGX agent

arXiv:2504.07853v2 Announce Type: replace Abstract: Light field microscopy (LFM) has gained significant attention due to its ability to capture snapshot-based, large-scale 3D fluorescence images. Howe

safetyarxiv-cs-cv
27 May 2026
Safety

VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models

DGX agent

arXiv:2510.17759v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) extend large language models with visual reasoning, but their multimodal design also introduces new, underexplor

safetyarxiv-cs-cl
27 May 2026
Safety

VR-DAgger: Immersive VR for Dexterous Data Collection and Uncertainty-Guided On-Policy Correction

DGX agent

arXiv:2605.27114v1 Announce Type: new Abstract: Learning from demonstrations is effective for robotic manipulation, but collecting sufficient task-specific data remains a major bottleneck. Under distr

safetyarxiv-cs-ro
27 May 2026
Safety

we will have the most science and math submissions ever in the next few years. but most ≠ best. to get to best, quality will have to beat sl…

DGX agent

we will have the most science and math submissions ever in the next few years. but most ≠ best. to get to best, quality will have to beat slop Terence Tao: AI is creating a “traffic jam” in math If AI

safetygary-marcus--x
27 May 2026
Safety

When Does LeJEPA Learn a World Model?

DGX agent

arXiv:2605.26379v1 Announce Type: cross Abstract: A representation that scrambles the true degrees of freedom of the world cannot support reliable planning or compositional generalization. We prove th

safetyarxiv-cs-lg
27 May 2026
Safety

When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Detection

DGX agent

arXiv:2605.27348v1 Announce Type: cross Abstract: Recent generative models have largely closed the gap on low-level artifacts - pixel fingerprints, frequency anomalies, upsampling traces - particularl

safetyarxiv-cs-ai
27 May 2026
Safety

Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and Solver-Grounded Reasoning

DGX agent

arXiv:2605.26530v1 Announce Type: new Abstract: Legal reasoning requires distinguishing changes that matter from those that do not. Legal AI should remain stable under legally irrelevant perturbations

safetyarxiv-cs-ai
27 May 2026
Safety

You reap what you sow.

DGX agent

You reap what you sow. OpenAI’s public image is becoming a bigger liability as political backlash against AI grows. The company has spoken with several communications executives but has yet to fill th

safetygary-marcus--x
27 May 2026
← Previous
1…180181182183184…302
Next →