AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,488 results
Safety

Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models

DGX agent

arXiv:2605.28306v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models have emerged as a dominant paradigm for efficient LLM scaling, yet adapting them to non-English downstream tasks remai

safetyarxiv-cs-ai
28 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

SA4Depth: Consistent Pose-Depth Scale Alignment for Self-Supervised Monocular Depth Estimation

DGX agent

arXiv:2605.28477v1 Announce Type: new Abstract: Self-supervised depth estimation from monocular sequences relies on the joint learning of a depth and a pose network. Despite abundant research done to

safetyarxiv-cs-cv
28 May 2026
Safety

SCALE-COMM: Shared, Contrastively-Aligned Latent Embeddings for MARL Communication

DGX agent

arXiv:2605.27532v1 Announce Type: new Abstract: Emergent communication enables partially observant Autonomous Mobile Robots (AMRs) to coordinate effectively in decentralized multi-agent reinforcement

safetyarxiv-cs-ro
28 May 2026
Safety

SEMAGIC: Learning Semantically Consistent Deformable 3D Representations from In-the-Wild Images

DGX agent

arXiv:2605.27938v1 Announce Type: new Abstract: Learning deformable 3D object models from single-view in-the-wild images has enabled impressive 3D shape reconstruction without supervision. However, it

safetyarxiv-cs-cv
28 May 2026
Safety

Semiparametrically Efficient Inference for Kernel Measures of Noise Heterogeneity

DGX agent

arXiv:2605.27526v1 Announce Type: cross Abstract: We develop semiparametrically efficient inference for kernel measures of noise heterogeneity in additive noise models. In many applications, the regre

safetyarxiv-cs-lg
28 May 2026
Safety

Sense Representations Are Inducible Interfaces

DGX agent

arXiv:2605.28669v1 Announce Type: cross Abstract: Sense representations (explicit, per-token meaning decompositions) are useful for disambiguation, steering, and cross-lingual alignment, but existing

safetyarxiv-cs-ai
28 May 2026
Safety

Singular Vectors of Attention Heads Align with Features

DGX agent

arXiv:2602.13524v2 Announce Type: replace-cross Abstract: Identifying feature representations in language models is a central task in mechanistic interpretability. Several recent studies have made the

safetyarxiv-cs-ai
28 May 2026
Safety

Skill-Conditioned Gated Self-Distillation for LLM Reasoning

DGX agent

arXiv:2605.28791v1 Announce Type: cross Abstract: On-policy self-distillation (SD) improves LLM reasoning by using teacher-side privileged information (PI) to turn sparse verifier outcomes into dense

safetyarxiv-cs-ai
28 May 2026
Safety

SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit Assignment

DGX agent

arXiv:2605.27899v1 Announce Type: new Abstract: Structured skill prompts improve exploration in long-horizon agentic reinforcement learning (RL). Skill-augmented RL methods retain external skills at i

safetyarxiv-cs-ai
28 May 2026
Safety

Smaller, Younger, and More Impactful: How AI-Assisted Writing Transforms Research Teams

DGX agent

arXiv:2605.27404v1 Announce Type: cross Abstract: The era of Big Science has long been defined by increasingly large and specialized research teams pushing the frontiers of knowledge. However, recent

safetyarxiv-cs-ai
28 May 2026
Safety

Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards

DGX agent

arXiv:2605.28561v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has improved language models in domains such as mathematics and code, where correctness can be che

safetyarxiv-cs-cl
28 May 2026
Safety

SPAR: Support-Preserving Action Rectification

DGX agent

arXiv:2605.27877v1 Announce Type: cross Abstract: Offline policy improvement faces an inherent conflict between maximizing value and fitting the data distribution. While in-sample weighted regression

safetyarxiv-cs-ai
28 May 2026
Safety

SPRINT: Efficient Spectral Priors for Humanoid Athletic Sprints

DGX agent

arXiv:2605.28549v1 Announce Type: cross Abstract: The pursuit of humanoid athletic sprints is hindered by a scarcity of humanoid-viable kinematic reference data and the inability of existing framework

safetyarxiv-cs-lg
28 May 2026
Safety

STARS: Spike Tail-Aware Relational Synthesis for ANN-to-SNN Data-Free Knowledge Distillation

DGX agent

arXiv:2605.27409v1 Announce Type: cross Abstract: SNNs promise energy-efficient and low-latency inference, but their performance still trails that of ANNs. ANN-to-SNN knowledge distillation helps narr

safetyarxiv-cs-ai
28 May 2026
Safety

Structure-Guided Visual Perturbation Neutralization for LVLMs

DGX agent

arXiv:2605.27927v1 Announce Type: new Abstract: Image inputs enable Large Vision Language Models (LVLMs) to perceive fine-grained visual information, but also introduce a pixel-level attack surface th

safetyarxiv-cs-cv
28 May 2026
Safety

Structured Agent Distillation for Large Language Model

DGX agent

arXiv:2505.13820v5 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strong capabilities as decision-making agents by interleaving reasoning and actions, as seen in ReAct-sty

safetyarxiv-cs-ai
28 May 2026
Safety

Supervised Distributional Reduction via Optimal Transport and Dependence Maximization

DGX agent

arXiv:2605.27619v1 Announce Type: cross Abstract: Learning representations that capture both intrinsic data geometry and target-relevant structure remains a fundamental challenge, particularly in sett

safetyarxiv-cs-ai
28 May 2026
Safety

Supervised Semantic Differential for Cross-Cultural Concept Analysis: A Case Study of Human Affect

DGX agent

arXiv:2605.28225v1 Announce Type: new Abstract: Cross-cultural comparison of psychological meaning requires methods that go beyond word-level translation and examine how semantic dimensions are organi

safetyarxiv-cs-cl
28 May 2026
Safety

SYNAPSE: Neuro-Symbolic Visual Thought-to-Text Decoding via Topological Semantic Denoising

DGX agent

arXiv:2605.27790v1 Announce Type: new Abstract: Recent advances in large language models have accelerated open-vocabulary EEG-to-imagined-text decoding, where non-invasive neural activity recorded dur

safetyarxiv-cs-lg
28 May 2026
Safety

Teacher-Student Representational Alignment for Reinforcement Learning-Driven Imitation Learning

DGX agent

arXiv:2605.28372v1 Announce Type: new Abstract: Imitation learning (IL) from a state-based reinforcement learning (RL) policy is a common approach to overcome the curse of dimensionality in complex an

safetyarxiv-cs-lg
28 May 2026
Safety

Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers

DGX agent

arXiv:2605.27686v1 Announce Type: cross Abstract: Transformers process images and videos by flattening space and time into long token sequences. While attention and KV caching preserve past features,

safetyarxiv-cs-ai
28 May 2026
Safety

Test-Time Collective Action: Proxy-Based Perturbations for Correcting Algorithmic Harms

DGX agent

arXiv:2605.27689v1 Announce Type: new Abstract: When machine learning systems under-perform for particular subgroups, affected users typically have no way to correct these disparities without relying

safetyarxiv-cs-lg
28 May 2026
Safety

The AI numbers are starting to look very ugly. Even under 'best case' assumptions, FT's own data shows Microsoft AI ROI at -9%, Google at -1…

DGX agent

The AI numbers are starting to look very ugly. Even under 'best case' assumptions, FT's own data shows Microsoft AI ROI at -9%, Google at -15%, Meta at -28%, Oracle at -35%. Only Amazon barely comes o

safetygary-marcus--x
28 May 2026
Safety

The Attentional White Bear Effect in Transformer Language Models

DGX agent

arXiv:2605.28639v1 Announce Type: cross Abstract: Instruction-based suppression is widely used to prevent language models from generating prohibited content, yet it remains unclear whether suppression

safetyarxiv-cs-ai
28 May 2026
Model Releases

The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

DGX agent

arXiv:2605.27901v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models. Howeve

model-releasesarxiv-cs-ai
28 May 2026
Safety

The Illusion of Opting in AI-Mediated Consequential Decisions

DGX agent

arXiv:2605.28210v1 Announce Type: new Abstract: Drawing on Ullmann-Margalit's concept of opting (transformative, irrevocable, and shadowed by foreclosed alternatives), we show that current AI systems

safetyarxiv-cs-ai
28 May 2026
Safety

The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes

DGX agent

arXiv:2602.15515v2 Announce Type: replace-cross Abstract: Training against white-box deception detectors has been proposed as a way to make AI systems honest. However, such training risks models learn

safetyarxiv-cs-ai
28 May 2026
Safety

The trial has concluded, the facts have been confirmed. As the prosecuting counsel put it, Digwa used his “trump card” by alleging he had be…

DGX agent

The trial has concluded, the facts have been confirmed. As the prosecuting counsel put it, Digwa used his “trump card” by alleging he had been the victim of racist abuse when police officers arrived.

safetyelon-musk--x
28 May 2026
Safety

this sh*t isn’t even funny anymore. it’s a trillion dollar embarrassment.

DGX agent

Gary Marcus critiques the current state of AI development as wasteful and problematic, arguing that the industry's trillion-dollar investment represents a significant failure or misallocation of resou

safetygary-marcus--x
28 May 2026
Safety

tokenmaxxing is officially over

DGX agent

tokenmaxxing is officially over Sources: Amazon has shut down an internal leaderboard that tracked employees' use of AI tools after workers tried to boost their scores with needless tasks (@rafeuddin_

safetygary-marcus--x
28 May 2026
Safety

Toward Robust Semi-supervised Regression via Dual-stream Knowledge Distillation

DGX agent

arXiv:2508.14082v3 Announce Type: replace Abstract: Semi-supervised regression (SSR), which aims to predict continuous scores for samples while reducing the reliance on large-scale labeled data, has r

safetyarxiv-cs-lg
28 May 2026
Safety

Towards automated data analysis: A guided framework for LLM-based risk estimation

DGX agent

arXiv:2603.04631v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly integrated into critical decision-making pipelines, a trend that raises the demand for robust and auto

safetyarxiv-cs-ai
28 May 2026
Safety

Trust Me, I'm an Expert: Decoding and Steering Authority Bias in Large Language Models

DGX agent

arXiv:2601.13433v3 Announce Type: replace Abstract: Prior research demonstrates that performance of language models on reasoning tasks can be influenced by suggestions, hints and endorsements. However

safetyarxiv-cs-cl
28 May 2026
Safety

Turning Video Models into Generalist Robot Policies

DGX agent

arXiv:2605.27817v1 Announce Type: cross Abstract: Video generative models have emerged as a promising robotics backbone, capable of generating videos that depict the completion of complex tasks across

safetyarxiv-cs-ai
28 May 2026
Safety

turns out i wasn’t wrong:

DGX agent

Gary Marcus reflects on a past prediction or stance that he held, asserting its correctness in retrospect, likely addressing criticisms or skepticism he previously faced regarding AI, cognitive scienc

safetygary-marcus--x
28 May 2026
Safety

Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models

DGX agent

arXiv:2605.27376v1 Announce Type: cross Abstract: While prompt-based text-to-speech (TTS) models enable natural language-driven speaking style control, they often provide limited fine-grained control

safetyarxiv-cs-ai
28 May 2026
Safety

Unsupervised Identification and Removal of Spurious Correlations During Fine-Tuning

DGX agent

arXiv:2605.27676v1 Announce Type: cross Abstract: Fine-tuning a pretrained language model on a curated dataset can produce spurious correlations between the fine-tuning task and unintended latent fact

safetyarxiv-cs-lg
28 May 2026
Safety

Utility-Aware Multimodal Contrastive Learning for Product Image Generation

DGX agent

arXiv:2605.28733v1 Announce Type: new Abstract: Product images strongly influence consumer decision-making in online marketplaces. Empowered by multimodal contrastive learning, generative AI can outpu

safetyarxiv-cs-ai
28 May 2026
Safety

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning

DGX agent

arXiv:2605.28023v1 Announce Type: cross Abstract: Visual captioning requires models to capture visual content faithfully while minimizing both omission and hallucination. As the dominant paradigm for

safetyarxiv-cs-ai
28 May 2026
Safety

Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs

DGX agent

arXiv:2605.28565v1 Announce Type: cross Abstract: Users of search-augmented LLMs rely on citations as evidence that responses are grounded in real sources, and rarely verify the cited pages themselves

safetyarxiv-cs-ai
28 May 2026
Safety

Visualizing Latent Phase Structures in Locomotion Policies: A Multi-Environment Study with Temporal Feature Extension

DGX agent

arXiv:2605.28186v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) has been shown to achieve high performance on locomotion control tasks in MuJoCo benchmarks such as HalfCheetah, Ant

safetyarxiv-cs-ai
28 May 2026
Safety

VLMs May Not Globally Enhance Human Alignment over LLMs During Natural Reading

DGX agent

arXiv:2605.28818v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly useful computational models of human language processing, but it remains unclear whether vision-la

safetyarxiv-cs-cl
28 May 2026
Safety

What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025

DGX agent

arXiv:2601.07648v2 Announce Type: replace Abstract: As Natural Language Generation (NLG) dominates modern NLP, scalable evaluation remains a critical bottleneck. Consequently, LLM-as-a-judge (LaaJ) ad

safetyarxiv-cs-cl
28 May 2026
Safety

What Frozen VLAs Already Know About Success: A Probing Study of Value-Like Structure in Foundation Robot Policies

DGX agent

arXiv:2605.28527v1 Announce Type: new Abstract: Vision--language--action (VLA) policies are trained to imitate actions; their loss never asks them to estimate reward, progress, or future success. Thei

safetyarxiv-cs-ro
28 May 2026
Safety

When pre-training hurts LoRA fine-tuning: a dynamical analysis via single-index models

DGX agent

arXiv:2602.02855v2 Announce Type: replace Abstract: Pre-training on a source task is usually expected to facilitate fine-tuning on similar downstream problems. In this work, we mathematically show tha

safetyarxiv-cs-lg
28 May 2026
Safety

Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR

DGX agent

arXiv:2605.28295v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) trains reasoning models without labeled trajectories, relying on grouped rollouts to expose the po

safetyarxiv-cs-ai
28 May 2026
Safety

Who Uses AI? Platform Selection and the Measurement of Occupational AI Exposure

DGX agent

arXiv:2605.21743v2 Announce Type: replace Abstract: Conversation logs from AI platforms are increasingly used to measure occupational exposure to artificial intelligence, but the users observed in the

safetyarxiv-cs-ai
28 May 2026
Safety

*Why* do you think (current) AI’s are conscious, @Grimezsz? And extra credit, what do you mean by “conscious”?

DGX agent

*Why* do you think (current) AI’s are conscious, @Grimezsz? And extra credit, what do you mean by “conscious”? My only issue with the Pope's encyclical is I think they are conscious and therefore dese

safetygary-marcus--x
28 May 2026
← Previous
1…177178179180181…302
Next →