AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,435 results
Safety

SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit Assignment

DGX agent

arXiv:2605.27899v1 Announce Type: new Abstract: Structured skill prompts improve exploration in long-horizon agentic reinforcement learning (RL). Skill-augmented RL methods retain external skills at i

safetyarxiv-cs-ai
28 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Smaller, Younger, and More Impactful: How AI-Assisted Writing Transforms Research Teams

DGX agent

arXiv:2605.27404v1 Announce Type: cross Abstract: The era of Big Science has long been defined by increasingly large and specialized research teams pushing the frontiers of knowledge. However, recent

safetyarxiv-cs-ai
28 May 2026
Safety

Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards

DGX agent

arXiv:2605.28561v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has improved language models in domains such as mathematics and code, where correctness can be che

safetyarxiv-cs-cl
28 May 2026
Safety

SPAR: Support-Preserving Action Rectification

DGX agent

arXiv:2605.27877v1 Announce Type: cross Abstract: Offline policy improvement faces an inherent conflict between maximizing value and fitting the data distribution. While in-sample weighted regression

safetyarxiv-cs-ai
28 May 2026
Safety

SPRINT: Efficient Spectral Priors for Humanoid Athletic Sprints

DGX agent

arXiv:2605.28549v1 Announce Type: cross Abstract: The pursuit of humanoid athletic sprints is hindered by a scarcity of humanoid-viable kinematic reference data and the inability of existing framework

safetyarxiv-cs-lg
28 May 2026
Safety

STARS: Spike Tail-Aware Relational Synthesis for ANN-to-SNN Data-Free Knowledge Distillation

DGX agent

arXiv:2605.27409v1 Announce Type: cross Abstract: SNNs promise energy-efficient and low-latency inference, but their performance still trails that of ANNs. ANN-to-SNN knowledge distillation helps narr

safetyarxiv-cs-ai
28 May 2026
Safety

Structure-Guided Visual Perturbation Neutralization for LVLMs

DGX agent

arXiv:2605.27927v1 Announce Type: new Abstract: Image inputs enable Large Vision Language Models (LVLMs) to perceive fine-grained visual information, but also introduce a pixel-level attack surface th

safetyarxiv-cs-cv
28 May 2026
Safety

Structured Agent Distillation for Large Language Model

DGX agent

arXiv:2505.13820v5 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strong capabilities as decision-making agents by interleaving reasoning and actions, as seen in ReAct-sty

safetyarxiv-cs-ai
28 May 2026
Safety

Supervised Distributional Reduction via Optimal Transport and Dependence Maximization

DGX agent

arXiv:2605.27619v1 Announce Type: cross Abstract: Learning representations that capture both intrinsic data geometry and target-relevant structure remains a fundamental challenge, particularly in sett

safetyarxiv-cs-ai
28 May 2026
Safety

Supervised Semantic Differential for Cross-Cultural Concept Analysis: A Case Study of Human Affect

DGX agent

arXiv:2605.28225v1 Announce Type: new Abstract: Cross-cultural comparison of psychological meaning requires methods that go beyond word-level translation and examine how semantic dimensions are organi

safetyarxiv-cs-cl
28 May 2026
Safety

SYNAPSE: Neuro-Symbolic Visual Thought-to-Text Decoding via Topological Semantic Denoising

DGX agent

arXiv:2605.27790v1 Announce Type: new Abstract: Recent advances in large language models have accelerated open-vocabulary EEG-to-imagined-text decoding, where non-invasive neural activity recorded dur

safetyarxiv-cs-lg
28 May 2026
Safety

Teacher-Student Representational Alignment for Reinforcement Learning-Driven Imitation Learning

DGX agent

arXiv:2605.28372v1 Announce Type: new Abstract: Imitation learning (IL) from a state-based reinforcement learning (RL) policy is a common approach to overcome the curse of dimensionality in complex an

safetyarxiv-cs-lg
28 May 2026
Safety

Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers

DGX agent

arXiv:2605.27686v1 Announce Type: cross Abstract: Transformers process images and videos by flattening space and time into long token sequences. While attention and KV caching preserve past features,

safetyarxiv-cs-ai
28 May 2026
Safety

Test-Time Collective Action: Proxy-Based Perturbations for Correcting Algorithmic Harms

DGX agent

arXiv:2605.27689v1 Announce Type: new Abstract: When machine learning systems under-perform for particular subgroups, affected users typically have no way to correct these disparities without relying

safetyarxiv-cs-lg
28 May 2026
Safety

The Attentional White Bear Effect in Transformer Language Models

DGX agent

arXiv:2605.28639v1 Announce Type: cross Abstract: Instruction-based suppression is widely used to prevent language models from generating prohibited content, yet it remains unclear whether suppression

safetyarxiv-cs-ai
28 May 2026
Model Releases

The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

DGX agent

arXiv:2605.27901v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models. Howeve

model-releasesarxiv-cs-ai
28 May 2026
Safety

The Illusion of Opting in AI-Mediated Consequential Decisions

DGX agent

arXiv:2605.28210v1 Announce Type: new Abstract: Drawing on Ullmann-Margalit's concept of opting (transformative, irrevocable, and shadowed by foreclosed alternatives), we show that current AI systems

safetyarxiv-cs-ai
28 May 2026
Safety

The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes

DGX agent

arXiv:2602.15515v2 Announce Type: replace-cross Abstract: Training against white-box deception detectors has been proposed as a way to make AI systems honest. However, such training risks models learn

safetyarxiv-cs-ai
28 May 2026
Safety

Toward Robust Semi-supervised Regression via Dual-stream Knowledge Distillation

DGX agent

arXiv:2508.14082v3 Announce Type: replace Abstract: Semi-supervised regression (SSR), which aims to predict continuous scores for samples while reducing the reliance on large-scale labeled data, has r

safetyarxiv-cs-lg
28 May 2026
Safety

Towards automated data analysis: A guided framework for LLM-based risk estimation

DGX agent

arXiv:2603.04631v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly integrated into critical decision-making pipelines, a trend that raises the demand for robust and auto

safetyarxiv-cs-ai
28 May 2026
Safety

Trust Me, I'm an Expert: Decoding and Steering Authority Bias in Large Language Models

DGX agent

arXiv:2601.13433v3 Announce Type: replace Abstract: Prior research demonstrates that performance of language models on reasoning tasks can be influenced by suggestions, hints and endorsements. However

safetyarxiv-cs-cl
28 May 2026
Safety

Turning Video Models into Generalist Robot Policies

DGX agent

arXiv:2605.27817v1 Announce Type: cross Abstract: Video generative models have emerged as a promising robotics backbone, capable of generating videos that depict the completion of complex tasks across

safetyarxiv-cs-ai
28 May 2026
Safety

Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models

DGX agent

arXiv:2605.27376v1 Announce Type: cross Abstract: While prompt-based text-to-speech (TTS) models enable natural language-driven speaking style control, they often provide limited fine-grained control

safetyarxiv-cs-ai
28 May 2026
Safety

Unsupervised Identification and Removal of Spurious Correlations During Fine-Tuning

DGX agent

arXiv:2605.27676v1 Announce Type: cross Abstract: Fine-tuning a pretrained language model on a curated dataset can produce spurious correlations between the fine-tuning task and unintended latent fact

safetyarxiv-cs-lg
28 May 2026
Safety

Utility-Aware Multimodal Contrastive Learning for Product Image Generation

DGX agent

arXiv:2605.28733v1 Announce Type: new Abstract: Product images strongly influence consumer decision-making in online marketplaces. Empowered by multimodal contrastive learning, generative AI can outpu

safetyarxiv-cs-ai
28 May 2026
Safety

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning

DGX agent

arXiv:2605.28023v1 Announce Type: cross Abstract: Visual captioning requires models to capture visual content faithfully while minimizing both omission and hallucination. As the dominant paradigm for

safetyarxiv-cs-ai
28 May 2026
Safety

Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs

DGX agent

arXiv:2605.28565v1 Announce Type: cross Abstract: Users of search-augmented LLMs rely on citations as evidence that responses are grounded in real sources, and rarely verify the cited pages themselves

safetyarxiv-cs-ai
28 May 2026
Safety

Visualizing Latent Phase Structures in Locomotion Policies: A Multi-Environment Study with Temporal Feature Extension

DGX agent

arXiv:2605.28186v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) has been shown to achieve high performance on locomotion control tasks in MuJoCo benchmarks such as HalfCheetah, Ant

safetyarxiv-cs-ai
28 May 2026
Safety

VLMs May Not Globally Enhance Human Alignment over LLMs During Natural Reading

DGX agent

arXiv:2605.28818v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly useful computational models of human language processing, but it remains unclear whether vision-la

safetyarxiv-cs-cl
28 May 2026
Safety

What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025

DGX agent

arXiv:2601.07648v2 Announce Type: replace Abstract: As Natural Language Generation (NLG) dominates modern NLP, scalable evaluation remains a critical bottleneck. Consequently, LLM-as-a-judge (LaaJ) ad

safetyarxiv-cs-cl
28 May 2026
Safety

What Frozen VLAs Already Know About Success: A Probing Study of Value-Like Structure in Foundation Robot Policies

DGX agent

arXiv:2605.28527v1 Announce Type: new Abstract: Vision--language--action (VLA) policies are trained to imitate actions; their loss never asks them to estimate reward, progress, or future success. Thei

safetyarxiv-cs-ro
28 May 2026
Safety

When pre-training hurts LoRA fine-tuning: a dynamical analysis via single-index models

DGX agent

arXiv:2602.02855v2 Announce Type: replace Abstract: Pre-training on a source task is usually expected to facilitate fine-tuning on similar downstream problems. In this work, we mathematically show tha

safetyarxiv-cs-lg
28 May 2026
Safety

Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR

DGX agent

arXiv:2605.28295v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) trains reasoning models without labeled trajectories, relying on grouped rollouts to expose the po

safetyarxiv-cs-ai
28 May 2026
Safety

Who Uses AI? Platform Selection and the Measurement of Occupational AI Exposure

DGX agent

arXiv:2605.21743v2 Announce Type: replace Abstract: Conversation logs from AI platforms are increasingly used to measure occupational exposure to artificial intelligence, but the users observed in the

safetyarxiv-cs-ai
28 May 2026
Safety

xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction

DGX agent

arXiv:2503.18893v2 Announce Type: replace Abstract: Long-context Large Language Models (LLMs) enable powerful applications but incur high memory costs due to the key-value states (KV-Cache). Recent st

safetyarxiv-cs-cl
28 May 2026
Safety

A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration

DGX agent

arXiv:2605.26174v1 Announce Type: cross Abstract: Production language-model systems answer a request by partitioning it across an invisible orchestration of worker agents that recompose one integrated

safetyarxiv-cs-ai
27 May 2026
Safety

Adversarial Dual On-Policy Distillation from Expressive Flow-based Teacher

DGX agent

arXiv:2605.27095v1 Announce Type: new Abstract: Learning from demonstrations in embodied control is often cast as behavioral cloning, and recent diffusion or flow-matching policies improve this paradi

safetyarxiv-cs-lg
27 May 2026
Safety

AI-Driven Contribution Evaluation and Conflict Resolution: A Framework & Design for Group Workload Investigation

DGX agent

arXiv:2511.07667v2 Announce Type: replace Abstract: The equitable assessment of individual contribution in teams remains a persistent challenge, where conflict and disparity in workload can result in

safetyarxiv-cs-ai
27 May 2026
Safety

AI-T2I: Aggregating-and-Isolating Cross-Attention to Diffusion Models for Text-to-Image Synthesis

DGX agent

arXiv:2605.25763v2 Announce Type: replace Abstract: Text-to-image synthesis has made significant progress, benefiting from the strong generative capabilities of diffusion models. However, these models

safetyarxiv-cs-cv
27 May 2026
Safety

AirCast-SR: A Foundation Model for Kilometer-Scale Atmospheric Super-Resolution via Latent Consistency Diffusion

DGX agent

arXiv:2605.26130v1 Announce Type: new Abstract: Operational weather prediction at kilometer scales remains computationally prohibitive for traditional numerical weather prediction (NWP) models, limiti

safetyarxiv-cs-lg
27 May 2026
Safety

Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment

DGX agent

arXiv:2511.16870v3 Announce Type: replace Abstract: Enforcing alignment between the internal representations of diffusion or flow-based generative models and those of pretrained self-supervised encode

safetyarxiv-cs-cv
27 May 2026
Safety

AlignEvoSkill: Towards Knowledge-Aware and Task-Aligned Agent Skill Evolution

DGX agent

arXiv:2506.23149v2 Announce Type: replace Abstract: Reusable skills play a key role in improving LLM-based agents, but existing skill-evolution methods often fail to ensure that evolved skills both co

safetyarxiv-cs-cl
27 May 2026
Safety

Aligning Few-Step Generative Models by Amortizing Sample-based Variational Inference

DGX agent

arXiv:2605.26552v1 Announce Type: cross Abstract: Aligning a few-step generative model is challenging, since existing alignment frameworks typically rely on restrictive assumptions: a tractable likeli

safetyarxiv-cs-ai
27 May 2026
Safety

Alignment Makes Language Models Normative, Not Descriptive

DGX agent

arXiv:2603.17218v2 Announce Type: replace-cross Abstract: Post-training alignment optimizes language models to match human preference signals, but this objective is not equivalent to modeling observed

safetyarxiv-cs-ai
27 May 2026
Safety

Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases

DGX agent

arXiv:2605.27355v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we

safetyarxiv-cs-ai
27 May 2026
Safety

Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelines

DGX agent

arXiv:2605.26442v1 Announce Type: cross Abstract: Much of the alignment tuning literature is organized around optimization objectives, while the construction of alignment data is often treated implici

safetyarxiv-cs-ai
27 May 2026
Safety

Annotator Positionality as Signal: Psychometric Weighting for Anti-Autistic Ableism Detection

DGX agent

arXiv:2605.26397v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in decision-making tasks where they can amplify or suppress perspectives, raising concerns in high-

safetyarxiv-cs-ai
27 May 2026
Safety

Aperiodic and Low-Frequency Spectral Bias in Reconstruction based EEG Foundation Models

DGX agent

arXiv:2605.26434v1 Announce Type: cross Abstract: EEG foundation models, pre-trained on large-scale unlabelled EEG data, have emerged as a promising direction towards learning generalizable EEG repres

safetyarxiv-cs-ai
27 May 2026
← Previous
1…155156157158159…260
Next →