AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,487 results
28 May 2026

Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction

SafetyDGX agent

arXiv:2605.27878v1 Announce Type: new Abstract: Large language models produce fluent fiction, yet their creative output is widely seen as flat. We ask where this quality originates in the training and

NCSAM Noise-Compensated Sharpness-Aware Minimization for Noisy Label Learning

SafetyDGX agent

arXiv:2601.19947v2 Announce Type: replace-cross Abstract: Learning from Noisy Labels (LNL) remains a fundamental challenge in deep learning because real-world datasets often contain corrupted annotati

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models

SafetyDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2603.01766v2 Announce Type: replace Abstract: Despite the rapid progress of vision-language-action (VLA) models, the prevailing practice of predicting action chunks as discrete waypoints remains

Off-Policy Learning to Reason Works Because It Is More Pessimistic Than You Think

SafetyDGX agent

arXiv:2605.28150v1 Announce Type: new Abstract: Large scale reinforcement learning has become a central tool for improving reasoning in large language models. At this scale, generation is often lagged

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning

SafetyDGX agent

arXiv:2604.18530v2 Announce Type: replace Abstract: Recent advancements in Reinforcement Learning with Verifiable Rewards (RLVR) have significantly improved Large Language Model (LLM) reasoning, yet m

@OpenAI Foundation just launched with a 25B commitment and equity in OpenAI valued at ~130B. Sounds like a philanthropy giant. But look cl…

SafetyDGX agent

@OpenAI Foundation just launched with a 25B commitment and equity in OpenAI valued at ~130B. Sounds like a philanthropy giant. But look closer: the big money goes to programs they control internally (

Operational AI Deployment Assurance: Governance-State Orchestration Under Threshold-Sensitive Deployment Conditions -- A Governance Framework for High-Stakes AI Systems

SafetyDGX agent

arXiv:2605.27827v1 Announce Type: new Abstract: AI governance frameworks increasingly emphasize fairness, transparency, accountability, and lifecycle risk management in high-stakes domains. However, m

Optimal Data Acquisition for Reinforcement Learning: A Large Deviations Perspective

SafetyDGX agent

arXiv:2605.28675v1 Announce Type: new Abstract: Data acquisition efficiency is a central challenge in deploying reinforcement learning in business and healthcare operations, where interactions are cos

Picid: A Modular Evaluation Infrastructure for Reproducible PHM Across Tasks and Domains

SafetyDGX agent

arXiv:2605.28345v1 Announce Type: new Abstract: Progress in Prognostics and Health Management (PHM) is hindered by the lack of standardized and reusable evaluation practices across tasks, datasets, an

Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents

Model ReleasesDGX agent

arXiv:2605.28201v1 Announce Type: new Abstract: Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into ext

Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning

SafetyDGX agent

arXiv:2602.01745v2 Announce Type: replace-cross Abstract: Token-level reweighting is a simple yet effective mechanism for controlling supervised fine-tuning, but common indicators are largely one-dime

Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups

SafetyDGX agent

arXiv:2510.06974v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in user-facing applications, raising concerns that they may reflect and amplify social biases

ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation

SafetyDGX agent

arXiv:2605.28293v1 Announce Type: cross Abstract: Proactive Recommender Systems (PRSs) aim to guide user preference shift toward target items by generating paths of intermediate recommendations. Reinf

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity

SafetyDGX agent

arXiv:2602.15894v2 Announce Type: replace Abstract: In many large language model (LLM) alignment applications, users expect not only high-quality outputs but also substantial diversity. However, exist

RE-TRIANGLE: Does TRIANGLE Enable Multimodal Alignment Beyond Cosine Similarity in Retrieval?

SafetyDGX agent

arXiv:2605.27436v1 Announce Type: cross Abstract: Multimodal alignment is critical for bridging the semantic gap in information retrieval. However, traditional pairwise strategies introduce a geometri

Reasoning Matters: Mitigate Hallucination in Multimodal Large Reasoning Models via Reasoning-Conditioned Preference Optimization

SafetyDGX agent

arXiv:2605.27906v1 Announce Type: new Abstract: Multimodal Large Reasoning Models introduce the reasoning paradigm, demonstrating strong capabilities on complex vision-language tasks. However, they st

Refining Multidimensional Video Reward Models via Disentangled Influence Functions

SafetyDGX agent

arXiv:2605.28203v1 Announce Type: new Abstract: As Text-to-Video (T2V) generation models continue to evolve, the complexity of video evaluation necessitates a fine-grained assessment across various ax

Restoring the Sweet Spot: Pass-Rate Weighted Self-Distillation for LLM Reasoning

SafetyDGX agent

arXiv:2605.27765v1 Announce Type: cross Abstract: Self-Distillation Policy Optimization (SDPO) provides dense token-level credit assignment for reinforcement learning with large language models by lev

Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?

SafetyDGX agent

arXiv:2605.27881v1 Announce Type: new Abstract: Search agents powered by large language models can autonomously decompose queries, retrieve information, and synthesize answers through multi-step reaso

Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure

SafetyDGX agent

arXiv:2605.27996v1 Announce Type: new Abstract: Single-axis mitigations of reward-model biases (e.g., reducing proxy reliance on length, sycophancy, or style) can rotate optimization pressure onto cor

Reward Transfer from Inverse Reinforcement Learning: A Coupled Minimax Approach

SafetyDGX agent

arXiv:2605.27834v1 Announce Type: new Abstract: We study the transfer of rewards learned using inverse reinforcement learning from expert demonstrations in one environment to reinforcement learning in

Right now, 20 British MPs are deciding which bills they'll introduce to Parliament. Sir Stephen Fry just asked them to bring forward our bil…

SafetyDGX agent

Right now, 20 British MPs are deciding which bills they'll introduce to Parliament. Sir Stephen Fry just asked them to bring forward our bill to ban superintelligence! Grateful to receive this strong

ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains

SafetyDGX agent

arXiv:2605.28014v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) improves the reasoning performance of large language models (LLMs) by providing dense token-level supervision for on-

Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models

SafetyDGX agent

arXiv:2605.28306v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models have emerged as a dominant paradigm for efficient LLM scaling, yet adapting them to non-English downstream tasks remai

SA4Depth: Consistent Pose-Depth Scale Alignment for Self-Supervised Monocular Depth Estimation

SafetyDGX agent

arXiv:2605.28477v1 Announce Type: new Abstract: Self-supervised depth estimation from monocular sequences relies on the joint learning of a depth and a pose network. Despite abundant research done to

SCALE-COMM: Shared, Contrastively-Aligned Latent Embeddings for MARL Communication

SafetyDGX agent

arXiv:2605.27532v1 Announce Type: new Abstract: Emergent communication enables partially observant Autonomous Mobile Robots (AMRs) to coordinate effectively in decentralized multi-agent reinforcement

SEMAGIC: Learning Semantically Consistent Deformable 3D Representations from In-the-Wild Images

SafetyDGX agent

arXiv:2605.27938v1 Announce Type: new Abstract: Learning deformable 3D object models from single-view in-the-wild images has enabled impressive 3D shape reconstruction without supervision. However, it

Semiparametrically Efficient Inference for Kernel Measures of Noise Heterogeneity

SafetyDGX agent

arXiv:2605.27526v1 Announce Type: cross Abstract: We develop semiparametrically efficient inference for kernel measures of noise heterogeneity in additive noise models. In many applications, the regre

Sense Representations Are Inducible Interfaces

SafetyDGX agent

arXiv:2605.28669v1 Announce Type: cross Abstract: Sense representations (explicit, per-token meaning decompositions) are useful for disambiguation, steering, and cross-lingual alignment, but existing

Singular Vectors of Attention Heads Align with Features

SafetyDGX agent

arXiv:2602.13524v2 Announce Type: replace-cross Abstract: Identifying feature representations in language models is a central task in mechanistic interpretability. Several recent studies have made the

Skill-Conditioned Gated Self-Distillation for LLM Reasoning

SafetyDGX agent

arXiv:2605.28791v1 Announce Type: cross Abstract: On-policy self-distillation (SD) improves LLM reasoning by using teacher-side privileged information (PI) to turn sparse verifier outcomes into dense

SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit Assignment

SafetyDGX agent

arXiv:2605.27899v1 Announce Type: new Abstract: Structured skill prompts improve exploration in long-horizon agentic reinforcement learning (RL). Skill-augmented RL methods retain external skills at i

Smaller, Younger, and More Impactful: How AI-Assisted Writing Transforms Research Teams

SafetyDGX agent

arXiv:2605.27404v1 Announce Type: cross Abstract: The era of Big Science has long been defined by increasingly large and specialized research teams pushing the frontiers of knowledge. However, recent

Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards

SafetyDGX agent

arXiv:2605.28561v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has improved language models in domains such as mathematics and code, where correctness can be che

SPAR: Support-Preserving Action Rectification

SafetyDGX agent

arXiv:2605.27877v1 Announce Type: cross Abstract: Offline policy improvement faces an inherent conflict between maximizing value and fitting the data distribution. While in-sample weighted regression

SPRINT: Efficient Spectral Priors for Humanoid Athletic Sprints

SafetyDGX agent

arXiv:2605.28549v1 Announce Type: cross Abstract: The pursuit of humanoid athletic sprints is hindered by a scarcity of humanoid-viable kinematic reference data and the inability of existing framework

STARS: Spike Tail-Aware Relational Synthesis for ANN-to-SNN Data-Free Knowledge Distillation

SafetyDGX agent

arXiv:2605.27409v1 Announce Type: cross Abstract: SNNs promise energy-efficient and low-latency inference, but their performance still trails that of ANNs. ANN-to-SNN knowledge distillation helps narr

Structure-Guided Visual Perturbation Neutralization for LVLMs

SafetyDGX agent

arXiv:2605.27927v1 Announce Type: new Abstract: Image inputs enable Large Vision Language Models (LVLMs) to perceive fine-grained visual information, but also introduce a pixel-level attack surface th

Structured Agent Distillation for Large Language Model

SafetyDGX agent

arXiv:2505.13820v5 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strong capabilities as decision-making agents by interleaving reasoning and actions, as seen in ReAct-sty

Supervised Distributional Reduction via Optimal Transport and Dependence Maximization

SafetyDGX agent

arXiv:2605.27619v1 Announce Type: cross Abstract: Learning representations that capture both intrinsic data geometry and target-relevant structure remains a fundamental challenge, particularly in sett

Supervised Semantic Differential for Cross-Cultural Concept Analysis: A Case Study of Human Affect

SafetyDGX agent

arXiv:2605.28225v1 Announce Type: new Abstract: Cross-cultural comparison of psychological meaning requires methods that go beyond word-level translation and examine how semantic dimensions are organi

SYNAPSE: Neuro-Symbolic Visual Thought-to-Text Decoding via Topological Semantic Denoising

SafetyDGX agent

arXiv:2605.27790v1 Announce Type: new Abstract: Recent advances in large language models have accelerated open-vocabulary EEG-to-imagined-text decoding, where non-invasive neural activity recorded dur

Teacher-Student Representational Alignment for Reinforcement Learning-Driven Imitation Learning

SafetyDGX agent

arXiv:2605.28372v1 Announce Type: new Abstract: Imitation learning (IL) from a state-based reinforcement learning (RL) policy is a common approach to overcome the curse of dimensionality in complex an

Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers

SafetyDGX agent

arXiv:2605.27686v1 Announce Type: cross Abstract: Transformers process images and videos by flattening space and time into long token sequences. While attention and KV caching preserve past features,

Test-Time Collective Action: Proxy-Based Perturbations for Correcting Algorithmic Harms

SafetyDGX agent

arXiv:2605.27689v1 Announce Type: new Abstract: When machine learning systems under-perform for particular subgroups, affected users typically have no way to correct these disparities without relying

The AI numbers are starting to look very ugly. Even under 'best case' assumptions, FT's own data shows Microsoft AI ROI at -9%, Google at -1…

SafetyDGX agent

The AI numbers are starting to look very ugly. Even under 'best case' assumptions, FT's own data shows Microsoft AI ROI at -9%, Google at -15%, Meta at -28%, Oracle at -35%. Only Amazon barely comes o

The Attentional White Bear Effect in Transformer Language Models

SafetyDGX agent

arXiv:2605.28639v1 Announce Type: cross Abstract: Instruction-based suppression is widely used to prevent language models from generating prohibited content, yet it remains unclear whether suppression

The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

Model ReleasesDGX agent

arXiv:2605.27901v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models. Howeve

The Illusion of Opting in AI-Mediated Consequential Decisions

SafetyDGX agent

arXiv:2605.28210v1 Announce Type: new Abstract: Drawing on Ullmann-Margalit's concept of opting (transformative, irrevocable, and shadowed by foreclosed alternatives), we show that current AI systems

The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes

SafetyDGX agent

arXiv:2602.15515v2 Announce Type: replace-cross Abstract: Training against white-box deception detectors has been proposed as a way to make AI systems honest. However, such training risks models learn

The trial has concluded, the facts have been confirmed. As the prosecuting counsel put it, Digwa used his “trump card” by alleging he had be…

SafetyDGX agent

The trial has concluded, the facts have been confirmed. As the prosecuting counsel put it, Digwa used his “trump card” by alleging he had been the victim of racist abuse when police officers arrived.

this sh*t isn’t even funny anymore. it’s a trillion dollar embarrassment.

SafetyDGX agent

Gary Marcus critiques the current state of AI development as wasteful and problematic, arguing that the industry's trillion-dollar investment represents a significant failure or misallocation of resou

tokenmaxxing is officially over

SafetyDGX agent

tokenmaxxing is officially over Sources: Amazon has shut down an internal leaderboard that tracked employees' use of AI tools after workers tried to boost their scores with needless tasks (@rafeuddin_

Toward Robust Semi-supervised Regression via Dual-stream Knowledge Distillation

SafetyDGX agent

arXiv:2508.14082v3 Announce Type: replace Abstract: Semi-supervised regression (SSR), which aims to predict continuous scores for samples while reducing the reliance on large-scale labeled data, has r

Towards automated data analysis: A guided framework for LLM-based risk estimation

SafetyDGX agent

arXiv:2603.04631v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly integrated into critical decision-making pipelines, a trend that raises the demand for robust and auto

Trust Me, I'm an Expert: Decoding and Steering Authority Bias in Large Language Models

SafetyDGX agent

arXiv:2601.13433v3 Announce Type: replace Abstract: Prior research demonstrates that performance of language models on reasoning tasks can be influenced by suggestions, hints and endorsements. However

Turning Video Models into Generalist Robot Policies

SafetyDGX agent

arXiv:2605.27817v1 Announce Type: cross Abstract: Video generative models have emerged as a promising robotics backbone, capable of generating videos that depict the completion of complex tasks across

turns out i wasn’t wrong:

SafetyDGX agent

Gary Marcus reflects on a past prediction or stance that he held, asserting its correctness in retrospect, likely addressing criticisms or skepticism he previously faced regarding AI, cognitive scienc

Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models

SafetyDGX agent

arXiv:2605.27376v1 Announce Type: cross Abstract: While prompt-based text-to-speech (TTS) models enable natural language-driven speaking style control, they often provide limited fine-grained control

Unsupervised Identification and Removal of Spurious Correlations During Fine-Tuning

SafetyDGX agent

arXiv:2605.27676v1 Announce Type: cross Abstract: Fine-tuning a pretrained language model on a curated dataset can produce spurious correlations between the fine-tuning task and unintended latent fact

← Previous
1…141142143144145…242
Next →