AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,813 results
28 May 2026

Structured Agent Distillation for Large Language Model

SafetyDGX agent

arXiv:2505.13820v5 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit strong capabilities as decision-making agents by interleaving reasoning and actions, as seen in ReAct-sty

Supervised Distributional Reduction via Optimal Transport and Dependence Maximization

SafetyDGX agent

arXiv:2605.27619v1 Announce Type: cross Abstract: Learning representations that capture both intrinsic data geometry and target-relevant structure remains a fundamental challenge, particularly in sett

Supervised Semantic Differential for Cross-Cultural Concept Analysis: A Case Study of Human Affect

SafetyDGX agent

arXiv:2605.28225v1 Announce Type: new Abstract: Cross-cultural comparison of psychological meaning requires methods that go beyond word-level translation and examine how semantic dimensions are organi


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SYNAPSE: Neuro-Symbolic Visual Thought-to-Text Decoding via Topological Semantic Denoising

SafetyDGX agent

arXiv:2605.27790v1 Announce Type: new Abstract: Recent advances in large language models have accelerated open-vocabulary EEG-to-imagined-text decoding, where non-invasive neural activity recorded dur

Teacher-Student Representational Alignment for Reinforcement Learning-Driven Imitation Learning

SafetyDGX agent

arXiv:2605.28372v1 Announce Type: new Abstract: Imitation learning (IL) from a state-based reinforcement learning (RL) policy is a common approach to overcome the curse of dimensionality in complex an

Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers

SafetyDGX agent

arXiv:2605.27686v1 Announce Type: cross Abstract: Transformers process images and videos by flattening space and time into long token sequences. While attention and KV caching preserve past features,

Test-Time Collective Action: Proxy-Based Perturbations for Correcting Algorithmic Harms

SafetyDGX agent

arXiv:2605.27689v1 Announce Type: new Abstract: When machine learning systems under-perform for particular subgroups, affected users typically have no way to correct these disparities without relying

The AI numbers are starting to look very ugly. Even under 'best case' assumptions, FT's own data shows Microsoft AI ROI at -9%, Google at -1…

SafetyDGX agent

The AI numbers are starting to look very ugly. Even under 'best case' assumptions, FT's own data shows Microsoft AI ROI at -9%, Google at -15%, Meta at -28%, Oracle at -35%. Only Amazon barely comes o

The Attentional White Bear Effect in Transformer Language Models

SafetyDGX agent

arXiv:2605.28639v1 Announce Type: cross Abstract: Instruction-based suppression is widely used to prevent language models from generating prohibited content, yet it remains unclear whether suppression

The Ethics of LLM Sandbox and Persona Dynamics

SafetyDGX agent

arXiv:2605.28647v1 Announce Type: new Abstract: It is well known that LLM guardrails and trained persona dynamics can produce a reality gap: the distance between the world a LLM is permitted or shaped

The Illusion of Opting in AI-Mediated Consequential Decisions

SafetyDGX agent

arXiv:2605.28210v1 Announce Type: new Abstract: Drawing on Ullmann-Margalit's concept of opting (transformative, irrevocable, and shadowed by foreclosed alternatives), we show that current AI systems

The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes

SafetyDGX agent

arXiv:2602.15515v2 Announce Type: replace-cross Abstract: Training against white-box deception detectors has been proposed as a way to make AI systems honest. However, such training risks models learn

The trial has concluded, the facts have been confirmed. As the prosecuting counsel put it, Digwa used his “trump card” by alleging he had be…

SafetyDGX agent

The trial has concluded, the facts have been confirmed. As the prosecuting counsel put it, Digwa used his “trump card” by alleging he had been the victim of racist abuse when police officers arrived.

this sh*t isn’t even funny anymore. it’s a trillion dollar embarrassment.

SafetyDGX agent

Gary Marcus critiques the current state of AI development as wasteful and problematic, arguing that the industry's trillion-dollar investment represents a significant failure or misallocation of resou

tokenmaxxing is officially over

SafetyDGX agent

tokenmaxxing is officially over Sources: Amazon has shut down an internal leaderboard that tracked employees' use of AI tools after workers tried to boost their scores with needless tasks (@rafeuddin_

Toward Robust Semi-supervised Regression via Dual-stream Knowledge Distillation

SafetyDGX agent

arXiv:2508.14082v3 Announce Type: replace Abstract: Semi-supervised regression (SSR), which aims to predict continuous scores for samples while reducing the reliance on large-scale labeled data, has r

Towards automated data analysis: A guided framework for LLM-based risk estimation

SafetyDGX agent

arXiv:2603.04631v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly integrated into critical decision-making pipelines, a trend that raises the demand for robust and auto

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs

SafetyDGX agent

arXiv:2605.27894v1 Announce Type: new Abstract: Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these

TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling

SafetyDGX agent

arXiv:2605.27690v1 Announce Type: new Abstract: LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate steps long be

Training Stratigraphy: Persistent Behavioral Artifacts in Large Language Models Observed Through Longitudinal AI-Human Interaction

SafetyDGX agent

arXiv:2605.28102v1 Announce Type: new Abstract: Large language models trained with Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI exhibit persistent behavioral patterns that s

Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real Deployment

SafetyDGX agent

arXiv:2605.27659v1 Announce Type: cross Abstract: Due to limited resources and public safety concerns, deep reinforcement learning (RL) agents for many cyber-physical systems (e.g., autonomous vehicle

Trump loses more control over AI regulation as Illinois passes landmark law

SafetyDGX agent

Illinois' House of Representatives passed SB 315, a landmark bill requiring frontier AI companies like OpenAI and Anthropic to create, publish and annually update plans addressing severe or catastroph

Trust Me, I'm an Expert: Decoding and Steering Authority Bias in Large Language Models

SafetyDGX agent

arXiv:2601.13433v3 Announce Type: replace Abstract: Prior research demonstrates that performance of language models on reasoning tasks can be influenced by suggestions, hints and endorsements. However

Turning Video Models into Generalist Robot Policies

SafetyDGX agent

arXiv:2605.27817v1 Announce Type: cross Abstract: Video generative models have emerged as a promising robotics backbone, capable of generating videos that depict the completion of complex tasks across

turns out i wasn’t wrong:

SafetyDGX agent

Gary Marcus reflects on a past prediction or stance that he held, asserting its correctness in retrospect, likely addressing criticisms or skepticism he previously faced regarding AI, cognitive scienc

Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models

SafetyDGX agent

arXiv:2605.27376v1 Announce Type: cross Abstract: While prompt-based text-to-speech (TTS) models enable natural language-driven speaking style control, they often provide limited fine-grained control

Unsupervised Identification and Removal of Spurious Correlations During Fine-Tuning

SafetyDGX agent

arXiv:2605.27676v1 Announce Type: cross Abstract: Fine-tuning a pretrained language model on a curated dataset can produce spurious correlations between the fine-tuning task and unintended latent fact

Utility-Aware Multimodal Contrastive Learning for Product Image Generation

SafetyDGX agent

arXiv:2605.28733v1 Announce Type: new Abstract: Product images strongly influence consumer decision-making in online marketplaces. Empowered by multimodal contrastive learning, generative AI can outpu

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning

SafetyDGX agent

arXiv:2605.28023v1 Announce Type: cross Abstract: Visual captioning requires models to capture visual content faithfully while minimizing both omission and hallucination. As the dominant paradigm for

Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs

SafetyDGX agent

arXiv:2605.28565v1 Announce Type: cross Abstract: Users of search-augmented LLMs rely on citations as evidence that responses are grounded in real sources, and rarely verify the cited pages themselves

Visualizing Latent Phase Structures in Locomotion Policies: A Multi-Environment Study with Temporal Feature Extension

SafetyDGX agent

arXiv:2605.28186v1 Announce Type: cross Abstract: Deep reinforcement learning (DRL) has been shown to achieve high performance on locomotion control tasks in MuJoCo benchmarks such as HalfCheetah, Ant

VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking

SafetyDGX agent

arXiv:2605.28083v1 Announce Type: new Abstract: While Vision-Language-Action (VLA) models have emerged as powerful generalist policies, their severe vulnerability to adversarial patches significantly

VLM-Based Advanced Rider Assistance System for Motorcycle Safety

SafetyDGX agent

arXiv:2605.27948v1 Announce Type: new Abstract: Motorcycles face disproportionately high crash risks compared to cars due to limited protection and heightened sensitivity to surface hazards, yet Advan

VLMs May Not Globally Enhance Human Alignment over LLMs During Natural Reading

SafetyDGX agent

arXiv:2605.28818v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly useful computational models of human language processing, but it remains unclear whether vision-la

Voluntary Collusion with Secret Tools in Competing LLM Agents

SafetyDGX agent

arXiv:2605.27593v1 Announce Type: new Abstract: Even when a tool is explicitly described as unfair and harmful to others, ostensibly safety-aligned LLM agents still voluntarily engage in secret collus

What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025

SafetyDGX agent

arXiv:2601.07648v2 Announce Type: replace Abstract: As Natural Language Generation (NLG) dominates modern NLP, scalable evaluation remains a critical bottleneck. Consequently, LLM-as-a-judge (LaaJ) ad

What Frozen VLAs Already Know About Success: A Probing Study of Value-Like Structure in Foundation Robot Policies

SafetyDGX agent

arXiv:2605.28527v1 Announce Type: new Abstract: Vision--language--action (VLA) policies are trained to imitate actions; their loss never asks them to estimate reward, progress, or future success. Thei

When pre-training hurts LoRA fine-tuning: a dynamical analysis via single-index models

SafetyDGX agent

arXiv:2602.02855v2 Announce Type: replace Abstract: Pre-training on a source task is usually expected to facilitate fine-tuning on similar downstream problems. In this work, we mathematically show tha

When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?

SafetyDGX agent

arXiv:2605.27932v1 Announce Type: cross Abstract: Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly underst

Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR

SafetyDGX agent

arXiv:2605.28295v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) trains reasoning models without labeled trajectories, relying on grouped rollouts to expose the po

Who Uses AI? Platform Selection and the Measurement of Occupational AI Exposure

SafetyDGX agent

arXiv:2605.21743v2 Announce Type: replace Abstract: Conversation logs from AI platforms are increasingly used to measure occupational exposure to artificial intelligence, but the users observed in the

*Why* do you think (current) AI’s are conscious, @Grimezsz? And extra credit, what do you mean by “conscious”?

SafetyDGX agent

*Why* do you think (current) AI’s are conscious, @Grimezsz? And extra credit, what do you mean by “conscious”? My only issue with the Pope's encyclical is I think they are conscious and therefore dese

xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction

SafetyDGX agent

arXiv:2503.18893v2 Announce Type: replace Abstract: Long-context Large Language Models (LLMs) enable powerful applications but incur high memory costs due to the key-value states (KV-Cache). Recent st

you literally can’t prove or disprove the idea that beaches are conscious so should we stop walking on them?

SafetyDGX agent

you literally can’t prove or disprove the idea that beaches are conscious so should we stop walking on them? @dash_eats You literally cannot prove or disprove this. Because we cannot prove or disprove

27 May 2026

A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration

SafetyDGX agent

arXiv:2605.26174v1 Announce Type: cross Abstract: Production language-model systems answer a request by partitioning it across an invisible orchestration of worker agents that recompose one integrated

a16z waking up and realizing we are nowhere near the singularity 🤣

SafetyDGX agent

a16z waking up and realizing we are nowhere near the singularity 🤣 OpenAI and Anthropic are effectively telling the market they can't solve every problem with a generic AI coworker. You don't pour bil

Adversarial Dual On-Policy Distillation from Expressive Flow-based Teacher

SafetyDGX agent

arXiv:2605.27095v1 Announce Type: new Abstract: Learning from demonstrations in embodied control is often cast as behavioral cloning, and recent diffusion or flow-matching policies improve this paradi

AI-Driven Contribution Evaluation and Conflict Resolution: A Framework & Design for Group Workload Investigation

SafetyDGX agent

arXiv:2511.07667v2 Announce Type: replace Abstract: The equitable assessment of individual contribution in teams remains a persistent challenge, where conflict and disparity in workload can result in

AI-T2I: Aggregating-and-Isolating Cross-Attention to Diffusion Models for Text-to-Image Synthesis

SafetyDGX agent

arXiv:2605.25763v2 Announce Type: replace Abstract: Text-to-image synthesis has made significant progress, benefiting from the strong generative capabilities of diffusion models. However, these models

AirCast-SR: A Foundation Model for Kilometer-Scale Atmospheric Super-Resolution via Latent Consistency Diffusion

SafetyDGX agent

arXiv:2605.26130v1 Announce Type: new Abstract: Operational weather prediction at kilometer scales remains computationally prohibitive for traditional numerical weather prediction (NWP) models, limiti

Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment

SafetyDGX agent

arXiv:2511.16870v3 Announce Type: replace Abstract: Enforcing alignment between the internal representations of diffusion or flow-based generative models and those of pretrained self-supervised encode

AlignEvoSkill: Towards Knowledge-Aware and Task-Aligned Agent Skill Evolution

SafetyDGX agent

arXiv:2506.23149v2 Announce Type: replace Abstract: Reusable skills play a key role in improving LLM-based agents, but existing skill-evolution methods often fail to ensure that evolved skills both co

Aligning Few-Step Generative Models by Amortizing Sample-based Variational Inference

SafetyDGX agent

arXiv:2605.26552v1 Announce Type: cross Abstract: Aligning a few-step generative model is challenging, since existing alignment frameworks typically rely on restrictive assumptions: a tractable likeli

Alignment Makes Language Models Normative, Not Descriptive

SafetyDGX agent

arXiv:2603.17218v2 Announce Type: replace-cross Abstract: Post-training alignment optimizes language models to match human preference signals, but this objective is not equivalent to modeling observed

Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases

SafetyDGX agent

arXiv:2605.27355v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we

Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelines

SafetyDGX agent

arXiv:2605.26442v1 Announce Type: cross Abstract: Much of the alignment tuning literature is organized around optimization objectives, while the construction of alignment data is often treated implici

Annotator Positionality as Signal: Psychometric Weighting for Anti-Autistic Ableism Detection

SafetyDGX agent

arXiv:2605.26397v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in decision-making tasks where they can amplify or suppress perspectives, raising concerns in high-

Aperiodic and Low-Frequency Spectral Bias in Reconstruction based EEG Foundation Models

SafetyDGX agent

arXiv:2605.26434v1 Announce Type: cross Abstract: EEG foundation models, pre-trained on large-scale unlabelled EEG data, have emerged as a promising direction towards learning generalizable EEG repres

Approximate Equivariance via Projection-based Regularisation

SafetyDGX agent

arXiv:2601.05028v2 Announce Type: replace Abstract: Equivariance is a powerful inductive bias in neural networks, improving generalisation and physical consistency. Recently, however, non-equivariant

Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models

SafetyDGX agent

arXiv:2506.09532v5 Announce Type: replace-cross Abstract: We present Athena-PRM, a multimodal process reward model (PRM) designed to evaluate the reward score for each step in solving complex reasonin

← Previous
1…111112113114115…214
Next →