AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,809 results
Safety

The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested

DGX agent

arXiv:2605.11496v1 Announce Type: cross Abstract: Recent published evidence from frontier laboratories shows that contemporary AI models can recognise evaluation contexts, latently represent them, and

safetyarxiv-cs-lg
13 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

The new era of SaMD: Why cloud infrastructure is the foundation for digital health in 2026

DGX agent

In the healthcare and life sciences industries, speed saves lives, but meeting regulatory requirements and other administrative burdens often pumps the brakes for manufacturers of software as a medica

safetygoogle-cloud-ai
13 May 2026
Safety

The tractability landscape of diffusion alignment: regularization, rewards, and computational primitives

DGX agent

arXiv:2605.11361v1 Announce Type: new Abstract: Inference-time reward alignment asks how to turn a pre-trained diffusion model with base law p into a sampler that favors a reward r while remaining clo

safetyarxiv-cs-lg
13 May 2026
Safety

there are many agent use cases locked behind 'what if' fears because after we give a tool to an agent, we're relying on the prompt to limit …

DGX agent

there are many agent use cases locked behind 'what if' fears because after we give a tool to an agent, we're relying on the prompt to limit behavior tools like @denieddotdev allow teams to manage capa

safetyyohei-nakajima--x
13 May 2026
Safety

TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignment

DGX agent

arXiv:2605.10983v1 Announce Type: cross Abstract: Reinforcement learning (RL) has shown extraordinary potential in aligning diffusion models to downstream tasks, yet most of them still suffer from sig

safetyarxiv-cs-cv
13 May 2026
Safety

TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning

DGX agent

arXiv:2605.12236v1 Announce Type: cross Abstract: Fine-tuning pre-trained robot policies with reinforcement learning (RL) often inherits the bottlenecks introduced by pre-training with behavioral clon

safetyarxiv-cs-lg
13 May 2026
Safety

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching

DGX agent

arXiv:2605.12288v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) is a widely used RL-free method for aligning language models from pairwise preferences, but it models preferences o

safetyarxiv-cs-cl
13 May 2026
Safety

Towards Fine-Grained Code-Switch Speech Translation with Semantic Space Alignment

DGX agent

arXiv:2511.10670v2 Announce Type: replace Abstract: Code-switching (CS) speech translation (ST) aims to translate speech that alternates between multiple languages into a target language text, posing

safetyarxiv-cs-cl
13 May 2026
Safety

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization

DGX agent

arXiv:2605.11974v1 Announce Type: new Abstract: Large Language Models (LLMs) suffer from order bias, where their performance is affected by the arrangement order of input elements. This unfairness lim

safetyarxiv-cs-lg
13 May 2026
Safety

Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness

DGX agent

arXiv:2503.16072v4 Announce Type: replace-cross Abstract: Toxicity detection has become core safety infrastructure for online moderation, dataset filtering, and deployed language-model systems. Yet mo

safetyarxiv-cs-cl
13 May 2026
Safety

Training Transformers for KV Cache Compressibility

DGX agent

arXiv:2605.05971v2 Announce Type: replace Abstract: Long-context language modeling is increasingly constrained by the Key-Value (KV) cache, whose memory and decode-time access costs scale linearly wit

safetyarxiv-cs-lg
13 May 2026
Safety

Trajectory First: A Curriculum for Discovering Diverse Policies

DGX agent

arXiv:2506.01568v3 Announce Type: replace Abstract: Being able to solve a task in diverse ways makes agents more robust to task variations and less prone to local optima. In this context, constrained

safetyarxiv-cs-lg
13 May 2026
Safety

Transferable Delay-Aware Reinforcement Learning via Implicit Causal Graph Modeling

DGX agent

arXiv:2605.12312v1 Announce Type: new Abstract: Random delays weaken the temporal correspondence between actions and subsequent state feedback, making it difficult for agents to identify the true prop

safetyarxiv-cs-lg
13 May 2026
Safety

Transformer-Based Autonomous Driving Models and Deployment-Oriented Compression: A Survey

DGX agent

arXiv:2304.10891v2 Announce Type: replace-cross Abstract: Transformer-based models are becoming a central paradigm in autonomous driving because they can capture long-range spatial dependencies, multi

safetyarxiv-cs-cv
13 May 2026
Safety

Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates

DGX agent

arXiv:2605.11020v1 Announce Type: new Abstract: Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching the distribution of expert trajectories. Classica

safetyarxiv-cs-lg
13 May 2026
Safety

Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training

DGX agent

arXiv:2605.12380v1 Announce Type: new Abstract: Reinforcement learning is structurally harder than supervised learning because the policy changes the data distribution it learns from. The resulting fr

safetyarxiv-cs-lg
13 May 2026
Safety

UGround: Towards Unified Visual Grounding with Unrolled Transformers

DGX agent

arXiv:2510.03853v4 Announce Type: replace Abstract: We present UGround, a extbf{U}nified visual extbf{Ground}ing paradigm that dynamically selects intermediate layers across extbf{U}nrolled transforme

safetyarxiv-cs-cv
13 May 2026
Safety

Understanding and Preventing Entropy Collapse in RLVR with On-Policy Entropy Flow Optimization

DGX agent

arXiv:2605.11491v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become an effective paradigm for improving the reasoning ability of large language models. How

safetyarxiv-cs-lg
13 May 2026
Safety

Understanding Sample Efficiency in Predictive Coding

DGX agent

arXiv:2605.11911v1 Announce Type: new Abstract: Predictive Coding (PC) is an influential account of cortical learning. Much of recent work has focused on comparing PC to Backpropagation (BP) to find w

safetyarxiv-cs-lg
13 May 2026
Safety

Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO

DGX agent

arXiv:2505.19770v5 Announce Type: replace-cross Abstract: We present a fine-grained theoretical analysis of the performance gap between two-stage reinforcement learning from human feedback~(RLHF) and

safetyarxiv-cs-cl
13 May 2026
Safety

UniFixer: A Universal Reference-Guided Fixer for Diffusion-Based View Synthesis

DGX agent

arXiv:2605.12169v1 Announce Type: new Abstract: With the recent surge of generative models, diffusion-based approaches have become mainstream for view synthesis tasks, either in an explicit depth-warp

safetyarxiv-cs-cv
13 May 2026
Safety

VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference

DGX agent

arXiv:2605.12325v1 Announce Type: new Abstract: Pursuing training-free open-vocabulary semantic segmentation in an efficient and generalizable manner remains challenging due to the deep-seated spatial

safetyarxiv-cs-cv
13 May 2026
Safety

VNDUQE: Information-Theoretic Novelty Detection using Deep Variational Information Bottleneck

DGX agent

arXiv:2605.11551v1 Announce Type: cross Abstract: Detecting out-of-distribution (OOD) samples is critical for safe deployment of neural networks in safety-critical applications. While maximum softmax

safetyarxiv-cs-cv
13 May 2026
Safety

way ahead of its time:

DGX agent

way ahead of its time: Three questions for @sama that the public deserves to better understand: 👉 What is current value of your indirect stake in OpenAI? (Note that you told the senate that you had no

safetygary-marcus--x
13 May 2026
Safety

What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization

DGX agent

arXiv:2605.12021v1 Announce Type: new Abstract: Many image understanding tasks involve identifying what is present and where it appears. However, tasks that address where, such as object discovery, de

safetyarxiv-cs-cv
13 May 2026
Safety

When Does ell_2-Boosting Overfit Benignly? High-Dimensional Risk Asymptotics and the ell_1 Implicit Bias

DGX agent

arXiv:2605.06314v2 Announce Type: replace Abstract: Benign overfitting is well-characterized in ell_2 geometries, but its behavior under the ell_1 implicit bias of greedy ensembles remains challenging

safetyarxiv-cs-lg
13 May 2026
Safety

When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy

DGX agent

arXiv:2605.12112v1 Announce Type: new Abstract: RLHF is widely used to align flow-matching text-to-image models with human preferences, but often leads to severe diversity collapse after fine-tuning.

safetyarxiv-cs-cv
13 May 2026
Safety

When to Ask a Question: Understanding Communication Strategies in Generative AI Tools

DGX agent

arXiv:2605.11240v1 Announce Type: cross Abstract: Generative AI models differ from traditional machine learning tools in that they allow users to provide as much or as little information as they choos

safetyarxiv-cs-lg
13 May 2026
Safety

World Action Models: The Next Frontier in Embodied AI

DGX agent

arXiv:2605.12090v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have achieved strong semantic generalization for embodied policy learning, yet they learn reactive observation-to-

safetyarxiv-cs-cl
13 May 2026
Safety

ZeroIDIR: Zero-Reference Illumination Degradation Image Restoration with Perturbed Consistency Diffusion Models

DGX agent

arXiv:2605.11435v1 Announce Type: new Abstract: In this paper, we propose a zero-reference diffusion-based framework, named ZeroIDIR, for illumination degradation image restoration, which decouples th

safetyarxiv-cs-cv
13 May 2026
Safety

A Cross-Layered Multi-Drone Coordination for Medical Supply Delivery during Disaster Response Management

DGX agent

arXiv:2605.09342v1 Announce Type: cross Abstract: Autonomous drone fleets have immense potential in medical supply delivery during disaster incident response. However, coordinating multiple drones in

safetyarxiv-cs-lg
12 May 2026
Safety

A Scalable Entity-Based Framework for Auditing Bias in LLMs

DGX agent

arXiv:2601.12374v2 Announce Type: replace-cross Abstract: Existing approaches to bias evaluation in large language models (LLMs) trade ecological validity for statistical control, relying either on ar

safetyarxiv-cs-ai
12 May 2026
Safety

A Single Deep Preference-Conditioned Policy for Learning Pareto Coverage Sets

DGX agent

arXiv:2605.08946v1 Announce Type: new Abstract: Preference-conditioned multi-objective reinforcement learning aims to learn a single policy that captures trade-offs across preferences, but under nonli

safetyarxiv-cs-lg
12 May 2026
Safety

A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Models

DGX agent

arXiv:2605.08513v1 Announce Type: cross Abstract: Safety alignment in language models operates through two mechanistically distinct systems: refusal neurons that gate whether harmful knowledge is expr

safetyarxiv-cs-ai
12 May 2026
Safety

A true exponential!

DGX agent

A true exponential! Oy. According to a new paper in The Lancet, the rate of made-up citations in biomedical papers has increased by more than 12x since 2023. https://www.thelancet.com/journals/lancet/

safetygary-marcus--x
12 May 2026
Safety

ActivationReasoning: Logical Reasoning in Latent Activation Spaces

DGX agent

arXiv:2510.18184v3 Announce Type: replace-cross Abstract: Large language models (LLMs) excel at generating fluent text, but their internal reasoning remains opaque and difficult to control. Sparse aut

safetyarxiv-cs-ai
12 May 2026
Safety

Active Tabular Augmentation via Policy-Guided Diffusion Inpainting

DGX agent

arXiv:2605.10315v1 Announce Type: cross Abstract: Generative tabular augmentation is appealing in data-scarce domains, yet the prevailing focus on distributional fidelity does not reliably translate i

safetyarxiv-cs-ai
12 May 2026
Safety

Adaptive Context Matters: Towards Provable Multi-Modality Guidance for Super-Resolution

DGX agent

arXiv:2605.10470v1 Announce Type: new Abstract: Super-resolution (SR) is a severely ill-posed problem with inherent ambiguity, as widely recognized in both empirical and theoretical studies. Although

safetyarxiv-cs-cv
12 May 2026
Safety

Adaptive Data Harvesting for Efficient Neural Network Learning with Universal Constraints

DGX agent

arXiv:2605.09707v1 Announce Type: cross Abstract: Training neural networks to satisfy universal constraints over continuous domains poses unique challenges. Common examples include Lyapunov Neural Net

safetyarxiv-cs-ai
12 May 2026
Safety

Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment

DGX agent

arXiv:2605.09902v1 Announce Type: new Abstract: Adversarial perturbations can mislead Multimodal Large Language Models (MLLMs) recognize a benign image as a specific target object, posing serious risk

safetyarxiv-cs-cv
12 May 2026
Safety

Agent-Omit: Adaptive Context Omission for Efficient LLM Agents

DGX agent

arXiv:2602.04284v2 Announce Type: replace Abstract: Managing agent context (e.g., thought and observation) during multi-turn agent-environment interactions is an emerging strategy to improve agent eff

safetyarxiv-cs-ai
12 May 2026
Safety

Agent-Sentry: Bounding LLM Agents via Execution Provenance

DGX agent

arXiv:2603.22868v2 Announce Type: replace-cross Abstract: Agentic computing systems, while immensely capable, raise serious security, privacy, and safety concerns. A key issue is that the full set of

safetyarxiv-cs-ai
12 May 2026
Safety

AgentReview: Exploring Peer Review Dynamics with LLM Agents

DGX agent

arXiv:2406.12708v3 Announce Type: replace Abstract: Peer review is fundamental to the integrity and advancement of scientific publication. Traditional methods of peer review analyses often rely on exp

safetyarxiv-cs-cl
12 May 2026
Safety

AI-Care: A Conversational Agentic System for Task Coordination in Alzheimer's Disease Care

DGX agent

arXiv:2605.08480v1 Announce Type: new Abstract: Individuals with Alzheimer's disease (AD) and Alzheimer's disease-related dementia (ADRD) experience memory and thinking changes that impact their abili

safetyarxiv-cs-ai
12 May 2026
Safety

AIPO: : Learning to Reason from Active Interaction

DGX agent

arXiv:2605.08401v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have demonstrated remarkable reasoning capabilities, largely stimulated by Reinforcement Learning with

safetyarxiv-cs-ai
12 May 2026
Safety

ALAM: Algebraically Consistent Latent Transitions for Vision-Language-Action Models

DGX agent

arXiv:2605.10819v1 Announce Type: cross Abstract: Vision-language-action (VLA) models remain constrained by the scarcity of action-labeled robot data, whereas action-free videos provide abundant evide

safetyarxiv-cs-ai
12 May 2026
Safety

Align and Shine: Building High-Quality Sentence-Aligned Corpora for Multilingual Text Simplification

DGX agent

arXiv:2605.09476v1 Announce Type: cross Abstract: Text simplification plays a crucial role in improving the accessibility and comprehensibility of written information for diverse audiences, including

safetyarxiv-cs-ai
12 May 2026
Safety

Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis

DGX agent

arXiv:2605.10415v1 Announce Type: new Abstract: Large language models for subjectivity analysis are typically trained with aggregated labels, which compress variations in human judgment into a single

safetyarxiv-cs-cl
12 May 2026
← Previous
1…185186187188189…267
Next →