AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,485 results
Safety

Enabling clinical use of foundation models for computational pathology

DGX agent

arXiv:2602.22347v2 Announce Type: replace Abstract: Foundation models for computational pathology are expected to facilitate the development of high-performing, generalisable deep learning systems. Ho

safetyarxiv-cs-cv
13 May 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Enabling Performant and Flexible Model-Internal Observability for LLM Inference

DGX agent

arXiv:2605.11093v1 Announce Type: new Abstract: Today's inference-time workloads increasingly depend on timely access to a model's internal states. We present DMI-Lib, a high-speed deep model inspecto

safetyarxiv-cs-lg
13 May 2026
Safety

Enforcing Constraints in Generative Sampling via Adaptive Correction Scheduling

DGX agent

arXiv:2605.11214v1 Announce Type: new Abstract: Hard constraints in generative sampling are typically enforced by projection, applied either once at the end of sampling or after every update. This bin

safetyarxiv-cs-lg
13 May 2026
Safety

Enhancing Multilingual Counterfactual Generation through Alignment-as-Preference Optimization

DGX agent

arXiv:2605.11632v1 Announce Type: new Abstract: Self-generated counterfactual explanations (SCEs) are minimally modified inputs (minimality) generated by large language models (LLMs) that flip their o

safetyarxiv-cs-cl
13 May 2026
Safety

Enhancing Target-Guided Proactive Dialogue Systems via Conversational Scenario Modeling and Intent-Keyword Bridging

DGX agent

arXiv:2605.11964v1 Announce Type: new Abstract: A target-guided proactive dialogue system aims to steer conversations proactively toward pre-defined targets, such as designated keywords or specific to

safetyarxiv-cs-cl
13 May 2026
Safety

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control

DGX agent

arXiv:2605.11775v1 Announce Type: cross Abstract: Policy entropy has emerged as a fundamental measure for understanding and controlling exploration in reinforcement learning with verifiable rewards (R

safetyarxiv-cs-cl
13 May 2026
Safety

Epistemic Uncertainty for Test-Time Discovery

DGX agent

arXiv:2605.11328v1 Announce Type: new Abstract: Automated scientific discovery using large language models relies on identifying genuinely novel solutions. Standard reinforcement learning penalizes hi

safetyarxiv-cs-lg
13 May 2026
Safety

Events as Triggers for Behavioral Diversity in Multi-Agent Reinforcement Learning

DGX agent

arXiv:2605.12388v1 Announce Type: cross Abstract: Effective multi-agent cooperation requires agents to adopt diverse behaviors as task conditions evolve-and to do so at the right moment. Yet, current

safetyarxiv-cs-lg
13 May 2026
Safety

EvoNav: Evolutionary Reward Function Design for Robot Navigation with Large Language Models

DGX agent

arXiv:2605.11859v1 Announce Type: new Abstract: Robot navigation is a crucial task with applications to social robots in dynamic human environments. While Reinforcement Learning (RL) has shown great p

safetyarxiv-cs-ro
13 May 2026
Safety

Expected Batch Optimal Transport Plans and Consequences for Flow Matching

DGX agent

arXiv:2605.12174v1 Announce Type: new Abstract: Solving optimal transport (OT) on random minibatches is a common surrogate for exact OT in large-scale learning. In flow matching (FM), this surrogate i

safetyarxiv-cs-lg
13 May 2026
Safety

Fair Conformal Classification via Learning Representation-Based Groups

DGX agent

arXiv:2605.12195v1 Announce Type: new Abstract: Conformal prediction methods provide statistically rigorous marginal coverage guarantees for machine learning models, but such guarantees fail to accoun

safetyarxiv-cs-lg
13 May 2026
Safety

Fed-BAC: Federated Bandit-Guided Additive Clustering in Hierarchical Federated Learning

DGX agent

arXiv:2605.11815v1 Announce Type: new Abstract: Hierarchical federated learning (HFL) leverages edge servers for partial aggregation in edge computing. Yet existing FL methods lack mechanisms for join

safetyarxiv-cs-lg
13 May 2026
Safety

FedOUI: OUI-Guided Client Weighting for Federated Aggregation

DGX agent

arXiv:2605.11571v1 Announce Type: new Abstract: Federated learning usually aggregates client updates using dataset size or gradient-level criteria, while overlooking internal signals about how each cl

safetyarxiv-cs-lg
13 May 2026
Safety

FedSurrogate: Backdoor Defense in Federated Learning via Layer Criticality and Surrogate Replacement

DGX agent

arXiv:2605.11122v1 Announce Type: cross Abstract: Federated Learning remains highly susceptible to backdoor attacks--malicious clients inject targeted behaviours into the global model. Existing defens

safetyarxiv-cs-lg
13 May 2026
Safety

Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models

DGX agent

arXiv:2605.12374v1 Announce Type: new Abstract: Visual latent reasoning lets a multimodal large language model (MLLM) create intermediate visual evidence as continuous tokens, avoiding external tools

safetyarxiv-cs-cv
13 May 2026
Safety

ForceFlow: Learning to Feel and Act via Contact-Driven Flow Matching

DGX agent

arXiv:2605.11048v1 Announce Type: new Abstract: Existing imitation learning methods enable robots to interact autonomously with the physical environment. However, contact-rich manipulation tasks remai

safetyarxiv-cs-ro
13 May 2026
Safety

From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation

DGX agent

arXiv:2605.11613v1 Announce Type: new Abstract: On-policy self-distillation has emerged as a promising paradigm for post-training language models, in which the model conditions on environment feedback

safetyarxiv-cs-lg
13 May 2026
Safety

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation

DGX agent

arXiv:2605.12167v1 Announce Type: cross Abstract: Video generation models offer a promising imagination mechanism for robot manipulation by predicting long-horizon future observations, but effectively

safetyarxiv-cs-cv
13 May 2026
Safety

From Message-Passing to Linearized Graph Sequence Models

DGX agent

arXiv:2605.12358v1 Announce Type: new Abstract: Message-passing based approaches form the default backbone of most learning architectures on graph-structured data. However, the rapid progress of moder

safetyarxiv-cs-lg
13 May 2026
Safety

Fused Gromov-Wasserstein Distance with Feature Selection

DGX agent

arXiv:2605.12161v1 Announce Type: new Abstract: Fused Gromov-Wasserstein (FGW) distances provide a principled framework for comparing objects by jointly aligning structure and node features. However,

safetyarxiv-cs-lg
13 May 2026
Safety

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation

DGX agent

arXiv:2605.11853v1 Announce Type: cross Abstract: Reinforcement learning has become a widely used post-training approach for LLM agents, where training commonly relies on outcome-level rewards that pr

safetyarxiv-cs-cl
13 May 2026
Safety

Generative AI has not made the world a better place.

DGX agent

Generative AI has not made the world a better place. Since 1893, Princeton professors have left the room when students take their final exams. The idea was that if you treat students honorably, they w

safetygary-marcus--x
13 May 2026
Safety

Generative climate downscaling enables high-resolution compound risk assessment by preserving multivariate dependencies

DGX agent

arXiv:2605.11531v1 Announce Type: cross Abstract: Physics-based climate projections using general circulation models are essential for assessing future risks, but their coarse resolution limits region

safetyarxiv-cs-lg
13 May 2026
Safety

Gradient-Free Noise Optimization for Reward Alignment in Generative Models

DGX agent

arXiv:2605.11347v1 Announce Type: cross Abstract: Existing reward alignment methods for diffusion and flow models rely on multi-step stochastic trajectories, making them difficult to extend to determi

safetyarxiv-cs-cv
13 May 2026
Safety

GRAFT: Graph-Tokenized LLMs for Tool Planning

DGX agent

arXiv:2605.11706v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to complete complex tasks by selecting and coordinating external tools across multiple steps. This re

safetyarxiv-cs-lg
13 May 2026
Safety

GRASP: Guided Residual Adapters with Sample-wise Partitioning

DGX agent

arXiv:2512.01675v2 Announce Type: replace Abstract: Text-to-image flow matching transformers degrade sharply in long-tail settings: tail-class outputs collapse in fidelity and diversity, limiting thei

safetyarxiv-cs-cv
13 May 2026
Safety

Hey @Elonmusk I laid out the core of your lawyer’s case against Altman’s credibility almost three years ago in this tweet. Aged well!

DGX agent

Hey @Elonmusk I laid out the core of your lawyer’s case against Altman’s credibility almost three years ago in this tweet. Aged well! Remember how Sam Altman told the US Senate he had no “direct” inve

safetygary-marcus--x
13 May 2026
Safety

Hindsight Hint Distillation: Scaffolded Reasoning for SWE Agents from CoT-free Answers

DGX agent

arXiv:2605.11556v1 Announce Type: cross Abstract: Solving complex long-horizon tasks requires strong planning and reasoning capabilities. Although datasets with explicit chain-of-thought (CoT) rationa

safetyarxiv-cs-lg
13 May 2026
Safety

How Does Differential Privacy Affect Social Bias in LLMs? A Systematic Evaluation

DGX agent

arXiv:2605.11195v1 Announce Type: new Abstract: Large language models (LLMs) trained on web-scale corpora can memorize sensitive training data, posing significant privacy risks. Differential privacy (

safetyarxiv-cs-cl
13 May 2026
Safety

How far can bias go? Tracing bias from pretraining data to alignment

DGX agent

arXiv:2411.19240v2 Announce Type: replace Abstract: As LLMs are increasingly integrated into user-facing applications, addressing biases that perpetuate societal inequalities is crucial. While much wo

safetyarxiv-cs-cl
13 May 2026
Safety

'I applied to be pope': Losing grip on reality while using ChatGPT. AFP spoke to members of a support group for people suffering from AI-ind…

DGX agent

'I applied to be pope': Losing grip on reality while using ChatGPT. AFP spoke to members of a support group for people suffering from AI-induced delusion or psychosis. All warned that the world has to

safetygary-marcus--x
13 May 2026
Safety

i just can’t get over the first sentence here; the absolute confidence with which is it said, and the high likelihood that is epically wrong…

DGX agent

i just can’t get over the first sentence here; the absolute confidence with which is it said, and the high likelihood that is epically wrong. hereby nominating it to the list of sentences that within

safetygary-marcus--x
13 May 2026
Safety

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation

DGX agent

arXiv:2605.12305v1 Announce Type: new Abstract: While recent advancements in multimodal language models have enabled image generation from expressive multi-image instructions, existing methods struggl

safetyarxiv-cs-cv
13 May 2026
Safety

In-Context Multi-Objective Optimization

DGX agent

arXiv:2512.11114v2 Announce Type: replace Abstract: Balancing competing objectives is omnipresent across disciplines, from drug design to autonomous systems. Multi-objective Bayesian optimization is a

safetyarxiv-cs-lg
13 May 2026
Safety

Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning

DGX agent

arXiv:2605.11889v1 Announce Type: new Abstract: Collaborative machine learning involves training high-quality models using datasets from a number of sources. To incentivize sources to share data, exis

safetyarxiv-cs-lg
13 May 2026
Safety

Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning

DGX agent

arXiv:2605.11235v1 Announce Type: new Abstract: In LLM Reinforcement Fine-Tuning (RFT), curriculum learning drives both efficiency and performance. Yet, current methods externalize curriculum judgment

safetyarxiv-cs-lg
13 May 2026
Safety

Intrinsic Vicarious Conditioning for Deep Reinforcement Learning

DGX agent

arXiv:2605.12224v1 Announce Type: new Abstract: Advancements in reinforcement learning have produced a variety of complex and useful intrinsic driving forces; crucially, these drivers operate under a

safetyarxiv-cs-lg
13 May 2026
Safety

Investigating Thinking Behaviours of Reasoning-Based Language Models for Social Bias Mitigation

DGX agent

arXiv:2510.17062v2 Announce Type: replace Abstract: While reasoning-based large language models excel at complex tasks through an internal, structured thinking process, a concerning phenomenon has eme

safetyarxiv-cs-cl
13 May 2026
Safety

Invisible failures in human-AI interactions

DGX agent

arXiv:2603.15423v2 Announce Type: replace Abstract: AI systems fail silently far more often than they fail visibly. In an analysis of 100K human-AI interactions from the WildChat dataset, we find that

safetyarxiv-cs-cl
13 May 2026
Safety

JACoP: Joint Alignment for Compliant Multi-Agent Prediction

DGX agent

arXiv:2605.11385v1 Announce Type: new Abstract: Stochastic Human Trajectory Prediction (HTP) using generative modeling has emerged as a significant area of research. Although state-of-the-art models e

safetyarxiv-cs-cv
13 May 2026
Safety

LA-Sign: Looped Transformers with Geometry-aware Alignment for Skeleton-based Sign Language Recognition

DGX agent

arXiv:2603.29057v2 Announce Type: replace Abstract: Skeleton-based isolated sign language recognition (ISLR) demands fine-grained understanding of articulated motion across multiple spatial scales, fr

safetyarxiv-cs-cv
13 May 2026
Safety

LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer?

DGX agent

arXiv:2605.11301v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have heterogeneous strengths across OCR, chart understanding, spatial reasoning, visual question answering, c

safetyarxiv-cs-cl
13 May 2026
Safety

LEAP: Unlocking dLLM Parallelism via Lookahead Early-Convergence Token Detection

DGX agent

arXiv:2605.10980v1 Announce Type: new Abstract: Diffusion Language Models (dLLMs) have garnered significant attention for their potential in highly parallel processing. The parallel capabilities of ex

safetyarxiv-cs-lg
13 May 2026
Safety

Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training

DGX agent

arXiv:2605.11931v1 Announce Type: new Abstract: Post-training with explicit reasoning traces is common to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, acqui

safetyarxiv-cs-cv
13 May 2026
Safety

Learning Adapter Rank via Symmetry Breaking

DGX agent

arXiv:2506.22809v4 Announce Type: replace-cross Abstract: Low-rank adaptation is effective partly because downstream updates lie in a low-dimensional subspace, but the latent rank coordinates of LoRA

safetyarxiv-cs-cl
13 May 2026
Safety

Learning Agentic Policy from Action Guidance

DGX agent

arXiv:2605.12004v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) for Large Language Models (LLMs) critically depends on the exploration capability of the base policy, as training si

safetyarxiv-cs-cl
13 May 2026
Safety

Learning Minimally Rigid Graphs with High Realization Counts

DGX agent

arXiv:2605.12427v1 Announce Type: new Abstract: For minimally rigid graphs, the same edge-length data can admit multiple realizations (up to translations and rotations). Finding graphs with exceptiona

safetyarxiv-cs-lg
13 May 2026
Safety

Leveraging Multimodal Large Language Models for All-in-One Image Restoration via a Mixture of Frequency Experts

DGX agent

arXiv:2605.11444v1 Announce Type: new Abstract: All-in-one image restoration seeks to recover clean images from inputs affected by diverse and unknown degradations using a unified framework. Recent me

safetyarxiv-cs-cv
13 May 2026
← Previous
1…213214215216217…302
Next →