AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,435 results
Safety

EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance

DGX agent

arXiv:2509.23730v2 Announce Type: replace Abstract: Large language models (LLMs) have recently advanced in reasoning when optimized with reinforcement learning (RL) under verifiable rewards. Existing

safetyarxiv-cs-ai
29 May 2026
Safety
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Emergent Semantic Representations in World Models through Physical Interaction without Linguistic Supervision

DGX agent

arXiv:2605.28865v1 Announce Type: cross Abstract: What does a world model learn from physical exploration, without any linguistic supervision? We argue the answer is organized by a single principle: t

safetyarxiv-cs-ai
29 May 2026
Safety

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models

DGX agent

arXiv:2605.29303v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) followed by reinforcement learning (RL) has become a standard post-training paradigm for large language models. This paradi

safetyarxiv-cs-ai
29 May 2026
Safety

EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance

DGX agent

arXiv:2505.21876v2 Announce Type: replace-cross Abstract: Recent approaches for video generation with camera control often create anchor videos (i.e., rendered videos that approximate desired camera m

safetyarxiv-cs-ai
29 May 2026
Safety

Evolutionary Refinement of Generative Graph Topologies: A Hybrid WGAN-GA Approach

DGX agent

arXiv:2605.29161v1 Announce Type: cross Abstract: Generating realistic graph-structured data is challenging due to discrete connectivity, varying graph sizes, and class-specific structural patterns. R

safetyarxiv-cs-ai
29 May 2026
Safety

EvoMD-LLM: Learning the Language of Species Evolution in Reactive Molecular Dynamics

DGX agent

arXiv:2605.29394v1 Announce Type: new Abstract: While large language models (LLMs) excel at static scientific reasoning, they struggle to model the temporal structure of dynamic physical processes. We

safetyarxiv-cs-ai
29 May 2026
Safety

EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation

DGX agent

arXiv:2605.29847v1 Announce Type: new Abstract: Reinforcement Learning (RL) has significantly advanced Large Language Models (LLMs) in verifiable domains, but aligning models for open-ended generation

safetyarxiv-cs-cl
29 May 2026
Safety

FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection

DGX agent

arXiv:2605.30062v1 Announce Type: new Abstract: The development of generative artificial intelligence technologies has propelled the visual realism of synthetic images to an unprecedented level. Altho

safetyarxiv-cs-cv
29 May 2026
Safety

Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using Language

DGX agent

arXiv:2605.29793v1 Announce Type: new Abstract: Given an untrimmed video and a sentence query, video moment retrieval using language (VMR) aims to locate a target query-relevant moment. Since the untr

safetyarxiv-cs-cv
29 May 2026
Safety

Fisher-Preserving Guidance: Training-Free Manifold Constraints for Safe Diffusion Control

DGX agent

arXiv:2605.29937v1 Announce Type: cross Abstract: Diffusion models are effective for waypoint prediction in visual navigation, but standard sampling and test time guidance can produce unreliable or in

safetyarxiv-cs-lg
29 May 2026
Safety

FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation

DGX agent

arXiv:2605.29461v1 Announce Type: new Abstract: LLM-conditioned segmentation has recently advanced rapidly by coupling large language models with iterative mask generation frameworks. However, we iden

safetyarxiv-cs-cv
29 May 2026
Safety

Follow-the-Perturbed-Leader for Decoupled Bandits: Best-of-Both-Worlds and Practicality

DGX agent

arXiv:2510.12152v2 Announce Type: replace-cross Abstract: We study the decoupled multi-armed bandit problem, where the learner separately selects one arm for exploration and one, possibly different, a

safetyarxiv-cs-lg
29 May 2026
Safety

From Context Shift to Stylistic Collapse: Why Training Objectives Matter More Than Scale

DGX agent

arXiv:2605.28826v1 Announce Type: new Abstract: In modern LLMs, linguistic features function not as stylistic artifacts but as probes of probability mass, allocated under training alignment objectives

safetyarxiv-cs-cl
29 May 2026
Safety

Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation

DGX agent

arXiv:2605.30083v1 Announce Type: new Abstract: Autoregressive (AR) video generation has emerged as a promising paradigm for long-horizon video synthesis, where each frame is generated conditioned on

safetyarxiv-cs-cv
29 May 2026
Safety

GAP3D: Generative Alignment of VLM Latents to Patch-Level Embeddings for 3D Generation

DGX agent

arXiv:2605.28995v1 Announce Type: new Abstract: Recent approaches integrating vision-language models (VLMs) as prompt encoders for generative model conditioning typically rely on expensive end-to-end

safetyarxiv-cs-cv
29 May 2026
Safety

GAPD: Gold-Action Policy Distillation for Agentic Reinforcement Learning in Knowledge Base Question Answering

DGX agent

arXiv:2605.29584v1 Announce Type: new Abstract: Reinforcement learning (RL) is a natural fit for agentic knowledge base question answering (KBQA), where a model must issue executable actions, observe

safetyarxiv-cs-cl
29 May 2026
Safety

GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation

DGX agent

arXiv:2602.17200v2 Announce Type: replace Abstract: Despite high semantic alignment, modern text-to-image (T2I) generative models still struggle to synthesize diverse images from a given prompt. In th

safetyarxiv-cs-cv
29 May 2026
Safety

Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation

DGX agent

arXiv:2605.30282v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently shown strong potential for robot learning by following language instructions. However, in practice, la

safetyarxiv-cs-ro
29 May 2026
Safety

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models

DGX agent

arXiv:2605.29398v1 Announce Type: cross Abstract: Reinforcement learning (RL) can be used to improve the policy (denoiser) of diffusion large language models (dLLMs), while being hindered by the intra

safetyarxiv-cs-ai
29 May 2026
Safety

Genetically Aligned Patient Representations Improve Hematological Diagnosis

DGX agent

arXiv:2605.29980v1 Announce Type: cross Abstract: Multimodal alignment of histopathology encoders with transcriptomic and genomic data has been shown to significantly improve performance in downstream

safetyarxiv-cs-ai
29 May 2026
Safety

Geometry-Guided Modeling of Foundation Features Enables Generalizable Object Shape Deformation Learning

DGX agent

arXiv:2605.29661v1 Announce Type: new Abstract: Monocular 3D shape recovery is fundamental to geometric understanding, yet achieving robust generalization across arbitrary viewpoints and unseen object

safetyarxiv-cs-cv
29 May 2026
Safety

Grammar-Aware Literate Generative Mathematical Programming with Compiler-in-the-Loop

DGX agent

arXiv:2601.17670v2 Announce Type: replace-cross Abstract: Mathematical programming is widely employed across various sectors - such as logistics, energy, and workforce planning - to model and solve in

safetyarxiv-cs-ai
29 May 2026
Safety

Graph-Enhanced Policy Optimization in LLM Agent Training

DGX agent

arXiv:2510.26270v2 Announce Type: replace Abstract: Multi-step LLM agents in interactive environments represent a crucial step toward long-horizon decision-making. To train such agents, group-based re

safetyarxiv-cs-ai
29 May 2026
Safety

GrepSeek: Training Search Agents for Direct Corpus Interaction

DGX agent

arXiv:2605.29307v1 Announce Type: cross Abstract: Large Language Model (LLM) search agents have shown strong promise for knowledge-intensive language tasks through multiple rounds of reasoning and inf

safetyarxiv-cs-ai
29 May 2026
Safety

Grounded 3D-Aware Spatial Vision-Language Modeling

DGX agent

arXiv:2605.30307v1 Announce Type: new Abstract: We present GR3D, a spatial vision language model equipped with three complementary grounding capabilities--explicit 2D grounding, implicit 2D grounding,

safetyarxiv-cs-cv
29 May 2026
Safety

GRPO is Secretly a Process Reward Model

DGX agent

arXiv:2509.21154v4 Announce Type: replace-cross Abstract: Process reward models (PRMs) allow for fine-grained credit assignment in reinforcement learning (RL), and seemingly contrast with outcome rewa

safetyarxiv-cs-ai
29 May 2026
Safety

GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German

DGX agent

arXiv:2605.30214v1 Announce Type: new Abstract: Third-person singular pronouns have long been used to study stereotypical biases in language models and to test their abilities to reason about referenc

safetyarxiv-cs-cl
29 May 2026
Safety

Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization

DGX agent

arXiv:2605.29198v1 Announce Type: new Abstract: Group-advantage-based reinforcement learning methods, such as GRPO and DAPO, have demonstrated strong performance across diverse domains, including math

safetyarxiv-cs-cv
29 May 2026
Safety

Harnessing non-adversarial robustness in large language models

DGX agent

arXiv:2605.29816v1 Announce Type: new Abstract: The work presents an approach for addressing the challenge of robustness in Large Language Models (LLMs) to alterations and potential errors caused by s

safetyarxiv-cs-ai
29 May 2026
Safety

How's it going? Reinforcement learning in language models recruits a functional welfare axis

DGX agent

arXiv:2605.30232v1 Announce Type: cross Abstract: How does reinforcement learning shape a language model's internal representations? We present evidence that RL recruits a pre-existing representation

safetyarxiv-cs-cl
29 May 2026
Safety

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime

DGX agent

arXiv:2605.30201v1 Announce Type: cross Abstract: We investigate a narrow but common failure mode of GRPO-style reinforcement learning in the context of sparse verifiable rewards: early updates contai

safetyarxiv-cs-ai
29 May 2026
Safety

Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning

DGX agent

arXiv:2605.29776v1 Announce Type: new Abstract: Vision-Language Models (VLMs) such as CLIP demonstrate strong zero-shot generalization, but their performance significantly degrades in cross-domain sce

safetyarxiv-cs-cv
29 May 2026
Safety

In-Context Reward Adaptation for Robust Preference Modeling

DGX agent

arXiv:2605.30323v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) typically relies on static reward models to align Large Language Models with human preferences. Howe

safetyarxiv-cs-ai
29 May 2026
Safety

Inferring Code Correctness from Specification

DGX agent

arXiv:2605.29822v1 Announce Type: cross Abstract: Large language models (LLMs) have become integral to modern software development, enabling automated code generation at scale. However, validating the

safetyarxiv-cs-ai
29 May 2026
Safety

Information-Directed Offline-to-Online Reinforcement Learning

DGX agent

arXiv:2605.29405v1 Announce Type: new Abstract: Decision-making from offline datasets typically warm-starts a policy or score model from fixed offline data and then refines it with limited online inte

safetyarxiv-cs-lg
29 May 2026
Safety

KGEdit: Ambiguity-Aware Knowledge Graphs for Training-Free Precise Video Generation and Editing

DGX agent

arXiv:2605.29509v1 Announce Type: new Abstract: In recent years, training-free video generation has progressed remarkably. However, when handling complex textual instructions, existing methods still s

safetyarxiv-cs-cv
29 May 2026
Safety

Learning A Simulation-based Visual Policy for Real-world Peg In Unseen Holes

DGX agent

arXiv:2205.04297v2 Announce Type: replace-cross Abstract: This paper proposes a learning-based visual peg-in-hole that enables training with several shapes in simulation, and adapting to arbitrary uns

safetyarxiv-cs-ai
29 May 2026
Safety

Learning to Choose: An Empowerment-Guided Multi-Agent System with semantic communication for Adaptive Method Selection

DGX agent

arXiv:2605.30042v1 Announce Type: new Abstract: Automating scientific computing workflows requires more than generating executable code: autonomous systems must also select appropriate computational s

safetyarxiv-cs-ai
29 May 2026
Safety

Less Is More: Elevating RAG via Performance-Driven Context Compression

DGX agent

arXiv:2508.19282v4 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) has emerged as a promising paradigm for improving the timeliness of knowledge updates and the factual acc

safetyarxiv-cs-ai
29 May 2026
Safety

LLM-Guided Future Hypotheses for Horizon-Aware Exploration in Multi-Step Robot Manipulation

DGX agent

arXiv:2605.29864v1 Announce Type: new Abstract: Multi-step robot manipulation requires acting under uncertainty about how the scene will evolve, making exploration and policy adaptation challenging. W

safetyarxiv-cs-ro
29 May 2026
Safety

LoMo: Local Modality Substitution for Deeper Vision-Language Fusion

DGX agent

arXiv:2605.30265v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have achieved substantial progress across a wide range of understanding and reasoning tasks, driven by large-scale image

safetyarxiv-cs-cl
29 May 2026
Safety

Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance

DGX agent

arXiv:2411.14279v2 Announce Type: replace-cross Abstract: Large vision-language models (LVLMs) have achieved impressive results in various vision-language tasks. However, despite showing promising per

safetyarxiv-cs-cl
29 May 2026
Safety

MARS Policy: Multimodality Only When It Matters

DGX agent

arXiv:2605.29766v1 Announce Type: new Abstract: Imitation learning has become a cornerstone for solving complex robotic manipulation tasks. In particular, multimodality, which enables robots to captur

safetyarxiv-cs-ro
29 May 2026
Safety

Mask the Target: A Plug-and-Play Regularizer Against LoRA Forgetting

DGX agent

arXiv:2605.29498v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become one of the most widely used fine-tuning mechanisms for adapting large language models to new domains, tasks, and u

safetyarxiv-cs-cl
29 May 2026
Safety

MATANet: A Multi-context Attention and Taxonomy-Aware Network for Fine-Grained Underwater Recognition of Marine Species

DGX agent

arXiv:2601.03729v2 Announce Type: replace Abstract: Fine-grained recognition of marine organisms is important for ecological research, biodiversity monitoring, habitat conservation, and evidence-based

safetyarxiv-cs-cv
29 May 2026
Safety

Mean-Field Diffuser: Scaling Offline MARL to Thousands of Agents

DGX agent

arXiv:2605.30190v1 Announce Type: new Abstract: Diffusion-based planning has achieved strong results in single-agent offline reinforcement learning, yet scaling to many-agent systems remains intractab

safetyarxiv-cs-lg
29 May 2026
Safety

MetaRanker: Human-in-the-loop Active Ranking for Metalens Image Quality

DGX agent

arXiv:2605.29212v1 Announce Type: new Abstract: Image quality in modern imaging systems emerges from the coupled effects of the sensor, optics, and computational reconstruction. Ultra-thin metalenses

safetyarxiv-cs-cv
29 May 2026
Safety

MIC: Maximizing Informational Capacity in Adaptive Representations via Isotropic Subspace Alignment

DGX agent

arXiv:2605.29987v1 Announce Type: cross Abstract: Although multi-scales representation learning enables elastic-dimension embeddings, nested subspaces often suffer from dimensional redundancy and spec

safetyarxiv-cs-cl
29 May 2026
← Previous
1…150151152153154…260
Next →