AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
12,314 results
Safety

PROXIMA: A Reliability Scoring Framework for Proxy Metrics in Online Controlled Experiments

DGX agent

arXiv:2604.14352v1 Announce Type: cross Abstract: Online A/B testing at scale relies on proxy metrics -- short-term, easily-measured signals used in place of slow-moving long-term outcomes. When the p

safetyarxiv-cs-lg
17 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options

DGX agent

arXiv:2604.14634v1 Announce Type: new Abstract: Multiple choice evaluation is widely used for benchmarking large language models, yet near ceiling accuracy in low option settings can be sustained by s

safetyarxiv-cs-cl
17 Apr 2026
Safety

QU-NLP at ArchEHR-QA 2026: Two-Stage QLoRA Fine-Tuning of Qwen3-4B for Patient-Oriented Clinical Question Answering and Evidence Sentence Alignment

DGX agent

arXiv:2604.14175v1 Announce Type: new Abstract: We present a unified system addressing both Subtask 3 (answer generation) and Subtask 4 (evidence sentence alignment) of the ArchEHR-QA Shared Task. For

safetyarxiv-cs-cl
17 Apr 2026
Safety

R3D: Revisiting 3D Policy Learning

DGX agent

arXiv:2604.15281v1 Announce Type: new Abstract: 3D policy learning promises superior generalization and cross-embodiment transfer, but progress has been hindered by training instabilities and severe o

safetyarxiv-cs-cv
17 Apr 2026
Safety

RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models

DGX agent

arXiv:2604.14951v1 Announce Type: cross Abstract: Tool learning with foundation models aims to endow AI systems with the ability to invoke external resources -- such as APIs, computational utilities,

safetyarxiv-cs-cl
17 Apr 2026
Safety

Reinforcement Learning via Value Gradient Flow

DGX agent

arXiv:2604.14265v1 Announce Type: new Abstract: We study behavior-regularized reinforcement learning (RL), where regularization toward a reference distribution (the dataset in offline RL or the base m

safetyarxiv-cs-lg
17 Apr 2026
Safety

Reward-Aware Trajectory Shaping for Few-step Visual Generation

DGX agent

arXiv:2604.14910v1 Announce Type: new Abstract: Achieving high-fidelity generation in extremely few sampling steps has long been a central goal of generative modeling. Existing approaches largely rely

safetyarxiv-cs-cv
17 Apr 2026
Safety

RoSLAC: Robust Simultaneous Localization and Calibration of Multiple Magnetometers

DGX agent

arXiv:2604.14353v1 Announce Type: new Abstract: Localization of autonomous mobile robots (AMRs) in enclosed or semi-enclosed environments such as offices, hotels, hospitals, indoor parking facilities,

safetyarxiv-cs-ro
17 Apr 2026
Safety

SatBLIP: Context Understanding and Feature Identification from Satellite Imagery with Vision-Language Learning

DGX agent

arXiv:2604.14373v1 Announce Type: new Abstract: Rural environmental risks are shaped by place-based conditions (e.g., housing quality, road access, land-surface patterns), yet standard vulnerability i

safetyarxiv-cs-cv
17 Apr 2026
Safety

Scouting By Reward: VLM-TO-IRL-Driven Player Selection For Esports

DGX agent

arXiv:2604.14474v1 Announce Type: new Abstract: Traditional esports scouting workflows rely heavily on manual video review and aggregate performance metrics, which often fail to capture the nuanced de

safetyarxiv-cs-lg
17 Apr 2026
Safety

SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models

DGX agent

arXiv:2604.14672v1 Announce Type: new Abstract: Large language models (LLMs) are being increasingly used in urban planning, but since gendered space theory highlights how gender hierarchies are embedd

safetyarxiv-cs-cl
17 Apr 2026
Safety

Step-level Denoising-time Diffusion Alignment with Multiple Objectives

DGX agent

arXiv:2604.14379v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as a powerful tool for aligning diffusion models with human preferences, typically by optimizing a single rewa

safetyarxiv-cs-cv
17 Apr 2026
Safety

StoryCoder: Narrative Reformulation for Structured Reasoning in LLM Code Generation

DGX agent

arXiv:2604.14631v1 Announce Type: new Abstract: Effective code generation requires both model capability and a problem representation that carefully structures how models reason and plan. Existing app

safetyarxiv-cs-cl
17 Apr 2026
Safety

The Devil Is in Gradient Entanglement: Energy-Aware Gradient Coordinator for Robust Generalized Category Discovery

DGX agent

arXiv:2604.14176v1 Announce Type: new Abstract: Generalized Category Discovery (GCD) leverages labeled data to categorize unlabeled samples from known or unknown classes. Most previous methods jointly

safetyarxiv-cs-lg
17 Apr 2026
Safety

The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows

DGX agent

arXiv:2604.14807v1 Announce Type: cross Abstract: The rapid integration of large language models (LLMs) into everyday workflows has transformed how individuals perform cognitive tasks such as writing,

safetyarxiv-cs-cl
17 Apr 2026
Safety

The PICCO Framework for Large Language Model Prompting: A Taxonomy and Reference Architecture for Prompt Structure

DGX agent

arXiv:2604.14197v1 Announce Type: new Abstract: Large language model (LLM) performance depends heavily on prompt design, yet prompt construction is often described and applied inconsistently. Our purp

safetyarxiv-cs-cl
17 Apr 2026
Safety

The Specification Trap: Why Static Value Alignment Alone Is Insufficient for Robust Alignment

DGX agent

arXiv:2512.03048v4 Announce Type: replace-cross Abstract: Static content-based AI value alignment is insufficient for robust alignment under capability scaling, distributional shift, and increasing au

safetyarxiv-cs-lg
17 Apr 2026
Safety

To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs

DGX agent

arXiv:2603.18373v2 Announce Type: replace Abstract: When VLMs answer correctly, do they genuinely rely on visual information or exploit language shortcuts? We introduce the Tri-Layer Diagnostic Framew

safetyarxiv-cs-cv
17 Apr 2026
Safety

Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithms

DGX agent

arXiv:2506.09457v3 Announce Type: replace Abstract: Direct Alignment Algorithms (DAAs), such as Direct Preference Optimization (DPO) and Simple Preference Optimization (SimPO), have emerged as efficie

safetyarxiv-cs-cl
17 Apr 2026
Safety

Towards Deploying VLA without Fine-Tuning: Plug-and-Play Inference-Time VLA Policy Steering via Embodied Evolutionary Diffusion

DGX agent

arXiv:2511.14178v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have demonstrated significant potential in real-world robotic manipulation. However, pre-trained VLA policies st

safetyarxiv-cs-ro
17 Apr 2026
Safety

Towards Fine-grained Temporal Perception: Post-Training Large Audio-Language Models with Audio-Side Time Prompt

DGX agent

arXiv:2604.13715v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) enable general audio understanding and demonstrate remarkable performance across various audio tasks. However, the

safetyarxiv-cs-ai
17 Apr 2026
Safety

Towards Trustworthy 6G Network Digital Twins: A Framework for Validating Counterfactual What-If Analysis in Edge Computing Resources

DGX agent

arXiv:2604.14787v1 Announce Type: cross Abstract: Network Digital Twins (NDTs) enable safe what-if analysis for 6G cloud-edge infrastructures, but adoption is often limited by fragmented workflows fro

safetyarxiv-cs-lg
17 Apr 2026
Safety

UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards

DGX agent

arXiv:2604.14967v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) extends Large Vision-Language Models (LVLMs) with external visual knowledge. However, existing visual RAG systems t

safetyarxiv-cs-cv
17 Apr 2026
Safety

Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization

DGX agent

arXiv:2604.15196v1 Announce Type: new Abstract: We propose a novel hierarchical spatiotemporal vector quantization framework for unsupervised skeleton-based temporal action segmentation. We first intr

safetyarxiv-cs-cv
17 Apr 2026
Safety

Wasserstein Formulation of Reinforcement Learning. An Optimal Transport Perspective on Policy Optimization

DGX agent

arXiv:2604.14765v1 Announce Type: new Abstract: We present a geometric framework for Reinforcement Learning (RL) that views policies as maps into the Wasserstein space of action probabilities. First,

safetyarxiv-cs-lg
17 Apr 2026
Safety

When Fairness Metrics Disagree: Evaluating the Reliability of Demographic Fairness Assessment in Machine Learning

DGX agent

arXiv:2604.15038v1 Announce Type: cross Abstract: The evaluation of fairness in machine learning systems has become a central concern in high-stakes applications, including biometric recognition, heal

safetyarxiv-cs-cv
17 Apr 2026
Safety

When Missing Becomes Structure: Intent-Preserving Policy Completion from Financial KOL Discourse

DGX agent

arXiv:2604.14333v1 Announce Type: new Abstract: Key Opinion Leader (KOL) discourse on social media is widely consumed as investment guidance, yet turning it into executable trading strategies without

safetyarxiv-cs-lg
17 Apr 2026
Safety

Why Do Vision Language Models Struggle To Recognize Human Emotions?

DGX agent

arXiv:2604.15280v1 Announce Type: new Abstract: Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans. Vision-language models (VLMs) have made trem

safetyarxiv-cs-cv
17 Apr 2026
Safety

Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception

DGX agent

arXiv:2510.23853v3 Announce Type: replace Abstract: Large language model (LLM) agents are increasingly used to interact with and execute tasks in dynamic environments. However, a critical yet overlook

safetyarxiv-cs-cl
17 Apr 2026
Safety

A High-Resolution Landscape Dataset for Concept-Based XAI With Application to Species Distribution Models

DGX agent

arXiv:2604.13240v1 Announce Type: new Abstract: Mapping the spatial distribution of species is essential for conservation policy and invasive species management. Species distribution models (SDMs) are

safetyarxiv-cs-cv
16 Apr 2026
Safety

A Mechanistic Analysis of Sim-and-Real Co-Training in Generative Robot Policies

DGX agent

arXiv:2604.13645v1 Announce Type: cross Abstract: Co-training, which combines limited in-domain real-world data with abundant surrogate data such as simulation or cross-embodiment robot data, is widel

safetyarxiv-cs-lg
16 Apr 2026
Safety

A Multi-Model Approach to English-Bangla Sentiment Classification of Government Mobile Banking App Reviews

DGX agent

arXiv:2604.13057v1 Announce Type: new Abstract: For millions of users in developing economies who depend on mobile banking as their primary gateway to financial services, app quality directly shapes f

safetyarxiv-cs-cl
16 Apr 2026
Safety

Action Images: End-to-End Policy Learning via Multiview Video Generation

DGX agent

arXiv:2604.06168v2 Announce Type: replace Abstract: World action models (WAMs) have emerged as a promising direction for robot policy learning, as they can leverage powerful video backbones to model t

safetyarxiv-cs-cv
16 Apr 2026
Safety

Activation-Guided Local Editing for Jailbreaking Attacks

DGX agent

arXiv:2508.00555v2 Announce Type: replace-cross Abstract: Jailbreaking is an essential adversarial technique for red-teaming these models to uncover and patch security flaws. However, existing jailbre

safetyarxiv-cs-cl
16 Apr 2026
Safety

ADP-DiT: Text-Guided Diffusion Transformer for Brain Image Generation in Alzheimer's Disease Progression

DGX agent

arXiv:2604.13495v1 Announce Type: new Abstract: Alzheimer's disease (AD) progresses heterogeneously across individuals, motivating subject-specific synthesis of follow-up magnetic resonance imaging (M

safetyarxiv-cs-cv
16 Apr 2026
Safety

Alignment as Institutional Design: From Behavioral Correction to Transaction Structure in Intelligent Systems

DGX agent

arXiv:2604.13079v1 Announce Type: cross Abstract: Current AI alignment paradigms rely on behavioral correction: external supervisors (e.g., RLHF) observe outputs, judge against preferences, and adjust

safetyarxiv-cs-lg
16 Apr 2026
Safety

Beyond Arrow's Impossibility: Fairness as an Emergent Property of Multi-Agent Collaboration

DGX agent

arXiv:2604.13705v1 Announce Type: new Abstract: Fairness in language models is typically studied as a property of a single, centrally optimized model. As large language models become increasingly agen

safetyarxiv-cs-cl
16 Apr 2026
Safety

Beyond State Consistency: Behavior Consistency in Text-Based World Models

DGX agent

arXiv:2604.13824v1 Announce Type: new Abstract: World models have been emerging as critical components for assessing the consequences of actions generated by interactive agents in online planning and

safetyarxiv-cs-lg
16 Apr 2026
Safety

Bias-Corrected Adaptive Conformal Inference for Multi-Horizon Time Series Forecasting

DGX agent

arXiv:2604.13253v1 Announce Type: new Abstract: Adaptive Conformal Inference (ACI) provides distribution-free prediction intervals with asymptotic coverage guarantees for time series under distributio

safetyarxiv-cs-lg
16 Apr 2026
Safety

Biased Federated Learning under Wireless Heterogeneity

DGX agent

arXiv:2503.06078v2 Announce Type: replace Abstract: Federated learning (FL) has emerged as a promising framework for distributed learning, enabling collaborative model training without sharing private

safetyarxiv-cs-lg
16 Apr 2026
Safety

CausalDisenSeg: A Causality-Guided Disentanglement Framework with Counterfactual Reasoning for Robust Brain Tumor Segmentation Under Missing Modalities

DGX agent

arXiv:2604.13409v1 Announce Type: new Abstract: In clinical practice, the robustness of deep learning models for multimodal brain tumor segmentation is severely compromised by incomplete MRI data. Thi

safetyarxiv-cs-cv
16 Apr 2026
Safety

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning

DGX agent

arXiv:2604.13804v1 Announce Type: new Abstract: The rapid evolution of multimodal large models has revolutionized the simulation of diverse characters in speech dialogue systems, enabling a novel inte

safetyarxiv-cs-lg
16 Apr 2026
Safety

Composite Silhouette: A Subsampling-based Aggregation Strategy

DGX agent

arXiv:2604.13816v1 Announce Type: new Abstract: Determining the number of clusters is a central challenge in unsupervised learning, where ground-truth labels are unavailable. The Silhouette coefficien

safetyarxiv-cs-lg
16 Apr 2026
Safety

Data-Efficient RLVR via Off-Policy Influence Guidance

DGX agent

arXiv:2510.26491v2 Announce Type: replace Abstract: Data selection is a critical aspect of Reinforcement Learning with Verifiable Rewards (RLVR) for enhancing the reasoning capabilities of large langu

safetyarxiv-cs-lg
16 Apr 2026
Safety

Debate to Align: Reliable Entity Alignment through Two-Stage Multi-Agent Debate

DGX agent

arXiv:2604.13551v1 Announce Type: new Abstract: Entity alignment (EA) aims to identify entities referring to the same real-world object across different knowledge graphs (KGs). Recent approaches based

safetyarxiv-cs-cl
16 Apr 2026
Safety

Depth-Aware Image and Video Orientation Estimation

DGX agent

arXiv:2604.13995v1 Announce Type: new Abstract: This paper introduces a novel approach for image and video orientation estimation by leveraging depth distribution in natural images. The proposed metho

safetyarxiv-cs-cv
16 Apr 2026
Safety

DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off

DGX agent

arXiv:2604.13902v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has catalyzed significant advances in the reasoning capabilities of Large Language Models (LLMs).

safetyarxiv-cs-lg
16 Apr 2026
Safety

Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching

DGX agent

arXiv:2509.21912v2 Announce Type: replace Abstract: Guidance provides a simple and effective framework for posterior sampling by steering the generation process towards the desired distribution. When

safetyarxiv-cs-lg
16 Apr 2026
← Previous
1…223224225226227…257
Next →