AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Tutorials

NoisePrints: Distortion-Free Watermarks for Authorship in Private Diffusion Models

DGX agent

arXiv:2510.13793v2 Announce Type: replace Abstract: With the rapid adoption of diffusion models for visual content generation, proving authorship and protecting copyright have become critical. This ch

tutorialsarxiv-cs-cv
15 Apr 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Nucleus-Image: Sparse MoE for Image Generation

DGX agent

arXiv:2604.12163v1 Announce Type: new Abstract: We present Nucleus-Image, a text-to-image generation model that establishes a new Pareto frontier in quality-versus-efficiency by matching or exceeding

model-releasesarxiv-cs-cv
15 Apr 2026
Model Releases

OFA-Diffusion Compression: Compressing Diffusion Model in One-Shot Manner

DGX agent

arXiv:2604.12668v1 Announce Type: new Abstract: The Diffusion Probabilistic Model (DPM) achieves remarkable performance in image generation, while its increasing parameter size and computational overh

model-releasesarxiv-cs-cv
15 Apr 2026
Research

OmniFood8K: Single-Image Nutrition Estimation via Hierarchical Frequency-Aligned Fusion

DGX agent

arXiv:2604.12356v1 Announce Type: new Abstract: Accurate estimation of food nutrition plays a vital role in promoting healthy dietary habits and personalized diet management. Most existing food datase

researcharxiv-cs-cv
15 Apr 2026
Applications

On Efficient Variants of Segment Anything Model: A Survey

DGX agent

arXiv:2410.04960v5 Announce Type: replace Abstract: The Segment Anything Model (SAM) is a foundational model for image segmentation tasks, known for its strong generalization across diverse applicatio

applicationsarxiv-cs-cv
15 Apr 2026
Applications

One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion

DGX agent

arXiv:2508.04559v3 Announce Type: replace Abstract: Recent diffusion-based approaches have made significant advances in image-based virtual try-on, enabling more realistic and end-to-end garment synth

applicationsarxiv-cs-cv
15 Apr 2026
Research

One View Is Enough! Monocular Training for In-the-Wild Novel View Generation

DGX agent

arXiv:2603.23488v2 Announce Type: replace Abstract: Monocular novel-view synthesis has long required multi-view image pairs for supervision, limiting training data scale and diversity. We argue it is

researcharxiv-cs-cv
15 Apr 2026
Local Ai

PC-MIL: Decoupling Feature Resolution from Supervision Scale in Whole-Slide Learning

DGX agent

arXiv:2604.12100v1 Announce Type: new Abstract: Whole-slide image (WSI) classification in computational pathology is commonly formulated as slide-level Multiple Instance Learning (MIL) with a single g

local-aiarxiv-cs-cv
15 Apr 2026
Applications

PDF-GS: Progressive Distractor Filtering for Robust 3D Gaussian Splatting

DGX agent

arXiv:2604.12580v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting (3DGS) have enabled impressive real-time photorealistic rendering. However, conventional training pipelines inh

applicationsarxiv-cs-cv
15 Apr 2026
Safety

Physics-Grounded Monocular Vehicle Distance Estimation Using Standardized License Plate Typography

DGX agent

arXiv:2604.12239v1 Announce Type: new Abstract: Accurate inter-vehicle distance estimation is a cornerstone of Advanced Driver Assistance Systems (ADAS) and autonomous driving. While LiDAR and radar p

safetyarxiv-cs-cv
15 Apr 2026
Applications

Pi-HOC: Pairwise 3D Human-Object Contact Estimation

DGX agent

arXiv:2604.12923v1 Announce Type: new Abstract: Resolving real-world human-object interactions in images is a many-to-many challenge, in which disentangling fine-grained concurrent physical contact is

applicationsarxiv-cs-cv
15 Apr 2026
Research

PianoFlow: Music-Aware Streaming Piano Motion Generation with Bimanual Coordination

DGX agent

arXiv:2604.12856v1 Announce Type: new Abstract: Audio-driven bimanual piano motion generation requires precise modeling of complex musical structures and dynamic cross-hand coordination. However, exis

researcharxiv-cs-cv
15 Apr 2026
Research

Point Prompting: Counterfactual Tracking with Video Diffusion Models

DGX agent

arXiv:2510.11715v2 Announce Type: replace Abstract: Trackers and video generators solve closely related problems: the former analyze motion, while the latter synthesize it. We show that this connectio

researcharxiv-cs-cv
15 Apr 2026
Research

Privacy-Preserving Structureless Visual Localization via Image Obfuscation

DGX agent

arXiv:2604.12068v1 Announce Type: new Abstract: Visual localization is the task of estimating the camera pose of an image relative to a scene representation. In practice, visual localization systems a

researcharxiv-cs-cv
15 Apr 2026
Safety

Probabilistic Feature Imputation and Uncertainty-Aware Multimodal Federated Aggregation

DGX agent

arXiv:2604.12970v1 Announce Type: cross Abstract: Multimodal federated learning enables privacy-preserving collaborative model training across healthcare institutions. However, a fundamental challenge

safetyarxiv-cs-cv
15 Apr 2026
Research

QMC-Net: Data-Aware Quantum Representations for Remote Sensing Image Classification

DGX agent

arXiv:2604.11817v1 Announce Type: cross Abstract: Hybrid quantum-classical models offer a promising route for learning from complex data; however, their application to multi-band remote sensing imager

researcharxiv-cs-cv
15 Apr 2026
Agents

Radar-Camera BEV Multi-Task Learning with Cross-Task Attention Bridge for Joint 3D Detection and Segmentation

DGX agent

arXiv:2604.12918v1 Announce Type: new Abstract: Bird's-eye-view (BEV) representations are the dominant paradigm for 3D perception in autonomous driving, providing a unified spatial canvas where detect

agentsarxiv-cs-cv
15 Apr 2026
Model Releases

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models

DGX agent

arXiv:2604.12371v1 Announce Type: new Abstract: We study typographic prompt injection attacks on vision-language models (VLMs), where adversarial text is rendered as images to bypass safety mechanisms

model-releasesarxiv-cs-cv
15 Apr 2026
Safety

Redefining Quality Criteria and Distance-Aware Score Modeling for Image Editing Assessment

DGX agent

arXiv:2604.12175v1 Announce Type: new Abstract: Recent advances in image editing have heightened the need for reliable Image Editing Quality Assessment (IEQA). Unlike traditional methods, IEQA require

safetyarxiv-cs-cv
15 Apr 2026
Research

ReefMapGS: Enabling Large-Scale Underwater Reconstruction by Closing the Loop Between Multimodal SLAM and Gaussian Splatting

DGX agent

arXiv:2604.11992v1 Announce Type: cross Abstract: 3D Gaussian Splatting is a powerful visual representation, providing high-quality and efficient 3D scene reconstruction, but it is crucially dependent

researcharxiv-cs-cv
15 Apr 2026
Safety

Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models

DGX agent

arXiv:2604.12582v1 Announce Type: new Abstract: Recent Video Large Language Models (Video-LLMs) have demonstrated strong capability in video understanding, yet they still suffer from hallucinations. E

safetyarxiv-cs-cv
15 Apr 2026
Research

Representing 3D Faces with Learnable B-Spline Volumes

DGX agent

arXiv:2604.12894v1 Announce Type: new Abstract: We present CUBE (Control-based Unified B-spline Encoding), a new geometric representation for human faces that combines B-spline volumes with learned fe

researcharxiv-cs-cv
15 Apr 2026
Research

Retrievals Can Be Detrimental: Unveiling the Backdoor Vulnerability of Retrieval-Augmented Diffusion Models

DGX agent

arXiv:2501.13340v4 Announce Type: replace Abstract: Diffusion models (DMs) have recently demonstrated remarkable generation capability. However, their training generally requires huge computational re

researcharxiv-cs-cv
15 Apr 2026
Safety

Risk-Calibrated Learning: Minimizing Fatal Errors in Medical AI

DGX agent

arXiv:2604.12693v1 Announce Type: new Abstract: Deep learning models often achieve expert-level accuracy in medical image classification but suffer from a critical flaw: semantic incoherence. These hi

safetyarxiv-cs-cv
15 Apr 2026
Model Releases

RPG-SAM: Reliability-Weighted Prototypes and Geometric Adaptive Threshold Selection for Training-Free One-Shot Polyp Segmentation

DGX agent

arXiv:2603.07436v2 Announce Type: replace Abstract: Training-free one-shot segmentation offers a scalable alternative to expert annotations where knowledge is often transferred from support images and

model-releasesarxiv-cs-cv
15 Apr 2026
Research

RSGMamba: Reliability-Aware Self-Gated State Space Model for Multimodal Semantic Segmentation

DGX agent

arXiv:2604.12319v1 Announce Type: new Abstract: Multimodal semantic segmentation has emerged as a powerful paradigm for enhancing scene understanding by leveraging complementary information from multi

researcharxiv-cs-cv
15 Apr 2026
Safety

SAM3-I: Segment Anything with Instructions

DGX agent

arXiv:2512.04585v3 Announce Type: replace Abstract: Segment Anything Model 3 (SAM3) advances open-vocabulary segmentation through promptable concept segmentation, enabling users to segment all instanc

safetyarxiv-cs-cv
15 Apr 2026
Hardware

Scalable Trajectory Generation for Whole-Body Mobile Manipulation

DGX agent

arXiv:2604.12565v1 Announce Type: cross Abstract: Robots deployed in unstructured environments must coordinate whole-body motion -- simultaneously moving a mobile base and arm -- to interact with the

hardwarearxiv-cs-cv
15 Apr 2026
Research

Scaling Exposes the Trigger: Input-Level Backdoor Detection in Text-to-Image Diffusion Models via Cross-Attention Scaling

DGX agent

arXiv:2604.12446v1 Announce Type: cross Abstract: Text-to-image (T2I) diffusion models have achieved remarkable success in image synthesis, but their reliance on large-scale data and open ecosystems i

researcharxiv-cs-cv
15 Apr 2026
Research

Scaling In-Context Segmentation with Hierarchical Supervision

DGX agent

arXiv:2604.12752v1 Announce Type: new Abstract: In-context learning (ICL) enables medical image segmentation models to adapt to new anatomical structures from limited examples, reducing the clinical a

researcharxiv-cs-cv
15 Apr 2026
Model Releases

SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents

DGX agent

arXiv:2510.10073v2 Announce Type: replace-cross Abstract: Large vision-language model (LVLM)-based web agents are emerging as powerful tools for automating complex online tasks. However, when deployed

model-releasesarxiv-cs-cv
15 Apr 2026
Model Releases

See, Point, Refine: Multi-Turn Approach to GUI Grounding with Visual Feedback

DGX agent

arXiv:2604.13019v1 Announce Type: new Abstract: Computer Use Agents (CUAs) fundamentally rely on graphical user interface (GUI) grounding to translate language instructions into executable screen acti

model-releasesarxiv-cs-cv
15 Apr 2026
Model Releases

Self-Adversarial One Step Generation via Condition Shifting

DGX agent

arXiv:2604.12322v1 Announce Type: new Abstract: The push for efficient text to image synthesis has moved the field toward one step sampling, yet existing methods still face a three way tradeoff among

model-releasesarxiv-cs-cv
15 Apr 2026
Model Releases

SinkSAM-Net: Knowledge-Driven Self-Supervised Sinkhole Segmentation Using Topographic Priors and Segment Anything Model

DGX agent

arXiv:2410.01473v2 Announce Type: replace Abstract: Soil sinkholes significantly influence soil degradation, infrastructure vulnerability, and landscape evolution. However, their irregular shapes, com

model-releasesarxiv-cs-cv
15 Apr 2026
Model Releases

SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks

DGX agent

arXiv:2506.14512v4 Announce Type: replace Abstract: Large Language Models (LLMs) have undergone rapid progress, largely attributed to reinforcement learning on complex reasoning tasks. In contrast, wh

model-releasesarxiv-cs-cv
15 Apr 2026
Model Releases

SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model

DGX agent

arXiv:2511.22039v3 Announce Type: replace Abstract: This paper introduces a novel architecture for trajectory-conditioned forecasting of future 3D scene occupancy. In contrast to methods that rely on

model-releasesarxiv-cs-cv
15 Apr 2026
Local Ai

Spatial-Spectral Adaptive Fidelity and Noise Prior Reduction Guided Hyperspectral Image Denoising

DGX agent

arXiv:2604.12600v1 Announce Type: new Abstract: The core challenge of hyperspectral image denoising is striking the right balance between data fidelity and noise prior modeling. Most existing methods

local-aiarxiv-cs-cv
15 Apr 2026
Tutorials

StructDiff: A Structure-Preserving and Spatially Controllable Diffusion Model for Single-Image Generation

DGX agent

arXiv:2604.12575v1 Announce Type: new Abstract: This paper introduces StructDiff, a generative framework based on a single-scale diffusion model for single-image generation. Single-image generation ai

tutorialsarxiv-cs-cv
15 Apr 2026
Model Releases

Style-Decoupled Adaptive Routing Network for Underwater Image Enhancement

DGX agent

arXiv:2604.12257v1 Announce Type: new Abstract: Underwater Image Enhancement (UIE) is essential for robust visual perception in marine applications. However, existing methods predominantly rely on uni

model-releasesarxiv-cs-cv
15 Apr 2026
Tutorials

SubFlow: Sub-mode Conditioned Flow Matching for Diverse One-Step Generation

DGX agent

arXiv:2604.12273v1 Announce Type: cross Abstract: Flow matching has emerged as a powerful generative framework, with recent few-step methods achieving remarkable inference acceleration. However, we id

tutorialsarxiv-cs-cv
15 Apr 2026
Model Releases

Subspace-Guided Feature Reconstruction for Unsupervised Anomaly Localization

DGX agent

arXiv:2309.13904v3 Announce Type: replace Abstract: Unsupervised anomaly localization aims to identify anomalous regions that deviate from normal sample patterns. Most recent methods perform feature m

model-releasesarxiv-cs-cv
15 Apr 2026
Model Releases

SynthPix: A lightspeed PIV image generator

DGX agent

arXiv:2512.09664v2 Announce Type: replace-cross Abstract: We describe SynthPix, a synthetic image generator for Particle Image Velocimetry (PIV) with a focus on performance and parallelism on accelera

model-releasesarxiv-cs-cv
15 Apr 2026
Model Releases

T2I-BiasBench: A Multi-Metric Framework for Auditing Demographic and Cultural Bias in Text-to-Image Models

DGX agent

arXiv:2604.12481v1 Announce Type: new Abstract: Text-to-image (T2I) generative models achieve impressive visual fidelity but inherit and amplify demographic imbalances and cultural biases embedded in

model-releasesarxiv-cs-cv
15 Apr 2026
Safety

Task Alignment: A simple and effective proxy for model merging in computer vision

DGX agent

arXiv:2604.12935v1 Announce Type: new Abstract: Efficiently merging several models fine-tuned for different tasks, but stemming from the same pretrained base model, is of great practical interest. Des

safetyarxiv-cs-cv
15 Apr 2026
Local Ai

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

DGX agent

arXiv:2512.03963v3 Announce Type: replace Abstract: Enhancing the temporal understanding of Multimodal Large Language Models (MLLMs) is essential for advancing long-form video analysis, enabling tasks

local-aiarxiv-cs-cv
15 Apr 2026
Local Ai

Time-reversed Flow Matching with Worst Transport in High-dimensional Latent Space for Image Anomaly Detection

DGX agent

arXiv:2508.05461v3 Announce Type: replace Abstract: Likelihood-based deep generative models have been widely investigated for Image Anomaly Detection (IAD), particularly Normalizing Flows, yet their s

local-aiarxiv-cs-cv
15 Apr 2026
Model Releases

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

DGX agent

arXiv:2604.12012v1 Announce Type: new Abstract: Recent progress in vision-language pretraining has enabled significant improvements to many downstream computer vision applications, such as classificat

model-releasesarxiv-cs-cv
15 Apr 2026
Local Ai

Toward Efficient and Robust Behavior Models for Multi-Agent Driving Simulation

DGX agent

arXiv:2512.05812v5 Announce Type: replace-cross Abstract: Scalable multi-agent driving simulation requires behavior models that are both realistic and computationally efficient. We address this by opt

local-aiarxiv-cs-cv
15 Apr 2026
← Previous
1…241242243244245…259
Next →