AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlog
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Tutorials

EAM: Enhancing Anything with Diffusion Transformers for Blind Super-Resolution

DGX agent

arXiv:2505.05209v4 Announce Type: replace Abstract: Utilizing pre-trained Text-to-Image (T2I) diffusion models to guide Blind Super-Resolution (BSR) has become a predominant approach in the field. Whi

tutorialsarxiv-cs-cv
12 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Local Ai

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing

DGX agent

arXiv:2605.08723v1 Announce Type: new Abstract: Weakly supervised Audio-Visual Video Parsing (AVVP) aims to recognize and temporally localize audio, visual, and audio-visual events in videos using onl

local-aiarxiv-cs-cv
12 May 2026
Research

EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs

DGX agent

arXiv:2605.10050v1 Announce Type: new Abstract: Long-form video understanding remains challenging for Video Large Language Models (VideoLLMs), as the dense frame sampling introduces massive visual tok

researcharxiv-cs-cv
12 May 2026
Local Ai

EditSleuth: A Dataset of Grounded Reasoning Chains for Image-Edit Forensics

DGX agent

arXiv:2605.08695v1 Announce Type: new Abstract: Forensic analysis of AI-edited images requires more than binary real-versus-fake prediction: a useful system should localize the edit, identify its sema

local-aiarxiv-cs-cv
12 May 2026
Hardware

Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching

DGX agent

arXiv:2602.05391v2 Announce Type: replace Abstract: Dataset distillation seeks to synthesize a highly compact dataset that achieves performance comparable to the original dataset on downstream tasks.

hardwarearxiv-cs-cv
12 May 2026
Local Ai

Efficient Hybrid CNN-GNN Architecture for Monocular Depth Estimation

DGX agent

arXiv:2605.10251v1 Announce Type: new Abstract: We present GraphDepth, a monocular depth estimation architecture that synergistically integrates Graph Neural Networks (GNNs) within a convolutional enc

local-aiarxiv-cs-cv
12 May 2026
Research

Egocentric Whole-Body Human Mesh Recovery with Prior-Guided Learning

DGX agent

arXiv:2605.08606v1 Announce Type: new Abstract: Egocentric human mesh recovery (HMR) from monocular head-mounted cameras is increasingly important for AR/VR applications, but remains challenging due t

researcharxiv-cs-cv
12 May 2026
Research

EMFormer: Efficient Multi-Scale Transformer for Accumulative Context Weather Forecasting

DGX agent

arXiv:2602.01194v2 Announce Type: replace Abstract: Long-term weather forecasting is critical for socioeconomic planning and disaster preparedness. While recent approaches employ finetuning to extend

researcharxiv-cs-cv
12 May 2026
Model Releases

EnergyLens: Interpretable Closed-Form Energy Models for Multimodal LLM Inference Serving

DGX agent

arXiv:2605.10556v1 Announce Type: new Abstract: As large language models span dense, mixture-of-experts, and state-space architectures and are deployed on heterogeneous accelerators under increasingly

model-releasesarxiv-cs-cv
12 May 2026
Agents

Enhancing Consistency Models for Multi-Agent Trajectory Prediction

DGX agent

arXiv:2605.08572v1 Announce Type: new Abstract: Diffusion models for multi-agent trajectory prediction are limited by iterative denoising, which causes inference latency that hinders their use in time

agentsarxiv-cs-cv
12 May 2026
Research

Enhancing Few-Shot Out-of-Distribution Detection via the Refinement of Foreground and Background

DGX agent

arXiv:2601.15065v2 Announce Type: replace Abstract: CLIP-based foreground-background (FG-BG) decomposition methods have demonstrated remarkable effectiveness in improving few-shot out-of-distribution

researcharxiv-cs-cv
12 May 2026
Applications

ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning

DGX agent

arXiv:2605.09982v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) enable large language models (LLMs) to process high-resolution images, significantly improving real

applicationsarxiv-cs-cv
12 May 2026
Applications

Establishing Robust Retinal Eye Tracking: A Weakly Supervised Algorithmic Framework

DGX agent

arXiv:2605.09181v1 Announce Type: new Abstract: Retinal image-based eye tracking is widely used in ophthalmic imaging and vision science, and is a promising path to deliver higher gaze accuracy than t

applicationsarxiv-cs-cv
12 May 2026
Research

Evidence-based Decision Modeling for Synthetic Face Detection with Uncertainty-driven Active Learning

DGX agent

arXiv:2605.09935v1 Announce Type: new Abstract: With the rapid development of deep generative models, forged facial images are massively exploited for illegal activities. Although existing synthetic f

researcharxiv-cs-cv
12 May 2026
Safety

Explanation-Aware Learning for Enhanced Interpretability in Biomedical Imaging

DGX agent

arXiv:2605.10054v1 Announce Type: new Abstract: Deep neural networks for medical image diagnosis often achieve high predictive accuracy while relying on spurious or clinically irrelevant visual cues,

safetyarxiv-cs-cv
12 May 2026
Research

Exploring 6D Object Pose Estimation with Deformation

DGX agent

arXiv:2604.06720v2 Announce Type: replace Abstract: We present DeSOPE, a large-scale dataset for 6DoF deformed objects. Most 6D object pose methods assume rigid or articulated objects, an assumption t

researcharxiv-cs-cv
12 May 2026
Research

Exploring and Exploiting Stability in Latent Flow Matching

DGX agent

arXiv:2605.08398v1 Announce Type: cross Abstract: In this work, we show that Latent Flow-Matching (LFM) models are robust to different types of perturbations, including data reduction and model capaci

researcharxiv-cs-cv
12 May 2026
Research

ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models

DGX agent

arXiv:2605.10045v1 Announce Type: new Abstract: Visual Autoregressive (VAR) models have emerged as a strong alternative to diffusion for image synthesis, yet their fixed training resolution prevents d

researcharxiv-cs-cv
12 May 2026
Model Releases

Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition

DGX agent

arXiv:2605.10127v1 Announce Type: new Abstract: Recent research work on fashion outfit generation focuses on promoting visual consistency of garments by leveraging key information from reference image

model-releasesarxiv-cs-cv
12 May 2026
Research

Fetal Brain Imaging: A Composite Neural Network Approach for Keyframe Detection in Ultrasound Videos

DGX agent

arXiv:2605.09750v1 Announce Type: new Abstract: This article presents a novel approach to keyframe detection in ultrasound videos, with a particular focus on fetal brain imaging. The proposed model is

researcharxiv-cs-cv
12 May 2026
Safety

Few-Click-Driven Interactive 3D Segmentation with Semantic Embedding

DGX agent

arXiv:2605.08925v1 Announce Type: new Abstract: Interactive segmentation allows efficient label generation by leveraging user-provided clicks to progressively refine predictions, which is critical whe

safetyarxiv-cs-cv
12 May 2026
Model Releases

Filtering Memorization from Parameter-Space in Diffusion Models

DGX agent

arXiv:2605.10439v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become a widely used mechanism for customizing diffusion models, enabling users to inject new visual concepts or styles t

model-releasesarxiv-cs-cv
12 May 2026
Safety

FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation

DGX agent

arXiv:2605.09430v1 Announce Type: new Abstract: Large-scale autoregressive models have demonstrated remarkable capabilities in image generation. However, their sequential raster-scan decoding relies o

safetyarxiv-cs-cv
12 May 2026
Model Releases

FlashClear: Ultra-Fast Image Content Removal via Efficient Step Distillation and Feature Caching

DGX agent

arXiv:2605.09003v1 Announce Type: new Abstract: Recently, diffusion-based object removal models have achieved impressive results in eliminating objects and their associated visual effects. However, th

model-releasesarxiv-cs-cv
12 May 2026
Research

FlowADMM: Plug-and-play ADMM with Flow-based Renoise-Denoise Priors

DGX agent

arXiv:2605.08640v1 Announce Type: new Abstract: Plug-and-play (PnP) methods for solving inverse problems have recently achieved strong performance by leveraging denoising priors based on powerful gene

researcharxiv-cs-cv
12 May 2026
Hardware

Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models

DGX agent

arXiv:2605.09681v1 Announce Type: new Abstract: Autoregressive (AR) video diffusion models adopt a streaming generation framework, enabling long-horizon video generation with real-time responsiveness,

hardwarearxiv-cs-cv
12 May 2026
Model Releases

FPGA-Based Hardware Architecture for Contrast Maximization in Event-Based Vision

DGX agent

arXiv:2605.09581v1 Announce Type: new Abstract: This paper presents a hardware architecture that implements the Contrast Maximization (CM) algorithm in Field-Programmable Gate Array (FPGA) resources f

model-releasesarxiv-cs-cv
12 May 2026
Safety

FrameTwin: Curve-Anchored Gaussian Alignment from Sparse Views for Adaptive Wireframe 3D Printing

DGX agent

arXiv:2605.09362v1 Announce Type: cross Abstract: We present FrameTwin, a curve-anchored Gaussian alignment framework that uses sparse-view images to close the control loop for adaptive wireframe 3D p

safetyarxiv-cs-cv
12 May 2026
Safety

Frequency Adapter with SAM for Generalized Medical Image Segmentation

DGX agent

arXiv:2605.09925v1 Announce Type: new Abstract: Medical image segmentation is a critical task in computer-aided diagnosis and treatment planning. However, deep learning models often struggle to genera

safetyarxiv-cs-cv
12 May 2026
Applications

FrequencyCT: Frequency domain pseudo-label generation for self-supervised low-dose CT denoising

DGX agent

arXiv:2605.10583v1 Announce Type: new Abstract: Despite extensive research on computed tomography (CT) denoising, few studies exploit projection-domain data characteristics to mitigate noise correlati

applicationsarxiv-cs-cv
12 May 2026
Model Releases

From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation

DGX agent

arXiv:2605.08712v1 Announce Type: new Abstract: Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimension

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

From Pixels to Concepts: Do Segmentation Models Understand What They Segment?

DGX agent

arXiv:2605.09591v1 Announce Type: new Abstract: Segmentation is a fundamental vision task underlying numerous downstream applications. Recent promptable segmentation models, such as Segment Anything M

model-releasesarxiv-cs-cv
12 May 2026
Research

From pre-training to downstream performance: Does domain-specific pre-training make sense?

DGX agent

arXiv:2605.08819v1 Announce Type: new Abstract: Deep learning techniques have revolutionised medical imaging, improving diagnostic accuracy and enabling both more accurate and earlier disease detectio

researcharxiv-cs-cv
12 May 2026
Research

FugSeg: Fast Uncertainty-aware Ground Segmentation for 3D Point Cloud

DGX agent

arXiv:2605.08952v1 Announce Type: new Abstract: In LiDAR-based environment perception systems, ground segmentation is a key preprocessing step supporting various applications such as mapping and navig

researcharxiv-cs-cv
12 May 2026
Safety

GemDepth: Geometry-Embedded Features for 3D-Consistent Video Depth

DGX agent

arXiv:2605.10525v1 Announce Type: new Abstract: Video depth estimation extends monocular prediction into the temporal domain to ensure coherence. However, existing methods often suffer from spatial bl

safetyarxiv-cs-cv
12 May 2026
Research

Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?

DGX agent

arXiv:2512.19115v2 Announce Type: replace Abstract: Despite the remarkable success of multimodal large language models (MLLMs) in generative tasks, we observe that they exhibit a counterintuitive defi

researcharxiv-cs-cv
12 May 2026
Model Releases

GenMed: A Pairwise Generative Reformulation of Medical Diagnostic Tasks

DGX agent

arXiv:2605.10645v1 Announce Type: new Abstract: Data-driven medical AI is traditionally formulated as a discriminative mapping from input X to output Y via a learned function f, which does not general

model-releasesarxiv-cs-cv
12 May 2026
Research

Geometric Flood Depth Estimation: Fusing Transformer-Based Segmentation with Digital Elevation Models

DGX agent

arXiv:2605.08521v1 Announce Type: new Abstract: Post-disaster situational awareness relies heavily on understanding both the extent and the volume of floodwaters. While 2D semantic segmentation provid

researcharxiv-cs-cv
12 May 2026
Research

Geometry-aware Prototype Learning for Cross-domain Few-shot Medical Image Segmentation

DGX agent

arXiv:2605.10885v1 Announce Type: new Abstract: Cross-domain few-shot medical image segmentation (CD-FSMIS) requires a model to generalise simultaneously to novel anatomical categories and unseen imag

researcharxiv-cs-cv
12 May 2026
Research

GLEAM: A Multimodal Imaging Dataset and HAMM for Glaucoma Classification

DGX agent

arXiv:2603.12800v2 Announce Type: replace-cross Abstract: We propose glaucoma lesion evaluation and analysis with multimodal imaging (GLEAM), the first publicly available tri-modal glaucoma dataset co

researcharxiv-cs-cv
12 May 2026
Agents

GSMap: 2D Gaussians for Online HD Mapping

DGX agent

arXiv:2605.09619v1 Announce Type: new Abstract: Accurate High-Definition (HD) map construction is critical for autonomous driving, yet existing methods face a fundamental trade-off: vectorization-base

agentsarxiv-cs-cv
12 May 2026
Model Releases

H-POPE: Hierarchical Polling-based Probing Evaluation of Hallucinations in Large Vision-Language Models

DGX agent

arXiv:2411.04077v2 Announce Type: replace Abstract: By leveraging both texts and images, large vision language models (LVLMs) have shown significant progress in various multi-modal tasks. Nevertheless

model-releasesarxiv-cs-cv
12 May 2026
Tutorials

HairGPT: Strand-as-Language Autoregressive Modeling for Realistic 3D Hairstyle Synthesis

DGX agent

arXiv:2605.08824v1 Announce Type: cross Abstract: Hair is a rich medium of visual and cultural expression, yet its digital modeling remains challenging due to the duality of fluidity and structure. Ma

tutorialsarxiv-cs-cv
12 May 2026
Agents

Halo Separation-guided Underwater Multi-scale Image Restoration

DGX agent

arXiv:2605.10374v1 Announce Type: new Abstract: Underwater images captured by Autonomous Underwater Vehicles (AUVs) are inevitably affected by artificial light sources, which often produce halos in th

agentsarxiv-cs-cv
12 May 2026
Research

Harmonized Feature Conditioning and Frequency-Prompt Personalization for Multi-Rater Medical Segmentation

DGX agent

arXiv:2605.08210v1 Announce Type: new Abstract: Multi-rater medical image segmentation captures the inherent ambiguity of clinical interpretation, where diagnostic boundaries vary across experts and i

researcharxiv-cs-cv
12 May 2026
Agents

Heteroscedastic Diffusion for Multi-Agent Trajectory Modeling

DGX agent

arXiv:2605.10717v1 Announce Type: cross Abstract: Multi-agent trajectory modeling traditionally focuses on forecasting, often neglecting more general tasks like trajectory completion, which is essenti

agentsarxiv-cs-cv
12 May 2026
Model Releases

HiDrive: A Closed-Loop Benchmark for High-Level Autonomous Driving

DGX agent

arXiv:2605.09972v1 Announce Type: cross Abstract: End-to-end autonomous driving has witnessed rapid progress, yet existing benchmarks are increasingly saturated, with state-of-the-art models achieving

model-releasesarxiv-cs-cv
12 May 2026
Research

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models

DGX agent

arXiv:2512.21815v2 Announce Type: replace Abstract: Vision-language models (VLMs) achieve remarkable performance but remain vulnerable to adversarial attacks. Entropy, as a measure of model uncertaint

researcharxiv-cs-cv
12 May 2026
← Previous
1…181182183184185…263
Next →