AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Research

Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools

DGX agent

arXiv:2510.08480v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable potential in bridging visual and textual reasoning, yet their reliance on text

researcharxiv-cs-cv
20 Apr 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions

DGX agent

arXiv:2604.15823v1 Announce Type: new Abstract: Embodied robotic agents often perceive movies through an egocentric screen-view interface rather than native cinematic footage, introducing domain shift

model-releasesarxiv-cs-cv
20 Apr 2026
Research

Weak-to-Strong Knowledge Distillation Accelerates Visual Learning

DGX agent

arXiv:2604.15451v1 Announce Type: new Abstract: Large-scale visual learning is increasingly limited by training cost. Existing knowledge distillation methods transfer from a stronger teacher to a weak

researcharxiv-cs-cv
20 Apr 2026
Model Releases

When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models

DGX agent

arXiv:2603.27759v3 Announce Type: replace Abstract: Visual-Language Models (VLMs) have demonstrated exceptional cross-modal understanding across various tasks, including zero-shot classification, imag

model-releasesarxiv-cs-cv
20 Apr 2026
Local Ai

Where Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization

DGX agent

arXiv:2604.16248v1 Announce Type: new Abstract: Image geolocalization has traditionally been addressed through retrieval-based place recognition or geometry-based visual localization pipelines. Recent

local-aiarxiv-cs-cv
20 Apr 2026
Research

Winner of CVPR2026 NTIRE Challenge on Image Shadow Removal: Semantic and Geometric Guidance for Shadow Removal via Cascaded Refinement

DGX agent

arXiv:2604.16177v1 Announce Type: new Abstract: We present a three-stage progressive shadow-removal pipeline for the CVPR2026 NTIRE WSRD+ challenge. Built on OmniSR, our method treats deshadowing as i

researcharxiv-cs-cv
20 Apr 2026
Research

3AM: 3egment Anything with Geometric Consistency in Videos

DGX agent

arXiv:2601.08831v5 Announce Type: replace Abstract: Video object segmentation methods like SAM2 achieve strong performance through memory-based architectures but struggle under large viewpoint changes

researcharxiv-cs-cv
17 Apr 2026
Research

3D Conditional Image Synthesis of Left Atrial LGE MRI from Composite Semantic Masks

DGX agent

arXiv:2601.04588v2 Announce Type: replace Abstract: Segmentation of the left atrial (LA) wall and endocardium from late gadolinium-enhanced (LGE) MRI is essential for quantifying atrial fibrosis in pa

researcharxiv-cs-cv
17 Apr 2026
Research

A deep learning framework for glomeruli segmentation with boundary attention

DGX agent

arXiv:2604.14263v1 Announce Type: cross Abstract: Accurate detection and segmentation of glomeruli in kidney tissue are essential for diagnostic applications. Traditional deep learning methods primari

researcharxiv-cs-cv
17 Apr 2026
Model Releases

AD4AD: Benchmarking Visual Anomaly Detection Models for Safer Autonomous Driving

DGX agent

arXiv:2604.15291v1 Announce Type: new Abstract: The reliability of a machine vision system for autonomous driving depends heavily on its training data distribution. When a vehicle encounters significa

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

AgentIAD: Agentic Industrial Anomaly Detection via Adaptive Memory Augmentation

DGX agent

arXiv:2512.13671v2 Announce Type: replace Abstract: Industrial anomaly detection (IAD) is challenging due to the subtle and highly localized nature of many defects, which single-pass vision--language

model-releasesarxiv-cs-cv
17 Apr 2026
Tutorials

All Changes May Have Invariant Principles: Improving Ever-Shifting Harmful Meme Detection via Design Concept Reproduction

DGX agent

arXiv:2601.04567v2 Announce Type: replace Abstract: Harmful memes are ever-shifting in the Internet communities, which are difficult to analyze due to their type-shifting and temporal-evolving nature.

tutorialsarxiv-cs-cv
17 Apr 2026
Research

An Analysis of Regularization and Fokker-Planck Residuals in Diffusion Models for Image Generation

DGX agent

arXiv:2604.15171v1 Announce Type: new Abstract: Recent work has shown that diffusion models trained with the denoising score matching (DSM) objective often violate the Fokker--Planck (FP) equation tha

researcharxiv-cs-cv
17 Apr 2026
Model Releases

AnimationBench: Are Video Models Good at Character-Centric Animation?

DGX agent

arXiv:2604.15299v1 Announce Type: new Abstract: Video generation has advanced rapidly, with recent methods producing increasingly convincing animated results. However, existing benchmarks-largely desi

model-releasesarxiv-cs-cv
17 Apr 2026
Research

ArrowGEV: Grounding Events in Video via Learning the Arrow of Time

DGX agent

arXiv:2601.06559v2 Announce Type: replace Abstract: Grounding events in videos serves as a fundamental capability in video analysis. While Vision Language Models (VLMs) are increasingly employed for t

researcharxiv-cs-cv
17 Apr 2026
Research

ASGNet: Adaptive Spectrum Guidance Network for Automatic Polyp Segmentation

DGX agent

arXiv:2604.14755v1 Announce Type: new Abstract: Early identification and removal of polyps can reduce the risk of developing colorectal cancer. However, the diverse morphologies, complex backgrounds a

researcharxiv-cs-cv
17 Apr 2026
Local Ai

Attention-Gated Convolutional Networks for Scanner-Agnostic Quality Assessment

DGX agent

arXiv:2604.15059v1 Announce Type: new Abstract: Motion artifacts present a significant challenge in structural MRI (sMRI), often compromising clinical diagnostics and large-scale automated analysis. W

local-aiarxiv-cs-cv
17 Apr 2026
Local Ai

Beyond Augmentation: Cross-Modal Transformer Fusion with Bi-directional Attention for Low-Data Aneurysm Screening

DGX agent

arXiv:2512.22185v2 Announce Type: replace Abstract: Intracranial aneurysm rupture causes subarachnoid hemorrhage with mortality near 50%, making early detection critical. Although CTA enables rapid sc

local-aiarxiv-cs-cv
17 Apr 2026
Applications

Beyond Independent Frames: Latent Attention Masked Autoencoders for Multi-View Echocardiography

DGX agent

arXiv:2604.15096v1 Announce Type: new Abstract: Echocardiography is a widely used modality for cardiac assessment due to its non-invasive and cost-effective nature, but the sparse and heterogeneous sp

applicationsarxiv-cs-cv
17 Apr 2026
Research

Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes

DGX agent

arXiv:2604.14914v1 Announce Type: new Abstract: Text-driven inversion of generative models is a core paradigm for manipulating 2D or 3D content, unlocking numerous applications such as text-based edit

researcharxiv-cs-cv
17 Apr 2026
Tutorials

Beyond Visual Cues: Semantic-Driven Token Filtering and Expert Routing for Anytime Person ReID

DGX agent

arXiv:2604.15090v1 Announce Type: new Abstract: Any-Time Person Re-identification (AT-ReID) necessitates the robust retrieval of target individuals under arbitrary conditions, encompassing both modali

tutorialsarxiv-cs-cv
17 Apr 2026
Research

Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo

DGX agent

arXiv:2604.15312v1 Announce Type: new Abstract: Conventional frame-based cameras capture rich contextual information but suffer from limited temporal resolution and motion blur in dynamic scenes. Even

researcharxiv-cs-cv
17 Apr 2026
Safety

Bird-SR: Bidirectional Reward-Guided Diffusion for Real-World Image Super-Resolution

DGX agent

arXiv:2602.07069v2 Announce Type: replace Abstract: Powered by multimodal text-to-image priors, diffusion-based super-resolution excels at synthesizing intricate details; however, models trained on sy

safetyarxiv-cs-cv
17 Apr 2026
Research

Boundary-Centric Active Learning for Temporal Action Segmentation

DGX agent

arXiv:2604.15173v1 Announce Type: new Abstract: Temporal action segmentation (TAS) demands dense temporal supervision, yet most of the annotation cost in untrimmed videos is spent identifying and refi

researcharxiv-cs-cv
17 Apr 2026
Model Releases

Building Extraction from Remote Sensing Imagery under Hazy and Low-light Conditions: Benchmark and Baseline

DGX agent

arXiv:2604.15088v1 Announce Type: new Abstract: Building extraction from optical Remote Sensing (RS) imagery suffers from performance degradation under real-world hazy and low-light conditions. Howeve

model-releasesarxiv-cs-cv
17 Apr 2026
Research

C2W-Tune: Cavity-to -Wall Transfer Learning for Thin Atrial Wall Segmentation in 3D Late Gadolinium-enhanced Magnetic Resonance

DGX agent

arXiv:2603.24992v2 Announce Type: replace Abstract: Accurate segmentation of the left atrial (LA) wall in 3D late gadolinium-enhanced MRI (LGE-MRI) is essential for wall thickness mapping and fibrosis

researcharxiv-cs-cv
17 Apr 2026
Model Releases

CaptionQA: Is Your Caption as Useful as the Image Itself?

DGX agent

arXiv:2511.21025v2 Announce Type: replace Abstract: Image captions serve as efficient surrogates for visual content in multimodal systems such as retrieval, recommendation, and multi-step agentic infe

model-releasesarxiv-cs-cv
17 Apr 2026
Tutorials

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding

DGX agent

arXiv:2604.14692v1 Announce Type: new Abstract: Video understanding requires identifying and reasoning over semantically discriminative visual objects across frames, yet existing object-agnostic solut

tutorialsarxiv-cs-cv
17 Apr 2026
Safety

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs

DGX agent

arXiv:2604.14520v1 Announce Type: new Abstract: Omni-modal Large Language Models (Omni-MLLMs) promise a unified integration of diverse sensory streams. However, recent evaluations reveal a critical pe

safetyarxiv-cs-cv
17 Apr 2026
Research

Chaotic CNN for Limited Data Image Classification

DGX agent

arXiv:2604.14645v1 Announce Type: new Abstract: Convolutional neural networks (CNNs) often exhibit poor generalisation in limited training data scenarios due to overfitting and insufficient feature di

researcharxiv-cs-cv
17 Apr 2026
Research

CI-CBM: Class-Incremental Concept Bottleneck Model for Interpretable Continual Learning

DGX agent

arXiv:2604.14519v1 Announce Type: cross Abstract: Catastrophic forgetting remains a fundamental challenge in continual learning, in which models often forget previous knowledge when fine-tuned on a ne

researcharxiv-cs-cv
17 Apr 2026
Model Releases

Class Unlearning via Depth-Aware Removal of Forget-Specific Directions

DGX agent

arXiv:2604.15166v1 Announce Type: new Abstract: Machine unlearning aims to remove targeted knowledge from a trained model without the cost of retraining from scratch. In class unlearning, however, red

model-releasesarxiv-cs-cv
17 Apr 2026
Research

CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation

DGX agent

arXiv:2604.14630v1 Announce Type: new Abstract: Recent advances in unsupervised video object segmentation have highlighted the potential of two-stream architectures that integrate appearance and motio

researcharxiv-cs-cv
17 Apr 2026
Research

Co-distilled attention guided masked image modeling with noisy teacher for self-supervised learning on medical images

DGX agent

arXiv:2604.14506v1 Announce Type: new Abstract: Masked image modeling (MIM) is a highly effective self-supervised learning (SSL) approach to extract useful feature representations from unannotated dat

researcharxiv-cs-cv
17 Apr 2026
Model Releases

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling

DGX agent

arXiv:2604.15086v1 Announce Type: cross Abstract: Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-grained cont

model-releasesarxiv-cs-cv
17 Apr 2026
Safety

Controllable Video Object Insertion via Multiview Priors

DGX agent

arXiv:2604.14556v1 Announce Type: new Abstract: Video object insertion is a critical task for dynamically inserting new objects into existing environments. Previous video generation methods focus prim

safetyarxiv-cs-cv
17 Apr 2026
Agents

CooperDrive: Enhancing Driving Decisions Through Cooperative Perception

DGX agent

arXiv:2604.14454v1 Announce Type: cross Abstract: Autonomous vehicles equipped with robust onboard perception, localization, and planning still face limitations in occlusion and non-line-of-sight (NLO

agentsarxiv-cs-cv
17 Apr 2026
Model Releases

Cross Paradigm Representation and Alignment Transformer for Image Deraining

DGX agent

arXiv:2504.16455v2 Announce Type: replace Abstract: Transformer-based networks have achieved strong performance in low-level vision tasks like image deraining by utilizing spatial or channel-wise self

model-releasesarxiv-cs-cv
17 Apr 2026
Safety

Crowdsourcing of Real-world Image Annotation via Visual Properties

DGX agent

arXiv:2604.14449v1 Announce Type: new Abstract: Recent advances in data-centric artificial intelligence highlight inherent limitations in object recognition datasets. One of the primary issues stems f

safetyarxiv-cs-cv
17 Apr 2026
Research

Data Synthesis Improves 3D Myotube Instance Segmentation

DGX agent

arXiv:2604.14720v1 Announce Type: new Abstract: Myotubes are multinucleated muscle fibers serving as key model systems for studying muscle physiology, disease mechanisms, and drug responses. Mechanist

researcharxiv-cs-cv
17 Apr 2026
Local Ai

Deepfake Detection Generalization with Diffusion Noise

DGX agent

arXiv:2604.14570v1 Announce Type: new Abstract: Deepfake detectors face growing challenges in generalization as new image synthesis techniques emerge. In particular, deepfakes generated by diffusion m

local-aiarxiv-cs-cv
17 Apr 2026
Research

Design and Validation of a Low-Cost Smartphone Based Fluorescence Detection Platform Compared with Conventional Microplate Readers

DGX agent

arXiv:2604.14527v1 Announce Type: new Abstract: A low cost fluorescence-based optical system is developed for detecting the presence of certain microorganisms and molecules within a diluted sample. A

researcharxiv-cs-cv
17 Apr 2026
Tutorials

DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts

DGX agent

arXiv:2604.14684v1 Announce Type: new Abstract: Visual prompted object detection enables interactive and flexible definition of target categories, thereby facilitating open-vocabulary detection. Since

tutorialsarxiv-cs-cv
17 Apr 2026
Model Releases

DeTracker: Motion-decoupled Vehicle Detection and Tracking in Unstabilized Satellite Videos

DGX agent

arXiv:2601.09240v2 Announce Type: replace Abstract: Satellite videos provide continuous observations of surface dynamics but pose significant challenges for multi-object tracking (MOT), especially und

model-releasesarxiv-cs-cv
17 Apr 2026
Safety

Diagnosing and Improving Diffusion Models by Estimating the Optimal Loss Value

DGX agent

arXiv:2506.13763v2 Announce Type: replace-cross Abstract: Diffusion models have achieved remarkable success in generative modeling. Despite more stable training, the loss of diffusion models is not in

safetyarxiv-cs-cv
17 Apr 2026
Local Ai

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA

DGX agent

arXiv:2511.22521v2 Announce Type: replace Abstract: Document visual question answering requires models not only to answer questions correctly, but also to precisely localize answers within complex doc

local-aiarxiv-cs-cv
17 Apr 2026
Applications

DVFace: Spatio-Temporal Dual-Prior Diffusion for Video Face Restoration

DGX agent

arXiv:2604.14560v1 Announce Type: new Abstract: Video face restoration aims to enhance degraded face videos into high-quality results with realistic facial details, stable identity, and temporal coher

applicationsarxiv-cs-cv
17 Apr 2026
Agents

EchoAgent: Towards Reliable Echocardiography Interpretation with 'Eyes','Hands' and 'Minds'

DGX agent

arXiv:2604.05541v2 Announce Type: replace Abstract: Reliable interpretation of echocardiography (Echo) is crucial for assessing cardiac function, which demands clinicians to synchronously orchestrate

agentsarxiv-cs-cv
17 Apr 2026
← Previous
1…235236237238239…261
Next →