AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
22 May 2026

MotionDPS: Motion-Compensated 3D Brain MRI Reconstruction

ResearchDGX agent

arXiv:2605.22121v1 Announce Type: new Abstract: Magnetic resonance imaging (MRI) is highly susceptible to patient motion due to its relatively long acquisition times and the fact that data are acquire

MOTOR: A Multimodal Dataset for Two-Wheeler Rider Behavior Understanding

Model ReleasesDGX agent

arXiv:2605.22550v1 Announce Type: new Abstract: Two-wheelers account for a disproportionately high share of road fatalities in the Global South. Research on two-wheeler rider behavior, however, lags f

MRecover: A Conditional Generative Model for Recovering Motion-Corrupted MR images Using AI Generated Contrast

ResearchDGX agent

arXiv:2605.21669v1 Announce Type: new Abstract: Hippocampal subfield segmentation requires high-resolution T2w turbo spin echo (TSE) MRI, yet this sequence is susceptible to motion artifacts, leading


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering

ResearchDGX agent

arXiv:2605.22269v1 Announce Type: new Abstract: Long streaming video QA remains challenging due to growing visual tokens and limited reasoning length of large language models (LLMs). KV-caching stores

Multi-scale interaction network for stereo image super-resolution

ResearchDGX agent

arXiv:2605.21913v1 Announce Type: new Abstract: Stereo image super-resolution aims to generate high-resolution images by leveraging complementary information from binocular systems. Although previous

No Pose, No Problem in 4D: Feed-Forward Dynamic Gaussians from Unposed Multi-View Videos

ResearchDGX agent

arXiv:2605.22190v1 Announce Type: new Abstract: Recent feed-forward 3D gaussian splatting methods have made dramatic progress on individual aspects of 3D scene reconstruction, but no existing method j

Not All Starting Points Are Equal: Pre-trained Priors and Their Outsized Impact on Person Identification

Model ReleasesDGX agent

arXiv:2507.17640v3 Announce Type: replace Abstract: Recent years have seen an explosion of diverse general purpose pre-training methodologies for computer vision. However, the impact that these pre-tr

One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2605.22144v1 Announce Type: new Abstract: Existing approaches for digital short-drama production typically rely on one-shot LLM generated scripts and loosely coupled pipelines, which fail to sat

OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization

AgentsDGX agent

arXiv:2605.22104v1 Announce Type: new Abstract: Real-world image restoration is challenging due to complex and interacting mixed degradations. Recent agent-based approaches address this problem by com

ORBIS: Output-Guided Token Reduction with Distribution-Aware Matching for Video Diffusion Acceleration

HardwareDGX agent

arXiv:2605.22015v1 Announce Type: new Abstract: Diffusion Transformer (DiT) has emerged as a powerful model architecture for generating high-quality images and videos. In the case of video DiT, 3D Spa

OSS: Open Suturing Skills Vision-Based Assessment Challenge 2024-2025

Model ReleasesDGX agent

arXiv:2605.22200v1 Announce Type: new Abstract: Achieving high levels of surgical skill through effective training is essential for optimal patient outcomes. Automated, data-driven skill assessment ho

PartCo: Part-Level Correspondence Priors Enhance Category Discovery

Model ReleasesDGX agent

arXiv:2509.22769v2 Announce Type: replace Abstract: Generalized Category Discovery (GCD) aims to identify both known and novel categories within unlabeled data by leveraging a set of labeled examples

Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?

Model ReleasesDGX agent

arXiv:2605.22109v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchm

Physiology and Anatomy Aware Inverse Inference of Myocardial Infarction for Cardiac Digital Twin

Model ReleasesDGX agent

arXiv:2605.22044v1 Announce Type: new Abstract: Accurate localization of myocardial infarction is essential for risk stratification. While LGE-MRI remains the gold standard, it is resource-intensive.

PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects

SafetyDGX agent

arXiv:2605.21572v1 Announce Type: new Abstract: Simulation-ready physical 3D assets have emerged as a promising direction owing to their broad applicability in downstream tasks. However, most existing

PIU: Proximity-guided Identity Unlearning in ID-Conditioned Diffusion Models

ResearchDGX agent

arXiv:2605.22311v1 Announce Type: new Abstract: Identity-conditioned diffusion models enable high-quality and identity-consistent face generation, but they also raise severe privacy concerns, as model

PointLLM-R: Enhancing 3D Point Cloud Reasoning via Chain-of-Thought

ApplicationsDGX agent

arXiv:2605.22013v1 Announce Type: new Abstract: Understanding 3D point clouds through language remains a fundamental challenge in computer graphics and visual computing, due to the irregular structure

Pre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts

Model ReleasesDGX agent

arXiv:2605.22446v1 Announce Type: new Abstract: While large vision-language-action (VLA) models and generative world models (WM) have advanced long-horizon embodied intelligence, their practical deplo

QuantSR+: Pushing the Limit of Quantized Image Super-Resolution Networks

SafetyDGX agent

arXiv:2605.22351v1 Announce Type: new Abstract: Low-bit quantization is widely used to compress super-resolution (SR) models and reduce storage and computation costs for deployment on resource-limited

REACH: Hand Pose Estimation from Room Corners

ResearchDGX agent

arXiv:2605.22231v1 Announce Type: new Abstract: We introduce a novel 3D hand pose estimator that can accurately recover the shape and pose of people's hands in a room from afar, typically from fixed c

Rethinking Noise-Robust Training for Frozen Vision Foundation Models: A Cross-Dataset Benchmark with a Case Study of Small-Loss Failure

Model ReleasesDGX agent

arXiv:2605.22591v1 Announce Type: new Abstract: Frozen Vision Foundation Models (VFMs) with lightweight classification heads are increasingly used in medical imaging because they offer efficient and r

Rethinking Token Reduction for Diffusion Models via Output-Similarity-Awareness

ResearchDGX agent

arXiv:2605.22011v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) achieve superior image generation quality but suffer from quadratic computational complexity relative to token count. Whil

Revisiting Integration of Image and Metadata for DICOM Series Classification: Cross-Attention and Dictionary Learning

ResearchDGX agent

arXiv:2602.23833v2 Announce Type: replace-cross Abstract: Automated identification of DICOM image series is essential for large-scale medical image analysis, quality control, protocol harmonization, a

RiT: Vanilla Diffusion Transformers Suffice in Representation Space

ResearchDGX agent

arXiv:2605.21981v1 Announce Type: new Abstract: Flow matching with x-prediction -- regressing the clean data point rather than the ambient velocity -- is known to exploit low-dimensional manifold stru

RobuQ: Pushing DiTs to W1.58A2 via Robust Activation Quantization

ResearchDGX agent

arXiv:2509.23582v2 Announce Type: replace Abstract: Diffusion Transformers (DiTs) have recently emerged as a powerful backbone for image generation, demonstrating superior scalability and performance

Robustness of breast lesion segmentation under MRI undersampling improves with k-space-aware deep learning

ResearchDGX agent

arXiv:2605.22327v1 Announce Type: new Abstract: Purpose: To assess whether breast lesion segmentation can be learned directly from acquired MRI k-space, and whether doing so improves robustness when d

SADGE: Structure and Appearance Domain Gap Estimation of Synthetic and Real Data

Model ReleasesDGX agent

arXiv:2605.22467v1 Announce Type: new Abstract: We propose SADGE, a quantitative similarity metric that predicts the performance of synthetic image datasets for common computer vision tasks without do

SceneAligner: 3D-Grounded Floorplan Localization in the Wild

Local AiDGX agent

arXiv:2605.22581v1 Announce Type: new Abstract: Many public buildings provide floorplans with a 'you are here' indicator to help visitors orient themselves. Floorplan localization seeks to computation

SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching

Model ReleasesDGX agent

arXiv:2605.21788v1 Announce Type: new Abstract: Zero-shot 3D visual grounding requires localizing objects in unstructured environments from free-form natural language. Recent vision-language model (VL

SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals

Model ReleasesDGX agent

arXiv:2605.21919v1 Announce Type: new Abstract: Assessing progress toward the Sustainable Development Goals (SDGs) requires multi-step reasoning over visual cues, contextual knowledge, and development

SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers

ResearchDGX agent

arXiv:2605.22668v1 Announce Type: new Abstract: Diffusion transformers (DiTs) have emerged as a dominant architecture for text-to-image generation, yet their performance drops when generating at resol

SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning Segmentation

Local AiDGX agent

arXiv:2605.22658v1 Announce Type: new Abstract: While large language models provide strong compositional reasoning, existing reasoning segmentation pipelines fail to transparently connect this reasoni

SegGuidedNet: Sub-Region-Aware Attention Supervision for Interpretable Brain Tumor Segmentation

Model ReleasesDGX agent

arXiv:2605.22572v1 Announce Type: new Abstract: Accurate segmentation of brain tumour sub-regions from multi-parametric MRI is critical for treatment planning yet remains challenging due to morphologi

Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking

TutorialsDGX agent

arXiv:2605.22538v1 Announce Type: new Abstract: Traditional visual object tracking (VOT) methods typically rely on task-specific supervised training, limiting their generalization to unseen objects an

Seizure-Semiology-Suite (S3): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding

Model ReleasesDGX agent

arXiv:2605.21852v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in general video understanding, their capacity to interpret invo

Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving

AgentsDGX agent

arXiv:2605.22809v1 Announce Type: new Abstract: Robust training and validation of Autonomous Driving Systems (ADS) require massive, diverse datasets. Proprietary data collected by Autonomous Vehicle (

SFN-YOLO: Towards Free-Range Poultry Detection via Scale-aware Fusion Networks

Model ReleasesDGX agent

arXiv:2509.17086v2 Announce Type: replace Abstract: Detecting and localizing poultry is essential for advancing smart poultry farming. Despite the progress of detection-centric methods, challenges per

Skarimva: Skeleton-based Action Recognition is a Multi-view Application

ResearchDGX agent

arXiv:2602.23231v2 Announce Type: replace Abstract: Human action recognition plays an important role when developing intelligent interactions between humans and machines. While there is a lot of activ

Slimmable ConvNeXt: Width-Adaptive Inference for Efficient Multi-Device Deployment

ResearchDGX agent

arXiv:2605.22677v1 Announce Type: new Abstract: Deploying vision models across devices with varying resource constraints, or even on a single device where available compute fluctuates due to battery s

SO-Mamba: State-Ownership Mamba for Unrolled MRI Reconstruction

ResearchDGX agent

arXiv:2605.22031v1 Announce Type: new Abstract: Accelerated MRI reconstruction requires recovering missing details while preserving anatomically coherent structures across large spatial regions. State

SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control

Model ReleasesDGX agent

arXiv:2511.07820v3 Announce Type: replace-cross Abstract: Despite the rise of billion-parameter foundation models trained across thousands of GPUs, similar scaling gains have not been shown for humano

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving

Model ReleasesDGX agent

arXiv:2512.10719v2 Announce Type: replace Abstract: End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual under

Spectral Tail Auxiliary Learning for AI-Generated Image Detection

ApplicationsDGX agent

arXiv:2605.22751v1 Announce Type: new Abstract: As generative image models evolve rapidly, the perceptual gap between generated and real images continues to narrow, making AI-generated image detection

SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents

ResearchDGX agent

arXiv:2603.08403v3 Announce Type: replace Abstract: Long-horizon action-conditioned video generation aims to synthesize temporally coherent videos that follow complex action instructions over extended

ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs

ResearchDGX agent

arXiv:2605.22158v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) face significant computational overhead when processing long videos due to the massive number of visual token

Supervised Classification Heads as Semantic Prototypes: Unlocking Vision-Language Alignment via Weight Recycling

SafetyDGX agent

arXiv:2605.22484v1 Announce Type: new Abstract: Vision-Language Models (VLMs) excel at tasks like zero-shot classification and cross-modal retrieval by mapping images and text to a shared space, but t

Swift Sampling: Selecting Temporal Surprises via Taylor Series

ResearchDGX agent

arXiv:2605.22678v1 Announce Type: new Abstract: While most frames in long-form video are redundant, the critical information resides in temporal surprises: moments where the actual visual features dev

Synthetic Data Alone is Enough? Rethinking Data Scarcity in Pediatric Rare Disease Recognition

ApplicationsDGX agent

arXiv:2605.22767v1 Announce Type: new Abstract: Children with rare genetic diseases often exhibit distinctive facial phenotypes, yet developing computer vision systems for early diagnosis remains chal

Tackle CSM in JPEG Steganalysis with Data Adaptation

Model ReleasesDGX agent

arXiv:2605.21523v1 Announce Type: cross Abstract: Steganalysis models excel on benchmark datasets but struggle in the wild when analyzed images are produced by a processing pipeline unseen during trai

TextTeacher: What Can Language Teach About Images?

TutorialsDGX agent

arXiv:2605.22098v1 Announce Type: new Abstract: The platonic representation hypothesis suggests that sufficiently large models converge to a shared representation geometry, even across modalities. Mot

The Neglected Baseline in Model Interpretation

ResearchDGX agent

arXiv:2605.22417v1 Announce Type: new Abstract: We observe that existing model interpretation methods generally ignore the baseline, and such neglect often results in imprecise or even incorrect inter

Thermo-VL: Extending Vision-Language Models to Thermal Infrared Perception

Model ReleasesDGX agent

arXiv:2605.21882v1 Announce Type: new Abstract: Vision-language models (VLMs) often fail under low illumination because their visual grounding is learned predominantly from RGB imagery, whereas therma

Time-varying rPPG signal separation via block-sparse signal model

ResearchDGX agent

arXiv:2605.22425v1 Announce Type: cross Abstract: Remote photoplethysmography (rPPG) enables non-contact measurement of cardiac pulse signals by analyzing subtle color changes in facial videos. Nevert

Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence

Model ReleasesDGX agent

arXiv:2605.22414v1 Announce Type: new Abstract: Visual Question Answering (VQA) holds great promise for clinical support, particularly in ophthalmology, where retinal fundus photography is essential f

Towards Initialization-free Calibrated Bundle Adjustment

ResearchDGX agent

arXiv:2506.23808v2 Announce Type: replace Abstract: A recent series of works has shown that initialization-free BA can be achieved using pseudo Object Space Error (pOSE) as a surrogate objective. The

Towards Selection of Large Multimodal Models as Engines for Burned-in Protected Health Information Detection in Medical Images

Model ReleasesDGX agent

arXiv:2511.02014v2 Announce Type: replace Abstract: The detection of Protected Health Information (PHI) in medical imaging is critical for safeguarding patient privacy and ensuring compliance with reg

Training-Free Fine-Grained Semantic Segmentations in Low Data Regimes: A FungiTastic Baseline

ResearchDGX agent

arXiv:2605.22492v1 Announce Type: new Abstract: Fine-grained semantic segmentation requires both precise localization and discrimination between visually similar classes. In FungiTastic, this problem

Translating Signals to Languages for sEMG-Based Activity Recognition

ResearchDGX agent

arXiv:2605.22403v1 Announce Type: new Abstract: Surface electromyography (sEMG) signal-based activity recognition has attracted increasing research attention in recent years. To develop accurate sEMG

Transporting Task Vectors across Different Architectures without Training

Model ReleasesDGX agent

arXiv:2602.12952v2 Announce Type: replace-cross Abstract: Adapting large pre-trained models to downstream tasks often produces task-specific parameter updates that are expensive to relearn for every m

TWINGS: Thin Plate Splines Warp-aligned Initialization for Sparse-View Gaussian Splatting

ResearchDGX agent

arXiv:2605.22069v1 Announce Type: new Abstract: Novel view synthesis from sparse-view inputs poses a significant challenge in 3D computer vision, particularly for achieving high-quality scene reconstr

← Previous
1…118119120121122…211
Next →