AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
22 May 2026

EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation

HardwareDGX agent

arXiv:2605.22051v1 Announce Type: new Abstract: Generating high-fidelity visual effects (VFX) typically demands massive datasets and prohibitive computational power due to the intricate coupling of sp

Echo4DIR: 4D Implicit Heart Reconstruction from 2D Echocardiography Videos

SafetyDGX agent

arXiv:2605.22066v1 Announce Type: new Abstract: Reconstructing 4D (3D+t) cardiac geometry from sparse 2D echocardiography is highly desirable yet fundamentally challenged by geometric ambiguity and te

Enhancing Gaze Reasoning in Vision Foundation Models for Gaze Following

ResearchDGX agent

arXiv:2605.22607v1 Announce Type: new Abstract: Gaze following requires both scene understanding and gaze reasoning to localize the gaze target of an in-scene person. Recently, vision foundation model


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Enhancing Multimodal Large Language Models for Safety-Critical Driving Video Analysis

SafetyDGX agent

arXiv:2605.22185v1 Announce Type: new Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding. However, thei

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding

ResearchDGX agent

arXiv:2605.22078v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have significantly advanced video understanding tasks, yet challenges remain in efficientl

Entropy-Guided Self-Supervised Learning for Medical Image Classification

ResearchDGX agent

arXiv:2605.21970v1 Announce Type: cross Abstract: Accurate and robust medical image classification is paramount for early disease diagnosis and treatment planning. However, challenges such as limited

Event-Illumination Collaborative Low-light Image Enhancement with a High-resolution Real-world Dataset

Model ReleasesDGX agent

arXiv:2605.22186v1 Announce Type: new Abstract: Event-based low-light image enhancement (LIE) methods mainly focus on incorporating high dynamic range (HDR) information from events while overlooking t

EventGait: Towards Robust Gait Recognition with Event Streams

Model ReleasesDGX agent

arXiv:2605.22139v1 Announce Type: new Abstract: Gait recognition enables non-intrusive, privacy-preserving identification but suffers in uncontrolled environments due to illumination and motion sensit

EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning

AgentsDGX agent

arXiv:2605.22208v1 Announce Type: new Abstract: Multimodal Large Language Model (MLLM)-driven image restoration agent demonstrates effectiveness in degradation coupling scenarios by flexibly selecting

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models

AgentsDGX agent

arXiv:2605.21931v1 Announce Type: new Abstract: Recent Video Large Language Models (Video-LLMs) have demonstrated strong capabilities in video reasoning through reinforcement learning (RL). However, e

Exposing Vulnerabilities in Visible-Infrared VLMs: A Unified Geometric Adversarial Framework with Cross-Task Transferability

ApplicationsDGX agent

arXiv:2605.22273v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong performance across diverse multimodal tasks, but their adversarial robustness in visible-infrared (VI

FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning

Model ReleasesDGX agent

arXiv:2605.22552v1 Announce Type: new Abstract: Fashion image retrieval is a cornerstone of modern e-commerce systems. A unified framework that supports diverse query formats and search intentions is

FastTab: A Fast Table Recognizer with a Tiny Recursive Module and 1D Transformers

Local AiDGX agent

arXiv:2605.22422v1 Announce Type: new Abstract: Table structure recognition (TSR) requires both table-level coherence (row/column counts, headers, spanning cells) and precise separator localization. W

Flow-based Gaussian Splatting for Continuous-Scale Remote Sensing Image Super-Resolution

ResearchDGX agent

arXiv:2605.22147v1 Announce Type: new Abstract: High-resolution remote sensing images (RSIs) are crucial for Earth observation applications, yet acquiring them is often limited by sensor constraints a

Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information Gain

SafetyDGX agent

arXiv:2602.17186v2 Announce Type: replace Abstract: Large Vision Language Models (LVLMs) have achieved remarkable progress, yet they often suffer from language bias, producing answers without relying

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding

ResearchDGX agent

arXiv:2605.21973v1 Announce Type: new Abstract: Current Video-LLM approaches for Video Temporal Grounding (VTG) typically rely on direct timestamp generation from an unstructured visual-token stream,

ForeSplat: Optimization-Aware Foresight for Feed-Forward 3D Gaussian Splatting

ResearchDGX agent

arXiv:2605.22020v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting (3DGS) models offer fast single-pass reconstruction,but scaling them to match per-scene optimization quality is funda

FRED: A Multi-Modal Autonomous Driving Dataset for Flooded Road Environments

Model ReleasesDGX agent

arXiv:2605.22018v1 Announce Type: new Abstract: The Flooded Road Environments Dataset (FRED) is, to our knowledge, the first multi-modal autonomous driving dataset specifically targeting the collectio

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model

SafetyDGX agent

arXiv:2605.22671v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior

From Baseline to Follow-Up: Counterfactual Spine DXA Image Synthesis in UK Biobank Using a Causal Hierarchical Variational Autoencoder

ResearchDGX agent

arXiv:2605.22649v1 Announce Type: new Abstract: Dual-energy X-ray absorptiometry (DXA) is widely used for large-scale skeletal assessment, yet learning controllable and interpretable factor-specific a

From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding

Model ReleasesDGX agent

arXiv:2605.22413v1 Announce Type: new Abstract: Extracting structured information from visual documents (Visual Information Extraction, VIE) is a cornerstone of business automation. While recent Multi

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation

AgentsDGX agent

arXiv:2605.22036v1 Announce Type: new Abstract: Despite significant progress in Vision-Language Navigation (VLN), existing approaches still rely on dense RGB videos that produce excessive patch tokens

GALAR-TemporalNet v2: Anatomy-Guided Dual-Branch Temporal Classification with Bidirectional Mamba and Dual-Graph GCN for Video Capsule Endoscopy -- after competition results

Local AiDGX agent

arXiv:2605.22209v1 Announce Type: new Abstract: Video Capsule Endoscopy (VCE) poses a challenging multi-label temporal classification problem, requiring simultaneous localization of 8 anatomical regio

GazePrior: Zero-Shot AR/VR Eye Tracking via Learned 3D Gaze Reconstruction

ResearchDGX agent

arXiv:2605.22359v1 Announce Type: new Abstract: Eye tracking (ET) is a foundational technology for advanced AR/VR applications. However, training ET models for every new ET device is challenging: real

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation

SafetyDGX agent

arXiv:2605.21605v1 Announce Type: new Abstract: Open-ended image generation is no longer a simple prompt-to-image problem. High-quality generation often requires an agent to combine a model's internal

GenHAR: Generalizing Cross-domain Human Activity Recognition for Last-mile Delivery

ApplicationsDGX agent

arXiv:2605.22086v1 Announce Type: new Abstract: Human Activity Recognition (HAR) has shown remarkable effectiveness in various applications, such as smart healthcare and intelligent manufacturing. How

GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning

Model ReleasesDGX agent

arXiv:2605.22558v1 Announce Type: new Abstract: Spatio-temporal reasoning in vision-language models requires visual representations that preserve physical geometry rather than merely semantic appearan

GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations

ApplicationsDGX agent

arXiv:2605.22812v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown strong potential for general-purpose robot manipulation by unifying perception and action. However, exi

GLeVE: Graph-Guided Lesion Grounding with Proposal Verification in 3D CT

SafetyDGX agent

arXiv:2605.22619v1 Announce Type: new Abstract: Grounding radiology report descriptions to 3D CT volumes is essential for verifiable clinical interpretation, yet remains challenging due to the semanti

Guided Trajectory Optimization with Sparse Scaling for Test-Time Diffusion

ResearchDGX agent

arXiv:2605.21907v1 Announce Type: new Abstract: The efficient Test-Time Scaling (TTS) paradigm offers a promising perspective for enhancing the generation performance of diffusion models. However, cur

H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning

Model ReleasesDGX agent

arXiv:2605.22629v1 Announce Type: new Abstract: Parametric human models capture global pose but cannot represent the non-rigid surface dynamics of clothing and soft tissue. Generic scene flow estimate

Hierarchical Variational Policies for Reward-Guided Diffusion

SafetyDGX agent

arXiv:2605.21661v1 Announce Type: cross Abstract: Adapting pretrained diffusion models to downstream objectives such as inverse problems often requires expensive test-time guidance or optimization. We

How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing

Model ReleasesDGX agent

arXiv:2602.01851v2 Announce Type: replace Abstract: Recent generative models have achieved remarkable progress in image editing. However, existing systems and benchmarks remain largely text-guided. In

HyperBench: Standardizing and Scaling Synthetic Evaluation for Hyperspectral Super-Resolution

ApplicationsDGX agent

arXiv:2605.21671v1 Announce Type: cross Abstract: Hyperspectral super-resolution (HSR) reconstructs a high-spatial-resolution hyperspectral image by fusing a low-resolution hyperspectral image (LR-HSI

Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors

ResearchDGX agent

arXiv:2605.22272v1 Announce Type: cross Abstract: Whole-body Humanoid-Object Interaction (HOI) is bottlenecked by the scarcity of high-fidelity 3D data. While video generative priors offer a promising

Impact of Atmospheric Turbulence and Pointing Error on Earth Observation

ApplicationsDGX agent

arXiv:2605.22268v1 Announce Type: cross Abstract: Earth Observation (EO) imagery is often degraded by atmospheric turbulence and pointing jitter; yet, these effects are rarely considered in datasets u

Improved DDIM Sampling with Moment Matching Gaussian Mixtures

ResearchDGX agent

arXiv:2311.04938v5 Announce Type: replace Abstract: We propose using a Gaussian Mixture Model (GMM) as reverse transition operator (kernel) within the Denoising Diffusion Implicit Models (DDIM) framew

Improving 3D Labeling in Self-Driving by Inferring Vehicle Information using Vision Language Models

ResearchDGX agent

arXiv:2605.21747v1 Announce Type: new Abstract: We present an approach to improve 3D vehicle labeling in self-driving applications through zero-shot inference of vehicle information, leveraging Vehicl

Improving Viewpoint-Invariance and Temporal Consistency for Action Detection

ResearchDGX agent

arXiv:2605.22695v1 Announce Type: new Abstract: Viewpoint change invariance and action temporal consistency are critical aspects for the effective deployment of human action detection of untrimmed vid

InfVSR: Breaking Length Limits of Generic Video Super-Resolution

Model ReleasesDGX agent

arXiv:2510.00948v2 Announce Type: replace Abstract: Real-world videos often extend over thousands of frames. Existing generative video super-resolution (VSR) approaches, however, face two persistent c

Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information Flow

ResearchDGX agent

arXiv:2605.21980v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) represent a significant leap towards empathetic agents, demonstrating remarkable capabilities in emotion understand

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

Model ReleasesDGX agent

arXiv:2605.22080v1 Announce Type: new Abstract: We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF material

Label tree semantic losses for rich multi-class medical image segmentation

ResearchDGX agent

arXiv:2507.15777v3 Announce Type: replace Abstract: Rich and accurate medical image segmentation is poised to underpin the next generation of AI-defined clinical practice by delineating critical anato

LACO: Adaptive Latent Communication for Collaborative Driving

SafetyDGX agent

arXiv:2605.22504v1 Announce Type: cross Abstract: Collaborative driving aims to improve safety and efficiency by enabling connected vehicles to coordinate under partial observability. Recent approache

Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models

Model ReleasesDGX agent

arXiv:2605.21861v1 Announce Type: new Abstract: Multi-modality medical vision (MV) foundation models (FM) are fundamentally challenged by pronounced Non-IID feature statistics across heterogeneous ima

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.21988v1 Announce Type: new Abstract: Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single-frame cues

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

Model ReleasesDGX agent

arXiv:2605.21573v1 Announce Type: new Abstract: We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with

LFX: Towards Unified Light Field Dense Semantic Segmentation and Salient Object Detection

ResearchDGX agent

arXiv:2503.00747v2 Announce Type: replace Abstract: Light field cameras capture multi-view observations within a single exposure. However, existing studies are typically tailored to specific LF repres

LongVT: Incentivizing 'Thinking with Long Videos' via Native Tool Calling

Model ReleasesDGX agent

arXiv:2511.20785v3 Announce Type: replace Abstract: Large multimodal models (LMMs) have shown great potential for video reasoning with textual Chain-of-Thought. However, they remain vulnerable to hall

Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming

Local AiDGX agent

arXiv:2605.21652v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have significantly advanced medical visual question answering, yet their performance in ultrasound remains suboptimal. In

LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model

Model ReleasesDGX agent

arXiv:2605.22089v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on sp

MagicFuse: Single Image Fusion for Visual and Semantic Reinforcement

TutorialsDGX agent

arXiv:2602.01760v2 Announce Type: replace Abstract: This paper focuses on a highly practical scenario: how to continue benefiting from the advantages of multi-modal image fusion under harsh conditions

Making the Discrete Continuous: Synthetic RAW Augmentations for Fine-Grained Evaluation of Person Detection Performance in Low Light

SafetyDGX agent

arXiv:2605.22455v1 Announce Type: new Abstract: Real-world deployment of AI vision models is both fueled and limited by the data available for training and testing. Real datasets are sparse and uneven

Mapping Tomato Cropping Systems in California Using AlphaEarth Geospatial Embeddings and Deep Learning Analysis

SafetyDGX agent

arXiv:2605.21804v1 Announce Type: cross Abstract: Field-scale crop maps support supply-chain forecasting and policy, yet statewide crop identification still often depends on retrospective surveys or r

MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation

Model ReleasesDGX agent

arXiv:2605.22469v1 Announce Type: new Abstract: Evaluating single-concept personalization in text-to-image diffusion requires measuring both concept preservation, which captures identity fidelity to a

Matching with Deliberation: Test-Time Evolutionary Hierarchical Multi-Agents for Zero-Shot Compositional Image Retrieval

Model ReleasesDGX agent

arXiv:2605.22478v1 Announce Type: new Abstract: Zero-Shot Compositional Image Retrieval (ZS-CIR) requires both preserving the visual continuity of the reference image and faithfully executing the sema

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

Model ReleasesDGX agent

arXiv:2605.21917v1 Announce Type: new Abstract: Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues

Model ReleasesDGX agent

arXiv:2605.21954v1 Announce Type: new Abstract: Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal la

Moment-Reenacting: Inverse Motion Degradation with Cross-shutter Guidance

ApplicationsDGX agent

arXiv:2605.22423v1 Announce Type: new Abstract: Motion degradation, manifested as blur in global shutter (GS) images or rolling shutter (RS) distortion in RS counterparts, remains a fundamental challe

MotiMotion: Motion-Controlled Video Generation with Visual Reasoning

Model ReleasesDGX agent

arXiv:2605.22818v1 Announce Type: new Abstract: Current motion-controlled image-to-video generation models rigidly follow user-provided trajectories that are often sparse, imprecise, and causally inco

← Previous
1…117118119120121…211
Next →