AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Safety

Case-Aware Medical Image Classification with Multimodal Knowledge Graphs and Reliability-Guided Refinement

DGX agent

arXiv:2605.22547v1 Announce Type: new Abstract: Deep learning has brought significant progress to medical image classification, yet most existing methods still rely on isolated visual evidence and can

safetyarxiv-cs-cv
22 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation

DGX agent

arXiv:2602.02214v3 Announce Type: replace Abstract: To achieve real-time interactive video generation, current methods distill pretrained bidirectional video diffusion models into few-step autoregress

researcharxiv-cs-cv
22 May 2026
Research

Cell Phantom Video Generation in Elliptical Fourier Descriptor Domain

DGX agent

arXiv:2605.22563v1 Announce Type: new Abstract: Training Deep Neural Networks for tracking individual cells in biomedical videos requires a large amount of annotated data. The annotation of videos for

researcharxiv-cs-cv
22 May 2026
Safety

Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models

DGX agent

arXiv:2505.16416v3 Announce Type: replace Abstract: Rotary Position Embedding (RoPE) is widely adopted in large language models, but when applied to vision-language models (VLMs) it couples text and i

safetyarxiv-cs-cv
22 May 2026
Model Releases

COCOTree: A Dataset and Benchmark for Open Tree-Structured Visual Decomposition

DGX agent

arXiv:2605.22068v1 Announce Type: new Abstract: We formalize and enable the task of open tree decomposition, which segments an image into hierarchical trees of visual components with unconstrained gra

model-releasesarxiv-cs-cv
22 May 2026
Tutorials

Conceptualizing Embeddings: Sparse Disentanglement for Vision-Language Models

DGX agent

arXiv:2605.22679v1 Announce Type: new Abstract: Vision-language models learn powerful multimodal embeddings, yet their internal semantics remain opaque. While sparse autoencoders (SAEs) can extract in

tutorialsarxiv-cs-cv
22 May 2026
Research

ConvNeXt-FD: A Fractal-Based Deep Model for Robust Biomedical Image Segmentation

DGX agent

arXiv:2605.22002v1 Announce Type: new Abstract: Biomedical image segmentation is a critical task in medical diagnosis and treatment planning, enabling precise delineation of anatomical structures and

researcharxiv-cs-cv
22 May 2026
Tutorials

Cross-Domain Human Action Recognition from Multiview Motion and Textual Descriptions

DGX agent

arXiv:2605.22697v1 Announce Type: new Abstract: Robustness to domain changes is a key capability for effective deployment of human action recognition systems in real-world scenarios, where action cate

tutorialsarxiv-cs-cv
22 May 2026
Model Releases

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models

DGX agent

arXiv:2605.21854v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have rapidly converged on a small set of architectural patterns: discrete-token autoregression (e.g. OpenVLA) and co

model-releasesarxiv-cs-cv
22 May 2026
Applications

CryoNet: A Deep Learning Framework for Multi-Modal Debris-Covered Glacier Mapping. A Case Study of the Poiqu Basin, Central Himalaya

DGX agent

arXiv:2605.21527v1 Announce Type: cross Abstract: Glaciers play a critical role as freshwater reserves and indicators of climate change, yet their automatic delineation, especially for debris-covered

applicationsarxiv-cs-cv
22 May 2026
Research

D3Seg: Dependency-Aware Diffusion for Brain Tumor Segmentation with Missing Modalities

DGX agent

arXiv:2605.22249v1 Announce Type: new Abstract: Accurate brain tumor segmentation using multiparametric MRI is critical for effective treatment planning. However, in clinical settings, complete acquis

researcharxiv-cs-cv
22 May 2026
Research

Decoupling Ego-Motion from Target Dynamics via Dual-Interval Motion Cues for UAV Detection

DGX agent

arXiv:2605.22605v1 Announce Type: cross Abstract: Object detection from Unmanned Aerial Vehicles (UAVs) is challenged by severe ego-motion, camera jitter, and large scale variations. While modern dete

researcharxiv-cs-cv
22 May 2026
Research

DecQ: Detail-Condensing Queries for Enhanced Reconstruction and Generation in Representation Autoencoders

DGX agent

arXiv:2605.22777v1 Announce Type: new Abstract: Representation Autoencoders (RAEs) leverage frozen vision foundation models (VFMs) as tokenizer encoders, providing robust high-level representations th

researcharxiv-cs-cv
22 May 2026
Applications

Demystifying Transition Matching: When and Why It Can Beat Flow Matching

DGX agent

arXiv:2510.17991v3 Announce Type: replace-cross Abstract: Flow Matching (FM) underpins many state-of-the-art generative models, yet recent results indicate that Transition Matching (TM) can achieve hi

applicationsarxiv-cs-cv
22 May 2026
Safety

Depth Augmented and FE Free 3D/2D Liver Registration for Laparoscopic Liver AR

DGX agent

arXiv:2602.17517v2 Announce Type: replace Abstract: Augmented reality (AR) guidance in laparoscopic liver surgery requires accurate registration of preoperative 3D models to intraoperative 2D video, b

safetyarxiv-cs-cv
22 May 2026
Research

Detection of Virus and Small Cell Patches in Foci Images Using Switchable Convolution and Feature Pyramid Networks

DGX agent

arXiv:2605.22290v1 Announce Type: new Abstract: Accurate detection and counting of virus patches in focus-forming unit (FFU) images, also known as foci images, are important for quantifying viral infe

researcharxiv-cs-cv
22 May 2026
Agents

Diffusion-guided Generalizable Enhancer for Urban Scene Reconstruction

DGX agent

arXiv:2605.22420v1 Announce Type: new Abstract: Urban scene reconstruction from real-world observations has emerged as a powerful tool for self-driving development and testing. While current neural re

agentsarxiv-cs-cv
22 May 2026
Research

Direct content-based retrieval from music scores images

DGX agent

arXiv:2605.22255v1 Announce Type: new Abstract: The digitization of musical scores plays a crucial role in their preservation and accessibility, yet information retrieval still depends mainly on metad

researcharxiv-cs-cv
22 May 2026
Model Releases

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

DGX agent

arXiv:2510.08759v2 Announce Type: replace Abstract: Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, exi

model-releasesarxiv-cs-cv
22 May 2026
Safety

Distributed Image Compression with Multimodal Side Information at Extremely Low Bitrates

DGX agent

arXiv:2605.22061v1 Announce Type: new Abstract: Distributed Image Compression (DIC) is crucial for multi-view transmission, especially when operating at extremely low bitrates (< 0.1 bpp). Its core ch

safetyarxiv-cs-cv
22 May 2026
Model Releases

Diverse Yet Consistent: Context-Guided Diffusion with Energy-Based Joint Refinement for Multi-Agent Motion Prediction

DGX agent

arXiv:2605.22017v1 Announce Type: new Abstract: Deepgenerative models havebecomeapromisingapproach for human motion prediction due to their ability to capture multimodal distributions and represent di

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Do Vision Models Encode Object-Level Semantic Relatedness? A Cognitive Psychology-Inspired Benchmark

DGX agent

arXiv:1709.03806v2 Announce Type: replace Abstract: Modern vision models have achieved strong object-recognition performance, yet it remains unclear whether their representations encode object-level s

model-releasesarxiv-cs-cv
22 May 2026
Hardware

Don't Collapse Your Features: Why CenterLoss Hurts OOD Detection and Multi-Scale Mahalanobis Wins

DGX agent

arXiv:2605.21493v1 Announce Type: cross Abstract: The ability to detect out-of-distribution (OOD) inputs is fundamental to safe deployment of machine learning systems. Yet, current methods often rely

hardwarearxiv-cs-cv
22 May 2026
Model Releases

Dual-Integrated Low-Latency Single-Lens Infrared Computational Imaging for Object Detection

DGX agent

arXiv:2605.21964v1 Announce Type: new Abstract: Computational imaging enables compact infrared systems, but deep-learning pipelines that combine image reconstruction and object detection often introdu

model-releasesarxiv-cs-cv
22 May 2026
Hardware

EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation

DGX agent

arXiv:2605.22051v1 Announce Type: new Abstract: Generating high-fidelity visual effects (VFX) typically demands massive datasets and prohibitive computational power due to the intricate coupling of sp

hardwarearxiv-cs-cv
22 May 2026
Safety

Echo4DIR: 4D Implicit Heart Reconstruction from 2D Echocardiography Videos

DGX agent

arXiv:2605.22066v1 Announce Type: new Abstract: Reconstructing 4D (3D+t) cardiac geometry from sparse 2D echocardiography is highly desirable yet fundamentally challenged by geometric ambiguity and te

safetyarxiv-cs-cv
22 May 2026
Research

Enhancing Gaze Reasoning in Vision Foundation Models for Gaze Following

DGX agent

arXiv:2605.22607v1 Announce Type: new Abstract: Gaze following requires both scene understanding and gaze reasoning to localize the gaze target of an in-scene person. Recently, vision foundation model

researcharxiv-cs-cv
22 May 2026
Safety

Enhancing Multimodal Large Language Models for Safety-Critical Driving Video Analysis

DGX agent

arXiv:2605.22185v1 Announce Type: new Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding. However, thei

safetyarxiv-cs-cv
22 May 2026
Research

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding

DGX agent

arXiv:2605.22078v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have significantly advanced video understanding tasks, yet challenges remain in efficientl

researcharxiv-cs-cv
22 May 2026
Research

Entropy-Guided Self-Supervised Learning for Medical Image Classification

DGX agent

arXiv:2605.21970v1 Announce Type: cross Abstract: Accurate and robust medical image classification is paramount for early disease diagnosis and treatment planning. However, challenges such as limited

researcharxiv-cs-cv
22 May 2026
Model Releases

Event-Illumination Collaborative Low-light Image Enhancement with a High-resolution Real-world Dataset

DGX agent

arXiv:2605.22186v1 Announce Type: new Abstract: Event-based low-light image enhancement (LIE) methods mainly focus on incorporating high dynamic range (HDR) information from events while overlooking t

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

EventGait: Towards Robust Gait Recognition with Event Streams

DGX agent

arXiv:2605.22139v1 Announce Type: new Abstract: Gait recognition enables non-intrusive, privacy-preserving identification but suffers in uncontrolled environments due to illumination and motion sensit

model-releasesarxiv-cs-cv
22 May 2026
Agents

EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning

DGX agent

arXiv:2605.22208v1 Announce Type: new Abstract: Multimodal Large Language Model (MLLM)-driven image restoration agent demonstrates effectiveness in degradation coupling scenarios by flexibly selecting

agentsarxiv-cs-cv
22 May 2026
Agents

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models

DGX agent

arXiv:2605.21931v1 Announce Type: new Abstract: Recent Video Large Language Models (Video-LLMs) have demonstrated strong capabilities in video reasoning through reinforcement learning (RL). However, e

agentsarxiv-cs-cv
22 May 2026
Applications

Exposing Vulnerabilities in Visible-Infrared VLMs: A Unified Geometric Adversarial Framework with Cross-Task Transferability

DGX agent

arXiv:2605.22273v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong performance across diverse multimodal tasks, but their adversarial robustness in visible-infrared (VI

applicationsarxiv-cs-cv
22 May 2026
Model Releases

FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning

DGX agent

arXiv:2605.22552v1 Announce Type: new Abstract: Fashion image retrieval is a cornerstone of modern e-commerce systems. A unified framework that supports diverse query formats and search intentions is

model-releasesarxiv-cs-cv
22 May 2026
Local Ai

FastTab: A Fast Table Recognizer with a Tiny Recursive Module and 1D Transformers

DGX agent

arXiv:2605.22422v1 Announce Type: new Abstract: Table structure recognition (TSR) requires both table-level coherence (row/column counts, headers, spanning cells) and precise separator localization. W

local-aiarxiv-cs-cv
22 May 2026
Research

Flow-based Gaussian Splatting for Continuous-Scale Remote Sensing Image Super-Resolution

DGX agent

arXiv:2605.22147v1 Announce Type: new Abstract: High-resolution remote sensing images (RSIs) are crucial for Earth observation applications, yet acquiring them is often limited by sensor constraints a

researcharxiv-cs-cv
22 May 2026
Safety

Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information Gain

DGX agent

arXiv:2602.17186v2 Announce Type: replace Abstract: Large Vision Language Models (LVLMs) have achieved remarkable progress, yet they often suffer from language bias, producing answers without relying

safetyarxiv-cs-cv
22 May 2026
Research

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding

DGX agent

arXiv:2605.21973v1 Announce Type: new Abstract: Current Video-LLM approaches for Video Temporal Grounding (VTG) typically rely on direct timestamp generation from an unstructured visual-token stream,

researcharxiv-cs-cv
22 May 2026
Research

ForeSplat: Optimization-Aware Foresight for Feed-Forward 3D Gaussian Splatting

DGX agent

arXiv:2605.22020v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting (3DGS) models offer fast single-pass reconstruction,but scaling them to match per-scene optimization quality is funda

researcharxiv-cs-cv
22 May 2026
Model Releases

FRED: A Multi-Modal Autonomous Driving Dataset for Flooded Road Environments

DGX agent

arXiv:2605.22018v1 Announce Type: new Abstract: The Flooded Road Environments Dataset (FRED) is, to our knowledge, the first multi-modal autonomous driving dataset specifically targeting the collectio

model-releasesarxiv-cs-cv
22 May 2026
Safety

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model

DGX agent

arXiv:2605.22671v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior

safetyarxiv-cs-cv
22 May 2026
Research

From Baseline to Follow-Up: Counterfactual Spine DXA Image Synthesis in UK Biobank Using a Causal Hierarchical Variational Autoencoder

DGX agent

arXiv:2605.22649v1 Announce Type: new Abstract: Dual-energy X-ray absorptiometry (DXA) is widely used for large-scale skeletal assessment, yet learning controllable and interpretable factor-specific a

researcharxiv-cs-cv
22 May 2026
Model Releases

From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding

DGX agent

arXiv:2605.22413v1 Announce Type: new Abstract: Extracting structured information from visual documents (Visual Information Extraction, VIE) is a cornerstone of business automation. While recent Multi

model-releasesarxiv-cs-cv
22 May 2026
Agents

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation

DGX agent

arXiv:2605.22036v1 Announce Type: new Abstract: Despite significant progress in Vision-Language Navigation (VLN), existing approaches still rely on dense RGB videos that produce excessive patch tokens

agentsarxiv-cs-cv
22 May 2026
Local Ai

GALAR-TemporalNet v2: Anatomy-Guided Dual-Branch Temporal Classification with Bidirectional Mamba and Dual-Graph GCN for Video Capsule Endoscopy -- after competition results

DGX agent

arXiv:2605.22209v1 Announce Type: new Abstract: Video Capsule Endoscopy (VCE) poses a challenging multi-label temporal classification problem, requiring simultaneous localization of 8 anatomical regio

local-aiarxiv-cs-cv
22 May 2026
Research

GazePrior: Zero-Shot AR/VR Eye Tracking via Learned 3D Gaze Reconstruction

DGX agent

arXiv:2605.22359v1 Announce Type: new Abstract: Eye tracking (ET) is a foundational technology for advanced AR/VR applications. However, training ET models for every new ET device is challenging: real

researcharxiv-cs-cv
22 May 2026
← Previous
1…146147148149150…263
Next →