AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
15 May 2026

From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes

Local AiDGX agent

arXiv:2603.03577v2 Announce Type: replace Abstract: Detecting and segmenting novel object instances in open-world environments is a fundamental problem in robotic perception. Given only a small set of

From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing

AgentsDGX agent

arXiv:2605.15181v1 Announce Type: new Abstract: Modern image editing models produce realistic results but struggle with abstract, multi step instructions (e.g., ``make this advertisement more vegetari

From Sparse to Dense: Spatio-Temporal Fusion for Multi-View 3D Human Pose Estimation with DenseWarper

ApplicationsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.14525v1 Announce Type: new Abstract: In multi-view 3D human pose estimation, models typically rely on images captured simultaneously from different camera views to predict a pose at a speci

From Street View to Visual Network: Mapping the Visibility of Urban Landmarks with Vision-Language Models

ApplicationsDGX agent

arXiv:2505.11809v3 Announce Type: replace Abstract: Visibility analysis in urban planning has traditionally relied on line-of-sight (LoS) simulations, which capture geometric occlusion. However, these

G-SHARP: Gaussian Surgical Hardware Accelerated Real-time Pipeline

Model ReleasesDGX agent

arXiv:2512.02482v2 Announce Type: replace Abstract: We propose G-SHARP, a commercially compatible, real-time surgical scene reconstruction framework designed for minimally invasive procedures that req

Generating HDR Video from SDR Video

ResearchDGX agent

arXiv:2605.14703v1 Announce Type: new Abstract: The high dynamic range (HDR) video ecosystem is approaching maturity, but the problem of upconverting legacy standard dynamic range (SDR) videos persist

Generative Deep Learning for Computational Destaining and Restaining of Unregistered Digital Pathology Images

SafetyDGX agent

arXiv:2605.14251v1 Announce Type: new Abstract: Conditional generative adversarial networks (cGANs) have enabled high-fidelity computational staining and destaining of hematoxylin and eosin (H&E) in d

GenExam: A Multidisciplinary Text-to-Image Exam

Model ReleasesDGX agent

arXiv:2509.14232v5 Announce Type: replace Abstract: Exams are a fundamental test of expert-level intelligence and require integrated understanding, reasoning, and generation. Existing exam-style bench

GeoViSTA: Geospatial Vision-Tabular Transformer for Multimodal Environment Representation

Local AiDGX agent

arXiv:2605.14406v1 Announce Type: cross Abstract: Large-scale pretraining on Earth observation imagery has yielded powerful representations of the natural and built environment. However, most existing

GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding

Local AiDGX agent

arXiv:2605.14475v1 Announce Type: new Abstract: Interpreting ultra-high-resolution (UHR) remote sensing images requires models to search for sparse and tiny visual evidence across large-scale scenes.

H-OmniStereo: Zero-Shot Omnidirectional Stereo Matching with Heading-Aligned Normal Priors

ApplicationsDGX agent

arXiv:2605.14963v1 Announce Type: new Abstract: Stereo matching on top-bottom equirectangular images provides an effective framework for full-surround perception, as vertically aligned epipolar lines

HDRFace: Rethinking Face Restoration with High-Dimensional Representation

Model ReleasesDGX agent

arXiv:2605.14821v1 Announce Type: new Abstract: Face restoration under complex degradations still remains an ill-posed inverse problem due to severe information loss. Although diffusion models benefit

HeatKV: Head-tuned KV-cache Compression for Visual Autoregressive Modeling

SafetyDGX agent

arXiv:2605.14877v1 Announce Type: new Abstract: Visual Autoregressive (VAR) models have recently demonstrated impressive image generation quality while maintaining low latency. However, they suffer fr

HERO: Hierarchical Extrapolation and Refresh for Efficient World Models

ResearchDGX agent

arXiv:2508.17588v2 Announce Type: replace Abstract: Generation-driven world models create immersive virtual environments but suffer slow inference due to the iterative nature of diffusion models. Whil

Hierarchical Image Tokenization for Multi-Scale Image Super Resolution

SafetyDGX agent

arXiv:2605.14891v1 Announce Type: new Abstract: We introduce a multi-scale Image Super Resolution (ISR) method building on recent advances in Visual Auto-Regressive (VAR) modeling. VAR models break im

HiSem: Hierarchical Semantic Disentangling for Remote Sensing Image Change Captioning

Model ReleasesDGX agent

arXiv:2605.15024v1 Announce Type: new Abstract: Remote sensing image change captioning (RSICC) aims to achieve high-level semantic understanding of genuine changes occurring between bi-temporal images

Hyperspectral Image Land Cover Captioning Dataset for Vision Language Models

Model ReleasesDGX agent

arXiv:2505.12217v2 Announce Type: replace Abstract: We introduce HyperCap, the first large-scale hyperspectral captioning dataset designed to enhance model performance and effectiveness in remote sens

IG-Diff: Complex Night Scene Restoration with Illumination-Guided Diffusion Model

TutorialsDGX agent

arXiv:2605.14337v1 Announce Type: new Abstract: In nighttime circumstances, it is challenging for individuals and machines to perceive their surroundings. While prevailing image restoration methods ad

ImmuVis: Hyperconvolutional Foundation Model for Imaging Mass Cytometry

ApplicationsDGX agent

arXiv:2602.04585v2 Announce Type: replace Abstract: We present ImmuVis, a family of efficient foundation models for imaging mass cytometry (IMC), a high-throughput multiplex imaging technology that ha

Implicit spatial-frequency fusion of hyperspectral and lidar data via kolmogorov-arnold networks

ResearchDGX agent

arXiv:2605.14239v1 Announce Type: new Abstract: Hyperspectral image (HSI) classification is challenging in complex scenes due to spectral ambiguity, spatial heterogeneity, and the strong coupling betw

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation

ResearchDGX agent

arXiv:2605.14333v1 Announce Type: new Abstract: Text and faces are among the most perceptually salient and practically important patterns in visual generation, yet they remain challenging for autoregr

Iskra: A System for Inverse Geometry Processing

ResearchDGX agent

arXiv:2602.12105v2 Announce Type: replace-cross Abstract: We propose a system for differentiating through solutions to geometry processing problems. Our system differentiates a broad class of geometri

JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation

Model ReleasesDGX agent

arXiv:2512.12772v2 Announce Type: replace-cross Abstract: Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models

Keyed Nonlinear Transform: Lightweight Privacy-Enhancing Feature Sharing for Medical Image Analysis

ResearchDGX agent

arXiv:2605.14123v1 Announce Type: cross Abstract: Feature sharing via split inference offers a lightweight alternative to federated learning for resource-constrained hospitals, but transmitted feature

KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration

SafetyDGX agent

arXiv:2605.14278v1 Announce Type: new Abstract: Aligning streaming autoregressive (AR) video generators with human preferences is challenging. Existing reinforcement learning methods predominantly rel

LATERN: Test-Time Context-Aware Explainable Video Anomaly Detection

SafetyDGX agent

arXiv:2605.15054v1 Announce Type: new Abstract: Vision-language models (VLMs) have recently emerged as a promising paradigm for video anomaly detection (VAD) due to their strong visual reasoning abili

Learning Direct Control Policies with Flow Matching for Autonomous Driving

AgentsDGX agent

arXiv:2605.14832v1 Announce Type: cross Abstract: We present a flow-matching planner for autonomous driving that directly outputs actionable control trajectories defined by acceleration and curvature

Learning Multimodal Embeddings for Traffic Accident Prediction and Causal Estimation

ResearchDGX agent

arXiv:2512.02920v3 Announce Type: replace-cross Abstract: We consider analyzing traffic accident patterns using both road network data and satellite images aligned to road graph nodes. Previous work f

Learning with Semantic Priors: Stabilizing Point-Supervised Infrared Small Target Detection via Hierarchical Knowledge Distillation

Local AiDGX agent

arXiv:2605.14346v1 Announce Type: new Abstract: Single-frame Infrared Small Target Detection (ISTD) aims to localize weak targets under heavy background clutter, yet dense pixel-wise annotations are e

LiWi: Layering in the Wild

AgentsDGX agent

arXiv:2605.14552v1 Announce Type: new Abstract: Recent advances in generative models have empowered impressive layered image generation, yet their success is largely confined to graphic design domains

Local Spatiotemporal Convolutional Network for Robust Gait Recognition

TutorialsDGX agent

arXiv:2605.14548v1 Announce Type: new Abstract: Gait recognition, as a promising biometric technology, identifies individuals through their unique walking patterns and offers distinctive advantages in

LPH-VTON: Resolving the Structure-Texture Dilemma of Virtual Try-On via Latent Process Handover

SafetyDGX agent

arXiv:2605.14874v1 Announce Type: new Abstract: Virtual Try-On (VTON) aims to synthesize photorealistic images of garments precisely aligned with a person's body and pose. Current diffusion-based meth

MambaRain: Multi-Scale Mamba-Attention Framework for 0-3 Hour Precipitation Nowcasting

ResearchDGX agent

arXiv:2605.14606v1 Announce Type: new Abstract: Accurate precipitation nowcasting over extended horizons (0-3 hours) is essential for disaster mitigation and operational decision-making, yet remains a

MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2605.14201v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are effective as end-to-end motion planners, but can be brittle when evaluated in closed-loop settings due to bein

Masked Autoencoders with Limited Data: Does It Work? A Fine-Grained Bioacoustics Case Study

ApplicationsDGX agent

arXiv:2605.14031v1 Announce Type: cross Abstract: Bioacoustic recognition requires fine-grained acoustic understanding to distinguish similar-sounding species. However, many large-scale data repositor

Masked Next-Scale Prediction for Self-supervised Scene Text Recognition

Model ReleasesDGX agent

arXiv:2605.14885v1 Announce Type: new Abstract: Scene Text Recognition requires modeling visual structures that evolve from coarse layouts to fine-grained character strokes. Training such models relie

MechVerse: Evaluating Physical Motion Consistency in Video Generation Models

Model ReleasesDGX agent

arXiv:2605.14843v1 Announce Type: new Abstract: Text- and image-conditioned video generation models have achieved strong visual fidelity and temporal coherence, but they often fail to generate motion

Med-DisSeg: Dispersion-Driven Representation Learning for Fine-Grained Medical Image Segmentation

ResearchDGX agent

arXiv:2605.14579v1 Announce Type: new Abstract: Accurate medical image segmentation is fundamental to precision medicine, yet robust delineation remains challenging under heterogeneous appearances, am

Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework

SafetyDGX agent

arXiv:2511.02271v2 Announce Type: replace Abstract: Medical Report Generation (MRG) is a key part of modern medical diagnostics, as it automatically generates reports from radiological images to reduc

MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.14906v1 Announce Type: new Abstract: Memory is essential for large vision-language models (LVLMs) to handle long, multimodal interactions, with two method directions providing this capabili

Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt

TutorialsDGX agent

arXiv:2510.15849v2 Announce Type: replace Abstract: Accurate tongue segmentation is crucial for reliable TCM analysis. Supervised models require large annotated datasets, while SAM-family models remai

Meschers: Geometry Processing of Impossible Objects

ResearchDGX agent

arXiv:2605.14960v1 Announce Type: cross Abstract: Impossible objects, geometric constructions that humans can perceive but that cannot exist in real life, have been a topic of intrigue in visual arts,

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models

SafetyDGX agent

arXiv:2605.14530v1 Announce Type: new Abstract: Large diffusion vision-language models (LDVLMs) have recently emerged as a promising alternative to autoregressive models, enabling parallel decoding fo

MiVE: Multiscale Vision-language features for reference-guided video Editing

ResearchDGX agent

arXiv:2605.14664v1 Announce Type: new Abstract: Reference-guided video editing takes a source video, a text instruction, and a reference image as inputs, requiring the model to faithfully apply the in

MonoPRIO: Adaptive Prior Conditioning for Unified Monocular 3D Object Detection

ResearchDGX agent

arXiv:2605.14781v1 Announce Type: new Abstract: Monocular 3D object detection remains challenging because metric size and depth are underdetermined by single-view evidence, particularly under occlusio

MoZoo:Unleashing Video Diffusion power in animal fur and muscle simulation

Model ReleasesDGX agent

arXiv:2605.13857v1 Announce Type: cross Abstract: The creation of cinematic-quality animal effects necessitates the precise modeling of muscle and fur dynamics, a process that remains both labor-inten

Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval

ResearchDGX agent

arXiv:2605.14838v1 Announce Type: new Abstract: This study focuses on weakly-supervised Video Moment Retrieval (VMR), aiming to identify a moment semantically similar to the given query within an untr

Multi-scale Coarse-to-fine Modeling for Test-time Human Motion Control

SafetyDGX agent

arXiv:2605.14935v1 Announce Type: new Abstract: We present MSCoT, a multi-scale, coarse-to-fine model for test-time human motion synthesis and control. Unlike recent approaches that rely on multiple i

MultiMat: Multimodal Program Synthesis for Procedural Materials using Large Multimodal Models

ApplicationsDGX agent

arXiv:2509.22151v3 Announce Type: replace Abstract: Material node graphs are programs that generate the 2D channels of procedural materials, including geometry such as roughness and displacement maps,

Multimodal Causal-Driven Representation Learning for Generalizable Medical Image Segmentation

ResearchDGX agent

arXiv:2508.05008v2 Announce Type: replace Abstract: Vision-Language Models (VLMs), such as CLIP, have demonstrated remarkable zero-shot capabilities in various computer vision tasks. However, their ap

Octopus: History-Free Gradient Orthogonalization for Continual Learning in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2605.14938v1 Announce Type: cross Abstract: Continual learning in multimodal large language models (MLLMs) aims to sequentially acquire knowledge while mitigating catastrophic forgetting, yet ex

PacTure: Efficient PBR Texture Generation on Packed Views with Visual Autoregressive Models

ResearchDGX agent

arXiv:2505.22394v2 Announce Type: replace Abstract: We present PacTure, a novel framework for generating physically-based rendering (PBR) material textures for an untextured 3D mesh from a text descri

PanoPlane: Plane-Aware Panoramic Completion for Sparse-View Indoor 3D Gaussian Splatting

ResearchDGX agent

arXiv:2605.14135v1 Announce Type: new Abstract: We present PanoPlane, an approach for high-fidelity sparse-view indoor novel view synthesis that reconstructs closed room geometry via panoramic scene c

Physics-Grounded Adversarial Stain Augmentation with Calibrated Coverage Guarantees

Model ReleasesDGX agent

arXiv:2605.13889v1 Announce Type: cross Abstract: Stain variation across hospitals degrades histopathology models at deployment. Existing augmentation methods perturb color spaces with arbitrary hyper

Probing into Camera Control of Video Models

Model ReleasesDGX agent

arXiv:2605.14815v1 Announce Type: new Abstract: Video is a rich and scalable source of 3D/4D visual observations, and camera control is a key capability for video generation models to produce geometri

PVRF: All-in-one Adverse Weather Removal via Prior-modulated and Velocity-constrained Rectified Flow

Model ReleasesDGX agent

arXiv:2605.14045v1 Announce Type: new Abstract: Adverse weather removal (AWR) in real-world images remains challenging due to heterogeneous and unseen degradations, while distortion-driven training of

RAM-W600: A Multi-Task Wrist Dataset and Benchmark for Rheumatoid Arthritis

Model ReleasesDGX agent

arXiv:2507.05193v4 Announce Type: replace-cross Abstract: Rheumatoid arthritis (RA) is a common autoimmune disease that has been the focus of research in computer-aided diagnosis (CAD) and disease mon

RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling

SafetyDGX agent

arXiv:2510.20206v2 Announce Type: replace Abstract: Prompt design plays a crucial role in text-to-video (T2V) generation, yet user-provided prompts are often short, unstructured, and misaligned with t

RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO

SafetyDGX agent

arXiv:2605.15190v1 Announce Type: new Abstract: Causal autoregressive video diffusion models support real-time streaming generation by extrapolating future chunks from previously generated content. Di

Real2Sim in HOI: Toward Physically Plausible HOI Reconstruction from Monocular Videos

Model ReleasesDGX agent

arXiv:2605.14462v1 Announce Type: new Abstract: Recovering 4D human-object interaction (HOI) from monocular video is a key step toward scalable 3D content creation, embodied AI, and simulation-based l

← Previous
1…135136137138139…211
Next →