AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Local Ai

From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes

DGX agent

arXiv:2603.03577v2 Announce Type: replace Abstract: Detecting and segmenting novel object instances in open-world environments is a fundamental problem in robotic perception. Given only a small set of

local-aiarxiv-cs-cv
15 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editing

DGX agent

arXiv:2605.15181v1 Announce Type: new Abstract: Modern image editing models produce realistic results but struggle with abstract, multi step instructions (e.g., ``make this advertisement more vegetari

agentsarxiv-cs-cv
15 May 2026
Applications

From Sparse to Dense: Spatio-Temporal Fusion for Multi-View 3D Human Pose Estimation with DenseWarper

DGX agent

arXiv:2605.14525v1 Announce Type: new Abstract: In multi-view 3D human pose estimation, models typically rely on images captured simultaneously from different camera views to predict a pose at a speci

applicationsarxiv-cs-cv
15 May 2026
Applications

From Street View to Visual Network: Mapping the Visibility of Urban Landmarks with Vision-Language Models

DGX agent

arXiv:2505.11809v3 Announce Type: replace Abstract: Visibility analysis in urban planning has traditionally relied on line-of-sight (LoS) simulations, which capture geometric occlusion. However, these

applicationsarxiv-cs-cv
15 May 2026
Model Releases

G-SHARP: Gaussian Surgical Hardware Accelerated Real-time Pipeline

DGX agent

arXiv:2512.02482v2 Announce Type: replace Abstract: We propose G-SHARP, a commercially compatible, real-time surgical scene reconstruction framework designed for minimally invasive procedures that req

model-releasesarxiv-cs-cv
15 May 2026
Research

Generating HDR Video from SDR Video

DGX agent

arXiv:2605.14703v1 Announce Type: new Abstract: The high dynamic range (HDR) video ecosystem is approaching maturity, but the problem of upconverting legacy standard dynamic range (SDR) videos persist

researcharxiv-cs-cv
15 May 2026
Safety

Generative Deep Learning for Computational Destaining and Restaining of Unregistered Digital Pathology Images

DGX agent

arXiv:2605.14251v1 Announce Type: new Abstract: Conditional generative adversarial networks (cGANs) have enabled high-fidelity computational staining and destaining of hematoxylin and eosin (H&E) in d

safetyarxiv-cs-cv
15 May 2026
Model Releases

GenExam: A Multidisciplinary Text-to-Image Exam

DGX agent

arXiv:2509.14232v5 Announce Type: replace Abstract: Exams are a fundamental test of expert-level intelligence and require integrated understanding, reasoning, and generation. Existing exam-style bench

model-releasesarxiv-cs-cv
15 May 2026
Local Ai

GeoViSTA: Geospatial Vision-Tabular Transformer for Multimodal Environment Representation

DGX agent

arXiv:2605.14406v1 Announce Type: cross Abstract: Large-scale pretraining on Earth observation imagery has yielded powerful representations of the natural and built environment. However, most existing

local-aiarxiv-cs-cv
15 May 2026
Local Ai

GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding

DGX agent

arXiv:2605.14475v1 Announce Type: new Abstract: Interpreting ultra-high-resolution (UHR) remote sensing images requires models to search for sparse and tiny visual evidence across large-scale scenes.

local-aiarxiv-cs-cv
15 May 2026
Applications

H-OmniStereo: Zero-Shot Omnidirectional Stereo Matching with Heading-Aligned Normal Priors

DGX agent

arXiv:2605.14963v1 Announce Type: new Abstract: Stereo matching on top-bottom equirectangular images provides an effective framework for full-surround perception, as vertically aligned epipolar lines

applicationsarxiv-cs-cv
15 May 2026
Model Releases

HDRFace: Rethinking Face Restoration with High-Dimensional Representation

DGX agent

arXiv:2605.14821v1 Announce Type: new Abstract: Face restoration under complex degradations still remains an ill-posed inverse problem due to severe information loss. Although diffusion models benefit

model-releasesarxiv-cs-cv
15 May 2026
Safety

HeatKV: Head-tuned KV-cache Compression for Visual Autoregressive Modeling

DGX agent

arXiv:2605.14877v1 Announce Type: new Abstract: Visual Autoregressive (VAR) models have recently demonstrated impressive image generation quality while maintaining low latency. However, they suffer fr

safetyarxiv-cs-cv
15 May 2026
Research

HERO: Hierarchical Extrapolation and Refresh for Efficient World Models

DGX agent

arXiv:2508.17588v2 Announce Type: replace Abstract: Generation-driven world models create immersive virtual environments but suffer slow inference due to the iterative nature of diffusion models. Whil

researcharxiv-cs-cv
15 May 2026
Safety

Hierarchical Image Tokenization for Multi-Scale Image Super Resolution

DGX agent

arXiv:2605.14891v1 Announce Type: new Abstract: We introduce a multi-scale Image Super Resolution (ISR) method building on recent advances in Visual Auto-Regressive (VAR) modeling. VAR models break im

safetyarxiv-cs-cv
15 May 2026
Model Releases

HiSem: Hierarchical Semantic Disentangling for Remote Sensing Image Change Captioning

DGX agent

arXiv:2605.15024v1 Announce Type: new Abstract: Remote sensing image change captioning (RSICC) aims to achieve high-level semantic understanding of genuine changes occurring between bi-temporal images

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

Hyperspectral Image Land Cover Captioning Dataset for Vision Language Models

DGX agent

arXiv:2505.12217v2 Announce Type: replace Abstract: We introduce HyperCap, the first large-scale hyperspectral captioning dataset designed to enhance model performance and effectiveness in remote sens

model-releasesarxiv-cs-cv
15 May 2026
Tutorials

IG-Diff: Complex Night Scene Restoration with Illumination-Guided Diffusion Model

DGX agent

arXiv:2605.14337v1 Announce Type: new Abstract: In nighttime circumstances, it is challenging for individuals and machines to perceive their surroundings. While prevailing image restoration methods ad

tutorialsarxiv-cs-cv
15 May 2026
Applications

ImmuVis: Hyperconvolutional Foundation Model for Imaging Mass Cytometry

DGX agent

arXiv:2602.04585v2 Announce Type: replace Abstract: We present ImmuVis, a family of efficient foundation models for imaging mass cytometry (IMC), a high-throughput multiplex imaging technology that ha

applicationsarxiv-cs-cv
15 May 2026
Research

Implicit spatial-frequency fusion of hyperspectral and lidar data via kolmogorov-arnold networks

DGX agent

arXiv:2605.14239v1 Announce Type: new Abstract: Hyperspectral image (HSI) classification is challenging in complex scenes due to spectral ambiguity, spatial heterogeneity, and the strong coupling betw

researcharxiv-cs-cv
15 May 2026
Research

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation

DGX agent

arXiv:2605.14333v1 Announce Type: new Abstract: Text and faces are among the most perceptually salient and practically important patterns in visual generation, yet they remain challenging for autoregr

researcharxiv-cs-cv
15 May 2026
Research

Iskra: A System for Inverse Geometry Processing

DGX agent

arXiv:2602.12105v2 Announce Type: replace-cross Abstract: We propose a system for differentiating through solutions to geometry processing problems. Our system differentiates a broad class of geometri

researcharxiv-cs-cv
15 May 2026
Model Releases

JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation

DGX agent

arXiv:2512.12772v2 Announce Type: replace-cross Abstract: Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models

model-releasesarxiv-cs-cv
15 May 2026
Research

Keyed Nonlinear Transform: Lightweight Privacy-Enhancing Feature Sharing for Medical Image Analysis

DGX agent

arXiv:2605.14123v1 Announce Type: cross Abstract: Feature sharing via split inference offers a lightweight alternative to federated learning for resource-constrained hospitals, but transmitted feature

researcharxiv-cs-cv
15 May 2026
Safety

KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Exploration

DGX agent

arXiv:2605.14278v1 Announce Type: new Abstract: Aligning streaming autoregressive (AR) video generators with human preferences is challenging. Existing reinforcement learning methods predominantly rel

safetyarxiv-cs-cv
15 May 2026
Safety

LATERN: Test-Time Context-Aware Explainable Video Anomaly Detection

DGX agent

arXiv:2605.15054v1 Announce Type: new Abstract: Vision-language models (VLMs) have recently emerged as a promising paradigm for video anomaly detection (VAD) due to their strong visual reasoning abili

safetyarxiv-cs-cv
15 May 2026
Agents

Learning Direct Control Policies with Flow Matching for Autonomous Driving

DGX agent

arXiv:2605.14832v1 Announce Type: cross Abstract: We present a flow-matching planner for autonomous driving that directly outputs actionable control trajectories defined by acceleration and curvature

agentsarxiv-cs-cv
15 May 2026
Research

Learning Multimodal Embeddings for Traffic Accident Prediction and Causal Estimation

DGX agent

arXiv:2512.02920v3 Announce Type: replace-cross Abstract: We consider analyzing traffic accident patterns using both road network data and satellite images aligned to road graph nodes. Previous work f

researcharxiv-cs-cv
15 May 2026
Local Ai

Learning with Semantic Priors: Stabilizing Point-Supervised Infrared Small Target Detection via Hierarchical Knowledge Distillation

DGX agent

arXiv:2605.14346v1 Announce Type: new Abstract: Single-frame Infrared Small Target Detection (ISTD) aims to localize weak targets under heavy background clutter, yet dense pixel-wise annotations are e

local-aiarxiv-cs-cv
15 May 2026
Agents

LiWi: Layering in the Wild

DGX agent

arXiv:2605.14552v1 Announce Type: new Abstract: Recent advances in generative models have empowered impressive layered image generation, yet their success is largely confined to graphic design domains

agentsarxiv-cs-cv
15 May 2026
Tutorials

Local Spatiotemporal Convolutional Network for Robust Gait Recognition

DGX agent

arXiv:2605.14548v1 Announce Type: new Abstract: Gait recognition, as a promising biometric technology, identifies individuals through their unique walking patterns and offers distinctive advantages in

tutorialsarxiv-cs-cv
15 May 2026
Safety

LPH-VTON: Resolving the Structure-Texture Dilemma of Virtual Try-On via Latent Process Handover

DGX agent

arXiv:2605.14874v1 Announce Type: new Abstract: Virtual Try-On (VTON) aims to synthesize photorealistic images of garments precisely aligned with a person's body and pose. Current diffusion-based meth

safetyarxiv-cs-cv
15 May 2026
Research

MambaRain: Multi-Scale Mamba-Attention Framework for 0-3 Hour Precipitation Nowcasting

DGX agent

arXiv:2605.14606v1 Announce Type: new Abstract: Accurate precipitation nowcasting over extended horizons (0-3 hours) is essential for disaster mitigation and operational decision-making, yet remains a

researcharxiv-cs-cv
15 May 2026
Safety

MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving

DGX agent

arXiv:2605.14201v1 Announce Type: cross Abstract: Vision-language-action (VLA) models are effective as end-to-end motion planners, but can be brittle when evaluated in closed-loop settings due to bein

safetyarxiv-cs-cv
15 May 2026
Applications

Masked Autoencoders with Limited Data: Does It Work? A Fine-Grained Bioacoustics Case Study

DGX agent

arXiv:2605.14031v1 Announce Type: cross Abstract: Bioacoustic recognition requires fine-grained acoustic understanding to distinguish similar-sounding species. However, many large-scale data repositor

applicationsarxiv-cs-cv
15 May 2026
Model Releases

Masked Next-Scale Prediction for Self-supervised Scene Text Recognition

DGX agent

arXiv:2605.14885v1 Announce Type: new Abstract: Scene Text Recognition requires modeling visual structures that evolve from coarse layouts to fine-grained character strokes. Training such models relie

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

MechVerse: Evaluating Physical Motion Consistency in Video Generation Models

DGX agent

arXiv:2605.14843v1 Announce Type: new Abstract: Text- and image-conditioned video generation models have achieved strong visual fidelity and temporal coherence, but they often fail to generate motion

model-releasesarxiv-cs-cv
15 May 2026
Research

Med-DisSeg: Dispersion-Driven Representation Learning for Fine-Grained Medical Image Segmentation

DGX agent

arXiv:2605.14579v1 Announce Type: new Abstract: Accurate medical image segmentation is fundamental to precision medicine, yet robust delineation remains challenging under heterogeneous appearances, am

researcharxiv-cs-cv
15 May 2026
Safety

Medical Report Generation: A Hierarchical Task Structure-Based Cross-Modal Causal Intervention Framework

DGX agent

arXiv:2511.02271v2 Announce Type: replace Abstract: Medical Report Generation (MRG) is a key part of modern medical diagnostics, as it automatically generates reports from radiological images to reduc

safetyarxiv-cs-cv
15 May 2026
Model Releases

MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models

DGX agent

arXiv:2605.14906v1 Announce Type: new Abstract: Memory is essential for large vision-language models (LVLMs) to handle long, multimodal interactions, with two method directions providing this capabili

model-releasesarxiv-cs-cv
15 May 2026
Tutorials

Memory-SAM: Human-Prompt-Free Tongue Segmentation via Retrieval-to-Prompt

DGX agent

arXiv:2510.15849v2 Announce Type: replace Abstract: Accurate tongue segmentation is crucial for reliable TCM analysis. Supervised models require large annotated datasets, while SAM-family models remai

tutorialsarxiv-cs-cv
15 May 2026
Research

Meschers: Geometry Processing of Impossible Objects

DGX agent

arXiv:2605.14960v1 Announce Type: cross Abstract: Impossible objects, geometric constructions that humans can perceive but that cannot exist in real life, have been a topic of intrigue in visual arts,

researcharxiv-cs-cv
15 May 2026
Safety

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models

DGX agent

arXiv:2605.14530v1 Announce Type: new Abstract: Large diffusion vision-language models (LDVLMs) have recently emerged as a promising alternative to autoregressive models, enabling parallel decoding fo

safetyarxiv-cs-cv
15 May 2026
Research

MiVE: Multiscale Vision-language features for reference-guided video Editing

DGX agent

arXiv:2605.14664v1 Announce Type: new Abstract: Reference-guided video editing takes a source video, a text instruction, and a reference image as inputs, requiring the model to faithfully apply the in

researcharxiv-cs-cv
15 May 2026
Research

MonoPRIO: Adaptive Prior Conditioning for Unified Monocular 3D Object Detection

DGX agent

arXiv:2605.14781v1 Announce Type: new Abstract: Monocular 3D object detection remains challenging because metric size and depth are underdetermined by single-view evidence, particularly under occlusio

researcharxiv-cs-cv
15 May 2026
Model Releases

MoZoo:Unleashing Video Diffusion power in animal fur and muscle simulation

DGX agent

arXiv:2605.13857v1 Announce Type: cross Abstract: The creation of cinematic-quality animal effects necessitates the precise modeling of muscle and fur dynamics, a process that remains both labor-inten

model-releasesarxiv-cs-cv
15 May 2026
Research

Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval

DGX agent

arXiv:2605.14838v1 Announce Type: new Abstract: This study focuses on weakly-supervised Video Moment Retrieval (VMR), aiming to identify a moment semantically similar to the given query within an untr

researcharxiv-cs-cv
15 May 2026
Safety

Multi-scale Coarse-to-fine Modeling for Test-time Human Motion Control

DGX agent

arXiv:2605.14935v1 Announce Type: new Abstract: We present MSCoT, a multi-scale, coarse-to-fine model for test-time human motion synthesis and control. Unlike recent approaches that rely on multiple i

safetyarxiv-cs-cv
15 May 2026
← Previous
1…169170171172173…263
Next →