AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
23 Jul 2026

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation

ApplicationsDGX agent

arXiv:2607.18789v1 Announce Type: new Abstract: Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture

MTVDiff: Multimodal Conditional Latent Diffusion for Enhanced Thermal-to-Visible Face Translation

SafetyDGX agent

arXiv:2607.19886v1 Announce Type: new Abstract: Thermal-to-visible face translation presents fundamental challenges including geometric discontinuities, semantic attribute mismatches, and identity deg

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.20284v1 Announce Type: new Abstract: The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU

MV-Bench: Benchmarking Multimodal Large Language Models for Coordinated Multi-View Interface Construction

Model ReleasesDGX agent

arXiv:2607.19910v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly expected to automate visualization development by generating code directly from visual designs

MVGD-Net: A Novel Motion-aware Video Glass Surface Detection Method

ResearchDGX agent

arXiv:2601.13715v2 Announce Type: replace Abstract: Glass surface ubiquitous in both daily life and professional environments presents a potential threat to vision-based systems, such as robot and dro

NGPS: GPS-Denied Aerial Geo-Localization and 2.5D Reconstruction via Deep Satellite Image Matching and Multi-Rate Sensor Fusion

HardwareDGX agent

arXiv:2607.18936v1 Announce Type: cross Abstract: We present NGPS (Next-Generation Positioning System), a visual geo-localization framework for high-altitude UAVs that provides GPS-free absolute posit

No Training, Better Flights: Test-Time Scaled VLMs for UAV Navigation

SafetyDGX agent

arXiv:2607.19288v1 Announce Type: new Abstract: Test-time scaling offers a promising method to improve the inference performance of Vision-Language Models (VLMs) without additional training. Existing

Norm or Direction? Decoding Vision Mambas for High-Resolution Vision

SafetyDGX agent

arXiv:2607.18625v1 Announce Type: new Abstract: Vision Mamba models replace quadratic self-attention with linear complexity selective state space models (SSMs), emerging as efficient visual backbones.

Not All Patches are Equal: Sampling Matters for Visible-Infrared Pre-Training

Model ReleasesDGX agent

arXiv:2607.20238v1 Announce Type: new Abstract: Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception. Most existing methods use uniform patch-wise contrast

Now We Know? A Systematic Comparison of TerraMind and THOR

ResearchDGX agent

arXiv:2607.18504v1 Announce Type: cross Abstract: Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ: how much

Occlusion-Aware Panoptic Segmentation with Joint Position Embedding and Occlusion-Level Attention

ResearchDGX agent

arXiv:2607.18112v2 Announce Type: replace Abstract: Panoptic segmentation in complex scenes remains challenging because of occlusions, yet modern approaches often neglect occlusion modelling. In this

OffNadirLoc: Benchmark and Framework for Challenging UAV-to-Satellite Geo-Localization under Large Off-Nadir Views

Model ReleasesDGX agent

arXiv:2607.19951v1 Announce Type: new Abstract: Cross-view geo-localization between UAV and satellite imagery remains a fundamental yet highly challenging task, especially under large off-nadir views

OmniReasoner: Thinking with Long Audio-Video via Native Tool Use

AgentsDGX agent

arXiv:2607.19339v1 Announce Type: new Abstract: Long audio-video reasoning is difficult for omnimodal LLMs because the decisive evidence is often sparse, cross-modal, and too expensive to preserve wit

OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation

Local AiDGX agent

arXiv:2607.18850v1 Announce Type: new Abstract: Large vision-language models (LVLMs) have recently shown strong potential for industrial anomaly detection (IAD) by providing image-level anomaly judgme

Open-Vocabulary Gaze Object Prediction: Benchmark and Method

Model ReleasesDGX agent

arXiv:2607.18827v1 Announce Type: new Abstract: Gaze Object Prediction (GOP) aims to localize and recognize the objects humans attend to, a task crucial for understanding human-centric interactions. H

Pain in 3D: Generating Controllable Synthetic Faces for Automated Pain Assessment

ResearchDGX agent

arXiv:2509.16727v5 Announce Type: replace Abstract: Automated pain assessment from facial expressions is crucial for non-communicative patient. Progress has been limited by two challenges: (i) existin

PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image

Model ReleasesDGX agent

arXiv:2607.19261v1 Announce Type: new Abstract: Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrating multi-scal

Pathologist Attention-Aligned Report Generation for Prostate Histopathology

SafetyDGX agent

arXiv:2607.19624v1 Announce Type: new Abstract: The allocation of visual attention by pathologists during cancer diagnosis is a highly selective process that critically shapes the information extracte

PathReportEval: A Systematic Benchmark for Pathology Report Generation

Model ReleasesDGX agent

arXiv:2607.18448v1 Announce Type: cross Abstract: Pathology report generation from whole-slide images (WSIs) is a rapidly growing multimodal learning problem, yet progress is difficult to measure beca

PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration

SafetyDGX agent

arXiv:2607.16602v2 Announce Type: replace Abstract: Action-conditioned world models are a key component of embodied AI, serving as scalable policy evaluators that reduce reliance on expensive real-wor

PC-Seg: Progressive Cross-View Consistency for 3D OCT Segmentation from Sparse 2D Annotations

TutorialsDGX agent

arXiv:2607.17718v2 Announce Type: replace Abstract: Volumetric segmentation of optical coherence tomography (OCT) images is essential for diagnosing ocular diseases but requires labor-intensive voxel-

PercepCap: Video Captioner with Structured Spatio-Temporal Perception

ResearchDGX agent

arXiv:2607.20389v1 Announce Type: new Abstract: Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal p

PerceptDrive: Perception Prior World-Action Modeling with Adaptive Expert Routing for End-to-End Autonomous Driving

AgentsDGX agent

arXiv:2607.20175v1 Announce Type: new Abstract: Frozen perception foundation models encode rich geometric, semantic, and dynamic knowledge. Yet narrow conditioning interfaces may attenuate task-releva

Physics Closure Matters for Machine Olfaction: A Maxwell--Stefan Graph Solver for Identifiable Dynamic Gas Unmixing

Local AiDGX agent

arXiv:2607.18544v1 Announce Type: new Abstract: Machine olfaction for gas unmixing is an underconstrained inverse problem in which gas compositions must be inferred from low-dimensional, delayed, and

Pixel-Space Diffusion Transformers

ResearchDGX agent

arXiv:2607.17585v2 Announce Type: replace Abstract: Latent diffusion models (LDMs) enable efficient high-resolution image synthesis by denoising in a VAE-compressed latent space. However, fixed visual

Point Ladder Tuning: Parameter-Efficient Hierarchical Adaptation for 3D Point Cloud Understanding

Model ReleasesDGX agent

arXiv:2607.19171v1 Announce Type: new Abstract: Fine-tuning pre-trained point-cloud backbones typically updates all parameters, resulting in substantial computation and memory overhead. More important

Point-Selection Fine-Tuning Framework for Robust Point Cloud Classification

Model ReleasesDGX agent

arXiv:2607.19711v1 Announce Type: new Abstract: Noisy and corrupted points can substantially degrade point cloud recognition performance, especially under challenging corruption settings. In particula

Pointing-Based Object Recognition

ResearchDGX agent

arXiv:2603.15403v2 Announce Type: replace Abstract: This paper presents a comprehensive pipeline for recognizing objects targeted by human pointing gestures using RGB images. As human-robot interactio

PoseIDON: 6DoF Pose Estimation with Foundation Model Features for Marine Sediment Burial Mapping

Local AiDGX agent

arXiv:2506.10386v2 Announce Type: replace Abstract: The burial state of anthropogenic objects on the seafloor provides insight into localized sedimentation dynamics and is also critical for assessing

Posterior Samplings are Missing Modalities Generators for Medical Image Translation

ApplicationsDGX agent

arXiv:2607.18763v1 Announce Type: new Abstract: Magnetic resonance imaging comes in various modality contrasts that provide complementary anatomical and pathological information. Complete multimodal a

PRiSM: Prototype Regularization for Few-Shot VLMs

Model ReleasesDGX agent

arXiv:2607.17820v2 Announce Type: replace Abstract: Training-free few-shot adaptation methods have gained significant attention recently in the context of Vision-language Models (VLMs). Yet, current b

Privileged Lesion-Context Relational Distillation for Mask-Free Skin Lesion Classification

SafetyDGX agent

arXiv:2607.18773v1 Announce Type: new Abstract: Accurate skin lesion classification can benefit from lesion segmentation masks, but requiring masks or an auxiliary segmentation model during inference

Rarity-Aware Discrete Diffusion with Spatially Consistent Decoding for Photo-Realistic Image Super-Resolution

Model ReleasesDGX agent

arXiv:2607.17612v2 Announce Type: replace Abstract: Continuous diffusion models have become the dominant paradigm for photo-realistic image Super-Resolution (SR), but they typically formulate reconstr

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models

Model ReleasesDGX agent

arXiv:2512.11899v2 Announce Type: replace Abstract: Large vision-language models (LVLMs) are vulnerable to typographic attacks, where misleading text inserted into an image can override visual underst

Real-Time EEG Cap Electrode Detection for Guided Point-of-Care Placement

Local AiDGX agent

arXiv:2607.20142v1 Announce Type: new Abstract: We present a two-stage vision system that detects EEG cap electrodes in a live webcam stream and validates their anatomical placement in real time. A si

Recti-Q: Feature-Space Rectification for Out-of-Distribution-Robust Quantized Perception in Edge Robotics

Model ReleasesDGX agent

arXiv:2607.18540v1 Announce Type: new Abstract: Robotic perception pipelines increasingly rely on large vision backbones deployed on SWaP-constrained edge platforms, making post-training quantization

ReFace: Reorganizing Facial Spatiotemporal Representations for Improved Pain Assessment

Model ReleasesDGX agent

arXiv:2607.19722v1 Announce Type: new Abstract: Automatic pain assessment from facial video remains challenging due to the spatial heterogeneity of pain-related facial cues. This study proposes ReFace

Reliability-Aware 3D Geometric Injection for Universal Person Re-identification

ApplicationsDGX agent

arXiv:2607.18863v1 Announce Type: new Abstract: Universal person re-identification (ReID) aims to retrieve pedestrian identities across diverse real-world scenarios, including severe occlusions, cloth

RIM: A Retrieval-In-Matching Framework for Cross-Domain Global Visual Localization of UAVs

Local AiDGX agent

arXiv:2607.20116v1 Announce Type: new Abstract: Global visual localization of unmanned aerial vehicles (UAVs) using remote-sensing reference maps has attracted increasing attention. However, acquisiti

Robust Activation Map Rectification for Weakly Supervised Volumetric Segmentation: Temporal Coherence as a Free Lunch

Local AiDGX agent

arXiv:2607.19877v1 Announce Type: new Abstract: Weakly supervised segmentation relies heavily on class activation maps (CAMs) to initially localize target regions. However, CAMs are often noisy and pr

Robust Multi-View Classification under Noisy Supervision via Global Anchor Consensus

ResearchDGX agent

arXiv:2607.18561v1 Announce Type: cross Abstract: In recent years, multi-view learning has attracted increasing attention, as it integrates the complementary information of heterogeneous views. Most e

ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling

Model ReleasesDGX agent

arXiv:2607.19332v1 Announce Type: cross Abstract: Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching. Along the way, the underlying techniques ha

RS-RIE-Bench: Benchmarking Reasoning-Guided Remote Sensing Image Editing

Model ReleasesDGX agent

arXiv:2607.20197v1 Announce Type: new Abstract: Remote sensing image editing aims to modify remote sensing images according to natural language instructions while preserving geographic rules and senso

SafeGen: Goal-Conditioned Video Diffusion of Safety-Critical Scenarios for VLM-Based Autonomous Driving

SafetyDGX agent

arXiv:2607.19701v1 Announce Type: new Abstract: VLMs are increasingly deployed in AD systems, creating an urgent need for rigorous safety evaluation under rare yet safety-critical scenarios. Among the

Sarus: Privacy-Preserving Multi-Vendor Perception Fusion via Homomorphic Encryption

AgentsDGX agent

arXiv:2607.19146v1 Announce Type: cross Abstract: Cooperative perception enables autonomous vehicles (AVs) to improve situational awareness by aggregating detection outputs from multiple agents and se

Seeing Before Generating: Object Perception Enhances Single-View 3D Reconstruction

Model ReleasesDGX agent

arXiv:2607.18630v1 Announce Type: new Abstract: The relationship between object perception and reconstruction is well established in human vision, yet remains underexplored in computer vision. In this

Self Gradient Forcing: Native Long Video Extrapolation

Model ReleasesDGX agent

arXiv:2607.20368v1 Announce Type: new Abstract: Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own ro

Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance

ResearchDGX agent

arXiv:2604.01848v4 Announce Type: replace Abstract: This work investigates the fundamental fragility of state-of-the-art Vision-Language Models (VLMs) under basic geometric transformations. While mode

SemICP: Semantic Non-Rigid Point Cloud Registration with Elastic Energy Regularization

SafetyDGX agent

arXiv:2503.00972v4 Announce Type: replace Abstract: Purpose: Accurate point cloud registration is essential in computer-aided interventions (CAI) to align multi-modal medical images for intraoperative

SHFormer: Dynamic Spectral Filtering Convolutional Neural Network and High-pass Kernel Generation Transformer for Adaptive MRI Reconstruction

Local AiDGX agent

arXiv:2607.20159v1 Announce Type: new Abstract: Attention Mechanism (AM) selectively focuses on essential information for imaging tasks and captures relationships between distant pixel neighborhoods t

Signed Rectified Flow: Negativity-Controlled Generation

ResearchDGX agent

arXiv:2607.18516v1 Announce Type: cross Abstract: We introduce Signed Rectified Flow (Signed RF), a generalization of Rectified Flow that targets the signed measure pi^{sign} = (1+alpha)pi^+ - alphapi

SIINR: Structurally Informed Implicit Neural Representations for super-resolution with uncertainty quantification of clinical quality diffusion MRI datasets

ResearchDGX agent

arXiv:2607.19943v1 Announce Type: new Abstract: Diffusion Magnetic Resonance Imaging (dMRI) is a powerful tool for probing brain microstructure, but clinical acquisitions are often limited by low out-

Single-Teacher View Augmentation: Enhancing Knowledge Distillation with Student-Guided Perturbations

Model ReleasesDGX agent

arXiv:2607.11557v2 Announce Type: replace Abstract: Knowledge distillation (KD) typically relies on the fixed perspective of a single teacher, limiting the diversity of supervisory signals. While mult

SkyEV: RGB-Event UAV detection and tracking dataset and baseline

ApplicationsDGX agent

arXiv:2607.18747v1 Announce Type: new Abstract: Detecting UAVs in air spaces has become increasingly important due to UAVs widespread availability and easy usage. However, due to their small size, the

STEREOFLOW: Progressive Stereo Matching with StereoDiT and Transition Flow Matching

SafetyDGX agent

arXiv:2607.19986v1 Announce Type: new Abstract: Stereo matching is a fundamental task in 3D reconstruction. Despite remarkable advances, the prevailing paradigms formulate stereo matching as a determi

Strength-Parity Ensembling with Parameter-Isolated Experts for Multi-Task Affect Recognition

Model ReleasesDGX agent

arXiv:2607.16290v2 Announce Type: replace Abstract: Leading entries on the multi-task track of the 11th ABAW challenge rely on heavy ensembling, yet which member is worth adding to an already strong e

StrokeSeg2: Stroke Lesion Segmentation in Clinical Research Workflows

Model ReleasesDGX agent

arXiv:2607.19901v1 Announce Type: new Abstract: Deep learning frameworks like nnU-Net achieve state-of-theart brain lesion segmentation performance but remain difficult to deploy in clinical research

Strong Gravitational Lensing Posterior Sampling in Pixel-Space Using Diffusion Models and Recurrent Inference Machines

ResearchDGX agent

arXiv:2607.19459v1 Announce Type: cross Abstract: Modeling galaxy-galaxy strong gravitational lenses to infer the brightness of the source galaxy and the mass distribution of the foreground galaxy is

STS-NET: Spatio-Temporal Stress Network for Self-Supervised Crop Stress Detection using Satellite Image Time Series

ApplicationsDGX agent

arXiv:2607.18791v1 Announce Type: new Abstract: Early and accurate detection of crop stress is essential to improve agricultural productivity and ensure global food security. However, collecting a lar

Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation

SafetyDGX agent

arXiv:2607.18508v1 Announce Type: new Abstract: Preference over model-generated emotion descriptions is emerging as a standard evaluation metric for multimodal emotion understanding, exemplified by th

← Previous
1…3738394041…209
Next →