AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
9 Jun 2026

See More, Match Better: Multi-Source Feature Fusion for Two-View Correspondence Learning

SafetyDGX agent

arXiv:2606.09262v1 Announce Type: new Abstract: Two-view correspondence learning aims to distinguish true correspondences (inliers) from false ones (outliers) in image pairs by leveraging their underl

SegmentAnyTreeV2: Scaling Transformer-Based Tree Instance Segmentation Across Sensors, Platforms, and Forests

Model ReleasesDGX agent

arXiv:2606.08206v1 Announce Type: new Abstract: We present SegmentAnyTreeV2, a sensor- and platform-agnostic framework for semantic and instance segmentation of forest point clouds. The model combines

Segmentation-Assisted Brain MRI Synthesis with Cross-Image Multi-Contrast Feature Memory Bank Retrieval Augmentation

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.08421v1 Announce Type: new Abstract: Multi-contrast brain MRI provide complementary soft-tissue characteristics that aid in the screening and diagnosis of diseases. However, limited scannin

Self-supervised Learning Matters: A Simple Ensemble Solution for Micro-Gesture Recognition

ResearchDGX agent

arXiv:2606.09261v1 Announce Type: new Abstract: In this paper, we present XInsight Lab's solution to the micro-gesture classification track of the 4th MiGA Challenge at IJCAI 2026, in which our soluti

Self-Supervised Learning with a Multi-Task Latent Space Objective

SafetyDGX agent

arXiv:2602.05845v2 Announce Type: replace Abstract: We propose a multi-task formulation of self-predictive Siamese SSL in which each spatial transformation defines a distinct latent-space alignment ta

SemDINO: A DINOv3-Driven Network for Cross-Temporal Semantic Alignment in Change Detection

SafetyDGX agent

arXiv:2606.09772v1 Announce Type: new Abstract: Semantic change detection (SCD) aims to simultaneously locate land-cover changes and identify semantic categories before and after transition. However,

Semi-supervised Source Detection in Astronomical Images: New Benchmark and Strong Baseline

Model ReleasesDGX agent

arXiv:2606.09219v1 Announce Type: new Abstract: Source detection in modern observational astronomy is a cornerstone for localizing and identifying stellar sources accurately. It is crucial for studies

Shift-Dependent Asymmetry: Orthogonal Inverse Low-Rank Adaptation for Federated Medical Segmentation

Model ReleasesDGX agent

arXiv:2606.08687v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) enables efficient federated fine-tuning of segmentation foundation models for medical imaging. However, most federated LoRA m

Simultaneous hyperkinetic movement disorders phenotyping: a cross-cohort pediatric transfer study using routine videos, markerless pose estimation and a tabular foundation model

ApplicationsDGX agent

arXiv:2606.07674v1 Announce Type: new Abstract: Objective: To develop and externally test a video-based framework for simultaneous detection of hyperkinetic MDs phenomenologies: dystonia, tremor, myoc

SMI: Efficient Self-Supervised Learning via Mutual-Information-Inspired Dependency Optimization

SafetyDGX agent

arXiv:2606.08332v1 Announce Type: new Abstract: Self-supervised learning (SSL) has achieved remarkable representation learning performance, but many existing methods rely on large batch sizes, memory

SoccerNet 2026 Player-Centric Ball-Action Spotting:Retraining and Post-Processing Extensions to the FOOTPASS Baselines

HardwareDGX agent

arXiv:2606.09679v1 Announce Type: new Abstract: We describe our system for the SoccerNet 2026 Player-Centric Ball-Action Spotting Challenge, which requires predicting who performs which action and whe

SOMA: From Surface Observations to Muscle Anatomy

ResearchDGX agent

arXiv:2606.09246v1 Announce Type: new Abstract: With the growing demand for realistic virtual humans, parametric body models have become a cornerstone of modern medicine, sports, and entertainment app

SPIRONet: Spatial-Frequency Learning and Graph-based Channel Interaction Network for Vessel Segmentation

Local AiDGX agent

arXiv:2406.19749v2 Announce Type: replace-cross Abstract: Automatic vessel segmentation plays a pivotal role in the development of next-generation interventional navigation systems for surgical roboti

SSAFE: Simple and Strong AI-Generated Image Detection via Frozen Vision Encoders

Model ReleasesDGX agent

arXiv:2606.08634v1 Announce Type: new Abstract: The rapid advancement of generative models has blurred the boundary between synthetic and real imagery, creating an urgent need for reliable deepfake de

Stabilizing On-Policy Distillation for MLLM Reasoning with Global Normalization

Model ReleasesDGX agent

arXiv:2606.09091v1 Announce Type: cross Abstract: On-policy distillation (OPD) has recently emerged as an important post-training paradigm. By using a stronger teacher model to provide dense, fine-gra

Stain-Aware Wavelet Regularization for Instant Adversarial Purification in Histopathology

SafetyDGX agent

arXiv:2606.08745v1 Announce Type: new Abstract: Deep learning has become prevalent in computational pathology pipelines that support tasks such as cancer screening and digital pathology analysis. Howe

Steer Where It Matters: Token-Level Visual-Sensitivity Steering for LVLMs Hallucination Mitigation

ResearchDGX agent

arXiv:2606.07647v1 Announce Type: new Abstract: Large vision language models (LVLMs) have made rapid advancements and are deployed across various applications, yet hallucinations remain a major challe

STGBD-Net: Spatio-temporal Gradient Basis Decomposition Network for Infrared Small Target Detection

ResearchDGX agent

arXiv:2512.03470v5 Announce Type: replace Abstract: A key challenge in infrared small target detection (IRSTD) is that weak target signal responses are easily obscured by strong background clutter, fr

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur?

Model ReleasesDGX agent

arXiv:2606.09547v1 Announce Type: new Abstract: Learning everyday skills, like cooking a dish, relies increasingly on instructional media such as online videos. This opens the door to the use of video

Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking

Model ReleasesDGX agent

arXiv:2606.07689v1 Announce Type: new Abstract: Deep research agents have attracted increasing attention for their ability to collect large-scale online information to acquire target knowledge, with r

SwiftVR: Real-Time One-Step Generative Video Restoration

HardwareDGX agent

arXiv:2606.09516v1 Announce Type: new Abstract: Real-time video restoration (VR) for live streams requires high-resolution outputs under strict per-frame latency constraints. Existing one-step diffusi

Taming Perception Jitter: Uncertainty-Aware LiDAR Object Detection for Reliable Motion Classification

AgentsDGX agent

arXiv:2606.09350v1 Announce Type: cross Abstract: Reliable motion classification is critical for autonomous driving, as false dynamic predictions of static objects can cascade into unnecessary planner

TBD-VLA: Temporal Block Diffusion Vision Language Action Model

ApplicationsDGX agent

arXiv:2606.07895v1 Announce Type: new Abstract: Discrete Vision-Language-Action (VLA) models typically formulate action generation as next-token prediction over discretized action spaces, conditioning

Temporal-Aware Reasoning Optimization for Video Temporal Grounding

Local AiDGX agent

arXiv:2606.09248v1 Announce Type: new Abstract: Multi-modal Large Language Models (MLLMs) have achieved remarkable progress in video temporal grounding with reinforcement learning for generating reaso

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning

ResearchDGX agent

arXiv:2606.08231v1 Announce Type: new Abstract: Test-time Scaling (TTS) has emerged as a pivotal research direction for enhancing model performance by dynamically allocating computational resources du

The Need for Neural ISP in the Small-Pixel Era: How Shrinking Pixels Push Optics to the Limit and Neural Restoration Pushes Back

Local AiDGX agent

arXiv:2606.07675v1 Announce Type: cross Abstract: Smartphone telephoto cameras are approaching a 'telephoto physics wall': as pixel pitches shrink toward sub-0.5 micron, the optics remain limited by g

Thinking Without Images: Internalizing Visual Manipulation with On-Policy Self-Distillation

Local AiDGX agent

arXiv:2606.08719v1 Announce Type: new Abstract: ''Thinking with Images'' has emerged as an effective paradigm for fine-grained visual reasoning: by explicitly zooming into relevant regions and reasoni

TIDE: Task-Isolated Diffusion for Unified Video Editing and Generation

ResearchDGX agent

arXiv:2606.08260v1 Announce Type: new Abstract: Recent advances in Diffusion Transformers have driven rapid progress in video generation and editing, yet these capabilities are still handled by separa

Toward Scalable Co-located Practical Learning: Assisting with Computer Vision and Multimodal Analytics

ResearchDGX agent

arXiv:2603.13679v2 Announce Type: replace-cross Abstract: Co-located practical learning leaves evidence in visible actions around patients, task resources and room zones, but these traces are often re

Towards Accurate Emotion-Attributed Video Captioning via Fine-grained Emotion-Cause Pair Extraction

SafetyDGX agent

arXiv:2606.08566v1 Announce Type: new Abstract: Emotional Video Captioning (EVC) is a challenging task that aims to generate factually accurate and emotionally rich descriptions for videos. Existing E

Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings

SafetyDGX agent

arXiv:2511.05017v2 Announce Type: replace Abstract: Hallucinations in Large Vision-Language Models (LVLMs) remain a persistent challenge, often stemming from inadequate integration of visual informati

Training-Free Generalized Few-Shot Segmentation through Open-Vocabulary Semantic Arbitration

Model ReleasesDGX agent

arXiv:2606.09474v1 Announce Type: new Abstract: Generalized Few-Shot Semantic Segmentation (GFSS) has traditionally been approached as a representation-learning problem, requiring task-specific adapta

Trajectory Optimization in Single and Dual-UAV Bearing-Only Target Localization

Local AiDGX agent

arXiv:2606.09188v1 Announce Type: cross Abstract: Bearing-only target localization is a fundamental problem in optical measurement and finds extensive applications in unmanned aerial vehicle (UAV) tec

Trustworthy Visual Predicates for Robust Manipulation Understanding under Degradation

ResearchDGX agent

arXiv:2606.08121v1 Announce Type: new Abstract: Manipulation understanding requires reliable relational evidence, such as contact, support, containment, motion coupling, grasp, release, and active-han

TUDSR: Twice Upsampling-Diffusion for Higher Super-Resolution

HardwareDGX agent

arXiv:2606.09608v1 Announce Type: new Abstract: Diffusion-based generative models have achieved remarkable success in real-world image super-resolution (SR). With tiled diffusion techniques, these mod

TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding

ResearchDGX agent

arXiv:2606.08464v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning has proven effective for enhancing problem-solving in large language models. However, when applied to multimodal LLMs (

UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models

ApplicationsDGX agent

arXiv:2602.18020v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models leverage pretrained Vision-Language Models (VLMs) as backbones to map images and instructions to actions, demons

Ultra Flash: Scaling Real-Time Streaming Video Generation to High Resolutions

HardwareDGX agent

arXiv:2606.09150v1 Announce Type: new Abstract: While recent autoregressive video diffusion models achieve remarkable streaming quality, they remain confined to low resolutions (e.g., 480P), leaving e

Uncertainty-Aware Hierarchical Re-Localization in OpenStreetMap via Semantic Alignment

SafetyDGX agent

arXiv:2603.01613v2 Announce Type: replace Abstract: Monocular re-localization enables robots to estimate camera poses from visual observations. However, many existing methods rely on dense maps or lar

UniADC: A Unified Framework for Anomaly Detection and Classification

ResearchDGX agent

arXiv:2511.06644v3 Announce Type: replace Abstract: In this paper, we introduce a novel task termed unified anomaly detection and classification, which aims to simultaneously detect anomalous regions

vesselFM-CT: Segmenting All Blood Vessels in CT Images for System-Level Cardiovascular Analysis

ResearchDGX agent

arXiv:2606.09400v1 Announce Type: new Abstract: The vascular network in the human body is characterized by blood vessels exhibiting drastic structural variations in radius, length, topological propert

VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation

Model ReleasesDGX agent

arXiv:2606.08091v1 Announce Type: new Abstract: Recent agent frameworks such as Claude Code, Codex, and OpenClaw are strong at tool use and orchestration, but whether they can handle long video genera

Virtual-point-based Solutions to Handle Generalized Absolute Pose Problem

AgentsDGX agent

arXiv:2606.09294v1 Announce Type: new Abstract: Multi-camera systems are increasingly adopted in robotics and autonomous navigation for their wide field of view, flexibility, and fault tolerance. Neve

Vision-Language Asymmetry in Bistable Image Captioning

SafetyDGX agent

arXiv:2606.08031v1 Announce Type: new Abstract: Wittgenstein's duck-rabbit poses a question for vision-language models: when a model captions an ambiguous image, where in the model is the commitment t

Vision-Language Guided Hyperspectral Object Tracking via Semantics Fusion and Contextual Template Updating

TutorialsDGX agent

arXiv:2606.09167v1 Announce Type: new Abstract: Hyperspectral object tracking (HOT) leverages the rich spectral information provided by hyperspectral videos (HSVs), offering substantial potential for

Vision-Language Work Zone Intelligence for Safety-Critical Speed Regulation of Mixed-Autonomy Vehicles in Dynamic Environments

SafetyDGX agent

arXiv:2606.08860v1 Announce Type: new Abstract: Temporary work-zone speed limits are communicated through visually inconsistent signage and are often missing from digital maps, creating safety risks f

Visual Para-Thinker++: A Single-Policy Multi-Agent Framework for Visual Reasoning

SafetyDGX agent

arXiv:2606.09290v1 Announce Type: new Abstract: Visual reasoning requires integrating evidence distributed across regions, attributes, and relations, making single-chain reasoning prone to early perce

Visual Template Inference for Data Extraction from Documents

Model ReleasesDGX agent

arXiv:2501.06659v2 Announce Type: replace-cross Abstract: Many templatized documents are programmatically generated from structured data following a visual template. Such documents include invoices, t

VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning?

Model ReleasesDGX agent

arXiv:2606.07872v1 Announce Type: new Abstract: When a multimodal large language model answers a visual reasoning question correctly, is the prediction actually supported by the task-critical visual e

WaveDiT: Distribution-Aware Wavelet Flow Matching for Efficient 3D Brain MRI Synthesis

SafetyDGX agent

arXiv:2606.08670v1 Announce Type: new Abstract: Large and demographically balanced datasets are essential for reliable neuroimaging biomarkers. Full-resolution 3D brain MRI synthesis can support data

What neurosurgeons need to see: synthetic intra-operative MRI from ultrasound for brain-shift compensation in brain tumour surgery

ResearchDGX agent

arXiv:2606.07658v1 Announce Type: new Abstract: Maximal safe resection is the primary objective in glioma surgery. Neuronavigation guidance is progressively degraded by brain shift after dural opening

When Vision Misleads, Let Location Speak: A Worldwide Image Geo-Localization Method via Location Attention Mechanism and Large Multimodal Models

Local AiDGX agent

arXiv:2606.08918v1 Announce Type: new Abstract: Worldwide image geo-localization aims to determine the capture location of an image on a global scale. Existing methods often mislocalize images by matc

Where Does the Answer Come From? Benchmarking View-Level Visual Evidence Identification in Multi-View MLLMs for Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.09644v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) achieve strong results on visual reasoning benchmarks, but answer accuracy alone does not indicate whether a

Where the Score Lives: A Wavelet View of Diffusion

ResearchDGX agent

arXiv:2606.08309v1 Announce Type: cross Abstract: Score-based generative models have had remarkable success over the last decade in generating a diverse set of visually plausible images. A variety of

Wispy to Voluminous: Prior-free Multi-view Capture of Strand-level Facial Hair

ApplicationsDGX agent

arXiv:2606.08041v1 Announce Type: cross Abstract: Facial hair is a defining trait of personal identity, yet remains a critical bottleneck for digital avatars. Recent volumetric methods achieve photore

X-Palm: Paired Multispectral-to-Smartphone Dataset for Cross-Domain Palmprint Authentication

ApplicationsDGX agent

arXiv:2606.08437v1 Announce Type: cross Abstract: Palmprint modality offers a privacy-preserving biometric solution, yet its deployment is hindered by the domain gap between controlled enrollment and

Zero-Parameter Geometric Gating for Temporally Stable Low-Altitude UAV Video Semantic Segmentation

Model ReleasesDGX agent

arXiv:2606.09162v1 Announce Type: new Abstract: Video semantic segmentation for low-altitude UAVs requires temporal consistency, yet dense optical flow introduces spatially structured noise in the pla

Zero-Shot Semantic Re-Identification for Autonomous Driving: A VLM Baseline Study

Model ReleasesDGX agent

arXiv:2606.09362v1 Announce Type: new Abstract: Re-Identification (ReID) in autonomous driving is typically formulated as a visual matching problem, where observations of vehicles, pedestrians, and cy

8 Jun 2026

3DMorph: Single-Image-Guided Local 3D Shape Editing and Morphing

Model ReleasesDGX agent

arXiv:2606.07115v1 Announce Type: new Abstract: Despite recent progress in 3D generation, intuitive editing of existing shapes remains limited. Unlike images, which benefit from well-established inpai

A Cross-view Fusion Framework for Robust 6-DoF Grasp Pose Estimation

Model ReleasesDGX agent

arXiv:2606.06878v1 Announce Type: cross Abstract: In this paper, we propose a cross-view fusion framework that enhances the robustness of 6-DoF grasp pose estimation in corner views. Our framework all

← Previous
1…9091929394…211
Next →