AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
8 Jun 2026

OPTIMUS-Prime: Minimal and Sufficient Concept Explanations for Deep Vision Models

Model ReleasesDGX agent

arXiv:2606.07180v1 Announce Type: new Abstract: The growing demand for transparency in automated decision-making has propelled eXplainable Artificial Intelligence (XAI) to the forefront of machine lea

PARSE: Part-Aware Relational Spatial Modeling

ResearchDGX agent

arXiv:2603.07704v2 Announce Type: replace Abstract: Inter-object relations underpin spatial intelligence, yet existing representations -- linguistic prepositions or object-level scene graphs -- are to

Physics-Driven Semantic Scattering Structure Understanding of Aircraft Target in SAR Images

Model ReleasesDGX agent

arXiv:2606.06847v1 Announce Type: cross Abstract: Synthetic aperture radar (SAR) has become indispensable for target interpretation owing to its all-day and all-weather observation capability. In SAR


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

polyDAG: Polynomial Acyclicity Constraints for Efficient Continuous Causal Discovery in Visual Semantic Graphs

ResearchDGX agent

arXiv:2606.06908v1 Announce Type: new Abstract: Modern image-analysis pipelines often convert images into structured semantic variables, such as facial attributes, object concepts, and scene descripto

Real-Time AttentionBender: Granular Interactive Network Bending of Video Diffusion Transformers

ResearchDGX agent

arXiv:2606.06497v1 Announce Type: cross Abstract: Generative video models have achieved remarkable visual fidelity, yet their prompt-only interface offers thin creative agency and obscures the model's

RealDocBench: A Benchmark for Field-Level QA and Layout Understanding on Real-World Regulated Documents

Model ReleasesDGX agent

arXiv:2606.07401v1 Announce Type: new Abstract: Document parsing systems are increasingly deployed in high-stakes, regulated workflows such as mortgage underwriting, financial reporting, supply-chain

Reconstructing Multi-Decadal Forest Disturbances: A Spatio-Temporal Transformer Approach

ResearchDGX agent

arXiv:2606.07249v1 Announce Type: new Abstract: Accurate monitoring of forest disturbances is essential for understanding carbon dynamics and land management, yet traditional approaches typically rely

RigPAPR: Rig-Based Animation of Static Neural Point Clouds from a Fixed-Viewpoint Video

ResearchDGX agent

arXiv:2606.06685v1 Announce Type: new Abstract: Static neural point reconstructions capture a subject at high fidelity from posed images. Given such a reconstruction, we aim to animate it to follow a

RISE: Single Static Radar-based Indoor Scene Understanding

Model ReleasesDGX agent

arXiv:2511.14019v3 Announce Type: replace Abstract: Robust and privacy-preserving indoor scene understanding remains a fundamental open problem. While optical sensors such as RGB and LiDAR offer high

Robotic Policy Adaptation via Weight-Space Meta-Learning

SafetyDGX agent

arXiv:2606.07217v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are emerging as a promising paradigm for robotic manipulation, enabling general-purpose policies trained from larg

RPC-GS: Gaussian Splatting with native RPC Rendering for Satellite Imagery

Model ReleasesDGX agent

arXiv:2606.06690v1 Announce Type: new Abstract: We present RPC-GS, the first Gaussian Splatting framework for satellite imagery that operates natively with Rational Polynomial Camera (RPC) models. The

S23DR 2026 Winning Solution

ResearchDGX agent

arXiv:2606.06695v1 Announce Type: new Abstract: This text presents the winning solution to the S23DR 2026 challenge for structured 3D wireframe reconstruction from sparse SfM, fitted depth, and semant

Seeing Without Exposing: Adaptive Privacy Control for Open-World, Context-Hungry MLLMs

Model ReleasesDGX agent

arXiv:2606.07175v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have raised new privacy challenges. On the data side, user-provided inputs often include unpredictable sensitiv

SegMoTE: Token-Level Mixture of Experts for Medical Image Segmentation

ResearchDGX agent

arXiv:2602.19213v2 Announce Type: replace Abstract: Medical image segmentation is vital for clinical diagnosis and quantitative analysis, yet remains challenging due to the heterogeneity of imaging mo

Semantic-Structural Alignment for Generative Pictorial Charts

SafetyDGX agent

arXiv:2606.06498v1 Announce Type: cross Abstract: Traditional statistical graphics are precise but often lack the visual appeal, memorability, and engagement of pictorial charts. We present a generati

Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning

Model ReleasesDGX agent

arXiv:2606.07436v1 Announce Type: new Abstract: This paper explores agentic 3D spatial understanding, i.e., MLLM agents performing 3D reasoning through tool use. Existing methods often misuse tools an

Spatial-Temporal Decoupled Adapter for Micro-gesture Online Recognition

Model ReleasesDGX agent

arXiv:2606.07355v1 Announce Type: new Abstract: Micro-gesture online recognition aims to temporally localize and classify subtle gestures in untrimmed videos. Owing to their extremely short duration,

Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors

Model ReleasesDGX agent

arXiv:2606.06891v1 Announce Type: new Abstract: Despite advances in 3D scene understanding, existing 3D Large Multimodal Models operate in offline settings, requiring complete scene observations or pr

Streaming Video Generation with Streaming Force Control

Local AiDGX agent

arXiv:2606.07508v1 Announce Type: new Abstract: We introduce StreamForce, a streaming video generation framework that enables physically grounded control through continuous force inputs. Unlike prior

SVHighlights: Towards Extremely Long Sport Video Highlight Detection

Model ReleasesDGX agent

arXiv:2606.06926v1 Announce Type: new Abstract: While highlight detection for long-form videos is of great practical importance, most existing methods remain limited to short-form content, largely due

T2LM: Long-Term 3D Human Motion Generation from Multiple Sentences

ResearchDGX agent

arXiv:2406.00636v2 Announce Type: replace Abstract: In this paper, we address the challenging problem of long-term 3D human motion generation. Specifically, we aim to generate a long sequence of smoot

The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models

ResearchDGX agent

arXiv:2603.22278v2 Announce Type: replace Abstract: Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their prop

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance

ResearchDGX agent

arXiv:2606.07161v1 Announce Type: new Abstract: Video Text Spotting (VTS) is essential for urban surveillance and intelligent transportation systems, enabling automated reading of street signs, vehicl

TrioPose: Native Triple-Stream Diffusion Transformers for Pose-Guided Text-to-Image Generation

SafetyDGX agent

arXiv:2606.07053v1 Announce Type: new Abstract: Pose-guided text-to-image generation often suffers from limb distortions and feature crosstalk in complex multi-person scenarios. While existing UNet-ba

Twin: Tuning Learning Rate and Weight Decay of Deep Homogeneous Classifiers without Validation

Model ReleasesDGX agent

arXiv:2403.05532v2 Announce Type: replace-cross Abstract: We introduce Tune without Validation (Twin), a simple and effective pipeline for tuning learning rate and weight decay of homogeneous classifi

Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information Flows

SafetyDGX agent

arXiv:2606.06875v1 Announce Type: new Abstract: Diffusion transformers (DiTs) equipped with multimodal attention (MM-Attn) have become a dominant paradigm for image generation. However, preventing the

UniSHARP: Universal Sharp Monocular View Synthesis

Model ReleasesDGX agent

arXiv:2606.07514v1 Announce Type: new Abstract: In this work, we focus on extending SHARP, the popular photorealistic view synthesis method, for universal monocular rendering across a continuum of cam

Unmixing ATR-{mu}FTIR spectroscopic images of cross-sections of historical oil paintings

ResearchDGX agent

arXiv:2603.06673v2 Announce Type: replace Abstract: Spectroscopic imaging (SI) has become central to heritage science because it enables non-invasive, spatially resolved characterisation of materials

Unregistered Spectral Image Fusion: Unmixing, Adversarial Learning, and Recoverability

ResearchDGX agent

arXiv:2603.21510v3 Announce Type: replace-cross Abstract: This paper addresses the fusion of a pair of spatially unregistered hyperspectral image (HSI) and multispectral image (MSI) covering roughly o

Unsupervised Learning Based Focal Stack Camera Depth Estimation

ResearchDGX agent

arXiv:2203.07904v3 Announce Type: replace-cross Abstract: We propose an unsupervised deep learning based method to estimate depth from focal stack camera images. On the NYU-v2 dataset, our method achi

USU-Corn-WeedDB: A UAV RGB Image Dataset for Multi-Species Weed Detection in Forage Corn

ApplicationsDGX agent

arXiv:2606.06709v1 Announce Type: new Abstract: Weed pressure in forage corn production causes yield losses of up to 31.5%, yet site-specific weed management (SSWM) systems built on UAV imagery and de

Varifold Moment Invariants for Sustainable and Explainable Contour Feature Extraction

ResearchDGX agent

arXiv:2606.07333v1 Announce Type: new Abstract: We introduce Varifold Moments Invariants (VMI) as a unifying framework for many previously introduced Moment Invariants. These invariants are deeply rel

VeriDrive: Verifiable Counterfactual Supervision for Cost-Efficient Vision-Language Planning

Model ReleasesDGX agent

arXiv:2606.07338v1 Announce Type: new Abstract: Vision-language driving models increasingly use reasoning supervision to bridge perception, prediction, and planning, but existing driving rationales ar

VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object Segmentation

Model ReleasesDGX agent

arXiv:2606.06819v1 Announce Type: new Abstract: Reasoning Video Object Segmentation (RVOS) demands a sophisticated integration of temporal dynamics, spatial details, and linguistic reasoning to achiev

VIRTUS-FPP: Virtual Sensor Modeling for Fringe Projection Profilometry in NVIDIA Isaac Sim

HardwareDGX agent

arXiv:2509.22685v2 Announce Type: replace-cross Abstract: Fringe projection profilometry (FPP) is a high-precision structured-light sensing technique for 3D surface reconstruction, yet its practical d

When CLIP Sees More, It Fights Back Harder: Multi-View Guided Adaptive Counterattacks for Test-Time Adversarial Robustness

ResearchDGX agent

arXiv:2606.06938v1 Announce Type: new Abstract: Vision-language models such as CLIP have achieved remarkable zero-shot recognition capabilities, yet their robustness against adversarial perturbations

When Recovery Matters: The Blind Spot of Surrogate Privacy in MLLM Editing

Model ReleasesDGX agent

arXiv:2606.07171v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) enable flexible instruction-driven image editing, but privacy risks arise when user images expose diverse and u

WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

Model ReleasesDGX agent

arXiv:2606.06538v1 Announce Type: new Abstract: In real-world applications, models are expected to perform reliably across diverse settings. Yet, many existing multimodal benchmarks expand task types

Zero-Shot Polygon Matching with Pre-trained Models for Pose Estimation and Polygon Cloud from Challenging Stereo

ResearchDGX agent

arXiv:2511.05949v2 Announce Type: replace Abstract: While stereo matching has achieved maturity for 0D point and 1D line primitives, establishing correspondences for 2D polygons remains largely unexpl

5 Jun 2026

A Vision-language Framework for Comparative Reasoning in Radiology

ResearchDGX agent

arXiv:2606.06407v1 Announce Type: new Abstract: Medical imaging artificial intelligence has achieved strong performance in isolated image interpretation, but remains poorly aligned with radiological p

ActiveMimic: Egocentric Video Pretraining with Active Perception

ApplicationsDGX agent

arXiv:2606.06194v1 Announce Type: cross Abstract: Egocentric human video offers a scalable alternative to robot data for pretraining, yet models pretrained on such video consistently underperform thos

Adaptive Tokenisation Via Temporal Redundancy Masking And Latent Inpainting

Model ReleasesDGX agent

arXiv:2606.06158v1 Announce Type: new Abstract: Adaptive video tokenisation seeks to dynamically allocate token budgets based on the underlying visual complexity of a sequence. Current continuous-regi

Adversarial Attacks Already Tell the Answer: Directional Bias-Guided Test-time Defense for Vision-Language Models

SafetyDGX agent

arXiv:2606.06186v1 Announce Type: new Abstract: Vision-Language Models (VLMs), such as CLIP, have shown strong zero-shot generalization but remain highly vulnerable to adversarial perturbations, posin

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding

Local AiDGX agent

arXiv:2606.06155v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models leverage the rich world knowledge of pretrained vision-language models (VLMs) to enable instruction-following robo

Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue

ApplicationsDGX agent

arXiv:2509.15061v2 Announce Type: cross Abstract: The ultimate goal of embodied agents is to create collaborators that can interact with humans, not mere executors that passively follow instructions.

ATT-CR: Adaptive Triangular Transformer for Cloud Removal

ResearchDGX agent

arXiv:2606.05999v1 Announce Type: new Abstract: Cloud removal aims to accurately reconstruct the ground objects obscured by clouds in remote sensing images. Existing Transformer-based methods utilizin

Beyond Absolute Scores: Relative Edit-induced Difference for Generalizable Image Aesthetic Assessment

ResearchDGX agent

arXiv:2606.05778v1 Announce Type: new Abstract: Traditional Image Aesthetic Assessment (IAA) methods mainly rely on regressing absolute Mean Opinion Scores (MOS). However, such a paradigm overlooks th

Biomazon: A Multimodal Dataset for 3D Forest Structure and Biomass Modeling in the Amazon Basin

Model ReleasesDGX agent

arXiv:2606.05368v1 Announce Type: new Abstract: Accurate, spatially explicit characterization of tropical forest structure is essential for carbon accounting and ecosystem monitoring, yet most ML pipe

BMCR: Adaptive Backbone Module Composition via Reinforcement Learning for Remote Sensing Object Detection

Local AiDGX agent

arXiv:2606.05586v1 Announce Type: new Abstract: In remote sensing object detection, Convolutional Neural Networks (CNNs) excel at capturing local details while Vision Transformers (ViTs) are better at

BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain

ResearchDGX agent

arXiv:2512.08560v3 Announce Type: replace Abstract: Understanding how the human brain represents visual concepts, and in which brain regions these representations are encoded, remains a long-standing

BRepCLIP: Contrastive Multimodal Pretraining on BRep Primitives for CAD Understanding

ResearchDGX agent

arXiv:2606.05515v1 Announce Type: new Abstract: Learning representations of CAD models is a largely open problem. While 3D representation learning has flourished around point clouds and meshes, the na

CamFlow+: Hybrid Motion Bases for 2D Camera Motion Estimation with Stabilization Applications

Model ReleasesDGX agent

arXiv:2606.05915v1 Announce Type: new Abstract: Estimating 2D camera motion is fundamental to computer vision and computational photography. Existing homography-based methods work well for planar scen

Can Language Models Learn to Listen?

TutorialsDGX agent

arXiv:2308.10897v2 Announce Type: replace Abstract: We present a framework for generating appropriate facial responses from a listener in dyadic social interactions based on the speaker's words. Given

Can We Predict The Human Preference For Text-to-Image Content Prior To Generation And Is It Even Useful To Do So?

ResearchDGX agent

arXiv:2606.05478v1 Announce Type: new Abstract: Diffusion Models (DM) have revolutionized text-driven generation by enabling the synthesis of high-quality, photorealistic visual content from user prom

CoFi-UCGen: Coarse-to-Fine Unsupervised Conditional Generation without Label Priors

TutorialsDGX agent

arXiv:2606.05652v1 Announce Type: new Abstract: Unsupervised conditional image generation (UCGen) aims to control generation without relying on manually annotated labels, yet remains challenging due t

Comparison of Deep Learning Frameworks For Rice Disease Mapping From UAV Multispectral Imaging

ResearchDGX agent

arXiv:2606.06359v1 Announce Type: new Abstract: In this study, UAV multispectral imagery is used to segment the severity of bacterial leaf blight (BLB) in rice using convolutional neural networks (CNN

Complexity-Balanced Diffusion Splitting

Local AiDGX agent

arXiv:2606.06477v1 Announce Type: new Abstract: Standard continuous-time generative models rely on monolithic architectures that must navigate vastly different signal regimes, from isotropic noise to

Computation-Aware Event-to-Frame Reconstruction via Selective Attention

ResearchDGX agent

arXiv:2606.06142v1 Announce Type: new Abstract: Event-to-frame (E2F) reconstruction bridges asynchronous event streams with frame-based vision pipelines, but existing methods often face a trade-off be

Cosine Misleads: Auxiliary Losses Reshape Vision Language Models, Not Their Latents

SafetyDGX agent

arXiv:2606.05753v1 Announce Type: new Abstract: Latent visual reasoning (LVR) inserts supervised latent tokens between perception and answer generation in vision-language models (VLMs). The field uses

Deep Learning-assisted AMD Staging based on OCT and OCT Angiography

ResearchDGX agent

arXiv:2606.05379v1 Announce Type: new Abstract: To develop and evaluate deep learning models for automated grading of age-related macular degeneration (AMD) severity using optical coherence tomography

← Previous
1…9293949596…211
Next →