AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
8 Jul 2026

Straight-Path Flow Matching for Incomplete Multi-View Clustering

SafetyDGX agent

arXiv:2607.06281v1 Announce Type: new Abstract: Incomplete Multi-View Clustering addresses the problem of clustering multi-modal data when certain views are missing. Recent end-to-end generative appro

Structured-Condensed Prompt Tuning in Vision-Language Models for Fine-grained Image Recognition

SafetyDGX agent

arXiv:2607.06185v1 Announce Type: new Abstract: Fine-grained image recognition poses a significant challenge due to the substantial expertise and effort required for manual annotation. Vision-language

Structured Data Extraction from Real Estate Documents using Clustering, Classification, and Large Language Models

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.06012v1 Announce Type: new Abstract: Real estate property listings expose structured metadata through the API. Still, the richest property-level information (i.e., legal status, structural

Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation

SafetyDGX agent

arXiv:2602.02401v2 Announce Type: replace Abstract: Human motion analysis tasks, such as temporal 3D pose estimation, motion prediction, and motion in-betweening, play an essential role in computer vi

Synthetic-to-Real Translation for Class-Agnostic Motion Prediction

SafetyDGX agent

arXiv:2607.06319v1 Announce Type: new Abstract: Motion understanding is critical for ensuring safety and robustness in autonomous driving systems, driving increasing interest in motion prediction. A k

Taxlifier: Leveraging Disease Taxonomy for Enhanced Multi-Label Classification in Chest Radiography

ResearchDGX agent

arXiv:2607.05628v1 Announce Type: new Abstract: Accurate and efficient classification of thoracic diseases in chest X-ray (CXR) images is crucial for timely diagnosis and treatment. However, the prese

Temporal Modeling of Optically Variable Devices in Identity Documents

ResearchDGX agent

arXiv:2607.06408v1 Announce Type: new Abstract: Robust remote verification of identity documents relies on analyzing faint, transparent security features like Optically Variable Devices (OVDs), or 'ho

TMF-RSE: Tri-Modal Fusion with Regional Semantics and Evidential Uncertainty for Lung Severity Scoring

ResearchDGX agent

arXiv:2607.06356v1 Announce Type: cross Abstract: Accurate quantification of lung disease severity from chest imaging is critical for clinical decision-making and resource allocation. We propose a tri

Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement

ApplicationsDGX agent

arXiv:2607.06370v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow matching-based V

TRIG: Trajectory-Rig Decoupled Metric Geometry Learning

AgentsDGX agent

arXiv:2607.05801v1 Announce Type: new Abstract: Vision-centric autonomous driving requires accurate metric geometry and ego-motion estimation from synchronized multi-camera observations. Recent visual

Tuned Reverse Distillation: Enhancing Multimodal Industrial Anomaly Detection with Crossmodal Tuners

TutorialsDGX agent

arXiv:2412.08949v4 Announce Type: replace Abstract: Knowledge distillation (KD) has been widely studied in unsupervised image Anomaly Detection (AD), but its application to unsupervised multimodal AD

Tuning-Free Latent Diffusion Models for Ultrahigh-Resolution Image Editing

HardwareDGX agent

arXiv:2607.06136v1 Announce Type: new Abstract: Recent diffusion-based generative models have shown impressive performance in image generation and editing. However, due to memory limitations and the h

UniField: A Unified Field-Aware MRI Enhancement Framework

SafetyDGX agent

arXiv:2603.09223v2 Announce Type: replace Abstract: Magnetic Resonance Imaging (MRI) field-strength enhancement holds immense value for both clinical diagnostics and advanced research. However, existi

Unlearnable Faces: Privacy Protection Surviving Extraction Pipeline

ResearchDGX agent

arXiv:2607.05996v1 Announce Type: new Abstract: Unlearnable examples keep publicly shared photos from being learned by unauthorized face-recognition models. An imperceptible perturbation, added before

Unsupervised Domain Adaptation for Calcification Classification in Mammography Across Multi-Site Datasets

ResearchDGX agent

arXiv:2607.06549v1 Announce Type: new Abstract: Deep learning-based computer-aided diagnosis (CAD) systems have shown strong performance in breast cancer diagnosis, particularly for classification tas

VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery

AgentsDGX agent

arXiv:2607.06374v1 Announce Type: new Abstract: Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact e

VEIL: How Visual Encoding Hijacking Induces Bias In Vision Models

SafetyDGX agent

arXiv:2607.05641v1 Announce Type: new Abstract: Rendering time series as chart images for CNN-based classification has become increasingly common in time-series classification (TSC). However, it remai

Verification of Dynamic Holographic Behavior in Identity Documents

Model ReleasesDGX agent

arXiv:2607.06466v1 Announce Type: new Abstract: This paper addresses the remote verification of the authenticity of Optically Variable Devices (commonly known as holograms) on identity documents. Typi

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models

Model ReleasesDGX agent

arXiv:2510.13808v2 Announce Type: replace Abstract: Large Vision Language Models (VLMs) excel at general visual reasoning but experience significant performance degradation when deployed in novel doma

Vision as Unified Multimodal Generation

ResearchDGX agent

arXiv:2607.06560v1 Announce Type: new Abstract: We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation sp

Visual graphs for image classification: does the structure affect performance?

ResearchDGX agent

arXiv:2607.06295v1 Announce Type: new Abstract: Deep learning models have emerged in machine learning and related fields, demonstrating astonishing performance in various visual tasks. Despite their g

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.06118v1 Announce Type: new Abstract: As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their nav

Why does Deep Learning Improve Visual SLAM?

Model ReleasesDGX agent

arXiv:2607.06023v1 Announce Type: new Abstract: Visual SLAM is a well-established technology utilized in a wide range of real-world applications. However, its performance still degrades under challeng

WING: A Window-Prior-Based Generative Network with Gated Inception for Cross-Modality CT Synthesis

TutorialsDGX agent

arXiv:2607.06234v1 Announce Type: new Abstract: Generating CT volumes from MRI and CBCT can improve treatment planning in adaptive radiotherapy while avoiding additional radiation exposure. However, d

WristMimic: Full-Body Humanoid Control with Wrist-Guided Manipulation

TutorialsDGX agent

arXiv:2607.06438v1 Announce Type: cross Abstract: Retargeting human object interaction demonstrations to physics based simulation requires reproducing not only body motion but also the object motion a

XRFormer: Multiscale Tokenization for XRF Representation Learning

Model ReleasesDGX agent

arXiv:2607.06424v1 Announce Type: new Abstract: X-ray fluorescence (XRF) spectroscopy is a key modality for material analysis in cultural heritage. However, automated learning from XRF spectra remains

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Model ReleasesDGX agent

arXiv:2502.09696v3 Announce Type: replace Abstract: Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children

7 Jul 2026

3DMPE: 3D Multi-Perspective Embedding

ResearchDGX agent

arXiv:2607.04898v1 Announce Type: new Abstract: We study 3D point cloud reconstruction from multiple partially observed 2D projections. Given two or more projections of an unknown 3D point cloud, toge

A Cost-Aware, Paired Protocol for Auditing Dynamic Tool Synthesis in Agentic Video Question Answering

AgentsDGX agent

arXiv:2607.01469v2 Announce Type: replace Abstract: Agentic Video Question Answering (VideoQA) systems invoke tools during inference, but their tool libraries are fixed, so recurring procedures are re

A Decomposable Probe for Few-Step Diffusion Models: Prompt, Latent, and Score Selectivity across Backbone Families and Distillation Paradigms

ResearchDGX agent

arXiv:2607.03256v1 Announce Type: new Abstract: Few-step distilled diffusion students cut text-to-image inference from ~50 to 1-8 network evaluations, but the quality gap is usually summarised by a si

A Large-Scale Dataset and a New Method for RemoteSensing Traffic Object Segmentation

Model ReleasesDGX agent

arXiv:2607.03945v1 Announce Type: new Abstract: Remote sensing imagery plays a crucial role in evaluating regional transportation capacity. However, existing segmentation datasets often lack diversity

A Multimodal Reasoning Typology for Grounding Chart-Image Coherence in Science Communication

Local AiDGX agent

arXiv:2607.05222v1 Announce Type: new Abstract: Charts and images appear together throughout scientific publications, yet most computational work does not characterize their coherence. We argue that a

A non-invasive video-based method for individual identification of wildlife using gait dynamics

ResearchDGX agent

arXiv:2607.04518v1 Announce Type: new Abstract: Gait is a distinctive behavioral characteristic that enables non-invasive individual identification without requiring physical interaction with an anima

A Reliable Context-Aware and Temporal Planning Framework for Autonomous Driving

Model ReleasesDGX agent

arXiv:2607.04689v1 Announce Type: cross Abstract: Safe operation of autonomous vehicles in dense urban traffic depends on perception and planning that remain reliable when onboard sensing is degraded.

A Simulation Framework for Electromagnetic Signal Injection Attacks on Image Sensors

ResearchDGX agent

arXiv:2408.05124v2 Announce Type: replace-cross Abstract: Image sensors are fundamental to many intelligent systems, allowing visual perception and AI-driven decision-making. However, their integrity

A Vision Based System for Guided and Collaborative Reconstruction of Fragmented Documents

ResearchDGX agent

arXiv:2607.03621v1 Announce Type: new Abstract: This paper presents the development and evaluation of a collaborative system for real-time reconstruction of fragmented paper documents in the context o

Adaptive Time-step Training for Enhancing Spike-Based Neural Radiance Fields

ResearchDGX agent

arXiv:2507.23033v2 Announce Type: replace Abstract: Spiking Neural Networks (SNNs) provide an energy-efficient computing paradigm for neural rendering, but existing spike-based Neural Radiance Field (

AdaptiveSplat:Texture Aware Controllable 3D Gaussian Allocation for Feed-Forward Reconstruction

ResearchDGX agent

arXiv:2607.04256v1 Announce Type: new Abstract: Current feed-forward 3D reconstruction methods predict pixel aligned Gaussian primitives, resulting in highly redundant representations. A natural solut

Agent-driven Long-tail Simulation for Autonomous Driving

Model ReleasesDGX agent

arXiv:2607.04331v1 Announce Type: cross Abstract: Evaluating autonomous driving systems in closed-loop settings requires realistic and interactive simulation, yet existing simulators largely rely on l

AgentFoX: LLM Agent-Guided Fusion with eXplainability for AI-Generated Image Detection

AgentsDGX agent

arXiv:2603.23115v2 Announce Type: replace Abstract: The realism of AI-generated images (AIGI) poses increasing challenges for reliable forensic detection, where heterogeneous expert detectors may prod

Agentic Very Long Video Understanding

AgentsDGX agent

arXiv:2601.18157v3 Announce Type: replace Abstract: The advent of always-on personal AI assistants, enabled by all-day wearable devices such as smart glasses, demands a new level of contextual underst

Aircraft Detection in Satellite Imagery using Deep Learning Object Detectors

TutorialsDGX agent

arXiv:2607.02699v1 Announce Type: new Abstract: The object detection in satellite imagery has garnered considerable attention due to its extensive real-world applications and the inherent challenges i

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO

ResearchDGX agent

arXiv:2602.06422v2 Announce Type: replace Abstract: Deploying GRPO on Flow Matching models has proven effective for text-to-image generation. However, existing paradigms typically propagate an outcome

An Automated Multimodal Glaucoma Detection Framework Using ViT and a Stacking-Based Ensemble

ResearchDGX agent

arXiv:2607.02692v1 Announce Type: new Abstract: Glaucoma is a progressive eye disease that can lead to irreversible vision loss if not detected at an early stage. Conventional diagnostic procedures ar

An event-driven framework for fly-inspired visual motion detection

AgentsDGX agent

arXiv:2607.05205v1 Announce Type: new Abstract: Fast and reliable motion detection is essential for machine vision and autonomous systems operating in dynamic environments. This work integrates emergi

AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis

SafetyDGX agent

arXiv:2512.11797v2 Announce Type: replace-cross Abstract: The collection of large-scale and diverse robot demonstrations remains a major bottleneck for imitation learning, as real-world data acquisiti

AnchorSplat: Fast and Structure Consistent Detail Synthesis for Gaussian Splatting

Model ReleasesDGX agent

arXiv:2607.01290v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful representation for high-fidelity rendering. However, existing assets often suffer from qualit

AnyDesign: Versatile Area Fashion Editing via Mask-Free Diffusion

ApplicationsDGX agent

arXiv:2408.11553v5 Announce Type: replace Abstract: Fashion image editing aims to modify a person's appearance based on a given instruction. Existing methods require auxiliary tools like segmenters an

AnyStyle: A Single LoRA is Sufficient for Image-Guided Style Transfer

ResearchDGX agent

arXiv:2607.04677v1 Announce Type: new Abstract: Image-guided style transfer aims to apply the artistic characteristics of a style image to a content image while preserving its semantic structure and l

Aperture-aware Dispersion 5-D Light-field Imaging Spectrometer

ApplicationsDGX agent

arXiv:2607.04635v1 Announce Type: new Abstract: Enhancing perceptual dimensions while miniaturizing imaging systems presents significant challenges for high-dimensional visual sensing. Conventionally,

AppAgent: Multimodal Agents as Smartphone Users

AgentsDGX agent

arXiv:2312.13771v3 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have led to the creation of intelligent agents capable of performing complex tasks. This paper i

AquaStereo: Enabling Underwater Stereo Matching via Depth-Conditioned Diffusion and Geometry Self-Distillation

SafetyDGX agent

arXiv:2607.04303v1 Announce Type: new Abstract: Learning-based stereo matching models struggle in underwater environments due to scarce in-domain data and the difficulty of extracting discriminative c

Atlas is Your Perfect Context: One-Shot Customization for Generalizable Foundational Medical Image Segmentation

ApplicationsDGX agent

arXiv:2512.18176v2 Announce Type: replace Abstract: Accurate segmentation of anatomical structures in medical images is essential for diagnosis and treatment planning. While recent interactive segment

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation

Model ReleasesDGX agent

arXiv:2606.12555v2 Announce Type: replace-cross Abstract: Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a

AULLM++: Structured-Token-Conditioned Large Language Models for Micro-Expression Action Unit Detection

Local AiDGX agent

arXiv:2603.08387v2 Announce Type: replace Abstract: Micro-expression Action Unit (AU) detection identifies localized AUs from subtle facial muscle activations, providing a foundation for decoding affe

Aura: Consistent Multi-Subject Video Generation via VLM-Grounded Semantic Alignment

SafetyDGX agent

arXiv:2607.04311v1 Announce Type: new Abstract: Subject-driven and multi-element video generation are central to controllable video synthesis, but existing methods still struggle to preserve identity

Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

ResearchDGX agent

arXiv:2607.02968v1 Announce Type: new Abstract: Massive Activations (MAs) have been widely observed in Transformer-based models, yet their structure and functional roles in Diffusion Transformers (DiT

BAT3R: Bootstrapping Articulated 3D Reconstruction from 2D Image Collections

ResearchDGX agent

arXiv:2607.03891v1 Announce Type: new Abstract: 3D reconstruction of articulated objects from a single image is challenging because large training datasets with paired image and 3D supervision are dif

Be Indiscrete: The Benefits of Learning Continuous Spine Degeneration Severity Scores

ResearchDGX agent

arXiv:2607.05090v1 Announce Type: new Abstract: Lumbar spine degeneration is a major contributor to chronic low back pain and is routinely assessed on MRI using ordinal grading systems, e.g. normal, m

Beyond Isolated Objects: Relationship-aware Open Vocabulary Scene Understanding via 3D Scene Graph Analysis

ResearchDGX agent

arXiv:2607.05348v1 Announce Type: new Abstract: Open-vocabulary 3D scene understanding aims to segment 3D scenes beyond predefined categories by transferring semantic knowledge from vision-language mo

← Previous
1…4748495051…209
Next →