AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Safety

Straight-Path Flow Matching for Incomplete Multi-View Clustering

DGX agent

arXiv:2607.06281v1 Announce Type: new Abstract: Incomplete Multi-View Clustering addresses the problem of clustering multi-modal data when certain views are missing. Recent end-to-end generative appro

safetyarxiv-cs-cv
8 Jul 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Structured-Condensed Prompt Tuning in Vision-Language Models for Fine-grained Image Recognition

DGX agent

arXiv:2607.06185v1 Announce Type: new Abstract: Fine-grained image recognition poses a significant challenge due to the substantial expertise and effort required for manual annotation. Vision-language

safetyarxiv-cs-cv
8 Jul 2026
Model Releases

Structured Data Extraction from Real Estate Documents using Clustering, Classification, and Large Language Models

DGX agent

arXiv:2607.06012v1 Announce Type: new Abstract: Real estate property listings expose structured metadata through the API. Still, the richest property-level information (i.e., legal status, structural

model-releasesarxiv-cs-cv
8 Jul 2026
Safety

Superman: Unifying Skeleton and Vision for Human Motion Perception and Generation

DGX agent

arXiv:2602.02401v2 Announce Type: replace Abstract: Human motion analysis tasks, such as temporal 3D pose estimation, motion prediction, and motion in-betweening, play an essential role in computer vi

safetyarxiv-cs-cv
8 Jul 2026
Safety

Synthetic-to-Real Translation for Class-Agnostic Motion Prediction

DGX agent

arXiv:2607.06319v1 Announce Type: new Abstract: Motion understanding is critical for ensuring safety and robustness in autonomous driving systems, driving increasing interest in motion prediction. A k

safetyarxiv-cs-cv
8 Jul 2026
Research

Taxlifier: Leveraging Disease Taxonomy for Enhanced Multi-Label Classification in Chest Radiography

DGX agent

arXiv:2607.05628v1 Announce Type: new Abstract: Accurate and efficient classification of thoracic diseases in chest X-ray (CXR) images is crucial for timely diagnosis and treatment. However, the prese

researcharxiv-cs-cv
8 Jul 2026
Research

Temporal Modeling of Optically Variable Devices in Identity Documents

DGX agent

arXiv:2607.06408v1 Announce Type: new Abstract: Robust remote verification of identity documents relies on analyzing faint, transparent security features like Optically Variable Devices (OVDs), or 'ho

researcharxiv-cs-cv
8 Jul 2026
Research

TMF-RSE: Tri-Modal Fusion with Regional Semantics and Evidential Uncertainty for Lung Severity Scoring

DGX agent

arXiv:2607.06356v1 Announce Type: cross Abstract: Accurate quantification of lung disease severity from chest imaging is critical for clinical decision-making and resource allocation. We propose a tri

researcharxiv-cs-cv
8 Jul 2026
Applications

Training-Free Acceleration for Vision-Language-Action Models with Action Caching and Refinement

DGX agent

arXiv:2607.06370v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a promising approach for generalizable robotic manipulations. In particular, flow matching-based V

applicationsarxiv-cs-cv
8 Jul 2026
Agents

TRIG: Trajectory-Rig Decoupled Metric Geometry Learning

DGX agent

arXiv:2607.05801v1 Announce Type: new Abstract: Vision-centric autonomous driving requires accurate metric geometry and ego-motion estimation from synchronized multi-camera observations. Recent visual

agentsarxiv-cs-cv
8 Jul 2026
Tutorials

Tuned Reverse Distillation: Enhancing Multimodal Industrial Anomaly Detection with Crossmodal Tuners

DGX agent

arXiv:2412.08949v4 Announce Type: replace Abstract: Knowledge distillation (KD) has been widely studied in unsupervised image Anomaly Detection (AD), but its application to unsupervised multimodal AD

tutorialsarxiv-cs-cv
8 Jul 2026
Hardware

Tuning-Free Latent Diffusion Models for Ultrahigh-Resolution Image Editing

DGX agent

arXiv:2607.06136v1 Announce Type: new Abstract: Recent diffusion-based generative models have shown impressive performance in image generation and editing. However, due to memory limitations and the h

hardwarearxiv-cs-cv
8 Jul 2026
Safety

UniField: A Unified Field-Aware MRI Enhancement Framework

DGX agent

arXiv:2603.09223v2 Announce Type: replace Abstract: Magnetic Resonance Imaging (MRI) field-strength enhancement holds immense value for both clinical diagnostics and advanced research. However, existi

safetyarxiv-cs-cv
8 Jul 2026
Research

Unlearnable Faces: Privacy Protection Surviving Extraction Pipeline

DGX agent

arXiv:2607.05996v1 Announce Type: new Abstract: Unlearnable examples keep publicly shared photos from being learned by unauthorized face-recognition models. An imperceptible perturbation, added before

researcharxiv-cs-cv
8 Jul 2026
Research

Unsupervised Domain Adaptation for Calcification Classification in Mammography Across Multi-Site Datasets

DGX agent

arXiv:2607.06549v1 Announce Type: new Abstract: Deep learning-based computer-aided diagnosis (CAD) systems have shown strong performance in breast cancer diagnosis, particularly for classification tas

researcharxiv-cs-cv
8 Jul 2026
Agents

VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery

DGX agent

arXiv:2607.06374v1 Announce Type: new Abstract: Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact e

agentsarxiv-cs-cv
8 Jul 2026
Safety

VEIL: How Visual Encoding Hijacking Induces Bias In Vision Models

DGX agent

arXiv:2607.05641v1 Announce Type: new Abstract: Rendering time series as chart images for CNN-based classification has become increasingly common in time-series classification (TSC). However, it remai

safetyarxiv-cs-cv
8 Jul 2026
Model Releases

Verification of Dynamic Holographic Behavior in Identity Documents

DGX agent

arXiv:2607.06466v1 Announce Type: new Abstract: This paper addresses the remote verification of the authenticity of Optically Variable Devices (commonly known as holograms) on identity documents. Typi

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models

DGX agent

arXiv:2510.13808v2 Announce Type: replace Abstract: Large Vision Language Models (VLMs) excel at general visual reasoning but experience significant performance degradation when deployed in novel doma

model-releasesarxiv-cs-cv
8 Jul 2026
Research

Vision as Unified Multimodal Generation

DGX agent

arXiv:2607.06560v1 Announce Type: new Abstract: We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation sp

researcharxiv-cs-cv
8 Jul 2026
Research

Visual graphs for image classification: does the structure affect performance?

DGX agent

arXiv:2607.06295v1 Announce Type: new Abstract: Deep learning models have emerged in machine learning and related fields, demonstrating astonishing performance in various visual tasks. Despite their g

researcharxiv-cs-cv
8 Jul 2026
Model Releases

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

DGX agent

arXiv:2607.06118v1 Announce Type: new Abstract: As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their nav

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

Why does Deep Learning Improve Visual SLAM?

DGX agent

arXiv:2607.06023v1 Announce Type: new Abstract: Visual SLAM is a well-established technology utilized in a wide range of real-world applications. However, its performance still degrades under challeng

model-releasesarxiv-cs-cv
8 Jul 2026
Tutorials

WING: A Window-Prior-Based Generative Network with Gated Inception for Cross-Modality CT Synthesis

DGX agent

arXiv:2607.06234v1 Announce Type: new Abstract: Generating CT volumes from MRI and CBCT can improve treatment planning in adaptive radiotherapy while avoiding additional radiation exposure. However, d

tutorialsarxiv-cs-cv
8 Jul 2026
Tutorials

WristMimic: Full-Body Humanoid Control with Wrist-Guided Manipulation

DGX agent

arXiv:2607.06438v1 Announce Type: cross Abstract: Retargeting human object interaction demonstrations to physics based simulation requires reproducing not only body motion but also the object motion a

tutorialsarxiv-cs-cv
8 Jul 2026
Model Releases

XRFormer: Multiscale Tokenization for XRF Representation Learning

DGX agent

arXiv:2607.06424v1 Announce Type: new Abstract: X-ray fluorescence (XRF) spectroscopy is a key modality for material analysis in cultural heritage. However, automated learning from XRF spectra remains

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

DGX agent

arXiv:2502.09696v3 Announce Type: replace Abstract: Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children

model-releasesarxiv-cs-cv
8 Jul 2026
Research

3DMPE: 3D Multi-Perspective Embedding

DGX agent

arXiv:2607.04898v1 Announce Type: new Abstract: We study 3D point cloud reconstruction from multiple partially observed 2D projections. Given two or more projections of an unknown 3D point cloud, toge

researcharxiv-cs-cv
7 Jul 2026
Agents

A Cost-Aware, Paired Protocol for Auditing Dynamic Tool Synthesis in Agentic Video Question Answering

DGX agent

arXiv:2607.01469v2 Announce Type: replace Abstract: Agentic Video Question Answering (VideoQA) systems invoke tools during inference, but their tool libraries are fixed, so recurring procedures are re

agentsarxiv-cs-cv
7 Jul 2026
Research

A Decomposable Probe for Few-Step Diffusion Models: Prompt, Latent, and Score Selectivity across Backbone Families and Distillation Paradigms

DGX agent

arXiv:2607.03256v1 Announce Type: new Abstract: Few-step distilled diffusion students cut text-to-image inference from ~50 to 1-8 network evaluations, but the quality gap is usually summarised by a si

researcharxiv-cs-cv
7 Jul 2026
Model Releases

A Large-Scale Dataset and a New Method for RemoteSensing Traffic Object Segmentation

DGX agent

arXiv:2607.03945v1 Announce Type: new Abstract: Remote sensing imagery plays a crucial role in evaluating regional transportation capacity. However, existing segmentation datasets often lack diversity

model-releasesarxiv-cs-cv
7 Jul 2026
Local Ai

A Multimodal Reasoning Typology for Grounding Chart-Image Coherence in Science Communication

DGX agent

arXiv:2607.05222v1 Announce Type: new Abstract: Charts and images appear together throughout scientific publications, yet most computational work does not characterize their coherence. We argue that a

local-aiarxiv-cs-cv
7 Jul 2026
Research

A non-invasive video-based method for individual identification of wildlife using gait dynamics

DGX agent

arXiv:2607.04518v1 Announce Type: new Abstract: Gait is a distinctive behavioral characteristic that enables non-invasive individual identification without requiring physical interaction with an anima

researcharxiv-cs-cv
7 Jul 2026
Model Releases

A Reliable Context-Aware and Temporal Planning Framework for Autonomous Driving

DGX agent

arXiv:2607.04689v1 Announce Type: cross Abstract: Safe operation of autonomous vehicles in dense urban traffic depends on perception and planning that remain reliable when onboard sensing is degraded.

model-releasesarxiv-cs-cv
7 Jul 2026
Research

A Simulation Framework for Electromagnetic Signal Injection Attacks on Image Sensors

DGX agent

arXiv:2408.05124v2 Announce Type: replace-cross Abstract: Image sensors are fundamental to many intelligent systems, allowing visual perception and AI-driven decision-making. However, their integrity

researcharxiv-cs-cv
7 Jul 2026
Research

A Vision Based System for Guided and Collaborative Reconstruction of Fragmented Documents

DGX agent

arXiv:2607.03621v1 Announce Type: new Abstract: This paper presents the development and evaluation of a collaborative system for real-time reconstruction of fragmented paper documents in the context o

researcharxiv-cs-cv
7 Jul 2026
Research

Adaptive Time-step Training for Enhancing Spike-Based Neural Radiance Fields

DGX agent

arXiv:2507.23033v2 Announce Type: replace Abstract: Spiking Neural Networks (SNNs) provide an energy-efficient computing paradigm for neural rendering, but existing spike-based Neural Radiance Field (

researcharxiv-cs-cv
7 Jul 2026
Research

AdaptiveSplat:Texture Aware Controllable 3D Gaussian Allocation for Feed-Forward Reconstruction

DGX agent

arXiv:2607.04256v1 Announce Type: new Abstract: Current feed-forward 3D reconstruction methods predict pixel aligned Gaussian primitives, resulting in highly redundant representations. A natural solut

researcharxiv-cs-cv
7 Jul 2026
Model Releases

Agent-driven Long-tail Simulation for Autonomous Driving

DGX agent

arXiv:2607.04331v1 Announce Type: cross Abstract: Evaluating autonomous driving systems in closed-loop settings requires realistic and interactive simulation, yet existing simulators largely rely on l

model-releasesarxiv-cs-cv
7 Jul 2026
Agents

AgentFoX: LLM Agent-Guided Fusion with eXplainability for AI-Generated Image Detection

DGX agent

arXiv:2603.23115v2 Announce Type: replace Abstract: The realism of AI-generated images (AIGI) poses increasing challenges for reliable forensic detection, where heterogeneous expert detectors may prod

agentsarxiv-cs-cv
7 Jul 2026
Agents

Agentic Very Long Video Understanding

DGX agent

arXiv:2601.18157v3 Announce Type: replace Abstract: The advent of always-on personal AI assistants, enabled by all-day wearable devices such as smart glasses, demands a new level of contextual underst

agentsarxiv-cs-cv
7 Jul 2026
Tutorials

Aircraft Detection in Satellite Imagery using Deep Learning Object Detectors

DGX agent

arXiv:2607.02699v1 Announce Type: new Abstract: The object detection in satellite imagery has garnered considerable attention due to its extensive real-world applications and the inherent challenges i

tutorialsarxiv-cs-cv
7 Jul 2026
Research

Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO

DGX agent

arXiv:2602.06422v2 Announce Type: replace Abstract: Deploying GRPO on Flow Matching models has proven effective for text-to-image generation. However, existing paradigms typically propagate an outcome

researcharxiv-cs-cv
7 Jul 2026
Research

An Automated Multimodal Glaucoma Detection Framework Using ViT and a Stacking-Based Ensemble

DGX agent

arXiv:2607.02692v1 Announce Type: new Abstract: Glaucoma is a progressive eye disease that can lead to irreversible vision loss if not detected at an early stage. Conventional diagnostic procedures ar

researcharxiv-cs-cv
7 Jul 2026
Agents

An event-driven framework for fly-inspired visual motion detection

DGX agent

arXiv:2607.05205v1 Announce Type: new Abstract: Fast and reliable motion detection is essential for machine vision and autonomous systems operating in dynamic environments. This work integrates emergi

agentsarxiv-cs-cv
7 Jul 2026
Safety

AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis

DGX agent

arXiv:2512.11797v2 Announce Type: replace-cross Abstract: The collection of large-scale and diverse robot demonstrations remains a major bottleneck for imitation learning, as real-world data acquisiti

safetyarxiv-cs-cv
7 Jul 2026
Model Releases

AnchorSplat: Fast and Structure Consistent Detail Synthesis for Gaussian Splatting

DGX agent

arXiv:2607.01290v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful representation for high-fidelity rendering. However, existing assets often suffer from qualit

model-releasesarxiv-cs-cv
7 Jul 2026
Applications

AnyDesign: Versatile Area Fashion Editing via Mask-Free Diffusion

DGX agent

arXiv:2408.11553v5 Announce Type: replace Abstract: Fashion image editing aims to modify a person's appearance based on a given instruction. Existing methods require auxiliary tools like segmenters an

applicationsarxiv-cs-cv
7 Jul 2026
← Previous
1…5960616263…261
Next →