AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
31 Jul 2026

Continual Learning with Vision-Language Models via Semantic-Geometry Preservation

ResearchDGX agent

arXiv:2603.12055v3 Announce Type: replace Abstract: Continual learning of pretrained vision-language models (VLMs) is prone to catastrophic forgetting, yet current approaches adapt to new tasks withou

Convolutional Neural Shading for High-Quality 3D Reconstruction from Multi-View Images

ResearchDGX agent

arXiv:2607.28132v1 Announce Type: new Abstract: We propose a convolutional neural shading (CNS), a novel pipeline to reconstruct high-quality 3D shapes from multi-view images. Several recent studies h

CoRE-UIR: Prior-guided common and residual experts for efficient all-in-one remote sensing image restoration

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.27898v1 Announce Type: new Abstract: Remote sensing images acquired by unmanned aerial vehicles (UAVs) and satellites are often degraded by adverse weather, illumination variation, and imag

Cross-Embodiment Transfer via Behavior-Aligned Representations

Model ReleasesDGX agent

arXiv:2607.27549v1 Announce Type: cross Abstract: Recent progress in large-scale imitation learning for robot manipulation has been driven by leveraging datasets across a wide range of robot embodimen

CXR-Retrieve: Compositional Text-to-Image Retrieval in Chest Radiography

Model ReleasesDGX agent

arXiv:2607.27779v1 Announce Type: new Abstract: Large chest radiography archives are difficult to search because most studies are paired only with free-text reports rather than structured clinical ann

DAS-PMVC: A Framework for Partial Multi-View Clustering via Dual Alignment and Structure Enhancement

SafetyDGX agent

arXiv:2607.27761v1 Announce Type: cross Abstract: In recent years, multi-view clustering has attracted widespread research interest. However, due to limitations in data collection devices, data across

DECODE: Tackling Representation and Decision Degradation in Continual AI-Generated Image Detection

SafetyDGX agent

arXiv:2607.27882v1 Announce Type: new Abstract: As generative models continue to evolve, AI-generated image detectors must incrementally adapt to emerging generative domains while preserving knowledge

Deep learning-based hierarchical insect classification using camera trap imagery

ResearchDGX agent

arXiv:2607.28005v1 Announce Type: new Abstract: Declining insect populations make reliable biodiversity monitoring increasingly urgent, yet monitoring of insect biodiversity is hampered by a lack of s

DinoLizer: Separating VAE and Diffusion Artifacts in Generative Inpainting Localization

Local AiDGX agent

arXiv:2511.20722v2 Announce Type: replace Abstract: We introduce DinoLizer, a DINOv2-based localizer of manipulated areas in generative inpainting. The model is trained to focus on semantically altere

Drawing-Recode: Annotation Grounding for Parametric CAD Code Generation from Raster 2D CAD Drawings

ApplicationsDGX agent

arXiv:2607.27558v1 Announce Type: new Abstract: Recovering Parametric CAD sequences from raster-format 2D Computer-Aided Design (CAD) drawings accumulated prior to digital transformation is important

DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis

Model ReleasesDGX agent

arXiv:2607.27763v1 Announce Type: new Abstract: We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with tw

EEG-EditBench: Probing Visual Information in EEG-Image Retrieval Models with Controlled Image Edits

Model ReleasesDGX agent

arXiv:2607.27857v1 Announce Type: new Abstract: Recent EEG-to-image retrieval models have achieved strong performance in identifying viewed images from semantically diverse candidates. Yet such succes

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

SafetyDGX agent

arXiv:2607.28243v1 Announce Type: new Abstract: Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodi

EgoGVAE: Ego-body Mesh Reconstruction via Guided Variational Autoencoder

Model ReleasesDGX agent

arXiv:2607.27755v1 Announce Type: new Abstract: We address the problem of recovering the full-body mesh from only the head pose. This task has become essential for various applications based on head-m

EHGCN: Hierarchical Euclidean-Hyperbolic Fusion via Motion-Aware GCN for Hybrid Event Stream Perception

Model ReleasesDGX agent

arXiv:2504.16616v4 Announce Type: replace Abstract: Event cameras, characterized by microsecond temporal resolution and very High Dynamic Range (HDR), emit high-speed event streams for perception task

EmoFeedback^2: Reinforcement of Continuous Emotional Image Generation via LVLM-based Reward and Textual Feedback

SafetyDGX agent

arXiv:2511.19982v3 Announce Type: replace Abstract: Continuous emotional image content generation (C-EICG) is emerging rapidly due to its ability to produce images aligned with both user descriptions

ENCORE: Event-Assisted Complementary Motion Refinement for Learned Video Compression

ResearchDGX agent

arXiv:2607.28020v1 Announce Type: new Abstract: Learned video compression relies on accurate temporal modeling to remove redundancy between adjacent frames. However, most existing codecs infer motion

Endo-NeRF++: Uncertainty-Aware Neural Rendering with Multi-Resolution Hash Encoding for Dynamic Surgical Scene Reconstruction

ResearchDGX agent

arXiv:2607.27825v1 Announce Type: cross Abstract: Reconstructing dynamic surgical scenes is crucial for robot-assisted minimally invasive surgery; however, it continues to be difficult because of tiss

Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models

ResearchDGX agent

arXiv:2603.05950v2 Announce Type: replace Abstract: Visual token reduction is critical for accelerating Vision-Language Models (VLMs), since visual inputs are represented as token sequences that intro

Enhancing Scene Transition Awareness in Video Generation via Post-Training

TutorialsDGX agent

arXiv:2507.18046v2 Announce Type: replace Abstract: Recent advances in AI-generated video have shown strong performance on text-to-video tasks, particularly for short clips depicting a single scene. H

Explaining Image Similarity with Automatically Extracted Concept Activation Vectors

Local AiDGX agent

arXiv:2607.28386v1 Announce Type: new Abstract: Image similarity underlies many computer vision applications, yet it is often unclear why two images receive a high or low similarity score. Existing ex

Face and Voice Cross-modal Association with Learning Convex Feature Embedding

ResearchDGX agent

arXiv:2607.28129v1 Announce Type: new Abstract: Face-and-voice association learning is one of the most challenging tasks in deep learning. In this paper, we propose a simple but powerful cross-modal f

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification

AgentsDGX agent

arXiv:2607.28225v1 Announce Type: new Abstract: Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, h

FDDWAN: A Frequency-Decoupled Diffusion Network for Watermarking Attack

ResearchDGX agent

arXiv:2607.27800v1 Announce Type: new Abstract: Existing invisible watermark removal methods often struggle to accurately capture the watermark-bearing features, leading to an unfavorable trade-off be

FeatFix: Reuse What You Verify through Local Exact-Feature Correction for Faster Cached Diffusion Inference

Local AiDGX agent

arXiv:2607.27842v1 Announce Type: new Abstract: Diffusion models are widely used to generate high-quality images and videos, but their iterative denoising process remains computationally intensive. A

Filling the Pareto-Optimal Front for Affordance Segmentation on Embedded Devices Using RGB-D Cameras

ApplicationsDGX agent

arXiv:2607.28293v1 Announce Type: new Abstract: While depth sensors have the potential to complement RGB data for affordance segmentation in wearable robots, their usage seems to remain underexplored.

Finding Change in Satellite Archives from Text: How to Combine Before-and-After Images Efficiently

TutorialsDGX agent

arXiv:2607.28571v1 Announce Type: new Abstract: Operational Earth observation increasingly calls for answering queries such as ``find the image pairs where a new building appeared.'' This means search

FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval

SafetyDGX agent

arXiv:2607.27959v1 Announce Type: new Abstract: Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significa

FlexiGrad: Adaptive Gradient Modulation for Hierarchical Fine-Grained Classification

Model ReleasesDGX agent

arXiv:2607.17563v2 Announce Type: replace Abstract: Many fine-grained recognition tasks contain hierarchical labels such as order, family and species. Although this supervision should be beneficial, j

FootprintNet: State-Transition-Guided Dynamic Footprint Learning for Multi-temporal Remote Sensing Change Detection

TutorialsDGX agent

arXiv:2607.27969v1 Announce Type: new Abstract: Despite substantial progress in remote sensing multi-temporal change detection (MTCD), most existing MTCD methods still represent the dynamic process at

GVR-Coder: A Visual-Feedback Framework for Structured SVG Generation in Complex Document and Meeting Scenarios

AgentsDGX agent

arXiv:2607.28073v1 Announce Type: cross Abstract: In demanding professional environments and meeting review scenarios, lengthy text often imposes a high cognitive load. To facilitate efficient informa

Hallucinations Leave a Grounding Signature:Verifier-Guided Decoding for Selective Object Correction

ResearchDGX agent

arXiv:2607.27823v1 Announce Type: new Abstract: Large vision-language models (LVLMs) often hallucinate objects that are absent from an image. Despite recent progress, existing mitigation methods still

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer

ResearchDGX agent

arXiv:2607.28394v1 Announce Type: new Abstract: Hand-object interaction (HOI) modeling remains challenging because it requires joint reasoning about hand articulation, object geometry, contact, semant

Human Mesh Modeling for Anny Body

ResearchDGX agent

arXiv:2511.03589v3 Announce Type: replace Abstract: Parametric body models provide the structural basis for many human-centric tasks, yet existing models often rely on costly 3D scans and learned shap

ID-Guard: A Universal Framework for Combating Facial Manipulation via Breaking Identification

ResearchDGX agent

arXiv:2409.13349v3 Announce Type: replace Abstract: The misuse of deep learning-based facial manipulation poses a serious threat to civil rights. To prevent such fraud at its source, proactive defense

IGME: Efficient Chained Method Ensemble for Transferable Semantic Segmentation Attacks

ResearchDGX agent

arXiv:2607.27465v1 Announce Type: new Abstract: Semantic segmentation models are vulnerable to transferable adversarial perturbations, yet evaluating transfer attacks on dense prediction models can be

Improved Classification of Nitrogen Stress Severity in Plants Under Combined Stress Conditions Using Spatio-Temporal Deep Learning Framework

ResearchDGX agent

arXiv:2509.06625v3 Announce Type: replace Abstract: Plants in their natural habitats endure an array of interacting stresses, both biotic and abiotic, that rarely occur in isolation. Nutrient stress-p

Inference-Time Agentic Decision Rules Beat Longer Evolving Search for Multi-Image Medical Reasoning

SafetyDGX agent

arXiv:2607.27564v1 Announce Type: new Abstract: Multi-image medical VQA is not merely a prompt-length problem; it is a fundamental challenge of agentic decision-making. Medical vision-language agents

Isolating to Harness: Cross-Division Distillation for Fully Unsupervised Anomaly Detection

TutorialsDGX agent

arXiv:2508.18007v2 Announce Type: replace Abstract: Fully Unsupervised Anomaly Detection (FUAD) addresses the practical scenario where training data is contaminated with unlabeled anomalies. This sett

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Model ReleasesDGX agent

arXiv:2607.27670v1 Announce Type: new Abstract: Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that creat

Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image Classification

Local AiDGX agent

arXiv:2607.28428v1 Announce Type: cross Abstract: We introduce Kohn--Sham Spectral Embedding (KSSE), a physics-inspired energy-based model replacing dense CNN classifiers with a sparse-graph spectral

Landmark shape spaces with induced metrics

ResearchDGX agent

arXiv:2607.28064v1 Announce Type: new Abstract: We present a unification of Kendall's landmark shape spaces, where rigid motions are factored out and scale fixed on landmark configurations equipped wi

Large scale cross-regional remote sensing flood monitoring framework for operative mapping and impact analysis

ResearchDGX agent

arXiv:2607.28401v1 Announce Type: new Abstract: Effective flood monitoring is critical for minimizing the impacts of flood disasters on populations and infrastructure. Yet reliable remote sensing acro

LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference

ResearchDGX agent

arXiv:2607.27952v1 Announce Type: new Abstract: Multimodal foundation models are reshaping edge-cloud visual intelligence from task-specific feature pipelines into token-based interfaces, where edge d

Learning Color Grading, No Photo Sharing: Federated Aesthetic Preference Learning for Personalized Image Enhancement

Model ReleasesDGX agent

arXiv:2607.27659v1 Announce Type: new Abstract: Personalized image enhancement should reflect individual aesthetic taste, yet learning such preferences commonly depends on private photos and ratings t

Learning to Understand Body Language from Flight through Robust 3D Avatar Placing

TutorialsDGX agent

arXiv:2607.27865v1 Announce Type: new Abstract: Perceiving human motion and intent at long range is a prerequisite for socially intelligent aerial robots, yet the data to learn it barely exists. We in

LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA

Model ReleasesDGX agent

arXiv:2607.27806v1 Announce Type: new Abstract: In clinical practice, patients often undergo multiple imaging examinations over successive visits, yielding longitudinal data. Modeling such temporal in

MarkushGlyph and OCSRGlyph: Improved Chemical Structure Recognition

ResearchDGX agent

arXiv:2607.28532v1 Announce Type: new Abstract: Chemical structures appear in patents and the scientific literature as images. For programmatic usage, such as indexing in databases or constructing mac

MedXplore: Towards Reliable and Unbiased Generalized Category Discovery in Medical Imaging

SafetyDGX agent

arXiv:2607.27620v1 Announce Type: new Abstract: Deep learning has shown strong potential in medical image analysis, but most existing methods rely on large-scale annotations and a closed-world assumpt

MeshFM: 2D Features Are All You Need for 3D Shape Understanding

ResearchDGX agent

arXiv:2607.27592v1 Announce Type: new Abstract: We present MeshFM, an efficient feedforward framework for extracting rich features from 3D inputs. Our method distills 2D features from visual foundatio

MetaRank: Task-Aware Metric Selection for Model Transferability Estimation

TutorialsDGX agent

arXiv:2511.21007v2 Announce Type: replace Abstract: Selecting an appropriate pre-trained source model is a critical, yet computationally expensive, task in transfer learning. Model Transferability Est

MIND: Multimodal Intent-Driven Network via Diffusion Transformers for Medical Image Fusion

SafetyDGX agent

arXiv:2607.28565v1 Announce Type: new Abstract: Medical image fusion aims to integrate complementary information from diverse imaging modalities to support clinical diagnosis. Existing methods typical

MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision Transformers

ResearchDGX agent

arXiv:2607.28589v1 Announce Type: new Abstract: Post-training quantization (PTQ) has emerged as an effective solution for deploying Vision Transformers (ViTs) on resource-constrained devices. However,

MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos

Model ReleasesDGX agent

arXiv:2607.27895v1 Announce Type: cross Abstract: Mental health understanding in long-form videos requires nuanced reasoning over observable behavior, interpersonal context, and latent psychological s

MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2607.27637v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or sh

mmRadarTwin: A Measurement-Calibrated Signal-Level Digital Twin Platform for Indoor mmWave Radar

ResearchDGX agent

arXiv:2607.28108v1 Announce Type: new Abstract: Indoor mmWave radar perception is difficult to reproduce because measured range-angle responses depend on scene geometry, material response, multipath,

MonoVoc: Decoupling Geometry and Semantics for Lightweight Monocular Open-Vocabulary 3D Gaussians

ResearchDGX agent

arXiv:2607.28300v1 Announce Type: new Abstract: Open vocabulary 3D scene understanding is essential for next-generation interactive systems, empowering users to intuitively query and navigate reconstr

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

Model ReleasesDGX agent

arXiv:2511.12449v3 Announce Type: replace Abstract: Recent Multimodal Large Language Models (MLLMs) have significantly advanced e-commerce product understanding. However, they still face three challen

Morphological Detection and Classification of Microplastics and Nanoplastics Emerged from Consumer Products by Deep Learning

ResearchDGX agent

arXiv:2409.13688v2 Announce Type: replace Abstract: Plastic pollution presents an escalating global issue, impacting health and environmental systems, with micro- and nanoplastics found across mediums

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

Model ReleasesDGX agent

arXiv:2607.27616v1 Announce Type: new Abstract: Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into share

← Previous
1…2122232425…207
Next →