AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
28 Jul 2026

Data Pyramid for Embodied Manipulation

SafetyDGX agent

arXiv:2607.24744v1 Announce Type: cross Abstract: Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require d

DDVT: Dynamic Dual-level Vision Transformer Fusion Network for Answer Grounding in Visual Question Answering

ResearchDGX agent

arXiv:2607.23921v1 Announce Type: new Abstract: Answer grounding in visual question answering aims to locate the region from a given natural language question associated with the visual content of an

Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos

ApplicationsDGX agent

arXiv:2501.13335v4 Announce Type: replace Abstract: We introduce a novel framework for modeling high-fidelity, animatable 3D human avatars from motion-blurred monocular video inputs. Motion blur is pr


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding

ResearchDGX agent

arXiv:2607.24554v1 Announce Type: cross Abstract: Advancing multimodal retrieval-augmented generation (RAG) for complex document understanding presents a formidable dual dilemma of accuracy and effici

Dementia Etiology Diagnosis via Collaborative Meta Knowledge Enhancement

SafetyDGX agent

arXiv:2607.22770v1 Announce Type: cross Abstract: Although artificial intelligence (AI) has shown promising performance in several medical tasks, accurate dementia etiology diagnosis with AI remains c

Denoising 3D images: robustness of persistent homology measures

ResearchDGX agent

arXiv:2607.24579v1 Announce Type: cross Abstract: When computing sub/super-level-set persistent homology (PH), the effect of noise may introduce millions of (short-lived) topological generators, prese

DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning

SafetyDGX agent

arXiv:2607.24159v1 Announce Type: cross Abstract: Generalizable robot manipulation requires policies that can anticipate how visual scenes evolve while executing language instructions. While recent Vi

Development of Vision-Language Model-based GNSS Spoofing Detection for Autonomous Vehicle Navigation

Local AiDGX agent

arXiv:2607.23962v1 Announce Type: new Abstract: Autonomous vehicles (AVs) depend on Global Navigation Satellite Systems (GNSS) for localization and navigation, making them vulnerable to spoofing attac

DINOv3-MIL: Per-Kidney Multi-Label Tumour and Cyst Detection from Foundation-Model Patch Tokens on KiTS23

ResearchDGX agent

arXiv:2607.22687v1 Announce Type: new Abstract: Foundation vision models trained on natural images transfer to medical tasks without domain pre-training, but volumetric classification requires aggrega

Direction-adaptive Mamba: Spatial-Frequency Dual-Domain Collaborative Learning for PolSAR Image Classification

ApplicationsDGX agent

arXiv:2607.23464v1 Announce Type: cross Abstract: Deep learning dominates polarimetric synthetic aperture radar (PolSAR) image classification, with Mamba architectures serving as favorable backbones d

DishSeg24k: A Large-Scale Benchmark for Food Segmentation with Stochastic Expert Decoding

Model ReleasesDGX agent

arXiv:2607.23070v1 Announce Type: new Abstract: Food segmentation is essential for applications such as intelligent catering, dietary assessment, and recommendation. However, existing benchmarks fail

DispatchRAG: Grounding Emergency Dispatch Decisions in Real-World Protocols from Traffic Accident Video

AgentsDGX agent

arXiv:2607.23132v1 Announce Type: new Abstract: Assessing the severity of a traffic accident scenario is important to decide which emergency service to dispatch. Missing an ambulance dispatch on a ped

DreamStyle3D: Efficient 3D Stylized Asset Generation via Dual-Attention Disentanglement

ResearchDGX agent

arXiv:2607.24721v1 Announce Type: new Abstract: With the growth of gaming, animation, and virtual reality industries, the demand for efficient generation of stylized 3D assets is rapidly increasing. H

DY-LUT: Depth-Aware YCbCr Lookup Tables for Real-Time Underwater Image Enhancement

Model ReleasesDGX agent

arXiv:2607.22801v1 Announce Type: cross Abstract: Underwater image enhancement is challenged by spatially non-uniform, wavelength-dependent attenuation. Propagation distance and wavelength govern this

EditCLEVR: A Paired-Scene Intervention Benchmark for Compositional Faithfulness of Object-Centric Representations

Model ReleasesDGX agent

arXiv:2607.22705v1 Announce Type: new Abstract: Object-centric learning aims to represent scenes as objects whose properties can be reused in new combinations. Existing evaluations usually score segme

Effect of User-Prompted Priors on Semi-Automated Cancer Lesion Segmentation in Whole-Body Computed Tomography

TutorialsDGX agent

arXiv:2607.24210v1 Announce Type: new Abstract: In clinical oncology studies, metastatic cancer is commonly evaluated using 'Response Evaluation Criteria in Solid Tumors' (RECIST), in which the diamet

Effective Receptive Field Ordering Matters for Infrared Small Target Detection

ResearchDGX agent

arXiv:2607.23994v1 Announce Type: new Abstract: In this work, we investigate a previously unexplored architectural dimension for infrared small target detection: the organization of effective receptiv

Embeddings based Anomaly Detection for Cleaning Global Crop Type Reference Datasets

Local AiDGX agent

arXiv:2607.23908v1 Announce Type: new Abstract: High quality reference data remain a critical bottleneck for crop-type mapping at any spatial and temporal scale. Operational systems such as WorldCerea

Energy Constrained Hierarchical Underwater Monitoring via Local Multi-Agent RAG

HardwareDGX agent

arXiv:2607.24313v1 Announce Type: cross Abstract: Marine life monitoring is limited by strict energy constraints, poor underwater connectivity, and the high cost of transmitting raw multimodal data fr

Evaluation of Blood Vessel Segmentation Methods on Hard-to-Detect Vascular Structures

Model ReleasesDGX agent

arXiv:2406.13128v2 Announce Type: replace Abstract: Due to the intricate structure of vascular trees, minor segmentation errors can significantly alter connectivity patterns and increase variability i

Event Driven Clustering Algorithm

ResearchDGX agent

arXiv:2602.00115v2 Announce Type: replace Abstract: This paper introduces a novel asynchronous, event-driven algorithm for real-time detection of small event clusters in event camera data. Similar to

Face Age Verification Vulnerabilities Under Simple Appearance Manipulations

SafetyDGX agent

arXiv:2607.24194v1 Announce Type: new Abstract: Online platforms increasingly rely on automated age estimation systems to enforce minimum-age policies. Focusing on vision-based models designed for thi

Farm-LightSeek: An Edge-centric Multimodal Agricultural IoT Data Analytics Framework with Lightweight LLMs

ApplicationsDGX agent

arXiv:2506.03168v2 Announce Type: replace Abstract: Amid the challenges posed by global population growth and climate change, traditional agricultural Internet of Things (IoT) systems is currently und

Fast Fourier Convolutional GAN for 30 m Clear-Sky Land Surface Temperature Gap-Free Reconstruction

ResearchDGX agent

arXiv:2607.22734v1 Announce Type: new Abstract: Satellite-derived Land Surface Temperature (LST) provides spatially comprehensive data that ground stations cannot match. However, its utility is freque

Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow

ResearchDGX agent

arXiv:2602.14021v2 Announce Type: replace Abstract: Reconstructing and tracking dynamic 3D scenes is a fundamental challenge in computer vision. Existing methods typically decouple geometry from motio

FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models

SafetyDGX agent

arXiv:2607.24522v1 Announce Type: cross Abstract: While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow

FogDrive: A Multi-Modal Synthetic Driving Dataset for Perception under Graded Fog

Model ReleasesDGX agent

arXiv:2607.22698v1 Announce Type: new Abstract: Perception under adverse weather remains a critical bottleneck for reliable autonomous driving, yet existing benchmarks lack the systematic multi-modal

Frequency-Aware Dual-Stream Learning for Balanced Realism and Fidelity in Electron Microscopy Imaging

ResearchDGX agent

arXiv:2607.22765v1 Announce Type: cross Abstract: Electron microscopy enables nanoscale cellular visualization but faces a trade-off between imaging resolution and acquisition speed. Existing learning

GaitFace: A Multimodal Dataset for Long-Range Person Identification

Model ReleasesDGX agent

arXiv:2607.23542v1 Announce Type: new Abstract: Efficient border control is becoming a significant global challenge, mainly due to severe congestion and extended passenger waiting times. To mitigate t

gamma-Bridge: A Look-Parametric Diffusion Bridge

Model ReleasesDGX agent

arXiv:2607.22719v1 Announce Type: new Abstract: Multiplicative Gamma noise is a signal-dependent degradation in coherent imaging; synthetic aperture radar (SAR) despeckling is its most prominent real-

Gaze-Anchored Social Net: Decoding Implicit Relations via Joint Modeling

Model ReleasesDGX agent

arXiv:2607.22847v1 Announce Type: new Abstract: Human gaze does more than point to visual targets; it serves as a subtle indicator of social intent within static images, whereas standard models typica

Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention

TutorialsDGX agent

arXiv:2607.23917v1 Announce Type: new Abstract: We introduce a novel learning problem: decoding gaze into natural language descriptions of human goals across diverse visual tasks. Unlike prior work, w

GazeLT: Visual attention-guided long-tailed disease classification in chest radiographs

TutorialsDGX agent

arXiv:2508.09478v2 Announce Type: replace Abstract: In this work, we present GazeLT, a human visual attention integration-disintegration approach for long-tailed disease classification. A radiologist'

GCR: Geometry-Consistent Routing for Task-Agnostic Continual Anomaly Detection

ResearchDGX agent

arXiv:2601.01856v3 Announce Type: replace Abstract: Feature-based anomaly detection is widely adopted in industrial inspection due to the strong representational power of large pre-trained vision enco

Generative Augmentation for EEG Motor Imagery Classification: A Class-Conditional VAE with Cycle-Consistent Decoder Refinement

ResearchDGX agent

arXiv:2607.22733v1 Announce Type: new Abstract: We investigate whether a generative model can supply useful synthetic motor-imagery (MI) electroencephalography (EEG) trials that improve the accuracy o

Generative Video Compression with Adaptive Score Distillation

Local AiDGX agent

arXiv:2607.22772v1 Announce Type: cross Abstract: Diffusion models provide strong generative capabilities for video compression at ultra-low bitrates. Existing diffusion-based video codecs adapt base

GenSplatCodec: Feed-Forward Gaussian Splatting Compression via One-Step Diffusion

ResearchDGX agent

arXiv:2607.24403v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting (3DGS) enables scalable scene reconstruction without per-scene optimization, yet produces dense Gaussians that are co

Geometry Meets Semantics: Fractional Gradient Stabilization for Semantic-Driven Bounding Box Optimization in Visual Detection Tasks

TutorialsDGX agent

arXiv:2607.23530v1 Announce Type: new Abstract: Bounding boxes are fundamental for object localization in visual detection tasks. Among them, oriented bounding boxes are widely used in visual detectio

GNM Head: A Generative aNthropometric Model of the human head

ResearchDGX agent

arXiv:2607.23687v1 Announce Type: new Abstract: Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animation, rendering, and reconstruction.

GRAPE: Graduated Routing for Articulated Portrait mesh Estimation

Model ReleasesDGX agent

arXiv:2607.23657v1 Announce Type: new Abstract: Articulated portrait mesh estimation is fundamental to 3D understanding, avatar generation, and immersive interaction. Existing approaches primarily rel

HandSCS: Structural Coordinate Space for Animatable Hand Gaussian Splatting

ResearchDGX agent

arXiv:2503.14736v3 Announce Type: replace Abstract: Photorealistic and animatable hand avatars are essential for applications such as AR/VR, gaming, and telepresence. Recent 3D Gaussian Splatting (3DG

Head Avatars with Dynamic Explicit Hair

ResearchDGX agent

arXiv:2607.23861v1 Announce Type: new Abstract: We present DynHair, a novel method for tracking and modeling dynamic hair for human head avatars. From video input, we reconstruct a dynamic head avatar

Hiding in Plain Sight: An Effective Physical Adversarial Patch Attack against Visual-Infrared Fused Face Detection

TutorialsDGX agent

arXiv:2607.23292v1 Announce Type: cross Abstract: Deep learning-based visual-infrared fused face detection models are increasingly deployed across a wide range of applications, yet they remain suscept

HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System

AgentsDGX agent

arXiv:2603.14807v3 Announce Type: replace Abstract: LLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) tasks. However, most zero-shot methods prima

HistoGPA: A Context-Conditioned Gene-Prior Attention Framework for Histology-Based Spatial Gene Expression Prediction

ResearchDGX agent

arXiv:2607.24364v1 Announce Type: new Abstract: Predicting spatial gene expression from routine hematoxylin and eosin (H&E) images provides a practical complement to experimental spatial transcriptomi

Histopathological Spectrum-Guided Prostate Stratification via Segmentation-Assisted Diagnostic Transformer

ResearchDGX agent

arXiv:2607.22703v1 Announce Type: new Abstract: Prostate cancer diagnosis with multiparametric MRI (mpMRI) is commonly based on PI-RADS assessment or binary classification, which suffer from subjectiv

ID-V2V: Identity-Preserving Video Restylization

ApplicationsDGX agent

arXiv:2607.22830v1 Announce Type: new Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance whil

IJCB-AFMFR 2026: Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data

SafetyDGX agent

arXiv:2607.24422v1 Announce Type: new Abstract: This paper presents a summary of the Competition on Adapting Foundation Models for Face Recognition Using Synthetic Training Data (AFMFR), held at the 2

Image Inpainting via Stochastic Dynamics

ResearchDGX agent

arXiv:2607.24140v1 Announce Type: new Abstract: Image inpainting aims to recover missing regions while preserving structural consistency. We propose a non-parametric method without network training ba

Improving Adversarial Robustness of Zero-Shot CLIP with Confidence-Aware Weighting

SafetyDGX agent

arXiv:2510.02913v2 Announce Type: replace Abstract: Vision-language models such as CLIP demonstrate impressive zero-shot generalization but remain highly vulnerable to adversarial attacks. Prior adver

Infrared Imaging Empowered by Artificial Intelligence for Pediatric Skeletal Triage: A Narrative Review and Future Perspectives

SafetyDGX agent

arXiv:2607.24727v1 Announce Type: new Abstract: Background. Pediatric musculoskeletal trauma represents up to 18% of pediatric ED visits, yet diagnosis still depends on ionizing radiography. Cumulativ

InnerGS: Internal Scenes Reconstruction and Segmentation via Factorized 3D Gaussian Splatting

HardwareDGX agent

arXiv:2508.13287v3 Announce Type: replace-cross Abstract: 3D Gaussian Splatting (3DGS) has recently gained popularity for efficient scene rendering by representing scenes as explicit sets of anisotrop

Inter-Reflective Gaussian Splatting for Robust and Efficient Inverse Rendering

ApplicationsDGX agent

arXiv:2607.22780v1 Announce Type: cross Abstract: Faithful inverse rendering requires visibility and indirect radiance to explain secondary illumination and inter-reflection, yet rasterization-oriente

InterOCF: Spatio-Temporal 2D-3D Interaction for Camera-Only 4D Occupancy Forecasting

SafetyDGX agent

arXiv:2607.24431v1 Announce Type: new Abstract: Camera-only 4D occupancy forecasting enables autonomous vehicles to predict future 3D semantic scenes solely from historical multi-view images, which is

Inverse Bayesian Inference for Extracting Lesion Dynamics from Longitudinal Spectral CT

Model ReleasesDGX agent

arXiv:2607.23078v1 Announce Type: new Abstract: Longitudinal medical imaging captures temporal evolution of lesions, yet extracting the underlying dynamical parameters governing this evolution remains

Investigating the Visual Cues of CNNs for Vascular Segmentation: A Case Study in Microscopy and Fundus Imaging

ApplicationsDGX agent

arXiv:2607.23371v1 Announce Type: cross Abstract: Vascular segmentation is a standard procedure for clinical diagnosis, yet the specific visual features determining model decisions remain poorly under

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

AgentsDGX agent

arXiv:2607.23588v1 Announce Type: new Abstract: Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high

JPEG AIC2026: A large-scale dataset for fine-grained assessment of image coding

Model ReleasesDGX agent

arXiv:2607.22783v1 Announce Type: cross Abstract: Recent advances in conventional and learning-based image coding have increased the demand for benchmark datasets that support fine-grained assessment

LabRobFail: A Benchmark for Robotic Failure Analysis in Chemical Self-driving Laboratories

Model ReleasesDGX agent

arXiv:2607.23704v1 Announce Type: cross Abstract: The deployment of embodied agents in self-driving laboratories could accelerate scientific discovery, yet their reliability is constrained by the irre

LanteRn: Latent Visual Structured Reasoning

ResearchDGX agent

arXiv:2603.25629v2 Announce Type: replace Abstract: While language reasoning models excel in many tasks, visual reasoning remains challenging for current large multimodal models (LMMs). As a result, m

← Previous
1…2930313233…209
Next →