AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Local Ai

Contrastive Joint-Embedding Prediction for Representation Learning in Structural MRI

DGX agent

arXiv:2607.11962v1 Announce Type: new Abstract: Self-supervised learning offers a compelling approach for medical imaging, where labeled data are scarce and acquisition costs are high. We present COJE

local-aiarxiv-cs-cv
15 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Controllable Generation of Diverse Dermatological Imagery for Fair and Efficient Malignancy Classification

DGX agent

arXiv:2607.12987v1 Announce Type: new Abstract: Accurate dermatological diagnosis naturally necessitates equitable performance across diverse populations, yet a systematic lack of expertly annotated i

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models

DGX agent

arXiv:2607.12786v1 Announce Type: new Abstract: Cross-image comparative reasoning remains challenging for vision-language models (VLMs), especially when correct prediction requires fine-grained attrib

model-releasesarxiv-cs-cv
15 Jul 2026
Research

CRC-HGD: A Histopathological Image Dataset for Grading Colorectal Cancer

DGX agent

arXiv:2607.12750v1 Announce Type: new Abstract: Colorectal cancer (CRC) is the third most common cancer worldwide and the second leading cause of cancer-related deaths globally, with approximately 1,9

researcharxiv-cs-cv
15 Jul 2026
Safety

Data Safety: Synthetic Data Quality Analysis Using CIFAKE Dataset

DGX agent

arXiv:2607.12165v1 Announce Type: new Abstract: Recently, the societal implementation of high-performance image classification models has expanded rapidly. While these models require vast amounts of t

safetyarxiv-cs-cv
15 Jul 2026
Model Releases

Decouple and Reason: Anatomically Guided Two-Stage Voxel-Level Grounding of Free-Text Findings in 3D Chest CT

DGX agent

arXiv:2607.12602v1 Announce Type: new Abstract: Automatic voxel-level grounding of free-text findings in 3D chest Computed Tomography (CT) is critical for clinical interpretability. However, this task

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

DeGuNet: Depth-Guided Ultra-Compact Backbones for Efficient LiDAR-Camera 3D Detection

DGX agent

arXiv:2607.12419v1 Announce Type: new Abstract: In autonomous driving perception, the fusion of LiDAR and camera modalities has become the dominant paradigm for 3D object detection. However, current m

model-releasesarxiv-cs-cv
15 Jul 2026
Applications

DermDepth: Toward Monocular Metric Scale 3D Reconstruction Models for Dermatology

DGX agent

arXiv:2607.13010v1 Announce Type: new Abstract: Dermatological practice routinely involves measuring and tracking lesion size, morphology and texture, as critical components of wound or skin cancer sc

applicationsarxiv-cs-cv
15 Jul 2026
Model Releases

DiTailed: Ensuring Visual Object Consistency in Text-Image-to-Image Flow Matching Models

DGX agent

arXiv:2607.12539v1 Announce Type: new Abstract: Despite remarkable progress in text-guided image editing, generative models frequently fail to preserve visual object consistency, defined as the preser

model-releasesarxiv-cs-cv
15 Jul 2026
Tutorials

DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery

DGX agent

arXiv:2607.12319v1 Announce Type: new Abstract: As vision-language models (VLMs) are increasingly deployed in geospatial question answering and visual scene understanding, improving their spatial cogn

tutorialsarxiv-cs-cv
15 Jul 2026
Research

Domain-Incremental Remote Sensing Change Detection via Difference-Guided Adaptation and Frequency-Decoupled Distillation

DGX agent

arXiv:2607.12934v1 Announce Type: new Abstract: Remote sensing change detection (RSCD) models are prone to catastrophic forgetting when incrementally adapted to new domains. Existing domain-incrementa

researcharxiv-cs-cv
15 Jul 2026
Research

DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs

DGX agent

arXiv:2607.12503v1 Announce Type: new Abstract: 4D spatio-temporal reasoning, jointly modeling 3D spatial structure and temporal evolution, is essential for understanding dynamic worlds and enabling e

researcharxiv-cs-cv
15 Jul 2026
Model Releases

Edge-Aware Thermal Infrared UAV Swarm Tracking

DGX agent

arXiv:2607.12544v1 Announce Type: new Abstract: Thermal infrared (TIR) imaging is essential for UAV swarm operations in visually degraded environments. However, tracking tiny UAVs remains challenging

model-releasesarxiv-cs-cv
15 Jul 2026
Agents

EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval

DGX agent

arXiv:2607.12764v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) has emerged as a critical paradigm for grounding Multimodal Large Language Models (MLLMs) in external knowledge. Re

agentsarxiv-cs-cv
15 Jul 2026
Research

Exact and Calibrated Diffusion Reconstruction for Digital Breast Tomosynthesis

DGX agent

arXiv:2607.12937v1 Announce Type: cross Abstract: Limited-angle digital breast tomosynthesis (DBT) reconstructs a volume from a few low-dose projections over a narrow arc. At a representative nine-vie

researcharxiv-cs-cv
15 Jul 2026
Safety

ExtraGS: Enhancing Endoscopic View Extrapolation via Diffusion-Guided 3D Gaussian Splatting

DGX agent

arXiv:2607.12785v1 Announce Type: new Abstract: Robot-assisted minimally invasive surgery (MIS) critically depends on reliable endoscopic perception for navigation and safety. However, conventional en

safetyarxiv-cs-cv
15 Jul 2026
Research

Fast and Accurate Image Restoration and Generation with Rank Enhanced Linear Attention

DGX agent

arXiv:2505.16157v2 Announce Type: replace Abstract: Transformer-based models have made remarkable progress in image restoration (IR) tasks. However, the quadratic complexity of self-attention in Trans

researcharxiv-cs-cv
15 Jul 2026
Agents

Filtering-out poor-quality images for data preparation

DGX agent

arXiv:2607.12352v1 Announce Type: new Abstract: Filtering noise is a fundamental part of data preparation that enhances image quality for applications such as object segmentation, detection, and recog

agentsarxiv-cs-cv
15 Jul 2026
Safety

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

DGX agent

arXiv:2607.13017v1 Announce Type: cross Abstract: World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveragin

safetyarxiv-cs-cv
15 Jul 2026
Local Ai

FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry

DGX agent

arXiv:2607.11588v2 Announce Type: replace Abstract: We present FoundationGeo, a two-stage framework that explicitly bridges relative and metric prediction via spatial calibration and principled data d

local-aiarxiv-cs-cv
15 Jul 2026
Research

From Hindsight to Foresight: Self-Encouraged Hindsight Distillation for Knowledge-based Visual Question Answering

DGX agent

arXiv:2511.11132v4 Announce Type: replace Abstract: Knowledge-based Visual Question Answering (KBVQA) necessitates external knowledge incorporation beyond cross-modal understanding. Existing KBVQA met

researcharxiv-cs-cv
15 Jul 2026
Model Releases

GAINS: Gaussian-based Inverse Rendering from Sparse Multi-View Captures

DGX agent

arXiv:2512.09925v2 Announce Type: replace Abstract: Recent advances in Gaussian Splatting-based inverse rendering extend Gaussian primitives with shading parameters and physically grounded light trans

model-releasesarxiv-cs-cv
15 Jul 2026
Research

Gaussian Mixture Modeling for Event-Aware Visual Allocation in Long Video Understanding

DGX agent

arXiv:2607.12557v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) face significant challenges in long video understanding due to the excessive computational cost and information los

researcharxiv-cs-cv
15 Jul 2026
Model Releases

GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents

DGX agent

arXiv:2607.11192v2 Announce Type: replace Abstract: A large share of day-to-day work in professional domains happens inside PDF files: benefits packets, leases, datasheets, clinical guidelines, constr

model-releasesarxiv-cs-cv
15 Jul 2026
Tutorials

GenDiff: A Dose and Anatomy Aware Diffusion Model with Structural Prior Refinement for Low-Dose CT Reconstruction and Generalization

DGX agent

arXiv:2607.11941v1 Announce Type: new Abstract: Computed tomography (CT) is a critical imaging modality for clinical diagnosis, but reducing radiation dose inevitably introduces severe noise and struc

tutorialsarxiv-cs-cv
15 Jul 2026
Research

Generalization and Memorization in Rectified Flow

DGX agent

arXiv:2603.13421v2 Announce Type: replace-cross Abstract: Generative models based on the Flow Matching objective, particularly Rectified Flow, have emerged as a dominant paradigm for efficient, high-f

researcharxiv-cs-cv
15 Jul 2026
Research

GeoSEAN: Explainable Country-Level Image Geolocation for ASEAN Regions

DGX agent

arXiv:2607.12284v1 Announce Type: new Abstract: Image geolocation aims to infer the geographic origin of an image from visual content alone. However, this task remains challenging in regions where cou

researcharxiv-cs-cv
15 Jul 2026
Model Releases

How to Realize Recursively Self-Improving Agents and Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture

DGX agent

arXiv:2607.12254v1 Announce Type: new Abstract: Large language model (LLM) agents can increasingly plan, use tools, maintain memory, and execute long-horizon tasks. These advances motivate two linked

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

Hy-Embodied-VLM-1.0: Efficient Physical-World Agents

DGX agent

arXiv:2607.12894v1 Announce Type: new Abstract: Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, a

model-releasesarxiv-cs-cv
15 Jul 2026
Research

Illuminant-Adaptive 3D Lookup Tables for Camera Color Correction

DGX agent

arXiv:2607.11681v2 Announce Type: replace Abstract: Color correction is a key component of camera image signal processing (ISP) pipelines, encompassing illuminant discounting and colorimetric mapping

researcharxiv-cs-cv
15 Jul 2026
Model Releases

Image Matching Filtering and Refinement by Planes and Beyond

DGX agent

arXiv:2411.09484v5 Announce Type: replace Abstract: This paper provides a consistent and extensive evaluation of state-of-the-art filtering and refinement methods on common image matching pipelines. U

model-releasesarxiv-cs-cv
15 Jul 2026
Research

Implicit 4D Gaussian Splatting for Fast Motion with Large Inter-Frame Displacements

DGX agent

arXiv:2607.12362v1 Announce Type: new Abstract: Recent 4D Gaussian Splatting (4DGS) methods often fail under fast motion with large inter-frame displacements, where Gaussian attributes are poorly lear

researcharxiv-cs-cv
15 Jul 2026
Research

Improved Robustness from Biologically Inspired Sparse Contrast Representations

DGX agent

arXiv:2509.24863v2 Announce Type: replace Abstract: Deep neural networks surpass humans on many vision benchmarks, yet remain far less robust to distribution shifts such as illumination and weather ch

researcharxiv-cs-cv
15 Jul 2026
Research

Inhibited Self-Attention: Sharpening Focus in Vision Transformers

DGX agent

arXiv:2607.12881v1 Announce Type: new Abstract: Vision Transformers (ViTs) have demonstrated remarkable performance in computer vision tasks. However, their self-attention mechanism often diffuses foc

researcharxiv-cs-cv
15 Jul 2026
Agents

Instance-Enriched Semantic Maps for Visual Language Navigation

DGX agent

arXiv:2607.12630v1 Announce Type: cross Abstract: Visual Language Navigation (VLN) aims to enable an embodied agent to navigate complex environments by following natural language instructions. Recent

agentsarxiv-cs-cv
15 Jul 2026
Applications

Kernel PCA for Out-of-Distribution Detection: Non-Linear Kernel Selection and Approximation

DGX agent

arXiv:2505.15284v2 Announce Type: replace-cross Abstract: Out-of-Distribution (OoD) detection is vital for the reliability of deep neural networks, the key of which lies in effectively characterizing

applicationsarxiv-cs-cv
15 Jul 2026
Model Releases

Label-Decoupled Style Augmentation for Domain Generalization in Multi-Label Remote Sensing Scene Classification

DGX agent

arXiv:2607.12704v1 Announce Type: new Abstract: Multi-label classification assigns several co-occurring labels to each aerial scene, yet deployed models often encounter data distributions different fr

model-releasesarxiv-cs-cv
15 Jul 2026
Agents

LARAD: Layout-Aware Road Anomaly Detection via Spatial-Logic Reasoning

DGX agent

arXiv:2607.12858v1 Announce Type: new Abstract: Accurate open-world obstacle detection is critical for autonomous driving. Current anomaly segmentation methods suffer from a fundamental blind spot: th

agentsarxiv-cs-cv
15 Jul 2026
Research

Learning from Complementary Ultrasound Representations for Liver Disease Classification

DGX agent

arXiv:2607.12062v1 Announce Type: new Abstract: Differentiating non-alcoholic steatohepatitis (NASH) from non-alcoholic fatty liver disease (NAFLD) using ultrasound remains challenging due to subtle t

researcharxiv-cs-cv
15 Jul 2026
Local Ai

LEGO-SLAM: Language-Embedded Gaussian Optimization SLAM

DGX agent

arXiv:2511.16144v2 Announce Type: replace Abstract: Recent advances in 3D Gaussian Splatting (3DGS) have enabled Simultaneous Localization and Mapping (SLAM) systems to build photorealistic maps. Howe

local-aiarxiv-cs-cv
15 Jul 2026
Research

Lesion Segmentation in Moderate to Severe Traumatic Brain Injury: An nnU-Net Based Approach with Adaptive Normalization in the AIMS-TBI 2025 Challenge

DGX agent

arXiv:2607.12684v1 Announce Type: new Abstract: The segmentation of lesions in Moderate to Severe Traumatic Brain Injury (msTBI) from T1-weighted MRI presents a significant clinical challenge due to t

researcharxiv-cs-cv
15 Jul 2026
Research

Let RGB Be the Language of Vision

DGX agent

arXiv:2607.12450v1 Announce Type: new Abstract: This work introduces a unified formulation for vision models, where diverse forms of visual information beyond natural images, such as masks, depth maps

researcharxiv-cs-cv
15 Jul 2026
Local Ai

Leveraging Prior Knowledge of Diffusion Model for Person Search

DGX agent

arXiv:2510.01841v2 Announce Type: replace Abstract: Person search aims to jointly perform person detection and re-identification by localizing and identifying a query person within a gallery of uncrop

local-aiarxiv-cs-cv
15 Jul 2026
Research

LVMark: Robust Watermark for Latent Video Diffusion Models

DGX agent

arXiv:2412.09122v4 Announce Type: replace Abstract: Rapid advancements in video diffusion models have enabled the creation of realistic videos, raising concerns about unauthorized use and driving the

researcharxiv-cs-cv
15 Jul 2026
Safety

M2I2HA: Multi-modal Object Detection Based on Intra- and Inter-Modal Hypergraph Attention

DGX agent

arXiv:2601.14776v3 Announce Type: replace Abstract: Recent advances in multi-modal detection have significantly improved detection accuracy in challenging environments (e.g., low light, overexposure).

safetyarxiv-cs-cv
15 Jul 2026
Local Ai

MAGE: Color-Invariant and Spatial Knowledge Distillation for Gastric Neoplasm Classification

DGX agent

arXiv:2607.12663v1 Announce Type: new Abstract: Accurate differentiation between gastric adenoma and carcinoma during endoscopy is critical for clinical decision-making. Yet, this task is highly chall

local-aiarxiv-cs-cv
15 Jul 2026
Research

MambaPSA: A Mamba-based Replacement for C2PSA in YOLO26

DGX agent

arXiv:2607.12681v1 Announce Type: new Abstract: State space models (SSMs), notably Mamba, have recently emerged as efficient alternatives to self-attention with linear computational complexity. We inv

researcharxiv-cs-cv
15 Jul 2026
Tutorials

MBTI: A Multi-Branch Efficient Fine-Tuning Framework for Hyperspectral Image Classification with Foundation Models

DGX agent

arXiv:2607.12782v1 Announce Type: new Abstract: Hyperspectral foundation models learn transferable spectral-spatial representations from large-scale unlabeled data. They provide an effective paradigm

tutorialsarxiv-cs-cv
15 Jul 2026
← Previous
1…5152535455…261
Next →