AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
15 Jul 2026

GenDiff: A Dose and Anatomy Aware Diffusion Model with Structural Prior Refinement for Low-Dose CT Reconstruction and Generalization

TutorialsDGX agent

arXiv:2607.11941v1 Announce Type: new Abstract: Computed tomography (CT) is a critical imaging modality for clinical diagnosis, but reducing radiation dose inevitably introduces severe noise and struc

Generalization and Memorization in Rectified Flow

ResearchDGX agent

arXiv:2603.13421v2 Announce Type: replace-cross Abstract: Generative models based on the Flow Matching objective, particularly Rectified Flow, have emerged as a dominant paradigm for efficient, high-f

GeoSEAN: Explainable Country-Level Image Geolocation for ASEAN Regions

ResearchDGX agent

arXiv:2607.12284v1 Announce Type: new Abstract: Image geolocation aims to infer the geographic origin of an image from visual content alone. However, this task remains challenging in regions where cou


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

How to Realize Recursively Self-Improving Agents and Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture

Model ReleasesDGX agent

arXiv:2607.12254v1 Announce Type: new Abstract: Large language model (LLM) agents can increasingly plan, use tools, maintain memory, and execute long-horizon tasks. These advances motivate two linked

Hy-Embodied-VLM-1.0: Efficient Physical-World Agents

Model ReleasesDGX agent

arXiv:2607.12894v1 Announce Type: new Abstract: Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, a

Illuminant-Adaptive 3D Lookup Tables for Camera Color Correction

ResearchDGX agent

arXiv:2607.11681v2 Announce Type: replace Abstract: Color correction is a key component of camera image signal processing (ISP) pipelines, encompassing illuminant discounting and colorimetric mapping

Image Matching Filtering and Refinement by Planes and Beyond

Model ReleasesDGX agent

arXiv:2411.09484v5 Announce Type: replace Abstract: This paper provides a consistent and extensive evaluation of state-of-the-art filtering and refinement methods on common image matching pipelines. U

Implicit 4D Gaussian Splatting for Fast Motion with Large Inter-Frame Displacements

ResearchDGX agent

arXiv:2607.12362v1 Announce Type: new Abstract: Recent 4D Gaussian Splatting (4DGS) methods often fail under fast motion with large inter-frame displacements, where Gaussian attributes are poorly lear

Improved Robustness from Biologically Inspired Sparse Contrast Representations

ResearchDGX agent

arXiv:2509.24863v2 Announce Type: replace Abstract: Deep neural networks surpass humans on many vision benchmarks, yet remain far less robust to distribution shifts such as illumination and weather ch

Inhibited Self-Attention: Sharpening Focus in Vision Transformers

ResearchDGX agent

arXiv:2607.12881v1 Announce Type: new Abstract: Vision Transformers (ViTs) have demonstrated remarkable performance in computer vision tasks. However, their self-attention mechanism often diffuses foc

Instance-Enriched Semantic Maps for Visual Language Navigation

AgentsDGX agent

arXiv:2607.12630v1 Announce Type: cross Abstract: Visual Language Navigation (VLN) aims to enable an embodied agent to navigate complex environments by following natural language instructions. Recent

Kernel PCA for Out-of-Distribution Detection: Non-Linear Kernel Selection and Approximation

ApplicationsDGX agent

arXiv:2505.15284v2 Announce Type: replace-cross Abstract: Out-of-Distribution (OoD) detection is vital for the reliability of deep neural networks, the key of which lies in effectively characterizing

Label-Decoupled Style Augmentation for Domain Generalization in Multi-Label Remote Sensing Scene Classification

Model ReleasesDGX agent

arXiv:2607.12704v1 Announce Type: new Abstract: Multi-label classification assigns several co-occurring labels to each aerial scene, yet deployed models often encounter data distributions different fr

LARAD: Layout-Aware Road Anomaly Detection via Spatial-Logic Reasoning

AgentsDGX agent

arXiv:2607.12858v1 Announce Type: new Abstract: Accurate open-world obstacle detection is critical for autonomous driving. Current anomaly segmentation methods suffer from a fundamental blind spot: th

Learning from Complementary Ultrasound Representations for Liver Disease Classification

ResearchDGX agent

arXiv:2607.12062v1 Announce Type: new Abstract: Differentiating non-alcoholic steatohepatitis (NASH) from non-alcoholic fatty liver disease (NAFLD) using ultrasound remains challenging due to subtle t

LEGO-SLAM: Language-Embedded Gaussian Optimization SLAM

Local AiDGX agent

arXiv:2511.16144v2 Announce Type: replace Abstract: Recent advances in 3D Gaussian Splatting (3DGS) have enabled Simultaneous Localization and Mapping (SLAM) systems to build photorealistic maps. Howe

Lesion Segmentation in Moderate to Severe Traumatic Brain Injury: An nnU-Net Based Approach with Adaptive Normalization in the AIMS-TBI 2025 Challenge

ResearchDGX agent

arXiv:2607.12684v1 Announce Type: new Abstract: The segmentation of lesions in Moderate to Severe Traumatic Brain Injury (msTBI) from T1-weighted MRI presents a significant clinical challenge due to t

Let RGB Be the Language of Vision

ResearchDGX agent

arXiv:2607.12450v1 Announce Type: new Abstract: This work introduces a unified formulation for vision models, where diverse forms of visual information beyond natural images, such as masks, depth maps

Leveraging Prior Knowledge of Diffusion Model for Person Search

Local AiDGX agent

arXiv:2510.01841v2 Announce Type: replace Abstract: Person search aims to jointly perform person detection and re-identification by localizing and identifying a query person within a gallery of uncrop

LVMark: Robust Watermark for Latent Video Diffusion Models

ResearchDGX agent

arXiv:2412.09122v4 Announce Type: replace Abstract: Rapid advancements in video diffusion models have enabled the creation of realistic videos, raising concerns about unauthorized use and driving the

M2I2HA: Multi-modal Object Detection Based on Intra- and Inter-Modal Hypergraph Attention

SafetyDGX agent

arXiv:2601.14776v3 Announce Type: replace Abstract: Recent advances in multi-modal detection have significantly improved detection accuracy in challenging environments (e.g., low light, overexposure).

MAGE: Color-Invariant and Spatial Knowledge Distillation for Gastric Neoplasm Classification

Local AiDGX agent

arXiv:2607.12663v1 Announce Type: new Abstract: Accurate differentiation between gastric adenoma and carcinoma during endoscopy is critical for clinical decision-making. Yet, this task is highly chall

MambaPSA: A Mamba-based Replacement for C2PSA in YOLO26

ResearchDGX agent

arXiv:2607.12681v1 Announce Type: new Abstract: State space models (SSMs), notably Mamba, have recently emerged as efficient alternatives to self-attention with linear computational complexity. We inv

MBTI: A Multi-Branch Efficient Fine-Tuning Framework for Hyperspectral Image Classification with Foundation Models

TutorialsDGX agent

arXiv:2607.12782v1 Announce Type: new Abstract: Hyperspectral foundation models learn transferable spectral-spatial representations from large-scale unlabeled data. They provide an effective paradigm

Medical Image Segmentation based on Deep Active Contour and Mean Curvature Loss Function

ResearchDGX agent

arXiv:2607.12586v1 Announce Type: cross Abstract: Medical image segmentation is a crucial task in the field of clinical analysis and applications. Though deep learning techniques recently play a cruci

MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors

ResearchDGX agent

arXiv:2607.12000v1 Announce Type: new Abstract: Current visual generation models are capable of producing high-quality content, yet they lack a coherent perception of the spatial structure. Existing g

Metric-Guided Synthetic Image Data Rendering for Deep Learning compatible with Agentic AI

Model ReleasesDGX agent

arXiv:2607.12874v1 Announce Type: new Abstract: Deep learning computer vision for scientific applications requires collecting and annotating large datasets in a laborious, expensive and error-prone pr

MobileSAM2: Lightweight Segment Anything for Spatial Intelligence

ResearchDGX agent

arXiv:2607.12297v1 Announce Type: new Abstract: The recent large video foundation model, SAM2, enables segment anything in both images and videos, serving as a powerful base model for various applicat

More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization

Local AiDGX agent

arXiv:2607.12429v1 Announce Type: new Abstract: Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same sc

MQAdapter: Multi-Modal Quantum Adapter for Coarse-to-Fine VLM Fine-tuning

Model ReleasesDGX agent

arXiv:2607.12418v1 Announce Type: new Abstract: Large-scale Vision-Language Models have demonstrated impressive transfer learning capabilities across a wide range of tasks. For few-shot classification

NEEDL-Bench: Dataset for Swiss Needle Cast and Stomata Detection in Microscopy Images

Model ReleasesDGX agent

arXiv:2607.12076v1 Announce Type: new Abstract: We present NEEDL-Bench, a microscopy detection benchmark for Swiss Needle Cast (SNC), a fungal disease of Douglas-fir trees. Douglas-fir is a keystone s

Open-KNEAD: Knowledge-grounded Nutrition Estimation via Agentic Decomposition

Local AiDGX agent

arXiv:2607.12911v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly used for dietary assessment from meal images, where retrieval-augmented grounding was shown to

Overview of Cross-Component In-loop Filters in Video Coding Standards

ResearchDGX agent

arXiv:2607.12186v1 Announce Type: new Abstract: In-loop filters have been comprehensively explored during the development of video coding standards due to their remarkable noise-reduction capability.

Physically Aware Radiomics Without Interpolation: Disentangling Voxel Geometry and Signal Modification in CT and MRI

ResearchDGX agent

arXiv:2607.12399v1 Announce Type: new Abstract: Objective: Radiomic texture features are usually computed in voxel-index neighborhoods, implicitly assuming isotropic spatial relationships. In anisotro

Pixel-Level Pavement Distress Assessment Using Instance Segmentation

Local AiDGX agent

arXiv:2605.26095v2 Announce Type: replace Abstract: Automated pavement distress assessment requires more than image-level classification or coarse bounding box detection, demanding precise localizatio

Point Tracking in Surgery--The 2025 Surgical Tattoos in Infrared Challenge (STIRC2025)

AgentsDGX agent

arXiv:2607.12939v1 Announce Type: new Abstract: Point tracking in surgery is crucial to enable applications in downstream tasks such as segmentation, 3D reconstruction, virtual tissue landmarking, aut

PoseAlign: Sculpting Pose-Consistent Meshes via Text-Guided Deformation

SafetyDGX agent

arXiv:2607.10560v2 Announce Type: replace-cross Abstract: Mesh deformation, the process of altering the vertex positions of a 3D mesh while preserving its topological structure, is a cornerstone of co

ProtoPointNet: Prototype-Based Interpretable Classification of 3D Dental Point Clouds with Verifiable Spatial Activations

ResearchDGX agent

arXiv:2607.12335v1 Announce Type: new Abstract: Prototype-based networks provide inherently interpretable classification by linking predictions to learned exemplars, but their use in 3D point clouds a

Rank-1 Identity Consensus Predicts Gallery Enrollment in 1:N Face Matching More Accurately than Score Thresholding

Model ReleasesDGX agent

arXiv:2607.12903v1 Announce Type: new Abstract: In operational 1:N face identification, a crucial question arises for each probe: is this person enrolled in the gallery or not? The stakes are high and

RealSkin: Spatio-Spectral Partial Neural Adjoint Maps for Image-to-3D Attribute Transfer

Model ReleasesDGX agent

arXiv:2607.12495v1 Announce Type: new Abstract: Creating photorealistic 3D assets requires bridging the appearance gap between real-world observations and synthetic models. A promising approach is to

ReflectVLN: Training Vision-Language Navigation Agents with Reflective Reasoning

AgentsDGX agent

arXiv:2607.12680v1 Announce Type: new Abstract: Existing vision-language navigation methods often couple a VLM with waypoint decoders to produce multi-step action plans, but they typically lack an exp

RegHead: Non-Humanoid Head Blendshapes via Feed-Forward Registration

Local AiDGX agent

arXiv:2607.12206v1 Announce Type: new Abstract: We present RegHead, a framework for constructing semantic blendshape sets for animatable non-humanoid head avatars. With a fixed expression vocabulary,

RFMSR: Residual Flow Matching for Image Super-Resolution

ResearchDGX agent

arXiv:2607.12753v1 Announce Type: new Abstract: Image super-resolution (ISR) has witnessed remarkable progress with diffusion models and flow matching. The dominant text-to-image (T2I) based approache

Rough Path Signature-Guided Geometry Augmentation for Few-Shot Industrial Surface Defect Detection

ResearchDGX agent

arXiv:2607.12245v1 Announce Type: new Abstract: Few-shot industrial defect detection remains difficult for standard supervised detectors, which achieve poor performance on boundary-dominated industria

Same Compression Principle, Different Geometry: Rate-Distortion Signatures Dissociate Biological and Artificial Visual Systems

ResearchDGX agent

arXiv:2603.01568v2 Announce Type: replace-cross Abstract: Efficient coding theory predicts that biological perceptual systems compress sensory input optimally under resource constraints, with the syst

Sat2RealCity: Geometry-Aware and Appearance-Controllable 3D Urban Generation from Satellite Imagery

SafetyDGX agent

arXiv:2511.11470v3 Announce Type: replace Abstract: 3D urban generation from satellite imagery is an important task for scalable digital twins and real-world simulation environments. Existing approach

SeamGen: Artist-Aligned UV Seam Generation via Graph Flow Matching

Local AiDGX agent

arXiv:2607.12379v1 Announce Type: new Abstract: UV seam placement is a critical yet labor-intensive step in 3D content creation, requiring artists to balance chart shape, seam concealment, and alignme

Seeing Globally, Refining Locally: Global Visual Guidance and Local Ultrasound Cues for Robust Freehand 3-D Ultrasound Reconstruction

Local AiDGX agent

arXiv:2607.12398v1 Announce Type: new Abstract: Freehand 3-D ultrasound (US) imaging has attracted increasing attention owing to its intuitive volumetric visualization, ease of use, and low cost. Howe

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

Model ReleasesDGX agent

arXiv:2607.12477v1 Announce Type: new Abstract: Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenar

Semantic-Edge Response Decoding of SAM3 for Zero-Shot Crack Segmentation

ResearchDGX agent

arXiv:2607.12292v1 Announce Type: new Abstract: Crack segmentation is essential for infrastructure inspection and structural health assessment, but existing high-performance methods typically require

Spherical-GOF: Geometry-Aware Panoramic Gaussian Opacity Fields for 3D Scene Reconstruction

Model ReleasesDGX agent

arXiv:2603.08503v2 Announce Type: replace Abstract: Omnidirectional images are increasingly used in robotics and vision due to their wide field of view. However, extending 3D Gaussian Splatting (3DGS)

SpikeDS: Dual Sparsity Spikformer for Perineural Invasion Prediction in 3D MRI

ResearchDGX agent

arXiv:2607.11986v1 Announce Type: new Abstract: Perineural invasion (PNI) is associated with poor prognosis in cholangiocarcinoma (CCA). However, its detection from 3D MRI remains challenging due to t

Statistical Non-linear Reconstruction Loss for Image Anomaly Detection

ResearchDGX agent

arXiv:2607.12866v1 Announce Type: new Abstract: Reconstruction-based methods are a cornerstone of unsupervised image anomaly detection, but they remain vulnerable to outlier leakage, where standard me

Steering Diffusion Models via Class-Contrastive Influence for Few-Shot Medical Classification

ResearchDGX agent

arXiv:2607.12464v1 Announce Type: new Abstract: When labeled data are scarce, off-the-shelf diffusion models can augment training sets for few-shot medical image classification, but not all generated

Structure-Semantic Co-optimized Latent Diffusion Model for Fast Visual Anagram Synthesis

SafetyDGX agent

arXiv:2606.16241v3 Announce Type: replace Abstract: Visual anagram is an intriguing form of art creation wherein a single image presents different conceptual interpretations under transformations such

SymbOmni: Evolving Agentic Omni Models via Symbolic Concept Learning

AgentsDGX agent

arXiv:2607.12042v1 Announce Type: new Abstract: Visual generation is increasingly ubiquitous in diverse domains, from text-to-image/video synthesis to multimodal interactive creation. Yet prevailing m

TerraLogic: A Benchmark for Hierarchical Geospatial Reasoning in Earth Observation

Model ReleasesDGX agent

arXiv:2607.12497v1 Announce Type: new Abstract: Beyond perception, reasoning is essential in remote sensing for advanced interpretation, inference, and decision-making. Recent advances in large langua

The GEST-Engine: From Event Graphs to Synthetic Video. A Full Technical Report

AgentsDGX agent

arXiv:2607.12231v1 Announce Type: new Abstract: We present the GEST-Engine, a complete system that goes from natural-language text to fully-annotated multi-actor video. At its core is an explicit worl

The Seriality Gap in Video Diffusion Models

ResearchDGX agent

arXiv:2607.13031v1 Announce Type: cross Abstract: When one ball strikes another, then another, video models should predict the consequences of each bounce. In controlled experiments on multi-ball hard

The TopCoW Challenge -- Topology-Aware Circle of Willis Segmentation for CT and MR Angiography

Model ReleasesDGX agent

arXiv:2312.17670v5 Announce Type: replace Abstract: The Circle of Willis (CoW) is an important network of arteries connecting major circulations of the brain. Its vascular architecture is believed to

← Previous
1…4142434445…209
Next →