AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
29 Apr 2026

BEVal: A Cross-dataset Evaluation Study of BEV Segmentation Models for Autonomous Driving

AgentsDGX agent

arXiv:2408.16322v4 Announce Type: replace Abstract: Current research in semantic bird's-eye view segmentation for autonomous driving focuses solely on optimizing neural network models using a single d

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models

SafetyDGX agent

arXiv:2604.25072v1 Announce Type: new Abstract: Unified Multimodal Models (uMMs) aim to support both visual understanding and visual generation within a shared representation. However, existing evalua

Beyond Fidelity: Semantic Similarity Assessment in Low-Level Image Processing

ResearchDGX agent

arXiv:2604.25408v1 Announce Type: new Abstract: Low-level image processing has long been evaluated mainly from the perspective of visual fidelity. However, with the rise of deep learning and generativ


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

BifDet: A 3D Bifurcation Detection Dataset for Airway-Tree Modeling

Model ReleasesDGX agent

arXiv:2604.24999v1 Announce Type: new Abstract: Thoracic Computed Tomography (CT) scans offer detailed insights into the intricate branching network of the airway tree, which is essential for understa

C3G: Learning Compact 3D Representations with 2K Gaussians

TutorialsDGX agent

arXiv:2512.04021v2 Announce Type: replace Abstract: Reconstructing and understanding 3D scenes from unposed sparse views in a feed-forward manner remains as a challenging task in 3D computer vision. R

Can We Change the Stroke Size for Easier Diffusion?

ResearchDGX agent

arXiv:2603.26783v2 Announce Type: replace Abstract: Diffusion models can be challenged in the low signal-to-noise regime, where they have to make pixel-level predictions despite the presence of high n

Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval

Model ReleasesDGX agent

arXiv:2604.25273v1 Announce Type: new Abstract: Despite significant progress in Unified Multimodal Retrieval (UMR) powered by Large Multimodal Models (LMMs), existing embedding methods primarily focus

COMPASS: COmpact Multi-channel Prior-map And Scene Signature for Floor-Plan-Based Visual Localization

Local AiDGX agent

arXiv:2604.25388v1 Announce Type: new Abstract: Architectural floor plans are widely available priors which contain not only geometry but also the semantic information of the environment, yet existing

Control Your Queries: Heterogeneous Query Interaction for Camera-Radar Fusion

AgentsDGX agent

arXiv:2604.25574v1 Announce Type: new Abstract: In autonomous driving, camera-radar fusion offers complementary sensing and low deployment cost. Existing methods perform fusion through input mixing, f

CoRE: Concept-Reasoning Expansion for Continual Brain Lesion Segmentation

Model ReleasesDGX agent

arXiv:2604.25376v1 Announce Type: new Abstract: Accurate brain lesion segmentation in MRI is vital for effective clinical diagnosis and treatment planning. Due to high annotation costs and strict data

CRC-SAM: SAM-Based Multi-Modal Segmentation and Quantification of Colorectal Cancer in CT, Colonoscopy, and Histology Images

ResearchDGX agent

arXiv:2604.24793v1 Announce Type: cross Abstract: We present CRC-SAM, a unified framework for colorectal cancer segmentation across colonoscopy, CT, and histopathology images. Unlike prior single-moda

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing

TutorialsDGX agent

arXiv:2604.25477v1 Announce Type: new Abstract: Recent image editing models have achieved strong visual fidelity but often struggle with tasks requiring complex reasoning. To investigate and enhance t

DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework

SafetyDGX agent

arXiv:2506.05199v3 Announce Type: replace Abstract: A core task in embodied intelligence is ego-centric 3D visual grounding. Existing methods typically adopt two-stage, heterogeneous pipelines that pa

DenseScout: Algorithm-System Co-design for Budgeted Tiny Object Selection on Edge Platforms

ResearchDGX agent

arXiv:2604.25300v1 Announce Type: new Abstract: Deploying tiny object perception on edge platforms is challenging because practical systems must satisfy both strict compute budgets and end-to-end late

Detecting Dental Landmarks from Intraoral 3D Scans: the 3DTeethLand challenge

Model ReleasesDGX agent

arXiv:2512.08323v2 Announce Type: replace Abstract: Teeth landmark detection is a key task in modern orthodontics, supporting advanced diagnosis, personalized treatment planning, and effective monitor

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

Model ReleasesDGX agent

arXiv:2603.29844v2 Announce Type: replace-cross Abstract: The development of Vision-Language-Action (VLA) models has been significantly accelerated by pre-trained Vision-Language Models (VLMs). Howeve

Diverse Image Priors for Black-box Data-free Knowledge Distillation

ResearchDGX agent

arXiv:2604.25794v1 Announce Type: cross Abstract: Knowledge distillation (KD) represents a vital mechanism to transfer expertise from complex teacher networks to efficient student models. However, in

DouC: Dual-Branch CLIP for Training-Free Open-Vocabulary Segmentation

Local AiDGX agent

arXiv:2604.24997v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation requires assigning pixel-level semantic labels while supporting an open and unrestricted set of categories. Traini

DualGeo: A Dual-View Framework for Worldwide Image Geo-localization

ResearchDGX agent

arXiv:2604.25533v1 Announce Type: new Abstract: Worldwide image geo-localization aims to infer the geographic location of an image captured anywhere on Earth, spanning street, city, regional, national

Edge-Cloud Collaborative Reconstruction via Structure-Aware Latent Diffusion for Downstream Remote Sensing Perception

ResearchDGX agent

arXiv:2604.25319v1 Announce Type: new Abstract: The exponential surge in high-resolution remote sensing data faces a severe bottleneck in satellite-to-ground transmission. Limited downlink bandwidth f

ESICA: A Scalable Framework for Text-Guided 3D Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2604.24876v1 Announce Type: new Abstract: Text guided 3D medical image segmentation offers a flexible alternative to class based and spatial prompt based models by allowing users to specify regi

Evaluating Computational Pathology Foundation Models for Prostate Cancer Grading under Distribution Shifts

Model ReleasesDGX agent

arXiv:2410.06723v2 Announce Type: replace-cross Abstract: Pathology foundation models (PFMs) have emerged as powerful pretrained encoders for computational pathology, but their robustness under clinic

Exploring Remote Photoplethysmography for Neonatal Pain Detection from Facial Videos

Model ReleasesDGX agent

arXiv:2604.25680v1 Announce Type: new Abstract: Unaddressed pain in neonates can lead to adverse effects, including delayed development and slower weight gain, emphasising the need for more objective

Exploring Time Conditioning in Diffusion Generative Models from Disjoint Noisy Data Manifolds

TutorialsDGX agent

arXiv:2604.25289v1 Announce Type: cross Abstract: Practically, training diffusion models typically requires explicit time conditioning to guide the network through the denoising sampling process. Espe

FCMBench-Video: Benchmarking Document Video Intelligence

Model ReleasesDGX agent

arXiv:2604.25186v1 Announce Type: new Abstract: Document understanding is a critical capability in financial credit review, onboarding, and remote verification, where both decision accuracy and eviden

Generalizable Human Gaussian Splatting via Multi-view Semantic Consistency

Model ReleasesDGX agent

arXiv:2604.25466v1 Announce Type: new Abstract: Recently, generalizable human Gaussian splatting from sparse-view inputs has been actively studied for the photorealistic human rendering. Most existing

GeoSearch: Augmenting Worldwide Geolocalization with Web-Scale Reverse Image Search and Image Matching

ResearchDGX agent

arXiv:2604.25390v1 Announce Type: cross Abstract: Worldwide image geolocalization, which aims to predict the GPS coordinates of any image on Earth, remains challenging due to global visual diversity.

Golden RPG: Confidence-Adaptive Region-Aware Noise for Compositional Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2604.25314v1 Announce Type: new Abstract: Compositional text-to-image (T2I) generation requires a model to honour multiple sub-prompts that describe distinct image regions. Recent work shows tha

GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment

Model ReleasesDGX agent

arXiv:2604.25370v1 Announce Type: new Abstract: The release of GPT-image-2 by OpenAI marks a watershed moment in AI-generated imagery: the boundary between photographic reality and synthetic content h

GramSR: Visual Feature Conditioning for Diffusion-Based Super-Resolution

ApplicationsDGX agent

arXiv:2604.25457v1 Announce Type: new Abstract: Despite recent advances, single-image super-resolution (SR) remains challenging, especially in real-world scenarios with complex degradations. Diffusion

High-Precision Dichotomous Image Segmentation via Depth Integrity-Prior and Fine-Grained Patch Strategy

ResearchDGX agent

arXiv:2503.06100v5 Announce Type: replace Abstract: High-precision dichotomous image segmentation (DIS) is a task of extracting fine-grained objects from high-resolution images. Existing methods trade

HuM-Eval: A Coarse-to-Fine Framework for Human-Centric Video Evaluation

Model ReleasesDGX agent

arXiv:2604.25361v1 Announce Type: new Abstract: Video generation models have developed rapidly in recent years, where generating natural human motion plays a pivotal role. However, accurately evaluati

I-INR: Iterative Implicit Neural Representations

SafetyDGX agent

arXiv:2504.17364v4 Announce Type: replace Abstract: Implicit Neural Representations (INRs) have revolutionized signal processing and computer vision by modeling signals as continuous, differentiable f

IAM: Identity-Aware Human Motion and Shape Joint Generation

ResearchDGX agent

arXiv:2604.25164v1 Announce Type: new Abstract: Recent advances in text-driven human motion generation enable models to synthesize realistic motion sequences from natural language descriptions. Howeve

Image Classification via Random Dilated Convolution with Multi-Branch Feature Extraction and Context Excitation

Model ReleasesDGX agent

arXiv:2604.25188v1 Announce Type: new Abstract: Image classification remains a fundamental yet challenging task in computer vision, particularly when fine-grained feature extraction and background noi

Image Compression with Bubble-Aware Frame Rate Adaptation for Energy-Efficient Video Capsule Endoscopy

ResearchDGX agent

arXiv:2604.25464v1 Announce Type: new Abstract: Video Capsule Endoscopy (VCE) is a promising method for improving the medical examination of the small intestine in the gastrointestinal tract. A key ch

Improving Diversity in Black-box Few-shot Knowledge Distillation

ResearchDGX agent

arXiv:2604.25795v1 Announce Type: new Abstract: Knowledge distillation (KD) is a well-known technique to effectively compress a large network (teacher) to a smaller network (student) with little sacri

Instruction-Evidence Contrastive Dual-Stream Decoding for Grounded Vision-Language Reasoning

ResearchDGX agent

arXiv:2604.25809v1 Announce Type: new Abstract: Vision-Language Models (VLMs) exhibit strong performance in instruction following and open-ended vision-language reasoning, yet they frequently generate

Interactive Episodic Memory with User Feedback

SafetyDGX agent

arXiv:2604.24893v1 Announce Type: new Abstract: In episodic memory with natural language queries (EM-NLQ), a user may ask a question (e.g., 'Where did I place the mug?') that requires searching a long

Is the Modality Gap a Bug or a Feature? A Robustness Perspective

ApplicationsDGX agent

arXiv:2603.29080v2 Announce Type: replace Abstract: Many modern multi-modal models (e.g. CLIP) seek an embedding space in which the two modalities are aligned. Somewhat surprisingly, almost all existi

Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization

SafetyDGX agent

arXiv:2604.24952v1 Announce Type: new Abstract: Human visual preferences are inherently multi-dimensional, encompassing aesthetics, detail fidelity, and semantic alignment. However, existing datasets

Learning Illumination Control in Diffusion Models

ResearchDGX agent

arXiv:2604.24877v1 Announce Type: new Abstract: Controlling illumination in images is essential for photography and visual content creation. While closed-source models have demonstrated impressive ill

Leveraging Previous-Traversal Point Cloud Map Priors for Camera-Based 3D Object Detection and Tracking

Local AiDGX agent

arXiv:2604.25405v1 Announce Type: new Abstract: Camera-based 3D object detection and tracking are central to autonomous driving, yet precise 3D object localization remains fundamentally constrained by

Lightweight Real-Time Rendering Parameter Optimization via XGBoost-Driven Lookup Tables

Model ReleasesDGX agent

arXiv:2604.25178v1 Announce Type: new Abstract: Achieving a desirable balance between rendering quality and real-time performance is a long-standing challenge in modern game and rendering engines, par

M^3-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering

Model ReleasesDGX agent

arXiv:2604.25122v1 Announce Type: new Abstract: We present M^3-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (ML

Magnification-Invariant Image Classification via Domain Generalization and Stable Sparse Embedding Signatures

ResearchDGX agent

arXiv:2604.25817v1 Announce Type: new Abstract: Magnification shift is a major obstacle to robust histopathology classification, because models trained on one imaging scale often generalize poorly to

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition

Model ReleasesDGX agent

arXiv:2512.07348v2 Announce Type: replace Abstract: In controllable image generation, synthesizing coherent and consistent images from multiple reference inputs, i.e., Multi-Image Composition (MICo),

MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding

Model ReleasesDGX agent

arXiv:2512.17492v2 Announce Type: replace Abstract: Geo-spatial analysis of our world benefits from a multimodal approach, as every single geographic location can be described in numerous ways (images

MTPano: Multi-Task Panoramic Scene Understanding via Label-Free Integration of Dense Prediction Priors

ResearchDGX agent

arXiv:2602.05330v2 Announce Type: replace Abstract: Comprehensive panoramic scene understanding is critical for immersive applications, yet it remains challenging due to the scarcity of high-resolutio

Multimodal Contextualized Support for Enhancing Video Retrieval System

ResearchDGX agent

arXiv:2412.07584v2 Announce Type: replace Abstract: Current video retrieval systems, especially those used in competitions, primarily focus on querying individual keyframes or images rather than encod

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation

ResearchDGX agent

arXiv:2604.25819v1 Announce Type: new Abstract: In this work, we propose Mutual Forcing, a framework for fast autoregressive audio-video generation with long-horizon audio-video synchronization. Our a

Natural Image Classification via Quasi-Cyclic Graph Ensembles and Random-Bond Ising Models at the Nishimori Temperature

ResearchDGX agent

arXiv:2508.18717v3 Announce Type: replace-cross Abstract: Modern multi-class image classification uses high-dimensional CNN features that incur large memory and computational costs and obscure the dat

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

Model ReleasesDGX agent

arXiv:2604.24954v1 Announce Type: cross Abstract: We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, i

NimbleReg: A light-weight deep-learning framework for diffeomorphic image registration

SafetyDGX agent

arXiv:2503.07768v2 Announce Type: replace Abstract: This paper presents NimbleReg, a light-weight deep-learning (DL) framework for diffeomorphic image registration leveraging surface representation of

No Pedestrian Left Behind: Real-Time Detection and Tracking of Vulnerable Road Users for Adaptive Traffic Signal Control

SafetyDGX agent

arXiv:2604.25887v1 Announce Type: new Abstract: Current pedestrian crossing signals operate on fixed timing without adjustment to pedestrian behavior, which can leave vulnerable road users (VRUs) such

Novel 3D Binary Indexed Tree for Volume Computation of 3D Reconstructed Models from Volumetric Data

ResearchDGX agent

arXiv:2412.10441v2 Announce Type: replace-cross Abstract: In the burgeoning field of medical imaging, precise computation of 3D volume holds a significant importance for subsequent qualitative analysi

OmniAlpha: Aligning Transparency-Aware Generation via Multi-Task Unified Reinforcement Learning

Local AiDGX agent

arXiv:2511.20211v2 Announce Type: replace Abstract: Transparency-aware generation requires modeling not only RGB appearance but also alpha-based opacity and cross-layer composition, which are essentia

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding

Local AiDGX agent

arXiv:2604.25276v1 Announce Type: new Abstract: Video Temporal Grounding (VTG), the task of localizing video segments from text queries, struggles in open-world settings due to limited dataset scale a

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations

Model ReleasesDGX agent

arXiv:2604.25102v1 Announce Type: new Abstract: Typographic prompt injection exploits vision language models' (VLMs) ability to read text rendered in images, posing a growing threat as VLMs power auto

OneThinker: All-in-one Reasoning Model for Image and Video

Model ReleasesDGX agent

arXiv:2512.03043v3 Announce Type: replace Abstract: Reinforcement learning (RL) has recently achieved remarkable success in eliciting visual reasoning within Multimodal Large Language Models (MLLMs).

← Previous
1…166167168169170…209
Next →