AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
23 Apr 2026

Amodal SAM: A Unified Amodal Segmentation Framework with Generalization

ApplicationsDGX agent

arXiv:2604.20748v1 Announce Type: new Abstract: Amodal segmentation is a challenging task that aims to predict the complete geometric shape of objects, including their occluded regions. Although exist

Automated Description Generation of Cytologic Findings for Lung Cytological Images Using a Pretrained Vision Model and Dual Text Decoders: Preliminary Study

ResearchDGX agent

arXiv:2403.18151v2 Announce Type: replace-cross Abstract: Objective: Cytology plays a crucial role in lung cancer diagnosis. Pulmonary cytology involves cell morphological characterization in the spec

Benchmarking ResNet for Short-Term Hypoglycemia Classification with DiaData

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2511.02849v2 Announce Type: replace-cross Abstract: Individualized therapy is driven forward by medical data analysis, which provides insight into the patient's context. In particular, for Type

Bio-inspired Color Constancy: From Gray Anchoring Theory to Gray Pixel Methods

ResearchDGX agent

arXiv:2604.20243v1 Announce Type: new Abstract: Color constancy is a fundamental ability of many biological visual systems and a crucial step in computer imaging systems. Bio-inspired modeling offers

Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens

TutorialsDGX agent

arXiv:2604.19954v1 Announce Type: new Abstract: Current text-to-image models struggle to provide precise camera control using natural language alone. In this work, we present a framework for precise c

CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs

Model ReleasesDGX agent

arXiv:2604.20460v1 Announce Type: new Abstract: Safety-critical traffic reasoning requires contrastive consistency: models must detect true hazards when an accident occurs, and reliably reject plausib

CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation

SafetyDGX agent

arXiv:2603.25383v3 Announce Type: replace Abstract: CLIP aligns image and text embeddings via contrastive learning and demonstrates strong zero-shot generalization. Its large-scale architecture requir

ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image Retrieval

Model ReleasesDGX agent

arXiv:2604.20358v1 Announce Type: new Abstract: The Composed Image Retrieval (CIR) task provides a flexible retrieval paradigm via a reference image and modification text, but it heavily relies on exp

Confidence-Based Mesh Extraction from 3D Gaussians

ResearchDGX agent

arXiv:2603.24725v2 Announce Type: replace Abstract: Recently, 3D Gaussian Splatting (3DGS) greatly accelerated mesh extraction from posed images due to its explicit representation and fast software ra

CoRe: Joint Optimization with Contrastive Learning for Medical Image Registration

SafetyDGX agent

arXiv:2603.23694v2 Announce Type: replace Abstract: Medical image registration is a fundamental task in medical image analysis, enabling the alignment of images from different modalities or time point

CrackForward: Context-Aware Severity Stage Crack Synthesis for Data Augmentation

ResearchDGX agent

arXiv:2604.19941v1 Announce Type: new Abstract: Reliable crack detection and segmentation are vital for structural health monitoring, yet the scarcity of well-annotated data constitutes a major challe

CXR-LanIC: Language-Grounded Interpretable Classifier for Chest X-Ray Diagnosis

ResearchDGX agent

arXiv:2510.21464v2 Announce Type: replace Abstract: Deep learning models have achieved remarkable accuracy in chest X-ray diagnosis, yet their widespread clinical adoption remains limited by the black

DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation

AgentsDGX agent

arXiv:2604.20841v1 Announce Type: new Abstract: Recent advances in video generative models enable the synthesis of realistic human-object interaction videos across a wide range of scenarios and object

Diagnosing Urban Street Vitality via a Visual-Semantic and Spatiotemporal Framework for Street-Level Economics

ResearchDGX agent

arXiv:2604.19798v1 Announce Type: cross Abstract: Micro-scale street-level economic assessment is fundamental for precision spatial resource allocation. While Street View Imagery (SVI) advances urban

DynamicRad: Content-Adaptive Sparse Attention for Long Video Diffusion

Local AiDGX agent

arXiv:2604.20470v1 Announce Type: new Abstract: Leveraging the natural spatiotemporal energy decay in video diffusion offers a path to efficiency, yet relying solely on rigid static masks risks losing

Efficient INT8 Single-Image Super-Resolution via Deployment-Aware Quantization and Teacher-Guided Training

ApplicationsDGX agent

arXiv:2604.20291v1 Announce Type: new Abstract: Efficient single-image super-resolution (SISR) requires balancing reconstruction fidelity, model compactness, and robustness under low-bit deployment, w

Energy-Based Open-Set Active Learning for Object Classification

ApplicationsDGX agent

arXiv:2604.20083v1 Announce Type: cross Abstract: Active learning (AL) has emerged as a crucial methodology for minimizing labeling costs in deep learning by selecting the most valuable samples from a

Excretion Detection in Pigsties Using Convolutional and Transformerbased Deep Neural Networks

ResearchDGX agent

arXiv:2412.00256v3 Announce Type: replace Abstract: Animal excretions in form of urine puddles and feces are a significant source of emissions in livestock farming. Automated detection of soiled floor

Exploring High-Order Self-Similarity for Video Understanding

TutorialsDGX agent

arXiv:2604.20760v1 Announce Type: new Abstract: Space-time self-similarity (STSS), which captures visual correspondences across frames, provides an effective way to represent temporal dynamics for vid

Exploring Spatial Intelligence from a Generative Perspective

Model ReleasesDGX agent

arXiv:2604.20570v1 Announce Type: new Abstract: Spatial intelligence is essential for multimodal large language models, yet current benchmarks largely assess it only from an understanding perspective.

FA-Seg: A Fast and Accurate Diffusion-Based Method for Open-Vocabulary Segmentation

Local AiDGX agent

arXiv:2506.23323v5 Announce Type: replace Abstract: Open-vocabulary semantic segmentation (OVSS) aims to segment objects from arbitrary text categories without requiring densely annotated datasets. Al

Fast Amortized Fitting of Scientific Signals Across Time and Ensembles via Transferable Neural Fields

ResearchDGX agent

arXiv:2604.19979v1 Announce Type: cross Abstract: Neural fields, also known as implicit neural representations (INRs), offer a powerful framework for modeling continuous geometry, but their effectiven

Fast-then-Fine: A Two-Stage Framework with Multi-Granular Representation for Cross-Modal Retrieval in Remote Sensing

Model ReleasesDGX agent

arXiv:2604.20429v1 Announce Type: new Abstract: Remote sensing (RS) image-text retrieval plays a critical role in understanding massive RS imagery. However, the dense multi-object distribution and com

FluSplat: Sparse-View 3D Editing without Test-Time Optimization

SafetyDGX agent

arXiv:2604.20038v1 Announce Type: new Abstract: Recent advances in text-guided image editing and 3D Gaussian Splatting (3DGS) have enabled high-quality 3D scene manipulation. However, existing pipelin

Fourier Series Coder: A Novel Perspective on Angle Boundary Discontinuity Problem for Oriented Object Detection

ResearchDGX agent

arXiv:2604.20281v1 Announce Type: new Abstract: With the rapid advancement of intelligent driving and remote sensing, oriented object detection has gained widespread attention. However, achieving high

From Competition to Synergy: Unlocking Reinforcement Learning for Subject-Driven Image Generation

ResearchDGX agent

arXiv:2510.18263v2 Announce Type: replace-cross Abstract: Subject-driven image generation models face a fundamental trade-off between identity preservation (fidelity) and prompt adherence (editability

From Diffusion to Flow: Efficient Motion Generation in MotionGPT3

ResearchDGX agent

arXiv:2603.26747v2 Announce Type: replace Abstract: Recent text-driven motion generation methods span both discrete token-based approaches and continuous-latent formulations. MotionGPT3 exemplifies th

From Ideal to Real: Stable Video Object Removal under Imperfect Conditions

TutorialsDGX agent

arXiv:2603.09283v2 Announce Type: replace Abstract: Removing objects from videos remains difficult in the presence of real-world imperfections such as shadows, abrupt motion, and defective masks. Exis

From Image to Music Language: A Two-Stage Structure Decoding Approach for Complex Polyphonic OMR

ResearchDGX agent

arXiv:2604.20522v1 Announce Type: cross Abstract: We propose a new approach for the second stage of a practical two-stage Optical Music Recognition (OMR) pipeline. Given symbol and event candidates fr

FurnSet: Exploiting Repeats for 3D Scene Reconstruction

SafetyDGX agent

arXiv:2604.20093v1 Announce Type: new Abstract: Single-view 3D scene reconstruction involves inferring both object geometry and spatial layout. Existing methods typically reconstruct objects independe

Gaussians on a Diet: High-Quality Memory-Bounded 3D Gaussian Splatting Training

HardwareDGX agent

arXiv:2604.20046v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has revolutionized novel view synthesis with high-quality rendering through continuous aggregations of millions of 3D Gauss

Generative Prior-Guided Neural Interface Reconstruction for 3D Electrical Impedance Tomography

ResearchDGX agent

arXiv:2505.16487v3 Announce Type: replace-cross Abstract: Reconstructing complex 3D interfaces from indirect measurements remains a grand challenge in scientific computing, particularly for ill-posed

GeoRect4D: Geometry-Compatible Generative Rectification for Dynamic Sparse-View 3D Reconstruction

ResearchDGX agent

arXiv:2604.20784v1 Announce Type: new Abstract: Reconstructing dynamic 3D scenes from sparse multi-view videos is highly ill-posed, often leading to geometric collapse, trajectory drift, and floating

GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers

ResearchDGX agent

arXiv:2604.20715v1 Announce Type: new Abstract: Relighting a person from a single photo is an attractive but ill-posed task, as a 2D image ambiguously entangles 3D geometry, intrinsic appearance, and

Global Offshore Wind Infrastructure: Deployment and Operational Dynamics from Dense Sentinel-1 Time Series

Model ReleasesDGX agent

arXiv:2604.20822v1 Announce Type: new Abstract: The offshore wind energy sector is expanding rapidly, increasing the need for independent, high-temporal-resolution monitoring of infrastructure deploym

GSCompleter: A Distillation-Free Plugin for Metric-Aware 3D Gaussian Splatting Completion in Seconds

ResearchDGX agent

arXiv:2604.20155v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) has revolutionized real-time rendering, its performance degrades significantly under sparse-view extrapolation, manif

Hallucination Early Detection in Diffusion Models

ResearchDGX agent

arXiv:2604.20354v1 Announce Type: new Abstract: Text-to-Image generation has seen significant advancements in output realism with the advent of diffusion models. However, diffusion models encounter di

Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders

ApplicationsDGX agent

arXiv:2508.18236v4 Announce Type: replace Abstract: The rapid development of generative AI has transformed content creation, communication, and human development. However, this technology raises profo

HumanScore: Benchmarking Human Motions in Generated Videos

ResearchDGX agent

arXiv:2604.20157v1 Announce Type: new Abstract: Recent advances in model architectures, compute, and data scale have driven rapid progress in video generation, producing increasingly realistic content

Hybrid Latent Reasoning with Decoupled Policy Optimization

SafetyDGX agent

arXiv:2604.20328v1 Announce Type: new Abstract: Chain-of-Thought (CoT) reasoning significantly elevates the complex problem-solving capabilities of multimodal large language models (MLLMs). However, a

i-WiViG: Interpretable Window Vision GNN

SafetyDGX agent

arXiv:2503.08321v2 Announce Type: replace Abstract: Vision graph neural networks have emerged as a popular approach for modeling the global and spatial context for image recognition. However, a signif

Improving Facial Emotion Recognition through Dataset Merging and Balanced Training Strategies

ResearchDGX agent

arXiv:2604.20307v1 Announce Type: new Abstract: In this paper, a deep learning framework is proposed for automatic facial emotion based on deep convolutional networks. In order to increase the general

Integrated AI Nodule Detection and Diagnosis for Lung Cancer Screening Beyond Size and Growth-Based Standards Compared with Radiologists and Leading Models

ResearchDGX agent

arXiv:2512.00281v2 Announce Type: replace Abstract: Early detection of malignant lung nodules remains limited by reliance on size- and growth-based screening criteria, which can delay diagnosis. We pr

Investigation of cardinality classification for bacterial colony counting using explainable artificial intelligence

ResearchDGX agent

arXiv:2604.20026v1 Announce Type: new Abstract: Automatic bacterial colony counting is a highly sought-after technology in modern biological laboratories because it eliminates manual counting effort.

KD-Judge: A Knowledge-Driven Automated Judge Framework for Functional Fitness Movements on Edge Devices

ResearchDGX agent

arXiv:2604.19834v1 Announce Type: new Abstract: Functional fitness movements are widely used in training, competition, and health-oriented exercise programs, yet consistently enforcing repetition (rep

Learn2Synth: Learning Optimal Data Synthesis Using Hypergradients for Brain Image Segmentation

ApplicationsDGX agent

arXiv:2411.16719v4 Announce Type: replace Abstract: Domain randomization through synthesis is a powerful strategy to train networks that are unbiased with respect to the domain of the input images. Ra

Learning Spatial-Temporal Coherent Correlations for Speech-Preserving Facial Expression Manipulation

Local AiDGX agent

arXiv:2604.20226v1 Announce Type: new Abstract: Speech-preserving facial expression manipulation (SPFEM) aims to modify facial emotions while meticulously maintaining the mouth animation associated wi

Learning to count small and clustered objects with application to bacterial colonies

SafetyDGX agent

arXiv:2604.20030v1 Announce Type: new Abstract: Automated bacterial colony counting from images is an important technique to obtain data required for the development of vaccines and antibiotics. Howev

LEXIS: LatEnt ProXimal Interaction Signatures for 3D HOI from an Image

ResearchDGX agent

arXiv:2604.20800v1 Announce Type: new Abstract: Reconstructing 3D Human-Object Interaction from an RGB image is essential for perceptive systems. Yet, this remains challenging as it requires capturing

Lifecycle-Aware Federated Continual Learning in Mobile Autonomous Systems

AgentsDGX agent

arXiv:2604.20745v1 Announce Type: cross Abstract: Federated continual learning (FCL) allows distributed autonomous fleets to adapt collaboratively to evolving terrain types across extended mission lif

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model

ResearchDGX agent

arXiv:2604.20796v1 Announce Type: new Abstract: We present LLaDA2.0-Uni, a unified discrete diffusion large language model (dLLM) that supports multimodal understanding and generation within a nativel

Lucky High Dynamic Range Smartphone Imaging

ResearchDGX agent

arXiv:2604.19976v1 Announce Type: new Abstract: While the human eye can perceive an impressive twenty stops of dynamic range, smartphone camera sensors remain limited to about twelve stops despite dec

MAPRPose: Mask-Aware Proposal and Amodal Refinement for Multi-Object 6D Pose Estimation

Model ReleasesDGX agent

arXiv:2604.20650v1 Announce Type: new Abstract: 6D object pose estimation in cluttered scenes remains challenging due to severe occlusion and sensor noise. We propose MAPRPose, a two-stage framework t

Maximum Likelihood Reconstruction for Multi-Look Digital Holography with Markov-Modeled Speckle Correlation

ResearchDGX agent

arXiv:2604.20154v1 Announce Type: cross Abstract: Multi-look acquisition is a widely used strategy for reducing speckle noise in coherent imaging systems such as digital holography. By acquiring multi

MD-Face: MoE-Enhanced Label-Free Disentangled Representation for Interactive Facial Attribute Editing

TutorialsDGX agent

arXiv:2604.20317v1 Announce Type: new Abstract: GAN-based facial attribute editing is widely used in virtual avatars and social media but often suffers from attribute entanglement, where modifying one

Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation

Model ReleasesDGX agent

arXiv:2604.20366v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) exhibit powerful generative capabilities but frequently produce hallucinations that compromise output reliability.

MLG-Stereo: ViT Based Stereo Matching with Multi-Stage Local-Global Enhancement

Model ReleasesDGX agent

arXiv:2604.20393v1 Announce Type: new Abstract: With the development of deep learning, ViT-based stereo matching methods have made significant progress due to their remarkable robustness and zero-shot

MSLAU-Net: A Hybrid CNN-Transformer Network for Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2505.18823v2 Announce Type: replace Abstract: Accurate medical image segmentation allows for the precise delineation of anatomical structures and pathological regions, which is essential for tre

Object Referring-Guided Scanpath Prediction with Perception-Enhanced Vision-Language Models

ResearchDGX agent

arXiv:2604.20361v1 Announce Type: new Abstract: Object Referring-guided Scanpath Prediction (ORSP) aims to predict the human attention scanpath when they search for a specific target object in a visua

On the Impact of Face Segmentation-Based Background Removal on Recognition and Morphing Attack Detection

ResearchDGX agent

arXiv:2604.20585v1 Announce Type: new Abstract: This study investigates the impact of face image background correction through segmentation on face recognition and morphing attack detection performanc

← Previous
1…175176177178179…209
Next →