AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
7 Jul 2026

Beyond Modality Fusion: Deep Ensembles for Multimodal Classification

Model ReleasesDGX agent

arXiv:2607.05019v1 Announce Type: cross Abstract: In multimodal classification, late-fusion approaches classify concatenated modality-specific features extracted by unimodal neural networks. When moda

Beyond Post-Quantization: Native Hash Learning with a Dedicated HASH Token

ResearchDGX agent

arXiv:2607.03328v1 Announce Type: new Abstract: Efficient large-scale image retrieval requires compact representations that preserve semantic similarity under fast Hamming-space search. Deep hashing i

Beyond Random Sampling: Distribution-Aware Alignment for Semi-Supervised Medical Image Segmentation

SafetyDGX agent

arXiv:2607.04249v1 Announce Type: new Abstract: Precise medical image segmentation is crucial for clinical diagnosis and treatment planning, yet relies heavily on expensive expert annotations. Semi-su


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Beyond Scene Priors: Fine-Grained Traffic Scene Reasoning with Benchmarking and Query-Guided Small-Object Focus

Model ReleasesDGX agent

arXiv:2607.04149v1 Announce Type: new Abstract: In safety-critical traffic scenarios, answering complex questions relies on minute, localized visual cues. However, standard Multimodal Large Language M

Binary Iterative Method for Non-targeted Adversarial Attack

TutorialsDGX agent

arXiv:2607.04145v1 Announce Type: cross Abstract: Adversarial attacks guide and provide additional training and test data for both adversarial training and adversarial robustness validation, and expos

Biomechanics-aware Multi-view Markerless Motion Capture of Dexterous Hand Movements

ResearchDGX agent

arXiv:2607.02796v1 Announce Type: new Abstract: Markerless motion capture (MMC) techniques have been widely beneficial in biomechanical analysis of human movement; however, application to complex moti

BiSLW: Bi-Spectral Latent Watermarking for Generative Diffusion Models

SafetyDGX agent

arXiv:2607.02643v1 Announce Type: new Abstract: Diffusion-based generative models have transformed visual content synthesis, yet they remain vulnerable to unauthorized usage and lack reliable attribut

BrainNormalizer: Anatomy-Informed Pseudo-Healthy Brain Reconstruction from Tumor MRI via Edge-Guided ControlNet

ResearchDGX agent

arXiv:2511.12853v2 Announce Type: replace-cross Abstract: Brain tumors induce complex structural deformations that obscure the patient' s original neuroanatomy, making it difficult to distinguish tumo

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception

ResearchDGX agent

arXiv:2607.03184v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) demonstrate impressive general capabilities, they struggle with fine-grained perception in ultra-high-res

C^3ASD: Multi-Level Consistency-Driven Representation Learning

TutorialsDGX agent

arXiv:2607.03018v1 Announce Type: new Abstract: Active Speaker Detection determines whether a visible person in a video is speaking at each moment. While recent audio-visual fusion methods perform wel

Cancelable Biometric Template Protection Based on Multi-Instance Fusion: A Contralateral Iris Approach

ResearchDGX agent

arXiv:2607.02860v1 Announce Type: new Abstract: Biometric templates are vulnerable to theft if stored without protection. Unlike passwords, a compromised iris cannot be reissued. Although existing can

Causal-RetiGraph: Cross-Cohort Retinal Support and Same-Subject Pathway Analysis for Diabetic Retinopathy

Local AiDGX agent

arXiv:2607.05204v1 Announce Type: new Abstract: Diabetic retinopathy (DR) is a local retinal lesion process and a visible manifestation of systemic microvascular injury. Modern retinal AI can grade im

CCFM: Collision-Constrained Flow Matching for Safety-Critical Scenario Generation

SafetyDGX agent

arXiv:2607.04451v1 Announce Type: new Abstract: Evaluation of autonomous vehicle (AV) planners in safety-critical closed-loop simulation is essential for real-world deployment. However, generating con

CDST: Color Disentangled Style Transfer for Universal Style Reference Customization

ResearchDGX agent

arXiv:2506.13770v2 Announce Type: replace Abstract: We introduce Color Disentangled Style Transfer (CDST), a novel and efficient two-stream style transfer training paradigm which completely isolates c

Cell as Point: One-Stage Framework for Efficient Cell Tracking

ResearchDGX agent

arXiv:2411.14833v4 Announce Type: replace-cross Abstract: Conventional multi-stage cell tracking approaches rely heavily on detection or segmentation in each frame as a prerequisite, requiring substan

CenSynCMB: Centre Maps and Physics-Guided Synthesis for Microbleed Detection

ResearchDGX agent

arXiv:2607.05325v1 Announce Type: new Abstract: Cerebral microbleeds (CMBs) are MRI markers of small vessel disease and the microbleed component of amyloid related imaging abnormalities (ARIA-H), but

ChatImage: Navigating Long-Form LLM Answers through Interactive Images

Model ReleasesDGX agent

arXiv:2607.05290v1 Announce Type: new Abstract: Large Language Models (LLMs) can produce detailed answers to complex queries, but these answers are typically presented as dense linear text, which make

CIPHER: Causal Intervention Pathways for Healthcare Equity and Robustness

ApplicationsDGX agent

arXiv:2607.02596v1 Announce Type: new Abstract: Deep learning models for medical diagnosis frequently exhibit substantial performance disparities across sensitive subgroups (e.g., race, sex), even whe

City landscape in sight: A crowdsourced framework for unlocking urban-scale window view perceptions from real estate imagery

ResearchDGX agent

arXiv:2606.15198v2 Announce Type: replace Abstract: City landscapes viewed through home windows influence quality of life, yet perceptions of actual window views at the urban scale remain understudied

City-Level 3D Surface Reconstruction with Viewpoint Orientation Partitioning and Scene Completion

ResearchDGX agent

arXiv:2607.03771v1 Announce Type: new Abstract: Multi-view 3D surface reconstruction is a longstanding challenge in computer vision. Although recent large-scale reconstruction methods based on 3D Gaus

CL-Anomaly: Layer-Adaptive Mixture-of-Experts with Multimodal Large Language Model for Continual Learning in Anomaly Detection

Model ReleasesDGX agent

arXiv:2607.02930v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) excel in diverse vision tasks, but full-parameter retraining is computationally expensive as real-world knowled

CLABTOOLKIT: An Open-Source Toolkit for Routine Processing, Manipulation, and Visualization of Neuroimaging Data

Model ReleasesDGX agent

arXiv:2607.02638v1 Announce Type: new Abstract: Neuroimaging research requires manipulating heterogeneous data structures, including raw MRI volumes, volumetric parcellations, cortical surface meshes,

Claim-Level Rubric Rewards for Video Caption Reinforcement Learning

SafetyDGX agent

arXiv:2607.05150v1 Announce Type: new Abstract: In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinfor

CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space

ResearchDGX agent

arXiv:2512.08029v3 Announce Type: replace-cross Abstract: Clinical decision-making in oncology requires predicting dynamic disease evolution, a task current static AI predictors cannot perform. While

Classroom Behavior Monitoring with YOLO An Empirical Study in Higher Education Settings

ApplicationsDGX agent

arXiv:2607.02580v1 Announce Type: new Abstract: Classroom behavior monitoring plays a vital role in evaluating student engagement and improving teaching effectiveness. Traditional observation methods

CLEAR: Closed-Loop Reinforcement Learning at Scale for End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2607.02841v1 Announce Type: cross Abstract: End-to-end autonomous driving (E2E-AD) aims to directly map raw sensor information to driving actions. Recently, with the rapid advancement of multi-m

CoDoL: Conditional Domain Prompt Learning for Out-of-Distribution Generalization

SafetyDGX agent

arXiv:2509.15330v2 Announce Type: replace Abstract: Recent advances in pre-training vision-language models (VLMs), e.g., contrastive language-image pre-training (CLIP) methods, have shown great potent

CogRad: A Cognitively-Inspired Multi-Agent Framework for Radiology Report Generation

AgentsDGX agent

arXiv:2607.03853v1 Announce Type: new Abstract: Automated radiology report generation (RRG) can ease radiologist workload, yet most existing systems produce a report in a single forward pass, with no

Combining Discrepancy-Confusion Uncertainty and Calibration Diversity for Active Fine-Grained Image Classification

ResearchDGX agent

arXiv:2509.24181v2 Announce Type: replace Abstract: Active learning (AL) aims to build high-quality labeled datasets by iteratively selecting the most informative samples from an unlabeled pool under

ComplexMimic: Human-Scene Interaction Imitation in Complex 3D Environments

Model ReleasesDGX agent

arXiv:2607.02034v2 Announce Type: replace Abstract: Physics-based Human-Scene Interaction (HSI) imitation learning is crucial for embodied intelligence as it bridges the gap between kinematic 3D motio

Compositional Generalization Requires Linear, Orthogonal Representations in Vision Embedding Models

ResearchDGX agent

arXiv:2602.24264v2 Announce Type: replace Abstract: Compositional generalization, the ability to recognize familiar parts in novel contexts, is a defining property of intelligent systems. Although mod

CompressedVQA-AEV: Full-Reference and No-Reference Quality Assessment Models for Asymmetric Encoded Videos

ResearchDGX agent

arXiv:2607.04606v1 Announce Type: cross Abstract: This report presents our solutions to the QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos, comprising a full-refe

Consistent and Editable: A Balanced Framework for Text-Guided Video Editing

ResearchDGX agent

arXiv:2607.05056v1 Announce Type: new Abstract: Recently, diffusion models have achieved considerable success in the text-guided video editing domain. However, existing works often struggle to balance

Continual Model Merging with Test-Time Adaptation for Whole-Slide Image Analysis

Model ReleasesDGX agent

arXiv:2607.04755v1 Announce Type: new Abstract: Model merging offers a practical alternative to conventional continual learning by integrating independently fine-tuned models without retaining previou

ContiStain: Cross-Domain Relation-Preserving Distillation for Continual Multi-Domain Virtual IHC Staining

Model ReleasesDGX agent

arXiv:2607.03851v1 Announce Type: new Abstract: A unified multiplex virtual staining model enables scalable and non-destructive multiplex analysis from H&E slides while promoting parameter efficiency,

ControlHair: Synergizing Physics Simulator and Video Diffusion for Controllable Dynamic Hair Rendering

ResearchDGX agent

arXiv:2509.21541v3 Announce Type: replace-cross Abstract: Hair simulation and rendering are challenging due to complex strand dynamics, diverse material properties, and intricate light-hair interactio

Conversational Human Audio-visual Talking Dialogue Generation

ResearchDGX agent

arXiv:2607.02799v1 Announce Type: new Abstract: Large-scale dyadic interactive audio-visual dialogue (DIAD) datasets provide fundamental data resources for developing humanoid interactive virtual agen

Coordinate Singularities Break Conformal Coverage for Gaze and Head Pose

ResearchDGX agent

arXiv:2607.02565v1 Announce Type: new Abstract: Conformal prediction provides distribution-free reliability guarantees for vision systems, but these guarantees depend on how prediction errors are meas

CORA: Generalizable coronary artery disease assessment and risk stratification from coronary CT angiography using pathology-centric representation learning

Local AiDGX agent

arXiv:2603.24847v2 Announce Type: replace Abstract: Coronary artery disease, a leading cause of cardiovascular mortality worldwide, can be assessed non-invasively by coronary computed tomography angio

CPR: Chained Perceptual Refinement for Coarse-to-Fine Medical Image Classification

HardwareDGX agent

arXiv:2607.02591v1 Announce Type: new Abstract: High resolution medical images contain fine grained, spatially sparse cues that are critical for diagnosis, yet preserving full resolution incurs substa

Cross-device Collaborative Test-time Adaptation with Zeroth-order Optimization and Model Merging

ApplicationsDGX agent

arXiv:2607.02988v1 Announce Type: new Abstract: Test-time adaptation (TTA) mitigates domain shifts by using incoming test data to update a model on the fly. The majority of TTA methods require resourc

Cross-Modal Fusion of OCT and OCT angiography enface for Improved Diagnostics of Diabetic Retinopathy

ResearchDGX agent

arXiv:2607.03959v1 Announce Type: cross Abstract: Diabetic retinopathy (DR) is a leading cause of vision impairment worldwide, highlighting the need for accurate and accessible screening tools. Optica

CTForensics: A Comprehensive Dataset and Method for AI-Generated CT Image Detection

Model ReleasesDGX agent

arXiv:2603.01878v2 Announce Type: replace Abstract: Recent advances in generative AI have made synthetic Computed Tomography (CT) images increasingly realistic, enabling promising applications in medi

Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations

Local AiDGX agent

arXiv:2506.20294v5 Announce Type: replace Abstract: Diffusion models generate conditional samples by progressively denoising Gaussian noise, yet the denoising trajectory can stall at visually plausibl

CURE: Controllable Unified Image Restoration for Complex Degradations

ResearchDGX agent

arXiv:2607.03044v1 Announce Type: new Abstract: The presence of composite degradations poses a significant challenge, since the underlying corruption factors exhibit complex and interdependent interac

CV-DCLR: Causal-Visual Dynamic Label Refinement for Robust Zero-Shot Learning

SafetyDGX agent

arXiv:2607.02601v1 Announce Type: new Abstract: Zero-Shot Learning (ZSL) facilitates knowledge transfer via shared semantic spaces. However, a critical bottleneck in this paradigm is Semantic Entangle

DC-Motion: Decoupling Structure and Details via Discrete-Continuous Tokens for Human Motion Generation

ResearchDGX agent

arXiv:2606.14721v2 Announce Type: replace-cross Abstract: Text-to-motion generation requires modeling both global action structure and fine-grained motion dynamics from natural language. Existing appr

Deep Learning-Based Characterization of Detonation-Cell Size Distributions in Soot-Foil Records

Model ReleasesDGX agent

arXiv:2607.03764v1 Announce Type: cross Abstract: The geometric size and regularity of detonation cells are key physical parameters for characterizing detonation waves. Traditional manual measurement

Deep Learning for Semen Analysis in Male Infertility: Computer Vision, Multimodal Fusion, and Clinical Translation

ApplicationsDGX agent

arXiv:2607.05311v1 Announce Type: new Abstract: Male infertility contributes substantially to the global infertility burden, and sperm analysis remains central to diagnosis, treatment planning, and as

Defending from GeoLocalization through Adversarial Road Trips

SafetyDGX agent

arXiv:2607.03277v1 Announce Type: new Abstract: Retrieval-based image geolocalization has emerged as a powerful technique for determining the location of a query image by matching it against a large,

Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models

Model ReleasesDGX agent

arXiv:2607.05390v1 Announce Type: cross Abstract: Predicting object dynamics (i.e., world modeling) is a fundamental challenge for robotic manipulation, and modeling deformable objects presents a part

DeGenseGS: Geometrically and Semantically Decoupled Surgical Scene Understanding in 4D Gaussian Splatting

SafetyDGX agent

arXiv:2607.04761v1 Announce Type: new Abstract: Real-time, text-promptable 4D reconstruction is indispensable for autonomous surgical interaction. Severe misalignment between semantic meaning and phys

DGSeg: Dynamic Gating of Semantic-Spatial Guided Predictions for Reasoning Segmentation

TutorialsDGX agent

arXiv:2607.04779v1 Announce Type: new Abstract: Reasoning segmentation aims to predict pixel-wise masks for targets given complex language queries. Existing approaches leverage Multimodal Large Langua

DiCE-CIR: Direct Composition Learning for Efficient Zero-Shot Composed Image Retrieval

SafetyDGX agent

arXiv:2607.04665v1 Announce Type: new Abstract: Zero-shot composed image retrieval (ZS-CIR) aims to retrieve a target image from a multimodal query consisting of a reference image and an edit text des

DICT: Data Injection and Contrastive Trajectory Refinement for Conditional Image Generation with Diffusion Models

SafetyDGX agent

arXiv:2607.03899v1 Announce Type: new Abstract: Diffusion models have become a dominant paradigm for conditional image generation, yet existing approaches generally follow two directions: task-specifi

Diffusion Models are Open-World Affordance Learners: Leveraging Generative Priors for 3D Affordance Learning

Model ReleasesDGX agent

arXiv:2508.01651v2 Announce Type: replace Abstract: 3D affordance grounding aims to understand how diverse objects can be manipulated, making it a cornerstone of embodied interaction. However, prior w

Direct Time-of-Flight Measurement Accuracy Improvement With Perimeter-Gated SPADs

ResearchDGX agent

arXiv:2607.02546v1 Announce Type: cross Abstract: Direct time of flight (dToF) measurements are susceptible to errors because of system-level and circuit-level timing jitters. In addition, device-leve

Displacement Preserving Relational Distillation for Robust Medical Segmentation

SafetyDGX agent

arXiv:2607.04599v1 Announce Type: new Abstract: Accurate 3D medical segmentation is limited by anatomical variability and high computational costs. While knowledge distillation (KD) offers a route for

DistillH-Mamba: A Hypergraph-Mamba-Based Knowledge Distillation Model for Efficient Impact Fall Detection

ResearchDGX agent

arXiv:2607.03156v1 Announce Type: new Abstract: Falls among the elderly represent a significant public health concern due to their prevalence, consequences, and societal burden. While deep learning ha

Distribution Matching Distillation Meets Reinforcement Learning

ResearchDGX agent

arXiv:2511.13649v5 Announce Type: replace Abstract: Distribution Matching Distillation (DMD) facilitates efficient inference by distilling multi-step diffusion models into few-step variants. Concurren

← Previous
1…4849505152…209
Next →