AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Tutorials

Did Models Learn Sufficiently? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation

DGX agent

arXiv:2511.12100v2 Announce Type: replace Abstract: In current visual model training, models often rely on only limited sufficient causes for their predictions, which makes them sensitive to distribut

tutorialsarxiv-cs-cv
25 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography

DGX agent

arXiv:2606.25546v1 Announce Type: new Abstract: Vision-language pre-training (VLP) holds great promise for general-purpose medical AI by leveraging radiology reports as rich textual supervision, yet e

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms

DGX agent

arXiv:2606.25066v1 Announce Type: cross Abstract: Visual search has been one of the most productive paradigms in the study of visual attention: the way reaction time scales with the number of items di

model-releasesarxiv-cs-cv
25 Jun 2026
Research

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation

DGX agent

arXiv:2606.26058v1 Announce Type: new Abstract: Open domain subject-driven text-to-video (S2V) generation has drawn significant interest in academia and industry. Open domain S2V mainly involves two s

researcharxiv-cs-cv
25 Jun 2026
Safety

DRM: Diffusion-based Reward Model With Step-wise Guidance

DGX agent

arXiv:2605.25661v2 Announce Type: replace Abstract: Current mainstream methods of aligning diffusion models with human preferences typically employ VLM-based reward models. However, these reward model

safetyarxiv-cs-cv
25 Jun 2026
Agents

DSP-SLAM++: A Unified Framework for Multi-Class, High-Fidelity Object SLAM in the Wild

DGX agent

arXiv:2606.25953v1 Announce Type: cross Abstract: Existing object-aware SLAM systems force a trade-off between real-time performance, multi-class support, and the generation of high-fidelity, semantic

agentsarxiv-cs-cv
25 Jun 2026
Safety

Dual Agreement Consistency Learning for Semi-Supervised Fetal Ultrasound Segmentation

DGX agent

arXiv:2606.25254v1 Announce Type: cross Abstract: Maternal-fetal US is the primary imaging modality for monitoring fetal development, yet accurate automated segmentation remains challenging due to the

safetyarxiv-cs-cv
25 Jun 2026
Model Releases

Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs

DGX agent

arXiv:2606.25758v1 Announce Type: new Abstract: While test-time adaptation (TTA) empowers vision-language models to adapt without costly retraining, it remains highly vulnerable to out-of-distribution

model-releasesarxiv-cs-cv
25 Jun 2026
Research

EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis

DGX agent

arXiv:2606.25465v1 Announce Type: new Abstract: While image stylization has been studied extensively, video stylization remains a critical and largely unsolved challenge in the field of intelligent co

researcharxiv-cs-cv
25 Jun 2026
Applications

Edges Before Embeddings: A Confidence-Aware Blur Gate for Vision-Language Pipelines

DGX agent

arXiv:2606.25838v1 Announce Type: new Abstract: Production vision pipelines silently degrade on blurry input, wasting compute on downstream OCR, retrieval, and vision-language model (VLM) calls that c

applicationsarxiv-cs-cv
25 Jun 2026
Research

Efficient Cross-Scale Invertible Hiding Network with Spatial-Frequency Collaboration and Non-Invertible Mechanism

DGX agent

arXiv:2606.25547v1 Announce Type: new Abstract: Image hiding aims to conceal image-level messages within cover images at the same resolution. Invertible neural networks (INN)-based image hiding has em

researcharxiv-cs-cv
25 Jun 2026
Model Releases

Efficient Real-World Dehazing via Physics-Inspired Global-Local Decoupling

DGX agent

arXiv:2606.25732v1 Announce Type: new Abstract: Real-world single image dehazing is highly ill-posed due to spatially and spectrally varying scattering, while practical deployment demands lightweight

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

Efficient Remote Sensing Instance Segmentation with Linear-Time State Space Distilled Visual Foundation Models

DGX agent

arXiv:2606.25324v1 Announce Type: new Abstract: The computational complexity of Transformers scales quadratically with the number of tokens, which significantly constrains the efficiency of vision mod

model-releasesarxiv-cs-cv
25 Jun 2026
Research

Energy-Efficient CNN Acceleration with MSDF Digit-Serial Arithmetic on FPGA

DGX agent

arXiv:2606.25562v1 Announce Type: cross Abstract: This paper presents an energy-efficient hardware acceleration of the convolutional layers in the U-Net architecture for image segmentation, implemente

researcharxiv-cs-cv
25 Jun 2026
Model Releases

Enhancing Brain MRI Anomaly Detection and Reasoning with ROI Rethink and Synthetic Data

DGX agent

arXiv:2606.25894v1 Announce Type: new Abstract: Medical vision-language models typically generate diagnoses through single-pass inference without indicating which image regions support their conclusio

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

Enhancing Pathological VLMs with Cross-scale Reasoning

DGX agent

arXiv:2606.17412v3 Announce Type: replace Abstract: Pathological images are inherently multi-scale, requiring pathologists to integrate evidence from global tissue architecture at low magnification to

model-releasesarxiv-cs-cv
25 Jun 2026
Research

Entropy-Controlled Flow Matching

DGX agent

arXiv:2602.22265v2 Announce Type: replace-cross Abstract: Modern vision generators transport a base distribution to data through time-indexed measures, implemented as deterministic flows (ODEs) or sto

researcharxiv-cs-cv
25 Jun 2026
Local Ai

ESMStereo: Enhanced ShuffleMixer Disparity Upsampling for Real-Time and Accurate Stereo Matching

DGX agent

arXiv:2506.21091v2 Announce Type: replace Abstract: Stereo matching has become an increasingly important component of modern autonomous systems. Developing deep learning-based stereo matching models t

local-aiarxiv-cs-cv
25 Jun 2026
Tutorials

ESTANet: Efficient Online Error Detection in Procedural Videos via Prediction Inconsistency

DGX agent

arXiv:2606.25317v1 Announce Type: new Abstract: An efficient and accurate system for detecting errors in procedural tasks is crucial for supporting human needs in daily life, as it can provide instant

tutorialsarxiv-cs-cv
25 Jun 2026
Safety

Evaluation Protocols and Validation for Cameras in Indoor Healthcare Monitoring

DGX agent

arXiv:2606.25284v1 Announce Type: new Abstract: Camera-based monitoring systems are increasingly adopted in healthcare settings for the continuous assessment of patient movement and activities. Howeve

safetyarxiv-cs-cv
25 Jun 2026
Model Releases

Evidential Perfusion Physics-Informed Neural Networks with Residual Uncertainty Quantification

DGX agent

arXiv:2603.09359v2 Announce Type: replace Abstract: Physics-informed neural networks (PINNs) have shown promise in addressing the ill-posed deconvolution problem in computed tomography perfusion (CTP)

model-releasesarxiv-cs-cv
25 Jun 2026
Research

Expresso-AI: Explainable Video-Based Deep Learning Models for Depression Diagnosis

DGX agent

arXiv:2606.25606v1 Announce Type: new Abstract: Given the widespread prevalence of depression and its consequential impact on individuals and society, it is crucial to obtain objective measures for ea

researcharxiv-cs-cv
25 Jun 2026
Model Releases

Falcon: Functional Assembly and Language for Compositional Reasoning in X-ray

DGX agent

arXiv:2606.25701v1 Announce Type: new Abstract: Conventional vision-language models are largely object-centric, focusing on detecting and describing individual entities. In safety-critical X-ray bagga

model-releasesarxiv-cs-cv
25 Jun 2026
Applications

fARfetch: Enabling Collocated AR-HRC in Large Visually Diverse Environments with VLM-Driven AR Content Adaptation

DGX agent

arXiv:2606.25162v1 Announce Type: cross Abstract: Augmented Reality (AR) can improve collocated human-robot collaboration by making robot state and intent visible and enabling intuitive control, yet l

applicationsarxiv-cs-cv
25 Jun 2026
Local Ai

FedReLa: Imbalanced Federated Learning via Re-Labeling

DGX agent

arXiv:2606.26037v1 Announce Type: cross Abstract: Federated learning has emerged as the foremost approach for decentralized model training with privacy preservation. The global class imbalance and cro

local-aiarxiv-cs-cv
25 Jun 2026
Research

FeVOS: Foresight Expression Video Object Segmentation

DGX agent

arXiv:2606.25585v1 Announce Type: new Abstract: Existing Referring Video Object Segmentation tasks focus on referring expressions describing events, actions or appearances of relevant objects within t

researcharxiv-cs-cv
25 Jun 2026
Model Releases

FlowID : Enhancing Forensic Identification with Latent Flow-Matching Models

DGX agent

arXiv:2603.29591v2 Announce Type: replace Abstract: Every day, many people die under violent circumstances, whether from crimes, war, migration, or climate disasters. Medico-legal and law enforcement

model-releasesarxiv-cs-cv
25 Jun 2026
Safety

Follow Your Track: Precise Skeleton Animation Controlled by 3D Trajectories

DGX agent

arXiv:2606.25344v1 Announce Type: new Abstract: 4D generation aims to animate 3D objects with realistic motion, holding great promise for applications. Existing methods typically decouple 3D asset gen

safetyarxiv-cs-cv
25 Jun 2026
Model Releases

FreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling

DGX agent

arXiv:2606.25079v1 Announce Type: new Abstract: Visual storytelling aims to generate image sequences that are both aligned with narrative prompts and consistent in character appearance across images.

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

From Sparse and Imperfect 2D Anchors to Consistent 3D Gaussian Street Scenes: Support-Aware Appearance

DGX agent

arXiv:2606.26007v1 Announce Type: new Abstract: Image priors can synthesize target conditions for 3D Gaussian street scenes, but independently edited views do not define a coherent 3D target. Direct f

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

FunPiQ: A New Benchmark for Pixel-Level Quality Assessment in Fundus Images

DGX agent

arXiv:2606.25915v1 Announce Type: new Abstract: Color fundus photography (CFP) is the most common ophthalmic imaging modality for large-scale screening. However, it is highly susceptible to degradatio

model-releasesarxiv-cs-cv
25 Jun 2026
Research

Gastroendoscopy View Synthesis: A New Real Dataset and Evaluation

DGX agent

arXiv:2606.25427v1 Announce Type: new Abstract: Novel view synthesis (NVS) is an active research topic in computer vision, owing to the success of neural radiance field (NeRF) and 3D Gaussian splattin

researcharxiv-cs-cv
25 Jun 2026
Safety

Geometry-Anchored Transport Framework for Exemplar-Free Class-Incremental Learning

DGX agent

arXiv:2606.25347v1 Announce Type: cross Abstract: Exemplar-free class-incremental learning (EFCIL) requires stable decision boundaries within a shifting feature space. While maintaining class-conditio

safetyarxiv-cs-cv
25 Jun 2026
Research

GeoRanker: Distance-Aware Ranking for Worldwide Image Geolocalization

DGX agent

arXiv:2505.13731v4 Announce Type: replace Abstract: Worldwide image geolocalization-the task of predicting GPS coordinates from images taken anywhere on Earth-poses a fundamental challenge due to the

researcharxiv-cs-cv
25 Jun 2026
Research

Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs

DGX agent

arXiv:2606.25842v1 Announce Type: new Abstract: Existing multi-modal large language models (MLLMs) face significant challenges in processing long video sequences due to strict input token limitations.

researcharxiv-cs-cv
25 Jun 2026
Model Releases

GroundSet: A Cadastral-Grounded Dataset for Spatial Understanding with Vector Data

DGX agent

arXiv:2603.14609v2 Announce Type: replace Abstract: Precise spatial understanding in Earth Observation is essential for translating raw aerial imagery into actionable insights for critical application

model-releasesarxiv-cs-cv
25 Jun 2026
Tutorials

H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks

DGX agent

arXiv:2606.25578v1 Announce Type: new Abstract: Hairstyle transfer has practical applications such as virtual try-on, yet remains challenging when the source and reference exhibit large head-pose disc

tutorialsarxiv-cs-cv
25 Jun 2026
Research

Heterogeneous and Adept Snapshot Distillation for 3D Semantic Segmentation

DGX agent

arXiv:2606.25278v1 Announce Type: new Abstract: Multi-modal fusion and multi-model ensembling are prevalent in enhancing the performance of 3D semantic segmentation. Despite the impressive performance

researcharxiv-cs-cv
25 Jun 2026
Model Releases

HG-Bench: A Benchmark for Multi-Page Handwritten Answer-Region Grounding in Automated Homework Assessment

DGX agent

arXiv:2606.25491v1 Announce Type: new Abstract: Automated homework assessment depends not only on recognizing student answers, but also on accurately locating where each answer and each intermediate r

model-releasesarxiv-cs-cv
25 Jun 2026
Applications

HiFiVe: High-Fidelity Vehicle Generation Leveraging Auto-Regressive 2D Generative Priors

DGX agent

arXiv:2606.25300v1 Announce Type: new Abstract: Existing 3D vehicle generation methods often suffer from low geometric fidelity and blurry textures, hindering their downstream applications. While rece

applicationsarxiv-cs-cv
25 Jun 2026
Local Ai

HiT-JEPA: A Hierarchical Self-supervised Trajectory Embedding Framework for Similarity Computation

DGX agent

arXiv:2507.00028v2 Announce Type: replace-cross Abstract: The representation of urban trajectory data plays a critical role in effectively analyzing spatial movement patterns. Despite considerable pro

local-aiarxiv-cs-cv
25 Jun 2026
Applications

Homomorphic Encryptions for Privacy Preserving Vision

DGX agent

arXiv:2606.25216v1 Announce Type: cross Abstract: Legal requirements might prevent organizations from sharing sensitive data like medical or financial details of consumers which prevents them from lev

applicationsarxiv-cs-cv
25 Jun 2026
Research

Hybrid deep learning-based phase diversity method for wavefront reconstruction

DGX agent

arXiv:2606.25855v1 Announce Type: cross Abstract: The efficiency of high-power laser systems is limited by wavefront distortions in the beam, particularly non-common path aberrations, which reduce the

researcharxiv-cs-cv
25 Jun 2026
Local Ai

Hypergraph Normal World Models for Logical Visual Anomaly Detection

DGX agent

arXiv:2606.25368v1 Announce Type: new Abstract: Visual anomaly detection is often deployed with only normal training images. Most one-class detectors map test patches or features to a normal reference

local-aiarxiv-cs-cv
25 Jun 2026
Research

ILV: Iterative Latent Volumes for Fast and Accurate Sparse-View CT Reconstruction

DGX agent

arXiv:2603.14915v2 Announce Type: replace Abstract: A long-term goal in CT imaging is to achieve fast and accurate 3D reconstruction from sparse-view projections, thereby reducing radiation exposure,

researcharxiv-cs-cv
25 Jun 2026
Model Releases

Improving Factuality of 3D Brain MRI Report Generation with Paired Image-domain Retrieval and Text-domain Augmentation

DGX agent

arXiv:2411.15490v2 Announce Type: replace Abstract: Acute ischemic stroke (AIS) requires time-critical decision-making, where inaccurate interpretation of neuroimaging findings can lead to irreversibl

model-releasesarxiv-cs-cv
25 Jun 2026
Research

In-context Region-based Drag: Drag Any Region to Any Shape

DGX agent

arXiv:2606.25907v1 Announce Type: new Abstract: Diffusion models have shown promise in drag-style editing. Previous works mainly focus on point-based drag, which is inherently ambiguous. This paper fo

researcharxiv-cs-cv
25 Jun 2026
Model Releases

In-Context World Modeling for Robotic Control

DGX agent

arXiv:2606.26025v1 Announce Type: cross Abstract: Modern Vision-Language-Action (VLA) models often fail to generalize to novel setups, such as altered camera viewpoints or robot morphologies, because

model-releasesarxiv-cs-cv
25 Jun 2026
← Previous
1…8990919293…263
Next →