AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
25 Jun 2026

Did Models Learn Sufficiently? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation

TutorialsDGX agent

arXiv:2511.12100v2 Announce Type: replace Abstract: In current visual model training, models often rely on only limited sufficient causes for their predictions, which makes them sensitive to distribut

Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography

Model ReleasesDGX agent

arXiv:2606.25546v1 Announce Type: new Abstract: Vision-language pre-training (VLP) holds great promise for general-purpose medical AI by leveraging radiology reports as rich textual supervision, yet e

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2606.25066v1 Announce Type: cross Abstract: Visual search has been one of the most productive paradigms in the study of visual attention: the way reaction time scales with the number of items di

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation

ResearchDGX agent

arXiv:2606.26058v1 Announce Type: new Abstract: Open domain subject-driven text-to-video (S2V) generation has drawn significant interest in academia and industry. Open domain S2V mainly involves two s

DRM: Diffusion-based Reward Model With Step-wise Guidance

SafetyDGX agent

arXiv:2605.25661v2 Announce Type: replace Abstract: Current mainstream methods of aligning diffusion models with human preferences typically employ VLM-based reward models. However, these reward model

DSP-SLAM++: A Unified Framework for Multi-Class, High-Fidelity Object SLAM in the Wild

AgentsDGX agent

arXiv:2606.25953v1 Announce Type: cross Abstract: Existing object-aware SLAM systems force a trade-off between real-time performance, multi-class support, and the generation of high-fidelity, semantic

Dual Agreement Consistency Learning for Semi-Supervised Fetal Ultrasound Segmentation

SafetyDGX agent

arXiv:2606.25254v1 Announce Type: cross Abstract: Maternal-fetal US is the primary imaging modality for monitoring fetal development, yet accurate automated segmentation remains challenging due to the

Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs

Model ReleasesDGX agent

arXiv:2606.25758v1 Announce Type: new Abstract: While test-time adaptation (TTA) empowers vision-language models to adapt without costly retraining, it remains highly vulnerable to out-of-distribution

EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis

ResearchDGX agent

arXiv:2606.25465v1 Announce Type: new Abstract: While image stylization has been studied extensively, video stylization remains a critical and largely unsolved challenge in the field of intelligent co

Edges Before Embeddings: A Confidence-Aware Blur Gate for Vision-Language Pipelines

ApplicationsDGX agent

arXiv:2606.25838v1 Announce Type: new Abstract: Production vision pipelines silently degrade on blurry input, wasting compute on downstream OCR, retrieval, and vision-language model (VLM) calls that c

Efficient Cross-Scale Invertible Hiding Network with Spatial-Frequency Collaboration and Non-Invertible Mechanism

ResearchDGX agent

arXiv:2606.25547v1 Announce Type: new Abstract: Image hiding aims to conceal image-level messages within cover images at the same resolution. Invertible neural networks (INN)-based image hiding has em

Efficient Real-World Dehazing via Physics-Inspired Global-Local Decoupling

Model ReleasesDGX agent

arXiv:2606.25732v1 Announce Type: new Abstract: Real-world single image dehazing is highly ill-posed due to spatially and spectrally varying scattering, while practical deployment demands lightweight

Efficient Remote Sensing Instance Segmentation with Linear-Time State Space Distilled Visual Foundation Models

Model ReleasesDGX agent

arXiv:2606.25324v1 Announce Type: new Abstract: The computational complexity of Transformers scales quadratically with the number of tokens, which significantly constrains the efficiency of vision mod

Energy-Efficient CNN Acceleration with MSDF Digit-Serial Arithmetic on FPGA

ResearchDGX agent

arXiv:2606.25562v1 Announce Type: cross Abstract: This paper presents an energy-efficient hardware acceleration of the convolutional layers in the U-Net architecture for image segmentation, implemente

Enhancing Brain MRI Anomaly Detection and Reasoning with ROI Rethink and Synthetic Data

Model ReleasesDGX agent

arXiv:2606.25894v1 Announce Type: new Abstract: Medical vision-language models typically generate diagnoses through single-pass inference without indicating which image regions support their conclusio

Enhancing Pathological VLMs with Cross-scale Reasoning

Model ReleasesDGX agent

arXiv:2606.17412v3 Announce Type: replace Abstract: Pathological images are inherently multi-scale, requiring pathologists to integrate evidence from global tissue architecture at low magnification to

Entropy-Controlled Flow Matching

ResearchDGX agent

arXiv:2602.22265v2 Announce Type: replace-cross Abstract: Modern vision generators transport a base distribution to data through time-indexed measures, implemented as deterministic flows (ODEs) or sto

ESMStereo: Enhanced ShuffleMixer Disparity Upsampling for Real-Time and Accurate Stereo Matching

Local AiDGX agent

arXiv:2506.21091v2 Announce Type: replace Abstract: Stereo matching has become an increasingly important component of modern autonomous systems. Developing deep learning-based stereo matching models t

ESTANet: Efficient Online Error Detection in Procedural Videos via Prediction Inconsistency

TutorialsDGX agent

arXiv:2606.25317v1 Announce Type: new Abstract: An efficient and accurate system for detecting errors in procedural tasks is crucial for supporting human needs in daily life, as it can provide instant

Evaluation Protocols and Validation for Cameras in Indoor Healthcare Monitoring

SafetyDGX agent

arXiv:2606.25284v1 Announce Type: new Abstract: Camera-based monitoring systems are increasingly adopted in healthcare settings for the continuous assessment of patient movement and activities. Howeve

Evidential Perfusion Physics-Informed Neural Networks with Residual Uncertainty Quantification

Model ReleasesDGX agent

arXiv:2603.09359v2 Announce Type: replace Abstract: Physics-informed neural networks (PINNs) have shown promise in addressing the ill-posed deconvolution problem in computed tomography perfusion (CTP)

Expresso-AI: Explainable Video-Based Deep Learning Models for Depression Diagnosis

ResearchDGX agent

arXiv:2606.25606v1 Announce Type: new Abstract: Given the widespread prevalence of depression and its consequential impact on individuals and society, it is crucial to obtain objective measures for ea

Falcon: Functional Assembly and Language for Compositional Reasoning in X-ray

Model ReleasesDGX agent

arXiv:2606.25701v1 Announce Type: new Abstract: Conventional vision-language models are largely object-centric, focusing on detecting and describing individual entities. In safety-critical X-ray bagga

fARfetch: Enabling Collocated AR-HRC in Large Visually Diverse Environments with VLM-Driven AR Content Adaptation

ApplicationsDGX agent

arXiv:2606.25162v1 Announce Type: cross Abstract: Augmented Reality (AR) can improve collocated human-robot collaboration by making robot state and intent visible and enabling intuitive control, yet l

FedReLa: Imbalanced Federated Learning via Re-Labeling

Local AiDGX agent

arXiv:2606.26037v1 Announce Type: cross Abstract: Federated learning has emerged as the foremost approach for decentralized model training with privacy preservation. The global class imbalance and cro

FeVOS: Foresight Expression Video Object Segmentation

ResearchDGX agent

arXiv:2606.25585v1 Announce Type: new Abstract: Existing Referring Video Object Segmentation tasks focus on referring expressions describing events, actions or appearances of relevant objects within t

FlowID : Enhancing Forensic Identification with Latent Flow-Matching Models

Model ReleasesDGX agent

arXiv:2603.29591v2 Announce Type: replace Abstract: Every day, many people die under violent circumstances, whether from crimes, war, migration, or climate disasters. Medico-legal and law enforcement

Follow Your Track: Precise Skeleton Animation Controlled by 3D Trajectories

SafetyDGX agent

arXiv:2606.25344v1 Announce Type: new Abstract: 4D generation aims to animate 3D objects with realistic motion, holding great promise for applications. Existing methods typically decouple 3D asset gen

FreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling

Model ReleasesDGX agent

arXiv:2606.25079v1 Announce Type: new Abstract: Visual storytelling aims to generate image sequences that are both aligned with narrative prompts and consistent in character appearance across images.

From Sparse and Imperfect 2D Anchors to Consistent 3D Gaussian Street Scenes: Support-Aware Appearance

Model ReleasesDGX agent

arXiv:2606.26007v1 Announce Type: new Abstract: Image priors can synthesize target conditions for 3D Gaussian street scenes, but independently edited views do not define a coherent 3D target. Direct f

FunPiQ: A New Benchmark for Pixel-Level Quality Assessment in Fundus Images

Model ReleasesDGX agent

arXiv:2606.25915v1 Announce Type: new Abstract: Color fundus photography (CFP) is the most common ophthalmic imaging modality for large-scale screening. However, it is highly susceptible to degradatio

Gastroendoscopy View Synthesis: A New Real Dataset and Evaluation

ResearchDGX agent

arXiv:2606.25427v1 Announce Type: new Abstract: Novel view synthesis (NVS) is an active research topic in computer vision, owing to the success of neural radiance field (NeRF) and 3D Gaussian splattin

Geometry-Anchored Transport Framework for Exemplar-Free Class-Incremental Learning

SafetyDGX agent

arXiv:2606.25347v1 Announce Type: cross Abstract: Exemplar-free class-incremental learning (EFCIL) requires stable decision boundaries within a shifting feature space. While maintaining class-conditio

GeoRanker: Distance-Aware Ranking for Worldwide Image Geolocalization

ResearchDGX agent

arXiv:2505.13731v4 Announce Type: replace Abstract: Worldwide image geolocalization-the task of predicting GPS coordinates from images taken anywhere on Earth-poses a fundamental challenge due to the

Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs

ResearchDGX agent

arXiv:2606.25842v1 Announce Type: new Abstract: Existing multi-modal large language models (MLLMs) face significant challenges in processing long video sequences due to strict input token limitations.

GroundSet: A Cadastral-Grounded Dataset for Spatial Understanding with Vector Data

Model ReleasesDGX agent

arXiv:2603.14609v2 Announce Type: replace Abstract: Precise spatial understanding in Earth Observation is essential for translating raw aerial imagery into actionable insights for critical application

H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks

TutorialsDGX agent

arXiv:2606.25578v1 Announce Type: new Abstract: Hairstyle transfer has practical applications such as virtual try-on, yet remains challenging when the source and reference exhibit large head-pose disc

Heterogeneous and Adept Snapshot Distillation for 3D Semantic Segmentation

ResearchDGX agent

arXiv:2606.25278v1 Announce Type: new Abstract: Multi-modal fusion and multi-model ensembling are prevalent in enhancing the performance of 3D semantic segmentation. Despite the impressive performance

HG-Bench: A Benchmark for Multi-Page Handwritten Answer-Region Grounding in Automated Homework Assessment

Model ReleasesDGX agent

arXiv:2606.25491v1 Announce Type: new Abstract: Automated homework assessment depends not only on recognizing student answers, but also on accurately locating where each answer and each intermediate r

HiFiVe: High-Fidelity Vehicle Generation Leveraging Auto-Regressive 2D Generative Priors

ApplicationsDGX agent

arXiv:2606.25300v1 Announce Type: new Abstract: Existing 3D vehicle generation methods often suffer from low geometric fidelity and blurry textures, hindering their downstream applications. While rece

HiT-JEPA: A Hierarchical Self-supervised Trajectory Embedding Framework for Similarity Computation

Local AiDGX agent

arXiv:2507.00028v2 Announce Type: replace-cross Abstract: The representation of urban trajectory data plays a critical role in effectively analyzing spatial movement patterns. Despite considerable pro

Homomorphic Encryptions for Privacy Preserving Vision

ApplicationsDGX agent

arXiv:2606.25216v1 Announce Type: cross Abstract: Legal requirements might prevent organizations from sharing sensitive data like medical or financial details of consumers which prevents them from lev

Hybrid deep learning-based phase diversity method for wavefront reconstruction

ResearchDGX agent

arXiv:2606.25855v1 Announce Type: cross Abstract: The efficiency of high-power laser systems is limited by wavefront distortions in the beam, particularly non-common path aberrations, which reduce the

Hypergraph Normal World Models for Logical Visual Anomaly Detection

Local AiDGX agent

arXiv:2606.25368v1 Announce Type: new Abstract: Visual anomaly detection is often deployed with only normal training images. Most one-class detectors map test patches or features to a normal reference

ILV: Iterative Latent Volumes for Fast and Accurate Sparse-View CT Reconstruction

ResearchDGX agent

arXiv:2603.14915v2 Announce Type: replace Abstract: A long-term goal in CT imaging is to achieve fast and accurate 3D reconstruction from sparse-view projections, thereby reducing radiation exposure,

Improving Factuality of 3D Brain MRI Report Generation with Paired Image-domain Retrieval and Text-domain Augmentation

Model ReleasesDGX agent

arXiv:2411.15490v2 Announce Type: replace Abstract: Acute ischemic stroke (AIS) requires time-critical decision-making, where inaccurate interpretation of neuroimaging findings can lead to irreversibl

In-context Region-based Drag: Drag Any Region to Any Shape

ResearchDGX agent

arXiv:2606.25907v1 Announce Type: new Abstract: Diffusion models have shown promise in drag-style editing. Previous works mainly focus on point-based drag, which is inherently ambiguous. This paper fo

In-Context World Modeling for Robotic Control

Model ReleasesDGX agent

arXiv:2606.26025v1 Announce Type: cross Abstract: Modern Vision-Language-Action (VLA) models often fail to generalize to novel setups, such as altered camera viewpoints or robot morphologies, because

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity

Model ReleasesDGX agent

arXiv:2606.25343v1 Announce Type: new Abstract: Vision Language Models have achieved near-human performance on single-document Visual Question Answering, yet their effectiveness degrades significantly

KidRisk: Benchmark Dataset for Children Dangerous Action Recognition

Model ReleasesDGX agent

arXiv:2606.25298v1 Announce Type: new Abstract: Children are naturally energetic, and during their spontaneous activities, they often encounter potentially dangerous situations, especially when lackin

Latent Space Analysis for Interpretable Uncertainty in Melanoma Classification

ResearchDGX agent

arXiv:2506.18414v3 Announce Type: replace Abstract: Melanoma is a highly aggressive skin cancer, making early and accurate diagnosis critical. While deep learning excels in skin lesion classification,

Learning Action Priors for Cross-embodiment Robot Manipulation

SafetyDGX agent

arXiv:2606.26095v1 Announce Type: cross Abstract: Most Vision-Language-Action (VLA) models build on a Vision-Language Model (VLM) backbone by attaching an action module and optimizing the full policy

LEVIRDet: A Million-Scale 159-Category Dataset and Foundation Model for Universal Remote Sensing Object Detection

Model ReleasesDGX agent

arXiv:2606.25312v1 Announce Type: new Abstract: Remote sensing object detection has advanced rapidly with the development of large-scale benchmarks and modern detection architectures. However, existin

LinStereo: Linear-Complexity Global Attention for Multi-Scale Iterative Stereo Matching

ApplicationsDGX agent

arXiv:2606.25437v1 Announce Type: new Abstract: Existing Vision Foundation Model (VFM)-based iterative stereo pipelines under-exploit three information pathways: multi-scale backbone features are coll

MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation

ResearchDGX agent

arXiv:2606.26016v1 Announce Type: new Abstract: Normalizing Flows (NFs) are powerful generative models capable of exact density estimation and sampling. However, their strict invertibility often force

Minimalist Preprocessing Approach for Image Synthesis Detection

ResearchDGX agent

arXiv:2606.25297v1 Announce Type: new Abstract: Generative models have significantly advanced image generation, resulting in synthesized images that are increasingly indistinguishable from authentic o

MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning

TutorialsDGX agent

arXiv:2606.25225v1 Announce Type: new Abstract: Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning. Since audio and visual strea

MRI2Rep: Autoregressive Structured Report Generation for 3D Liver MRI

ApplicationsDGX agent

arXiv:2606.25279v1 Announce Type: new Abstract: Manual reporting of 3D MRI studies is time-consuming, yet end-to-end structured report generation for 3D liver MRI remains underexplored due to volumetr

MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation

ResearchDGX agent

arXiv:2606.26087v1 Announce Type: new Abstract: Synthesizing a novel-view video from a monocular reference video along a target camera trajectory requires both geometric consistency and motion fidelit

Naturalness Predicts but Does Not Cause Transferability in Image Encodings of Real-World Streams

ApplicationsDGX agent

arXiv:2606.25844v1 Announce Type: new Abstract: A common practice converts a one-dimensional signal into an image so that a vision backbone pretrained on natural photographs can be reused for recognit

← Previous
1…7172737475…211
Next →