AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
23 Jul 2026

SUPER Module for Detail-Sensitive and Cost-Efficient U-Net Variant Decoders

ResearchDGX agent

arXiv:2511.11015v2 Announce Type: replace Abstract: Skip-connected U-Net variants are widely used for dense inverse problems, yet their decoders commonly recover resolution through spatial upscaling,

Surprise Forcing: What to Remember, When to Skip in Long Video Generation

ResearchDGX agent

arXiv:2607.18436v1 Announce Type: new Abstract: Streaming autoregressive diffusion makes minute-scale video synthesis practical, but its bounded context and fixed denoising schedule allocate resources

SWITi: Quantifying and Reducing Tiling Artifacts with Sliding Window Inner Tiling

TutorialsDGX agent

arXiv:2607.18990v1 Announce Type: new Abstract: SWITi is a test-time method for reducing artifacts in tiled predictions, particularly for neural networks that learn posterior distributions from which


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SynGallery: A Synthetic Gallery of Real Paintings for Instance-Level Artwork Recognition

Model ReleasesDGX agent

arXiv:2607.18907v1 Announce Type: new Abstract: Instance-level artwork recognition requires matching a handheld visitor photograph to a specific work in a large museum collection. This is challenging

Synthetic and Derived Training Images for Campus Waste Detection: A Multi-Seed Evaluation with YOLOv8n

ResearchDGX agent

arXiv:2607.19535v2 Announce Type: new Abstract: Incorrect disposal can contaminate campus recycling streams, and a bin-mounted camera could provide feedback as an item is discarded. We evaluated wheth

TAP-RAG: Task-Aware Policy Control for Long-Document Multimodal Question Answering

SafetyDGX agent

arXiv:2607.18917v1 Announce Type: new Abstract: Long-document multimodal question answering requires more than retrieving relevant chunks from a large document. Different queries require different evi

Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach

ResearchDGX agent

arXiv:2607.14702v2 Announce Type: replace Abstract: Automatic recognition of ambivalence and hesitancy is challenging because these states may be expressed through inconsistent linguistic, acoustic, f

Test-Time Registers as Global Priors for Tokenized Image Generation

SafetyDGX agent

arXiv:2607.16824v2 Announce Type: replace Abstract: Attention-based models often develop attention sinks, where a small number of tokens repeatedly attract attention and accumulate unusually large act

Text-conditioned Segmentation for Tomato Phenotyping via Procedural Synthetic Data

ApplicationsDGX agent

arXiv:2607.18576v1 Announce Type: new Abstract: Vision-based automation is an excellent candidate for reducing manual labor in greenhouse crop production and phenotyping. However, progress is constrai

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers

ResearchDGX agent

arXiv:2607.19139v1 Announce Type: new Abstract: Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation during denoising remains poorly unders

The JEPA Predictor: A Transferable Operator for Occluded Feature Completion

ResearchDGX agent

arXiv:2607.16274v2 Announce Type: replace Abstract: Joint-Embedding Predictive Architectures (JEPAs) train a predictor jointly with their encoder, but downstream deployment discards the predictor and

The PAR dataset: Prostate biopsy whole slide images from an underrepresented Middle Eastern population

ResearchDGX agent

arXiv:2512.03854v2 Announce Type: replace Abstract: Artificial intelligence (AI) is increasingly used in digital pathology. Publicly available histopathology datasets remain scarce, and those that do

Think Sparse, Predict Dense: Continuous Thought Machines for Image Super-Resolution

Model ReleasesDGX agent

arXiv:2607.18856v1 Announce Type: new Abstract: Continuous Thought Machines introduce an internal temporal dimension in which neuron-level histories and synchronization-derived representations evolve

Timeripple: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space

ResearchDGX agent

arXiv:2511.12035v2 Announce Type: replace-cross Abstract: The recent surge in video generation has shown the growing demand for high-quality video synthesis using large vision models. Existing video g

Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation

Model ReleasesDGX agent

arXiv:2509.24739v4 Announce Type: replace Abstract: Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (A

Toward Seasonal Guidelines for Robust Deep-Learning Sentinel-2 Building Detection in Different Area Types

ResearchDGX agent

arXiv:2607.19994v2 Announce Type: new Abstract: Sentinel-2 imagery offers open access, global coverage, and frequent revisit times, making it attractive for practical building mapping at scale; howeve

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning

ResearchDGX agent

arXiv:2607.19790v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has substantially improved language-model reasoning, yet its extension to vision-language models r

Trusted Multi-View Deep Learning Classification of Fetal Congenital Heart Disease with Feature-level and Decision-level Fusion

Model ReleasesDGX agent

arXiv:2606.15265v2 Announce Type: replace Abstract: Congenital heart disease (CHD) refers to the abnormal anatomical structure caused by the abnormal development of the heart and great vessels during

Universality Reconsidered: Rethinking the Validation of Foundation Models for General-Purpose 3D Medical Segmentation

Model ReleasesDGX agent

arXiv:2602.07643v2 Announce Type: replace Abstract: Foundation models have emerged as a transformative paradigm in 3D medical imaging, with the promise of unified quantitative analysis across diverse

UVFaceFusion: Fast Multi-view Topologically Consistent Face Reconstruction in the Wild via UV-space Neural Fusion

ResearchDGX agent

arXiv:2607.18798v1 Announce Type: new Abstract: Reconstructing high-fidelity facial geometry with an assigned topology is essential for digital avatar creation and animation, yet existing automated me

Vera: Identity-Faithful Human Subject-to-Video Generation

ResearchDGX agent

arXiv:2607.20247v1 Announce Type: new Abstract: Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across diverse categories, yet generic subject consiste

VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers

ResearchDGX agent

arXiv:2607.19575v1 Announce Type: new Abstract: Vector Quantization (VQ) underpins modern discrete visual tokenization. However, training quantization modules for state-of-the-art VQ-based models requ

WanSong v1.0 Technical Report

ResearchDGX agent

arXiv:2607.14749v3 Announce Type: replace-cross Abstract: Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fi

WASABI: Whole-graph Assignment-based Stabilizer for lAne topology By Inter-frame tracking

AgentsDGX agent

arXiv:2607.19781v1 Announce Type: new Abstract: Autonomous driving requires understanding the road as a graph of drivable lanes and their connectivity, beyond the ego lane alone, to follow routes thro

Wave2Body: Rethinking mmWave Human Pose Estimation as Radar-to-Body Token Translation

SafetyDGX agent

arXiv:2607.18875v1 Announce Type: new Abstract: Millimeter-wave (mmWave) radar enables privacy-friendly human sensing, but its sparse point clouds are physical measurements of view-dependent electroma

Wavefront Parallelization for Efficient Learned Image Compression

ResearchDGX agent

arXiv:2607.19082v1 Announce Type: cross Abstract: Autoregressive context models are foundational for learned image compression,but they suffer from slow serial inference. Existing acceleration methods

Weakly Supervised Pathology-Informed Representation Learning for PET-Based Content Retrieval of Intra-Tumour Heterogeneity

TutorialsDGX agent

arXiv:2607.18762v1 Announce Type: new Abstract: We propose a weakly supervised 18FFDG PET representation-learning framework for content based medical image retrieval, using H&E derived information dur

WearWow: Native 2K Multi-Garment Virtual Try-On via Adaptive Token Packing and Preference Alignment

SafetyDGX agent

arXiv:2607.19923v1 Announce Type: new Abstract: Synthesizing native 2K multi-garment virtual try-on is a formidable frontier in digital fashion, critically bottlenecked by two fundamental limitations:

WHU-PCPR: A cross-platform heterogeneous point cloud dataset for place recognition in complex urban scenes

Model ReleasesDGX agent

arXiv:2601.06442v2 Announce Type: replace Abstract: Point Cloud-based Place Recognition (PCPR) demonstrates considerable potential in applications such as autonomous driving, robot localization and na

ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting

ApplicationsDGX agent

arXiv:2607.18801v1 Announce Type: new Abstract: Recent advancements in 3D Gaussian Splatting (3DGS) have enabled language-guided scene understanding. However, existing Referring 3D Gaussian Splatting

16 Jul 2026

2D Rotary Position Embedding for Scene Text Recognition with Transformers

Model ReleasesDGX agent

arXiv:2607.13458v1 Announce Type: new Abstract: Scene Text Recognition (STR) remains challenging due to the diversity of text appearances, including curvature, rotation, and perspective distortion. Re

A Comparative Evaluation of Large Vision-Language Models for 2D Object Detection under SOTIF Conditions

Model ReleasesDGX agent

arXiv:2601.22830v2 Announce Type: replace Abstract: Reliable environmental perception remains one of the main obstacles for safe operation of automated vehicles. Safety of the Intended Functionality (

A Masked Autoencoder Approach to Unsupervised Steel Surface Defect Recognition

TutorialsDGX agent

arXiv:2607.13178v1 Announce Type: new Abstract: Automated visual inspection of steel surface defects is a recurring quality control task in which labeled defect data is scarce and costly to obtain, wh

A novel unsupervised machine learning strategy to handle multimodal cardiac PET/MRI data

ResearchDGX agent

arXiv:2607.13936v1 Announce Type: new Abstract: Arrhythmogenic left ventricular cardiomyopathy is a genetic myocardial disease difficult to diagnose due to the lack of gold standard criteria. Simultan

A Space-Time Transformer for Precipitation Nowcasting

ResearchDGX agent

arXiv:2511.11090v3 Announce Type: replace Abstract: Until recently, numerical weather prediction (NWP) models have stood rivalless in operational forecasting despite a few limitations. Namely, physica

Active Learning for Efficient Annotation of Surgical Videos with Weak Supervision

Local AiDGX agent

arXiv:2607.13237v1 Announce Type: new Abstract: Precise spatial-temporal annotation of laparoscopic videos is time-consuming and requires expert knowledge. We propose a human-in-the-loop knowledge acq

AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow

ResearchDGX agent

arXiv:2607.13250v1 Announce Type: new Abstract: We present extbf{AffectFlow-DINO}, a multi-task learning system for the 11th ABAW challenge that extends a standard deterministic architecture with a co

AnomExpert: Identifying and Selecting Anatomical Planes for Prenatal Ultrasound Anomaly Diagnosis

Model ReleasesDGX agent

arXiv:2607.13409v1 Announce Type: new Abstract: Life-limiting congenital anomalies require accurate prenatal diagnosis for appropriate clinical decision-making. Prenatal ultrasound (US) examinations i

AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization

SafetyDGX agent

arXiv:2607.13805v1 Announce Type: new Abstract: Contrastive Language-Image Pretraining learns a shared representation space through large-scale contrastive learning. However, existing methods that enf

Attention-Free and Lightweight Token Reduction for Efficient Vision-Language Models

ApplicationsDGX agent

arXiv:2607.13500v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have achieved strong performance in multimodal understanding, yet remain challenging to deploy on resource-constrained edg

Attentive multilayer fusion for vision transformers

ResearchDGX agent

arXiv:2601.09322v2 Announce Type: replace Abstract: With the rise of large-scale foundation models, efficiently adapting them to downstream tasks remains a central challenge. Linear probing, which fre

Audio-Text Cross-Attention with Psycholinguistic Support Features for Ambivalence/Hesitancy Recognition

ResearchDGX agent

arXiv:2607.13345v1 Announce Type: new Abstract: We present an audio-text system for the Ambivalence/Hesitancy Video Recognition Challenge of the 11th ABAW Competition. The method excludes visual frame

Bake It Till You Make It: Ultrafast Spatial Texture-Atlas Splatting

HardwareDGX agent

arXiv:2607.13808v1 Announce Type: new Abstract: Recent extensions of 3D Gaussian Splatting (3DGS) capture fine color details using hash-grid-based appearance parameterization but incur high computatio

BenthiCat: An opti-acoustic dataset for advancing benthic classification and habitat mapping

Model ReleasesDGX agent

arXiv:2510.04876v3 Announce Type: replace Abstract: Benthic habitat mapping is fundamental for understanding marine ecosystems, guiding conservation efforts, and supporting sustainable resource manage

Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents

Model ReleasesDGX agent

arXiv:2511.18685v4 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) show promising results as decision-making engines for embodied agents operating in complex, physical enviro

BiLoG-Net: A Bi-Context Location-Guided Network for Breast Mass Segmentation and Malignancy Classification in Mammography

Local AiDGX agent

arXiv:2607.10188v2 Announce Type: replace Abstract: Breast cancer remains the most commonly diagnosed malignancy among women worldwide, yet accurate detection and characterization of breast masses in

Bring Music The Horizon: Music-Driven 360^irc Video Generation

ResearchDGX agent

arXiv:2607.13471v1 Announce Type: new Abstract: Music visualization offers a powerful way to enhance listeners' understanding and experience of music by translating auditory signals into visual forms.

C-Norm: Cell-Distribution Normalization Enables Precision Recognition of Medical-Cell Image

SafetyDGX agent

arXiv:2607.13116v1 Announce Type: new Abstract: ThinPrep Cytologic Test (TCT) enables early cervical cancer screening, but manual reading is time-consuming and yields inconsistent diagnostic results a

Calibrated Closed-Form Uncertainty for Radiative Gaussian Splatting in Sparse-View CT

Model ReleasesDGX agent

arXiv:2607.13682v1 Announce Type: new Abstract: Radiative Gaussian splatting has made sparse-view CT reconstruction fast, but existing methods output point estimates with no notion of where the recons

CASA-SDF: Curriculum-Aware Spatial Adaptation with Curvature-Guided Density for Neural Implicit Surface Reconstruction

Model ReleasesDGX agent

arXiv:2607.13492v1 Announce Type: new Abstract: Neural implicit representations have emerged as a powerful paradigm for 3D reconstruction. However, high-fidelity indoor surface reconstruction remains

CF-Net: Conflict Fusion with Speaker Normalisation and Certainty Weighting for Ambivalence/Hesitancy Recognition

ResearchDGX agent

arXiv:2607.13976v1 Announce Type: new Abstract: Detecting ambivalence and hesitancy (AH) in unconstrained video is challenging because the target signal is inherently ambiguous and expressed through s

CLIP-Guided Label-Free Discriminative Region Scoring for Fine-Grained Classification

ResearchDGX agent

arXiv:2607.13437v1 Announce Type: new Abstract: Recent vision models such as CLIP and SAM enable training-free segmentation and semantic encoding for fine-grained classification. A common approach is

ClusIR: Towards Cluster-Guided All-in-One Image Restoration

ResearchDGX agent

arXiv:2512.10948v2 Announce Type: replace Abstract: All-in-One Image Restoration (AiOIR) aims to recover high-quality images from diverse degradations within a unified framework. However, existing met

Cyclone: Diffusion Model for Cycle-Consistent Weather Editing from Unpaired Driving Data

AgentsDGX agent

arXiv:2607.13927v1 Announce Type: new Abstract: Reliable perception under diverse weather conditions remains a major challenge for autonomous driving systems. A common strategy to improve robustness i

Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization

ResearchDGX agent

arXiv:2607.13336v1 Announce Type: new Abstract: Recent diffusion-based video generation models have enabled high-quality personalized video customization through both tuning-based pipelines, which fin

Detector Confidence Signals Presence Rather Than Occlusion in Cluttered Manipulation

Model ReleasesDGX agent

arXiv:2607.13361v1 Announce Type: new Abstract: Occlude a named object until about an eighth of it remains visible, and an open-vocabulary detector's confidence that the object is present barely chang

Differentiable Polarized Path Tracing

Local AiDGX agent

arXiv:2607.13265v1 Announce Type: new Abstract: Physically based differentiable rendering has proven to be a powerful tool for inverse rendering problems (e.g., 3D reconstruction, reflectance estimati

DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation

ResearchDGX agent

arXiv:2607.13365v1 Announce Type: new Abstract: Existing 3D generative models predominantly rely on implicit volumetric representations, which enforce watertight topology and struggle to represent thi

DNA: Dual-stage Native Attribution for Generated Image Source Tracing

Model ReleasesDGX agent

arXiv:2607.13685v1 Announce Type: new Abstract: The rapid evolution of image generation has produced numerous within-family variants, making source-model attribution of suspect images increasingly imp

DP-BOA: Dirichlet-Process Birth-or-Assign for On-the-Fly Category Discovery

ResearchDGX agent

arXiv:2607.13504v1 Announce Type: new Abstract: On-the-fly category discovery requires deciding for each incoming test sample whether to assign it to an existing category or spawn a new one. Existing

← Previous
1…3839404142…209
Next →