AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
5 Aug 2026

Modeling Scientific Experiment Scenes: Dataset and Model

ApplicationsDGX agent

arXiv:2608.02892v1 Announce Type: new Abstract: Scene Graph Generation (SGG) is fundamental to structured visual understanding, yet existing benchmarks focus mainly on daily life images and overlook s

MoECa: Aligning Feature Reuse with Expert Decomposition in Diffusion Transformers

ResearchDGX agent

arXiv:2606.15615v2 Announce Type: replace-cross Abstract: Diffusion Transformers with Mixture-of-Experts (DiT-MoE) improve model capacity under sparse activation, but diffusion inference is still bott

Morphology-Aware Implicit Super-Resolution Network for Pathological Images

Model ReleasesDGX agent

arXiv:2608.03664v1 Announce Type: new Abstract: Accurate diagnosis in Digital Pathology (DP) relies on high-resolution whole-slide images, yet clinical deployment is often limited by hardware costs. S


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

MSTAR: Multi-Scale Backbone Architecture Search for Timeseries Classification

ResearchDGX agent

arXiv:2402.13822v2 Announce Type: replace Abstract: Most of the previous approaches to Time Series Classification (TSC) highlight the significance of receptive fields and frequencies while overlooking

MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification

Model ReleasesDGX agent

arXiv:2608.03474v1 Announce Type: new Abstract: Recent advances in Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in web UI generation. However, existing benchmarks pre

MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding

Model ReleasesDGX agent

arXiv:2608.03708v1 Announce Type: new Abstract: Text-to-image diffusion models enable personalization of specific visual concepts from a small number of reference images. However, generating a single

Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis

ResearchDGX agent

arXiv:2608.03109v1 Announce Type: new Abstract: Plant root phenotyping is fundamental to understanding below-ground structures, optimizing crop management, and improving agricultural sustainability. T

MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization

Model ReleasesDGX agent

arXiv:2608.03885v1 Announce Type: new Abstract: Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-ti

NanoMorph-3D: An End-to-End Physics-Driven Unrolling Framework for Nanomaterial Reconstruction

Local AiDGX agent

arXiv:2608.03257v1 Announce Type: new Abstract: Precise 3D characterization of nanomaterials is essential for unlocking structure-property relationships. However, standard electron tomography is funda

NCGR: Noise-Conditional Gated Rectification for Camera Extrinsic Perturbations in BEV 3D Object Detection

ResearchDGX agent

arXiv:2608.03895v1 Announce Type: new Abstract: Camera-based bird's-eye-view (BEV) 3D detection typically assumes accurate and fixed camera extrinsics. In detectors using spatial cross-attention (SCA)

NearID: Identity Representation Learning via Near-identity Distractors

Model ReleasesDGX agent

arXiv:2604.01973v2 Announce Type: replace Abstract: When evaluating identity-focused tasks such as personalized generation and image editing, existing vision encoders entangle object identity with bac

Non-Destructive Quantification of Urea Adulteration in Bovine Milk Using Transmittance Multispectral Imaging

ResearchDGX agent

arXiv:2608.03113v1 Announce Type: new Abstract: Adulteration of bovine milk using urea remains a major food quality and health concern, motivating the development of rapid and quantitative screening t

Oh Deer, How Should I Handle This? Seasonal Priors for Selective Wildlife Annotation and Classification

TutorialsDGX agent

arXiv:2608.02762v1 Announce Type: new Abstract: Fine-grained wildlife classification in aerial imagery is limited not only by model performance, but also by unreliable labels: animals occupy few pixel

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

ResearchDGX agent

arXiv:2608.03812v1 Announce Type: new Abstract: Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly re

Open-Linguistic Concept Unified Learning for Cross-Site Interpretable Dermatology Image Diagnosis

ResearchDGX agent

arXiv:2608.03225v1 Announce Type: new Abstract: Human-interpretable computer-aided diagnosis is crucial for clinical decision making. Concept-based models excel by providing transparent reasoning and

Perceptual Anchoring: Prototype-Guided Text Calibration for Training-free Open-Vocabulary Semantic Segmentation

SafetyDGX agent

arXiv:2608.03991v1 Announce Type: new Abstract: Training-free open-vocabulary semantic segmentation (OVSS) partitions an image into semantically distinct regions based on arbitrary text descriptions,

PixelUp: Zero-Shot Semantic Feature Upsampling for Fine-Grained Vision Tasks

ResearchDGX agent

arXiv:2608.02792v1 Announce Type: new Abstract: Self-supervised Vision Foundation Models (VFMs) have become essential backbones for downstream tasks due to their strong and transferable visual represe

PLS-Calib: A Partial Least Squares Framework for Event Camera and Odometry Calibration under Ground Motion Constraints

ApplicationsDGX agent

arXiv:2608.03296v1 Announce Type: cross Abstract: Accurate extrinsic rotation calibration between sensors is fundamental to the performance of robotic perception systems. However, most existing calibr

Poisoning Prompt-Guided Sampling in Video Large Language Models

SafetyDGX agent

arXiv:2509.20851v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) are increasingly deployed as automated moderators on user-generated video platforms, where a few unwatched s

PolyLayout: Multi-room Manhattan Layout Estimation

ResearchDGX agent

arXiv:2608.03323v1 Announce Type: new Abstract: Estimating room layouts from multi-view imagery is a core task for indoor scene understanding. Existing methods are typically limited either by poor gen

Predictive Enhancement Calibration for Latent Breast MRI Virtual Contrast Enhancement

Model ReleasesDGX agent

arXiv:2608.03612v1 Announce Type: cross Abstract: Virtual contrast enhancement (VCE) synthesizes enhanced breast MR images from pre-contrast acquisitions. Modern latent generators offer strong image p

Progressive Learning of a Diffusion-based Inpainting Model for Separating Overlapped Fingerprints

TutorialsDGX agent

arXiv:2608.03937v1 Announce Type: new Abstract: Overlapped friction ridge patterns are a recurring problem in latent fingerprints recovered from crime scenes and in live-scan scenarios where residual

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

Model ReleasesDGX agent

arXiv:2608.02980v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to f

RealWeather: Realistic and Scene-Faithful Weather Translation with Driving World Models

SafetyDGX agent

arXiv:2608.02953v1 Announce Type: new Abstract: Realistic weather translation is valuable for developing and evaluating autonomous driving systems, yet collecting paired videos of the same scenes unde

ReCamDriving: LiDAR-Free Camera-Controlled Video Synthesis for Novel Trajectories

AgentsDGX agent

arXiv:2512.03621v3 Announce Type: replace Abstract: Synthesizing multi-pass videos is important for autonomous driving. While current repair-based methods often struggle with out-of-distribution artif

Recurrent Contrastive Learning for Imbalanced Medical Image Classification

ResearchDGX agent

arXiv:2608.03304v1 Announce Type: new Abstract: Medical image classification often suffers from class imbalance due to the inherent disparities in disease incidence. Existing approaches, such as class

Residual Flow Matching with Dynamic Cross-Interaction for 3D Multi-Person Motion Prediction

AgentsDGX agent

arXiv:2608.03379v1 Announce Type: new Abstract: 3D multi-person motion prediction requires modeling both individual kinematics and inter-person interactions. While Flow Matching is effective for multi

Rethinking Uncertainty Quantification and Entanglement in Image Segmentation

SafetyDGX agent

arXiv:2603.18792v2 Announce Type: replace Abstract: Uncertainty quantification (UQ) is crucial in safety-critical applications such as medical image segmentation. Total uncertainty is typically decomp

RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing

SafetyDGX agent

arXiv:2608.03059v1 Announce Type: new Abstract: Inversion-free flow-based image editing avoids latent inversion, but still requires a target-side state at every editing step. The widely used equal-dis

S^3-Diff: Structural Semantic Synergy Diffusion Model for High Fidelity Super Resolution of Pathological Images

ResearchDGX agent

arXiv:2608.03540v1 Announce Type: new Abstract: Digital pathology relies on high-resolution whole slide images for accurate diagnosis, yet limitations in imaging devices, storage, and transmission oft

SAMSEM -- A Generic and Scalable Approach for IC Metal Line Segmentation

ApplicationsDGX agent

arXiv:2603.16548v2 Announce Type: replace-cross Abstract: In light of globalized hardware supply chains, the assurance of hardware components has gained significant interest, particularly in cryptogra

SeCo-SBIR: Semantically Consistent Prompt Learning for Zero-Shot Sketch-Based Image Retrieval

ResearchDGX agent

arXiv:2608.03120v1 Announce Type: new Abstract: Adapting CLIP for zero-shot sketch-based image retrieval (ZS-SBIR) via prompt learning faces a fundamental tension: the model must bridge the sketch-pho

SEER: A Self-Grounded Evidence Interface for Controlled Spatial Relation Classification

Local AiDGX agent

arXiv:2608.03631v1 Announce Type: new Abstract: Spatial relation questions require a model to identify the queried subject and object before comparing their layout. Yet a VLM can recognize both entiti

Self-Supervised Uncalibrated Multi-View Video Anonymization in the Operating Room

Local AiDGX agent

arXiv:2602.02850v3 Announce Type: replace Abstract: Privacy preservation is a prerequisite for using video data in Operating Room (OR) research. Effective anonymization relies on the exhaustive locali

SGFormer: Structure-Guided Transformer for Robust Local Feature Matching

Local AiDGX agent

arXiv:2608.03423v1 Announce Type: new Abstract: Local feature matching is a fundamental component of photogrammetry, enabling accurate image correspondence critical for tasks such as 3D reconstruction

SLAMFormer-infty: Infinite SLAM Transformer for Unbounded Frontend and Backend Processing

ResearchDGX agent

arXiv:2608.03429v1 Announce Type: new Abstract: We introduce the Infinite SLAM Transformer (SLAMFormer-infty), the first geometric transformer capable of supporting both long-range frontend and backen

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models

Model ReleasesDGX agent

arXiv:2608.03580v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have demonstrated remarkable performance in processing and understanding both text and images, their large parameter

SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference

SafetyDGX agent

arXiv:2608.03335v1 Announce Type: new Abstract: Video diffusion transformers (vDiTs) generate high quality but pay quadratic self-attention cost, making inference prohibitive at video-token scales. Th

SpreadMark: Robust Image Watermarking via Spread-Spectrum Embedding

ResearchDGX agent

arXiv:2608.03165v1 Announce Type: cross Abstract: Invisible image watermarks are increasingly used for deepfake detection and provenance tracking, where they must survive not only incidental distortio

SRAP: SVD-Refined Adversarial Perturbations for Imperceptible Face-Swap Defense

ResearchDGX agent

arXiv:2608.03395v1 Announce Type: new Abstract: Deepfake technologies pose increasing threats to facial privacy and identity security, motivating proactive defenses that protect facial images before m

StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision

Model ReleasesDGX agent

arXiv:2603.29368v2 Announce Type: replace Abstract: Driven by the advancement of 3D devices, stereo vision tasks including stereo matching and stereo conversion have emerged as a critical research fro

Stop Replacing Noise with Noise: Two-Source Reliability Assessment for Label Correction and Sample Reweighting in Label-Noise Learning

ApplicationsDGX agent

arXiv:2608.03432v1 Announce Type: cross Abstract: Refurbishment-based noisy-label learning mixes an observed label with a model-derived pseudo target, typically using one sample-wise cleanliness score

StreamDAM: Presence-Aware Memory for Real-Time Streaming Video Object Segmentation

SafetyDGX agent

arXiv:2608.03912v1 Announce Type: new Abstract: Quality-tier video object segmentation (VOS) trackers such as DAM4SAM top accuracy leaderboards, but they are measured offline, one frame at a time with

Style-Aware Gloss Control for Generative Non-Photorealistic Rendering

TutorialsDGX agent

arXiv:2602.16611v3 Announce Type: replace-cross Abstract: Humans can infer material characteristics of objects from their visual appearance, and this ability extends to artistic depictions, where simi

Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation

ResearchDGX agent

arXiv:2608.03158v1 Announce Type: new Abstract: Daily activities require humans to coordinate whole-body motion with the motion of surrounding objects. Despite recent progress in human-object interact

SUV: Future Scene Understanding as Video Generation for End-to-End Driving

Model ReleasesDGX agent

arXiv:2608.03084v1 Announce Type: new Abstract: End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output form

T2VAttack: Adversarial Attack on Text-to-Video Diffusion Models

SafetyDGX agent

arXiv:2512.23953v2 Announce Type: replace Abstract: The rapid evolution of Text-to-Video (T2V) diffusion models has driven remarkable advancements in generating high-quality, temporally coherent video

TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models

ResearchDGX agent

arXiv:2608.03057v1 Announce Type: new Abstract: Static quantization assigns one weight precision to every denoising step. To preserve quality, that precision must accommodate the most quantization-sen

TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding

Local AiDGX agent

arXiv:2608.03763v1 Announce Type: new Abstract: Zero-shot 3D visual grounding aims to localize specific objects based on textual descriptions and 3D visual input. However, the effectiveness of existin

Test Time Adaptation Methods for Point Cloud Registration in Laparoscopic Surgery

SafetyDGX agent

arXiv:2608.02883v1 Announce Type: new Abstract: 3D point cloud registration in laparoscopic surgery estimates the transformation between an intraoperative organ reconstructed from video and its preope

Test-Time Augmentation for Tabular-to-Image Classifiers under Distribution Shifts

Model ReleasesDGX agent

arXiv:2608.03557v1 Announce Type: new Abstract: Tabular-to-image methods that convert tabular data into visual representations have emerged as a novel paradigm for leveraging the high performance of d

Toward Visual Grounding: A Survey

ResearchDGX agent

arXiv:2412.20206v4 Announce Type: replace Abstract: Visual Grounding, also known as Referring Expression Comprehension and Phrase Grounding, aims to ground the specific region(s) within the image(s) b

Towards Reliable and Reproducible Fetal Brain Biometry: A Deep Learning Approach Using MRI

ResearchDGX agent

arXiv:2608.03724v1 Announce Type: new Abstract: Fetal brain biometry is essential for quantitative assessment of brain development, supporting gestational age estimation, developmental monitoring, and

Two-Way Garment Transfer: Unified Diffusion Framework for Dressing and Undressing Synthesis

ResearchDGX agent

arXiv:2508.04551v2 Announce Type: replace Abstract: While recent advances in virtual try-on (VTON) have achieved realistic garment transfer to human subjects, its inverse task, virtual try-off (VTOFF)

UniEvo-RS: Omni-Prompt Unified Remote Sensing Segmentation with Representative Exemplar-Driven Prototype Evolution

ResearchDGX agent

arXiv:2608.03911v1 Announce Type: new Abstract: Prompt-driven vision-language models (VLMs) hold immense promise for accelerating dense remote sensing (RS) annotation, but static models suffer from se

UniWorld-Design: From Pixel Generation to Layer-Native Design

Model ReleasesDGX agent

arXiv:2608.03971v1 Announce Type: new Abstract: We introduce UniWorld-Design, a framework that redefines image generation from flat pixel synthesis to structured visual composition, with semantic RGBA

Unsupervised Adversarial Domain Adaptation for Uterine layer Segmentation: From Labeled Cine to Unlabeled Dynamic EPI MRI

ResearchDGX agent

arXiv:2608.03762v1 Announce Type: cross Abstract: Uterine peristalsis is a key physiological phenomenon responsible for various functions across the menstrual cycle, intimately linked to uterine wall

VLC Fusion: Vision-Language Conditioned Sensor Fusion for Robust Object Detection

AgentsDGX agent

arXiv:2505.12715v2 Announce Type: replace Abstract: Although fusing multiple sensor modalities can enhance object detection performance, existing fusion approaches often overlook subtle variations in

What is the Right Embedding Space for Contrastive Learning in REC?

SafetyDGX agent

arXiv:2505.22850v2 Announce Type: replace Abstract: Referring Expression Counting (REC) requires distinguishing visually similar objects described by fine-grained text cues. Existing methods tackle th

When Classes Evolve: A Benchmark and Framework for Stage-Aware Class-Incremental Learning

Model ReleasesDGX agent

arXiv:2602.00573v2 Announce Type: replace-cross Abstract: Class-Incremental Learning (CIL) aims to sequentially learn new classes while mitigating catastrophic forgetting of previously learned knowled

← Previous
1…1213141516…207
Next →