AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
27 May 2026

HydraPrompt: An Adaptive and Asymmetric Framework of Vision-Language Models for Synthetic Image Detection

ResearchDGX agent

arXiv:2605.26421v1 Announce Type: new Abstract: The rapid evolution of generative models has precipitated a proliferation of fabricated content, posing significant challenges to existing Synthetic Ima

I2PRef: Image-Driven Point Completion with Iterative Refinement

TutorialsDGX agent

arXiv:2605.26914v1 Announce Type: new Abstract: We present an image-conditioned point cloud completion approach that treats images as the primary geometric source rather than a secondary guide. To thi

Image Thresholding: Understanding Bias of Evaluation Metrics towards Specific Evaluation Functions

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.27132v1 Announce Type: new Abstract: Multilevel image thresholding is widely used for segmentation in applications ranging from medical imaging to remote sensing. Classical objective functi

ImViD: Immersive Volumetric Videos for Enhanced VR Engagement

Model ReleasesDGX agent

arXiv:2503.14359v2 Announce Type: replace Abstract: User engagement is greatly enhanced by fully immersive multi-modal experiences that combine visual and auditory stimuli. Consequently, the next fron

Innovative Silicosis and Pneumonia Classification: Leveraging Graph Transformer Post-hoc Modeling and Ensemble Techniques

ResearchDGX agent

arXiv:2501.00520v2 Announce Type: replace Abstract: This paper presents a comprehensive study on the classification and detection of Silicosis-related lung inflammation. Our main contributions include

Intuitions of Machine Learning Researchers about Transfer Learning for Medical Image Classification

SafetyDGX agent

arXiv:2510.00902v2 Announce Type: replace Abstract: Transfer learning is crucial for medical imaging, yet the selection of source datasets often relies on researchers' intuition rather than systematic

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams

Model ReleasesDGX agent

arXiv:2605.27074v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) achieve strong performance on reactive question answering, but real-world streaming assistants require p

Is an Image Also Worth 16x16=256 Superpixels? A Framework for Attentional Image Classification

ResearchDGX agent

arXiv:2605.27144v1 Announce Type: new Abstract: Superpixel-based image classification has traditionally leveraged graph neural networks (GNNs) for processing irregular image representations. Recent ad

ISTASTrack: Bridging ANN and SNN via ISTA Adapter for RGB-Event Tracking

ResearchDGX agent

arXiv:2509.09977v2 Announce Type: replace Abstract: RGB-Event tracking has become a promising trend in visual object tracking to leverage the complementary strengths of both RGB images and dynamic spi

JLT: Clean-Latent Prediction in Latent Diffusion Transformers

ResearchDGX agent

arXiv:2605.27102v1 Announce Type: new Abstract: Flow matching with clean-data prediction has shown that regressing the clean point can exploit low-dimensional structure more effectively than predictin

Joint 2D-3D Segmentation and Association in Street-level Imaging

TutorialsDGX agent

arXiv:2605.26725v1 Announce Type: new Abstract: Accurate interpretation of street-level imagery is essential for large-scale urban mapping and the creation of Spatial Digital Twin (SDT) environments.

Joint Instance Segmentation and Geometric Attribute Regression for Roof Structures in Aerial Imagery

ResearchDGX agent

arXiv:2605.26370v1 Announce Type: new Abstract: We present a method for jointly predicting instance-level roof segment masks together with three continuous geometric attributes -- building height, roo

LDP-Slicing: Local Differential Privacy for Images via Randomized Bit-Plane Slicing

Local AiDGX agent

arXiv:2603.03711v3 Announce Type: replace Abstract: Local Differential Privacy (LDP) is the gold standard trust model for privacy-preserving machine learning by guaranteeing privacy at the data source

Learning Reference-Guided Exposure Correction with Hybrid Illumination Characteristics

ResearchDGX agent

arXiv:2605.26729v1 Announce Type: new Abstract: We present HICNet, a reference-guided exposure correction framework. A lightweight, content-agnostic encoder distills each image into a compact illumina

Leveraging Text-to-Image Diffusion Models for Unsupervised Visual Object Tracking

TutorialsDGX agent

arXiv:2605.26933v1 Announce Type: new Abstract: Unsupervised visual object tracking is a challenging task that requires following arbitrary targets in videos without training on ground-truth annotatio

Leveraging Visual Signals for Robust Token-Level Uncertainty in Vision-Language Generation

ApplicationsDGX agent

arXiv:2605.27136v1 Announce Type: new Abstract: Uncertainty quantification (UQ) remains a critical challenge in Large Vision Language Models (LVLMs) for reliable predictions and real-world deployment.

Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models

TutorialsDGX agent

arXiv:2506.11253v2 Announce Type: replace Abstract: Machine unlearning removes certain training data points and their influence from AI models (e.g., when a data owner revokes their consent to allow m

LiM-YOLO: Less is More with Pyramid Level Shift for Ship Detection in Optical Remote Sensing

ResearchDGX agent

arXiv:2512.09700v3 Announce Type: replace Abstract: General-purpose object detectors face fundamental structural limitations when applied to ship detection in satellite imagery, where the ship scale d

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV

Model ReleasesDGX agent

arXiv:2605.26244v1 Announce Type: new Abstract: Audio-visual generation is rapidly advancing from short clips to minute-long content, while existing evaluation protocols remain largely confined to sho

LongCat-Video-Avatar 1.5 Technical Report

Model ReleasesDGX agent

arXiv:2605.26486v1 Announce Type: new Abstract: Despite advances in audio-driven video generation, achieving commercial-grade stability remains challenging. We present LongCat-Video-Avatar 1.5, an upg

LuxRemix: Lighting Decomposition and Remixing for Indoor Scenes

ApplicationsDGX agent

arXiv:2601.15283v2 Announce Type: replace Abstract: We present a novel approach for interactive light editing in indoor scenes from a single multi-view scene capture. Our method leverages a generative

Memory-Distilled Selection for Noise-Robust Anomaly Detection

ResearchDGX agent

arXiv:2605.26676v1 Announce Type: new Abstract: Anomaly detection (AD) under data contamination is critical for deploying unsupervised defect detection in industrial environments, where curating perfe

METATR: A Multilingual, Evolving Benchmark for Automatic Text Recognition

Model ReleasesDGX agent

arXiv:2605.26712v1 Announce Type: new Abstract: Benchmarks that reflect the diversity and complexity of real-world documents are essential for accurately evaluating Automatic Text Recognition (ATR) sy

Model discovery for dynamical systems with complex-valued product units

Model ReleasesDGX agent

arXiv:2605.27158v1 Announce Type: new Abstract: Discovering the governing equations of a dynamical system from observed trajectories provides deeper insight into its structure than mere prediction of

MotionPRO: Exploring the Role of Pressure in Human MoCap and Beyond

ApplicationsDGX agent

arXiv:2504.05046v2 Announce Type: replace Abstract: Existing human Motion Capture (MoCap) methods mostly focus on the visual similarity while neglecting the physical plausibility. As a result, downstr

MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale

Model ReleasesDGX agent

arXiv:2605.27235v1 Announce Type: new Abstract: Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of generated visual content, an

MSCGC-KAN: Multi-scale Causal Graph Convolution and Kolmogorov-Arnold Feature Mapping for EEG Emotion Recognition

ResearchDGX agent

arXiv:2605.26624v1 Announce Type: new Abstract: Electroencephalogram (EEG)-based emotion recognition is an important affective computing task, and recent EEG foundation models provide useful generic r

Multi-Modal Building Inspection via Perceiver IO Fusion of Satellite and Street-Level Imagery

ResearchDGX agent

arXiv:2605.26381v1 Announce Type: new Abstract: We present a multi-modal classification framework that fuses satellite and street-level imagery through a Perceiver IO architecture operating on spatial

MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation

SafetyDGX agent

arXiv:2602.09878v2 Announce Type: replace Abstract: World-model-based imagine-then-act becomes a promising paradigm for robotic manipulation, yet existing approaches typically support either purely im

Natural Human Motion Recovery by Aligning High-Order Temporal Dynamics from Monocular Videos

ResearchDGX agent

arXiv:2605.26879v1 Announce Type: new Abstract: Human motion recovered from monocular videos often appears overly smooth or dynamically inconsistent, even when joint positions are numerically accurate

NeR-SC: Adapting Neural Video Representation to Screen Content

ApplicationsDGX agent

arXiv:2605.27024v1 Announce Type: new Abstract: Implicit neural representations have emerged as a promising paradigm for video compression, with recent methods achieving competitive performance on nat

No Data? No Problem: Robust Vision-Tabular Learning with Missing Values

ApplicationsDGX agent

arXiv:2512.19602v2 Announce Type: replace Abstract: Large-scale medical biobanks provide imaging data complemented by extensive tabular information, such as clinical measurements or demographics. Howe

Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos

ResearchDGX agent

arXiv:2605.26232v1 Announce Type: new Abstract: Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, dept

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding

Model ReleasesDGX agent

arXiv:2605.26584v1 Announce Type: new Abstract: Omnimodal large language models enable unified audio video understanding, but long joint token sequences make inference costly, and existing benchmarks

Object Pose and Shape Estimation for Grasping: Does it Work?

ResearchDGX agent

arXiv:2605.26944v1 Announce Type: cross Abstract: The problem of object pose and shape estimation has seen key advancements lately. Encoder-decoder (e.g., SAM3D, LRM, CRISP) and diffusion-based models

ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection

Model ReleasesDGX agent

arXiv:2508.01253v2 Announce Type: replace Abstract: Existing studies typically investigate domain shift and category shift as independent problems, however, in real-world scenarios, the two types of s

OmniGF: A Dual-Branch Vision-Language Framework for Unified Gaze Following

Local AiDGX agent

arXiv:2605.26399v1 Announce Type: new Abstract: Understanding human gaze behavior is essential for complex scene comprehension and human-computer interaction. Traditional gaze following models are typ

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation

Model ReleasesDGX agent

arXiv:2605.26641v1 Announce Type: new Abstract: Unified multimodal embedding spaces have become the standard interface for cross-modal retrieval and multimodal RAG, and recent audio-video-text (AVT) e

On the Robustness of Machine Unlearning for Vision-Language Models

ResearchDGX agent

arXiv:2605.26992v1 Announce Type: new Abstract: Vision-language models (VLMs) may memorize undesirable information from training data, motivating growing interest in machine unlearning. In this work,

Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning

ResearchDGX agent

arXiv:2605.26761v1 Announce Type: new Abstract: Multimodal instruction tuning is the de facto recipe for adapting vision language models (VLMs), yet instruction data are highly redundant, making data

OSMa-Bench++: Toward Open-Ended Benchmarking of Semantic Mapping for Manipulation with Prompt-Generated Synthetic Scenes

Model ReleasesDGX agent

arXiv:2605.26831v1 Announce Type: new Abstract: Semantic mapping methods are increasingly used as intermediate scene representations for downstream robotic reasoning and manipulation, yet their evalua

PARE: Pruning and Adaptive Routing for Efficient Video Generation

ResearchDGX agent

arXiv:2605.27336v1 Announce Type: new Abstract: Video Diffusion Transformers (DiTs) generate high-quality videos but demand substantial compute due to wide blocks, deep architectures, and iterative sa

PILOT: A Data-Free Continual Learning Approach for Real-Time Semantic Segmentation via Boundary Guidance

TutorialsDGX agent

arXiv:2605.27128v1 Announce Type: new Abstract: Real-time semantic segmentation models offer an excellent balance between accuracy and inference speed. However, deploying these models in dynamic real

PlayClass: Automated Play Behaviour Classification in Poultry

ResearchDGX agent

arXiv:2605.27304v1 Announce Type: new Abstract: Automated monitoring of animal welfare has largely targeted negative indicators, leaving positive welfare behaviours such as play underexplored. To addr

PRBench: A Standardized Probabilistic Robustness Benchmark

Model ReleasesDGX agent

arXiv:2511.01724v3 Announce Type: replace Abstract: Deep learning models are notoriously vulnerable to imperceptible perturbations. Most existing research centers on adversarial robustness (AR), which

Prototyping an End-to-End Multi-Modal Tiny-CNN for Cardiovascular Sensor Patches

Local AiDGX agent

arXiv:2510.18668v2 Announce Type: replace-cross Abstract: The vast majority of cardiovascular diseases may be preventable if early signs and risk factors are detected. Cardiovascular monitoring with b

Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation

ResearchDGX agent

arXiv:2507.16116v2 Announce Type: replace Abstract: The rapid advancement of video diffusion models has been hindered by fundamental limitations in temporal modeling, particularly the rigid synchroniz

PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation

SafetyDGX agent

arXiv:2508.02806v3 Announce Type: replace Abstract: Recently, a significant improvement in the accuracy of 3D human pose estimation has been achieved by combining convolutional neural networks (CNNs)

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning

ResearchDGX agent

arXiv:2605.27318v1 Announce Type: new Abstract: Video spatial reasoning requires accumulating viewpoint-dependent evidence over time while retaining information useful to the question being asked. Exi

R^3: 3D Reconstruction via Relative Regression

ResearchDGX agent

arXiv:2605.26519v1 Announce Type: new Abstract: Recent feed-forward geometry foundation models have demonstrated impressive generalization by recovering depth and poses in a single forward pass. Howev

RadarSim: Simulating Single-Chip Radar via Multimodal Neural Fields

ResearchDGX agent

arXiv:2605.26328v1 Announce Type: new Abstract: Radars are an ideal complement to cameras: both are inexpensive, solid-state sensors, with cameras offering fine angular resolution, while radars provid

Re-M3Dr: Rebalanced MultiModal Mean Deviation Regression

ResearchDGX agent

arXiv:2605.26513v1 Announce Type: new Abstract: Mean Deviation (MD) is a critical metric for assessing visual field loss in ophthalmology. While previous work has focused solely on predicting MD from

Receipt Replay OOD: A Small Benchmark for Screen Replay Detection Under Domain Shift

Model ReleasesDGX agent

arXiv:2605.26855v1 Announce Type: new Abstract: Public datasets such as DLC-2021, SynID, and KID34K have significantly contributed to research on presentation attack detection for identity documents,

Resolving Ambiguity in Composed Image Retrieval via Calibrated Interaction

Model ReleasesDGX agent

arXiv:2605.24634v2 Announce Type: replace Abstract: Composed image retrieval (CIR) searches a corpus with a reference image and a text describing how to modify it. Despite rapid progress from triplet-

Revealing the core dimensions underlying representations in brains, behavior and AI

ResearchDGX agent

arXiv:2605.26921v1 Announce Type: new Abstract: The study of representations is widespread across fields, including neuroscience, psychology, and artificial intelligence. While representations are oft

REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization

Local AiDGX agent

arXiv:2605.26861v1 Announce Type: new Abstract: Image geo-localization aims to determine where a photograph was taken, a task that often requires more than recognizing visible landmarks. Human experts

RoadGIE: Towards A Global-Scale Aerial Benchmark for Generalizable Interactive Road Extraction

Model ReleasesDGX agent

arXiv:2605.26862v1 Announce Type: new Abstract: Accurate road segmentation from aerial imagery is fundamental to many geospatial applications. However, existing datasets often suffer from limited scen

RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation

ResearchDGX agent

arXiv:2605.26241v1 Announce Type: new Abstract: Success in generative modeling across language, image, and video demonstrates that large, well-curated datasets are the key driver for building capable

Scheduled Style Injection: Expanding the Style-Content Pareto Frontier in Training-Free Diffusion-based Style Transfer

Model ReleasesDGX agent

arXiv:2605.26538v1 Announce Type: new Abstract: Style transfer with pre-trained diffusion models has advanced rapidly, but a core question remains underexplored: where in the model should style inject

SCKAN: Structural Consensus-based KAN Prototype Learning for Semi-Supervised Pancreas Segmentation

SafetyDGX agent

arXiv:2605.27032v1 Announce Type: new Abstract: Accurate pancreas segmentation is critical for early cancer diagnosis, where annotation scarcity necessitates Semi-Supervised Learning (SSL). However, d

← Previous
1…113114115116117…211
Next →