AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
5 May 2026

SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy

SafetyDGX agent

arXiv:2603.29962v3 Announce Type: replace Abstract: Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes. C

SVGS: Enhancing Gaussian Splatting Using Primitives with Spatially Varying Colors

ApplicationsDGX agent

arXiv:2411.18966v2 Announce Type: replace Abstract: Gaussian Splatting demonstrates impressive results in multi-view reconstruction based on Gaussian explicit representations. However, the current Gau

SwiftPie: Lightning-fast Subject-driven Image Personalization via One step Diffusion

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.01510v1 Announce Type: new Abstract: Diffusion models have achieved remarkable success in high-quality image synthesis, sparking interest in image-guided generation tasks such as subject-dr

Synergistic Perception and Generative Recomposition: A Multi-Agent Orchestration for Expert-Level Building Inspection

AgentsDGX agent

arXiv:2603.20143v2 Announce Type: replace Abstract: Building facade defect inspection is fundamental to structural health monitoring and sustainable urban maintenance, yet it remains a formidable chal

SynPAIN: A Synthetic Dataset of Pain and Non-Pain Facial Expressions

SafetyDGX agent

arXiv:2507.19673v3 Announce Type: replace Abstract: Accurate pain assessment in patients with limited ability to communicate, such as older adults with severe dementia, represents a critical healthcar

Synthetic Designed Experiments for Diagnosing Vision Model Failure

ResearchDGX agent

arXiv:2605.00832v1 Announce Type: new Abstract: Current synthetic data pipelines for computer vision generate images without diagnosing what the downstream model actually needs. This open-loop paradig

Task-Driven Subspace Decomposition for Knowledge Sharing and Isolation in LoRA-based Continual Learning

Model ReleasesDGX agent

arXiv:2603.00191v2 Announce Type: replace-cross Abstract: Continual Learning (CL) requires models to sequentially adapt to new tasks without forgetting old knowledge. Recently, Low-Rank Adaptation (Lo

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective

SafetyDGX agent

arXiv:2506.01097v2 Announce Type: replace Abstract: Existing Multimodal Large Language Models (MLLMs) process a large number of visual tokens, leading to significant computational costs and inefficien

Temporally Consistent Object 6D Pose Estimation for Robot Control

ResearchDGX agent

arXiv:2605.02708v1 Announce Type: cross Abstract: Single-view RGB object pose estimators have reached a level of precision and efficiency that makes them good candidates for vision-based robot control

TemPose-TF-ASF: Two-Stage Bidirectional Stroke Context Fusion for Badminton Stroke Classification

TutorialsDGX agent

arXiv:2605.02558v1 Announce Type: new Abstract: Accurate badminton stroke prediction is crucial for fine-grained sports analysis and tactical decision support. However, existing methods struggle to mo

The Multi-View Paradigm Shift in MRI Radiomics: Predicting MGMT Methylation in Glioblastoma

ResearchDGX agent

arXiv:2512.22331v2 Announce Type: replace Abstract: Non-invasive inference of molecular tumor characteristics from medical imaging is a central goal of radiogenomics, particularly in glioblastoma (GBM

Thermal Imaging for Contactless Cardiorespiratory and Sudomotor Response Monitoring

SafetyDGX agent

arXiv:2602.12361v2 Announce Type: replace Abstract: Human-machine interfaces in industrial automation need sensing modules that monitor operator actions and physiological state. This is important in f

TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images

Local AiDGX agent

arXiv:2603.07119v2 Announce Type: replace Abstract: Recent text-to-image models have improved global realism, but text rendering remains a persistent failure mode: images may look convincing overall,

TOC-SR: Task-Optimal Compact diffusion for Image Super Resolution

Model ReleasesDGX agent

arXiv:2605.02767v1 Announce Type: new Abstract: Diffusion models have recently demonstrated strong performance for image restoration tasks, including super-resolution. However, their large model size

Toward a Scientific Discovery Engine for Weather and Climate Data: A Visual Analytics Workbench for Embedding-Based Exploration

SafetyDGX agent

arXiv:2605.00972v1 Announce Type: cross Abstract: Earth system science is producing increasingly large, high-dimensional datasets from physics based Earth system models to AI-based weather and climate

Toward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization

Model ReleasesDGX agent

arXiv:2605.02223v1 Announce Type: cross Abstract: Recent advances in voice cloning and text-to-speech synthesis have made partial speech manipulation - where an adversary replaces a few words within a

Towards High Fidelity Face Swapping: A Comprehensive Survey and New Benchmark

Model ReleasesDGX agent

arXiv:2605.00883v1 Announce Type: new Abstract: Face swapping has witnessed significant progress in recent years, largely driven by advances in deep generative models such as GANs and diffusion models

Towards Lightest Low-Light Image Enhancement Architecture for Mobile Devices

Model ReleasesDGX agent

arXiv:2507.04277v2 Announce Type: replace Abstract: Real-time low-light image enhancement on mobile and embedded devices requires models that balance visual quality and computational efficiency. Exist

Towards Visual Query Localization in the 3D World

Model ReleasesDGX agent

arXiv:2605.01498v1 Announce Type: new Abstract: Visual query localization (VQL) aims to predict the spatio-temporal response of the most recent occurrence in a sequence given a query. Currently, most

Training-Free Adaptive 360-degree Video Streaming via Semantic Potential Fields

SafetyDGX agent

arXiv:2603.20999v2 Announce Type: replace-cross Abstract: Adaptive 360{eg} video streaming for teleoperation faces two coupled challenges: viewport prediction under uncertain gaze patterns and bitrate

TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigation

TutorialsDGX agent

arXiv:2605.01700v1 Announce Type: new Abstract: Existing zero-shot Object Goal Navigation (ObjectNav) methods often exploit commonsense knowledge from large language or vision-language models to guide

TrajShield: Trajectory-Level Safety Mediation for Defending Text-to-Video Models Against Jailbreak Attacks

SafetyDGX agent

arXiv:2605.01761v1 Announce Type: new Abstract: Text-to-Video (T2V) models have demonstrated remarkable capability in generating temporally coherent videos from natural language prompts, yet they also

TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness

ResearchDGX agent

arXiv:2506.20588v2 Announce Type: replace Abstract: The increasing ubiquity of video content and the corresponding demand for efficient access to meaningful information have elevated video summarizati

TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning

ApplicationsDGX agent

arXiv:2605.01659v1 Announce Type: new Abstract: The rapid growth of video content across domains such as surveillance, education, and social media has made efficient content understanding increasingly

TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation

Model ReleasesDGX agent

arXiv:2605.00907v1 Announce Type: new Abstract: Large language models (LLMs) and multimodal large models (MLLMs) are increasingly used for transportation tasks such as regulation question answering, t

Triple Spectral Fusion for Sensor-based Human Activity Recognition

Model ReleasesDGX agent

arXiv:2605.02743v1 Announce Type: cross Abstract: The field of sensor-based human activity recognition (HAR) mainly uses posture, motion and context data of Inertial Measurement Units (IMUs) to identi

TT4D: A Pipeline and Dataset for Table Tennis 4D Reconstruction From Monocular Videos

ResearchDGX agent

arXiv:2605.01234v1 Announce Type: new Abstract: We present TT4D, a large-scale, high-fidelity table tennis dataset. It provides 140+ hours of reconstructed singles and doubles gameplay from monocular

Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Environments

Model ReleasesDGX agent

arXiv:2510.04142v2 Announce Type: replace Abstract: This paper identifies a critical yet underexplored challenge in reasoning alignment from multiple multi-modal large language models (MLLMs): In non-

TwistNet-2D: Learning Second-Order Channel Interactions via Spiral Twisting for Texture Recognition

Model ReleasesDGX agent

arXiv:2602.07262v3 Announce Type: replace Abstract: Second-order feature statistics are central to texture recognition, yet existing mechanisms exhibit a structural tension: bilinear pooling and Gram

Two-Pass Zero-Shot Temporal-Spatial Grounding of Rare Traffic Events in Surveillance Video

Model ReleasesDGX agent

arXiv:2605.01512v1 Announce Type: new Abstract: Grounding traffic accidents in real CCTV footage is a rare-event problem where training on labeled accident video is often prohibited, yet accurate join

UCATSC: Uncertainty-Aware Constrained Traffic Signal Control Under Vision-Based Partial Observability

SafetyDGX agent

arXiv:2602.07784v3 Announce Type: replace Abstract: Camera-based adaptive traffic signal control is inherently partially observable: detections can be missed, vehicle speeds and distances can be noisy

Ultrasound Vision-Language Alignment via Contrastive Learning

SafetyDGX agent

arXiv:2605.02126v1 Announce Type: new Abstract: Ultrasound foundation models have achieved strong performance on structured prediction tasks but remain exclusively vision-based, limiting zero-shot and

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis

Model ReleasesDGX agent

arXiv:2605.00826v1 Announce Type: cross Abstract: Text-to-video retrieval enables users to find relevant video content using natural language queries, a task that has grown increasingly important with

UnGAP: Uncertainty-Guided Affine Prompting for Real-Time Crack Segmentation

Local AiDGX agent

arXiv:2605.02380v1 Announce Type: new Abstract: Real-time crack segmentation is vital for structural health monitoring but is plagued by aleatoric uncertainties arising from varying lighting, blur, an

Unified Map Prior Encoder for Mapping and Planning

SafetyDGX agent

arXiv:2605.02762v1 Announce Type: new Abstract: Online mapping and end-to-end (E2E) planning in autonomous driving remain largely sensor-centric, leaving rich map priors, including HD/SD vector maps,

Unifying Deep Stochastic Processes for Image Enhancement

ResearchDGX agent

arXiv:2605.01568v1 Announce Type: new Abstract: Deep stochastic processes have recently become a central paradigm for image enhancement, with many methods explicitly conditioning the stochastic trajec

Unsupervised Learning of Robust Spectral Shape Matching

ResearchDGX agent

arXiv:2304.14419v2 Announce Type: replace Abstract: We propose a novel learning-based approach for robust 3D shape matching. Our method builds upon deep functional maps and can be trained in a fully u

Validation of an AI-based end-to-end model for prostate pathology using long-term archived routine samples

ResearchDGX agent

arXiv:2605.02614v1 Announce Type: new Abstract: Artificial intelligence (AI) is becoming a clinical tool for prostate pathology, but generalization across variations in sample preparation and preserva

Validation of Whole-Slide Foundation Models for Image Retrieval in TCGA Data

ResearchDGX agent

arXiv:2605.00902v1 Announce Type: new Abstract: Foundation models are reshaping computational histopathology, yet their value for whole-slide image retrieval relative to strong patch-based and supervi

Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision

ResearchDGX agent

arXiv:2604.04934v2 Announce Type: replace Abstract: We present Vanast, a unified framework that generates garment-transferred human animation videos directly from a single human image, garment images,

VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector Animation

Model ReleasesDGX agent

arXiv:2605.01517v1 Announce Type: new Abstract: Scalable Vector Graphics (SVG) animation generation is pivotal for professional design due to their structural editability and resolution independence.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.01662v1 Announce Type: new Abstract: Large vision-language models (VLMs) have advanced multimodal tasks such as video question answering (QA). However, VLMs face the challenge of selecting

Video Generation with Predictive Latents

TutorialsDGX agent

arXiv:2605.02134v1 Announce Type: new Abstract: Video Variational Autoencoder (VAE) enables latent video generative modeling by mapping the visual world into compact spatiotemporal latent spaces, impr

VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation

SafetyDGX agent

arXiv:2601.23286v2 Announce Type: replace Abstract: While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, o

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition

Model ReleasesDGX agent

arXiv:2605.02834v1 Announce Type: new Abstract: Videos are unique in their ability to capture actions which transcend multiple frames. Accordingly, for many years action recognition was the quintessen

ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking

ResearchDGX agent

arXiv:2605.02638v1 Announce Type: new Abstract: Cross-view Referring Multi-Object Tracking (CRMOT) aims to track multiple objects specified by natural language across multiple camera views, with globa

ViM-Q: Scalable Algorithm-Hardware Co-Design for Vision Mamba Model Inference on FPGA

HardwareDGX agent

arXiv:2605.01935v1 Announce Type: cross Abstract: Vision Mamba (ViM) models offer a compelling efficiency advantage over Transformers by leveraging the linear complexity of State Space Models (SSMs),

Virtual Scanning for NSCLC Histology: Investigating the Discriminatory Power of Synthetic PET

ResearchDGX agent

arXiv:2605.02746v1 Announce Type: new Abstract: Accurate histological differentiation between adenocarcinoma (ADC) and squamous cell carcinoma (SCC) is critical for personalized treatment in non-small

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

Model ReleasesDGX agent

arXiv:2605.01391v1 Announce Type: new Abstract: Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute

Visual Chart Representations for Cryptocurrency Regime Prediction: A Systematic Deep Learning Study

ResearchDGX agent

arXiv:2605.00875v1 Announce Type: new Abstract: Technical traders have long relied on visual analysis of candlestick charts to identify market patterns and predict price movements. While deep learning

Visual Implicit Autoregressive Modeling

Model ReleasesDGX agent

arXiv:2605.01220v1 Announce Type: new Abstract: Visual Autoregressive Modeling (VAR) based on next-scale prediction achieves strong generation quality, but their explicit deep stacks fix the amount of

VoxAfford: Multi-Scale Voxel-Token Fusion for Open-Vocabulary 3D Affordance Detection

ResearchDGX agent

arXiv:2605.01365v1 Announce Type: new Abstract: Open-vocabulary 3D affordance detection requires localizing interaction regions on point clouds given novel affordance descriptions. Recent methods exte

VRGaussianAvatar: Integrating 3D Gaussian Avatars into VR

ResearchDGX agent

arXiv:2602.01674v2 Announce Type: replace Abstract: We present VRGaussianAvatar, an integrated system that enables real-time full-body 3D Gaussian Splatting (3DGS) avatars in virtual reality using onl

Watch Your Step: Information Injection in Diffusion Models via Shadow Timestep Embedding

ResearchDGX agent

arXiv:2605.00935v1 Announce Type: cross Abstract: Diffusion models have become the foundation of modern generative systems, with most research focusing primarily on improving generation efficiency and

When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2605.00911v1 Announce Type: new Abstract: Industrial Retrieval-Augmented Generation (RAG) systems depend on optical character recognition (OCR) to transform visual documents into text. Existing

When Less Is More: Simplicity Beats Complexity for Physics-Constrained InSAR Phase Unwrapping

Model ReleasesDGX agent

arXiv:2605.00896v1 Announce Type: new Abstract: Operational phase unwrapping is the primary computational bottleneck in InSAR-based volcanic and seismic monitoring. We challenge the industry trend of

When To Adapt? Adapting the Model or Data in Federated Medical Imaging

ResearchDGX agent

arXiv:2605.00892v1 Announce Type: new Abstract: Federated learning enables collaborative model training across medical institutions without sharing raw data, but its performance is often limited by do

WILD SAM: A Simulated-and-Real Data Augmentation for Autonomous Driving Perception under Challenging Weather

Model ReleasesDGX agent

arXiv:2605.01081v1 Announce Type: new Abstract: The performance of state-of-the-art object detectors degrades significantly under adverse weather, causing a safety-critical domain shift problem for au

WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild

Model ReleasesDGX agent

arXiv:2605.01018v1 Announce Type: new Abstract: Using multimodal foundation models to analyze table images is a high-value yet challenging application in consumer and enterprise scenarios. Despite its

X2SAM: Any Segmentation in Images and Videos

Model ReleasesDGX agent

arXiv:2605.00891v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong image-level visual understanding and reasoning, yet their pixel-level perception acros

← Previous
1…159160161162163…209
Next →