AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Safety

SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy

DGX agent

arXiv:2603.29962v3 Announce Type: replace Abstract: Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes. C

safetyarxiv-cs-cv
5 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Applications

SVGS: Enhancing Gaussian Splatting Using Primitives with Spatially Varying Colors

DGX agent

arXiv:2411.18966v2 Announce Type: replace Abstract: Gaussian Splatting demonstrates impressive results in multi-view reconstruction based on Gaussian explicit representations. However, the current Gau

applicationsarxiv-cs-cv
5 May 2026
Safety

SwiftPie: Lightning-fast Subject-driven Image Personalization via One step Diffusion

DGX agent

arXiv:2605.01510v1 Announce Type: new Abstract: Diffusion models have achieved remarkable success in high-quality image synthesis, sparking interest in image-guided generation tasks such as subject-dr

safetyarxiv-cs-cv
5 May 2026
Agents

Synergistic Perception and Generative Recomposition: A Multi-Agent Orchestration for Expert-Level Building Inspection

DGX agent

arXiv:2603.20143v2 Announce Type: replace Abstract: Building facade defect inspection is fundamental to structural health monitoring and sustainable urban maintenance, yet it remains a formidable chal

agentsarxiv-cs-cv
5 May 2026
Safety

SynPAIN: A Synthetic Dataset of Pain and Non-Pain Facial Expressions

DGX agent

arXiv:2507.19673v3 Announce Type: replace Abstract: Accurate pain assessment in patients with limited ability to communicate, such as older adults with severe dementia, represents a critical healthcar

safetyarxiv-cs-cv
5 May 2026
Research

Synthetic Designed Experiments for Diagnosing Vision Model Failure

DGX agent

arXiv:2605.00832v1 Announce Type: new Abstract: Current synthetic data pipelines for computer vision generate images without diagnosing what the downstream model actually needs. This open-loop paradig

researcharxiv-cs-cv
5 May 2026
Model Releases

Task-Driven Subspace Decomposition for Knowledge Sharing and Isolation in LoRA-based Continual Learning

DGX agent

arXiv:2603.00191v2 Announce Type: replace-cross Abstract: Continual Learning (CL) requires models to sequentially adapt to new tasks without forgetting old knowledge. Recently, Low-Rank Adaptation (Lo

model-releasesarxiv-cs-cv
5 May 2026
Safety

Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective

DGX agent

arXiv:2506.01097v2 Announce Type: replace Abstract: Existing Multimodal Large Language Models (MLLMs) process a large number of visual tokens, leading to significant computational costs and inefficien

safetyarxiv-cs-cv
5 May 2026
Research

Temporally Consistent Object 6D Pose Estimation for Robot Control

DGX agent

arXiv:2605.02708v1 Announce Type: cross Abstract: Single-view RGB object pose estimators have reached a level of precision and efficiency that makes them good candidates for vision-based robot control

researcharxiv-cs-cv
5 May 2026
Tutorials

TemPose-TF-ASF: Two-Stage Bidirectional Stroke Context Fusion for Badminton Stroke Classification

DGX agent

arXiv:2605.02558v1 Announce Type: new Abstract: Accurate badminton stroke prediction is crucial for fine-grained sports analysis and tactical decision support. However, existing methods struggle to mo

tutorialsarxiv-cs-cv
5 May 2026
Research

The Multi-View Paradigm Shift in MRI Radiomics: Predicting MGMT Methylation in Glioblastoma

DGX agent

arXiv:2512.22331v2 Announce Type: replace Abstract: Non-invasive inference of molecular tumor characteristics from medical imaging is a central goal of radiogenomics, particularly in glioblastoma (GBM

researcharxiv-cs-cv
5 May 2026
Safety

Thermal Imaging for Contactless Cardiorespiratory and Sudomotor Response Monitoring

DGX agent

arXiv:2602.12361v2 Announce Type: replace Abstract: Human-machine interfaces in industrial automation need sensing modules that monitor operator actions and physiological state. This is important in f

safetyarxiv-cs-cv
5 May 2026
Local Ai

TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images

DGX agent

arXiv:2603.07119v2 Announce Type: replace Abstract: Recent text-to-image models have improved global realism, but text rendering remains a persistent failure mode: images may look convincing overall,

local-aiarxiv-cs-cv
5 May 2026
Model Releases

TOC-SR: Task-Optimal Compact diffusion for Image Super Resolution

DGX agent

arXiv:2605.02767v1 Announce Type: new Abstract: Diffusion models have recently demonstrated strong performance for image restoration tasks, including super-resolution. However, their large model size

model-releasesarxiv-cs-cv
5 May 2026
Safety

Toward a Scientific Discovery Engine for Weather and Climate Data: A Visual Analytics Workbench for Embedding-Based Exploration

DGX agent

arXiv:2605.00972v1 Announce Type: cross Abstract: Earth system science is producing increasingly large, high-dimensional datasets from physics based Earth system models to AI-based weather and climate

safetyarxiv-cs-cv
5 May 2026
Model Releases

Toward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization

DGX agent

arXiv:2605.02223v1 Announce Type: cross Abstract: Recent advances in voice cloning and text-to-speech synthesis have made partial speech manipulation - where an adversary replaces a few words within a

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Towards High Fidelity Face Swapping: A Comprehensive Survey and New Benchmark

DGX agent

arXiv:2605.00883v1 Announce Type: new Abstract: Face swapping has witnessed significant progress in recent years, largely driven by advances in deep generative models such as GANs and diffusion models

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Towards Lightest Low-Light Image Enhancement Architecture for Mobile Devices

DGX agent

arXiv:2507.04277v2 Announce Type: replace Abstract: Real-time low-light image enhancement on mobile and embedded devices requires models that balance visual quality and computational efficiency. Exist

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Towards Visual Query Localization in the 3D World

DGX agent

arXiv:2605.01498v1 Announce Type: new Abstract: Visual query localization (VQL) aims to predict the spatio-temporal response of the most recent occurrence in a sequence given a query. Currently, most

model-releasesarxiv-cs-cv
5 May 2026
Safety

Training-Free Adaptive 360-degree Video Streaming via Semantic Potential Fields

DGX agent

arXiv:2603.20999v2 Announce Type: replace-cross Abstract: Adaptive 360{eg} video streaming for teleoperation faces two coupled challenges: viewport prediction under uncertain gaze patterns and bitrate

safetyarxiv-cs-cv
5 May 2026
Tutorials

TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigation

DGX agent

arXiv:2605.01700v1 Announce Type: new Abstract: Existing zero-shot Object Goal Navigation (ObjectNav) methods often exploit commonsense knowledge from large language or vision-language models to guide

tutorialsarxiv-cs-cv
5 May 2026
Safety

TrajShield: Trajectory-Level Safety Mediation for Defending Text-to-Video Models Against Jailbreak Attacks

DGX agent

arXiv:2605.01761v1 Announce Type: new Abstract: Text-to-Video (T2V) models have demonstrated remarkable capability in generating temporally coherent videos from natural language prompts, yet they also

safetyarxiv-cs-cv
5 May 2026
Research

TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness

DGX agent

arXiv:2506.20588v2 Announce Type: replace Abstract: The increasing ubiquity of video content and the corresponding demand for efficient access to meaningful information have elevated video summarizati

researcharxiv-cs-cv
5 May 2026
Applications

TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning

DGX agent

arXiv:2605.01659v1 Announce Type: new Abstract: The rapid growth of video content across domains such as surveillance, education, and social media has made efficient content understanding increasingly

applicationsarxiv-cs-cv
5 May 2026
Model Releases

TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation

DGX agent

arXiv:2605.00907v1 Announce Type: new Abstract: Large language models (LLMs) and multimodal large models (MLLMs) are increasingly used for transportation tasks such as regulation question answering, t

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Triple Spectral Fusion for Sensor-based Human Activity Recognition

DGX agent

arXiv:2605.02743v1 Announce Type: cross Abstract: The field of sensor-based human activity recognition (HAR) mainly uses posture, motion and context data of Inertial Measurement Units (IMUs) to identi

model-releasesarxiv-cs-cv
5 May 2026
Research

TT4D: A Pipeline and Dataset for Table Tennis 4D Reconstruction From Monocular Videos

DGX agent

arXiv:2605.01234v1 Announce Type: new Abstract: We present TT4D, a large-scale, high-fidelity table tennis dataset. It provides 140+ hours of reconstructed singles and doubles gameplay from monocular

researcharxiv-cs-cv
5 May 2026
Model Releases

Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Environments

DGX agent

arXiv:2510.04142v2 Announce Type: replace Abstract: This paper identifies a critical yet underexplored challenge in reasoning alignment from multiple multi-modal large language models (MLLMs): In non-

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

TwistNet-2D: Learning Second-Order Channel Interactions via Spiral Twisting for Texture Recognition

DGX agent

arXiv:2602.07262v3 Announce Type: replace Abstract: Second-order feature statistics are central to texture recognition, yet existing mechanisms exhibit a structural tension: bilinear pooling and Gram

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Two-Pass Zero-Shot Temporal-Spatial Grounding of Rare Traffic Events in Surveillance Video

DGX agent

arXiv:2605.01512v1 Announce Type: new Abstract: Grounding traffic accidents in real CCTV footage is a rare-event problem where training on labeled accident video is often prohibited, yet accurate join

model-releasesarxiv-cs-cv
5 May 2026
Safety

UCATSC: Uncertainty-Aware Constrained Traffic Signal Control Under Vision-Based Partial Observability

DGX agent

arXiv:2602.07784v3 Announce Type: replace Abstract: Camera-based adaptive traffic signal control is inherently partially observable: detections can be missed, vehicle speeds and distances can be noisy

safetyarxiv-cs-cv
5 May 2026
Safety

Ultrasound Vision-Language Alignment via Contrastive Learning

DGX agent

arXiv:2605.02126v1 Announce Type: new Abstract: Ultrasound foundation models have achieved strong performance on structured prediction tasks but remain exclusively vision-based, limiting zero-shot and

safetyarxiv-cs-cv
5 May 2026
Model Releases

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis

DGX agent

arXiv:2605.00826v1 Announce Type: cross Abstract: Text-to-video retrieval enables users to find relevant video content using natural language queries, a task that has grown increasingly important with

model-releasesarxiv-cs-cv
5 May 2026
Local Ai

UnGAP: Uncertainty-Guided Affine Prompting for Real-Time Crack Segmentation

DGX agent

arXiv:2605.02380v1 Announce Type: new Abstract: Real-time crack segmentation is vital for structural health monitoring but is plagued by aleatoric uncertainties arising from varying lighting, blur, an

local-aiarxiv-cs-cv
5 May 2026
Safety

Unified Map Prior Encoder for Mapping and Planning

DGX agent

arXiv:2605.02762v1 Announce Type: new Abstract: Online mapping and end-to-end (E2E) planning in autonomous driving remain largely sensor-centric, leaving rich map priors, including HD/SD vector maps,

safetyarxiv-cs-cv
5 May 2026
Research

Unifying Deep Stochastic Processes for Image Enhancement

DGX agent

arXiv:2605.01568v1 Announce Type: new Abstract: Deep stochastic processes have recently become a central paradigm for image enhancement, with many methods explicitly conditioning the stochastic trajec

researcharxiv-cs-cv
5 May 2026
Research

Unsupervised Learning of Robust Spectral Shape Matching

DGX agent

arXiv:2304.14419v2 Announce Type: replace Abstract: We propose a novel learning-based approach for robust 3D shape matching. Our method builds upon deep functional maps and can be trained in a fully u

researcharxiv-cs-cv
5 May 2026
Research

Validation of an AI-based end-to-end model for prostate pathology using long-term archived routine samples

DGX agent

arXiv:2605.02614v1 Announce Type: new Abstract: Artificial intelligence (AI) is becoming a clinical tool for prostate pathology, but generalization across variations in sample preparation and preserva

researcharxiv-cs-cv
5 May 2026
Research

Validation of Whole-Slide Foundation Models for Image Retrieval in TCGA Data

DGX agent

arXiv:2605.00902v1 Announce Type: new Abstract: Foundation models are reshaping computational histopathology, yet their value for whole-slide image retrieval relative to strong patch-based and supervi

researcharxiv-cs-cv
5 May 2026
Research

Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision

DGX agent

arXiv:2604.04934v2 Announce Type: replace Abstract: We present Vanast, a unified framework that generates garment-transferred human animation videos directly from a single human image, garment images,

researcharxiv-cs-cv
5 May 2026
Model Releases

VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector Animation

DGX agent

arXiv:2605.01517v1 Announce Type: new Abstract: Scalable Vector Graphics (SVG) animation generation is pivotal for professional design due to their structural editability and resolution independence.

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models

DGX agent

arXiv:2605.01662v1 Announce Type: new Abstract: Large vision-language models (VLMs) have advanced multimodal tasks such as video question answering (QA). However, VLMs face the challenge of selecting

model-releasesarxiv-cs-cv
5 May 2026
Tutorials

Video Generation with Predictive Latents

DGX agent

arXiv:2605.02134v1 Announce Type: new Abstract: Video Variational Autoencoder (VAE) enables latent video generative modeling by mapping the visual world into compact spatiotemporal latent spaces, impr

tutorialsarxiv-cs-cv
5 May 2026
Safety

VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation

DGX agent

arXiv:2601.23286v2 Announce Type: replace Abstract: While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, o

safetyarxiv-cs-cv
5 May 2026
Model Releases

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition

DGX agent

arXiv:2605.02834v1 Announce Type: new Abstract: Videos are unique in their ability to capture actions which transcend multiple frames. Accordingly, for many years action recognition was the quintessen

model-releasesarxiv-cs-cv
5 May 2026
Research

ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking

DGX agent

arXiv:2605.02638v1 Announce Type: new Abstract: Cross-view Referring Multi-Object Tracking (CRMOT) aims to track multiple objects specified by natural language across multiple camera views, with globa

researcharxiv-cs-cv
5 May 2026
Hardware

ViM-Q: Scalable Algorithm-Hardware Co-Design for Vision Mamba Model Inference on FPGA

DGX agent

arXiv:2605.01935v1 Announce Type: cross Abstract: Vision Mamba (ViM) models offer a compelling efficiency advantage over Transformers by leveraging the linear complexity of State Space Models (SSMs),

hardwarearxiv-cs-cv
5 May 2026
Research

Virtual Scanning for NSCLC Histology: Investigating the Discriminatory Power of Synthetic PET

DGX agent

arXiv:2605.02746v1 Announce Type: new Abstract: Accurate histological differentiation between adenocarcinoma (ADC) and squamous cell carcinoma (SCC) is critical for personalized treatment in non-small

researcharxiv-cs-cv
5 May 2026
← Previous
1…199200201202203…261
Next →