AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
11 Jun 2026

LAST: Bridging Vision-Language and Action Manifolds via Gromov-Wasserstein Alignment

SafetyDGX agent

arXiv:2606.11221v1 Announce Type: new Abstract: We take a Gromov-Wasserstein perspective on Vision-Language-Action (VLA) learning, where the goal is to make the relational geometry of action represent

Learning Instance-Adaptive Low-Rank Orthogonal Subspaces for Clothes-Changing Person Re-Identification

SafetyDGX agent

arXiv:2606.11661v1 Announce Type: new Abstract: Clothes-changing person re-identification (CC-ReID) aims to recognize individuals despite drastic appearance changes caused by clothing variation. While

MFEN:Multi-Frequency Expert Network for Visible-Infrared Person Re-ID

ResearchDGX agent

arXiv:2606.12051v1 Announce Type: new Abstract: Visible-infrared person re-identification (VI-ReID) is challenging due to the large modality discrepancy between visible and infrared images. We contend


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

MLT-Dedup: Efficient Large-Scale Online Video Deduplication via Multi-Level Representations and Spatial-Temporal Matching

SafetyDGX agent

arXiv:2606.12215v1 Announce Type: new Abstract: The explosive growth of user-generated video content on online platforms is accompanied by the emergence of numerous near-duplicate videos--videos that

Motion Reinforces Appearance: RGB-Skeleton Gated Residual Fusion for Micro-Gesture Online Recognition

Local AiDGX agent

arXiv:2606.11645v1 Announce Type: new Abstract: Micro-gesture analysis attracts increasing attention for inferring spontaneous emotion from subtle body movements. Micro-gesture online recognition, whi

NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization

ResearchDGX agent

arXiv:2606.11363v1 Announce Type: new Abstract: Vector quantization is central to modern generative modeling pipelines, but large-codebook VQ models often suffer from codebook collapse. We identify en

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning

Model ReleasesDGX agent

arXiv:2606.11602v1 Announce Type: new Abstract: Audio-visual Generalized Zero-shot Learning (AV-GZSL) is a challenging task that aims to classify both seen and unseen objects or scenes by integrating

OSCS-SupCon: Orthogonal Sigmoid-based Common and Style Supervised Contrastive Learning for Robust Feature Disentanglement

Model ReleasesDGX agent

arXiv:2606.11233v1 Announce Type: new Abstract: Supervised Contrastive Learning (SupCon) has achieved strong performance by explicitly modeling pairwise relationships among samples. However, existing

Parameter-Efficient Adapter Tuning for Tabular-Image Multimodal Learning

Model ReleasesDGX agent

arXiv:2606.11682v1 Announce Type: new Abstract: Tabular-image multimodal learning aims to improve predictive modeling by jointly using structured tabular attributes and visual data. Although pretraine

ParseFixer: An Agentic Framework for Document Parsing via Selective Multimodal Correction

Model ReleasesDGX agent

arXiv:2606.11977v1 Announce Type: new Abstract: In this report, we present our third-place solution for the DataMFM Challenge Track 1: Document Parsing. This track requires models to recover structure

Performance Analysis of YOLOv11 and YOLOv8 for Mixed Traffic Object Detection under Adverse Weather Conditions in Developing Countries

SafetyDGX agent

arXiv:2606.12066v1 Announce Type: new Abstract: In modern vehicular systems, robust performance under harsh conditions has become a critical problem of autonomous driving. Our study delivers a compreh

Periodic-MAE: Periodic Video Masked Autoencoder for rPPG Estimation

Model ReleasesDGX agent

arXiv:2506.21855v2 Announce Type: replace Abstract: In this paper, we propose Periodic-MAE, a self-supervised framework for learning generalizable spatio-temporal representations of periodic physiolog

Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection

ResearchDGX agent

arXiv:2510.08073v2 Announce Type: replace Abstract: AI-generated videos have achieved near-perfect visual realism (e.g., Sora), urgently necessitating reliable detection mechanisms. However, detecting

PIGEON: VLM-Driven Object Navigation via Points of Interest Selection

Local AiDGX agent

arXiv:2511.13207v2 Announce Type: replace-cross Abstract: Object navigation in unseen indoor environments requires agents to perform semantic search under partial observability. Vision-language models

Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding

SafetyDGX agent

arXiv:2606.11838v1 Announce Type: new Abstract: Reward models for text-to-video (T2V) generation guide post-training but often fail at fine-grained semantic alignment. We trace this to two structural

Precision-Aware Illumination-Disentangled Vision Transformer for Spacecraft 6D Pose Estimation

Model ReleasesDGX agent

arXiv:2606.11619v1 Announce Type: new Abstract: Vision sensors provide a lightweight solution for spacecraft proximity operations, but monocular spacecraft 6D pose estimation remains difficult under i

PT-WNO: Point Transformer with Wavelet Neural Operator for 3D Point Cloud Semantic Segmentation

ResearchDGX agent

arXiv:2606.11466v1 Announce Type: new Abstract: Point cloud semantic segmentation requires architectures that capture both fine-grained local geometry and broad global scene structure. Transformer-bas

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding

Model ReleasesDGX agent

arXiv:2606.12125v1 Announce Type: new Abstract: Long-video understanding remains challenging for multimodal large language models, because temporally extended videos often contain thousands of frames

RankVR: Low-Rank Structure Perception and Value Recalibration for Robust Composed Image Retrieval

Model ReleasesDGX agent

arXiv:2606.11689v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) constitutes a pivotal paradigm requiring models to perform joint reasoning on reference images and modification texts. Ho

ReMoT: Reinforcement Learning with Motion Contrast Triplets

Model ReleasesDGX agent

arXiv:2603.00461v3 Announce Type: replace Abstract: We present ReMoT, a unified training paradigm to systematically address the fundamental shortcomings of VLMs in spatio-temporal consistency -- a cri

RSTR: Reducing SpatioTemporal Redundancy in Diffusion Transformers

Model ReleasesDGX agent

arXiv:2512.14096v2 Announce Type: replace Abstract: Diffusion Transformers (DiTs) have achieved remarkable success in image generation, yet their deployment is hindered by high computational costs. We

Scene-Adaptive Nonlinear Tone Curves for Pseudo Ground-Truth Generation in Low-Light 3D Gaussian Splatting

ResearchDGX agent

arXiv:2606.11841v1 Announce Type: new Abstract: Low-light novel view synthesis is challenging because dark multi-view images contain noise, weak structural detail, and compressed dynamic range. Recent

SceneMiner: Identity-Preserving Multi-Task Fine-Tuning for Unified BEV Scene Mining

Model ReleasesDGX agent

arXiv:2606.11507v1 Announce Type: new Abstract: Mining hard, safety-critical scenes from driving logs is bottlenecked by the absence of difficulty labels, and no single proxy, collision risk, trajecto

Seeing What Matters: Perceptual Wrapper with Common Randomness for 3D Gaussian Splatting

Local AiDGX agent

arXiv:2606.11782v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) achieves impressive real-time rendering, it frequently struggles to synthesize high-frequency textures, a limitation

Semantic Segmentation of Node and Edge Diagrams for Assistive Technology

ResearchDGX agent

arXiv:2606.11320v1 Announce Type: new Abstract: In this paper, we present a novel set of related models for semantic segmentation of node-link diagrams. These diagrams are frequently used to represent

Semantically-Aware Diver Activity Recognition Framework for Effective Underwater Multi-Human-Robot Collaboration

SafetyDGX agent

arXiv:2606.12374v1 Announce Type: cross Abstract: Effective multi-human-robot collaboration is essential for expanding human-led operations in the challenging and high-risk underwater environment. For

SG2Loc: Sequential Visual Localization on 3D Scene Graphs

AgentsDGX agent

arXiv:2606.11880v1 Announce Type: new Abstract: Visual localization in complex indoor environments remains a critical challenge for robotics and AR applications. Sequential localization, where pose es

SheafStain: Sheaf-Theoretic Schrodinger Bridge for Spatially and Biologically Coherent Virtual Staining

Model ReleasesDGX agent

arXiv:2606.11846v1 Announce Type: new Abstract: Current virtual staining approaches offer the potential for time- and cost-efficient biomarker quantification in cancer diagnostics and prognostics. How

SHERPA: Seam-aware Harmonized ERP Adaptation for Open-Domain 360^irc Panorama Generation

ResearchDGX agent

arXiv:2606.12213v1 Announce Type: new Abstract: Panoramic imagery is increasingly used in world-generation, games, and simulation, where users may need not only photorealistic scenes but also stylized

Slots, Transitions, Loops: Learning Composable World Models for ARC

ResearchDGX agent

arXiv:2606.12316v1 Announce Type: new Abstract: ARC tests in-context rule induction: given a few input-output demonstrations, a model must infer the hidden rule and apply it to a new query. While many

Spatially Coupled Phase-to-Depth Calibration for Fringe Projection Profilometry

Model ReleasesDGX agent

arXiv:2606.11601v1 Announce Type: new Abstract: In fringe projection profilometry (FPP), depth is commonly recovered by fitting a phase-to-depth relation independently at each camera pixel. Although s

SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation

Local AiDGX agent

arXiv:2606.11969v1 Announce Type: new Abstract: Flow Matching has enabled robust text-to-video generation via latent ODE sampling. However, velocity approximation and numerical discretization errors i

SpikeTAD: Spiking Neural Networks for End-to-End Temporal Action Detection

ApplicationsDGX agent

arXiv:2606.12033v1 Announce Type: new Abstract: Video understanding is a crucial part of computer vision, with numerous application scenarios. With the increasing popularity of mobile devices, an incr

STEAM: Squeeze and Transform Enhanced Attention Module

Model ReleasesDGX agent

arXiv:2412.09023v3 Announce Type: replace Abstract: Channel and spatial attention mechanisms introduced in earlier work enhance the representational capabilities of deep convolutional neural networks

Tac-DINO: Learning Vision-Tactile Features with Patch Alignment

Model ReleasesDGX agent

arXiv:2606.12069v1 Announce Type: new Abstract: Touch is the primary medium through which humans interact with the environment. Currently, tactile learning mainly focuses on image-level pretraining or

Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks

ResearchDGX agent

arXiv:2409.18478v2 Announce Type: replace Abstract: With the development of video understanding, there is a proliferation of tasks for clip-level temporal video analysis, including temporal action det

The N-Body Problem: Parallel Execution from Single-Person Egocentric Video

Model ReleasesDGX agent

arXiv:2512.11393v2 Announce Type: replace Abstract: Humans can intuitively parallelise complex activities, but can a model predict this from observing a single person? Given one egocentric video, we i

Time-Conditioned and Multi-Time Survival Prediction from 2D PET/CT Projections in Lung Cancer

ResearchDGX agent

arXiv:2606.12140v1 Announce Type: new Abstract: Accurate prediction of overall survival (OS) from positron emission tomography/computed tomography (PET/CT) can support personalized treatment and follo

TopoCap: Learning Topology-Agnostic Motion Priors for Monocular Video-to-Animation

TutorialsDGX agent

arXiv:2606.12153v1 Announce Type: new Abstract: The explosion of generative 3D assets has created a massive demand for animation, yet current motion capture methods remain brittle, restricted to speci

Towards Conditional Feature Alignment for Cross-Domain Counting

SafetyDGX agent

arXiv:2506.17137v3 Announce Type: replace Abstract: Object counting models often degrade under cross-domain deployment because density composition varies across domains and is itself task-relevant. St

Traits Run Deeper: Trait-Specific Asymmetric Fusion for Personality Assessment

SafetyDGX agent

arXiv:2606.11269v1 Announce Type: new Abstract: Personality assessment aims to infer stable personality traits from dynamic behaviors across language, voice, and facial cues. Since different personali

TRON: Tracing Rays to Orchestrate a Neural Renderer for 3D Gaussian Reconstructions

ApplicationsDGX agent

arXiv:2606.11314v1 Announce Type: new Abstract: We introduce TRON, a rendering framework that combines 3D Gaussian ray tracing with neural rendering to enable realistic and controllable rendering of r

Understanding Cross-Sensor Feature Variations for Generalizable 3D Perception

ResearchDGX agent

arXiv:2606.11573v1 Announce Type: new Abstract: Radar-camera BEV perception often suffers from degraded performance when evaluated across datasets, as changes in driving scenes, sensor configurations,

Vision Transformers for Face Recognition Need More Registers

ResearchDGX agent

arXiv:2606.12036v1 Announce Type: new Abstract: Recent advances in Vision Transformers (ViTs) for face recognition (FR) have moved beyond the standard CLS-token paradigm. In this paradigm, a special c

ViT-FREE: Efficient Face Recognition via Early Exiting and Synthetic Adaptation

SafetyDGX agent

arXiv:2606.12023v1 Announce Type: new Abstract: Vision Transformers (ViTs) have gained significant attention in computer vision and shown strong potential for face recognition (FR). However, their hig

VL-DINO: Leveraging CLIP Vision-Language Knowledge for Open-Vocabulary Object Detectio

Model ReleasesDGX agent

arXiv:2606.11546v1 Announce Type: new Abstract: Vision-language models like CLIP can provide rich semantic priors for open-vocabulary object detection. However, jointly integrating both textual and vi

VLGA: Vision-Language-Geometry-Action Models for Autonomous Driving

SafetyDGX agent

arXiv:2606.12396v1 Announce Type: new Abstract: Vision-language-action (VLA) models can describe scenes and reason about them in language, yet still struggle to ground their actions in the dense 3D wo

VOID: Defeating Unauthorized Mimicry in Latent Diffusion Models

ResearchDGX agent

arXiv:2606.12263v1 Announce Type: new Abstract: While Latent Diffusion Models (LDMs) have revolutionized visual synthesis, they are increasingly exploited for unauthorized mimicry of individuals. Exis

Wild3R: Feed-Forward 3D Gaussian Splatting from Unconstrained Sparse Photo Collection

ApplicationsDGX agent

arXiv:2606.11894v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting (3DGS) removes the need for time-consuming per-scene optimization required by traditional 3DGS. However, existing fee

World Model Self-Distillation: Training World Models to Solve General Tasks

Model ReleasesDGX agent

arXiv:2606.12072v1 Announce Type: new Abstract: Pretrained video generators are promising visual world models that exhibit emergent task-solving abilities; however, their reliance on detailed textual

XPR: An Extensible Cross-Platform Point-Based Differentiable Renderer

ResearchDGX agent

arXiv:2606.11529v1 Announce Type: cross Abstract: Point-based differentiable rendering underpins modern 3D reconstruction, novel-view synthesis, and learning-based graphics pipelines, but developing n

10 Jun 2026

3D-CoS: A New 3D Reconstruction Paradigm Based on VLM Code Synthesis

AgentsDGX agent

arXiv:2606.10478v1 Announce Type: new Abstract: Most recent 3D reconstruction and editing systems operate on implicit and explicit representations such as NeRF, point clouds, or meshes. While these re

5% > 100%: Flatness Preference is All You Need for Multimodal Parameter-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2606.10488v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) methods provide a streamlined and efficient tool for adapting large models to domain-specific multimodal downstre

A fine-grained attention and geometric correspondence model for musculoskeletal risk classification in athletes using multimodal visual and skeletal features

SafetyDGX agent

arXiv:2509.05913v3 Announce Type: replace Abstract: Musculoskeletal disorders pose significant risks to athletes, and early risk assessment is essential for prevention. However, most existing methods

A Large Scale Open-Source Image and Video Dataset for Robust Wildfire Detection and Classification

Model ReleasesDGX agent

arXiv:2606.10174v1 Announce Type: new Abstract: Wildfire detection and monitoring are critical for mitigating fire spread and reducing environmental and infrastructural damage. In this work, we introd

A Multimodal RGB and Events Dataset for Hand Detection in First-Person View

ResearchDGX agent

arXiv:2606.10790v1 Announce Type: new Abstract: Existing hand detection algorithms work on images and the detection rate is restricted by the frame rate of the camera. In hand detection applications f

ABot-Earth 0.5: Generative 3D Earth Model

ApplicationsDGX agent

arXiv:2606.09967v1 Announce Type: new Abstract: We present ABot-Earth 0.5, a generative 3D framework designed to synthesize vast, seamless 3D environments from ubiquitous, geospatially referenced sate

Advancing Wood Identification in the Philippines: Utilizing the Xylorix Platform for Efficient AI Model Development and Deployment for Five Key Species

ResearchDGX agent

arXiv:2606.10876v1 Announce Type: new Abstract: Illegal logging and timber trade continue to pose significant challenges in the Philippines, where accurate wood species identification is essential for

An Uncertainty Estimation Framework for Dose Accumulation in Adaptive Radiotherapy: Application to CBCT-Guided Radiotherapy for Cervical Cancer

ResearchDGX agent

arXiv:2606.11012v1 Announce Type: new Abstract: Background and purpose: oART enables daily plan adaptation to interfraction anatomical variations, but cumulative dose estimation remains limited by DIR

Analyzing Training-Free Corruption Detection for Object Detection Datasets

ApplicationsDGX agent

arXiv:2606.10666v1 Announce Type: new Abstract: Annotation errors are widespread in computer vision datasets and can significantly degrade the performance of systems trained on them, particularly in c

← Previous
1…8485868788…211
Next →