AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Safety

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models

DGX agent

arXiv:2602.08735v3 Announce Type: replace Abstract: While multimodal large language models (MLLMs) have made substantial progress in single-image spatial reasoning, multi-image spatial reasoning, whic

safetyarxiv-cs-cv
11 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

From Nominal Intensity to Equivalent Rainfall: A Path-Based Credibility Evaluation Framework for Simulated Rainfall in Autonomous-Driving Perception Tests

DGX agent

arXiv:2606.11989v1 Announce Type: new Abstract: Credible simulated-rainfall conditions are essential for identifying perception-system boundaries and supporting SOTIF-oriented risk assessment in autom

agentsarxiv-cs-cv
11 Jun 2026
Hardware

From Simulation to Real-World: An In-Field 6D Pose Dataset and Baseline for Robotic Strawberry Harvesting

DGX agent

arXiv:2606.11381v1 Announce Type: new Abstract: Robotic strawberry harvesting requires precise 6D pose estimation; however, collecting 6D pose ground truth in real agricultural fields is inherently ch

hardwarearxiv-cs-cv
11 Jun 2026
Local Ai

Frozen Foundation-Model Embeddings Discard Small-Lesion Signal in Chest Radiography: Implications for Pre-Deployment Evaluation

DGX agent

arXiv:2606.11606v1 Announce Type: new Abstract: Frozen vision-transformer (ViT) foundation-model embeddings increasingly serve as the substrate for downstream chest-radiography (CXR) pipelines, yet wh

local-aiarxiv-cs-cv
11 Jun 2026
Research

Higher order PCA-like rotation-invariant features for detailed shape descriptors modulo rotation

DGX agent

arXiv:2601.03326v2 Announce Type: replace Abstract: PCA can be used for rotation invariant features, describing a shape with its p_{ab}=E[(x_i-E[x_a])(x_b-E[x_b])] covariance matrix approximating shap

researcharxiv-cs-cv
11 Jun 2026
Model Releases

How Auxiliary Reasoning Unleashes GUI Grounding in VLMs

DGX agent

arXiv:2509.11548v2 Announce Type: replace Abstract: Graphical user interface (GUI) grounding is a fundamental task for building GUI agents. However, general vision-language models (VLMs) struggle with

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

How Seemingly Inconsequential Design Choices Dictate Performance of LLMs in Pathology

DGX agent

arXiv:2606.12407v1 Announce Type: new Abstract: General-purpose large language models (LLMs) are routinely used as baselines when evaluating specialized pathology models on whole-slide images (WSIs).

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models

DGX agent

arXiv:2606.11289v1 Announce Type: new Abstract: Diffusion models have consistently driven progress in text-to-image generation. However, it is challenging to attribute recent progress to specific mode

model-releasesarxiv-cs-cv
11 Jun 2026
Research

Image Quality Assessment of Identity Cards Using Measures from Open Face Image Quality

DGX agent

arXiv:2606.11884v1 Announce Type: new Abstract: This paper addresses the challenge of assessing image quality in ID cards in remote verification systems by applying capture-related quality measures fr

researcharxiv-cs-cv
11 Jun 2026
Local Ai

Intelligent Skin Cancer Detection Using a Multispectral Metasurface and a Hybrid

DGX agent

arXiv:2606.11287v1 Announce Type: cross Abstract: Skin cancer is among the most prevalent malignancies worldwiAdbe satnradcitts early detection is essential for improving patient survival and reducing

local-aiarxiv-cs-cv
11 Jun 2026
Safety

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning

DGX agent

arXiv:2606.12195v1 Announce Type: new Abstract: Recent progress in foundation models has shifted toward agentic behavior involving multi-step reasoning and tool use. However, open-source efforts large

safetyarxiv-cs-cv
11 Jun 2026
Safety

ISAP-3D: Identity-Slot Aligned Part-Aware 3D Generation

DGX agent

arXiv:2606.12099v1 Announce Type: new Abstract: Part-aware 3D generation aims to synthesize structured objects with semantically meaningful components, yet often suffers from structural ambiguity due

safetyarxiv-cs-cv
11 Jun 2026
Safety

LAST: Bridging Vision-Language and Action Manifolds via Gromov-Wasserstein Alignment

DGX agent

arXiv:2606.11221v1 Announce Type: new Abstract: We take a Gromov-Wasserstein perspective on Vision-Language-Action (VLA) learning, where the goal is to make the relational geometry of action represent

safetyarxiv-cs-cv
11 Jun 2026
Safety

Learning Instance-Adaptive Low-Rank Orthogonal Subspaces for Clothes-Changing Person Re-Identification

DGX agent

arXiv:2606.11661v1 Announce Type: new Abstract: Clothes-changing person re-identification (CC-ReID) aims to recognize individuals despite drastic appearance changes caused by clothing variation. While

safetyarxiv-cs-cv
11 Jun 2026
Research

MFEN:Multi-Frequency Expert Network for Visible-Infrared Person Re-ID

DGX agent

arXiv:2606.12051v1 Announce Type: new Abstract: Visible-infrared person re-identification (VI-ReID) is challenging due to the large modality discrepancy between visible and infrared images. We contend

researcharxiv-cs-cv
11 Jun 2026
Safety

MLT-Dedup: Efficient Large-Scale Online Video Deduplication via Multi-Level Representations and Spatial-Temporal Matching

DGX agent

arXiv:2606.12215v1 Announce Type: new Abstract: The explosive growth of user-generated video content on online platforms is accompanied by the emergence of numerous near-duplicate videos--videos that

safetyarxiv-cs-cv
11 Jun 2026
Local Ai

Motion Reinforces Appearance: RGB-Skeleton Gated Residual Fusion for Micro-Gesture Online Recognition

DGX agent

arXiv:2606.11645v1 Announce Type: new Abstract: Micro-gesture analysis attracts increasing attention for inferring spontaneous emotion from subtle body movements. Micro-gesture online recognition, whi

local-aiarxiv-cs-cv
11 Jun 2026
Research

NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization

DGX agent

arXiv:2606.11363v1 Announce Type: new Abstract: Vector quantization is central to modern generative modeling pipelines, but large-codebook VQ models often suffer from codebook collapse. We identify en

researcharxiv-cs-cv
11 Jun 2026
Model Releases

On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning

DGX agent

arXiv:2606.11602v1 Announce Type: new Abstract: Audio-visual Generalized Zero-shot Learning (AV-GZSL) is a challenging task that aims to classify both seen and unseen objects or scenes by integrating

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

OSCS-SupCon: Orthogonal Sigmoid-based Common and Style Supervised Contrastive Learning for Robust Feature Disentanglement

DGX agent

arXiv:2606.11233v1 Announce Type: new Abstract: Supervised Contrastive Learning (SupCon) has achieved strong performance by explicitly modeling pairwise relationships among samples. However, existing

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

Parameter-Efficient Adapter Tuning for Tabular-Image Multimodal Learning

DGX agent

arXiv:2606.11682v1 Announce Type: new Abstract: Tabular-image multimodal learning aims to improve predictive modeling by jointly using structured tabular attributes and visual data. Although pretraine

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

ParseFixer: An Agentic Framework for Document Parsing via Selective Multimodal Correction

DGX agent

arXiv:2606.11977v1 Announce Type: new Abstract: In this report, we present our third-place solution for the DataMFM Challenge Track 1: Document Parsing. This track requires models to recover structure

model-releasesarxiv-cs-cv
11 Jun 2026
Safety

Performance Analysis of YOLOv11 and YOLOv8 for Mixed Traffic Object Detection under Adverse Weather Conditions in Developing Countries

DGX agent

arXiv:2606.12066v1 Announce Type: new Abstract: In modern vehicular systems, robust performance under harsh conditions has become a critical problem of autonomous driving. Our study delivers a compreh

safetyarxiv-cs-cv
11 Jun 2026
Model Releases

Periodic-MAE: Periodic Video Masked Autoencoder for rPPG Estimation

DGX agent

arXiv:2506.21855v2 Announce Type: replace Abstract: In this paper, we propose Periodic-MAE, a self-supervised framework for learning generalizable spatio-temporal representations of periodic physiolog

model-releasesarxiv-cs-cv
11 Jun 2026
Research

Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection

DGX agent

arXiv:2510.08073v2 Announce Type: replace Abstract: AI-generated videos have achieved near-perfect visual realism (e.g., Sora), urgently necessitating reliable detection mechanisms. However, detecting

researcharxiv-cs-cv
11 Jun 2026
Local Ai

PIGEON: VLM-Driven Object Navigation via Points of Interest Selection

DGX agent

arXiv:2511.13207v2 Announce Type: replace-cross Abstract: Object navigation in unseen indoor environments requires agents to perform semantic search under partial observability. Vision-language models

local-aiarxiv-cs-cv
11 Jun 2026
Safety

Plan-and-Verify Video Reward Reasoning with Spatio-Temporal Scene Graph Grounding

DGX agent

arXiv:2606.11838v1 Announce Type: new Abstract: Reward models for text-to-video (T2V) generation guide post-training but often fail at fine-grained semantic alignment. We trace this to two structural

safetyarxiv-cs-cv
11 Jun 2026
Model Releases

Precision-Aware Illumination-Disentangled Vision Transformer for Spacecraft 6D Pose Estimation

DGX agent

arXiv:2606.11619v1 Announce Type: new Abstract: Vision sensors provide a lightweight solution for spacecraft proximity operations, but monocular spacecraft 6D pose estimation remains difficult under i

model-releasesarxiv-cs-cv
11 Jun 2026
Research

PT-WNO: Point Transformer with Wavelet Neural Operator for 3D Point Cloud Semantic Segmentation

DGX agent

arXiv:2606.11466v1 Announce Type: new Abstract: Point cloud semantic segmentation requires architectures that capture both fine-grained local geometry and broad global scene structure. Transformer-bas

researcharxiv-cs-cv
11 Jun 2026
Model Releases

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding

DGX agent

arXiv:2606.12125v1 Announce Type: new Abstract: Long-video understanding remains challenging for multimodal large language models, because temporally extended videos often contain thousands of frames

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

RankVR: Low-Rank Structure Perception and Value Recalibration for Robust Composed Image Retrieval

DGX agent

arXiv:2606.11689v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) constitutes a pivotal paradigm requiring models to perform joint reasoning on reference images and modification texts. Ho

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

ReMoT: Reinforcement Learning with Motion Contrast Triplets

DGX agent

arXiv:2603.00461v3 Announce Type: replace Abstract: We present ReMoT, a unified training paradigm to systematically address the fundamental shortcomings of VLMs in spatio-temporal consistency -- a cri

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

RSTR: Reducing SpatioTemporal Redundancy in Diffusion Transformers

DGX agent

arXiv:2512.14096v2 Announce Type: replace Abstract: Diffusion Transformers (DiTs) have achieved remarkable success in image generation, yet their deployment is hindered by high computational costs. We

model-releasesarxiv-cs-cv
11 Jun 2026
Research

Scene-Adaptive Nonlinear Tone Curves for Pseudo Ground-Truth Generation in Low-Light 3D Gaussian Splatting

DGX agent

arXiv:2606.11841v1 Announce Type: new Abstract: Low-light novel view synthesis is challenging because dark multi-view images contain noise, weak structural detail, and compressed dynamic range. Recent

researcharxiv-cs-cv
11 Jun 2026
Model Releases

SceneMiner: Identity-Preserving Multi-Task Fine-Tuning for Unified BEV Scene Mining

DGX agent

arXiv:2606.11507v1 Announce Type: new Abstract: Mining hard, safety-critical scenes from driving logs is bottlenecked by the absence of difficulty labels, and no single proxy, collision risk, trajecto

model-releasesarxiv-cs-cv
11 Jun 2026
Local Ai

Seeing What Matters: Perceptual Wrapper with Common Randomness for 3D Gaussian Splatting

DGX agent

arXiv:2606.11782v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) achieves impressive real-time rendering, it frequently struggles to synthesize high-frequency textures, a limitation

local-aiarxiv-cs-cv
11 Jun 2026
Research

Semantic Segmentation of Node and Edge Diagrams for Assistive Technology

DGX agent

arXiv:2606.11320v1 Announce Type: new Abstract: In this paper, we present a novel set of related models for semantic segmentation of node-link diagrams. These diagrams are frequently used to represent

researcharxiv-cs-cv
11 Jun 2026
Safety

Semantically-Aware Diver Activity Recognition Framework for Effective Underwater Multi-Human-Robot Collaboration

DGX agent

arXiv:2606.12374v1 Announce Type: cross Abstract: Effective multi-human-robot collaboration is essential for expanding human-led operations in the challenging and high-risk underwater environment. For

safetyarxiv-cs-cv
11 Jun 2026
Agents

SG2Loc: Sequential Visual Localization on 3D Scene Graphs

DGX agent

arXiv:2606.11880v1 Announce Type: new Abstract: Visual localization in complex indoor environments remains a critical challenge for robotics and AR applications. Sequential localization, where pose es

agentsarxiv-cs-cv
11 Jun 2026
Model Releases

SheafStain: Sheaf-Theoretic Schrodinger Bridge for Spatially and Biologically Coherent Virtual Staining

DGX agent

arXiv:2606.11846v1 Announce Type: new Abstract: Current virtual staining approaches offer the potential for time- and cost-efficient biomarker quantification in cancer diagnostics and prognostics. How

model-releasesarxiv-cs-cv
11 Jun 2026
Research

SHERPA: Seam-aware Harmonized ERP Adaptation for Open-Domain 360^irc Panorama Generation

DGX agent

arXiv:2606.12213v1 Announce Type: new Abstract: Panoramic imagery is increasingly used in world-generation, games, and simulation, where users may need not only photorealistic scenes but also stylized

researcharxiv-cs-cv
11 Jun 2026
Research

Slots, Transitions, Loops: Learning Composable World Models for ARC

DGX agent

arXiv:2606.12316v1 Announce Type: new Abstract: ARC tests in-context rule induction: given a few input-output demonstrations, a model must infer the hidden rule and apply it to a new query. While many

researcharxiv-cs-cv
11 Jun 2026
Model Releases

Spatially Coupled Phase-to-Depth Calibration for Fringe Projection Profilometry

DGX agent

arXiv:2606.11601v1 Announce Type: new Abstract: In fringe projection profilometry (FPP), depth is commonly recovered by fitting a phase-to-depth relation independently at each camera pixel. Although s

model-releasesarxiv-cs-cv
11 Jun 2026
Local Ai

SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation

DGX agent

arXiv:2606.11969v1 Announce Type: new Abstract: Flow Matching has enabled robust text-to-video generation via latent ODE sampling. However, velocity approximation and numerical discretization errors i

local-aiarxiv-cs-cv
11 Jun 2026
Applications

SpikeTAD: Spiking Neural Networks for End-to-End Temporal Action Detection

DGX agent

arXiv:2606.12033v1 Announce Type: new Abstract: Video understanding is a crucial part of computer vision, with numerous application scenarios. With the increasing popularity of mobile devices, an incr

applicationsarxiv-cs-cv
11 Jun 2026
Model Releases

STEAM: Squeeze and Transform Enhanced Attention Module

DGX agent

arXiv:2412.09023v3 Announce Type: replace Abstract: Channel and spatial attention mechanisms introduced in earlier work enhance the representational capabilities of deep convolutional neural networks

model-releasesarxiv-cs-cv
11 Jun 2026
Model Releases

Tac-DINO: Learning Vision-Tactile Features with Patch Alignment

DGX agent

arXiv:2606.12069v1 Announce Type: new Abstract: Touch is the primary medium through which humans interact with the environment. Currently, tactile learning mainly focuses on image-level pretraining or

model-releasesarxiv-cs-cv
11 Jun 2026
Research

Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks

DGX agent

arXiv:2409.18478v2 Announce Type: replace Abstract: With the development of video understanding, there is a proliferation of tasks for clip-level temporal video analysis, including temporal action det

researcharxiv-cs-cv
11 Jun 2026
← Previous
1…105106107108109…263
Next →