AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
1 Jun 2026

Vanilla ViT for Automotive Point Cloud Semantic Segmentation

Local AiDGX agent

arXiv:2605.31177v1 Announce Type: new Abstract: Plain Transformers have become the de-facto architecture for processing text, audio, image, and video, offering a unified backbone for multimodal learni

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding

ResearchDGX agent

arXiv:2508.20478v2 Announce Type: replace Abstract: Long-form video understanding, characterized by long-range temporal dependencies and multiple events, remains a challenge. Existing methods often re

Vision-Based Localization in Dense Urban Environments: A Case Study of an Urban Village in China

Local AiDGX agent

arXiv:2605.30714v1 Announce Type: new Abstract: Urban villages, the widespread informal settlements which have emerged as a result of rapid urbanization, are now major residential hubs for migrant wor


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning

ApplicationsDGX agent

arXiv:2605.31457v1 Announce Type: new Abstract: With the rapid advancement of large multimodal models (LMMs), inference-time overhead has become a key bottleneck for real-world deployment. Existing me

VLM-GLoc: Vision-Language Model Enhanced Monte Carlo Localization for Robust Semantic Global Localization in Cluttered Quasi-Static Environments

ApplicationsDGX agent

arXiv:2605.30506v1 Announce Type: cross Abstract: Global localization in geometrically aliased, quasi-static environments such as grocery stores, offices, schools, and hospitals poses a significant ch

VolFill: Single-View Amodal 3D Scene Reconstruction with Volumetric Flow Matching

ResearchDGX agent

arXiv:2605.31466v1 Announce Type: new Abstract: Reconstructing the complete geometry of a scene from a single RGB image remains challenging - especially when inferring hidden structures where visual e

What is Missing? Explaining Neurons Activated by Absent Concepts

ResearchDGX agent

arXiv:2603.09787v2 Announce Type: replace Abstract: Explainable artificial intelligence (XAI) aims to provide human-interpretable insights into the behavior of deep neural networks (DNNs), typically b

World2Act: Latent Action Post-Training from World Model Dynamics

SafetyDGX agent

arXiv:2603.10422v2 Announce Type: replace Abstract: World Models (WMs) offer a promising mechanism for post-training Vision-Language-Action (VLA) policies by providing dynamics priors that improve gen

WristCompass: Kinematic Coupling as a Learnable Visual Concept for Ego-Camera Orientation

Model ReleasesDGX agent

arXiv:2605.30671v1 Announce Type: new Abstract: Recovering ego-camera orientation from manipulation video is a prerequisite for disentangling hand motion from camera motion, a key step in imitation le

YARD: Y-Architecture Register Decoding for Efficient Hallucination Mitigation in Large Vision-Language Models

Local AiDGX agent

arXiv:2605.31429v1 Announce Type: new Abstract: Contrastive decoding (CD) seeks to mitigate hallucinations in Large Vision-Language Models (LVLMs) by contrasting the output distributions of a standard

29 May 2026

3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding

ResearchDGX agent

arXiv:2605.29416v1 Announce Type: cross Abstract: Vision-Language-Action models have achieved remarkable progress in robotic manipulation, yet they suffer from a critical limitation: a lack of 3D scen

4DPC^2hat: Towards Dynamic Point Cloud Understanding with Failure-Aware Bootstrapping

ResearchDGX agent

arXiv:2602.03890v2 Announce Type: replace Abstract: Point clouds provide a compact and expressive representation of 3D objects, and have recently been integrated into multimodal large language models

A Deep Learning Iterative Framework for Sentinel-1 Stripmap Enhancement Based on Azimuth Doppler Decomposition

ResearchDGX agent

arXiv:2605.29088v1 Announce Type: new Abstract: Synthetic Aperture Radar (SAR) imagery enables all-weather, day-and-night Earth observation; however, it remains difficult to interpret due to speckle n

A Geometric View of SRC: Learning Representations for Stable Residual Inference

SafetyDGX agent

arXiv:2605.29673v1 Announce Type: cross Abstract: Reconstruction-based inference assigns a class by comparing class-wise reconstruction residuals; Sparse Representation Classification (SRC) is a canon

Accelerating HEVC Intra Partitioning via a CNN-Hierarchical Attention Transformer Hybrid

Local AiDGX agent

arXiv:2605.29063v1 Announce Type: cross Abstract: The recursive quad-tree partitioning in High Efficiency Video Coding (HEVC) incurs considerable computational overhead, with exhaustive rate-distortio

AdaState: Self-Evolving Anchors for Streaming Video Generation

ResearchDGX agent

arXiv:2605.30349v1 Announce Type: new Abstract: Autoregressive video diffusion models generate streaming video by producing frames sequentially, conditioning each chunk on previously generated content

AgentCVR: Active Multi-Agent Cross-Video Reasoning via Script-Simulated Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.29643v1 Announce Type: new Abstract: Cross-Video Reasoning (CVR) has emerged as a critical frontier in multimodal intelligence, requiring models to retrieve, align, and aggregate evidence d

AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation

Model ReleasesDGX agent

arXiv:2512.01334v2 Announce Type: replace Abstract: Text-guided image-to-video generation has made substantial progress, yet it still struggles to execute text-specified edits that require substantial

Ambient-robust Inverse Rendering using Active RGB-NIR Imaging

ResearchDGX agent

arXiv:2605.30250v1 Announce Type: new Abstract: Inverse rendering aims to reconstruct geometry and reflectance of objects from images. Despite recent progress, existing methods often produces inaccura

An Approach for Thyroid Nodule Analysis Using Thermographic Images

AgentsDGX agent

arXiv:2605.29221v1 Announce Type: new Abstract: Thyroid cancer is said to be the second most common type of cancer in female individuals and the third in males by 2030, according to projections. In ge

AnomalyAgent: Training-Free Agentic Models for Zero-/Few-Shot Anomaly Detection

AgentsDGX agent

arXiv:2605.30140v1 Announce Type: new Abstract: Benefiting from generalizability of vision-language models (VLMs) such as CLIP, many zero-/few-shot anomaly detection (AD) approaches have achieved impr

Audio Deepfake Detection with Half-Truth Localisation Using Cross-Attentive Feature Fusion

Model ReleasesDGX agent

arXiv:2605.29531v1 Announce Type: cross Abstract: Audio deepfake detection is well-studied as a binary problem, but partially manipulated speech, where a short synthesised segment is spliced into an o

Auditing Training-Free 3D Shape Retrieval with Diffused Geodesic Moments

Model ReleasesDGX agent

arXiv:2605.29004v1 Announce Type: new Abstract: Reported retrieval scores for training-free shape descriptors conflate local signal design, normalization, aggregation, codebook fitting, and metric cho

BadBlocks: Low-Cost and Stealthy Backdoor Attacks Tailored for Text-to-Image Diffusion Models

HardwareDGX agent

arXiv:2508.03221v5 Announce Type: replace-cross Abstract: Despite the remarkable progress of diffusion models in image generation, recent studies reveal their vulnerability to backdoor attacks via cov

Benchmarking Single-Factor Physical Video-to-Audio Generation

Model ReleasesDGX agent

arXiv:2605.30339v1 Announce Type: new Abstract: Generative video-to-audio (V2A) models produce highly plausible soundtracks, but it remains unclear whether they capture the underlying physical process

BitC-3DGS: High-Capacity 3D Gaussian Splatting Watermarking via Bit Compression

ResearchDGX agent

arXiv:2605.29583v1 Announce Type: new Abstract: High-capacity watermarking is necessary for 3D Gaussian Splatting (3DGS) assets to embed rich information (e.g., ownership, provenance, and authenticati

Boosting Image Quality Assessment Performance: Unsupervised Score Fusion by Deep Maximum a Posteriori Estimation

ResearchDGX agent

arXiv:2605.30269v1 Announce Type: new Abstract: Over the past decades, numerous Image Quality Assessment (IQA) models have emerged, aiming to predict the perceptual quality of images. However, individ

Boosting Zero-Shot 3D Style Transfer with 2D Pre-trained Priors

ResearchDGX agent

arXiv:2605.30065v1 Announce Type: new Abstract: In this work, we focus on zero-shot 3D style transfer that can generate multi-view consistent stylized views of the 3D scene given an arbitrary style im

Building and Road Recognition in Dense Urban Informal Settlements: A Dataset and Benchmark

Model ReleasesDGX agent

arXiv:2605.29856v1 Announce Type: new Abstract: As a widespread form of informal settlements, urban villages present significant challenges for sustainable urban development and governance. Precise ma

BullingerDB: A Dataset for Handwritten Text Recognition and Writer Retrieval

Model ReleasesDGX agent

arXiv:2605.30235v1 Announce Type: new Abstract: We present BullingerDB, a large-scale benchmark dataset for historical document analysis based on the correspondence of Heinrich Bullinger (1504-1575).

CamC2V: Context-aware Controllable Video Generation

TutorialsDGX agent

arXiv:2504.06022v3 Announce Type: replace Abstract: Recently, image-to-video (I2V) diffusion models have demonstrated impressive scene understanding and generative quality, incorporating image conditi

CapTalk: Text-Guided Stylization and Speech-Driven 3D Head Animation

ResearchDGX agent

arXiv:2605.29316v1 Announce Type: new Abstract: Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing met

Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive Smoothing

ResearchDGX agent

arXiv:2605.29809v1 Announce Type: cross Abstract: Large-scale text-to-image (T2I) diffusion models have enabled unprecedented creative applications, but their unauthorized use has raised serious intel

Ciphera: A Decentralised Biometric Identity Framework

ApplicationsDGX agent

arXiv:2605.29868v1 Announce Type: cross Abstract: Centralised biometric identity systems expose users to single points of failure, opaque verification processes, and irreversible biometric compromise.

CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning

ResearchDGX agent

arXiv:2605.29602v1 Announce Type: new Abstract: Multi-modal Retrieval-Augmented Generation (MMRAG) has emerged as a powerful paradigm for enhancing Multimodal Large Language Models in knowledge-intens

Colored Noise Diffusion Sampling

SafetyDGX agent

arXiv:2605.30332v1 Announce Type: new Abstract: Diffusion models achieve state-of-the-art image synthesis, with their generative trajectories fundamentally exhibiting a spectral bias, resolving low-fr

Comparative evaluation of photogrammetric reconstruction methods and 3D Gaussian Splatting for road surface roughness analysis

SafetyDGX agent

arXiv:2605.29452v1 Announce Type: new Abstract: Image-based 3D reconstruction offers a low-cost alternative to traditional sensor-based techniques for road surface assessment. This study compares four

Constructing efficient channels for ideal observers using the conjugate gradient method

ResearchDGX agent

arXiv:2605.29415v1 Announce Type: cross Abstract: Task-based assessment of image quality (IQ) is critically important for the design and optimization of medical imaging systems. Ideal observers, inclu

Cycle Consistency in Video Object-Centric Learning

SafetyDGX agent

arXiv:2605.30211v1 Announce Type: new Abstract: Self-supervised video Object-Centric Learning (OCL) aims to discover distinct objects and associate them across time, whereas self-supervised Multi-Obje

Deep Psychovisual Image Representations

TutorialsDGX agent

arXiv:2605.29260v1 Announce Type: new Abstract: Psychovisual models suggest human vision decouples low-level feature extraction from higher cognition by first forming intermediate abstractions. In con

DeepFake Forensics AI: A Multi-Modal Detection and Blockchain-Anchored Evidence Management Platform

ApplicationsDGX agent

arXiv:2605.29353v1 Announce Type: cross Abstract: The proliferation of AI-generated synthetic media poses a critical threat to the integrity of digital evidence in legal and forensic contexts. Existin

DefSynUS: Real-time Patient-specific Intrahepatic Vessel Identification via Deformation-Aware CT-US Domain Adaptation

SafetyDGX agent

arXiv:2605.29570v1 Announce Type: new Abstract: Purpose: Laparoscopic ultrasound (LUS) enhances the safety of liver surgery by visualizing intrahepatic vessels in real-time. Still, vessel identificati

Deja View: Looping Transformers for Multi-View 3D Reconstruction

SafetyDGX agent

arXiv:2605.30215v1 Announce Type: new Abstract: Recent feed-forward 3D reconstruction transformers have scaled to over a billion parameters, following the broader trend of increasing model capacity in

Detecting Unknown Objects via Energy-based Separation for Open World Object Detection

TutorialsDGX agent

arXiv:2603.29954v2 Announce Type: replace Abstract: In this work, we tackle the problem of Open World Object Detection (OWOD). This challenging scenario requires the detector to incrementally learn to

DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding

AgentsDGX agent

arXiv:2605.29879v1 Announce Type: new Abstract: Integrating open-vocabulary semantic information into dynamic 3D scene representations is essential for long-term embodied scene understanding. However,

DMC-CF: Dynamic Multimodal CounterFactual QA benchmark for Causal Reasoning

Model ReleasesDGX agent

arXiv:2605.29339v1 Announce Type: new Abstract: With the rapid advancement of multimodal large language models (MLLMs), models have demonstrated increasingly powerful multimodal capabilities. However,

DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark

Model ReleasesDGX agent

arXiv:2605.30027v1 Announce Type: new Abstract: Multimodal documents contain diverse elements, such as tables, figures, and layouts, which can complicate retrieval tasks. While current approaches typi

Domain-Agnostic Feature Modulation for Semi-Supervised Domain Generalization

ResearchDGX agent

arXiv:2503.20897v2 Announce Type: replace Abstract: Semi-supervised domain generalization (SSDG) leverages a small fraction of labeled data alongside unlabeled data to enhance model generalization. Mo

Dual Quaternion SE(3) Synchronization with Recovery Guarantees

ApplicationsDGX agent

arXiv:2602.00324v2 Announce Type: replace-cross Abstract: Synchronization over the special Euclidean group SE(3) aims to recover absolute poses from noisy pairwise relative transformations and is a co

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

SafetyDGX agent

arXiv:2510.27607v3 Announce Type: replace Abstract: Augmenting vision-language-action models (VLAs) with world models is promising for robotic policy learning but faces challenges in jointly predictin

DVSM: Decoder-only View Synthesis Model Done Right

ResearchDGX agent

arXiv:2605.29891v1 Announce Type: new Abstract: Recent Large View Synthesis Models (LVSMs) advocate an encoder-decoder architecture that separates reconstruction and rendering into distinct networks.

EarlyTom: Early Token Compression Completes Fast Video Understanding

HardwareDGX agent

arXiv:2605.30010v1 Announce Type: new Abstract: Video large language models (Video-LLMs) have demonstrated strong capabilities in video understanding tasks. However, their practical deployment is stil

EarthShift: a benchmark for measuring robustness to real-world distribution shifts in Earth observation

Model ReleasesDGX agent

arXiv:2605.29330v1 Announce Type: new Abstract: Current Earth observation benchmarks focus on measuring performance on diverse tasks and applications, typically measuring generalization in-distributio

Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets

ResearchDGX agent

arXiv:2605.29720v1 Announce Type: new Abstract: We propose Intrinsic Quality (IQ), a validation-free metric designed to estimate the inherent potential of face recognition (FR) datasets to produce hig

Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models

Model ReleasesDGX agent

arXiv:2605.29074v1 Announce Type: new Abstract: Are current Vision Language Models (VLMs) ready to comprehend and reason about complex embodied interactions in 3D environments? We introduce Embodied3D

ESAM++: Efficient Online 3D Perception on the Edge

HardwareDGX agent

arXiv:2605.29505v1 Announce Type: new Abstract: Online 3D scene perception in real time is essential for robotics, AR/VR, and autonomous systems, particularly in edge computing scenarios where computa

Eulerian Gaussian Splatting using Hashed Probability Pyramids

ResearchDGX agent

arXiv:2605.29136v1 Announce Type: new Abstract: We introduce a probabilistic splat-based radiance field framework that retains the fast rasterization and test-time efficiency of 3D Gaussian Splatting

Evaluation of Conversational Agents: Understanding Culture, Context and Environment in Emotion Detection

ApplicationsDGX agent

arXiv:2605.30099v1 Announce Type: new Abstract: Valuable decisions and highly prioritized analysis now depend on applications such as facial biometrics, social media photo tagging, and human robots in

EVL-ECG: Efficient ECG Interpretation With Multi-Aspect Heterogeneous Knowledge Distillation

Model ReleasesDGX agent

arXiv:2605.29977v1 Announce Type: new Abstract: High-fidelity ECG interpretation is increasingly reliant on massive foundation models, yet their deployment in clinical edge-care remains hindered by ex

F-RNG: Feed-Forward Relightable Neural Gaussians

ApplicationsDGX agent

arXiv:2605.25975v2 Announce Type: replace-cross Abstract: Capturing relightable 3D assets from real-world objects is a widely researched problem. Several per-scene optimization-based methods, based on

← Previous
1…107108109110111…211
Next →