AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Local Ai

Vanilla ViT for Automotive Point Cloud Semantic Segmentation

DGX agent

arXiv:2605.31177v1 Announce Type: new Abstract: Plain Transformers have become the de-facto architecture for processing text, audio, image, and video, offering a unified backbone for multimodal learni

local-aiarxiv-cs-cv
1 Jun 2026
Research
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding

DGX agent

arXiv:2508.20478v2 Announce Type: replace Abstract: Long-form video understanding, characterized by long-range temporal dependencies and multiple events, remains a challenge. Existing methods often re

researcharxiv-cs-cv
1 Jun 2026
Local Ai

Vision-Based Localization in Dense Urban Environments: A Case Study of an Urban Village in China

DGX agent

arXiv:2605.30714v1 Announce Type: new Abstract: Urban villages, the widespread informal settlements which have emerged as a result of rapid urbanization, are now major residential hubs for migrant wor

local-aiarxiv-cs-cv
1 Jun 2026
Applications

VisionPulse: Dynamic Visual Sparsity for Efficient Multimodal Reasoning

DGX agent

arXiv:2605.31457v1 Announce Type: new Abstract: With the rapid advancement of large multimodal models (LMMs), inference-time overhead has become a key bottleneck for real-world deployment. Existing me

applicationsarxiv-cs-cv
1 Jun 2026
Applications

VLM-GLoc: Vision-Language Model Enhanced Monte Carlo Localization for Robust Semantic Global Localization in Cluttered Quasi-Static Environments

DGX agent

arXiv:2605.30506v1 Announce Type: cross Abstract: Global localization in geometrically aliased, quasi-static environments such as grocery stores, offices, schools, and hospitals poses a significant ch

applicationsarxiv-cs-cv
1 Jun 2026
Research

VolFill: Single-View Amodal 3D Scene Reconstruction with Volumetric Flow Matching

DGX agent

arXiv:2605.31466v1 Announce Type: new Abstract: Reconstructing the complete geometry of a scene from a single RGB image remains challenging - especially when inferring hidden structures where visual e

researcharxiv-cs-cv
1 Jun 2026
Research

What is Missing? Explaining Neurons Activated by Absent Concepts

DGX agent

arXiv:2603.09787v2 Announce Type: replace Abstract: Explainable artificial intelligence (XAI) aims to provide human-interpretable insights into the behavior of deep neural networks (DNNs), typically b

researcharxiv-cs-cv
1 Jun 2026
Safety

World2Act: Latent Action Post-Training from World Model Dynamics

DGX agent

arXiv:2603.10422v2 Announce Type: replace Abstract: World Models (WMs) offer a promising mechanism for post-training Vision-Language-Action (VLA) policies by providing dynamics priors that improve gen

safetyarxiv-cs-cv
1 Jun 2026
Model Releases

WristCompass: Kinematic Coupling as a Learnable Visual Concept for Ego-Camera Orientation

DGX agent

arXiv:2605.30671v1 Announce Type: new Abstract: Recovering ego-camera orientation from manipulation video is a prerequisite for disentangling hand motion from camera motion, a key step in imitation le

model-releasesarxiv-cs-cv
1 Jun 2026
Local Ai

YARD: Y-Architecture Register Decoding for Efficient Hallucination Mitigation in Large Vision-Language Models

DGX agent

arXiv:2605.31429v1 Announce Type: new Abstract: Contrastive decoding (CD) seeks to mitigate hallucinations in Large Vision-Language Models (LVLMs) by contrasting the output distributions of a standard

local-aiarxiv-cs-cv
1 Jun 2026
Research

3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding

DGX agent

arXiv:2605.29416v1 Announce Type: cross Abstract: Vision-Language-Action models have achieved remarkable progress in robotic manipulation, yet they suffer from a critical limitation: a lack of 3D scen

researcharxiv-cs-cv
29 May 2026
Research

4DPC^2hat: Towards Dynamic Point Cloud Understanding with Failure-Aware Bootstrapping

DGX agent

arXiv:2602.03890v2 Announce Type: replace Abstract: Point clouds provide a compact and expressive representation of 3D objects, and have recently been integrated into multimodal large language models

researcharxiv-cs-cv
29 May 2026
Research

A Deep Learning Iterative Framework for Sentinel-1 Stripmap Enhancement Based on Azimuth Doppler Decomposition

DGX agent

arXiv:2605.29088v1 Announce Type: new Abstract: Synthetic Aperture Radar (SAR) imagery enables all-weather, day-and-night Earth observation; however, it remains difficult to interpret due to speckle n

researcharxiv-cs-cv
29 May 2026
Safety

A Geometric View of SRC: Learning Representations for Stable Residual Inference

DGX agent

arXiv:2605.29673v1 Announce Type: cross Abstract: Reconstruction-based inference assigns a class by comparing class-wise reconstruction residuals; Sparse Representation Classification (SRC) is a canon

safetyarxiv-cs-cv
29 May 2026
Local Ai

Accelerating HEVC Intra Partitioning via a CNN-Hierarchical Attention Transformer Hybrid

DGX agent

arXiv:2605.29063v1 Announce Type: cross Abstract: The recursive quad-tree partitioning in High Efficiency Video Coding (HEVC) incurs considerable computational overhead, with exhaustive rate-distortio

local-aiarxiv-cs-cv
29 May 2026
Research

AdaState: Self-Evolving Anchors for Streaming Video Generation

DGX agent

arXiv:2605.30349v1 Announce Type: new Abstract: Autoregressive video diffusion models generate streaming video by producing frames sequentially, conditioning each chunk on previously generated content

researcharxiv-cs-cv
29 May 2026
Model Releases

AgentCVR: Active Multi-Agent Cross-Video Reasoning via Script-Simulated Reinforcement Learning

DGX agent

arXiv:2605.29643v1 Announce Type: new Abstract: Cross-Video Reasoning (CVR) has emerged as a critical frontier in multimodal intelligence, requiring models to retrieve, align, and aggregate evidence d

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation

DGX agent

arXiv:2512.01334v2 Announce Type: replace Abstract: Text-guided image-to-video generation has made substantial progress, yet it still struggles to execute text-specified edits that require substantial

model-releasesarxiv-cs-cv
29 May 2026
Research

Ambient-robust Inverse Rendering using Active RGB-NIR Imaging

DGX agent

arXiv:2605.30250v1 Announce Type: new Abstract: Inverse rendering aims to reconstruct geometry and reflectance of objects from images. Despite recent progress, existing methods often produces inaccura

researcharxiv-cs-cv
29 May 2026
Agents

An Approach for Thyroid Nodule Analysis Using Thermographic Images

DGX agent

arXiv:2605.29221v1 Announce Type: new Abstract: Thyroid cancer is said to be the second most common type of cancer in female individuals and the third in males by 2030, according to projections. In ge

agentsarxiv-cs-cv
29 May 2026
Agents

AnomalyAgent: Training-Free Agentic Models for Zero-/Few-Shot Anomaly Detection

DGX agent

arXiv:2605.30140v1 Announce Type: new Abstract: Benefiting from generalizability of vision-language models (VLMs) such as CLIP, many zero-/few-shot anomaly detection (AD) approaches have achieved impr

agentsarxiv-cs-cv
29 May 2026
Model Releases

Audio Deepfake Detection with Half-Truth Localisation Using Cross-Attentive Feature Fusion

DGX agent

arXiv:2605.29531v1 Announce Type: cross Abstract: Audio deepfake detection is well-studied as a binary problem, but partially manipulated speech, where a short synthesised segment is spliced into an o

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Auditing Training-Free 3D Shape Retrieval with Diffused Geodesic Moments

DGX agent

arXiv:2605.29004v1 Announce Type: new Abstract: Reported retrieval scores for training-free shape descriptors conflate local signal design, normalization, aggregation, codebook fitting, and metric cho

model-releasesarxiv-cs-cv
29 May 2026
Hardware

BadBlocks: Low-Cost and Stealthy Backdoor Attacks Tailored for Text-to-Image Diffusion Models

DGX agent

arXiv:2508.03221v5 Announce Type: replace-cross Abstract: Despite the remarkable progress of diffusion models in image generation, recent studies reveal their vulnerability to backdoor attacks via cov

hardwarearxiv-cs-cv
29 May 2026
Model Releases

Benchmarking Single-Factor Physical Video-to-Audio Generation

DGX agent

arXiv:2605.30339v1 Announce Type: new Abstract: Generative video-to-audio (V2A) models produce highly plausible soundtracks, but it remains unclear whether they capture the underlying physical process

model-releasesarxiv-cs-cv
29 May 2026
Research

BitC-3DGS: High-Capacity 3D Gaussian Splatting Watermarking via Bit Compression

DGX agent

arXiv:2605.29583v1 Announce Type: new Abstract: High-capacity watermarking is necessary for 3D Gaussian Splatting (3DGS) assets to embed rich information (e.g., ownership, provenance, and authenticati

researcharxiv-cs-cv
29 May 2026
Research

Boosting Image Quality Assessment Performance: Unsupervised Score Fusion by Deep Maximum a Posteriori Estimation

DGX agent

arXiv:2605.30269v1 Announce Type: new Abstract: Over the past decades, numerous Image Quality Assessment (IQA) models have emerged, aiming to predict the perceptual quality of images. However, individ

researcharxiv-cs-cv
29 May 2026
Research

Boosting Zero-Shot 3D Style Transfer with 2D Pre-trained Priors

DGX agent

arXiv:2605.30065v1 Announce Type: new Abstract: In this work, we focus on zero-shot 3D style transfer that can generate multi-view consistent stylized views of the 3D scene given an arbitrary style im

researcharxiv-cs-cv
29 May 2026
Model Releases

Building and Road Recognition in Dense Urban Informal Settlements: A Dataset and Benchmark

DGX agent

arXiv:2605.29856v1 Announce Type: new Abstract: As a widespread form of informal settlements, urban villages present significant challenges for sustainable urban development and governance. Precise ma

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

BullingerDB: A Dataset for Handwritten Text Recognition and Writer Retrieval

DGX agent

arXiv:2605.30235v1 Announce Type: new Abstract: We present BullingerDB, a large-scale benchmark dataset for historical document analysis based on the correspondence of Heinrich Bullinger (1504-1575).

model-releasesarxiv-cs-cv
29 May 2026
Tutorials

CamC2V: Context-aware Controllable Video Generation

DGX agent

arXiv:2504.06022v3 Announce Type: replace Abstract: Recently, image-to-video (I2V) diffusion models have demonstrated impressive scene understanding and generative quality, incorporating image conditi

tutorialsarxiv-cs-cv
29 May 2026
Research

CapTalk: Text-Guided Stylization and Speech-Driven 3D Head Animation

DGX agent

arXiv:2605.29316v1 Announce Type: new Abstract: Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing met

researcharxiv-cs-cv
29 May 2026
Research

Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive Smoothing

DGX agent

arXiv:2605.29809v1 Announce Type: cross Abstract: Large-scale text-to-image (T2I) diffusion models have enabled unprecedented creative applications, but their unauthorized use has raised serious intel

researcharxiv-cs-cv
29 May 2026
Applications

Ciphera: A Decentralised Biometric Identity Framework

DGX agent

arXiv:2605.29868v1 Announce Type: cross Abstract: Centralised biometric identity systems expose users to single points of failure, opaque verification processes, and irreversible biometric compromise.

applicationsarxiv-cs-cv
29 May 2026
Research

CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning

DGX agent

arXiv:2605.29602v1 Announce Type: new Abstract: Multi-modal Retrieval-Augmented Generation (MMRAG) has emerged as a powerful paradigm for enhancing Multimodal Large Language Models in knowledge-intens

researcharxiv-cs-cv
29 May 2026
Safety

Colored Noise Diffusion Sampling

DGX agent

arXiv:2605.30332v1 Announce Type: new Abstract: Diffusion models achieve state-of-the-art image synthesis, with their generative trajectories fundamentally exhibiting a spectral bias, resolving low-fr

safetyarxiv-cs-cv
29 May 2026
Safety

Comparative evaluation of photogrammetric reconstruction methods and 3D Gaussian Splatting for road surface roughness analysis

DGX agent

arXiv:2605.29452v1 Announce Type: new Abstract: Image-based 3D reconstruction offers a low-cost alternative to traditional sensor-based techniques for road surface assessment. This study compares four

safetyarxiv-cs-cv
29 May 2026
Research

Constructing efficient channels for ideal observers using the conjugate gradient method

DGX agent

arXiv:2605.29415v1 Announce Type: cross Abstract: Task-based assessment of image quality (IQ) is critically important for the design and optimization of medical imaging systems. Ideal observers, inclu

researcharxiv-cs-cv
29 May 2026
Safety

Cycle Consistency in Video Object-Centric Learning

DGX agent

arXiv:2605.30211v1 Announce Type: new Abstract: Self-supervised video Object-Centric Learning (OCL) aims to discover distinct objects and associate them across time, whereas self-supervised Multi-Obje

safetyarxiv-cs-cv
29 May 2026
Tutorials

Deep Psychovisual Image Representations

DGX agent

arXiv:2605.29260v1 Announce Type: new Abstract: Psychovisual models suggest human vision decouples low-level feature extraction from higher cognition by first forming intermediate abstractions. In con

tutorialsarxiv-cs-cv
29 May 2026
Applications

DeepFake Forensics AI: A Multi-Modal Detection and Blockchain-Anchored Evidence Management Platform

DGX agent

arXiv:2605.29353v1 Announce Type: cross Abstract: The proliferation of AI-generated synthetic media poses a critical threat to the integrity of digital evidence in legal and forensic contexts. Existin

applicationsarxiv-cs-cv
29 May 2026
Safety

DefSynUS: Real-time Patient-specific Intrahepatic Vessel Identification via Deformation-Aware CT-US Domain Adaptation

DGX agent

arXiv:2605.29570v1 Announce Type: new Abstract: Purpose: Laparoscopic ultrasound (LUS) enhances the safety of liver surgery by visualizing intrahepatic vessels in real-time. Still, vessel identificati

safetyarxiv-cs-cv
29 May 2026
Safety

Deja View: Looping Transformers for Multi-View 3D Reconstruction

DGX agent

arXiv:2605.30215v1 Announce Type: new Abstract: Recent feed-forward 3D reconstruction transformers have scaled to over a billion parameters, following the broader trend of increasing model capacity in

safetyarxiv-cs-cv
29 May 2026
Tutorials

Detecting Unknown Objects via Energy-based Separation for Open World Object Detection

DGX agent

arXiv:2603.29954v2 Announce Type: replace Abstract: In this work, we tackle the problem of Open World Object Detection (OWOD). This challenging scenario requires the detector to incrementally learn to

tutorialsarxiv-cs-cv
29 May 2026
Agents

DGSG-Mind: Dynamic 3D Gaussian Scene Graphs for Long-Term Scene Understanding and Grounding

DGX agent

arXiv:2605.29879v1 Announce Type: new Abstract: Integrating open-vocabulary semantic information into dynamic 3D scene representations is essential for long-term embodied scene understanding. However,

agentsarxiv-cs-cv
29 May 2026
Model Releases

DMC-CF: Dynamic Multimodal CounterFactual QA benchmark for Causal Reasoning

DGX agent

arXiv:2605.29339v1 Announce Type: new Abstract: With the rapid advancement of multimodal large language models (MLLMs), models have demonstrated increasingly powerful multimodal capabilities. However,

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

DocRetriever: A Plug-and-Play Framework for Multimodal Document Retrieval with Comprehensive Benchmark

DGX agent

arXiv:2605.30027v1 Announce Type: new Abstract: Multimodal documents contain diverse elements, such as tables, figures, and layouts, which can complicate retrieval tasks. While current approaches typi

model-releasesarxiv-cs-cv
29 May 2026
Research

Domain-Agnostic Feature Modulation for Semi-Supervised Domain Generalization

DGX agent

arXiv:2503.20897v2 Announce Type: replace Abstract: Semi-supervised domain generalization (SSDG) leverages a small fraction of labeled data alongside unlabeled data to enhance model generalization. Mo

researcharxiv-cs-cv
29 May 2026
← Previous
1…134135136137138…263
Next →