AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Local Ai

Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion

DGX agent

arXiv:2605.16579v1 Announce Type: new Abstract: Autoregressive (AR) video diffusion is a powerful paradigm for streaming and interactive video generation. However, its reliance on softmax self-attenti

local-aiarxiv-cs-cv
19 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

Attention-Aware Transformer-Based Aggregation Network for Video Periocular Recognition

DGX agent

arXiv:2605.16550v1 Announce Type: new Abstract: Video periocular recognition is the task of recognizing an individual's identity based on the region around an individual's eyes. The periocular area is

researcharxiv-cs-cv
19 May 2026
Model Releases

Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring

DGX agent

arXiv:2605.16386v1 Announce Type: new Abstract: Multimodal large language models (LLMs) are increasingly explored as automated evaluators in clinical settings, yet their scoring behavior on ordinal cl

model-releasesarxiv-cs-cv
19 May 2026
Local Ai

Aurora: Unified Video Editing with a Tool-Using Agent

DGX agent

arXiv:2605.18748v1 Announce Type: new Abstract: Recent video editing models have converged on a unified conditioning design: a single diffusion transformer jointly consumes text, source video, and ref

local-aiarxiv-cs-cv
19 May 2026
Local Ai

Axial-Relation Guided Fusion State Space Model for Optical-Elevation Sensing Image Segmentation

DGX agent

arXiv:2605.16768v1 Announce Type: new Abstract: Semantic segmentation of multi-source remote sensing images is a fundamental task for Earth observation applications. Existing methods often struggle wi

local-aiarxiv-cs-cv
19 May 2026
Model Releases

Benchmarking Recurrent Event-Based Object Detection for Industrial Multi-Class Recognition on MTevent

DGX agent

arXiv:2603.21787v2 Announce Type: replace Abstract: Event cameras are attractive for industrial robotics because they provide high temporal resolution, high dynamic range, and reduced motion blur. How

model-releasesarxiv-cs-cv
19 May 2026
Safety

Benchmarking transferability of SSL pretraining to same and different modality segmentation tasks

DGX agent

arXiv:2605.18491v1 Announce Type: new Abstract: Methods: Nine SSL methods spanning four pretext-task families were pretrained from scratch using the same 10{,}412 3D CT scans (1.89~M 2D axial slices)

safetyarxiv-cs-cv
19 May 2026
Research

Best Segmentation Buddies for Image-Shape Correspondence

DGX agent

arXiv:2605.18193v1 Announce Type: new Abstract: Finding correspondences is a fundamental and extensively researched problem in computer vision and graphics. In this work, we examine the underexplored

researcharxiv-cs-cv
19 May 2026
Research

Better Together: Evaluating the Complementarity of Earth Embedding Models

DGX agent

arXiv:2605.18667v1 Announce Type: new Abstract: Earth embedding models transform Earth observation data into embeddings uniquely tied to locations on the Earth's surface. These models are typically ev

researcharxiv-cs-cv
19 May 2026
Model Releases

Beyond Detection: A Structure-Aware Framework for Scene Text Tracking

DGX agent

arXiv:2605.17270v1 Announce Type: new Abstract: Modern visual object trackers show impressive results on general targets, yet their performance drops substantially when dealing with scene text. Althou

model-releasesarxiv-cs-cv
19 May 2026
Research

Beyond Euclidean Prototypes: Spectral Disentanglement and Geodesic Matching for Few-Shot Medical Image Segmentation

DGX agent

arXiv:2605.17904v1 Announce Type: new Abstract: Few-Shot Medical Image Segmentation (FSMIS) aims to delineate novel anatomical targets from one or a few annotated support images, addressing the annota

researcharxiv-cs-cv
19 May 2026
Research

Beyond MMSE: Enhancing PnP Restoration with ProxiMAP

DGX agent

arXiv:2605.16396v1 Announce Type: new Abstract: Plug-and-Play (PnP) methods have become standard tools for solving imaging inverse problems by replacing the intractable maximum a posteriori (MAP) deno

researcharxiv-cs-cv
19 May 2026
Research

Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation

DGX agent

arXiv:2601.01593v2 Announce Type: replace Abstract: Manual font design is an intricate process that transforms a stylistic visual concept into a coherent glyph set. This challenge persists in automate

researcharxiv-cs-cv
19 May 2026
Model Releases

Beyond Point-Wise Matching: Structural Representation Alignment for Accelerating Diffusion Transformers

DGX agent

arXiv:2605.16949v1 Announce Type: new Abstract: Recent advances in Diffusion Transformers (DiTs) demonstrate that aligning noisy latent states with well-trained semantic features-as pioneered by Repre

model-releasesarxiv-cs-cv
19 May 2026
Safety

BIDO: A Biometric Identity Online Authentication Framework

DGX agent

arXiv:2605.16908v1 Announce Type: cross Abstract: Security systems demand continuous, cryptograph- ically robust identity verification without requiring subjects to carry physical tokens, smart cards,

safetyarxiv-cs-cv
19 May 2026
Agents

Bio-Inspired Event-Based Visual Servoing for Ground Robots

DGX agent

arXiv:2603.23672v2 Announce Type: replace-cross Abstract: Biological sensory systems are inherently adaptive, filtering out constant stimuli and prioritizing relative changes, likely enhancing computa

agentsarxiv-cs-cv
19 May 2026
Model Releases

Brain-inspired spike-timing plasticity for reliable label-efficient event-camera vision

DGX agent

arXiv:2605.17686v1 Announce Type: new Abstract: Deploying event-camera object detectors is constrained by per-frame labeling requirements and GPU compute demands. This work introduces three local spik

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

Breaking Annotation Barriers: Generalized Video Quality Assessment via Ranking-based Self-Supervision

DGX agent

arXiv:2505.03631v4 Announce Type: replace Abstract: Video quality assessment (VQA) is essential for quantifying perceptual quality in various video processing workflows, spanning from camera capture s

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

Bridging Data Trials and Task Barriers: A Unified Framework for Sketch Biometric Identification

DGX agent

arXiv:2605.17367v1 Announce Type: new Abstract: Different from existing cross-modality identification tasks (e.g., heterogeneous face recognition, sketch re-identification, etc.), we introduce a novel

model-releasesarxiv-cs-cv
19 May 2026
Safety

Bridging the Intention-Expression Gap: Aligning Multi-Dimensional Preferences via Hierarchical Relevance Feedback in Text-to-Image Diffusion

DGX agent

arXiv:2603.14936v3 Announce Type: replace Abstract: Users often possess a clear visual intent but struggle to articulate it precisely in language. This intention-expression gap makes aligning generate

safetyarxiv-cs-cv
19 May 2026
Research

Bridging the Modality Bottleneck in Pathology MIL through Virtual Molecular Staining

DGX agent

arXiv:2605.16392v1 Announce Type: cross Abstract: Multiple instance learning (MIL) is the dominant framework for whole-slide image analysis in computational pathology, typically combining a frozen pat

researcharxiv-cs-cv
19 May 2026
Hardware

Bundle Adjustment in the Eager Mode

DGX agent

arXiv:2409.12190v4 Announce Type: replace-cross Abstract: Bundle adjustment (BA) is a critical technique in various robotic applications such as simultaneous localization and mapping (SLAM), augmented

hardwarearxiv-cs-cv
19 May 2026
Research

CAB: Accelerating Flow and Diffusion Sampling via Rectification and Corrected Adams-Bashforth

DGX agent

arXiv:2605.16736v1 Announce Type: new Abstract: Flow and diffusion models achieve high-fidelity, high-resolution image synthesis, but often require many function evaluations (NFEs) at sampling time. E

researcharxiv-cs-cv
19 May 2026
Research

CADS: Conformal Adaptive Decision System for Cost-Efficient Image Classification

DGX agent

arXiv:2605.16401v1 Announce Type: new Abstract: While high-capacity AI models have advanced state-of-the-art performance, their practical deployment is often hindered by high inference costs, environm

researcharxiv-cs-cv
19 May 2026
Model Releases

Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate

DGX agent

arXiv:2605.18754v1 Announce Type: new Abstract: Multiview 3D evaluation assumes that the images being scored are observations of one static 3D scene. This assumption can fail in NVS and sparse-view re

model-releasesarxiv-cs-cv
19 May 2026
Local Ai

CanViT: Toward Active-Vision Foundation Models

DGX agent

arXiv:2603.22570v2 Announce Type: replace Abstract: Active computer vision promises efficient, biologically plausible perception through sequential, localized glimpses, but lacks scalable general-purp

local-aiarxiv-cs-cv
19 May 2026
Research

CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything Model

DGX agent

arXiv:2605.16901v1 Announce Type: new Abstract: Segment Anything Models (SAMs) are extensively used in computer vision for universal image segmentation, but deploying them on resource-constrained devi

researcharxiv-cs-cv
19 May 2026
Research

CATRF: Codec-Adaptive TriPlane Radiance Fields for Volumetric Content Delivery

DGX agent

arXiv:2605.18054v1 Announce Type: cross Abstract: Volumetric media promises next-generation content delivery applications, but its bandwidth demand remains a key bottleneck. Implicit and hybrid volume

researcharxiv-cs-cv
19 May 2026
Local Ai

Causal Attribution via Activation Patching

DGX agent

arXiv:2603.13652v2 Announce Type: replace Abstract: Attribution methods for Vision Transformers (ViTs) aim to identify image regions that influence model predictions, but producing faithful and well-l

local-aiarxiv-cs-cv
19 May 2026
Research

ChronoSC: Task-Oriented Semantic Communication via Temporal-to-Color Encoding

DGX agent

arXiv:2605.16388v1 Announce Type: new Abstract: Semantic communication (SC) aims to reduce transmission overhead by conveying task-relevant information rather than raw data. However, existing SC appro

researcharxiv-cs-cv
19 May 2026
Applications

CineMatte: Background Matting for Virtual Production and Beyond

DGX agent

arXiv:2605.18328v1 Announce Type: new Abstract: LED Virtual Production (VP) uses large LED volumes to render backgrounds in real time, enabling in-camera visual effects but making post-shot changes la

applicationsarxiv-cs-cv
19 May 2026
Research

CLEAR-HPV: Interpretable Concept Discovery for HPV-Associated Morphology in Whole-Slide Histologyhttps://arxiv.org/submit/7596892/preview

DGX agent

arXiv:2602.05126v2 Announce Type: replace Abstract: Human papillomavirus (HPV) status is a critical determinant of prognosis and treatment response in head and neck and cervical cancers. Although atte

researcharxiv-cs-cv
19 May 2026
Local Ai

CMAG: Concept-Scaffolded Retrieval for Marketplace Avatar Generation

DGX agent

arXiv:2605.18680v1 Announce Type: new Abstract: Metaverse platforms rely on creator-driven marketplaces where avatars are assembled from discrete, taxonomy-labeled 3D assets (e.g., tops, bottoms, shoe

local-aiarxiv-cs-cv
19 May 2026
Research

Coarse Semantic Injection for LLM-Conditioned Structured Indoor Prediction

DGX agent

arXiv:2605.16832v1 Announce Type: new Abstract: Large language models (LLMs) have recently been used as structured decoders for indoor understanding from 3D point-token inputs. However, point cloud en

researcharxiv-cs-cv
19 May 2026
Model Releases

Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis

DGX agent

arXiv:2605.18451v1 Announce Type: new Abstract: Designing realistic and functional 3D indoor rooms is essential for a wide range of applications, including interior design, virtual reality, gaming, an

model-releasesarxiv-cs-cv
19 May 2026
Research

CogBlender: Towards Continuous Cognitive Intervention in Text-to-Image Generation

DGX agent

arXiv:2603.09286v2 Announce Type: replace Abstract: Beyond conveying semantic information, images also possess cognitive properties that elicit specific psychological responses from viewers, such as m

researcharxiv-cs-cv
19 May 2026
Safety

Collaborative Learning for Semi-Supervised LiDAR Semantic Segmentation

DGX agent

arXiv:2605.17135v1 Announce Type: new Abstract: Annotating large-scale LiDAR point clouds for 3D semantic segmentation is costly and time-consuming, which motivates the use of semi-supervised learning

safetyarxiv-cs-cv
19 May 2026
Tutorials

Collision-Resistant Single-Pass Method for Unsupervised Fine-Grained Image Hashing

DGX agent

arXiv:2605.18288v1 Announce Type: new Abstract: Unsupervised fine-grained image hashing aims to learn compact binary codes that preserve subtle visual differences among highly similar instances withou

tutorialsarxiv-cs-cv
19 May 2026
Research

Color as the Impetus: Transforming Few-Shot Learner

DGX agent

arXiv:2507.22136v3 Announce Type: replace Abstract: Humans possess innate meta-learning capabilities, partly attributable to their exceptional color perception. In this paper, we pioneer an innovative

researcharxiv-cs-cv
19 May 2026
Model Releases

CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects

DGX agent

arXiv:2604.02060v2 Announce Type: replace Abstract: When told to 'cut the cake,' a robot must choose the knife over nearby scissors, despite both objects affording the same cutting function. In real-w

model-releasesarxiv-cs-cv
19 May 2026
Research

Compositional Adversarial Training for Robust Visual Watermarking

DGX agent

arXiv:2605.16720v1 Announce Type: new Abstract: Robust watermarking is typically trained with random post-processing augmentation, but random sampling under-covers the combinatorial space of realistic

researcharxiv-cs-cv
19 May 2026
Research

Concepts Worth Having: Refining VLM-Guided Concept Bottleneck Models with Minimal Annotations

DGX agent

arXiv:2605.16405v1 Announce Type: new Abstract: Concept-bottleneck models (CBMs) are neural classifiers that compute predictions from high-level concepts extracted from the input. CBMs ensure stakehol

researcharxiv-cs-cv
19 May 2026
Model Releases

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting

DGX agent

arXiv:2508.04227v2 Announce Type: replace Abstract: Vision-language models (VLMs) and the recent surge of Multimodal Large Language Models (MLLMs) have revolutionized artificial intelligence with unpr

model-releasesarxiv-cs-cv
19 May 2026
Safety

Contrastive-SDXL: Annotation-Preserving Night-Time Augmentation for Pedestrian Detection

DGX agent

arXiv:2605.16406v1 Announce Type: new Abstract: Night-time pedestrian detection remains challenging because labelled night-time data are limited and large illumination differences make daytime-only tr

safetyarxiv-cs-cv
19 May 2026
Model Releases

Controlla: Learning Controllability via Graph-Constrained Latent Geometry

DGX agent

arXiv:2605.16603v1 Announce Type: new Abstract: Controllable multimodal generation is commonly formulated as an inference-time conditioning problem using prompts, guidance, or auxiliary modules. While

model-releasesarxiv-cs-cv
19 May 2026
Safety

Controlling Decision Drift in Multimodal Sentiment Analysis with Missing Modalities

DGX agent

arXiv:2605.16889v1 Announce Type: new Abstract: Multimodal sentiment analysis relies on textual, acoustic, and visual signals, yet real-world data often suffer from modality missing and quality imbala

safetyarxiv-cs-cv
19 May 2026
Research

Counting Machine Parts

DGX agent

arXiv:2605.17952v1 Announce Type: new Abstract: Counting objects in an image is a task applicable across many domains. For instance, crowd counting, inventory counting, and cell counting have been the

researcharxiv-cs-cv
19 May 2026
Safety

Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models

DGX agent

arXiv:2605.18413v1 Announce Type: new Abstract: Automated structural health monitoring is essential to prevent catastrophic infrastructure failures. Precise, pixel-level defect segmentation is needed

safetyarxiv-cs-cv
19 May 2026
← Previous
1…158159160161162…263
Next →