AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlog
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Model Releases

Tabular GANs for uneven distribution

DGX agent

arXiv:2010.00638v2 Announce Type: replace-cross Abstract: Generative models for tabular data have evolved rapidly beyond Generative Adversarial Networks (GANs). While GANs pioneered synthetic tabular

model-releasesarxiv-cs-cv
10 Apr 2026
Research

Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation

DGX agent
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

arXiv:2604.07916v1 Announce Type: new Abstract: Referring Expression Segmentation (RES) aims to segment image regions described by natural-language expressions, serving as a bridge between vision and

researcharxiv-cs-cv
10 Apr 2026
Model Releases

Tensor-Augmented Convolutional Neural Networks: Enhancing Expressivity with Generic Tensor Kernels

DGX agent

arXiv:2604.08072v1 Announce Type: new Abstract: Convolutional Neural Networks (CNNs) excel at extracting local features hierarchically, but their performance in capturing complex correlations hinges h

model-releasesarxiv-cs-cv
10 Apr 2026
Research

The Persistence of Cultural Memory: Investigating Multimodal Iconicity in Diffusion Models

DGX agent

arXiv:2511.11435v3 Announce Type: replace Abstract: The ambiguity between generalization and memorization in TTI diffusion models becomes pronounced when prompts invoke culturally shared visual refere

researcharxiv-cs-cv
10 Apr 2026
Research

The Weaponization of Computer Vision: Tracing Military-Surveillance Ties through Conference Sponsorship

DGX agent

arXiv:2604.07803v1 Announce Type: cross Abstract: Computer vision, a core domain of artificial intelligence (AI), is the field that enables the computational analysis, understanding, and generation of

researcharxiv-cs-cv
10 Apr 2026
Research

Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM Decoding

DGX agent

arXiv:2503.10183v4 Announce Type: replace Abstract: Existing vision-language models (VLMs) often suffer from visual hallucination, where the generated responses contain inaccuracies that are not groun

researcharxiv-cs-cv
10 Apr 2026
Model Releases

Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval

DGX agent

arXiv:2512.08410v2 Announce Type: replace Abstract: Due to excessive memory overhead, most Multimodal Large Language Models (MLLMs) can only process videos of limited frames. In this paper, we propose

model-releasesarxiv-cs-cv
10 Apr 2026
Research

Training-free Spatially Grounded Geometric Shape Encoding (Technical Report)

DGX agent

arXiv:2604.07522v1 Announce Type: new Abstract: Positional encoding has become the de facto standard for grounding deep neural networks on discrete point-wise positions, and it has achieved remarkable

researcharxiv-cs-cv
10 Apr 2026
Research

U-CECE: A Universal Multi-Resolution Framework for Conceptual Counterfactual Explanations

DGX agent

arXiv:2604.08295v1 Announce Type: cross Abstract: As AI models grow more complex, explainability is essential for building trust, yet concept-based counterfactual methods still face a trade-off betwee

researcharxiv-cs-cv
10 Apr 2026
Tutorials

Understanding Task Transfer in Vision-Language Models

DGX agent

arXiv:2511.18787v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) perform well on multimodal benchmarks but lag behind humans and specialized models on visual perception tasks like dep

tutorialsarxiv-cs-cv
10 Apr 2026
Research

Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator

DGX agent

arXiv:2604.08121v1 Announce Type: new Abstract: Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher co

researcharxiv-cs-cv
10 Apr 2026
Applications

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models

DGX agent

arXiv:2602.20231v2 Announce Type: replace-cross Abstract: Latent action representations learned from unlabeled videos have recently emerged as a promising paradigm for pretraining vision-language-acti

applicationsarxiv-cs-cv
10 Apr 2026
Model Releases

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding

DGX agent

arXiv:2604.08522v1 Announce Type: new Abstract: Video temporal grounding (VTG) is typically tackled with dataset-specific models that transfer poorly across domains and query styles. Recent efforts to

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents

DGX agent

arXiv:2603.15118v2 Announce Type: replace Abstract: We introduce VAREX (VARied-schema EXtraction), a benchmark for evaluating multimodal foundation models on structured data extraction from government

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs

DGX agent

arXiv:2509.08016v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) face a critical bottleneck: increasing the number of input frames to capture fine-grained temporal detail le

model-releasesarxiv-cs-cv
10 Apr 2026
Applications

Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment

DGX agent

arXiv:2604.08212v1 Announce Type: new Abstract: General-purpose vision-language models demonstrate strong performance in everyday domains but struggle with specialized technical fields requiring preci

applicationsarxiv-cs-cv
10 Apr 2026
Model Releases

Visually-grounded Humanoid Agents

DGX agent

arXiv:2604.08509v1 Announce Type: new Abstract: Digital human generation has been studied for decades and supports a wide range of real-world applications. However, most existing systems are passively

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

VSAS-BENCH: Real-Time Evaluation of Visual Streaming Assistant Models

DGX agent

arXiv:2604.07634v1 Announce Type: new Abstract: Streaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core

model-releasesarxiv-cs-cv
10 Apr 2026
Research

Weakly-Supervised Lung Nodule Segmentation via Training-Free Guidance of 3D Rectified Flow

DGX agent

arXiv:2604.08313v1 Announce Type: new Abstract: Dense annotations, such as segmentation masks, are expensive and time-consuming to obtain, especially for 3D medical images where expert voxel-wise labe

researcharxiv-cs-cv
10 Apr 2026
Research

Weight Group-wise Post-Training Quantization for Medical Foundation Model

DGX agent

arXiv:2604.07674v1 Announce Type: new Abstract: Foundation models have achieved remarkable results in medical image analysis. However, its large network architecture and high computational complexity

researcharxiv-cs-cv
10 Apr 2026
Research

When Fine-Tuning Changes the Evidence: Architecture-Dependent Semantic Drift in Chest X-Ray Explanations

DGX agent

arXiv:2604.08513v1 Announce Type: new Abstract: Transfer learning followed by fine-tuning is widely adopted in medical image classification due to consistent gains in diagnostic performance. However,

researcharxiv-cs-cv
10 Apr 2026
Safety

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

DGX agent

arXiv:2604.08546v1 Announce Type: new Abstract: Text-to-video diffusion models have enabled open-ended video synthesis, but often struggle with generating the correct number of objects specified in a

safetyarxiv-cs-cv
10 Apr 2026
Research

WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models

DGX agent

arXiv:2604.07957v1 Announce Type: cross Abstract: Vision-language models (VLMs) and generative world models are opening new opportunities for embodied navigation. VLMs are increasingly used as direct

researcharxiv-cs-cv
10 Apr 2026
Research

WUTDet: A 100K-Scale Ship Detection Dataset and Benchmarks with Dense Small Objects

DGX agent

arXiv:2604.07759v1 Announce Type: new Abstract: Ship detection for navigation is a fundamental perception task in intelligent waterway transportation systems. However, existing public ship detection d

researcharxiv-cs-cv
10 Apr 2026
Model Releases

You Point, I Learn: Online Adaptation of Interactive Segmentation Models for Handling Distribution Shifts in Medical Imaging

DGX agent

arXiv:2503.06717v3 Announce Type: replace Abstract: Interactive segmentation uses real-time user inputs, such as mouse clicks, to iteratively refine model predictions. Although not originally designed

model-releasesarxiv-cs-cv
10 Apr 2026
Hardware

ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training

DGX agent

arXiv:2603.04385v3 Announce Type: replace Abstract: Feed-forward transformer models have driven rapid progress in 3D vision, but state-of-the-art methods such as VGGT and pi^3 have a computational

hardwarearxiv-cs-cv
10 Apr 2026
Syntheses

Wiki Lint Report — 2026-04-19

DGX agent

Automated lint: 43 errors, 9 warnings, 3 info

linthealth-checkautomated
19 Apr 2026
Syntheses

Synthesis: Arxiv-Cs-Ai

DGX agent

Auto-generated synthesis of 1623 entries about arxiv-cs-ai

synthesisarxiv-cs-aiauto-generated
16 Apr 2026
Syntheses

Synthesis: Arxiv-Cs-Cl

DGX agent

Auto-generated synthesis of 505 entries about arxiv-cs-cl

synthesisarxiv-cs-clauto-generated
16 Apr 2026
Syntheses

Synthesis: Arxiv-Cs-Lg

DGX agent

Auto-generated synthesis of 663 entries about arxiv-cs-lg

synthesisarxiv-cs-lgauto-generated
16 Apr 2026
← Previous
1…257258259
Next →