AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
10 Apr 2026

SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses

Model ReleasesDGX agent

arXiv:2602.22683v2 Announce Type: replace Abstract: The rapid advancement of AI-powered smart glasses-one of the hottest wearable devices-has unlocked new frontiers for multimodal interaction, with Vi

SurfelSplat: Learning Efficient and Generalizable Gaussian Surfel Representations for Sparse-View Surface Reconstruction

ResearchDGX agent

arXiv:2604.08370v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has demonstrated impressive performance in 3D scene reconstruction. Beyond novel view synthesis, it shows great potential f

SVGFusion: A VAE-Diffusion Transformer for Vector Graphic Generation

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2412.10437v3 Announce Type: replace Abstract: Generating high-quality Scalable Vector Graphics (SVGs) from text remains a significant challenge. Existing LLM-based models that generate SVG code

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation

ResearchDGX agent

arXiv:2604.08405v1 Announce Type: new Abstract: Diffusion-based audio-driven talking-head generation enables realistic portrait animation, but also introduces risks of misuse, such as fraud and misinf

T-Gated Adapter: A Lightweight Temporal Adapter for Vision-Language Medical Segmentation

Model ReleasesDGX agent

arXiv:2604.08167v1 Announce Type: new Abstract: Medical image segmentation traditionally relies on fully supervised 3D architectures that demand a large amount of dense, voxel-level annotations from c

Tabular GANs for uneven distribution

Model ReleasesDGX agent

arXiv:2010.00638v2 Announce Type: replace-cross Abstract: Generative models for tabular data have evolved rapidly beyond Generative Adversarial Networks (GANs). While GANs pioneered synthetic tabular

Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation

ResearchDGX agent

arXiv:2604.07916v1 Announce Type: new Abstract: Referring Expression Segmentation (RES) aims to segment image regions described by natural-language expressions, serving as a bridge between vision and

Tensor-Augmented Convolutional Neural Networks: Enhancing Expressivity with Generic Tensor Kernels

Model ReleasesDGX agent

arXiv:2604.08072v1 Announce Type: new Abstract: Convolutional Neural Networks (CNNs) excel at extracting local features hierarchically, but their performance in capturing complex correlations hinges h

The Persistence of Cultural Memory: Investigating Multimodal Iconicity in Diffusion Models

ResearchDGX agent

arXiv:2511.11435v3 Announce Type: replace Abstract: The ambiguity between generalization and memorization in TTI diffusion models becomes pronounced when prompts invoke culturally shared visual refere

The Weaponization of Computer Vision: Tracing Military-Surveillance Ties through Conference Sponsorship

ResearchDGX agent

arXiv:2604.07803v1 Announce Type: cross Abstract: Computer vision, a core domain of artificial intelligence (AI), is the field that enables the computational analysis, understanding, and generation of

Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM Decoding

ResearchDGX agent

arXiv:2503.10183v4 Announce Type: replace Abstract: Existing vision-language models (VLMs) often suffer from visual hallucination, where the generated responses contain inaccuracies that are not groun

Towards Effective Long Video Understanding of Multimodal Large Language Models via One-shot Clip Retrieval

Model ReleasesDGX agent

arXiv:2512.08410v2 Announce Type: replace Abstract: Due to excessive memory overhead, most Multimodal Large Language Models (MLLMs) can only process videos of limited frames. In this paper, we propose

Training-free Spatially Grounded Geometric Shape Encoding (Technical Report)

ResearchDGX agent

arXiv:2604.07522v1 Announce Type: new Abstract: Positional encoding has become the de facto standard for grounding deep neural networks on discrete point-wise positions, and it has achieved remarkable

U-CECE: A Universal Multi-Resolution Framework for Conceptual Counterfactual Explanations

ResearchDGX agent

arXiv:2604.08295v1 Announce Type: cross Abstract: As AI models grow more complex, explainability is essential for building trust, yet concept-based counterfactual methods still face a trade-off betwee

Understanding Task Transfer in Vision-Language Models

TutorialsDGX agent

arXiv:2511.18787v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) perform well on multimodal benchmarks but lag behind humans and specialized models on visual perception tasks like dep

Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator

ResearchDGX agent

arXiv:2604.08121v1 Announce Type: new Abstract: Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher co

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models

ApplicationsDGX agent

arXiv:2602.20231v2 Announce Type: replace-cross Abstract: Latent action representations learned from unlabeled videos have recently emerged as a promising paradigm for pretraining vision-language-acti

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding

Model ReleasesDGX agent

arXiv:2604.08522v1 Announce Type: new Abstract: Video temporal grounding (VTG) is typically tackled with dataset-specific models that transfer poorly across domains and query styles. Recent efforts to

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents

Model ReleasesDGX agent

arXiv:2603.15118v2 Announce Type: replace Abstract: We introduce VAREX (VARied-schema EXtraction), a benchmark for evaluating multimodal foundation models on structured data extraction from government

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs

Model ReleasesDGX agent

arXiv:2509.08016v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) face a critical bottleneck: increasing the number of input frames to capture fine-grained temporal detail le

Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment

ApplicationsDGX agent

arXiv:2604.08212v1 Announce Type: new Abstract: General-purpose vision-language models demonstrate strong performance in everyday domains but struggle with specialized technical fields requiring preci

Visually-grounded Humanoid Agents

Model ReleasesDGX agent

arXiv:2604.08509v1 Announce Type: new Abstract: Digital human generation has been studied for decades and supports a wide range of real-world applications. However, most existing systems are passively

VSAS-BENCH: Real-Time Evaluation of Visual Streaming Assistant Models

Model ReleasesDGX agent

arXiv:2604.07634v1 Announce Type: new Abstract: Streaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core

Weakly-Supervised Lung Nodule Segmentation via Training-Free Guidance of 3D Rectified Flow

ResearchDGX agent

arXiv:2604.08313v1 Announce Type: new Abstract: Dense annotations, such as segmentation masks, are expensive and time-consuming to obtain, especially for 3D medical images where expert voxel-wise labe

Weight Group-wise Post-Training Quantization for Medical Foundation Model

ResearchDGX agent

arXiv:2604.07674v1 Announce Type: new Abstract: Foundation models have achieved remarkable results in medical image analysis. However, its large network architecture and high computational complexity

When Fine-Tuning Changes the Evidence: Architecture-Dependent Semantic Drift in Chest X-Ray Explanations

ResearchDGX agent

arXiv:2604.08513v1 Announce Type: new Abstract: Transfer learning followed by fine-tuning is widely adopted in medical image classification due to consistent gains in diagnostic performance. However,

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

SafetyDGX agent

arXiv:2604.08546v1 Announce Type: new Abstract: Text-to-video diffusion models have enabled open-ended video synthesis, but often struggle with generating the correct number of objects specified in a

WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models

ResearchDGX agent

arXiv:2604.07957v1 Announce Type: cross Abstract: Vision-language models (VLMs) and generative world models are opening new opportunities for embodied navigation. VLMs are increasingly used as direct

WUTDet: A 100K-Scale Ship Detection Dataset and Benchmarks with Dense Small Objects

ResearchDGX agent

arXiv:2604.07759v1 Announce Type: new Abstract: Ship detection for navigation is a fundamental perception task in intelligent waterway transportation systems. However, existing public ship detection d

You Point, I Learn: Online Adaptation of Interactive Segmentation Models for Handling Distribution Shifts in Medical Imaging

Model ReleasesDGX agent

arXiv:2503.06717v3 Announce Type: replace Abstract: Interactive segmentation uses real-time user inputs, such as mouse clicks, to iteratively refine model predictions. Although not originally designed

ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training

HardwareDGX agent

arXiv:2603.04385v3 Announce Type: replace Abstract: Feed-forward transformer models have driven rapid progress in 3D vision, but state-of-the-art methods such as VGGT and pi^3 have a computational

19 Apr 2026

Wiki Lint Report — 2026-04-19

SynthesesDGX agent

Automated lint: 43 errors, 9 warnings, 3 info

16 Apr 2026

Synthesis: Arxiv-Cs-Ai

SynthesesDGX agent

Auto-generated synthesis of 1623 entries about arxiv-cs-ai

Synthesis: Arxiv-Cs-Cl

SynthesesDGX agent

Auto-generated synthesis of 505 entries about arxiv-cs-cl

Synthesis: Arxiv-Cs-Lg

SynthesesDGX agent

Auto-generated synthesis of 663 entries about arxiv-cs-lg

← Previous
1…207208209
Next →