AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
15 Apr 2026

RPG-SAM: Reliability-Weighted Prototypes and Geometric Adaptive Threshold Selection for Training-Free One-Shot Polyp Segmentation

Model ReleasesDGX agent

arXiv:2603.07436v2 Announce Type: replace Abstract: Training-free one-shot segmentation offers a scalable alternative to expert annotations where knowledge is often transferred from support images and

RSGMamba: Reliability-Aware Self-Gated State Space Model for Multimodal Semantic Segmentation

ResearchDGX agent

arXiv:2604.12319v1 Announce Type: new Abstract: Multimodal semantic segmentation has emerged as a powerful paradigm for enhancing scene understanding by leveraging complementary information from multi

SAM3-I: Segment Anything with Instructions

SafetyDGX agent

arXiv:2512.04585v3 Announce Type: replace Abstract: Segment Anything Model 3 (SAM3) advances open-vocabulary segmentation through promptable concept segmentation, enabling users to segment all instanc


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Scalable Trajectory Generation for Whole-Body Mobile Manipulation

HardwareDGX agent

arXiv:2604.12565v1 Announce Type: cross Abstract: Robots deployed in unstructured environments must coordinate whole-body motion -- simultaneously moving a mobile base and arm -- to interact with the

Scaling Exposes the Trigger: Input-Level Backdoor Detection in Text-to-Image Diffusion Models via Cross-Attention Scaling

ResearchDGX agent

arXiv:2604.12446v1 Announce Type: cross Abstract: Text-to-image (T2I) diffusion models have achieved remarkable success in image synthesis, but their reliance on large-scale data and open ecosystems i

Scaling In-Context Segmentation with Hierarchical Supervision

ResearchDGX agent

arXiv:2604.12752v1 Announce Type: new Abstract: In-context learning (ICL) enables medical image segmentation models to adapt to new anatomical structures from limited examples, reducing the clinical a

SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents

Model ReleasesDGX agent

arXiv:2510.10073v2 Announce Type: replace-cross Abstract: Large vision-language model (LVLM)-based web agents are emerging as powerful tools for automating complex online tasks. However, when deployed

See, Point, Refine: Multi-Turn Approach to GUI Grounding with Visual Feedback

Model ReleasesDGX agent

arXiv:2604.13019v1 Announce Type: new Abstract: Computer Use Agents (CUAs) fundamentally rely on graphical user interface (GUI) grounding to translate language instructions into executable screen acti

Self-Adversarial One Step Generation via Condition Shifting

Model ReleasesDGX agent

arXiv:2604.12322v1 Announce Type: new Abstract: The push for efficient text to image synthesis has moved the field toward one step sampling, yet existing methods still face a three way tradeoff among

SinkSAM-Net: Knowledge-Driven Self-Supervised Sinkhole Segmentation Using Topographic Priors and Segment Anything Model

Model ReleasesDGX agent

arXiv:2410.01473v2 Announce Type: replace Abstract: Soil sinkholes significantly influence soil degradation, infrastructure vulnerability, and landscape evolution. However, their irregular shapes, com

SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks

Model ReleasesDGX agent

arXiv:2506.14512v4 Announce Type: replace Abstract: Large Language Models (LLMs) have undergone rapid progress, largely attributed to reinforcement learning on complex reasoning tasks. In contrast, wh

SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model

Model ReleasesDGX agent

arXiv:2511.22039v3 Announce Type: replace Abstract: This paper introduces a novel architecture for trajectory-conditioned forecasting of future 3D scene occupancy. In contrast to methods that rely on

Spatial-Spectral Adaptive Fidelity and Noise Prior Reduction Guided Hyperspectral Image Denoising

Local AiDGX agent

arXiv:2604.12600v1 Announce Type: new Abstract: The core challenge of hyperspectral image denoising is striking the right balance between data fidelity and noise prior modeling. Most existing methods

StructDiff: A Structure-Preserving and Spatially Controllable Diffusion Model for Single-Image Generation

TutorialsDGX agent

arXiv:2604.12575v1 Announce Type: new Abstract: This paper introduces StructDiff, a generative framework based on a single-scale diffusion model for single-image generation. Single-image generation ai

Style-Decoupled Adaptive Routing Network for Underwater Image Enhancement

Model ReleasesDGX agent

arXiv:2604.12257v1 Announce Type: new Abstract: Underwater Image Enhancement (UIE) is essential for robust visual perception in marine applications. However, existing methods predominantly rely on uni

SubFlow: Sub-mode Conditioned Flow Matching for Diverse One-Step Generation

TutorialsDGX agent

arXiv:2604.12273v1 Announce Type: cross Abstract: Flow matching has emerged as a powerful generative framework, with recent few-step methods achieving remarkable inference acceleration. However, we id

Subspace-Guided Feature Reconstruction for Unsupervised Anomaly Localization

Model ReleasesDGX agent

arXiv:2309.13904v3 Announce Type: replace Abstract: Unsupervised anomaly localization aims to identify anomalous regions that deviate from normal sample patterns. Most recent methods perform feature m

SynthPix: A lightspeed PIV image generator

Model ReleasesDGX agent

arXiv:2512.09664v2 Announce Type: replace-cross Abstract: We describe SynthPix, a synthetic image generator for Particle Image Velocimetry (PIV) with a focus on performance and parallelism on accelera

T2I-BiasBench: A Multi-Metric Framework for Auditing Demographic and Cultural Bias in Text-to-Image Models

Model ReleasesDGX agent

arXiv:2604.12481v1 Announce Type: new Abstract: Text-to-image (T2I) generative models achieve impressive visual fidelity but inherit and amplify demographic imbalances and cultural biases embedded in

Task Alignment: A simple and effective proxy for model merging in computer vision

SafetyDGX agent

arXiv:2604.12935v1 Announce Type: new Abstract: Efficiently merging several models fine-tuned for different tasks, but stemming from the same pretrained base model, is of great practical interest. Des

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

Local AiDGX agent

arXiv:2512.03963v3 Announce Type: replace Abstract: Enhancing the temporal understanding of Multimodal Large Language Models (MLLMs) is essential for advancing long-form video analysis, enabling tasks

Time-reversed Flow Matching with Worst Transport in High-dimensional Latent Space for Image Anomaly Detection

Local AiDGX agent

arXiv:2508.05461v3 Announce Type: replace Abstract: Likelihood-based deep generative models have been widely investigated for Image Anomaly Detection (IAD), particularly Normalizing Flows, yet their s

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

Model ReleasesDGX agent

arXiv:2604.12012v1 Announce Type: new Abstract: Recent progress in vision-language pretraining has enabled significant improvements to many downstream computer vision applications, such as classificat

Toward Efficient and Robust Behavior Models for Multi-Agent Driving Simulation

Local AiDGX agent

arXiv:2512.05812v5 Announce Type: replace-cross Abstract: Scalable multi-agent driving simulation requires behavior models that are both realistic and computationally efficient. We address this by opt

Towards Interpretable Foundation Models for Retinal Fundus Images

Local AiDGX agent

arXiv:2603.18846v2 Announce Type: replace Abstract: Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically via self-supervised learning (SSL

Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priors

ResearchDGX agent

arXiv:2604.12309v1 Announce Type: new Abstract: We present a novel method for generating geometrically realistic and consistent orbital videos from a single image of an object. Existing video generati

Ultra-low-light computer vision using trained photon correlations

TutorialsDGX agent

arXiv:2604.11993v1 Announce Type: new Abstract: Illumination using correlated photon sources has been established as an approach to allowing high-fidelity images to be reconstructed from noisy camera

Uncertainty-Aware Image Classification In Biomedical Imaging Using Spectral-normalized Neural Gaussian Processes

SafetyDGX agent

arXiv:2602.02370v2 Announce Type: replace Abstract: Accurate histopathologic interpretation is key for clinical decision-making; however, current deep learning models for digital pathology are often o

UniMark: Unified Adaptive Multi-bit Watermarking for Autoregressive Image Generators

ResearchDGX agent

arXiv:2604.11843v1 Announce Type: new Abstract: Invisible watermarking for autoregressive (AR) image generation has recently gained attention as a means of protecting image ownership and tracing AI-ge

Unlocking the Potential of Grounding DINO in Videos: Parameter-Efficient Adaptation for Limited-Data Spatial-Temporal Localization

Model ReleasesDGX agent

arXiv:2604.12346v1 Announce Type: new Abstract: Spatio-temporal video grounding (STVG) aims to localize queried objects within dynamic video segments. Prevailing fully-trained approaches are notorious

V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos

Model ReleasesDGX agent

arXiv:2604.11913v1 Announce Type: new Abstract: Nutrition estimation of meals from visual data is an important problem for dietary monitoring and computational health, but existing approaches largely

Variational Autoencoding Discrete Diffusion with Enhanced Dimensional Correlations Modeling

ResearchDGX agent

arXiv:2505.17384v2 Announce Type: replace-cross Abstract: Discrete diffusion models have recently shown great promise for modeling complex discrete data, with masked diffusion models (MDMs) offering a

VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization

Local AiDGX agent

arXiv:2604.12887v1 Announce Type: new Abstract: Visual tokenizers map high-dimensional raw pixels into a compressed representation for downstream modeling. Beyond compression, tokenizers dictate what

VidTAG: Temporally Aligned Video to GPS Geolocalization with Denoising Sequence Prediction at a Global Scale

ResearchDGX agent

arXiv:2604.12159v1 Announce Type: new Abstract: The task of video geolocalization aims to determine the precise GPS coordinates of a video's origin and map its trajectory; with applications in forensi

ViLL-E: Video LLM Embeddings for Retrieval

Local AiDGX agent

arXiv:2604.12148v1 Announce Type: new Abstract: Video Large Language Models (VideoLLMs) excel at video understanding tasks where outputs are textual, such as Video Question Answering and Video Caption

Vision Transformers Need More Than Registers

ResearchDGX agent

arXiv:2602.22394v2 Announce Type: replace Abstract: Vision Transformers (ViTs), when pre-trained on large-scale data, provide general-purpose representations for diverse downstream tasks. However, art

Visual Diffusion Models are Geometric Solvers

ResearchDGX agent

arXiv:2510.21697v2 Announce Type: replace Abstract: In this paper we show that visual diffusion models can serve as effective geometric solvers: they can directly reason about geometric problems by wo

VPTracker: Global Vision-Language Tracking via Visual Prompt

Local AiDGX agent

arXiv:2512.22799v2 Announce Type: replace Abstract: Vision-Language Tracking aims to continuously localize objects described by a visual template and a language description. Existing methods, however,

Whole-Body Mobile Manipulation using Offline Reinforcement Learning on Sub-optimal Controllers

SafetyDGX agent

arXiv:2604.12509v1 Announce Type: cross Abstract: Mobile Manipulation (MoMa) of articulated objects, such as opening doors, drawers, and cupboards, demands simultaneous, whole-body coordination betwee

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding

ResearchDGX agent

arXiv:2604.12358v1 Announce Type: new Abstract: Recently, visual token pruning has been studied to handle the vast number of visual tokens in Multimodal Large Language Models. However, we observe that

14 Apr 2026

3D Multi-View Stylization with Pose-Free Correspondences Matching for Robust 3D Geometry Preservation

SafetyDGX agent

arXiv:2604.09639v1 Announce Type: new Abstract: Artistic style transfer is well studied for images and videos, but extending it to multi-view 3D scenes remains difficult because stylization can disrup

3DTV: A Feedforward Interpolation Network for Real-Time View Synthesis

ResearchDGX agent

arXiv:2604.11211v1 Announce Type: new Abstract: Real-time free-viewpoint rendering requires balancing multi-camera redundancy with the latency constraints of interactive applications. We address this

A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation

Model ReleasesDGX agent

arXiv:2604.10456v1 Announce Type: new Abstract: The surging demand for adapting long-form cinematic content into short videos has motivated the need for versatile automatic video compilation systems.

A Comparative Study of Modern Object Detectors for Robust Apple Detection in Orchard Imagery

Model ReleasesDGX agent

arXiv:2604.09996v1 Announce Type: new Abstract: Accurate apple detection in orchard images is important for yield prediction, fruit counting, robotic harvesting, and crop monitoring. However, changing

A Comparison of Multi-View Stereo Methods for Photogrammetric 3D Reconstruction: From Traditional to Learning-Based Approaches

ResearchDGX agent

arXiv:2604.10246v1 Announce Type: new Abstract: Photogrammetric 3D reconstruction has long relied on traditional Structure-from-Motion (SfM) and Multi-View Stereo (MVS) methods, which provide high acc

A Data-driven Loss Weighting Scheme across Heterogeneous Tasks for Image Denoising

TutorialsDGX agent

arXiv:2301.06081v4 Announce Type: replace-cross Abstract: In a variational denoising model, weight in the data fidelity term plays the role of enhancing the noise-removal capability. It is profoundly

A Deep Equilibrium Network for Hyperspectral Unmixing

ApplicationsDGX agent

arXiv:2604.11279v1 Announce Type: new Abstract: Hyperspectral unmixing (HU) is crucial for analyzing hyperspectral imagery, yet achieving accurate unmixing remains challenging. While traditional metho

A Faster Path to Continual Learning

ResearchDGX agent

arXiv:2604.11064v1 Announce Type: cross Abstract: Continual Learning (CL) aims to train neural networks on a dynamic stream of tasks without forgetting previously learned knowledge. Among optimization

A Modular Zero-Shot Pipeline for Accident Detection, Localization, and Classification in Traffic Surveillance Video

Local AiDGX agent

arXiv:2604.09685v1 Announce Type: new Abstract: We describe a zero-shot pipeline developed for the ACCIDENT @ CVPR 2026 challenge. The challenge requires predicting when, where, and what type of traff

A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation

ResearchDGX agent

arXiv:2508.09977v4 Announce Type: replace Abstract: In the context of novel view synthesis, 3D Gaussian Splatting (3DGS) has recently emerged as an efficient and competitive counterpart to Neural Radi

A Survey on Deep Learning Techniques for Action Anticipation

AgentsDGX agent

arXiv:2309.17257v2 Announce Type: replace Abstract: The ability to anticipate possible future human actions is essential for a wide range of applications, including autonomous driving and human-robot

A Two-Stage Dual-Modality Model for Facial Emotional Expression Recognition

Local AiDGX agent

arXiv:2603.12221v2 Announce Type: replace Abstract: This paper addresses the expression (EXPR) recognition challenge in the 10th Affective Behavior Analysis in-the-Wild (ABAW) workshop and competition

A3-FPN: Asymptotic Content-Aware Pyramid Attention Network for Dense Visual Prediction

ResearchDGX agent

arXiv:2604.10210v1 Announce Type: new Abstract: Learning multi-scale representations is the common strategy to tackle object scale variation in dense prediction tasks. Although existing feature pyrami

ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents

Local AiDGX agent

arXiv:2604.10096v1 Announce Type: new Abstract: Current embodied intelligent systems still face a substantial gap between high-level reasoning and low-level physical execution in open-world environmen

AC-MIL: Weakly Supervised Atrial LGE-MRI Quality Assessment via Adversarial Concept Disentanglement

Model ReleasesDGX agent

arXiv:2604.10303v1 Announce Type: new Abstract: High-quality Late Gadolinium Enhancement (LGE) MRI can be helpful for atrial fibrillation management, yet scan quality is frequently compromised by pati

ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models

TutorialsDGX agent

arXiv:2511.18082v3 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) models have shown impressive flexibility and generalization, yet their deployment in robotic manipulation remain

Active Diffusion Matching: Score-based Iterative Alignment of Cross-Modal Retinal Images

SafetyDGX agent

arXiv:2604.10084v1 Announce Type: new Abstract: Objective: The study aims to address the challenge of aligning Standard Fundus Images (SFIs) and Ultra-Widefield Fundus Images (UWFIs), which is difficu

Adversarial Video Promotion Against Text-to-Video Retrieval

ResearchDGX agent

arXiv:2508.06964v3 Announce Type: replace Abstract: Thanks to the development of cross-modal models, text-to-video retrieval (T2VR) is advancing rapidly, but its robustness remains largely unexamined.

Affostruction: 3D Affordance Grounding with Generative Reconstruction

ResearchDGX agent

arXiv:2601.09211v2 Announce Type: replace Abstract: This paper addresses the problem of affordance grounding from RGBD images of an object, which aims to localize surface regions corresponding to a te

Agentic Video Generation: From Text to Executable Event Graphs via Tool-Constrained LLM Planning

SafetyDGX agent

arXiv:2604.10383v1 Announce Type: new Abstract: Existing multi-agent video generation systems use LLM agents to orchestrate neural video generators, producing visually impressive but semantically unre

← Previous
1…193194195196197…207
Next →