AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
21 Apr 2026

Privatar: Scalable Privacy-preserving Multi-user VR via Secure Offloading

Local AiDGX agent

arXiv:2604.17476v1 Announce Type: cross Abstract: Multi-user virtual reality enables immersive interaction. However, rendering avatars for numerous participants on each headset incurs prohibitive comp

Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions

Model ReleasesDGX agent

arXiv:2604.18459v1 Announce Type: new Abstract: Visual agents operating in the wild must respond to queries precisely when sufficient evidence first appears in a video stream, a critical capability th

Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.17126v1 Announce Type: new Abstract: Vision-language models enable open-vocabulary object grounding through natural language queries, under the implicit assumption that semantically equival

Prompting Foundation Models for Zero-Shot Ship Instance Segmentation in SAR Imagery

Model ReleasesDGX agent

arXiv:2604.17920v1 Announce Type: new Abstract: Synthetic Aperture Radar (SAR) plays a critical role in maritime surveillance, yet deep learning for SAR analysis is limited by the lack of pixel-level

ProtoCLIP: Prototype-Aligned Latent Refinement for Robust Zero-Shot Chest X-Ray Classification

SafetyDGX agent

arXiv:2604.18444v1 Announce Type: cross Abstract: Zero-shot vision-language models (VLMs) have shown promise for chest radiograph classification, but their performance is often limited by confounding

Q-DeepSight: Incentivizing Thinking with Images for Image Quality Assessment and Refinement

Local AiDGX agent

arXiv:2604.16858v1 Announce Type: new Abstract: Image Quality Assessment (IQA) models are increasingly deployed as perceptual critics to guide generative models and image restoration. This role demand

QuadSync: Quadrifocal Tensor Synchronization via Tucker Decomposition

ResearchDGX agent

arXiv:2602.22639v2 Announce Type: replace Abstract: In structure from motion, quadrifocal tensors capture more information than their pairwise counterparts (essential matrices), yet they have often be

R-FLoRA: Residual-Statistic-Gated Low-Rank Adaptation for Single-Image Face Morphing Attack Detection

Local AiDGX agent

arXiv:2604.17321v1 Announce Type: new Abstract: Face morphing attacks pose a substantial risk to the reliability of face recognition systems used in passport issuance, border control, and digital iden

R3D2: Realistic 3D Asset Insertion via Diffusion for Autonomous Driving Simulation

SafetyDGX agent

arXiv:2506.07826v2 Announce Type: replace Abstract: Validating autonomous driving (AD) systems requires diverse and safety-critical testing, making photorealistic virtual environments essential. Tradi

RainFusion2.0: Temporal-Spatial Awareness and Hardware-Efficient Block-wise Sparse Attention

HardwareDGX agent

arXiv:2512.24086v2 Announce Type: replace Abstract: In video and image generation tasks, Diffusion Transformer (DiT) models incur extremely high computational costs due to attention mechanisms, which

Re^2MoGen: Open-Vocabulary Motion Generation via LLM Reasoning and Physics-Aware Refinement

ResearchDGX agent

arXiv:2604.17807v1 Announce Type: new Abstract: Text-to-motion (T2M) generation aims to control the behavior of a target character via textual descriptions. Leveraging text-motion paired datasets, exi

Real-Time Cellist Postural Evaluation With On-Device Computer Vision

Local AiDGX agent

arXiv:2604.17530v1 Announce Type: cross Abstract: Posture is a critical factor for beginning instrumental learners. Most students receive instruction only once a week, and during the intervals between

Real-Time Visual Attribution Streaming in Thinking Model

ResearchDGX agent

arXiv:2604.16587v1 Announce Type: new Abstract: We present an amortized framework for real-time visual attribution streaming in multimodal thinking models. When these models generate code from a scree

ReCap: Lightweight Referential Grounding for Coherent Story Visualization

Model ReleasesDGX agent

arXiv:2604.18575v1 Announce Type: new Abstract: Story Visualization aims to generate a sequence of images that faithfully depicts a textual narrative that preserve character identity, spatial configur

Reducing Peak Memory Usage for Modern Multimodal Large Language Model Pipelines

ResearchDGX agent

arXiv:2604.16734v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have recently demonstrated strong capabilities in understanding and generating responses from diverse visual in

ReFineVLA: Multimodal Reasoning-Aware Generalist Robotic Policies via Teacher-Guided Fine-Tuning

Model ReleasesDGX agent

arXiv:2604.17800v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have gained much attention from the research community thanks to their strength in translating multimodal observat

Region-Affinity Attention for Whole-Slide Breast Cancer Classification in Deep Ultraviolet Imaging

ResearchDGX agent

arXiv:2604.17222v1 Announce Type: new Abstract: Breast cancer diagnosis demands rapid and precise tools, yet traditional histopathological methods often fall short in intra-operative settings. Deep Ul

Region-Grounded Report Generation for 3D Medical Imaging: A Fine-Grained Dataset and Graph-Enhanced Framework

Local AiDGX agent

arXiv:2604.18145v1 Announce Type: new Abstract: Automated medical report generation for 3D PET/CT imaging is fundamentally challenged by the high-dimensional nature of volumetric data and a critical s

Relative State Estimation using Event-Based Propeller Sensing

AgentsDGX agent

arXiv:2604.18289v1 Announce Type: cross Abstract: Autonomous swarms of multi-Unmanned Aerial Vehicle (UAV) system requires an accurate and fast relative state estimation. Although monocular frame-base

RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation

SafetyDGX agent

arXiv:2604.17243v1 Announce Type: new Abstract: A robust Multimodal Large Language Model (MLLM) for Earth Observation should maintain consistent interpretation and reasoning under realistic input vari

Residual Diffusion Bridge Model for Image Restoration

ResearchDGX agent

arXiv:2510.23116v3 Announce Type: replace Abstract: Diffusion bridge models establish probabilistic paths between arbitrary paired distributions and exhibit great potential for universal image restora

Rethinking Cross-Dose PET Denoising: Mitigating Averaging Effects via Residual Noise Learning

TutorialsDGX agent

arXiv:2604.16925v1 Announce Type: new Abstract: Cross-dose denoising for low-dose positron emission tomography (LDPET) has been proposed to address the limited generalization of models trained at a si

Rethinking Post-Unlearning Behavior of Large Vision-Language Models

ResearchDGX agent

arXiv:2506.02541v2 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) can recognize individuals in images and disclose sensitive personal information about them, raising criti

ReTrack: Evidence-Driven Dual-Stream Directional Anchor Calibration Network for Composed Video Retrieval

Model ReleasesDGX agent

arXiv:2604.17898v1 Announce Type: new Abstract: With the rapid growth of video data, Composed Video Retrieval (CVR) has emerged as a novel paradigm in video retrieval and is receiving increasing atten

Revisiting Change VQA in Remote Sensing with Structured and Native Multimodal Qwen Models

Model ReleasesDGX agent

arXiv:2604.18429v1 Announce Type: new Abstract: Change visual question answering (Change VQA) addresses the problem of answering natural-language questions about semantic changes between bi-temporal r

Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models

SafetyDGX agent

arXiv:2604.17415v1 Announce Type: cross Abstract: Reward-based fine-tuning aims to steer a pretrained diffusion or flow-based generative model toward higher-reward samples while remaining close to the

Rewind-IL: Online Failure Detection and State Respawning for Imitation Learning

Local AiDGX agent

arXiv:2604.16683v1 Announce Type: cross Abstract: Imitation learning has enabled robots to acquire complex visuomotor manipulation skills from demonstrations, but deployment failures remain a major ob

Robust Diabetic Retinopathy Grading Using Dual-Resolution Attention-Based Deep Learning with Ordinal Regression

ResearchDGX agent

arXiv:2604.17341v1 Announce Type: new Abstract: Diabetic retinopathy (DR) is a leading cause of vision impairment worldwide, and automated grading systems play a crucial role in large-scale screening

RS-HyRe-R1: A Hybrid Reward Mechanism to Overcome Perceptual Inertia for Remote Sensing Images Understanding

Local AiDGX agent

arXiv:2604.17504v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training substantially improves remote sensing vision-language models (RS-VLMs). However, when handling complex remote

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification

Model ReleasesDGX agent

arXiv:2604.16993v1 Announce Type: cross Abstract: As embodied AI transitions to real-world deployment, the success of the Vision-and-Language Navigation (VLN) task tends to evolve from mere reachabili

S2H-DPO: Hardness-Aware Preference Optimization for Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.18512v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated remarkable progress in single-image understanding, yet effective reasoning across multiple images remain

Saccade Attention Networks: Using Transfer Learning of Attention to Reduce Network Sizes

TutorialsDGX agent

arXiv:2604.16485v1 Announce Type: new Abstract: One of the limitations of transformer networks is the sequence length due to the quadratic nature of the attention matrix. Classical self attention uses

SAND: The Challenge on Speech Analysis for Neurodegenerative Disease Assessment

ResearchDGX agent

arXiv:2604.16445v1 Announce Type: cross Abstract: Recent advances in Artificial Intelligence (AI) and the exploration of noninvasive, objective biomarkers, such as speech signals, have encouraged the

ScenarioControl: Vision-Language Controllable Vectorized Latent Scenario Generation

ResearchDGX agent

arXiv:2604.17147v1 Announce Type: new Abstract: We introduce ScenarioControl, the first vision-language control mechanism for learned driving scenario generation. Given a text prompt or an input image

SciDraw-6K: A Multilingual Scientific Illustration Dataset Generated by Google Gemini

Model ReleasesDGX agent

arXiv:2604.17206v1 Announce Type: new Abstract: We present SciDraw-6K, a curated dataset of 6,291 scientific illustrations synthesized by Google Gemini image-generation models, each paired with prompt

Score-Based Matching with Target Guidance for Cryo-EM Denoising

ResearchDGX agent

arXiv:2604.17734v1 Announce Type: new Abstract: Cryo-electron microscopy (cryo-EM) enables single-particle analysis of biological macromolecules under strict low-dose imaging conditions, but the resul

See Through the Noise: Improving Domain Generalization in Gaze Estimation

SafetyDGX agent

arXiv:2604.16562v1 Announce Type: new Abstract: Generalizable gaze estimation methods have garnered increasing attention due to their critical importance in real-world applications and have achieved s

SegTTA: Training-Free Test-Time Augmentation for Zero-Shot Medical Imaging Segmentation

ResearchDGX agent

arXiv:2604.17451v1 Announce Type: new Abstract: Increasingly advanced data augmentation techniques have greatly aided clinical medical research, increasing data diversity and improving model generaliz

Self-Reasoning Agentic Framework for Narrative Product Grid-Collage Generation

AgentsDGX agent

arXiv:2604.16958v1 Announce Type: new Abstract: Narrative-driven product photography has become a prevalent paradigm in modern marketing, as coherent visual storytelling helps convey product value and

Self-Supervised Super-Resolution for Sentinel-5P Hyperspectral Images

ResearchDGX agent

arXiv:2604.17652v1 Announce Type: new Abstract: Sentinel-5P (S5P) plays a critical role in atmospheric monitoring; however, its spatial resolution limits fine-scale analysis. Existing super-resolution

Semantically Stable Image Composition Analysisvia Saliency and Gradient Vector Flow Fusion

Model ReleasesDGX agent

arXiv:2604.16500v1 Announce Type: new Abstract: The reliable computational assessment of photographic composition requires features that are discriminative of spatial layout yet robust to semantic con

SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection

SafetyDGX agent

arXiv:2604.18476v1 Announce Type: new Abstract: Camera-only 3D object detection has emerged as a cost-effective and scalable alternative to LiDAR for autonomous driving, yet existing methods primarily

SemMorph3D: Unsupervised Semantic-Aware 3D Morphing via Mesh-Guided Gaussians

Model ReleasesDGX agent

arXiv:2510.02034v2 Announce Type: replace Abstract: We introduce METHODNAME, a novel framework for semantic-aware 3D shape and texture morphing directly from multi-view images. While 3D Gaussian Splat

SentiAvatar: Towards Expressive and Interactive Digital Humans

ResearchDGX agent

arXiv:2604.02908v2 Announce Type: replace Abstract: We present SentiAvatar, a framework for building expressive interactive 3D digital humans, and use it to create SuSu, a virtual character that speak

SetFlow: Generating Structured Sets of Representations for Multiple Instance Learning

Model ReleasesDGX agent

arXiv:2604.16362v1 Announce Type: cross Abstract: Data scarcity and weak supervision continue to limit the performance of machine learning models in many real-world applications, such as mammography,

Sharpening Lightweight Models for Generalized Polyp Segmentation: A Boundary Guided Distillation from Foundation Models

SafetyDGX agent

arXiv:2604.17865v1 Announce Type: new Abstract: Automated polyp segmentation is critical for early colorectal cancer detection and its prevention, yet remains challenging due to weak boundaries, large

SHRUG-FM: Reliability-Aware Foundation Models for Earth Observation

Model ReleasesDGX agent

arXiv:2511.10370v2 Announce Type: replace Abstract: Geospatial foundation models (GFMs) for Earth observation often fail to perform reliably in environments underrepresented during pretraining. We int

SIF: Semantically In-Distribution Fingerprints for Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.17041v1 Announce Type: new Abstract: The public accessibility of large vision-language models (LVLMs) raises serious concerns about unauthorized model reuse and intellectual property infrin

Sky2Ground: A Benchmark for Site Modeling under Varying Altitude

Model ReleasesDGX agent

arXiv:2603.13740v3 Announce Type: replace Abstract: We introduce Sky2Ground, a three-view dataset designed for varying altitude camera localization, correspondence learning, and reconstruction. The da

SMILE-UHURA Challenge -- Small Vessel Segmentation at Mesoscopic Scale from Ultra-High Resolution 7T Magnetic Resonance Angiograms

ResearchDGX agent

arXiv:2411.09593v2 Announce Type: replace-cross Abstract: The human brain receives nutrients and oxygen through an intricate network of blood vessels. Pathology affecting small vessels, at the mesosco

Soft Label Pruning and Quantization for Large-Scale Dataset Distillation

SafetyDGX agent

arXiv:2604.18135v1 Announce Type: new Abstract: Large-scale dataset distillation requires storing auxiliary soft labels that can be 30-40x larger on ImageNet-1K and 200x larger on ImageNet-21K than th

Source-Free Domain Adaptation with Vision-Language Prior

SafetyDGX agent

arXiv:2604.17748v1 Announce Type: new Abstract: Source-Free Domain Adaptation (SFDA) seeks to adapt a source model, which is pre-trained on a supervised source domain, for a target domain, with only a

(Sparse) Attention to the Details: Preserving Spectral Fidelity in ML-based Weather Forecasting Models

SafetyDGX agent

arXiv:2604.16429v1 Announce Type: cross Abstract: We introduce Mosaic, a probabilistic weather forecasting model that addresses two principal sources of spectral degradation in ML-based weather predic

Spatial-Regularization-Aware Dual-Branch Collaborative Inference for Training-Free OVSS in Remote Sensing Imagery

Local AiDGX agent

arXiv:2601.21159v2 Announce Type: replace Abstract: High-resolution remote sensing images contain densely distributed objects with pronounced scale variations and complex boundaries, which impose high

SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning

ResearchDGX agent

arXiv:2604.17385v1 Announce Type: new Abstract: Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge fo

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

Local AiDGX agent

arXiv:2603.27437v2 Announce Type: replace Abstract: Large vision-language models (VLMs) still struggle with reliable 3D spatial reasoning, a core capability for embodied and physical AI systems. This

Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models

Model ReleasesDGX agent

arXiv:2604.17873v1 Announce Type: new Abstract: Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational

Spectral Forensics of Diffusion Attention Graphs for Copy-Move Forgery Detection

ResearchDGX agent

arXiv:2604.17287v1 Announce Type: new Abstract: Copy-move forgery, where a region within an image is duplicated to hide or fabricate content, remains a persistent threat to visual media integrity. We

Speculative Decoding for Autoregressive Video Generation

ResearchDGX agent

arXiv:2604.17397v1 Announce Type: new Abstract: Autoregressive video diffusion is emerging as a promising paradigm for streaming video synthesis, with step distillation serving as the primary means of

Spike-NVPT: Learning Robust Visual Prompts via Bio-Inspired Temporal Filtering and Discretization

Model ReleasesDGX agent

arXiv:2604.18284v1 Announce Type: new Abstract: Pre-trained vision models have found widespread application across diverse domains. Prompt tuning-based methods have emerged as a parameter-efficient pa

← Previous
1…184185186187188…209
Next →