AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
12 May 2026

PIDNet: Progressive Implicit Decouple Network for Multimodal Action Quality Assessment

ResearchDGX agent

arXiv:2605.08945v1 Announce Type: new Abstract: Action quality assessment (AQA) aims to automatically quantify the execution quality of human actions in videos and is valuable for applications such as

Pix2Fact: When Vision Is Not Enough -- Benchmarking Fine-Grained VQA with Web Verification on High-Resolution Real-World Scenes

Model ReleasesDGX agent

arXiv:2602.00593v2 Announce Type: replace Abstract: Despite progress on general tasks, vision-language models (VLMs) still struggle with challenges that demand both fine-grained visual grounding and e

Pixal3D: Pixel-Aligned 3D Generation from Images

ResearchDGX agent

arXiv:2605.10922v1 Announce Type: new Abstract: Recent advances in 3D generative models have rapidly improved image-to-3D synthesis quality, enabling higher-resolution geometry and more realistic appe


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

PixelFlowCast: Latent-Free Precipitation Nowcasting via Pixel Mean Flows

ApplicationsDGX agent

arXiv:2605.10046v1 Announce Type: new Abstract: Precipitation nowcasting aims to forecast short-term radar echo sequences for extreme weather warning, where both prediction fidelity and inference effi

PolarVSR: A Unified Framework and Benchmark for Continuous Space-Time Polarization Video Reconstruction

Model ReleasesDGX agent

arXiv:2605.10275v1 Announce Type: new Abstract: Polarimetric imaging captures surface polarization characteristics, such as the Degree of Linear Polarization (DoLP) and the Angle of Polarization (AoP)

Polygon-mamba: Retinal vessel segmentation using polygon scanning mamba and space-frequency collaborative attention

Local AiDGX agent

arXiv:2605.10581v1 Announce Type: new Abstract: Retinal vessel segmentation is crucial for diagnosis and assessment of ocular diseases. Notably, segmentation of small retinal vessels has been consiste

Position: Life-Logging Video Streams Make the Privacy-Utility Trade-off Inevitable

ResearchDGX agent

arXiv:2605.10404v1 Announce Type: new Abstract: With the growing prevalence of always-on hardware such as smart glasses, body cameras, and home security systems, life-logging visual sensing is becomin

Post-hoc Selective Classification for Reliable Synthetic Image Detection

ResearchDGX agent

arXiv:2605.08574v1 Announce Type: new Abstract: As synthetic images become increasingly realistic, reliable synthetic image detection techniques are of pressing need to prevent their misuse. Despite s

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping

Local AiDGX agent

arXiv:2605.10937v1 Announce Type: new Abstract: Recently, post-training methods based on reinforcement learning, with a particular focus on Group Relative Policy Optimization (GRPO), have emerged as t

Predicting 3D structure by latent posterior sampling

TutorialsDGX agent

arXiv:2605.10830v1 Announce Type: new Abstract: The remarkable achievements of both generative models of 2D images and neural field representations for 3D scenes present a compelling opportunity to in

Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional Drift

Model ReleasesDGX agent

arXiv:2505.19519v3 Announce Type: replace Abstract: Personalizing text-to-image diffusion models involves integrating novel visual concepts from a small set of reference images while retaining the mod

Probability-Flow Distillation: Exact Wasserstein Gradient Flow for High-Fidelity 3D Generation

ResearchDGX agent

arXiv:2605.09071v1 Announce Type: new Abstract: Score Distillation Sampling (SDS) and its variants have been widely used for text-to-3D generation by distilling 2D image diffusion priors. However, the

ProDG: Prototypes for Data-Free Generative Post-Hoc Explainability

ResearchDGX agent

arXiv:2605.08858v1 Announce Type: new Abstract: Ante-hoc interpretability methods based on prototypes provide highly accurate explanations by utilizing the intuitive 'this looks like that' reasoning p

Product-of-Gaussian-Mixture Diffusion Models for Joint Nonlinear MRI Reconstruction

Model ReleasesDGX agent

arXiv:2605.10629v1 Announce Type: new Abstract: Recently, diffusion models have attracted considerable attention for magnetic resonance image reconstruction due to their high sample quality. However,

Progressive Photorealistic Simplification

ResearchDGX agent

arXiv:2605.10409v1 Announce Type: new Abstract: Existing image simplification techniques often rely on Non-Photorealistic Rendering (NPR), transforming photographs into stylized sketches, cartoons, or

Prompt Estimation from Prototypes for Federated Prompt Tuning of Vision Transformers

Model ReleasesDGX agent

arXiv:2510.25372v2 Announce Type: replace Abstract: Visual Prompt Tuning (VPT) of pre-trained Vision Transformers (ViTs) has proven highly effective as a parameter-efficient fine-tuning technique for

QueST: Persistent Queries as Semantic Monitors for Drift Suppression in Long-Horizon Tracking

ResearchDGX agent

arXiv:2605.09513v1 Announce Type: new Abstract: Tracking points in videos is typically formulated as frame-to-frame correspondence, where each point is matched locally to the next frame. While this wo

Qwen-Image-2.0 Technical Report

Model ReleasesDGX agent

arXiv:2605.10730v1 Announce Type: new Abstract: We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a si

R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object Detection

Model ReleasesDGX agent

arXiv:2603.11566v2 Announce Type: replace Abstract: 4D radar-camera sensing configuration has gained increasing importance in autonomous driving. However, existing 3D object detection methods that fus

Radar-Guided Polynomial Fitting for Metric Depth Estimation

ResearchDGX agent

arXiv:2503.17182v4 Announce Type: replace Abstract: We propose POLAR, a novel radar-guided depth estimation method that introduces polynomial fitting to efficiently transform scaleless depth predictio

RadThinking: A Dataset for Longitudinal Clinical Reasoning in Radiology

Model ReleasesDGX agent

arXiv:2605.10761v1 Announce Type: new Abstract: Cancer screening is a reasoning task. A radiologist observes findings, compares them to prior scans, integrates clinical context, and reaches a diagnost

Rapid Forest Fuel Load Estimation via Virtual Remote Sensing and Metric-Scale Feed-Forward 3D Reconstruction

ResearchDGX agent

arXiv:2605.10789v1 Announce Type: new Abstract: Accurate quantification of forest coverage and combustible biomass (fuel load) is critical for wildfire risk assessment and ecosystem management. Howeve

Raster2Seq: Polygon Sequence Generation for Floorplan Reconstruction

ResearchDGX agent

arXiv:2602.09016v2 Announce Type: replace Abstract: Reconstructing a structured vector-graphics representation from a rasterized floorplan image is typically an important prerequisite for computationa

ReaMOT: A Benchmark and Framework for Reasoning-based Multi-Object Tracking

Model ReleasesDGX agent

arXiv:2505.20381v4 Announce Type: replace Abstract: Referring Multi-Object Tracking (RMOT) aims to track targets specified by language instructions. However, existing RMOT paradigms heavily rely on ex

Reducing Annotation Burden for Femoral Cartilage Segmentation in Knee MRI via Cross-Sequence Transfer Learning

ResearchDGX agent

arXiv:2605.09067v1 Announce Type: new Abstract: Purpose: To develop and evaluate cross-sequence transfer learning for automatic femoral cartilage segmentation, testing bidirectional transfer between d

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning

SafetyDGX agent

arXiv:2605.09614v1 Announce Type: new Abstract: Long chain-of-thought (CoT) reasoning improves large vision--language models, but visual information often fades during generation, limiting long-horizo

Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow-Matching Models

ResearchDGX agent

arXiv:2605.10759v1 Announce Type: cross Abstract: Diffusion and flow-matching models scale because pretraining is supervised regression: a clean sample is noised analytically, and a model regresses ag

Relightable Gaussian Splatting for Virtual Production Using Image-Based Illumination

Local AiDGX agent

arXiv:2605.09024v1 Announce Type: new Abstract: Virtual production (VP) use LED walls to provide both background imagery and image-based lighting. While this enables on-set compositing, it couples lig

ReorgGS: Equivalent Distribution Reorganization for 3D Gaussian Splatting

Model ReleasesDGX agent

arXiv:2605.08739v1 Announce Type: new Abstract: A converged 3D Gaussian Splatting (3DGS) model may approximate the target scene while remaining poorly parameterized for further optimization. We identi

Restoration-Aligned Generative Flow Models for Blind Motion Deblurring

ResearchDGX agent

arXiv:2605.08854v1 Announce Type: new Abstract: Generative flow models offer powerful priors learned from large-scale natural images, but directly adapting them to restoration tasks such as motion deb

Rethinking Event-Based Object Dtection through Representation-Level Temporal Aggregation and Model-Level Hypergraph Reasoning

ResearchDGX agent

arXiv:2605.08825v1 Announce Type: new Abstract: Event cameras provide microsecond-level temporal resolution, low latency, and high dynamic range, offering potential for perception under fast motion an

Revitalizing the Beginning: Avoiding Storage Dependency for Model Merging in Continual Learning

SafetyDGX agent

arXiv:2605.08311v1 Announce Type: cross Abstract: Model merging provides a compelling paradigm for integrating specialized expertise into a unified multi-task model, a goal that aligns naturally with

S2FT: Parameter-Efficient Fine-Tuning in Sparse Spectrum Domain

Model ReleasesDGX agent

arXiv:2605.08589v1 Announce Type: new Abstract: Parameter Efficient Fine-Tuning (PEFT) is a key technique for adapting a large pretrained model to downstream tasks by fine-tuning only a small number o

SABER: A Scalable Action-Based Embodied Dataset for Real-World VLA Adaptation

ApplicationsDGX agent

arXiv:2605.09613v1 Announce Type: cross Abstract: Robotic deployment in real-world environments depends on rich, domain-specific action data as much as on strong model architecture. General-purpose ro

SAMOFT: Robust Multi-Object Tracking via Region and Flow

ResearchDGX agent

arXiv:2605.09417v1 Announce Type: new Abstract: Multi-object tracking (MOT) is a fundamental task in computer vision that requires continuously tracking multiple targets while maintaining consistent i

SciLT: Long-tailed Image Classification under Scientific Image Domains

ResearchDGX agent

arXiv:2604.03687v2 Announce Type: replace Abstract: Long-tailed recognition has benefited from foundation models and fine-tuning paradigms, yet existing studies and benchmarks are mainly confined to n

SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation

Model ReleasesDGX agent

arXiv:2605.10187v1 Announce Type: new Abstract: Scientific reasoning is a key aspect of human intelligence, requiring the integration of multimodal inputs, domain expertise, and multi-step inference a

SeasonScapes: Learning Large-scale Re-lightable 3D Landscapes with Seasonal Variation from Sparse Webcams

ResearchDGX agent

arXiv:2605.09039v1 Announce Type: new Abstract: We introduce SeasonScapes framework and a the SeasonScapes dataset: Swiss Sparse-view Mountain Scenes with Seasonal Changes that covers over 50 km x 60

Segment Anything with Robust Uncertainty-Accuracy Correlation

SafetyDGX agent

arXiv:2605.10603v1 Announce Type: new Abstract: Despite strong zero-shot performance, SAM is unreliable under domain shift due to Mask-level Confidence Confusion (MCC), where a single IoU-based mask s

SegSTRONG-C: Segmenting Surgical Tools Robustly On Non-adversarial Generated Corruptions -- An EndoVis'24 Challenge

ResearchDGX agent

arXiv:2407.11906v3 Announce Type: replace Abstract: Surgical data science has seen rapid advancement with the excellent performance of end-to-end deep neural networks (DNNs). Despite their successes,

SEIS: Subspace-based Equivariance and Invariance Scores for Neural Representations

ResearchDGX agent

arXiv:2602.04054v2 Announce Type: replace-cross Abstract: Understanding how neural representations respond to geometric transformations is essential for evaluating whether learned features preserve me

Semantic Alignment in Hyperbolic Space for Open-Vocabulary Semantic Segmentation

SafetyDGX agent

arXiv:2605.08874v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation requires adapting image-level vision-language models such as CLIP to dense pixel-level prediction, which is challe

Sens-VisualNews: A Benchmark Dataset for Sensational Image Detection

Model ReleasesDGX agent

arXiv:2605.10394v1 Announce Type: new Abstract: The detection of sensational content in media items can be a critical filtering mechanism for identifying check-worthy content and flagging potential di

Set-Based Groupwise Registration for Variable-Length, Variable-Contrast Cardiac MRI

ResearchDGX agent

arXiv:2605.10571v1 Announce Type: cross Abstract: Quantitative cardiac magnetic resonance imaging (MRI) enables non-invasive myocardial tissue characterization but relies on robust motion correction w

simpleposter: a simple baseline for product poster generation

Model ReleasesDGX agent

arXiv:2605.08784v1 Announce Type: new Abstract: Product poster generation poses distinct challenges beyond general poster design, requiring both faithful preservation of product appearance and precise

Simultaneous Monitoring of Shape and Surface Color via 4D Point Clouds: A Registration-free Approach

ApplicationsDGX agent

arXiv:2605.08753v1 Announce Type: new Abstract: Advanced manufacturing technologies allow for the production of intricate parts featuring high shape complexity and spatially-varying material compositi

SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation

Model ReleasesDGX agent

arXiv:2605.10376v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have advanced rapidly in multimodal perception and language understanding, yet it remains unclear whether they can reliabl

Slum Detection and Density Mapping with AlphaEarth Foundations: A Representation Learning Evaluation Across 12 Global Cities

ResearchDGX agent

arXiv:2605.10029v1 Announce Type: new Abstract: Pixel-level slum mapping has long been constrained by limited cross-city generalisation, the absence of continuous density estimation, and weak global c

Smart Railway Obstruction Detection System using IoT and Computer Vision

Local AiDGX agent

arXiv:2605.08246v1 Announce Type: new Abstract: Railway track intrusions pose a critical safety challenge for Indian Railways, encompassing wildlife incursions and deliberate malicious obstructions. T

SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy

Model ReleasesDGX agent

arXiv:2605.09598v1 Announce Type: new Abstract: Vision-language models (VLMs) have recently shown strong potential in soccer video understanding. However, given the high complexity of soccer videos du

SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation

ApplicationsDGX agent

arXiv:2605.10079v1 Announce Type: new Abstract: Video generation has advanced rapidly, producing photorealistic videos from text or image prompts. Meanwhile, film production and social robotics increa

SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs

ResearchDGX agent

arXiv:2605.09449v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) have made remarkable progress in visual understanding and language-based reasoning, yet they lack a pers

Sparsity Hurts: Simple Linear Adapter Can Boost Generalized Category Discovery

SafetyDGX agent

arXiv:2605.08183v1 Announce Type: new Abstract: Generalized Category Discovery (GCD) seeks to identify novel categories from unlabeled data while retaining the classification ability of seen categorie

Spatial-Frequency Gated Swin Transformer for Remote Sensing Single-Image Super-Resolution

ResearchDGX agent

arXiv:2605.09687v1 Announce Type: new Abstract: Remote Sensing (RS) single-image super-resolution aims to reconstruct high-resolution imagery from low-resolution observations while preserving fine spa

SPECTRA-Net: Scalable Pipeline for Explainable Cross-domain Tensor Representations for AI-generated Images Detection

ApplicationsDGX agent

arXiv:2605.08226v1 Announce Type: new Abstract: The rapid proliferation of AI-generated images (AIGI) presents a significant challenge to digital information integrity. While human observers and exist

Spectrally-Guided Diffusion Noise Schedules

ResearchDGX agent

arXiv:2603.19222v2 Announce Type: replace Abstract: Denoising diffusion models are widely used for high-quality image and video generation. Their performance depends on noise schedules, which define t

Stealthy Patch-Wise Backdoor Attack in 3D Point Cloud via Curvature Awareness

ResearchDGX agent

arXiv:2503.09336v4 Announce Type: replace Abstract: Backdoor attacks pose a severe threat to deep neural networks (DNNs) by implanting hidden backdoors that can be activated with predefined triggers t

StereoPolicy: Improving Robotic Manipulation Policies via Stereo Perception

SafetyDGX agent

arXiv:2605.09989v1 Announce Type: cross Abstract: Recent advances in robot imitation learning have yielded powerful visuomotor policies capable of manipulating a wide variety of objects directly from

STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering

SafetyDGX agent

arXiv:2604.01824v2 Announce Type: replace Abstract: We introduce STRIVE (SpatioTemporal Reinforcement with Importance-aware Variant Exploration), a structured reinforcement learning framework for vide

Supersampling Stable Diffusion and More: An Approach for Interpolating Neural Networks Using Common Interpolation Methods

ResearchDGX agent

arXiv:2605.08698v1 Announce Type: new Abstract: Stable Diffusion (SD) has evolved DDPM (Denoising Diffusion Probabilistic Model) based image generation significantly by denoising in latent space inste

← Previous
1…145146147148149…209
Next →