AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlog
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

Qwen-Image-2.0 Technical Report

DGX agent

arXiv:2605.10730v1 Announce Type: new Abstract: We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a si

model-releasesarxiv-cs-cv
12 May 2026
Model Releases

R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object Detection

DGX agent
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

arXiv:2603.11566v2 Announce Type: replace Abstract: 4D radar-camera sensing configuration has gained increasing importance in autonomous driving. However, existing 3D object detection methods that fus

model-releasesarxiv-cs-cv
12 May 2026
Research

Radar-Guided Polynomial Fitting for Metric Depth Estimation

DGX agent

arXiv:2503.17182v4 Announce Type: replace Abstract: We propose POLAR, a novel radar-guided depth estimation method that introduces polynomial fitting to efficiently transform scaleless depth predictio

researcharxiv-cs-cv
12 May 2026
Model Releases

RadThinking: A Dataset for Longitudinal Clinical Reasoning in Radiology

DGX agent

arXiv:2605.10761v1 Announce Type: new Abstract: Cancer screening is a reasoning task. A radiologist observes findings, compares them to prior scans, integrates clinical context, and reaches a diagnost

model-releasesarxiv-cs-cv
12 May 2026
Research

Rapid Forest Fuel Load Estimation via Virtual Remote Sensing and Metric-Scale Feed-Forward 3D Reconstruction

DGX agent

arXiv:2605.10789v1 Announce Type: new Abstract: Accurate quantification of forest coverage and combustible biomass (fuel load) is critical for wildfire risk assessment and ecosystem management. Howeve

researcharxiv-cs-cv
12 May 2026
Research

Raster2Seq: Polygon Sequence Generation for Floorplan Reconstruction

DGX agent

arXiv:2602.09016v2 Announce Type: replace Abstract: Reconstructing a structured vector-graphics representation from a rasterized floorplan image is typically an important prerequisite for computationa

researcharxiv-cs-cv
12 May 2026
Model Releases

ReaMOT: A Benchmark and Framework for Reasoning-based Multi-Object Tracking

DGX agent

arXiv:2505.20381v4 Announce Type: replace Abstract: Referring Multi-Object Tracking (RMOT) aims to track targets specified by language instructions. However, existing RMOT paradigms heavily rely on ex

model-releasesarxiv-cs-cv
12 May 2026
Research

Reducing Annotation Burden for Femoral Cartilage Segmentation in Knee MRI via Cross-Sequence Transfer Learning

DGX agent

arXiv:2605.09067v1 Announce Type: new Abstract: Purpose: To develop and evaluate cross-sequence transfer learning for automatic femoral cartilage segmentation, testing bidirectional transfer between d

researcharxiv-cs-cv
12 May 2026
Safety

Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning

DGX agent

arXiv:2605.09614v1 Announce Type: new Abstract: Long chain-of-thought (CoT) reasoning improves large vision--language models, but visual information often fades during generation, limiting long-horizo

safetyarxiv-cs-cv
12 May 2026
Research

Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow-Matching Models

DGX agent

arXiv:2605.10759v1 Announce Type: cross Abstract: Diffusion and flow-matching models scale because pretraining is supervised regression: a clean sample is noised analytically, and a model regresses ag

researcharxiv-cs-cv
12 May 2026
Local Ai

Relightable Gaussian Splatting for Virtual Production Using Image-Based Illumination

DGX agent

arXiv:2605.09024v1 Announce Type: new Abstract: Virtual production (VP) use LED walls to provide both background imagery and image-based lighting. While this enables on-set compositing, it couples lig

local-aiarxiv-cs-cv
12 May 2026
Model Releases

ReorgGS: Equivalent Distribution Reorganization for 3D Gaussian Splatting

DGX agent

arXiv:2605.08739v1 Announce Type: new Abstract: A converged 3D Gaussian Splatting (3DGS) model may approximate the target scene while remaining poorly parameterized for further optimization. We identi

model-releasesarxiv-cs-cv
12 May 2026
Research

Restoration-Aligned Generative Flow Models for Blind Motion Deblurring

DGX agent

arXiv:2605.08854v1 Announce Type: new Abstract: Generative flow models offer powerful priors learned from large-scale natural images, but directly adapting them to restoration tasks such as motion deb

researcharxiv-cs-cv
12 May 2026
Research

Rethinking Event-Based Object Dtection through Representation-Level Temporal Aggregation and Model-Level Hypergraph Reasoning

DGX agent

arXiv:2605.08825v1 Announce Type: new Abstract: Event cameras provide microsecond-level temporal resolution, low latency, and high dynamic range, offering potential for perception under fast motion an

researcharxiv-cs-cv
12 May 2026
Safety

Revitalizing the Beginning: Avoiding Storage Dependency for Model Merging in Continual Learning

DGX agent

arXiv:2605.08311v1 Announce Type: cross Abstract: Model merging provides a compelling paradigm for integrating specialized expertise into a unified multi-task model, a goal that aligns naturally with

safetyarxiv-cs-cv
12 May 2026
Model Releases

S2FT: Parameter-Efficient Fine-Tuning in Sparse Spectrum Domain

DGX agent

arXiv:2605.08589v1 Announce Type: new Abstract: Parameter Efficient Fine-Tuning (PEFT) is a key technique for adapting a large pretrained model to downstream tasks by fine-tuning only a small number o

model-releasesarxiv-cs-cv
12 May 2026
Applications

SABER: A Scalable Action-Based Embodied Dataset for Real-World VLA Adaptation

DGX agent

arXiv:2605.09613v1 Announce Type: cross Abstract: Robotic deployment in real-world environments depends on rich, domain-specific action data as much as on strong model architecture. General-purpose ro

applicationsarxiv-cs-cv
12 May 2026
Research

SAMOFT: Robust Multi-Object Tracking via Region and Flow

DGX agent

arXiv:2605.09417v1 Announce Type: new Abstract: Multi-object tracking (MOT) is a fundamental task in computer vision that requires continuously tracking multiple targets while maintaining consistent i

researcharxiv-cs-cv
12 May 2026
Research

SciLT: Long-tailed Image Classification under Scientific Image Domains

DGX agent

arXiv:2604.03687v2 Announce Type: replace Abstract: Long-tailed recognition has benefited from foundation models and fine-tuning paradigms, yet existing studies and benchmarks are mainly confined to n

researcharxiv-cs-cv
12 May 2026
Model Releases

SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation

DGX agent

arXiv:2605.10187v1 Announce Type: new Abstract: Scientific reasoning is a key aspect of human intelligence, requiring the integration of multimodal inputs, domain expertise, and multi-step inference a

model-releasesarxiv-cs-cv
12 May 2026
Research

SeasonScapes: Learning Large-scale Re-lightable 3D Landscapes with Seasonal Variation from Sparse Webcams

DGX agent

arXiv:2605.09039v1 Announce Type: new Abstract: We introduce SeasonScapes framework and a the SeasonScapes dataset: Swiss Sparse-view Mountain Scenes with Seasonal Changes that covers over 50 km x 60

researcharxiv-cs-cv
12 May 2026
Safety

Segment Anything with Robust Uncertainty-Accuracy Correlation

DGX agent

arXiv:2605.10603v1 Announce Type: new Abstract: Despite strong zero-shot performance, SAM is unreliable under domain shift due to Mask-level Confidence Confusion (MCC), where a single IoU-based mask s

safetyarxiv-cs-cv
12 May 2026
Research

SegSTRONG-C: Segmenting Surgical Tools Robustly On Non-adversarial Generated Corruptions -- An EndoVis'24 Challenge

DGX agent

arXiv:2407.11906v3 Announce Type: replace Abstract: Surgical data science has seen rapid advancement with the excellent performance of end-to-end deep neural networks (DNNs). Despite their successes,

researcharxiv-cs-cv
12 May 2026
Research

SEIS: Subspace-based Equivariance and Invariance Scores for Neural Representations

DGX agent

arXiv:2602.04054v2 Announce Type: replace-cross Abstract: Understanding how neural representations respond to geometric transformations is essential for evaluating whether learned features preserve me

researcharxiv-cs-cv
12 May 2026
Safety

Semantic Alignment in Hyperbolic Space for Open-Vocabulary Semantic Segmentation

DGX agent

arXiv:2605.08874v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation requires adapting image-level vision-language models such as CLIP to dense pixel-level prediction, which is challe

safetyarxiv-cs-cv
12 May 2026
Model Releases

Sens-VisualNews: A Benchmark Dataset for Sensational Image Detection

DGX agent

arXiv:2605.10394v1 Announce Type: new Abstract: The detection of sensational content in media items can be a critical filtering mechanism for identifying check-worthy content and flagging potential di

model-releasesarxiv-cs-cv
12 May 2026
Research

Set-Based Groupwise Registration for Variable-Length, Variable-Contrast Cardiac MRI

DGX agent

arXiv:2605.10571v1 Announce Type: cross Abstract: Quantitative cardiac magnetic resonance imaging (MRI) enables non-invasive myocardial tissue characterization but relies on robust motion correction w

researcharxiv-cs-cv
12 May 2026
Model Releases

simpleposter: a simple baseline for product poster generation

DGX agent

arXiv:2605.08784v1 Announce Type: new Abstract: Product poster generation poses distinct challenges beyond general poster design, requiring both faithful preservation of product appearance and precise

model-releasesarxiv-cs-cv
12 May 2026
Applications

Simultaneous Monitoring of Shape and Surface Color via 4D Point Clouds: A Registration-free Approach

DGX agent

arXiv:2605.08753v1 Announce Type: new Abstract: Advanced manufacturing technologies allow for the production of intricate parts featuring high shape complexity and spatially-varying material compositi

applicationsarxiv-cs-cv
12 May 2026
Model Releases

SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigation

DGX agent

arXiv:2605.10376v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have advanced rapidly in multimodal perception and language understanding, yet it remains unclear whether they can reliabl

model-releasesarxiv-cs-cv
12 May 2026
Research

Slum Detection and Density Mapping with AlphaEarth Foundations: A Representation Learning Evaluation Across 12 Global Cities

DGX agent

arXiv:2605.10029v1 Announce Type: new Abstract: Pixel-level slum mapping has long been constrained by limited cross-city generalisation, the absence of continuous density estimation, and weak global c

researcharxiv-cs-cv
12 May 2026
Local Ai

Smart Railway Obstruction Detection System using IoT and Computer Vision

DGX agent

arXiv:2605.08246v1 Announce Type: new Abstract: Railway track intrusions pose a critical safety challenge for Indian Railways, encompassing wildlife incursions and deliberate malicious obstructions. T

local-aiarxiv-cs-cv
12 May 2026
Model Releases

SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy

DGX agent

arXiv:2605.09598v1 Announce Type: new Abstract: Vision-language models (VLMs) have recently shown strong potential in soccer video understanding. However, given the high complexity of soccer videos du

model-releasesarxiv-cs-cv
12 May 2026
Applications

SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation

DGX agent

arXiv:2605.10079v1 Announce Type: new Abstract: Video generation has advanced rapidly, producing photorealistic videos from text or image prompts. Meanwhile, film production and social robotics increa

applicationsarxiv-cs-cv
12 May 2026
Research

SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs

DGX agent

arXiv:2605.09449v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) have made remarkable progress in visual understanding and language-based reasoning, yet they lack a pers

researcharxiv-cs-cv
12 May 2026
Safety

Sparsity Hurts: Simple Linear Adapter Can Boost Generalized Category Discovery

DGX agent

arXiv:2605.08183v1 Announce Type: new Abstract: Generalized Category Discovery (GCD) seeks to identify novel categories from unlabeled data while retaining the classification ability of seen categorie

safetyarxiv-cs-cv
12 May 2026
Research

Spatial-Frequency Gated Swin Transformer for Remote Sensing Single-Image Super-Resolution

DGX agent

arXiv:2605.09687v1 Announce Type: new Abstract: Remote Sensing (RS) single-image super-resolution aims to reconstruct high-resolution imagery from low-resolution observations while preserving fine spa

researcharxiv-cs-cv
12 May 2026
Applications

SPECTRA-Net: Scalable Pipeline for Explainable Cross-domain Tensor Representations for AI-generated Images Detection

DGX agent

arXiv:2605.08226v1 Announce Type: new Abstract: The rapid proliferation of AI-generated images (AIGI) presents a significant challenge to digital information integrity. While human observers and exist

applicationsarxiv-cs-cv
12 May 2026
Research

Spectrally-Guided Diffusion Noise Schedules

DGX agent

arXiv:2603.19222v2 Announce Type: replace Abstract: Denoising diffusion models are widely used for high-quality image and video generation. Their performance depends on noise schedules, which define t

researcharxiv-cs-cv
12 May 2026
Research

Stealthy Patch-Wise Backdoor Attack in 3D Point Cloud via Curvature Awareness

DGX agent

arXiv:2503.09336v4 Announce Type: replace Abstract: Backdoor attacks pose a severe threat to deep neural networks (DNNs) by implanting hidden backdoors that can be activated with predefined triggers t

researcharxiv-cs-cv
12 May 2026
Safety

StereoPolicy: Improving Robotic Manipulation Policies via Stereo Perception

DGX agent

arXiv:2605.09989v1 Announce Type: cross Abstract: Recent advances in robot imitation learning have yielded powerful visuomotor policies capable of manipulating a wide variety of objects directly from

safetyarxiv-cs-cv
12 May 2026
Safety

STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering

DGX agent

arXiv:2604.01824v2 Announce Type: replace Abstract: We introduce STRIVE (SpatioTemporal Reinforcement with Importance-aware Variant Exploration), a structured reinforcement learning framework for vide

safetyarxiv-cs-cv
12 May 2026
Research

Supersampling Stable Diffusion and More: An Approach for Interpolating Neural Networks Using Common Interpolation Methods

DGX agent

arXiv:2605.08698v1 Announce Type: new Abstract: Stable Diffusion (SD) has evolved DDPM (Denoising Diffusion Probabilistic Model) based image generation significantly by denoising in latent space inste

researcharxiv-cs-cv
12 May 2026
Research

Survey on Disaster Management Datasets for Remote Sensing Based Emergency Applications

DGX agent

arXiv:2605.08196v1 Announce Type: new Abstract: Recent natural disasters have highlighted the urgent need for efficient data-driven approaches to disaster management. Machine learning (ML) and deep le

researcharxiv-cs-cv
12 May 2026
Hardware

SwiftI2V: Efficient High-Resolution Image-to-Video Generation via Conditional Segment-wise Generation

DGX agent

arXiv:2605.06356v2 Announce Type: replace Abstract: High-resolution image-to-video (I2V) generation aims to synthesize realistic temporal dynamics while preserving fine-grained appearance details of t

hardwarearxiv-cs-cv
12 May 2026
Model Releases

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

DGX agent

arXiv:2605.08412v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have made rapid progress in single-video understanding, yet their ability to reason across multiple independent

model-releasesarxiv-cs-cv
12 May 2026
Safety

SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment

DGX agent

arXiv:2605.08724v1 Announce Type: new Abstract: Unifying multimodal understanding and generation is a compelling frontier that is beginning to emerge in the medical field. However, the limited existin

safetyarxiv-cs-cv
12 May 2026
Research

TARO: Temporal Adversarial Rectification Optimization Using Diffusion Models as Purifiers

DGX agent

arXiv:2605.08440v1 Announce Type: cross Abstract: Adversarial purification with diffusion models seeks to project adversarial examples back toward the data manifold, but balancing semantic preservatio

researcharxiv-cs-cv
12 May 2026
← Previous
1…184185186187188…263
Next →