AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

Enabling Progressive Whole-slide Image Analysis with Multi-scale Pyramidal Network

DGX agent

arXiv:2602.01951v2 Announce Type: replace Abstract: Multiple-instance Learning (MIL) is commonly used for computational pathology (CPath), where multi-scale features are essential for capturing both f

model-releasesarxiv-cs-cv
10 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Agents

Envision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous Driving

DGX agent

arXiv:2606.10656v1 Announce Type: new Abstract: Forecasting the future evolution of dynamic scenes is crucial in autonomous driving. However, existing feed-forward paradigms are primarily designed for

agentsarxiv-cs-cv
10 Jun 2026
Research

FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion

DGX agent

arXiv:2606.10671v1 Announce Type: new Abstract: Autoregressive video generators synthesize long videos by generating successive temporal segments, but their historical KV cache grows with video length

researcharxiv-cs-cv
10 Jun 2026
Research

Few-step Generative Models as Lossy Compression

DGX agent

arXiv:2606.10450v1 Announce Type: new Abstract: DiffC provides a principled way to reuse pre-trained diffusion models for lossy compression, but its encoding and decoding procedures remain slow becaus

researcharxiv-cs-cv
10 Jun 2026
Research

FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models

DGX agent

arXiv:2509.16518v2 Announce Type: replace Abstract: Using diffusion transformers for media generation may require evaluating attention over extremely long sequences, with attention layers accounting f

researcharxiv-cs-cv
10 Jun 2026
Research

FlexPath: Learned Semantic Path Priors for Image-Based Planning

DGX agent

arXiv:2606.10167v1 Announce Type: new Abstract: Recent learning-based path planners use neural networks to process visual map representations and approximate heuristics for classical search algorithms

researcharxiv-cs-cv
10 Jun 2026
Applications

FoA-SR: Faithful or Aesthetic? Profile-Aware Preference Optimization for Real-World Image Super-Resolution

DGX agent

arXiv:2606.10275v1 Announce Type: new Abstract: Real-world image super-resolution (SR) is often designed with a single restoration objective, despite the current capacity of generative models to produ

applicationsarxiv-cs-cv
10 Jun 2026
Model Releases

From Patches to Patients: A study of the tile-to-slide performance transferability in Digital Pathology

DGX agent

arXiv:2606.10778v1 Announce Type: new Abstract: Foundation Models (FMs) have recently redefined the state-of-the-art in histopathology by providing robust representations for whole-slide image (WSI) a

model-releasesarxiv-cs-cv
10 Jun 2026
Research

FSS-Net: Frequency-Spatial Synergy Network with Wavelet Attention for Carotid Artery Ultrasound Segmentation

DGX agent

arXiv:2606.10378v1 Announce Type: new Abstract: Accurate segmentation of carotid arteries in ultrasound imaging is critical for stroke risk assessment. However, speckle noise, low contrast, and blurre

researcharxiv-cs-cv
10 Jun 2026
Research

Fusing Satellite Imagery and Planimetric Maps for Cross-View Localization

DGX agent

arXiv:2606.10166v1 Announce Type: new Abstract: Current cross-view localization methods predominantly rely on satellite imagery as the aerial modality. Although recent work explores planimetric maps (

researcharxiv-cs-cv
10 Jun 2026
Research

GaussTrace: Provenance Analysis of 3D Gaussian Splatting Models with Evidence-based LLM Reasoning

DGX agent

arXiv:2606.10612v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) is a powerful technique for creating high-fidelity 3D assets. However, the widespread sharing and iterative modification of

researcharxiv-cs-cv
10 Jun 2026
Tutorials

GeoLoom: High-quality Geometric Diagram Generation from Textual Input

DGX agent

arXiv:2512.08180v2 Announce Type: replace Abstract: High-quality geometric diagram generation presents both a challenge and an opportunity: it demands strict spatial accuracy while offering well-defin

tutorialsarxiv-cs-cv
10 Jun 2026
Local Ai

Geometric Coastline Localization using Vision-Language Models

DGX agent

arXiv:2606.10468v1 Announce Type: new Abstract: Coastline detection in remote sensing imagery is commonly formulated as a pixel-wise segmentation problem, where the final coastline is extracted from a

local-aiarxiv-cs-cv
10 Jun 2026
Model Releases

Geometry-Aware Reinforcement Learning for 2D Irregular Nesting

DGX agent

arXiv:2606.10611v1 Announce Type: cross Abstract: Traditional heuristic solvers for the 2D irregular nesting problem share a fundamental limitation: they are blind to polygon geometry, relying on guid

model-releasesarxiv-cs-cv
10 Jun 2026
Safety

GHOST: Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation

DGX agent

arXiv:2606.10025v1 Announce Type: cross Abstract: We present GHOST, a framework for learning visuomotor manipulation policies that generalize beyond the training distribution. GHOST factorizes control

safetyarxiv-cs-cv
10 Jun 2026
Safety

Globally Localizing Lunar Rover in Pixels via Graph Alignment

DGX agent

arXiv:2606.10602v1 Announce Type: new Abstract: Precise rover localization is a prerequisite for autonomous lunar exploration, yet the absence of Global Navigation Satellite System (GNSS) signals and

safetyarxiv-cs-cv
10 Jun 2026
Research

Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves

DGX agent

arXiv:2603.20850v2 Announce Type: replace Abstract: Understanding hand-object interaction (HOI) is fundamental to computer vision, robotics, and AR/VR. However, conventional hand videos often lack ess

researcharxiv-cs-cv
10 Jun 2026
Research

GRAR: Glass-induced Reflection Artifact Removal in LiDAR Point Clouds

DGX agent

arXiv:2606.10541v1 Announce Type: new Abstract: Terrestrial Laser Scanning (TLS) point clouds captured in urban environments frequently suffer from glass-induced reflection artifacts, severely degradi

researcharxiv-cs-cv
10 Jun 2026
Safety

GUI-AC: Enhancing Continual Learning in GUI Agents

DGX agent

arXiv:2606.10522v1 Announce Type: new Abstract: Graphical User Interfaces (GUIs) serve as the dominant medium for human-computer interaction, yet building GUI agents that generalize across the vast di

safetyarxiv-cs-cv
10 Jun 2026
Model Releases

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation

DGX agent

arXiv:2606.10839v1 Announce Type: new Abstract: Current identity-consistent video generation methods struggle to preserve appearance fidelity under large viewpoint changes. While introducing multi-vie

model-releasesarxiv-cs-cv
10 Jun 2026
Safety

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder

DGX agent

arXiv:2606.11096v1 Announce Type: new Abstract: Built on pretrained vision foundation models (VFMs), representation autoencoders (RAEs) have recently emerged as a promising approach for constructing s

safetyarxiv-cs-cv
10 Jun 2026
Safety

IMPACT: Learning Internal-Model Predictive Control for Forceful Robotic Manipulation

DGX agent

arXiv:2606.10818v1 Announce Type: cross Abstract: Real-world robotic manipulation tasks often involve forceful interactions with the environment, such as using tools of varying weights, transporting o

safetyarxiv-cs-cv
10 Jun 2026
Local Ai

Improving PET/CT-Based Whole-Body Lesion Segmentation Using Prediction Uncertainty-Augmented Models

DGX agent

arXiv:2606.10115v1 Announce Type: new Abstract: Accurate lesion segmentation from whole-body Positron Emission Tomography (PET)/Computed Tomography (CT) scans is essential for cancer staging and treat

local-aiarxiv-cs-cv
10 Jun 2026
Model Releases

Interpretable Temporal Facial-Region Motion Analysis for In-the-Wild Parkinson's Disease Video Classification

DGX agent

arXiv:2606.10088v1 Announce Type: new Abstract: Reduced facial expressivity is a common motor manifestation of Parkinson's disease (PD), often described as hypomimia or facial bradykinesia. This paper

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

IPSM-Bench: A New Intermediate Phase Segmentation Benchmark in Microstructure Images of Zinc-Based Absorbable Biomaterials

DGX agent

arXiv:2606.11001v1 Announce Type: new Abstract: Zinc-based alloys are indispensable emerging absorbable metallic biomaterials, and their macroscopic performance is governed by microstructural characte

model-releasesarxiv-cs-cv
10 Jun 2026
Research

Is Task-Specific Training Necessary for Anomaly Detection?

DGX agent

arXiv:2601.22763v3 Announce Type: replace Abstract: Current state-of-the-art multi-class unsupervised anomaly detection (MUAD) methods rely on training encoder--decoder models to reconstruct anomaly-f

researcharxiv-cs-cv
10 Jun 2026
Model Releases

iSAGE: A Human-in-the-Loop Framework for Remote Sensing Semantic Segmentation via Sparse Point Supervision

DGX agent

arXiv:2606.10136v1 Announce Type: new Abstract: Semantic segmentation in remote sensing requires costly pixel-level annotations, and nearly every problem demands a new dataset since models rarely tran

model-releasesarxiv-cs-cv
10 Jun 2026
Model Releases

Kwai Keye-VL-2.0 Technical Report

DGX agent

arXiv:2606.10651v1 Announce Type: new Abstract: We introduce Kwai Keye-VL-2.0-30B-A3B, an open-source Mixture-of-Experts (MoE) multimodal foundation model designed to advance long-video understanding

model-releasesarxiv-cs-cv
10 Jun 2026
Safety

LAFP: Preserving Latent Action Structure in Latent Policy Learning via Flow Matching

DGX agent

arXiv:2606.10517v1 Announce Type: new Abstract: Learning high-quality latent actions from large-scale unlabeled videos, coupled with limited real-world interaction data for training an action decoder,

safetyarxiv-cs-cv
10 Jun 2026
Applications

LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning

DGX agent

arXiv:2504.18424v2 Announce Type: replace Abstract: We present Layered Ray Intersections (LaRI), a fully supervised method for occluded geometry reasoning from a single image. Unlike conventional dept

applicationsarxiv-cs-cv
10 Jun 2026
Tutorials

Leveraging Metric Depth for Relative Depth Prediction

DGX agent

arXiv:2606.10628v1 Announce Type: new Abstract: We present our solution to the 2025 SoccerNet Monocular Depth Estimation Competition Challenge. Predicting the relative depth in football scenarios is c

tutorialsarxiv-cs-cv
10 Jun 2026
Safety

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization

DGX agent

arXiv:2606.11180v1 Announce Type: new Abstract: Diffusion-based lip synchronization models achieve strong visual quality and audio-visual alignment, but full-sequence bidirectional attention and many

safetyarxiv-cs-cv
10 Jun 2026
Tutorials

Listen, Look, and Learn: Learning Without Forgetting through SAM-Audio

DGX agent

arXiv:2606.10887v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) aims to continuously learn new classes without forgetting previously acquired knowledge. While recent CIL advances have

tutorialsarxiv-cs-cv
10 Jun 2026
Safety

ManiSplat: Manipulation Trajectory Synthesis from Monocular Video via Decoupled 3D Gaussian Splatting

DGX agent

arXiv:2606.10645v1 Announce Type: new Abstract: Reconstructing dynamic and interactive 3D scenes from real-world observations remains a fundamental challenge in computer vision and robotics. While rec

safetyarxiv-cs-cv
10 Jun 2026
Research

Maximum Matching Accuracy: An Instance Segmentation Evaluation Metric Utilizing Globally Optimal Matching

DGX agent

arXiv:2606.10107v1 Announce Type: new Abstract: Reliable evaluation of instance segmentation models requires metrics that accurately and consistently reflect segmentation quality. However, the metrics

researcharxiv-cs-cv
10 Jun 2026
Safety

Mean Flow Distillation: Robust and Stable Distillation for Flow Matching Models

DGX agent

arXiv:2606.11155v1 Announce Type: new Abstract: Flow Matching models have demonstrated strong performance across a wide range of generative tasks. However, their reliance on ODE-based iterative sampli

safetyarxiv-cs-cv
10 Jun 2026
Safety

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment

DGX agent

arXiv:2512.06628v3 Announce Type: replace-cross Abstract: Scalable embodied intelligence is constrained by the scarcity of diverse, long-horizon robotic manipulation data. Existing video world models

safetyarxiv-cs-cv
10 Jun 2026
Local Ai

MinhwaNet: Faithful but Insufficient Object Grounding in Korean Folk Painting

DGX agent

arXiv:2606.09855v1 Announce Type: cross Abstract: Korean folk painting (minhwa) is built from a small vocabulary of auspicious symbols, a tiger for protection, a pair of birds for marital harmony, a p

local-aiarxiv-cs-cv
10 Jun 2026
Tutorials

MOFA-VTON: More Fashion Possibilities with Fine-Grained Adaptations in Virtual Try-On

DGX agent

arXiv:2606.11148v1 Announce Type: new Abstract: Virtual try-on aims to fit an in-shop clothing image onto a specific human body. An optimal virtual try-on method should provide diverse and flexible dr

tutorialsarxiv-cs-cv
10 Jun 2026
Research

Multi-Angular Reflectance Anisotropy Observed from UAV Multispectral Imagery

DGX agent

arXiv:2606.10350v1 Announce Type: new Abstract: UAV multispectral imagery naturally contains multi-angular observations due to low flight altitude and wide field-of-view imaging, which may introduce g

researcharxiv-cs-cv
10 Jun 2026
Research

Multimodal Brain Tumour Classification Using Feature Fusion

DGX agent

arXiv:2606.11107v1 Announce Type: cross Abstract: Clinicians diagnose brain tumors by synthesizing patient symptoms, medical history, and quantitative imaging data from modalities such as MRI and CT s

researcharxiv-cs-cv
10 Jun 2026
Model Releases

Next Forcing: Causal World Modeling with Multi-Chunk Prediction

DGX agent

arXiv:2606.11187v1 Announce Type: new Abstract: Autoregressive video generation has emerged as a powerful paradigm for World Action Models (WAMs). However, existing approaches suffer from slow trainin

model-releasesarxiv-cs-cv
10 Jun 2026
Research

NoiseSDF2NoiseSDF: Learning Clean Neural Fields from Noisy Supervision

DGX agent

arXiv:2507.13595v3 Announce Type: replace Abstract: Reconstructing accurate implicit surface representations from point clouds remains a challenging task, particularly when data is captured using low-

researcharxiv-cs-cv
10 Jun 2026
Agents

ObjSplat: Geometry-Aware Gaussian Surfels for Active Object Reconstruction

DGX agent

arXiv:2601.06997v2 Announce Type: replace-cross Abstract: Autonomous high-fidelity object reconstruction is fundamental for creating digital assets and bridging the simulation-to-reality gap in roboti

agentsarxiv-cs-cv
10 Jun 2026
Safety

On the Controllability-Fidelity Frontier in Diffusion Editing

DGX agent

arXiv:2606.09901v1 Announce Type: cross Abstract: Diffusion-based generative models enable powerful image editing capabilities, but achieving precise control while maintaining fidelity and safety rema

safetyarxiv-cs-cv
10 Jun 2026
Hardware

One-Step Residual Shifting Diffusion for Image Super-Resolution via Distillation

DGX agent

arXiv:2503.13358v5 Announce Type: replace Abstract: Diffusion models for super-resolution (SR) produce high-quality visual results but require expensive computational costs. Despite the development of

hardwarearxiv-cs-cv
10 Jun 2026
Research

Overlapped Wavelet Diffusion for Low-Light Image Enhancement

DGX agent

arXiv:2606.10280v1 Announce Type: cross Abstract: In this study, we propose an overlapped wavelet diffusion framework for Low-Light Image Enhancement (LLIE), which incorporates two complementary compo

researcharxiv-cs-cv
10 Jun 2026
Model Releases

P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning

DGX agent

arXiv:2606.11152v1 Announce Type: new Abstract: Multimodal large language models can write code to produce complex programs as well as use programs to do 3D modeling, which opens up a new avenue for 3

model-releasesarxiv-cs-cv
10 Jun 2026
← Previous
1…107108109110111…263
Next →