AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
10 Jun 2026

AnimaSpark: A Feed-Forward Method for Animating Arbitrary 3D Objects

SafetyDGX agent

arXiv:2606.10988v1 Announce Type: new Abstract: While recent advancements in generative AI have substantially accelerated static 3D model creation workflows, the synthesis of category-agnostic 3D anim

AnyMod-LLVE: Low-Light Video Enhancement with Modality-Agnostic Inference

ApplicationsDGX agent

arXiv:2606.11186v1 Announce Type: new Abstract: Low-light video enhancement (LLVE) remains a challenging task due to severe information degradation under low-illumination conditions. Recent multimodal

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations

SafetyDGX agent

arXiv:2606.11188v1 Announce Type: new Abstract: This paper introduces ARM, a discrete representation-based AutoRegressive Model that unifies image understanding, generation, and editing within a next-


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning

ResearchDGX agent

arXiv:2606.10533v1 Announce Type: new Abstract: Audio-visual captioning generates natural language descriptions from video and audio content. Multimodal LLMs have advanced this task, but both modaliti

Automatic Labelling for Low-Light Pedestrian Detection

SafetyDGX agent

arXiv:2507.02513v4 Announce Type: replace Abstract: Pedestrian detection in RGB images is a key task in pedestrian safety, as the most common sensor in autonomous vehicles and advanced driver assistan

Benchmarking stereo reconstruction for 3D printable Martian terrain models

Model ReleasesDGX agent

arXiv:2606.10364v1 Announce Type: new Abstract: Reconstructing printable 3D models from Mars rover imagery is challenging because Martian terrain is low-texture, irregular, and partially observed. We

Beyond Model Size: Probing the Gaps in Visual in-Context Learning by Training a Tiny Model

Model ReleasesDGX agent

arXiv:2606.10905v1 Announce Type: new Abstract: Visual in-Context Learning (VICL) aims at making progress towards adaptive vision models, that can -- based on a few examples -- adapt to a new task at

Breaking the Curse of Dimensionality: Diffusion Models Efficiently Learn Low-Dimensional Distributions

TutorialsDGX agent

arXiv:2409.02426v5 Announce Type: replace-cross Abstract: Despite their empirical success across a wide range of generative tasks, the fundamental principles underlying the ability of diffusion models

CapStARE: Capsule-based Sequential Architecture for Robust and Efficient Gaze Estimation

Local AiDGX agent

arXiv:2509.19936v2 Announce Type: replace Abstract: Human gaze estimation is essential for applications such as human-computer interaction, social robotics, and assistive systems. However, achieving a

ChartLens: A Dual-Branch Framework for Chart Data Correction and Factual Summary Refinement

Model ReleasesDGX agent

arXiv:2606.10640v1 Announce Type: new Abstract: In this report, we present our champion solution for the DataMFM Challenge Track 2: Chart Understanding. This track requires models to recover structure

ClinReadNet: A clinical reading-inspired network for low-dose abdominal CT image quality assessment

ResearchDGX agent

arXiv:2606.10372v1 Announce Type: new Abstract: In abdominal CT imaging, developing a low-dose, no-reference image quality assessment (No-reference IQA) model that mimics doctors' reading habits for e

CoCoSI: Collaborative Cognitive Map Construction for Spatial Intelligence

Model ReleasesDGX agent

arXiv:2606.10401v1 Announce Type: new Abstract: Spatial intelligence is a key frontier for multimodal large language models (MLLMs), enabling them to reason about the physical world from visual experi

Continuous Neural Reparameterization as a Deep Geometric Prior for Robust Fixed-Chart UV Repair

Model ReleasesDGX agent

arXiv:2606.10050v1 Announce Type: cross Abstract: Traditional UV unwrapping relies on direct optimization of geometric distortion energies and can fail through invalid initialization, local minima, or

Contrastive Spectral Rectification: Test-Time Defense towards Zero-shot Adversarial Robustness of CLIP

SafetyDGX agent

arXiv:2601.19210v2 Announce Type: replace Abstract: Vision-language models (VLMs) such as CLIP have demonstrated remarkable zero-shot generalization, yet remain highly vulnerable to adversarial exampl

Cost-Aware Routing for Efficient Text-To-Image Generation

ResearchDGX agent

arXiv:2506.14753v3 Announce Type: replace Abstract: Diffusion models are well known for their ability to generate a high-fidelity image for an input prompt through an iterative denoising process. Unfo

Cyst-X: A Multi-Center MRI Benchmark and Federated Learning Framework for Malignancy-Risk Stratification of Pancreatic Cystic Neoplasm

Model ReleasesDGX agent

arXiv:2507.22017v4 Announce Type: replace-cross Abstract: Pancreatic cancer is projected to be the second-deadliest cancer by 2030, making early detection critical. Intraductal papillary mucinous neop

DB-3DME: From Dataset to Benchmark for Human-aligned Automatic 3D Mesh Evaluation

Model ReleasesDGX agent

arXiv:2606.10142v1 Announce Type: new Abstract: Recent advances in 3D generation have led to substantial improvements in realism, controllability, and efficiency, yet the evaluation of 3D assets remai

DD-INR: Dynamics-Driven Implicit Neural Representation for Accelerated Whole-Brain Functional MRI Reconstruction

ResearchDGX agent

arXiv:2606.10756v1 Announce Type: new Abstract: Accelerated acquisition of fMRI enables enhanced detection of neurovascular (BOLD) activity in the brain, but image reconstruction becomes challenging w

Deep learning for echo sounder data

ResearchDGX agent

arXiv:2606.10811v1 Announce Type: new Abstract: There is no doubt that over the last decade, techniques from the field of machine learning have revolutionized how we process and interpret data, especi

Dexterous Point Policy: Learning Point-based Dexterous Hand Policies from Human Demonstrations

SafetyDGX agent

arXiv:2606.10614v1 Announce Type: cross Abstract: Robotic foundation models pre-trained on human demonstration videos have shown promise, but a significant embodiment gap remains when the resulting po

Dissect and Prune: Enhancing Robustness in AI-Generated Image Detection

SafetyDGX agent

arXiv:2606.10309v1 Announce Type: new Abstract: While existing AI-generated image detectors report high performance, we identify that this is largely driven by a critical prediction asymmetry: a bias

Don't waste SAM

Model ReleasesDGX agent

arXiv:2606.10696v1 Announce Type: new Abstract: Meta AI has recently released the Segment Anything Model (SAM), which demonstrates exceptional zero-shot image segmentation performance across various t

Dual-stream attention-guided learning for weakly supervised whole slide image classification

Local AiDGX agent

arXiv:2505.23341v3 Announce Type: replace Abstract: Whole slide images (WSIs) play a crucial role in cancer diagnosis due to their ultra-high resolution and rich morphological information, and multipl

Efficient RWKV-based Representation Learning for 3D Point Clouds

Model ReleasesDGX agent

arXiv:2606.10395v1 Announce Type: new Abstract: The recent receptance weighted key value (RWKV) model combines RNN-style recurrence, offering a linear-complexity alternative to Transformers' quadratic

Enabling Progressive Whole-slide Image Analysis with Multi-scale Pyramidal Network

Model ReleasesDGX agent

arXiv:2602.01951v2 Announce Type: replace Abstract: Multiple-instance Learning (MIL) is commonly used for computational pathology (CPath), where multi-scale features are essential for capturing both f

Envision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous Driving

AgentsDGX agent

arXiv:2606.10656v1 Announce Type: new Abstract: Forecasting the future evolution of dynamic scenes is crucial in autonomous driving. However, existing feed-forward paradigms are primarily designed for

FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion

ResearchDGX agent

arXiv:2606.10671v1 Announce Type: new Abstract: Autoregressive video generators synthesize long videos by generating successive temporal segments, but their historical KV cache grows with video length

Few-step Generative Models as Lossy Compression

ResearchDGX agent

arXiv:2606.10450v1 Announce Type: new Abstract: DiffC provides a principled way to reuse pre-trained diffusion models for lossy compression, but its encoding and decoding procedures remain slow becaus

FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models

ResearchDGX agent

arXiv:2509.16518v2 Announce Type: replace Abstract: Using diffusion transformers for media generation may require evaluating attention over extremely long sequences, with attention layers accounting f

FlexPath: Learned Semantic Path Priors for Image-Based Planning

ResearchDGX agent

arXiv:2606.10167v1 Announce Type: new Abstract: Recent learning-based path planners use neural networks to process visual map representations and approximate heuristics for classical search algorithms

FoA-SR: Faithful or Aesthetic? Profile-Aware Preference Optimization for Real-World Image Super-Resolution

ApplicationsDGX agent

arXiv:2606.10275v1 Announce Type: new Abstract: Real-world image super-resolution (SR) is often designed with a single restoration objective, despite the current capacity of generative models to produ

From Patches to Patients: A study of the tile-to-slide performance transferability in Digital Pathology

Model ReleasesDGX agent

arXiv:2606.10778v1 Announce Type: new Abstract: Foundation Models (FMs) have recently redefined the state-of-the-art in histopathology by providing robust representations for whole-slide image (WSI) a

FSS-Net: Frequency-Spatial Synergy Network with Wavelet Attention for Carotid Artery Ultrasound Segmentation

ResearchDGX agent

arXiv:2606.10378v1 Announce Type: new Abstract: Accurate segmentation of carotid arteries in ultrasound imaging is critical for stroke risk assessment. However, speckle noise, low contrast, and blurre

Fusing Satellite Imagery and Planimetric Maps for Cross-View Localization

ResearchDGX agent

arXiv:2606.10166v1 Announce Type: new Abstract: Current cross-view localization methods predominantly rely on satellite imagery as the aerial modality. Although recent work explores planimetric maps (

GaussTrace: Provenance Analysis of 3D Gaussian Splatting Models with Evidence-based LLM Reasoning

ResearchDGX agent

arXiv:2606.10612v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) is a powerful technique for creating high-fidelity 3D assets. However, the widespread sharing and iterative modification of

GeoLoom: High-quality Geometric Diagram Generation from Textual Input

TutorialsDGX agent

arXiv:2512.08180v2 Announce Type: replace Abstract: High-quality geometric diagram generation presents both a challenge and an opportunity: it demands strict spatial accuracy while offering well-defin

Geometric Coastline Localization using Vision-Language Models

Local AiDGX agent

arXiv:2606.10468v1 Announce Type: new Abstract: Coastline detection in remote sensing imagery is commonly formulated as a pixel-wise segmentation problem, where the final coastline is extracted from a

Geometry-Aware Reinforcement Learning for 2D Irregular Nesting

Model ReleasesDGX agent

arXiv:2606.10611v1 Announce Type: cross Abstract: Traditional heuristic solvers for the 2D irregular nesting problem share a fundamental limitation: they are blind to polygon geometry, relying on guid

GHOST: Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation

SafetyDGX agent

arXiv:2606.10025v1 Announce Type: cross Abstract: We present GHOST, a framework for learning visuomotor manipulation policies that generalize beyond the training distribution. GHOST factorizes control

Globally Localizing Lunar Rover in Pixels via Graph Alignment

SafetyDGX agent

arXiv:2606.10602v1 Announce Type: new Abstract: Precise rover localization is a prerequisite for autonomous lunar exploration, yet the absence of Global Navigation Satellite System (GNSS) signals and

Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves

ResearchDGX agent

arXiv:2603.20850v2 Announce Type: replace Abstract: Understanding hand-object interaction (HOI) is fundamental to computer vision, robotics, and AR/VR. However, conventional hand videos often lack ess

GRAR: Glass-induced Reflection Artifact Removal in LiDAR Point Clouds

ResearchDGX agent

arXiv:2606.10541v1 Announce Type: new Abstract: Terrestrial Laser Scanning (TLS) point clouds captured in urban environments frequently suffer from glass-induced reflection artifacts, severely degradi

GUI-AC: Enhancing Continual Learning in GUI Agents

SafetyDGX agent

arXiv:2606.10522v1 Announce Type: new Abstract: Graphical User Interfaces (GUIs) serve as the dominant medium for human-computer interaction, yet building GUI agents that generalize across the vast di

HarmoView: Harmonizing Multi-View Constraints for Identity-Consistent Video Generation

Model ReleasesDGX agent

arXiv:2606.10839v1 Announce Type: new Abstract: Current identity-consistent video generation methods struggle to preserve appearance fidelity under large viewpoint changes. While introducing multi-vie

IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder

SafetyDGX agent

arXiv:2606.11096v1 Announce Type: new Abstract: Built on pretrained vision foundation models (VFMs), representation autoencoders (RAEs) have recently emerged as a promising approach for constructing s

IMPACT: Learning Internal-Model Predictive Control for Forceful Robotic Manipulation

SafetyDGX agent

arXiv:2606.10818v1 Announce Type: cross Abstract: Real-world robotic manipulation tasks often involve forceful interactions with the environment, such as using tools of varying weights, transporting o

Improving PET/CT-Based Whole-Body Lesion Segmentation Using Prediction Uncertainty-Augmented Models

Local AiDGX agent

arXiv:2606.10115v1 Announce Type: new Abstract: Accurate lesion segmentation from whole-body Positron Emission Tomography (PET)/Computed Tomography (CT) scans is essential for cancer staging and treat

Interpretable Temporal Facial-Region Motion Analysis for In-the-Wild Parkinson's Disease Video Classification

Model ReleasesDGX agent

arXiv:2606.10088v1 Announce Type: new Abstract: Reduced facial expressivity is a common motor manifestation of Parkinson's disease (PD), often described as hypomimia or facial bradykinesia. This paper

IPSM-Bench: A New Intermediate Phase Segmentation Benchmark in Microstructure Images of Zinc-Based Absorbable Biomaterials

Model ReleasesDGX agent

arXiv:2606.11001v1 Announce Type: new Abstract: Zinc-based alloys are indispensable emerging absorbable metallic biomaterials, and their macroscopic performance is governed by microstructural characte

Is Task-Specific Training Necessary for Anomaly Detection?

ResearchDGX agent

arXiv:2601.22763v3 Announce Type: replace Abstract: Current state-of-the-art multi-class unsupervised anomaly detection (MUAD) methods rely on training encoder--decoder models to reconstruct anomaly-f

iSAGE: A Human-in-the-Loop Framework for Remote Sensing Semantic Segmentation via Sparse Point Supervision

Model ReleasesDGX agent

arXiv:2606.10136v1 Announce Type: new Abstract: Semantic segmentation in remote sensing requires costly pixel-level annotations, and nearly every problem demands a new dataset since models rarely tran

Kwai Keye-VL-2.0 Technical Report

Model ReleasesDGX agent

arXiv:2606.10651v1 Announce Type: new Abstract: We introduce Kwai Keye-VL-2.0-30B-A3B, an open-source Mixture-of-Experts (MoE) multimodal foundation model designed to advance long-video understanding

LAFP: Preserving Latent Action Structure in Latent Policy Learning via Flow Matching

SafetyDGX agent

arXiv:2606.10517v1 Announce Type: new Abstract: Learning high-quality latent actions from large-scale unlabeled videos, coupled with limited real-world interaction data for training an action decoder,

LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning

ApplicationsDGX agent

arXiv:2504.18424v2 Announce Type: replace Abstract: We present Layered Ray Intersections (LaRI), a fully supervised method for occluded geometry reasoning from a single image. Unlike conventional dept

Leveraging Metric Depth for Relative Depth Prediction

TutorialsDGX agent

arXiv:2606.10628v1 Announce Type: new Abstract: We present our solution to the 2025 SoccerNet Monocular Depth Estimation Competition Challenge. Predicting the relative depth in football scenarios is c

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization

SafetyDGX agent

arXiv:2606.11180v1 Announce Type: new Abstract: Diffusion-based lip synchronization models achieve strong visual quality and audio-visual alignment, but full-sequence bidirectional attention and many

Listen, Look, and Learn: Learning Without Forgetting through SAM-Audio

TutorialsDGX agent

arXiv:2606.10887v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) aims to continuously learn new classes without forgetting previously acquired knowledge. While recent CIL advances have

ManiSplat: Manipulation Trajectory Synthesis from Monocular Video via Decoupled 3D Gaussian Splatting

SafetyDGX agent

arXiv:2606.10645v1 Announce Type: new Abstract: Reconstructing dynamic and interactive 3D scenes from real-world observations remains a fundamental challenge in computer vision and robotics. While rec

Maximum Matching Accuracy: An Instance Segmentation Evaluation Metric Utilizing Globally Optimal Matching

ResearchDGX agent

arXiv:2606.10107v1 Announce Type: new Abstract: Reliable evaluation of instance segmentation models requires metrics that accurately and consistently reflect segmentation quality. However, the metrics

Mean Flow Distillation: Robust and Stable Distillation for Flow Matching Models

SafetyDGX agent

arXiv:2606.11155v1 Announce Type: new Abstract: Flow Matching models have demonstrated strong performance across a wide range of generative tasks. However, their reliance on ODE-based iterative sampli

← Previous
1…8586878889…211
Next →