AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Applications

Enhancing Computer Vision Model Generalization in Warehouse Facilities: A Case Study on Anomaly Detection in Vertical Material Handling Systems

DGX agent

arXiv:2605.31487v1 Announce Type: new Abstract: Deploying computer vision models in Warehouse Facilities traditionally requires extensive resources for camera mounting, image collection, annotation, t

applicationsarxiv-cs-cv
1 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Equivariant Latent Alignment via Flow Matching under Group Symmetries

DGX agent

arXiv:2605.30705v1 Announce Type: new Abstract: Geometry-aware generative models and novel view synthesis approaches have shown strong potential in visual fidelity and consistency. In parallel, equiva

safetyarxiv-cs-cv
1 Jun 2026
Local Ai

Fixed-Point Masked Generative Modeling

DGX agent

arXiv:2605.31215v1 Announce Type: cross Abstract: Masked Generative Models (MGMs) enable parallel decoding and achieve strong performance across modalities, but require full-sequence bidirectional tra

local-aiarxiv-cs-cv
1 Jun 2026
Research

Foundation VAEs for 3D CT Reconstruction, Augmentation, and Generation

DGX agent

arXiv:2605.30893v1 Announce Type: new Abstract: Variational autoencoders (VAEs) compress high resolution CT volumes into compact latents while preserving clinically relevant structure. However, traini

researcharxiv-cs-cv
1 Jun 2026
Research

From Local Geometry to Global Pseudo Labeling for Robust Positive Unlabeled Learning under Covariate Shift

DGX agent

arXiv:2605.31187v1 Announce Type: new Abstract: Detecting covariate shift is critical for building reliable vision systems. While most prior work focuses on improving robustness to shift, explicitly d

researcharxiv-cs-cv
1 Jun 2026
Model Releases

FSM-Net: An Efficient Frequency-Spatial Network for Real-World Deblurring

DGX agent

arXiv:2605.31400v1 Announce Type: new Abstract: Real-world image deblurring demands both high-fidelity restoration and computational efficiency, a balance existing methods often struggle to achieve. I

model-releasesarxiv-cs-cv
1 Jun 2026
Tutorials

Function2Scene: 3D Indoor Scene Layout from Functional Specifications

DGX agent

arXiv:2605.30819v1 Announce Type: new Abstract: Most text-driven 3D indoor scene synthesis methods generate rooms from object-centric prompts, asking what furniture should be placed rather than how th

tutorialsarxiv-cs-cv
1 Jun 2026
Applications

GGT-100K: Generative Ground Truth for Generalizable Real-World Image Restoration

DGX agent

arXiv:2605.31039v1 Announce Type: new Abstract: Real-world image restoration (IR) is bottlenecked by the scarcity of high-quality paired training data. Synthetic datasets are abundant but often fail t

applicationsarxiv-cs-cv
1 Jun 2026
Model Releases

GUI-C^2: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning

DGX agent

arXiv:2605.30884v1 Announce Type: new Abstract: Existing agentic reinforcement learning methods for GUI grounding have limitations at two levels. At the data level, current approaches typically treat

model-releasesarxiv-cs-cv
1 Jun 2026
Tutorials

Guidance for Low-Level Perceptual Editing in Unconditional Diffusion Models

DGX agent

arXiv:2605.31162v1 Announce Type: new Abstract: Unconditional diffusion models offer powerful generative priors, yet steering them toward aesthetically enhanced outputs remains largely unexplored. We

tutorialsarxiv-cs-cv
1 Jun 2026
Research

HiERO-StepG @ Ego4D Step Grounding Challenge: hierarchical activity understanding enables zero-shot step grounding

DGX agent

arXiv:2605.31227v1 Announce Type: new Abstract: Procedural activities follow well-defined structures: whether we consider a cooking recipe or a mechanic repairing a car, these activities naturally dec

researcharxiv-cs-cv
1 Jun 2026
Tutorials

How can embedding models bind concepts?

DGX agent

arXiv:2605.31503v1 Announce Type: new Abstract: Humans easily determine which color belongs to which shape in multi-object scenes, an ability known as concept binding. Vision-language embedding models

tutorialsarxiv-cs-cv
1 Jun 2026
Safety

HQ-JEPA: Hybrid Quantum Joint-Embedding Predictive Architecture for Cross-Modal Remote Sensing Representation Learning

DGX agent

arXiv:2605.31068v1 Announce Type: new Abstract: We introduce HQ-JEPA, a hybrid quantum-classical joint-embedding predictive architecture for cross-modal remote sensing representation learning. The pro

safetyarxiv-cs-cv
1 Jun 2026
Research

HUNT: High-Speed UAV Navigation and Tracking in Unstructured Environments via Instantaneous Relative Frames

DGX agent

arXiv:2509.19452v4 Announce Type: replace-cross Abstract: Search and rescue operations require unmanned aerial vehicles to both traverse unknown unstructured environments at high speed and track targe

researcharxiv-cs-cv
1 Jun 2026
Local Ai

Hyperspectral Image Classification using Spectral-Spatial Mixer Network

DGX agent

arXiv:2511.15692v2 Announce Type: replace Abstract: This paper introduces SS-MixNet, a lightweight and effective deep learning model for hyperspectral image (HSI) classification. The architecture inte

local-aiarxiv-cs-cv
1 Jun 2026
Agents

IAF-Net: Illumination-Adaptive Fusion for Low-Light Urban Road Segmentation

DGX agent

arXiv:2605.30939v1 Announce Type: new Abstract: Semantic road segmentation is important for autonomous driving, but existing methods suffer severe performance degradation under low-light conditions. M

agentsarxiv-cs-cv
1 Jun 2026
Research

Immuno-VLM: Immunizing Large Vision-Language Models via Generative Semantic Antibodies for Open-World Trustworthiness

DGX agent

arXiv:2605.30745v1 Announce Type: new Abstract: Large Vision-Language Models have achieved unprecedented success in zero-shot recognition by aligning visual features with broad semantic concepts. Howe

researcharxiv-cs-cv
1 Jun 2026
Model Releases

Inference-Free Multimodal Learned Sparse Retrieval for Production-Scale Visual Document Search

DGX agent

arXiv:2605.30917v1 Announce Type: cross Abstract: As large-scale visual-document corpora such as arXiv papers and enterprise PDFs continue to grow, visual-document retrieval has gained increasing atte

model-releasesarxiv-cs-cv
1 Jun 2026
Tutorials

Internalizing Temporal Consistency in Video Object-Centric Learning without Explicit Regularization

DGX agent

arXiv:2605.31508v1 Announce Type: new Abstract: Video Object-Centric Learning (OCL) aims to represent objects as extit{slot} vectors and maintain their consistency across frames. Slot-Slot Contrastive

tutorialsarxiv-cs-cv
1 Jun 2026
Tutorials

Interpretability Without Tradeoffs: Disentangling Polysemanticity At Equal Predictive Performance

DGX agent

arXiv:2605.31304v1 Announce Type: cross Abstract: Deep neural networks (DNNs) are widely used, but interpreting what they actually learn remains difficult. A major obstacle is that individual neurons

tutorialsarxiv-cs-cv
1 Jun 2026
Safety

IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal Alignment

DGX agent

arXiv:2603.19862v2 Announce Type: replace Abstract: Vision-Language Models like CLIP are extensively used for inter-modal tasks which involve both visual and text modalities. However, when the individ

safetyarxiv-cs-cv
1 Jun 2026
Research

Iterative Framework For Data Augmentation Of Segmented Fingerprints

DGX agent

arXiv:2605.31001v1 Announce Type: new Abstract: Infant biometrics presents unique challenges due to the physiological differences between infants and adults, compounded by the scarcity of available da

researcharxiv-cs-cv
1 Jun 2026
Local Ai

iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning

DGX agent

arXiv:2605.31096v1 Announce Type: new Abstract: While visually grounded Chain-of-Thought (CoT) has emerged as a promising paradigm to enhance fine-grained perception in multimodal large language model

local-aiarxiv-cs-cv
1 Jun 2026
Research

Joint Multi-Camera LiDAR Extrinsic Calibration via Learned Pairwise Initialization and Geometric Refinement

DGX agent

arXiv:2605.31576v1 Announce Type: new Abstract: Most learning-based camera-LiDAR calibration methods treat each camera-LiDAR pair independently, ignoring the rigid geometric coupling in multi-camera p

researcharxiv-cs-cv
1 Jun 2026
Research

KLIP: localized distribution shift detection via KL-divergence with diffusion priors in Inverse Problems

DGX agent

arXiv:2605.31596v1 Announce Type: new Abstract: Diffusion models have shown promising performance as data-driven priors for computational imaging, as well as some capacity to detect out-of-distributio

researcharxiv-cs-cv
1 Jun 2026
Model Releases

LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation

DGX agent

arXiv:2602.02220v2 Announce Type: replace Abstract: Language-conditioned goal navigation (LGN) requires agents to locate user-specified targets without step-by-step guidance. However, existing benchma

model-releasesarxiv-cs-cv
1 Jun 2026
Research

Latent Geometric Chords for Query-Efficient Decision-Based Adversarial Attacks

DGX agent

arXiv:2605.31219v1 Announce Type: new Abstract: While decision-based black-box adversarial attacks present a severe security threat, current methodologies suffer from fundamental limitations. Pixel-wi

researcharxiv-cs-cv
1 Jun 2026
Research

Learning Global Motion with Compact Gaussians for Feed-Forward 4D Reconstruction

DGX agent

arXiv:2605.31595v1 Announce Type: new Abstract: Dynamic scene reconstruction from monocular video remains a fundamental challenge in computer vision. Existing feed-forward methods predict 3D Gaussians

researcharxiv-cs-cv
1 Jun 2026
Model Releases

LegSegNet: A Public Deep Learning System for Lower Extremity CT Tissue Segmentation and Quantification

DGX agent

arXiv:2605.30829v1 Announce Type: new Abstract: Lower extremity computed tomography (CT) contains clinically relevant information for body composition analysis, sarcopenia assessment, and musculoskele

model-releasesarxiv-cs-cv
1 Jun 2026
Safety

LiftNav: Path Planning via Semantic Lifting in TSDF-Guided Gaussian Splatting

DGX agent

arXiv:2605.31376v1 Announce Type: cross Abstract: Autonomous robots in unknown indoor environments require both reliable collision avoidance and object-level understanding. Classical representations s

safetyarxiv-cs-cv
1 Jun 2026
Local Ai

Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models

DGX agent

arXiv:2605.31158v1 Announce Type: new Abstract: Interactive video world models generate video chunk by chunk in response to user-controlled camera movements, enabling applications such as real-time ga

local-aiarxiv-cs-cv
1 Jun 2026
Research

Lightweight SAR Ship Detection via Contrastive Distillation

DGX agent

arXiv:2605.30380v1 Announce Type: new Abstract: Deep convolutional and transformer-based detectors achieve strong performance for SAR ship detection but are often computationally prohibitive for real-

researcharxiv-cs-cv
1 Jun 2026
Research

Linear Scaling Video VLMs for Long Video Understanding

DGX agent

arXiv:2605.31598v1 Announce Type: new Abstract: Video vision-language models (VLMs) are increasingly used in long-horizon and streaming settings, yet most video encoders still rely on spatiotemporal s

researcharxiv-cs-cv
1 Jun 2026
Tutorials

LPTR-AFLNet: Lightweight Integrated Chinese License Plate Rectification and Recognition Network

DGX agent

arXiv:2507.16362v3 Announce Type: replace Abstract: Chinese License Plate Recognition (CLPR) faces numerous challenges in unconstrained and complex environments, particularly due to perspective distor

tutorialsarxiv-cs-cv
1 Jun 2026
Safety

LVSA: Training-Free Sparse Attention for Long Video Diffusion

DGX agent

arXiv:2605.31057v1 Announce Type: new Abstract: Dense self-attention is the compute and quality bottleneck of long-video diffusion inference: cost grows quadratically with the sequence length, and bey

safetyarxiv-cs-cv
1 Jun 2026
Research

Mathematical Morphology in Machine Learning

DGX agent

arXiv:2605.30700v1 Announce Type: new Abstract: This work introduces mathematical morphology-an established visual computing theory-into machine learning to exploit shape and density aspects often ove

researcharxiv-cs-cv
1 Jun 2026
Safety

MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging

DGX agent

arXiv:2605.30904v1 Announce Type: new Abstract: Most visual tokenizers for image generation are bifurcated into two families with complementary limitations: continuous VAEs offer high-fidelity reconst

safetyarxiv-cs-cv
1 Jun 2026
Research

Mitigating Content Shift and Hallucination in GenAI Image Editing via Structural Refinement

DGX agent

arXiv:2605.30437v1 Announce Type: new Abstract: Generative AI (GenAI) image editors, such as Nano Banana, produce visually compelling results for retouching tasks, enabling non-experts to edit images

researcharxiv-cs-cv
1 Jun 2026
Research

MoE-dqINR: A Unified Mixture-of-Experts Implicit Neural Representation Framework for Scan-Specific Dynamic and Quantitative MRI Reconstruction

DGX agent

arXiv:2605.31302v1 Announce Type: cross Abstract: Undersampled magnetic resonance imaging (MRI) reconstruction seeks to recover temporally or contrast-varying image series from incomplete multicoil k-

researcharxiv-cs-cv
1 Jun 2026
Research

MultiAct: Text-to-Motion Generation from Composite Text via Tailored Attention Guidance

DGX agent

arXiv:2605.30925v1 Announce Type: new Abstract: Text-to-motion generation has progressed rapidly in recent years, offering an expressive interface for animation and human-computer interaction. However

researcharxiv-cs-cv
1 Jun 2026
Research

Multimodal Fusion via Self-Consistent Task-Gradient Fields

DGX agent

arXiv:2410.15475v2 Announce Type: replace Abstract: Multimodal learning aims to preserve as much task-related information as possible from different inputs. However, current fusion designs often disto

researcharxiv-cs-cv
1 Jun 2026
Model Releases

MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models

DGX agent

arXiv:2511.16940v3 Announce Type: replace Abstract: Modern Vision-Language Models (VLMs) pose significant individual-level privacy risks by linking fragmented multimodal data to identifiable individua

model-releasesarxiv-cs-cv
1 Jun 2026
Hardware

Neurosim: A Fast Simulator for Neuromorphic Robot Perception

DGX agent

arXiv:2602.15018v2 Announce Type: replace-cross Abstract: Neurosim is a fast, real-time, high-performance library for simulating sensors such as dynamic vision sensors, RGB cameras, depth sensors, and

hardwarearxiv-cs-cv
1 Jun 2026
Research

Non-Parametric Probabilistic Robustness: A Conservative Risk Estimator under Unknown Perturbation Distributions

DGX agent

arXiv:2511.17380v2 Announce Type: replace Abstract: Deep learning (DL) models, despite their remarkable success, remain vulnerable to small input perturbations that can cause erroneous outputs, motiva

researcharxiv-cs-cv
1 Jun 2026
Agents

NTR: Neural Token Reconstruction for Scene Token Bottleneck in End-to-End Driving

DGX agent

arXiv:2605.31116v1 Announce Type: new Abstract: Recent perception-free end-to-end (E2E) autonomous driving methods bypass explicit perception outputs by compressing dense image patch tokens into compa

agentsarxiv-cs-cv
1 Jun 2026
Model Releases

nuReasoning: A Reasoning-Centric Dataset and Benchmark for Long-Tail Autonomous Driving

DGX agent

arXiv:2605.31572v1 Announce Type: new Abstract: Reasoning is essential for autonomous driving (AD) in long-tail scenarios, where vehicles must apply commonsense knowledge, understand spatial relations

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Omni-Supervised Motion Editing: Balancing Change and Invariance through Positive-Negative Learning

DGX agent

arXiv:2605.30969v1 Announce Type: new Abstract: Text-based human motion editing aims to modify existing motion sequences according to natural language instructions while maintaining the consistency of

model-releasesarxiv-cs-cv
1 Jun 2026
Safety

OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation

DGX agent

arXiv:2605.30519v1 Announce Type: new Abstract: Autoregressive (AR) video generation extends videos by producing latent chunks sequentially, but scaling to long videos requires repeated access to a gr

safetyarxiv-cs-cv
1 Jun 2026
← Previous
1…132133134135136…263
Next →