AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

MambaRaw: Selective State Space Modeling for Efficient 4K Raw Image Reconstruction

DGX agent

arXiv:2606.24479v1 Announce Type: new Abstract: In-camera JPEG previews are ubiquitous in raw image formats and provide an sRGB reference at negligible storage cost. Although existing metadata-based r

model-releasesarxiv-cs-cv
24 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Applications

MATCH: Flow Matching for Multi-View Anomaly Detection

DGX agent

arXiv:2606.24375v1 Announce Type: new Abstract: Detecting anomalies in industrial objects is an important topic for increasing production efficiency. More complex objects often require the analysis of

applicationsarxiv-cs-cv
24 Jun 2026
Agents

MM-TRELLIS: Point-Cloud Guided Multi-Modal 3D Vehicle Generation in Autonomous Driving

DGX agent

arXiv:2606.24301v1 Announce Type: new Abstract: Recovering realistic 3D vehicle models from autonomous driving scenes is crucial for synthesizing training data and building simulation environment. How

agentsarxiv-cs-cv
24 Jun 2026
Model Releases

Modality-Aware Out-of-Distribution Detection for Multi-Modal Action Recognition

DGX agent

arXiv:2606.24404v1 Announce Type: new Abstract: The incorporation of additional modalities into action recognition models increases their performance across a wide range of settings. However, how this

model-releasesarxiv-cs-cv
24 Jun 2026
Research

MorVess: Morphology-Aware Pulmonary Vessel Segmentation Network

DGX agent

arXiv:2606.24214v1 Announce Type: new Abstract: Accurate pulmonary vessel segmentation remains challenging due to the sparse, tortuous, and multi-scale nature of vascular structures, where small branc

researcharxiv-cs-cv
24 Jun 2026
Research

MotifGen: Spatiotemporal interpolation of misaligned satellite images via multi-source generative modeling, in an application to tropical cyclones

DGX agent

arXiv:2606.24263v1 Announce Type: new Abstract: Microwave satellite imagery plays a crucial role in monitoring tropical cyclone precipitation and intensity worldwide, but suffers from long revisit tim

researcharxiv-cs-cv
24 Jun 2026
Safety

MSPL: Multi-Step Pseudo-Labeling for Open-Vocabulary Object Detection

DGX agent

arXiv:2510.14792v4 Announce Type: replace Abstract: Open-vocabulary object detection (OVD) aims to recognize and localize object categories beyond the training set. Recent approaches leverage vision-l

safetyarxiv-cs-cv
24 Jun 2026
Research

Multilevel Stochastic Plug-and-Play for Sparse-View CT Reconstruction

DGX agent

arXiv:2606.24567v1 Announce Type: new Abstract: Sparse-view computed tomography (SVCT) reduces radiation exposure and acquisition time, but the limited number of projection views makes the reconstruct

researcharxiv-cs-cv
24 Jun 2026
Agents

NavWM: A Unified Navigation World Model for Foresight-Driven Planning

DGX agent

arXiv:2606.24101v1 Announce Type: cross Abstract: Conventional visual navigation policies often struggle with myopic decision-making and mode collapse in complex environments. While world models offer

agentsarxiv-cs-cv
24 Jun 2026
Local Ai

Neural Particle Automata: Learning Self-Organizing Particle Dynamics

DGX agent

arXiv:2601.16096v2 Announce Type: replace-cross Abstract: We introduce Neural Particle Automata (NPA), a Lagrangian generalization of Neural Cellular Automata (NCA) from static lattices to dynamic par

local-aiarxiv-cs-cv
24 Jun 2026
Agents

ObsGraph: Hierarchical Observation Representation for Embodied Reasoning and Exploration

DGX agent

arXiv:2606.24068v1 Announce Type: new Abstract: Embodied reasoning and exploration are increasingly considered crucial abilities for robots operating in complex and unfamiliar environments. To accompl

agentsarxiv-cs-cv
24 Jun 2026
Agents

Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints

DGX agent

arXiv:2606.24353v1 Announce Type: new Abstract: Bird's-eye view (BEV) perception fuses multi-camera images into a unified top-down representation for autonomous driving. Despite recent progress, state

agentsarxiv-cs-cv
24 Jun 2026
Research

P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling

DGX agent

arXiv:2606.24447v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have revolutionized document parsing by enabling end-to-end mapping from images to structured text, imposing a significant

researcharxiv-cs-cv
24 Jun 2026
Model Releases

PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments

DGX agent

arXiv:2606.24564v1 Announce Type: new Abstract: Reconstructing realistic, physically plausible garments from a single image remains a fundamental challenge. Template-free methods capture surface geome

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

Performance and Interpretability of Convolutional, Transformer, and Hybrid Deep Learning Models in Colorectal Histology Classification

DGX agent

arXiv:2606.23744v1 Announce Type: cross Abstract: Deep learning has become an important tool in computational pathology, enabling automated analysis of histopathological images. While convolutional ne

model-releasesarxiv-cs-cv
24 Jun 2026
Agents

Pocket-SLAM: Rendering-Area-Aware Pruning for Memory-Efficient 3DGS-SLAM

DGX agent

arXiv:2606.24796v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has garnered significant attention in Simultaneous Localization and Mapping (SLAM) due to its advances in capturing fine-gr

agentsarxiv-cs-cv
24 Jun 2026
Model Releases

Point-Voxel Absorbing Graph Representation Learning for Event Stream based Recognition

DGX agent

arXiv:2306.05239v3 Announce Type: replace Abstract: Sampled point and voxel methods are usually employed to downsample the dense events into sparse ones. After that, one popular way is to leverage a g

model-releasesarxiv-cs-cv
24 Jun 2026
Tutorials

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought

DGX agent

arXiv:2606.24539v1 Announce Type: new Abstract: Pointing-based visual grounding requires models to precisely locate target objects by deciphering complex spatial relationships between the visual scene

tutorialsarxiv-cs-cv
24 Jun 2026
Research

Predicting brain tumour enhancement from non-contrast MR imaging with artificial intelligence: a multi-cohort retrospective diagnostic accuracy study

DGX agent

arXiv:2508.16650v3 Announce Type: replace-cross Abstract: Brain tumour MRI typically requires both pre- and post-contrast imaging, but gadolinium is not always desirable (frequent follow-up, renal imp

researcharxiv-cs-cv
24 Jun 2026
Local Ai

Progressive Pixel-Neighborhood Deformable Cross-Attention for Multispectral Object Detection

DGX agent

arXiv:2606.24092v1 Announce Type: new Abstract: Effective cross-modal feature alignment and interaction are central challenges in multispectral object detection. Although global cross-attention provid

local-aiarxiv-cs-cv
24 Jun 2026
Local Ai

Quantum CT via Dynamic Interval Encoding and Prior-Balanced QUBO Reconstruction

DGX agent

arXiv:2606.24561v1 Announce Type: new Abstract: Quadratic unconstrained binary optimization (QUBO)-based quantum computed tomography (CT) casts reconstruction as a binary quadratic problem for quantum

local-aiarxiv-cs-cv
24 Jun 2026
Model Releases

REALM: A Unified Red-Teaming Benchmark for Physical-World VLMs

DGX agent

arXiv:2606.23892v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used as perception-reasoning backbones for embodied intelligence in safety-critical physical systems, whe

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

REDI-Match: Rotation-Equivariant Distillation for Efficient and Robust Dense Matching

DGX agent

arXiv:2606.24330v1 Announce Type: new Abstract: Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

Revealing Training Data Exposure in Vision Language Large Models via Parameter Gradients

DGX agent

arXiv:2606.24774v1 Announce Type: new Abstract: Vision-Language Large Models (VLLMs) trained on massive crawled corpora raise pressing copyright and data-provenance concerns. These concerns are partic

model-releasesarxiv-cs-cv
24 Jun 2026
Research

S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing

DGX agent

arXiv:2606.24441v1 Announce Type: new Abstract: We present S1-Omni-Image, an open-weight unified multimodal model for scientific image understanding, generation, and editing. Unlike general-purpose im

researcharxiv-cs-cv
24 Jun 2026
Applications

Sat2City v2: Native 3D City Asset Generation from a Single Satellite Image

DGX agent

arXiv:2606.24138v1 Announce Type: new Abstract: Generating explicit 3D city assets from a single satellite image is important for digital twins, urban simulation, and geospatial intelligence. Unlike s

applicationsarxiv-cs-cv
24 Jun 2026
Safety

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation

DGX agent

arXiv:2606.18610v2 Announce Type: replace-cross Abstract: Evaluating generalist robot manipulation policies in the real world is expensive, slow, and difficult to scale. Action-conditioned video world

safetyarxiv-cs-cv
24 Jun 2026
Research

Segmentation and Classification of Pap Smear Images for Cervical Cancer Detection Using Deep Learning

DGX agent

arXiv:2508.17728v2 Announce Type: replace Abstract: Cervical cancer remains a significant global health concern and a leading cause of cancer-related deaths among women. Early detection through Pap sm

researcharxiv-cs-cv
24 Jun 2026
Local Ai

SENTRY: SAM2-Enhanced Neighbor-Aware and Temporally Reasoned Memory for Visual Tracking

DGX agent

arXiv:2606.24449v1 Announce Type: new Abstract: We revisit the memory update mechanism in SAM2-based visual object tracking and identify confidence-only mask selection as the dominant cause of drift u

local-aiarxiv-cs-cv
24 Jun 2026
Model Releases

SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards

DGX agent

arXiv:2606.24726v1 Announce Type: new Abstract: Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Alth

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks

DGX agent

arXiv:2606.24361v1 Announce Type: new Abstract: Sign language models are typically trained on datasets captured under constrained conditions, with limited viewpoint, background, and signer-identity di

model-releasesarxiv-cs-cv
24 Jun 2026
Tutorials

Solving Semi-Supervised Few-Shot Learning from an Auto-Annotation Perspective

DGX agent

arXiv:2512.10244v2 Announce Type: replace Abstract: Semi-supervised few-shot learning (SSFSL) resembles real-world applications such as auto-annotation, as it aims to learn a model from a few labeled

tutorialsarxiv-cs-cv
24 Jun 2026
Safety

Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models

DGX agent

arXiv:2606.24165v1 Announce Type: new Abstract: Reducing visual token redundancy is critical for accelerating Multimodal Large Language Models (MLLMs) without degrading cross-modal reasoning performan

safetyarxiv-cs-cv
24 Jun 2026
Research

Spherical-to-ERP Epipolar Rectification for Single-Axis Disparity in 360 Stereo

DGX agent

arXiv:2606.24847v1 Announce Type: new Abstract: Omnidirectional stereo images provide full-surround perception but violate the geometric assumptions of classical disparity estimation: in spherical or

researcharxiv-cs-cv
24 Jun 2026
Model Releases

Systematic Exploration of 4-Expert Heterogeneous Mixture-of-Experts via Automated Pipeline Search

DGX agent

arXiv:2606.23739v1 Announce Type: cross Abstract: We present an automated large-scale search pipeline for heterogeneous 4-Expert Mixture-of-Experts (MoE4) architectures within the LEMUR neural network

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration

DGX agent

arXiv:2606.24336v1 Announce Type: new Abstract: Face Video Restoration (FVR) aims to recover high-fidelity facial videos from degraded input while preserving identity and semantic consistency across f

model-releasesarxiv-cs-cv
24 Jun 2026
Safety

Token-to-Token Alignment of Text Embeddings for Semantic Blending

DGX agent

arXiv:2606.24021v1 Announce Type: new Abstract: In modern generative models, images are specified and controlled through text prompts. In practice, images are generated from sequences of tokens derive

safetyarxiv-cs-cv
24 Jun 2026
Model Releases

Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling

DGX agent

arXiv:2606.24187v1 Announce Type: new Abstract: Long video understanding remains a daunting challenge for Multimodal Large Language Models (MLLMs) due to the excessive computation and memory footprint

model-releasesarxiv-cs-cv
24 Jun 2026
Research

Training-free Cross-domain Few-shot Segmentation via Robust Semantic Representation and Matching

DGX agent

arXiv:2606.24297v1 Announce Type: new Abstract: Cross-domain Few-shot Segmentation (CD-FSS) aims to transfer knowledge learned from source domain to distinct target domains, segmenting unseen target c

researcharxiv-cs-cv
24 Jun 2026
Model Releases

Tri-Efficient Transfer Learning for Point Cloud Videos

DGX agent

arXiv:2606.24175v1 Announce Type: new Abstract: While point cloud foundation models have significantly advanced point cloud video understanding, existing parameter-efficient fine-tuning (PEFT) methods

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

Trimming the Long-Tail of Visual World Modeling Evaluation

DGX agent

arXiv:2606.24256v1 Announce Type: new Abstract: Physical interactions follow a long-tailed distribution: a set of common and regular interactions dominates human experience and visual data, while a br

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

TrOCR for Medieval HTR: A Systematic Ablation Study with Cross-Dataset Validation

DGX agent

arXiv:2606.24302v1 Announce Type: new Abstract: Fine-tuning transformer-based handwritten text recognition (HTR) models on medieval manuscripts is challenging because these models are pre-trained on m

model-releasesarxiv-cs-cv
24 Jun 2026
Research

Trustworthy Image Authentication using Forensic Knowledge Graphs

DGX agent

arXiv:2606.23917v1 Announce Type: new Abstract: Advances in generative AI have made image falsification highly realistic, demanding trustworthy authentication systems. Existing forensic detectors can

researcharxiv-cs-cv
24 Jun 2026
Applications

TuringViT: Making SOTA Vision Transformers Accessible to All

DGX agent

arXiv:2606.24253v1 Announce Type: new Abstract: Modern VLMs and VLA systems commonly adopt off-the-shelf ViTs such as SigLIP2 as visual encoders, but diverse downstream requirements in latency, tempor

applicationsarxiv-cs-cv
24 Jun 2026
Research

Understanding Deep Representation Learning via Layerwise Feature Compression and Discrimination

DGX agent

arXiv:2311.02960v5 Announce Type: replace-cross Abstract: Over the past decade, deep learning has proven to be a highly effective tool for learning meaningful features from raw data. However, it remai

researcharxiv-cs-cv
24 Jun 2026
Model Releases

UniRED: Unified RGB-D Video Frame Interpolation with Event Guidance

DGX agent

arXiv:2606.24282v1 Announce Type: new Abstract: High frame-rate RGB-D videos are crucial for a variety of downstream tasks, including motion analysis, dynamic scene understanding, and 3D reconstructio

model-releasesarxiv-cs-cv
24 Jun 2026
Safety

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation

DGX agent

arXiv:2606.24333v1 Announce Type: new Abstract: In-Image Machine Translation (IIMT) aims to translate scene text in an image and render the translated text back into the original regions while preserv

safetyarxiv-cs-cv
24 Jun 2026
Agents

Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent

DGX agent

arXiv:2606.24094v1 Announce Type: new Abstract: Unifying image clustering across different clustering scenarios remains challenging due to fundamental gaps among tasks. We introduce a Guideline-Driven

agentsarxiv-cs-cv
24 Jun 2026
← Previous
1…9394959697…263
Next →