AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
24 Jun 2026

High-Fidelity Synthetic Transmission Electron Microscopy Image Generation Using Diffusion Probabilistic Models for Data-Limited Semiconductor Metrology

ResearchDGX agent

arXiv:2606.24817v1 Announce Type: new Abstract: Advanced semiconductor nodes drastically increased demand for Transmission Electron Microscopy (TEM), yet destructive sample preparation, slow imaging a

Hybrid Event Frame Sensors: Modeling, Calibration, and Simulation

ResearchDGX agent

arXiv:2511.18037v2 Announce Type: replace Abstract: Hybrid event-frame sensors integrate an Event Vision Sensor (EVS) and an Active Pixel Sensor (APS) within a single chip, combining the high dynamic

Ill-Posed by Design: Probing Evidence Use in VLMs

Local AiDGX agent

arXiv:2606.24335v1 Announce Type: new Abstract: Counterfactual analysis is widely used to study evidence use in vision-language models, but its diagnostic value is limited on well-posed tasks: when se


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Ingredient-Level Food Image Segmentation for Nutrition Awareness

ResearchDGX agent

arXiv:2606.24059v1 Announce Type: new Abstract: Food images often contain several visible ingredients, so assigning one dish label to an entire image hides important visual structure. This work studie

Jolia: Concept-Level Vision-Language Alignment for 3D CT Contrastive Learning

Model ReleasesDGX agent

arXiv:2606.24570v1 Announce Type: new Abstract: Vision-language contrastive pretraining has become the dominant recipe for 3D medical foundation models, leveraging the large volumes of paired scans an

Latent Visual States for Efficient Multimodal Reasoning

SafetyDGX agent

arXiv:2606.24233v1 Announce Type: new Abstract: The integration of visual evidence has significantly enhanced the capabilities of large multimodal models. However, this integration predominantly relie

Lite Any Stereo V2: Faster and Stronger Efficient Zero-Shot Stereo Matching

ApplicationsDGX agent

arXiv:2606.24457v1 Announce Type: new Abstract: Recent advances in stereo matching have achieved remarkable accuracy, but often rely on large models, heavy computation, or additional foundation-model

LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation

Model ReleasesDGX agent

arXiv:2509.17773v2 Announce Type: replace Abstract: The rapid progress of image-guided video generation (I2V) has raised concerns about its potential misuse in misinformation and fraud, underscoring t

M^2C-EvDet: Multi-Domain Multi-Order Cross-Modal Knowledge Distillation for Event-based Object Detection

ResearchDGX agent

arXiv:2606.24248v1 Announce Type: new Abstract: Event-based object Detection (EvDet), as a biologically inspired visual perception paradigm, demonstrates superior performance in scenarios demanding hi

M4-SAR: A Multi-Resolution, Multi-Polarization, Multi-Scene, Multi-Source Dataset and Benchmark for optical-SAR Object Detection

Model ReleasesDGX agent

arXiv:2505.10931v4 Announce Type: replace Abstract: Single-source remote sensing object detection using optical or SAR images struggles in complex environments. Optical images offer rich textural deta

Machine Learning Modeling for Real-Time Melt Pool Monitoring in Laser Powder Bed Fusion Additive Manufacturing: A Hybrid Approach

Model ReleasesDGX agent

arXiv:2606.23851v1 Announce Type: cross Abstract: This work investigates the implementation of artificial intelligence and machine learning (AI/ML) for real-time monitoring in laser powder bed fusion

Mamba-FSCIL: Dynamic Adaptation with Selective State Space Model for Few-Shot Class-Incremental Learning

Model ReleasesDGX agent

arXiv:2407.06136v4 Announce Type: replace Abstract: Few-shot class-incremental learning (FSCIL) aims to incrementally learn novel classes from limited examples while preserving knowledge of previously

MambaRaw: Selective State Space Modeling for Efficient 4K Raw Image Reconstruction

Model ReleasesDGX agent

arXiv:2606.24479v1 Announce Type: new Abstract: In-camera JPEG previews are ubiquitous in raw image formats and provide an sRGB reference at negligible storage cost. Although existing metadata-based r

MATCH: Flow Matching for Multi-View Anomaly Detection

ApplicationsDGX agent

arXiv:2606.24375v1 Announce Type: new Abstract: Detecting anomalies in industrial objects is an important topic for increasing production efficiency. More complex objects often require the analysis of

MM-TRELLIS: Point-Cloud Guided Multi-Modal 3D Vehicle Generation in Autonomous Driving

AgentsDGX agent

arXiv:2606.24301v1 Announce Type: new Abstract: Recovering realistic 3D vehicle models from autonomous driving scenes is crucial for synthesizing training data and building simulation environment. How

Modality-Aware Out-of-Distribution Detection for Multi-Modal Action Recognition

Model ReleasesDGX agent

arXiv:2606.24404v1 Announce Type: new Abstract: The incorporation of additional modalities into action recognition models increases their performance across a wide range of settings. However, how this

MorVess: Morphology-Aware Pulmonary Vessel Segmentation Network

ResearchDGX agent

arXiv:2606.24214v1 Announce Type: new Abstract: Accurate pulmonary vessel segmentation remains challenging due to the sparse, tortuous, and multi-scale nature of vascular structures, where small branc

MotifGen: Spatiotemporal interpolation of misaligned satellite images via multi-source generative modeling, in an application to tropical cyclones

ResearchDGX agent

arXiv:2606.24263v1 Announce Type: new Abstract: Microwave satellite imagery plays a crucial role in monitoring tropical cyclone precipitation and intensity worldwide, but suffers from long revisit tim

MSPL: Multi-Step Pseudo-Labeling for Open-Vocabulary Object Detection

SafetyDGX agent

arXiv:2510.14792v4 Announce Type: replace Abstract: Open-vocabulary object detection (OVD) aims to recognize and localize object categories beyond the training set. Recent approaches leverage vision-l

Multilevel Stochastic Plug-and-Play for Sparse-View CT Reconstruction

ResearchDGX agent

arXiv:2606.24567v1 Announce Type: new Abstract: Sparse-view computed tomography (SVCT) reduces radiation exposure and acquisition time, but the limited number of projection views makes the reconstruct

NavWM: A Unified Navigation World Model for Foresight-Driven Planning

AgentsDGX agent

arXiv:2606.24101v1 Announce Type: cross Abstract: Conventional visual navigation policies often struggle with myopic decision-making and mode collapse in complex environments. While world models offer

Neural Particle Automata: Learning Self-Organizing Particle Dynamics

Local AiDGX agent

arXiv:2601.16096v2 Announce Type: replace-cross Abstract: We introduce Neural Particle Automata (NPA), a Lagrangian generalization of Neural Cellular Automata (NCA) from static lattices to dynamic par

ObsGraph: Hierarchical Observation Representation for Embodied Reasoning and Exploration

AgentsDGX agent

arXiv:2606.24068v1 Announce Type: new Abstract: Embodied reasoning and exploration are increasingly considered crucial abilities for robots operating in complex and unfamiliar environments. To accompl

Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints

AgentsDGX agent

arXiv:2606.24353v1 Announce Type: new Abstract: Bird's-eye view (BEV) perception fuses multi-camera images into a unified top-down representation for autonomous driving. Despite recent progress, state

P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling

ResearchDGX agent

arXiv:2606.24447v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have revolutionized document parsing by enabling end-to-end mapping from images to structured text, imposing a significant

PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments

Model ReleasesDGX agent

arXiv:2606.24564v1 Announce Type: new Abstract: Reconstructing realistic, physically plausible garments from a single image remains a fundamental challenge. Template-free methods capture surface geome

Performance and Interpretability of Convolutional, Transformer, and Hybrid Deep Learning Models in Colorectal Histology Classification

Model ReleasesDGX agent

arXiv:2606.23744v1 Announce Type: cross Abstract: Deep learning has become an important tool in computational pathology, enabling automated analysis of histopathological images. While convolutional ne

Pocket-SLAM: Rendering-Area-Aware Pruning for Memory-Efficient 3DGS-SLAM

AgentsDGX agent

arXiv:2606.24796v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has garnered significant attention in Simultaneous Localization and Mapping (SLAM) due to its advances in capturing fine-gr

Point-Voxel Absorbing Graph Representation Learning for Event Stream based Recognition

Model ReleasesDGX agent

arXiv:2306.05239v3 Announce Type: replace Abstract: Sampled point and voxel methods are usually employed to downsample the dense events into sparse ones. After that, one popular way is to leverage a g

PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought

TutorialsDGX agent

arXiv:2606.24539v1 Announce Type: new Abstract: Pointing-based visual grounding requires models to precisely locate target objects by deciphering complex spatial relationships between the visual scene

Predicting brain tumour enhancement from non-contrast MR imaging with artificial intelligence: a multi-cohort retrospective diagnostic accuracy study

ResearchDGX agent

arXiv:2508.16650v3 Announce Type: replace-cross Abstract: Brain tumour MRI typically requires both pre- and post-contrast imaging, but gadolinium is not always desirable (frequent follow-up, renal imp

Progressive Pixel-Neighborhood Deformable Cross-Attention for Multispectral Object Detection

Local AiDGX agent

arXiv:2606.24092v1 Announce Type: new Abstract: Effective cross-modal feature alignment and interaction are central challenges in multispectral object detection. Although global cross-attention provid

Quantum CT via Dynamic Interval Encoding and Prior-Balanced QUBO Reconstruction

Local AiDGX agent

arXiv:2606.24561v1 Announce Type: new Abstract: Quadratic unconstrained binary optimization (QUBO)-based quantum computed tomography (CT) casts reconstruction as a binary quadratic problem for quantum

REALM: A Unified Red-Teaming Benchmark for Physical-World VLMs

Model ReleasesDGX agent

arXiv:2606.23892v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used as perception-reasoning backbones for embodied intelligence in safety-critical physical systems, whe

REDI-Match: Rotation-Equivariant Distillation for Efficient and Robust Dense Matching

Model ReleasesDGX agent

arXiv:2606.24330v1 Announce Type: new Abstract: Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing

Revealing Training Data Exposure in Vision Language Large Models via Parameter Gradients

Model ReleasesDGX agent

arXiv:2606.24774v1 Announce Type: new Abstract: Vision-Language Large Models (VLLMs) trained on massive crawled corpora raise pressing copyright and data-provenance concerns. These concerns are partic

S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing

ResearchDGX agent

arXiv:2606.24441v1 Announce Type: new Abstract: We present S1-Omni-Image, an open-weight unified multimodal model for scientific image understanding, generation, and editing. Unlike general-purpose im

Sat2City v2: Native 3D City Asset Generation from a Single Satellite Image

ApplicationsDGX agent

arXiv:2606.24138v1 Announce Type: new Abstract: Generating explicit 3D city assets from a single satellite image is important for digital twins, urban simulation, and geospatial intelligence. Unlike s

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation

SafetyDGX agent

arXiv:2606.18610v2 Announce Type: replace-cross Abstract: Evaluating generalist robot manipulation policies in the real world is expensive, slow, and difficult to scale. Action-conditioned video world

Segmentation and Classification of Pap Smear Images for Cervical Cancer Detection Using Deep Learning

ResearchDGX agent

arXiv:2508.17728v2 Announce Type: replace Abstract: Cervical cancer remains a significant global health concern and a leading cause of cancer-related deaths among women. Early detection through Pap sm

SENTRY: SAM2-Enhanced Neighbor-Aware and Temporally Reasoned Memory for Visual Tracking

Local AiDGX agent

arXiv:2606.24449v1 Announce Type: new Abstract: We revisit the memory update mechanism in SAM2-based visual object tracking and identify confidence-only mask selection as the dominant cause of drift u

SER: Learning to Ground Video Reasoning with Semantic Evidence Rewards

Model ReleasesDGX agent

arXiv:2606.24726v1 Announce Type: new Abstract: Video MLLMs often struggle with fine-grained spatio-temporal reasoning, sometimes generating correct answers based on irrelevant frames or objects. Alth

SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks

Model ReleasesDGX agent

arXiv:2606.24361v1 Announce Type: new Abstract: Sign language models are typically trained on datasets captured under constrained conditions, with limited viewpoint, background, and signer-identity di

Solving Semi-Supervised Few-Shot Learning from an Auto-Annotation Perspective

TutorialsDGX agent

arXiv:2512.10244v2 Announce Type: replace Abstract: Semi-supervised few-shot learning (SSFSL) resembles real-world applications such as auto-annotation, as it aims to learn a model from a few labeled

Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models

SafetyDGX agent

arXiv:2606.24165v1 Announce Type: new Abstract: Reducing visual token redundancy is critical for accelerating Multimodal Large Language Models (MLLMs) without degrading cross-modal reasoning performan

Spherical-to-ERP Epipolar Rectification for Single-Axis Disparity in 360 Stereo

ResearchDGX agent

arXiv:2606.24847v1 Announce Type: new Abstract: Omnidirectional stereo images provide full-surround perception but violate the geometric assumptions of classical disparity estimation: in spherical or

Systematic Exploration of 4-Expert Heterogeneous Mixture-of-Experts via Automated Pipeline Search

Model ReleasesDGX agent

arXiv:2606.23739v1 Announce Type: cross Abstract: We present an automated large-scale search pipeline for heterogeneous 4-Expert Mixture-of-Experts (MoE4) architectures within the LEMUR neural network

TIGER: Taming Identity, Geometry, and Generative Priors for High-Quality Face Video Restoration

Model ReleasesDGX agent

arXiv:2606.24336v1 Announce Type: new Abstract: Face Video Restoration (FVR) aims to recover high-fidelity facial videos from degraded input while preserving identity and semantic consistency across f

Token-to-Token Alignment of Text Embeddings for Semantic Blending

SafetyDGX agent

arXiv:2606.24021v1 Announce Type: new Abstract: In modern generative models, images are specified and controlled through text prompts. In practice, images are generated from sequences of tokens derive

Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling

Model ReleasesDGX agent

arXiv:2606.24187v1 Announce Type: new Abstract: Long video understanding remains a daunting challenge for Multimodal Large Language Models (MLLMs) due to the excessive computation and memory footprint

Training-free Cross-domain Few-shot Segmentation via Robust Semantic Representation and Matching

ResearchDGX agent

arXiv:2606.24297v1 Announce Type: new Abstract: Cross-domain Few-shot Segmentation (CD-FSS) aims to transfer knowledge learned from source domain to distinct target domains, segmenting unseen target c

Tri-Efficient Transfer Learning for Point Cloud Videos

Model ReleasesDGX agent

arXiv:2606.24175v1 Announce Type: new Abstract: While point cloud foundation models have significantly advanced point cloud video understanding, existing parameter-efficient fine-tuning (PEFT) methods

Trimming the Long-Tail of Visual World Modeling Evaluation

Model ReleasesDGX agent

arXiv:2606.24256v1 Announce Type: new Abstract: Physical interactions follow a long-tailed distribution: a set of common and regular interactions dominates human experience and visual data, while a br

TrOCR for Medieval HTR: A Systematic Ablation Study with Cross-Dataset Validation

Model ReleasesDGX agent

arXiv:2606.24302v1 Announce Type: new Abstract: Fine-tuning transformer-based handwritten text recognition (HTR) models on medieval manuscripts is challenging because these models are pre-trained on m

Trustworthy Image Authentication using Forensic Knowledge Graphs

ResearchDGX agent

arXiv:2606.23917v1 Announce Type: new Abstract: Advances in generative AI have made image falsification highly realistic, demanding trustworthy authentication systems. Existing forensic detectors can

TuringViT: Making SOTA Vision Transformers Accessible to All

ApplicationsDGX agent

arXiv:2606.24253v1 Announce Type: new Abstract: Modern VLMs and VLA systems commonly adopt off-the-shelf ViTs such as SigLIP2 as visual encoders, but diverse downstream requirements in latency, tempor

Understanding Deep Representation Learning via Layerwise Feature Compression and Discrimination

ResearchDGX agent

arXiv:2311.02960v5 Announce Type: replace-cross Abstract: Over the past decade, deep learning has proven to be a highly effective tool for learning meaningful features from raw data. However, it remai

UniRED: Unified RGB-D Video Frame Interpolation with Event Guidance

Model ReleasesDGX agent

arXiv:2606.24282v1 Announce Type: new Abstract: High frame-rate RGB-D videos are crucial for a variety of downstream tasks, including motion analysis, dynamic scene understanding, and 3D reconstructio

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation

SafetyDGX agent

arXiv:2606.24333v1 Announce Type: new Abstract: In-Image Machine Translation (IIMT) aims to translate scene text in an image and render the translated text back into the original regions while preserv

Universal Guideline-Driven Image Clustering via a Hybrid LLM Agent

AgentsDGX agent

arXiv:2606.24094v1 Announce Type: new Abstract: Unifying image clustering across different clustering scenarios remains challenging due to fundamental gaps among tasks. We introduce a Guideline-Driven

← Previous
1…7475767778…211
Next →