AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
4 Jun 2026

INTACT: Ego-Guided Typed Sparse Evidence Retrieval for Heterogeneous Collaborative Perception

AgentsDGX agent

arXiv:2606.04437v1 Announce Type: new Abstract: Collaborative perception extends the perceptual range of autonomous vehicles by sharing information across agents, but heterogeneous sensors and percept

Intra-Modal Neighbors Never Lie: Rectifying Inter-Modal Noisy Correspondence via Graph-Based Intra-Modal Reasoning

ResearchDGX agent

arXiv:2606.04061v1 Announce Type: new Abstract: Large-scale web-harvested datasets have fueled the progress of cross-modal retrieval but inevitably suffer from noisy correspondence, which severely deg

IRIS-GAN: Staged Specialist Detection of Deepfake Faces

ResearchDGX agent

arXiv:2606.04863v1 Announce Type: new Abstract: We introduce IRIS-GAN, a specialist forensic detector for synthetic face images under cross-generator shift. Rather than addressing universal synthetic-


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

J-RAS: Mutual Adaptation for Medical Image Segmentation via Contrastive Retrieval-Augmented Joint Optimization

TutorialsDGX agent

arXiv:2510.09953v3 Announce Type: replace Abstract: Manual medical image segmentation by clinicians, though accurate, is time-consuming and variable across experts, whereas AI-based models automate th

Label-Efficient 3D Forest Mapping: Self-Supervised and Transfer Learning for Instance Segmentation, Semantic Segmentation, and Species Classification

ResearchDGX agent

arXiv:2511.06331v2 Announce Type: replace Abstract: Detailed structural and species information on individual tree level is increasingly important to support precision forestry, biodiversity conservat

Learning Association via Track-Detection Matching for Multi-Object Tracking

TutorialsDGX agent

arXiv:2512.22105v2 Announce Type: replace Abstract: Multi-object tracking aims to maintain object identities over time by associating detections across video frames. Two dominant paradigms exist in li

MaCo-GAN: Manifold-Contrastive Adversarial Learning for Single Image Super-Resolution

ResearchDGX agent

arXiv:2606.05068v1 Announce Type: new Abstract: Conventional Generative Adversarial Networks (GANs) for Single Image Super-Resolution (SISR) often struggle with hallucinated artifacts, largely because

MAOAM: Unified Object and Material Selection with Vision-Language Models

ResearchDGX agent

arXiv:2606.04880v1 Announce Type: new Abstract: Selection is a core operation in interactive image editing. To be practical, a user should be able to specify and disambiguate the desired selection reg

MATCH: Multi-faceted Adaptive Topo-Consistency for Semi-Supervised Histopathology Segmentation

SafetyDGX agent

arXiv:2510.01532v2 Announce Type: replace Abstract: In semi-supervised segmentation, capturing meaningful semantic structures from unlabeled data is essential. This is particularly challenging in hist

Measuring Model Robustness via Fisher Information: Spectral Bounds, Theoretical Guarantees, and Practical Algorithms

SafetyDGX agent

arXiv:2606.04767v1 Announce Type: cross Abstract: The robustness of deep neural networks is crucial for safety-critical deployments, yet existing evaluation methods are often attack-dependent and lack

Med-Banana: Learning Quality-Controlled Medical Image Editing from Success-and-Failure Trajectories

ResearchDGX agent

arXiv:2511.00801v4 Announce Type: replace Abstract: Text-guided medical image editing must satisfy the requested pathology while preserving anatomy, modality-specific appearance, and clinical plausibi

MeshFlow: Efficient Artistic Mesh Generation via MeshVAE and Flow-based Diffusion Transformer

ResearchDGX agent

arXiv:2606.04621v1 Announce Type: new Abstract: We present MeshFlow, a new method for generating artist-like 3D meshes. Current mesh generators often adopt Auto-Regressive (AR) next-token prediction,

MeshWeaver: Sparse-Voxel-Guided Surface Weaving for Autoregressive Mesh Generation

Local AiDGX agent

arXiv:2606.04688v1 Announce Type: new Abstract: Autoregressive mesh generation has gained attention by tokenizing meshes into sequences and training models in a language-modeling fashion. However, exi

MetaPoint: Unlocking Precise Spatial Control in Agentic Visual Generation

AgentsDGX agent

arXiv:2606.05031v1 Announce Type: new Abstract: Generative visual models fundamentally struggle with precise spatial control. This arises from a core disconnect: models can process textual description

Motion-Guided Causal Disentanglement for Robust Multi-View Cine Cardiac MRI Diagnosis

TutorialsDGX agent

arXiv:2606.04414v1 Announce Type: new Abstract: Multi-view cardiac magnetic resonance (CMR) imaging provides complementary anatomical information and is widely used for noninvasive disease assessment.

Multi-Camera AR Guidance System for Surgical Instrument Handling and Assembly: Investigating Workload and Efficiency

ApplicationsDGX agent

arXiv:2606.04992v1 Announce Type: new Abstract: The handling and assembly of instruments during surgery imposes high cognitive demands on scrub nurses, particularly when instruments are unfamiliar. We

Optimal Transport Flow Matching by Design

SafetyDGX agent

arXiv:2606.04092v1 Announce Type: new Abstract: Flow matching models learn to transport samples from a simple prior distribution to a complex data distribution. When prior-data pairs are coupled via o

Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment

Model ReleasesDGX agent

arXiv:2606.04737v1 Announce Type: new Abstract: Large-scale video generation models have made remarkable progress in semantic consistency and visual quality, producing videos that are increasingly coh

Pinpoint: Grounded Worldwide Image Geolocation via Cross-Source Retrieval and Reranking

ResearchDGX agent

arXiv:2606.04133v1 Announce Type: new Abstract: Image geolocation aims to estimate where a photograph was taken from its visual content. At worldwide scale, this remains challenging because visual evi

Plug-and-Play Diffusion Meets ADMM: Dual-Variable Coupling for Robust Medical Image Reconstruction

SafetyDGX agent

arXiv:2602.23214v2 Announce Type: replace Abstract: Plug-and-Play diffusion prior (PnPDP) frameworks have emerged as a powerful paradigm for solving imaging inverse problems by treating pretrained gen

Prospective Dynamic 3D MRI Reconstruction via Latent-Space Motion Tracking from Single Measurement

TutorialsDGX agent

arXiv:2606.04249v1 Announce Type: new Abstract: Prospective reconstruction is crucial in many clinical applications such as MRI-guided radiotherapy, which demands accurate image reconstruction and fas

PureLight: Learning Complex Luminaires with Light Tracing

ResearchDGX agent

arXiv:2606.04319v1 Announce Type: cross Abstract: We propose a neural formulation for estimating the appearance of complex luminaires. We focus on challenging luminaires with complex light transport (

Radiomic Feature Selection Using Gradient Loss of Deep Neural Network for Lung Cancer Stage Detection

ResearchDGX agent

arXiv:2606.04453v1 Announce Type: new Abstract: Radiomics enables extraction of quantitative imaging biomarkers from medical images and has become an important tool for computer-aided cancer diagnosis

Recent Advances and Trends in Learning-based 3D Representations

AgentsDGX agent

arXiv:2606.04871v1 Announce Type: new Abstract: The selection of an appropriate 3D representation is a fundamental design decision that dictates the efficiency, quality, and capabilities of modern com

ReConFuse: Reconstruction-Error Guided Semantic Fusion for AI-Generated Video Detection

ResearchDGX agent

arXiv:2606.04706v1 Announce Type: new Abstract: AI-generated videos are becoming increasingly realistic, raising serious concerns about misinformation, content authenticity, and media trust. Reliable

Reflection Separation from a Single Image via Joint Latent Diffusion

ApplicationsDGX agent

arXiv:2606.04107v1 Announce Type: new Abstract: Single-image reflection separation is highly challenging under extreme conditions like glare or weak reflections. Existing methods often struggle to rec

Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models

SafetyDGX agent

arXiv:2502.01576v2 Announce Type: replace Abstract: Multi-modal Large Language Models (MLLMs) excel in vision-language tasks but remain vulnerable to visual adversarial perturbations that can induce h

Robust Multi-view Clustering against Imperfect Information

Model ReleasesDGX agent

arXiv:2606.04343v1 Announce Type: new Abstract: Real-world multi-view data always suffer from imperfect information problem, where the view-specific observations are absent (i.e., Incomplete Views, IV

SBP-Net: Learning Thin Structure Reconstruction with Sliding-Box Projections

Local AiDGX agent

arXiv:2606.04251v1 Announce Type: new Abstract: Reconstructing thin 3D structures is challenging due to their sparsity, scale variation, and complex geometry. Such structures arise in a wide range of

Scalable Event Cloud Network for Event-based Classification

ResearchDGX agent

arXiv:2412.20803v2 Announce Type: replace Abstract: Event cameras are biologically inspired sensors garnering significant attention from both industry and academia. Mainstream methods favor frame and

Scene-Centric Unsupervised Video Panoptic Segmentation

Model ReleasesDGX agent

arXiv:2606.04925v1 Announce Type: new Abstract: Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regio

SharpNet: Enhancing MLPs to Represent Functions with Controlled Non-differentiability

ResearchDGX agent

arXiv:2601.19683v2 Announce Type: replace Abstract: Multi-layer perceptrons (MLPs) are a standard tool for learning and function approximation, but they inherently produce globally smooth outputs. Con

Shifting the Breaking Point of Flow Matching for Multi-Instance Editing

Model ReleasesDGX agent

arXiv:2602.08749v3 Announce Type: replace Abstract: Flow matching models have recently emerged as an efficient alternative to diffusion, especially for text-guided image generation and editing, offeri

Spatial Artifact Coherence Determines Codec Robustness in Patch-Based rPPG

ResearchDGX agent

arXiv:2606.04198v1 Announce Type: new Abstract: Remote photoplethysmography (rPPG) achieves low heart-rate error on uncompressed benchmarks yet is deployed over compressed video channels in telehealth

Spatially Grounded Concept Bottleneck Models via Part-Factorized Attention

ResearchDGX agent

arXiv:2606.04364v1 Announce Type: new Abstract: Concept bottleneck models (CBMs) predict a layer of human-named attributes before predicting a class, which makes their decisions auditable. On fine-gra

StrokeTimer: Robust Representation Learning for Ischemic Stroke Onset-Time Estimation from Non-contrast CT

ApplicationsDGX agent

arXiv:2606.04722v1 Announce Type: new Abstract: Ischemic stroke is a major global disease. Treatment decisions are highly time-sensitive, as eligibility for reperfusion therapies relies on the interva

Take a Peek: Efficient Encoder Adaptation for Few-Shot Semantic Segmentation via LoRA

ResearchDGX agent

arXiv:2512.10521v2 Announce Type: replace Abstract: Few-shot semantic segmentation (FSS) aims to segment novel classes in query images using only a small annotated support set. While prior research ha

TGSD: Topology-Guided State-Space Diffusion for EEG Spatial Super-Resolution

TutorialsDGX agent

arXiv:2606.03998v1 Announce Type: cross Abstract: Low-density EEG is more suitable for wearable and IoT-based brain sensing, but sparse electrode sampling often lacks sufficient spatial information to

Toward Multi-Domain and Long-Tailed Quantization via Feature Alignment and Scaling

SafetyDGX agent

arXiv:2606.04920v1 Announce Type: cross Abstract: Quantizing deep neural networks is essential for efficient inference on resource-constrained devices. However, most existing methods are designed for

Toward Trustworthy Portrait Editing: Evaluation of Demographic Misrepresentation in I2I Models

Model ReleasesDGX agent

arXiv:2602.16149v2 Announce Type: replace Abstract: Instruction-guided image-to-image (I2I) editors are increasingly used in consumer and professional visual workflows, where trustworthiness depends n

Towards Evaluating the Robustness of Visual State Space Models

ApplicationsDGX agent

arXiv:2406.09407v3 Announce Type: replace Abstract: Vision State Space Models (VSSMs), a novel architecture that combines the strengths of recurrent neural networks and latent variable models, have de

Transferable Multi-Bit Watermarking Across Frozen Diffusion Models via Latent Consistency Bridges

SafetyDGX agent

arXiv:2603.20304v2 Announce Type: replace Abstract: As generative AI advances, global governance frameworks increasingly mandate verifiable content provenance. However, existing watermarking technique

Ultra-Fast Neural Video Compression

ApplicationsDGX agent

arXiv:2606.04410v1 Announce Type: new Abstract: While neural video codecs (NVCs) have demonstrated superior compression ratio, their prohibitive computational complexity remains a critical barrier to

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation

ResearchDGX agent

arXiv:2606.04264v1 Announce Type: new Abstract: Recent years have seen remarkable progress in unified vision-language models handling both multimodal understanding and generation within a single archi

ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models

ResearchDGX agent

arXiv:2512.14099v3 Announce Type: replace Abstract: Motivated by discrete diffusion's success in language-vision modeling, we explore its potential for multi-view generation, a task dominated by conti

Vision Transformer Finetuning Benefits from Non-Smooth Components

ResearchDGX agent

arXiv:2602.06883v3 Announce Type: replace-cross Abstract: The smoothness of the transformer architecture has been extensively studied in the context of generalization, training stability, and adversar

VT-3DAD: Cross-Category 3D Anomaly Detection via Visual-Text Normal Space Alignment

SafetyDGX agent

arXiv:2606.04369v1 Announce Type: new Abstract: Few-shot cross-category 3D anomaly detection aims to determine whether an unknown point cloud belongs to a target normal category using only a few norma

Weakly Supervised Incremental Segmentation via Semantic Anchors and Spatial Arbitration

ResearchDGX agent

arXiv:2606.04060v1 Announce Type: new Abstract: Weakly Incremental Learning for Semantic Segmentation (WILSS) suffers from the continuous introduction of noisy supervision, which progressively corrupt

When Detectors Forget Forensics: Blocking Semantic Shortcuts for Generalizable AI-Generated Image Detection

ResearchDGX agent

arXiv:2603.09242v2 Announce Type: replace Abstract: The growing realism of generative models has blurred the boundary between real and synthetic content, posing significant challenges to reliable AI-g

When Seeing Is Not Believing -- A Benchmark for Search-Grounded Video Misinformation Detection

Model ReleasesDGX agent

arXiv:2606.04098v1 Announce Type: new Abstract: Video misinformation increasingly operates at the semantic and evidential level: authentic footage may be selectively edited, temporally reordered, spli

XSSR: Cross-Domain Self-Supervised Representative Selection for Efficient Annotation in Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2606.04301v1 Announce Type: new Abstract: Acquiring labeled medical image data is resource-intensive and a challenge further exacerbated in cross-domain scenarios where source and target dataset

Z-FLoc: Zero-Shot Floorplan Localization via Geometric Primitives

ApplicationsDGX agent

arXiv:2606.04788v1 Announce Type: new Abstract: Visual localization -- estimating a camera pose within a pre-existing map -- is a fundamental problem in computer vision. Floorplans are an attractive m

ZipSplat: Fewer Gaussians, Better Splats

ResearchDGX agent

arXiv:2606.05102v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting methods reconstruct a scene from posed or pose-free images in a single forward pass, yet current approaches predict o

3 Jun 2026

A Benchmark for Semi-supervised Multi-modal Crowd Counting

Model ReleasesDGX agent

arXiv:2606.03646v1 Announce Type: new Abstract: This paper constructs the first benchmark on semi-supervised multi-modal crowd counting. To lay the foundation for this unexplored task, we first formul

A Fast Methane Detection Pipeline on Board Satellites Based on Mag1c-SAS and LinkNet

Model ReleasesDGX agent

arXiv:2606.03675v1 Announce Type: new Abstract: Methane is a potent greenhouse gas, and detecting leaks early via hyperspectral satellite imagery can help climate change mitigation efforts. Meanwhile,

A unified multi-task framework enables interpretable chest radiograph analysis

Local AiDGX agent

arXiv:2606.03417v1 Announce Type: new Abstract: While multimodal deep learning has advanced medical imaging analysis, existing black-box systems extcolor{black}{may remain confined to isolated tasks,

A^2: Smaller Self-Supervised ViTs Localize Better than Larger Ones

Local AiDGX agent

arXiv:2606.03148v1 Announce Type: new Abstract: Robust visual classification often depends on localizing the main foreground objects in an image while ignoring contextual distractors. Surprisingly, we

AAD-1: Asymmetric Adversarial Distillation for One-Step Autoregressive Video Generation

ResearchDGX agent

arXiv:2606.03972v1 Announce Type: new Abstract: We present AAD-1, an Asymmetric Adversarial Distillation framework for One-step autoregressive image-to-video generation. State-of-the-art methods adopt

Act Like a Pathologist: Tissue-Aware Whole Slide Image Reasoning

ResearchDGX agent

arXiv:2603.00667v3 Announce Type: replace Abstract: Computational pathology has advanced rapidly in recent years, driven by domain-specific image encoders and growing interest in using vision-language

Adaptive Causal Alignment for High-Confidence Adversarial Training

SafetyDGX agent

arXiv:2606.03925v1 Announce Type: new Abstract: Inverse adversarial training leverages high-confidence predictions to stabilize robust learning, yet we uncover a critical paradox: high confidence ofte

← Previous
1…96979899100…211
Next →