AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Applications

BEV-Denoise: Learning Intrinsic Noise for Accurate Bird's-Eye-View Semantic Segmentation

DGX agent

arXiv:2606.22931v1 Announce Type: new Abstract: In this paper, we present a framework dubbed extbf{BEV-Denoise} that estimates and removes intrinsic noise from learned Bird's-Eye-View (BEV) features t

applicationsarxiv-cs-cv
23 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Applications

Beyond a Single Light: A Large-Scale Aerial Dataset for Urban Scene Reconstruction Under Varying Illumination

DGX agent

arXiv:2512.14200v2 Announce Type: replace Abstract: Recent advances in Neural Radiance Fields and 3D Gaussian Splatting have demonstrated strong potential for large-scale UAV-based 3D reconstruction t

applicationsarxiv-cs-cv
23 Jun 2026
Research

Beyond Damage Assessment: Recyclable Material Detection in Aerial Disaster Imagery Using a Lightweight Patch-Based Framework

DGX agent

arXiv:2606.21279v1 Announce Type: new Abstract: Nowadays, more and more disasters of different natures are appearing. Several disaster assessment approaches have been developed in order to identify da

researcharxiv-cs-cv
23 Jun 2026
Research

Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification

DGX agent

arXiv:2606.21838v1 Announce Type: new Abstract: Multimodal contrastive learning has enabled zero-shot visual classification by aligning images with textual categories. However, in hierarchically struc

researcharxiv-cs-cv
23 Jun 2026
Research

Beyond ROC-AUC: Operating-Point Performance Reporting for Biometric Verification

DGX agent

arXiv:2606.20680v1 Announce Type: new Abstract: A biometric verifier is often deployed with a strict false match budget, so only a narrow, low false match rate (FMR) slice of the score range is used.

researcharxiv-cs-cv
23 Jun 2026
Tutorials

Beyond Templates: Revisiting Zero-Shot Remote Sensing through Meta-Prompting

DGX agent

arXiv:2606.20702v1 Announce Type: new Abstract: Vision-language models (VLMs) have sparked growing interest in zero-shot Earth Observation (EO) downstream tasks, with further gains enabled by remote-s

tutorialsarxiv-cs-cv
23 Jun 2026
Model Releases

Beyond the LUMIR challenge: The pathway to foundational registration models

DGX agent

arXiv:2505.24160v3 Announce Type: replace-cross Abstract: Medical image challenges have played a transformative role in advancing the field, catalyzing innovation and establishing new performance benc

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation

DGX agent

arXiv:2511.22973v2 Announce Type: replace Abstract: Long video generation is a critical step toward building realistic world models, requiring both high visual fidelity and long-range interaction cons

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Biological Sex Determination in Cadavers Using Deep Learning Algorithms from Computed Tomography Images of Pelvis and Skull

DGX agent

arXiv:2606.22515v1 Announce Type: new Abstract: Sexual identification of decomposed cadavers challenges traditional methods dependent on visual anthropological analysis. This study evaluates state-of-

researcharxiv-cs-cv
23 Jun 2026
Model Releases

Black-Box Continual Learning for Vision-Language Models

DGX agent

arXiv:2606.22999v1 Announce Type: new Abstract: The rapid deployment of Vision-Language Models (VLMs) in dynamic environments necessitates the ability to learn continuously without forgetting. However

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Boosting Neural Video Codec via Scale-Driven Online Flow Refinement

DGX agent

arXiv:2606.23023v1 Announce Type: new Abstract: Although state-of-the-art neural video codecs (NVCs) have achieved remarkable performance, they suffer from limited generalization when encountering com

researcharxiv-cs-cv
23 Jun 2026
Research

Boundary-by-Mask: Few-Shot Instance Segmentation with Mask-Conditioned Boundary Learning for Texture-Poor Industrial Parts

DGX agent

arXiv:2606.21594v1 Announce Type: new Abstract: Recent advances in large pre-trained models have led to remarkable progress in instance segmentation on general images. However, industrial scenarios re

researcharxiv-cs-cv
23 Jun 2026
Tutorials

BoxCtrl: 3D-Aware Visual Prompting for Geometric Image Editing

DGX agent

arXiv:2606.23270v1 Announce Type: new Abstract: As instruction-based editing models and multimodal large language models advance, diverse image editing tasks have become feasible. However, achieving p

tutorialsarxiv-cs-cv
23 Jun 2026
Research

Brain-Adapter: A Dual-Stream Vision-Language MIL Framework for Comprehensive 3D CT Diagnosis of Acute Intracranial Pathologies

DGX agent

arXiv:2606.23494v1 Announce Type: new Abstract: Automated diagnosis of 3D brain CT scans is essential for critical care, yet it remains challenging due to the heavy reliance on manual annotations and

researcharxiv-cs-cv
23 Jun 2026
Tutorials

Brain-Inspired Stochastic Joint Embedding Representation Learning

DGX agent

arXiv:2505.11129v2 Announce Type: replace Abstract: Representation learning is one of the key research topics in machine learning, and the framework of self-supervised learning (SSL) has revolutionize

tutorialsarxiv-cs-cv
23 Jun 2026
Applications

Bridging Single Distortion Artifacts and Multifactorial Clinical Quality: Few-shot Biparametric MRI Quality Assessment via Distortion-trained Prototypical Networks

DGX agent

arXiv:2606.18872v2 Announce Type: replace Abstract: Clinical prostate multi-parametric MRI relies heavily on high-quality diffusion-weighted imaging (DWI), yet reading DWI is frequently compromised by

applicationsarxiv-cs-cv
23 Jun 2026
Local Ai

Build Once, Monitor Continuously: Persistent Semantic Mapping via Autonomous Exploration and Open-Vocabulary Object Updates

DGX agent

arXiv:2409.15493v4 Announce Type: replace-cross Abstract: Persistent semantic monitoring of indoor spaces such as warehouses, hospitals, and offices requires a robot to repeatedly monitor an environme

local-aiarxiv-cs-cv
23 Jun 2026
Research

C^2GR: Coupled Comprehensive Generative Replay for a Continually Learnable Universal Segmentation Model

DGX agent

arXiv:2606.23473v1 Announce Type: new Abstract: Universal segmentation models exhibit significant potential for diverse tasks involving different imaging modalities and segmentation objectives. Task-I

researcharxiv-cs-cv
23 Jun 2026
Research

Can Single-View Mesh Reconstruction Generalize to Robot Camera Rotation?

DGX agent

arXiv:2606.22987v1 Announce Type: new Abstract: Single-view mesh reconstruction predicts object meshes and spatial layouts from a single observation, making it attractive for fast robot spatial reason

researcharxiv-cs-cv
23 Jun 2026
Model Releases

CAOA -- Completion-Assisted Object-CAD Alignment

DGX agent

arXiv:2606.18429v2 Announce Type: replace Abstract: Accurately aligning CAD models to their corresponding objects in indoor RGB-D scans is a central challenge in 3D semantic reconstruction. The task r

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales

DGX agent

arXiv:2606.21949v1 Announce Type: new Abstract: Accurate and comprehensive video captions with consistent subject references are critical for downstream understanding and generation tasks. However, fe

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Catching Lies Without Sending the Video: Privacy-Preserving Multimodal Deception Detection

DGX agent

arXiv:2606.22699v1 Announce Type: new Abstract: Frontier multimodal models can guess whether a person is lying from a testimony video. To do so, they stream that raw face and voice to a third-party mo

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Category-Adaptive Cross-Modal Semantic Refinement and Transfer for Open-Vocabulary Multi-Label Recognition

DGX agent

arXiv:2412.06190v2 Announce Type: replace Abstract: Benefiting from the generalization capability of CLIP, recent vision language pre-training (VLP) models have demonstrated the ability to capture a w

researcharxiv-cs-cv
23 Jun 2026
Model Releases

CDER-SME: A Cross-Device Event-RGB Micro-Expression Dataset under Multi-Level Stress Induction

DGX agent

arXiv:2606.20715v1 Announce Type: new Abstract: Micro-expression recognition (MER) in realistic scenarios demands high temporal sensitivity and ecological validity, yet existing benchmarks are largely

model-releasesarxiv-cs-cv
23 Jun 2026
Safety

CFPO: Counterfactual Policy Optimization for Multimodal Reasoning

DGX agent

arXiv:2606.23206v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in multimodal reasoning. However, prevailing reinforcement learning (RL)

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

Chains That See, Answers That Don't: A Multi-Aspect Evaluation Recipe for Forced Chain-of-Thought on Video-MME

DGX agent

arXiv:2606.22862v1 Announce Type: new Abstract: Forced chain-of-thought (CoT) is widely assumed to make vision-language models more reliable on video question answering. We propose a small three-probe

model-releasesarxiv-cs-cv
23 Jun 2026
Applications

Changing Modalities: Adapting Remote Sensing Models to New Satellites and Sensors

DGX agent

arXiv:2606.23356v1 Announce Type: new Abstract: Machine learning models for remote sensing are trained and deployed on a static set of modalities. However, as we equip newer satellites with novel sens

applicationsarxiv-cs-cv
23 Jun 2026
Model Releases

Chehre: An Emoji-Prompted Video Dataset for Perceptually Diverse Facial Expression Recognition

DGX agent

arXiv:2606.21657v1 Announce Type: new Abstract: Facial expressions are nonverbal social signals used in human interaction, but facial expression recognition datasets often focus on static images, basi

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

CheXpercept: A Benchmark for Evaluating Expert-Level Lesion Perception in Chest X-rays

DGX agent

arXiv:2606.21020v1 Announce Type: new Abstract: The evaluation of vision-language models (VLMs) for chest X-ray (CXR) analysis has largely been limited to disease-presence classification without visua

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

ChronoLock: Protecting Videos from Unauthorized Text-to-Video Personalization

DGX agent

arXiv:2606.21146v1 Announce Type: new Abstract: Text-to-video (T2V) diffusion models have made it increasingly easy to synthesize realistic and temporally coherent videos, while recent personalization

model-releasesarxiv-cs-cv
23 Jun 2026
Safety

CLAR: Learning 3D Representations for Robotic Manipulation by Fusing Masked Reconstruction with Multi-Level Contrastive Alignment

DGX agent

arXiv:2507.08262v2 Announce Type: replace-cross Abstract: The spatial information inherent in 3D point clouds is crucial for robotic manipulation. However, existing 3D pre-training methods face a fund

safetyarxiv-cs-cv
23 Jun 2026
Research

CLoE: Expert Consistency Learning for Robust Missing Modality Segmentation

DGX agent

arXiv:2603.09316v2 Announce Type: replace Abstract: Multimodal medical image segmentation often faces missing modalities at inference, which induces disagreement among modality experts and makes fusio

researcharxiv-cs-cv
23 Jun 2026
Local Ai

CMDS-AD: Cross-Modal Dual-Stream Decoupling for Few-Shot Anomaly Detection

DGX agent

arXiv:2606.20300v2 Announce Type: replace Abstract: Few-shot anomaly detection remains challenging due to limited training data. Multi-modal anomaly detection (MAD) offers a viable solution, leveragin

local-aiarxiv-cs-cv
23 Jun 2026
Model Releases

CodePercept: Code-Grounded Visual STEM Perception for MLLMs

DGX agent

arXiv:2603.10757v2 Announce Type: replace Abstract: When MLLMs fail at Science, Technology, Engineering, and Mathematics (STEM) visual reasoning, a fundamental question arises: is it due to perceptual

model-releasesarxiv-cs-cv
23 Jun 2026
Local Ai

CoDMD: Copula-aware Distribution Matching Distillation for Fast Video Generation

DGX agent

arXiv:2606.21982v1 Announce Type: new Abstract: Few-step distillation for video diffusion models has attracted significant attention, driven by the urgent demand for efficient deployment in real-world

local-aiarxiv-cs-cv
23 Jun 2026
Research

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models

DGX agent

arXiv:2606.20970v1 Announce Type: new Abstract: Omni-modal models can ingest video, audio, and text, but unified access to multiple modalities does not guarantee that a model uses the right evidence.

researcharxiv-cs-cv
23 Jun 2026
Agents

Compressing Observation History into Agent Memory: Distilling Transformers into Recurrent Transformers

DGX agent

arXiv:2606.21562v1 Announce Type: new Abstract: Transformers are AI's workhorse with strong performance in modeling sequential data, but their computational cost becomes prohibitive when processing lo

agentsarxiv-cs-cv
23 Jun 2026
Research

Compression and Retrieval: Implicit Memory Retrieval for Video World Models

DGX agent

arXiv:2606.23105v1 Announce Type: new Abstract: Video world models hold promise for simulating interactive environments, yet maintaining consistent long-term memory across complex camera trajectories

researcharxiv-cs-cv
23 Jun 2026
Safety

Concept Alignment Contrast and Long-Short Prompt Memory for Test-Time Adaptation of SAM3 in Medical Image Segmentation

DGX agent

arXiv:2606.22963v1 Announce Type: new Abstract: Concept segmentation models like Segment Anything Model 3 (SAM3) show strong generalization on natural images, yet their performance degrades in medical

safetyarxiv-cs-cv
23 Jun 2026
Safety

Confidence-Uncertainty Boundary Calibration for Bayesian Deep Learning in Medical Image Analysis

DGX agent

arXiv:2602.11973v2 Announce Type: replace Abstract: In critical decision support systems based on medical imaging, the reliability of AI-assisted decision-making is as relevant as predictive accuracy.

safetyarxiv-cs-cv
23 Jun 2026
Local Ai

Configurable Algorithms for Histopathologic Cancer Detection on Quantum Hardware

DGX agent

arXiv:2606.21752v1 Announce Type: cross Abstract: Histopathologic cancer detection is challenging due to tissue variability, staining differences, and subtle visual distinctions between disease classe

local-aiarxiv-cs-cv
23 Jun 2026
Model Releases

ConnectomeBench2: A Unified Benchmark for Automated Connectomic Proofreading

DGX agent

arXiv:2606.21116v1 Announce Type: new Abstract: Proofreading--correcting segmentation errors in 3D brain reconstructions--is the rate-limiting step in synapse-resolution connectomics. We release Conne

model-releasesarxiv-cs-cv
23 Jun 2026
Applications

Context-Aware Autoregressive Diffusion for Gloss-Wise Sign Language Production

DGX agent

arXiv:2606.21234v1 Announce Type: new Abstract: To generate natural and accurate sentence-level sign language, synthesizing the 'gloss', the fundamental semantic unit, is essential. However, most curr

applicationsarxiv-cs-cv
23 Jun 2026
Tutorials

Contrastive and Adaptive Multi-modal Masked Autoencoder for Spatial Transcriptomics

DGX agent

arXiv:2606.21156v1 Announce Type: new Abstract: The high cost of spatial transcriptomics (ST) has driven extensive studies into predicting gene expression directly from H&E histology images. However,

tutorialsarxiv-cs-cv
23 Jun 2026
Research

Controllable Texture Tiling with Transformed RoPE-Enhanced Diffusion Models

DGX agent

arXiv:2606.22945v1 Announce Type: cross Abstract: Realistic integration of user-specified textures into scene images is a fundamental task in computer graphics and image editing. While existing materi

researcharxiv-cs-cv
23 Jun 2026
Model Releases

CoSA: Correlation-Guided Change Attention with Learnable Residual Gating for Remote Sensing Change Detection

DGX agent

arXiv:2606.21932v1 Announce Type: new Abstract: Remote sensing change detection (CD) from bi-temporal imagery is critical for applications such as urban monitoring, disaster assessment, and environmen

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

CourseBlueprint: A Structured Pipeline for Adaptive Pedagogical Video Generation Grounded in Course Corpora

DGX agent

arXiv:2606.20608v1 Announce Type: cross Abstract: Generative text-to-video systems can produce visually fluent educational clips, but they rarely encode the pedagogical content knowledge (PCK) needed

model-releasesarxiv-cs-cv
23 Jun 2026
Local Ai

CoVStream: Edge-Cloud Collaboration for Understanding of Long Video Streams

DGX agent

arXiv:2606.22804v1 Announce Type: new Abstract: Long, continuous video streams are an increasingly critical driver of multimedia intelligence. Existing efforts often handle long videos with a sample-e

local-aiarxiv-cs-cv
23 Jun 2026
← Previous
1…9596979899…263
Next →