AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
23 Jun 2026

BoxCtrl: 3D-Aware Visual Prompting for Geometric Image Editing

TutorialsDGX agent

arXiv:2606.23270v1 Announce Type: new Abstract: As instruction-based editing models and multimodal large language models advance, diverse image editing tasks have become feasible. However, achieving p

Brain-Adapter: A Dual-Stream Vision-Language MIL Framework for Comprehensive 3D CT Diagnosis of Acute Intracranial Pathologies

ResearchDGX agent

arXiv:2606.23494v1 Announce Type: new Abstract: Automated diagnosis of 3D brain CT scans is essential for critical care, yet it remains challenging due to the heavy reliance on manual annotations and

Brain-Inspired Stochastic Joint Embedding Representation Learning

TutorialsDGX agent

arXiv:2505.11129v2 Announce Type: replace Abstract: Representation learning is one of the key research topics in machine learning, and the framework of self-supervised learning (SSL) has revolutionize


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Bridging Single Distortion Artifacts and Multifactorial Clinical Quality: Few-shot Biparametric MRI Quality Assessment via Distortion-trained Prototypical Networks

ApplicationsDGX agent

arXiv:2606.18872v2 Announce Type: replace Abstract: Clinical prostate multi-parametric MRI relies heavily on high-quality diffusion-weighted imaging (DWI), yet reading DWI is frequently compromised by

Build Once, Monitor Continuously: Persistent Semantic Mapping via Autonomous Exploration and Open-Vocabulary Object Updates

Local AiDGX agent

arXiv:2409.15493v4 Announce Type: replace-cross Abstract: Persistent semantic monitoring of indoor spaces such as warehouses, hospitals, and offices requires a robot to repeatedly monitor an environme

C^2GR: Coupled Comprehensive Generative Replay for a Continually Learnable Universal Segmentation Model

ResearchDGX agent

arXiv:2606.23473v1 Announce Type: new Abstract: Universal segmentation models exhibit significant potential for diverse tasks involving different imaging modalities and segmentation objectives. Task-I

Can Single-View Mesh Reconstruction Generalize to Robot Camera Rotation?

ResearchDGX agent

arXiv:2606.22987v1 Announce Type: new Abstract: Single-view mesh reconstruction predicts object meshes and spatial layouts from a single observation, making it attractive for fast robot spatial reason

CAOA -- Completion-Assisted Object-CAD Alignment

Model ReleasesDGX agent

arXiv:2606.18429v2 Announce Type: replace Abstract: Accurately aligning CAD models to their corresponding objects in indoor RGB-D scans is a central challenge in 3D semantic reconstruction. The task r

CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales

Model ReleasesDGX agent

arXiv:2606.21949v1 Announce Type: new Abstract: Accurate and comprehensive video captions with consistent subject references are critical for downstream understanding and generation tasks. However, fe

Catching Lies Without Sending the Video: Privacy-Preserving Multimodal Deception Detection

Model ReleasesDGX agent

arXiv:2606.22699v1 Announce Type: new Abstract: Frontier multimodal models can guess whether a person is lying from a testimony video. To do so, they stream that raw face and voice to a third-party mo

Category-Adaptive Cross-Modal Semantic Refinement and Transfer for Open-Vocabulary Multi-Label Recognition

ResearchDGX agent

arXiv:2412.06190v2 Announce Type: replace Abstract: Benefiting from the generalization capability of CLIP, recent vision language pre-training (VLP) models have demonstrated the ability to capture a w

CDER-SME: A Cross-Device Event-RGB Micro-Expression Dataset under Multi-Level Stress Induction

Model ReleasesDGX agent

arXiv:2606.20715v1 Announce Type: new Abstract: Micro-expression recognition (MER) in realistic scenarios demands high temporal sensitivity and ecological validity, yet existing benchmarks are largely

CFPO: Counterfactual Policy Optimization for Multimodal Reasoning

SafetyDGX agent

arXiv:2606.23206v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in multimodal reasoning. However, prevailing reinforcement learning (RL)

Chains That See, Answers That Don't: A Multi-Aspect Evaluation Recipe for Forced Chain-of-Thought on Video-MME

Model ReleasesDGX agent

arXiv:2606.22862v1 Announce Type: new Abstract: Forced chain-of-thought (CoT) is widely assumed to make vision-language models more reliable on video question answering. We propose a small three-probe

Changing Modalities: Adapting Remote Sensing Models to New Satellites and Sensors

ApplicationsDGX agent

arXiv:2606.23356v1 Announce Type: new Abstract: Machine learning models for remote sensing are trained and deployed on a static set of modalities. However, as we equip newer satellites with novel sens

Chehre: An Emoji-Prompted Video Dataset for Perceptually Diverse Facial Expression Recognition

Model ReleasesDGX agent

arXiv:2606.21657v1 Announce Type: new Abstract: Facial expressions are nonverbal social signals used in human interaction, but facial expression recognition datasets often focus on static images, basi

CheXpercept: A Benchmark for Evaluating Expert-Level Lesion Perception in Chest X-rays

Model ReleasesDGX agent

arXiv:2606.21020v1 Announce Type: new Abstract: The evaluation of vision-language models (VLMs) for chest X-ray (CXR) analysis has largely been limited to disease-presence classification without visua

ChronoLock: Protecting Videos from Unauthorized Text-to-Video Personalization

Model ReleasesDGX agent

arXiv:2606.21146v1 Announce Type: new Abstract: Text-to-video (T2V) diffusion models have made it increasingly easy to synthesize realistic and temporally coherent videos, while recent personalization

CLAR: Learning 3D Representations for Robotic Manipulation by Fusing Masked Reconstruction with Multi-Level Contrastive Alignment

SafetyDGX agent

arXiv:2507.08262v2 Announce Type: replace-cross Abstract: The spatial information inherent in 3D point clouds is crucial for robotic manipulation. However, existing 3D pre-training methods face a fund

CLoE: Expert Consistency Learning for Robust Missing Modality Segmentation

ResearchDGX agent

arXiv:2603.09316v2 Announce Type: replace Abstract: Multimodal medical image segmentation often faces missing modalities at inference, which induces disagreement among modality experts and makes fusio

CMDS-AD: Cross-Modal Dual-Stream Decoupling for Few-Shot Anomaly Detection

Local AiDGX agent

arXiv:2606.20300v2 Announce Type: replace Abstract: Few-shot anomaly detection remains challenging due to limited training data. Multi-modal anomaly detection (MAD) offers a viable solution, leveragin

CodePercept: Code-Grounded Visual STEM Perception for MLLMs

Model ReleasesDGX agent

arXiv:2603.10757v2 Announce Type: replace Abstract: When MLLMs fail at Science, Technology, Engineering, and Mathematics (STEM) visual reasoning, a fundamental question arises: is it due to perceptual

CoDMD: Copula-aware Distribution Matching Distillation for Fast Video Generation

Local AiDGX agent

arXiv:2606.21982v1 Announce Type: new Abstract: Few-step distillation for video diffusion models has attracted significant attention, driven by the urgent demand for efficient deployment in real-world

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models

ResearchDGX agent

arXiv:2606.20970v1 Announce Type: new Abstract: Omni-modal models can ingest video, audio, and text, but unified access to multiple modalities does not guarantee that a model uses the right evidence.

Compressing Observation History into Agent Memory: Distilling Transformers into Recurrent Transformers

AgentsDGX agent

arXiv:2606.21562v1 Announce Type: new Abstract: Transformers are AI's workhorse with strong performance in modeling sequential data, but their computational cost becomes prohibitive when processing lo

Compression and Retrieval: Implicit Memory Retrieval for Video World Models

ResearchDGX agent

arXiv:2606.23105v1 Announce Type: new Abstract: Video world models hold promise for simulating interactive environments, yet maintaining consistent long-term memory across complex camera trajectories

Concept Alignment Contrast and Long-Short Prompt Memory for Test-Time Adaptation of SAM3 in Medical Image Segmentation

SafetyDGX agent

arXiv:2606.22963v1 Announce Type: new Abstract: Concept segmentation models like Segment Anything Model 3 (SAM3) show strong generalization on natural images, yet their performance degrades in medical

Confidence-Uncertainty Boundary Calibration for Bayesian Deep Learning in Medical Image Analysis

SafetyDGX agent

arXiv:2602.11973v2 Announce Type: replace Abstract: In critical decision support systems based on medical imaging, the reliability of AI-assisted decision-making is as relevant as predictive accuracy.

Configurable Algorithms for Histopathologic Cancer Detection on Quantum Hardware

Local AiDGX agent

arXiv:2606.21752v1 Announce Type: cross Abstract: Histopathologic cancer detection is challenging due to tissue variability, staining differences, and subtle visual distinctions between disease classe

ConnectomeBench2: A Unified Benchmark for Automated Connectomic Proofreading

Model ReleasesDGX agent

arXiv:2606.21116v1 Announce Type: new Abstract: Proofreading--correcting segmentation errors in 3D brain reconstructions--is the rate-limiting step in synapse-resolution connectomics. We release Conne

Context-Aware Autoregressive Diffusion for Gloss-Wise Sign Language Production

ApplicationsDGX agent

arXiv:2606.21234v1 Announce Type: new Abstract: To generate natural and accurate sentence-level sign language, synthesizing the 'gloss', the fundamental semantic unit, is essential. However, most curr

Contrastive and Adaptive Multi-modal Masked Autoencoder for Spatial Transcriptomics

TutorialsDGX agent

arXiv:2606.21156v1 Announce Type: new Abstract: The high cost of spatial transcriptomics (ST) has driven extensive studies into predicting gene expression directly from H&E histology images. However,

Controllable Texture Tiling with Transformed RoPE-Enhanced Diffusion Models

ResearchDGX agent

arXiv:2606.22945v1 Announce Type: cross Abstract: Realistic integration of user-specified textures into scene images is a fundamental task in computer graphics and image editing. While existing materi

CoSA: Correlation-Guided Change Attention with Learnable Residual Gating for Remote Sensing Change Detection

Model ReleasesDGX agent

arXiv:2606.21932v1 Announce Type: new Abstract: Remote sensing change detection (CD) from bi-temporal imagery is critical for applications such as urban monitoring, disaster assessment, and environmen

CourseBlueprint: A Structured Pipeline for Adaptive Pedagogical Video Generation Grounded in Course Corpora

Model ReleasesDGX agent

arXiv:2606.20608v1 Announce Type: cross Abstract: Generative text-to-video systems can produce visually fluent educational clips, but they rarely encode the pedagogical content knowledge (PCK) needed

CoVStream: Edge-Cloud Collaboration for Understanding of Long Video Streams

Local AiDGX agent

arXiv:2606.22804v1 Announce Type: new Abstract: Long, continuous video streams are an increasingly critical driver of multimedia intelligence. Existing efforts often handle long videos with a sample-e

Cross-Modal Corroboration for Annotation-Free Wildlife Monitoring

SafetyDGX agent

arXiv:2606.21613v1 Announce Type: new Abstract: Scaling wildlife monitoring for real-world conservation deployments requires automated analysis of smart sensors that operate under severe annotation sc

Cross-View Yaw Estimation in Location Uncertainty with Line-Aligning Yaw Scoring

ResearchDGX agent

arXiv:2606.22094v1 Announce Type: new Abstract: Accurate yaw estimation is a bottleneck in cross-view localization between ground view and Bird's Eye View (BEV). Existing methods couple yaw with trans

CuDi: Curve Distillation for Efficient and Controllable Exposure Adjustment

Local AiDGX agent

arXiv:2207.14273v2 Announce Type: replace Abstract: We present Curve Distillation, CuDi, for efficient and controllable exposure adjustment without the requirement of paired or unpaired data during tr

Cultural Counterfactuals: Evaluating Cultural Biases in Large Vision-Language Models with Counterfactual Examples

ResearchDGX agent

arXiv:2603.02370v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) have grown increasingly powerful in recent years, but can also exhibit harmful biases. Prior studies investigat

Curvature-Adaptive Consistency Flow Matching: Autonomous Trajectory Optimization via Reinforcement Learning

AgentsDGX agent

arXiv:2606.22394v1 Announce Type: new Abstract: Consistency distillation has significantly accelerated the inference of diffusion models. In this work, we reveal an intriguing asymmetry: while Logit-N

Curvature-aware 3D length estimation of greenhouse cucumbers using RGB-D imaging and cubic spline arc-length integration

Model ReleasesDGX agent

arXiv:2606.22439v1 Announce Type: new Abstract: Commercial greenhouse cucumber production is graded by fruit length, which drives harvest scheduling, labour allocation, and logistics. Manual measureme

CurvSegFlow: Time-Conditioned Flow Matching for Robust Segmentation of Curvilinear Structures in Noisy Biomedical Images

TutorialsDGX agent

arXiv:2606.21608v1 Announce Type: new Abstract: Accurate segmentation of curvilinear structures remains challenging in biomedical imaging due to their thin geometry, complex topology, and sensitivity

Customizing Video Portraits via Identity-ActionDecoupling

SafetyDGX agent

arXiv:2606.22347v1 Announce Type: new Abstract: Identity-Preserving Text-to-Video Generation (IPT2V) seeks to synthesize a temporally coherent video from a reference image and a textual description, w

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming

Model ReleasesDGX agent

arXiv:2606.22476v1 Announce Type: new Abstract: Humans can effortlessly reason about scenes across different viewpoints, yet it remains unclear whether Vision-Language Models (VLMs) possess similar cr

D2HDMap: Non-visible Driveline Map Prior for Online Vectorized HD Map Prediction

SafetyDGX agent

arXiv:2606.20725v1 Announce Type: new Abstract: Accurate, up-to-date representations of road structures are critical for the safe operation of autonomous vehicles. Existing systems rely either on cost

DamageArbiter: A Multimodal Arbitration Framework for Disaster Damage Assessment from Street-View Imagery

ResearchDGX agent

arXiv:2603.14837v2 Announce Type: replace Abstract: Analyzing street-view imagery with computer vision models offers a promising approach for rapid, hyperlocal disaster damage assessment, but existing

Data-Driven Image Registration and Deformation Modeling for Image-Guided Neurosurgery: A Systematic Review

SafetyDGX agent

arXiv:2602.10155v2 Announce Type: replace-cross Abstract: Accurate compensation of brain deformation is critical for reliable image-guided neurosurgery. Surgical manipulation and tumor resection induc

Data Selection Through Iterative Self-Filtering for Vision-Language Settings

ResearchDGX agent

arXiv:2606.23611v1 Announce Type: new Abstract: The availability of large amounts of clean data is paramount to training neural networks. However, at large scales, manual oversight is impractical, res

DBT-Bleed: Dual-Branch Temporal Modeling with Key-Frame Selection for Surgical Bleeding Detection

SafetyDGX agent

arXiv:2606.22829v1 Announce Type: new Abstract: Intraoperative Adverse Events (IAEs) detection is critical for improving surgical safety, with bleeding being among the most frequent events across many

DE-FIVE: Detecting Malicious Image Prompts via Fourier Features and Image Vector Embeddings

ResearchDGX agent

arXiv:2606.22779v1 Announce Type: cross Abstract: Vision language models (VLMs) employ both visual and textual modalities to enable advanced vision-language inference. However, incorporating visual mo

Decoupling the Declarative from the Procedural in Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.21496v1 Announce Type: cross Abstract: Deploying generalist robotic agents in the real world requires transferable skills. Specifically, a policy trained to clone a behavior from object-spe

Deep EM with Hierarchical Latent Label Modelling for Multi-Site Prostate Lesion Segmentation

Local AiDGX agent

arXiv:2603.14418v2 Announce Type: replace Abstract: Label variability is a major challenge for prostate lesion segmentation. In multi-site datasets, annotations often reflect centre-specific contourin

Deep Unrolled Networks in Representation Space Applied to MRI Reconstruction

TutorialsDGX agent

arXiv:2606.21602v1 Announce Type: cross Abstract: Deep unrolled networks (DUNs) integrate physical forward models with learned regularization in cascaded network architectures, achieving exceptional p

Delta-Diffusion: Modeling Longitudinal Brain Amyloid-PET Trajectories via Conditional Poisson Diffusion Bridge

SafetyDGX agent

arXiv:2606.22216v1 Announce Type: cross Abstract: While longitudinal brain PET imaging is the gold standard for quantifying the spatiotemporal accumulation of Beta-amyloid, its widespread clinical uti

Democratizing and accelerating AI-driven pathology research through agentic intelligence

AgentsDGX agent

arXiv:2606.20677v1 Announce Type: cross Abstract: Computational pathology has advanced rapidly with the emergence of foundation models, yet widespread adoption remains limited by substantial technical

Denoising-Enhanced Coarse-to-Fine Infrared Small Target Detection with Attention Prior-Guided Knowledge Distillation

ResearchDGX agent

arXiv:2606.21956v1 Announce Type: new Abstract: Infrared small target detection (IRSTD) in high-resolution images is crucial for many practical applications, such as surveillance of unmanned aerial ve

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views

SafetyDGX agent

arXiv:2606.23557v1 Announce Type: new Abstract: Multi-view 3D Visual Question Answering (MV3D-VQA) requires integrating partial observations into a coherent 3D scene representation and selecting infor

Detail++: Training-Free Detail Enhancer for T2I Diffusion Models

Model ReleasesDGX agent

arXiv:2507.17853v3 Announce Type: replace Abstract: Recent advances in text-to-image (T2I) generation have led to impressive visual results. However, these models still face significant challenges whe

Diffusion Integrated Gradients: Controllable Path Generation for Flexible Feature Attribution

TutorialsDGX agent

arXiv:2606.22314v1 Announce Type: cross Abstract: Path-based attribution methods such as Integrated Gradients (IG) are widely adopted for their strong axiomatic properties and effectiveness in attribu

← Previous
1…7677787980…211
Next →