AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
5 Jun 2026

MS-DKC: A Dataset Knowledge Card Framework for Designing and Adapting Medical Image Segmentation Models

ResearchDGX agent

arXiv:2606.06103v1 Announce Type: new Abstract: Medical image segmentation is often framed as a search for stronger architectures, but this can obscure a more fundamental question: what does the datas

Multi-Task Crack Foundation Model for Engineering-Reliable Crack Representation and Topology Preservation in Civil Infrastructure

ResearchDGX agent

arXiv:2606.05641v1 Announce Type: new Abstract: Reliable crack assessment requires not only accurate pixel-level masks but also connected crack geometry and confidence estimates that remain stable und

Multimodal Sexism Identification and Characterization using Large Language Models and Gradient Boosting

Research

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2606.05997v1 Announce Type: new Abstract: We present the AILS-NTUA submission to the EXIST 2026 Lab at CLEF, addressing multimodal sexism identification and characterization in memes (Task 2) an

Next-Generation Parallel Decoder for LPDR: Architectural Optimization and Class-Balanced GAN-Augmentation

ResearchDGX agent

arXiv:2606.05785v1 Announce Type: new Abstract: Real-Time License Plate Detection and Recognition (LPDR) forms the backbone of modern smart cities. Although the YOLOV5-PDLPR model substantially improv

NIV: Neural Axis Variations for Variable Font Generation

ResearchDGX agent

arXiv:2606.05261v1 Announce Type: new Abstract: Variable fonts enable continuous variation of glyph geometry along semantic design axes such as weight, width, slant, and optical size. However, constru

Noise-Adaptive Regularization for Robust Multi-Label Remote Sensing Image Classification

ResearchDGX agent

arXiv:2601.08446v2 Announce Type: replace Abstract: The development of reliable methods for multi-label classification (MLC) has become a prominent research direction in remote sensing (RS). As the sc

Noise-Aware Visual Representation Learning for Medical Visual Question Answering

Model ReleasesDGX agent

arXiv:2606.05535v1 Announce Type: new Abstract: Medical visual question answering (Med-VQA) has strong potential for clinical decision support by enabling AI models to interpret medical images and ans

Oklch+: A Three-Parameter Extension of Oklab for Improved Color Difference Prediction

Model ReleasesDGX agent

arXiv:2606.05255v1 Announce Type: cross Abstract: Oklab and its cylindrical representation Oklch are widely adopted in interpolation and design workflows as perceptually motivated color spaces, but th

ORACLE-CT: Anatomy-Aware Support Pooling for CT Classification

ResearchDGX agent

arXiv:2606.05460v1 Announce Type: new Abstract: Abdominal CT disease classification is challenging because each scan is a large 3D volume with many possible findings, while diagnostic evidence is ofte

PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding

ResearchDGX agent

arXiv:2606.06485v1 Announce Type: new Abstract: Recent advances in 3D multimodal large language models (3D-MLLMs) have enabled unified solutions for 3D scene understanding tasks, including visual ques

Parallel Jacobi Decoding for Fast Autoregressive Image Generation

ResearchDGX agent

arXiv:2606.05703v1 Announce Type: new Abstract: Autoregressive (AR) models have demonstrated remarkable performance in generating high-fidelity images. However, their inherently sequential next-token

PathWISE: Multi-Agent Cancer Pathway Triaging Ontology Learning from Clinical Flowcharts

AgentsDGX agent

arXiv:2605.25970v2 Announce Type: replace Abstract: Clinical pathways are disseminated as visual flowcharts where spatial topology, arrow direction, colour coding, and font weight encode critical tria

PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation

SafetyDGX agent

arXiv:2503.14295v3 Announce Type: replace Abstract: Recent advancements in audio-driven talking face generation have made great progress in lip synchronization. However, current methods often lack suf

Personal AI Agent for Camera Roll VQA

AgentsDGX agent

arXiv:2606.05275v1 Announce Type: new Abstract: We study the personal camera roll visual question answering setting. In this setting, a conversational AI assistant can access a user's personal camera

Physics-Guided Deep Unfolding for Blind Cross-Sensor Spectral Super-Resolution via Learning the Spectral Transformation Function

Model ReleasesDGX agent

arXiv:2606.05759v1 Announce Type: new Abstract: Hyperspectral imaging provides rich spectral information for quantitative remote sensing, yet hyperspectral sensors remain costly and thus unavailable i

Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them

ResearchDGX agent

arXiv:2606.06361v1 Announce Type: new Abstract: Image-to-Video diffusion models leverage input images to generate visually stunning content, yet frequently produce motion that violates physical laws.

RadiusFPS: Efficient Farthest Point Sampling on CPUs and GPUs via Spherical Voxel Pruning

Local AiDGX agent

arXiv:2606.06255v1 Announce Type: cross Abstract: Point clouds are a primary sensory representation for robotic perception, underpinning LiDAR-based autonomous driving, simultaneous localization and m

RAPTOR+: A Visually Grounded Vision-Language Framework to Improve Clinical Trust and Auditability in Automated Cancer Referral Processing

Model ReleasesDGX agent

arXiv:2605.25956v2 Announce Type: replace Abstract: Urgent suspected colorectal cancer (CRC) referrals create operational bottlenecks because semi-structured clinical documents often require manual re

Real-Time Threat Detection from Surveillance Cameras using Machine Learning

SafetyDGX agent

arXiv:2606.05708v1 Announce Type: new Abstract: Ensuring public safety in densely populated urban environments remains a critical challenge, necessitating the deployment of intelligent and automated v

ReCache: Learning Budget-Aware Caching Schedules for Diffusion Models via REINFORCE

SafetyDGX agent

arXiv:2606.06060v1 Announce Type: new Abstract: Modern diffusion models generate high-quality images and videos, but their iterative denoising process makes inference expensive. Feature caching accele

Recovering Physically Plausible Human-Object Interactions from Monocular Videos

SafetyDGX agent

arXiv:2606.05359v1 Announce Type: new Abstract: In this paper, we propose RePHO, a method to reconstruct physically plausible human-object interactions (HOI) from monocular videos. While existing kine

ReSAGE-PAR: Representational Similarity Assessment for Generative Expansion in Pedestrian Attribute Recognition

SafetyDGX agent

arXiv:2606.06020v1 Announce Type: new Abstract: To address the limited diversity and data scarcity in Pedestrian Attribute Recognition (PAR), we explore image synthesis using diffusion models guided b

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind

AgentsDGX agent

arXiv:2606.05896v1 Announce Type: new Abstract: Creating lifelike digital humans with genuine social intelligence requires unifying cognitive reasoning and multimodal generation within a coherent fram

RhymeFlow: Training-Free Acceleration for Video Generation with Asynchronous Denoising Flow Scheduling

TutorialsDGX agent

arXiv:2606.06309v1 Announce Type: new Abstract: Video generation models based on Diffusion Transformers (DiTs) have achieved remarkable performance in video synthesis, yet they suffer from high infere

Robust Scene Transfer for PointGoal Navigation via Privileged Sensor Guided Contrastive Learning

SafetyDGX agent

arXiv:2606.05506v1 Announce Type: new Abstract: We propose a sensor-guided adaptive contrastive learning framework for visual representation learning in PointGoal navigation. During training, privileg

RoCA: Robust Cross-Domain End-to-End Autonomous Driving

AgentsDGX agent

arXiv:2506.10145v3 Announce Type: replace Abstract: End-to-end (E2E) autonomous driving has recently emerged as a new paradigm, offering significant potential. However, few studies have looked into th

RQUL-UIE: Revitalizing Quality-Unstable Labels for Underwater Image Enhancement via In-Dataset Self-Supervision

ResearchDGX agent

arXiv:2606.06176v1 Announce Type: new Abstract: Underwater Image Enhancement (UIE) is essential for mitigating degradations caused by water medium. Although learning-based methods have advanced signif

SAM-Flow: Source-Anchored Masked Flow for Training-Free Image Editing

ResearchDGX agent

arXiv:2606.06228v1 Announce Type: new Abstract: Training-free image editing has recently attracted increasing attention due to its ability to modify real images using powerful pre-trained diffusion an

SC-MFJ: A Simple Haptic Quality Metric for Medical Image Segmentation

ResearchDGX agent

arXiv:2606.06199v1 Announce Type: new Abstract: Standard segmentation metrics such as Dice and Hausdorff distance measure geometric overlap but say nothing about whether a segmented surface is suitabl

Second-order Gaussian directional derivative representations for image high-resolution corner detection

TutorialsDGX agent

arXiv:2601.08182v2 Announce Type: replace Abstract: Corner detection is widely used in various computer vision tasks, such as image matching and 3D reconstruction. Our research indicates that there ar

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.05702v1 Announce Type: cross Abstract: Recent advancements in Vision-Language Models (VLMs) have significantly enhanced their ability to interpret complex visual semantics, yet their capaci

Self-Learning Expression Deformations for Data-Efficient Gaussian Avatars

ResearchDGX agent

arXiv:2606.05912v1 Announce Type: new Abstract: Modeling dynamic facial expressions using 3D Gaussian representations remains challenging due to their unstructured nature. Conventional Gaussian avatar

Self-supervised Feature Disentanglement and Augmentation Network for One-class Face Anti-spoofing

ResearchDGX agent

arXiv:2503.22929v3 Announce Type: replace Abstract: Face anti-spoofing (FAS) techniques aim to enhance the security of facial identity authentication by distinguishing authentic live faces from decept

Semantic-decoupled Spatial Partition Guided Point-supervised Oriented Object Detection

HardwareDGX agent

arXiv:2506.10601v2 Announce Type: replace Abstract: Given its ability to reduce annotation costs, weakly supervised learning based on single-point annotations has emerged as a research focus in orient

ShotCrop^3: Cropping Human-Centric Images into Cinematic Triple-Shot Compositions

Model ReleasesDGX agent

arXiv:2606.05635v1 Announce Type: new Abstract: Prior work on aesthetic composition typically produces a single aesthetically pleasing crop, overlooking the narrative value of composing multiple shots

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset

Model ReleasesDGX agent

arXiv:2606.06338v1 Announce Type: new Abstract: Video question answering (VideoQA) aims to answer questions about given videos. While existing approaches excel on factoid VideoQA, they struggle with d

Symb-xMIL: Symbolic Explanations for Multiple Instance Learning in Digital Pathology

SafetyDGX agent

arXiv:2606.06224v1 Announce Type: new Abstract: Explanations of multiple instance learning (MIL) models are widely used for validation and discovery in digital histopathology. Existing methods primari

Synthetic Data Generation and Vision-based Wrinkle and Keypoint Detection for Bimanual Cloth Manipulation

ApplicationsDGX agent

arXiv:2606.06292v1 Announce Type: new Abstract: Robotic manipulation of textiles remains challenging because continuous deformation and self-occlusions hinder the robust visual perception required to

T-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality Segmentation

Local AiDGX agent

arXiv:2606.05975v1 Announce Type: new Abstract: Open-vocabulary 3D functionality segmentation enables robots to localize functional object components in 3D scenes. It is a challenging task that requir

T-SAR-JEPA: Self-Supervised Temporal Anomaly Detection in SAR Amplitude Stacks via Latent Prediction

Local AiDGX agent

arXiv:2606.05700v1 Announce Type: new Abstract: We present T-SAR-JEPA, a self-supervised framework for temporal anomaly detection in SAR amplitude stacks via latent prediction. A ViT-Base/16 encoder f

Tamaththul3D: High-Fidelity 3D Saudi Sign Language Avatars from Monocular Video

ResearchDGX agent

arXiv:2605.05367v2 Announce Type: replace Abstract: Existing 3D sign language avatar reconstruction methods are developed and evaluated exclusively on Western sign languages, and no 3D parametric anno

Texture-preserving implicit neural representation for Cone beam CT truncated reconstruction

SafetyDGX agent

arXiv:2606.06039v1 Announce Type: new Abstract: Cone-beam computed tomography (CBCT) frequently suffers from data truncation, which introduces severe artifacts and limits the effective field of view (

TextWand: A Unified Framework for Scene Text Editing

Model ReleasesDGX agent

arXiv:2606.05730v1 Announce Type: new Abstract: We propose TextWand, a general-purpose framework that unifies scene text removal, generation, and replacement into a single model. By decomposing comple

The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show

ResearchDGX agent

arXiv:2606.05328v1 Announce Type: cross Abstract: Modern video diffusion models generate increasingly realistic and temporally coherent videos, motivating their use as candidate world simulators. Yet

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

Model ReleasesDGX agent

arXiv:2606.06476v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have shown strong visual reasoning capabilities, their spatial reasoning abilities remain largely constrained to the

Three-Dimensional Retinal Microvasculature Restoration in OCT Angiography

ResearchDGX agent

arXiv:2606.05375v1 Announce Type: new Abstract: Optical coherence tomographic angiography (OCTA) is a powerful technique for imaging retinal microvasculature. However, acquiring reliable quantificatio

TopoPult-SSL: Gland-Mask-Free Cross-Device Meibomian Gland Segmentation via Self-Distilled Weak Clinical Priors

Model ReleasesDGX agent

arXiv:2606.05347v1 Announce Type: new Abstract: Every new clinical imaging device creates a domain shift where dense gland masks are expensive yet cheap clinical signals -- eyelid outlines, Pult grade

Towards Accurate Heart Rate Measurement from Ultra-Short Video Clips via Periodicity-Guided rPPG Estimation and Signal Reconstruction

Model ReleasesDGX agent

arXiv:2506.22078v2 Announce Type: replace Abstract: Many remote Heart Rate (HR) measurement methods focus on estimating remote photoplethysmography (rPPG) signals from video clips lasting around 10 se

Towards Label-Noise Resistant Learning via Optimal Brain Damage Masking

ApplicationsDGX agent

arXiv:2508.09697v3 Announce Type: replace-cross Abstract: Noisy labels are inevitable in real-world scenarios. Due to the strong capacity of deep neural networks to memorize corrupted labels, these no

Towards One-to-Many Temporal Grounding

Model ReleasesDGX agent

arXiv:2606.06294v1 Announce Type: new Abstract: Temporal Grounding (TG) aims to localize video segments corresponding to a textual query. Prior research predominantly focuses on single-segment retriev

Two-Way Is Better Than One: Bidirectional Alignment with Cycle Consistency for Exemplar-Free Class-Incremental Learning

SafetyDGX agent

arXiv:2606.05675v1 Announce Type: cross Abstract: Continual learning (CL) seeks models that acquire new skills without erasing prior knowledge. In exemplar-free class-incremental learning (EFCIL), thi

UltraVR: A Diagnostic Ultra-Resolution Image-VQA Benchmark for Evidence-Grounded Reasoning

Model ReleasesDGX agent

arXiv:2606.05576v1 Announce Type: new Abstract: Vision-language models (VLMs) excel on visual question answering and multimodal reasoning benchmarks. Yet their capability on ultra-resolution images -

Uncertainty-Aware Adaptive Sensor Fusion for Autonomous Navigation

HardwareDGX agent

arXiv:2606.05437v1 Announce Type: cross Abstract: This work introduces a hybrid deep learning approach integrated with an Unscented Kalman Filter (UKF) to enhance pose estimation accuracy in Visual-In

UnHype: CLIP-Guided Hypernetworks for Dynamic LoRA Unlearning

ResearchDGX agent

arXiv:2602.03410v2 Announce Type: replace Abstract: Recent advances in large-scale diffusion models have intensified concerns about their potential misuse, particularly in generating realistic yet har

Unifying Dataset Pruning and Distillation for Efficient Large-scale Compression

Model ReleasesDGX agent

arXiv:2502.06434v2 Announce Type: replace Abstract: Dataset pruning (DP) and dataset distillation (DD) fundamentally differ in their outputs: DP selects original image subsets, while DD generates synt

UniPixie: Unified and Probabilistic 3D Physics Learning via Flow Matching

Model ReleasesDGX agent

arXiv:2606.05399v1 Announce Type: new Abstract: Existing feed-forward networks excel at predicting a single set of physical properties from visual appearance, but this point-estimate paradigm fundamen

Unpaired RGB-Thermal Gaussian-Splatting Using Visual Geometric Transformers

SafetyDGX agent

arXiv:2606.05491v1 Announce Type: new Abstract: Multi-modal novel view synthesis (NVS) combining RGB and thermal imagery enables precise 3D scene reconstruction with visual and thermal information. Ho

Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors

ResearchDGX agent

arXiv:2507.12336v2 Announce Type: replace Abstract: Most existing 3D keypoint estimation methods rely on manual annotations or calibrated multi-view images, both of which are expensive to collect. Thi

Unveiling the Unknown: Open Vocabulary Object Detection with Scene Graphs

SafetyDGX agent

arXiv:2606.05916v1 Announce Type: new Abstract: Open-vocabulary object detection seeks to identify novel object categories that were not part of the training data. Many knowledge distillation-based ap

Using street view images and visual LLMs to predict heritage values for governance support: Risks, ethics, and policy implications

SafetyDGX agent

arXiv:2601.06056v2 Announce Type: replace-cross Abstract: During 2025 and 2026, the Energy Performance of Buildings Directive is being implemented in the European Union member states, requiring all me

← Previous
1…9495969798…211
Next →