AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
27 Apr 2026

Different Strokes for Different Folks: Writer Identification for Historical Arabic Manuscripts

Model ReleasesDGX agent

arXiv:2604.22515v1 Announce Type: new Abstract: Handwritten Arabic manuscripts preserve the Arab world's intellectual and cultural heritage, and writer identification supports provenance, authenticity

Distilling Vision Transformers for Distortion-Robust Representation Learning

TutorialsDGX agent

arXiv:2604.22529v1 Announce Type: new Abstract: Self-supervised learning has achieved remarkable success in learning visual representations from clean data, yet remains challenging when clean observat

DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.22281v1 Announce Type: new Abstract: Recent advances in vision-language models have demonstrated remarkable performance across diverse multi-modal tasks, including document question answeri

Edit-aware RAW Reconstruction

ResearchDGX agent

arXiv:2512.05859v2 Announce Type: replace Abstract: Users frequently edit camera images post-capture to achieve their preferred photofinishing style. While editing in the RAW domain provides greater a

Efficient Diffusion Distillation via Embedding Loss

ResearchDGX agent

arXiv:2604.22379v1 Announce Type: new Abstract: Recent advances in distilling expensive diffusion models into efficient few-step generators show significant promise. However, these methods typically d

EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges

Model ReleasesDGX agent

arXiv:2604.22595v1 Announce Type: new Abstract: CLIP has demonstrated strong generalization in visual domains through natural language supervision, even for video action recognition. However, most exi

Evaluation of image simulation open source solutions for simulation of synthetic images in lunar environment

AgentsDGX agent

arXiv:2604.22296v1 Announce Type: new Abstract: Synthetic image generation is one of the crucial input for planetary missions. It enables researchers and engineers to visualize planned planetary missi

EvFlow-GS: Event Enhanced Motion Deblurring with Optical Flow for 3D Gaussian Splatting

ResearchDGX agent

arXiv:2604.22183v1 Announce Type: new Abstract: Achieving sharp 3D reconstruction from motion-blurred images alone becomes challenging, motivating recent methods to incorporate event cameras, benefiti

Evolving Thematic Map Design in Academic Cartography: A Thirty-Year Study Based on Multilingual Journals

ResearchDGX agent

arXiv:2604.22539v1 Announce Type: new Abstract: Thematic maps play a central role in academic communication, yet their large-scale design evolution has rarely been examined empirically. This study pre

FeudalNav: A Simple Framework for Visual Navigation

ResearchDGX agent

arXiv:2602.06974v2 Announce Type: replace-cross Abstract: Visual navigation for robotics is inspired by the human ability to navigate environments using visual cues and memory, eliminating the need fo

FILTR: Extracting Topological Features from Pretrained 3D Models

Model ReleasesDGX agent

arXiv:2604.22334v1 Announce Type: new Abstract: Recent advances in pretraining 3D point cloud encoders (e.g., Point-BERT, Point-MAE) have produced powerful models, whose abilities are typically evalua

FLARE-BO: Fused Luminance and Adaptive Retinex Enhancement via Bayesian Optimisation for Low-Light Robotic Vision

Model ReleasesDGX agent

arXiv:2604.22093v1 Announce Type: new Abstract: Reliable visual perception under low illumination remains a core challenge for autonomous robotic systems, where degraded image quality directly comprom

Flow4DGS-SLAM: Optical Flow-Guided 4D Gaussian Splatting SLAM

Local AiDGX agent

arXiv:2604.22339v1 Announce Type: new Abstract: Handling the dynamic environments is a significant research challenge in Visual Simultaneous Localization and Mapping (SLAM). Recent research combines 3

FlowAnchor: Stabilizing the Editing Signal for Inversion-Free Video Editing

SafetyDGX agent

arXiv:2604.22586v1 Announce Type: new Abstract: We propose FlowAnchor, a training-free framework for stable and efficient inversion-free, flow-based video editing. Inversion-free editing methods have

Forecasting Solar Energy Using a Single Image

ResearchDGX agent

arXiv:2604.21982v1 Announce Type: new Abstract: Solar panels are increasingly deployed in cities on rooftops, walls, and urban infrastructure. Although the panel costs have fallen in recent years, the

Generative Modeling of Neurodegenerative Brain Anatomy with 4D Longitudinal Diffusion Model

ResearchDGX agent

arXiv:2604.22700v1 Announce Type: new Abstract: Understanding and predicting the progression of neurodegenerative diseases remains a major challenge in medical AI, with significant implications for ea

GOSPA and T-GOSPA quasi-metrics for evaluation of multi-object tracking algorithms

TutorialsDGX agent

arXiv:2507.13706v2 Announce Type: replace Abstract: This paper introduces two quasi-metrics for performance assessment of multi-object tracking (MOT) algorithms. One quasi-metric is an extension of th

HFS-TriNet: A Three-Branch Collaborative Feature Learning Network for Prostate Cancer Classification from TRUS Videos

ResearchDGX agent

arXiv:2604.22388v1 Announce Type: new Abstract: Transrectal ultrasound (TRUS) imaging is a cost-effective and non-invasive modality widely used in the diagnosis of prostate cancer. The computer-aided

Holo360D: A Large-Scale Real-World Dataset with Continuous Trajectories for Advancing Panoramic 3D Reconstruction and Beyond

Model ReleasesDGX agent

arXiv:2604.22482v1 Announce Type: new Abstract: While feed-forward 3D reconstruction models have advanced rapidly, they still exhibit degraded performance on panoramas due to spherical distortions. Mo

How Many Visual Levers Drive Urban Perception? Interventional Counterfactuals via Multiple Localised Edits

SafetyDGX agent

arXiv:2604.22103v1 Announce Type: cross Abstract: Street-view perception models predict subjective attributes such as safety at scale, but remain correlational: they do not identify which localized vi

ICPR 2026 Competition on Low-Resolution License Plate Recognition

ApplicationsDGX agent

arXiv:2604.22506v1 Announce Type: new Abstract: Low-Resolution License Plate Recognition (LRLPR) remains a challenging problem in real-world surveillance scenarios, where long capture distances, compr

Improving Driver Drowsiness Detection via Personalized EAR/MAR Thresholds and CNN-Based Classification

SafetyDGX agent

arXiv:2604.22479v1 Announce Type: new Abstract: Driver drowsiness is a major cause of traffic accidents worldwide, posing a serious threat to public safety. Vision-based driver monitoring systems ofte

Inter-Stance: A Dyadic Multimodal Corpus for Conversational Stance Analysis

ResearchDGX agent

arXiv:2604.22739v1 Announce Type: new Abstract: Social interactions dominate our perceptions of the world and shape our daily behavior by attaching social meaning to acts as simple and spontaneous as

Knowledge Visualization: A Benchmark and Method for Knowledge-Intensive Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2604.22302v1 Announce Type: new Abstract: Recent text-to-image (T2I) models have demonstrated impressive capabilities in photorealistic synthesis and instruction following. However, their reliab

Learning Reactive Human Motion Generation from Paired Interaction Data Using Transformer-Based Models

AgentsDGX agent

arXiv:2604.22164v1 Announce Type: new Abstract: Recent advances in deep learning have enabled the generation of videos from textual descriptions as well as the prediction of future sequences from inpu

Long-tail Internet photo reconstruction

Model ReleasesDGX agent

arXiv:2604.22714v1 Announce Type: new Abstract: Internet photo collections exhibit an extremely long-tailed distribution: a few famous landmarks are densely photographed and easily reconstructed in 3D

LTBs-KAN: Linear-Time B-splines Kolmogorov-Arnold Networks

Model ReleasesDGX agent

arXiv:2604.22034v1 Announce Type: cross Abstract: Kolmogorov-Arnold Networks (KANs) are a recent neural network architecture offering an alternative to Multilayer Perceptrons (MLPs) with improved expl

MTT-Bench: Predicting Social Dominance in Mice via Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2604.22492v1 Announce Type: cross Abstract: Understanding social dominance in animal behavior is critical for neuroscience and behavioral studies. In this work, we explore the capability of Mult

Multimodal Diffusion to Mutually Enhance Polarized Light and Low Resolution EBSD Data

TutorialsDGX agent

arXiv:2604.22212v1 Announce Type: cross Abstract: In spite of the utility of 3-D electron back-scattered diffraction (EBSD) microscopy, the data collection process can be time-consuming with serial-se

Non-Minimal Sampling and Consensus for Prohibitively Large Datasets

ResearchDGX agent

arXiv:2604.22518v1 Announce Type: new Abstract: We introduce NONSAC (Non-Minimal Sampling and Consensus), a general framework for robust and scalable model estimation from arbitrarily large datasets c

NRGS: Neural Regularization for Robust 3D Semantic Gaussian Splatting

ResearchDGX agent

arXiv:2604.22439v1 Announce Type: new Abstract: We propose a neural regularization method that refines the noisy 3D semantic field produced by lifting multi-view inconsistent 2D features, in order to

Nuclear Diffusion Models for Low-Rank Background Suppression in Videos

ApplicationsDGX agent

arXiv:2509.20886v2 Announce Type: replace Abstract: Video sequences often contain structured noise and background artifacts that obscure dynamic content, posing challenges for accurate analysis and re

OccDirector: Language-Guided Behavior and Interaction Generation in 4D Occupancy Space

Model ReleasesDGX agent

arXiv:2604.22240v1 Announce Type: new Abstract: Generative world models increasingly rely on 4D occupancy for realistic autonomous driving simulation. However, existing generation frameworks depend on

One Shot Learning for Edge Detection on Point Clouds

TutorialsDGX agent

arXiv:2604.22354v1 Announce Type: new Abstract: Each scanner possesses its unique characteristics and exhibits its distinct sampling error distribution. Training a network on a dataset that includes d

PAGaS: Pixel-Aligned 1DoF Gaussian Splatting for Depth Refinement

ResearchDGX agent

arXiv:2604.22129v1 Announce Type: new Abstract: Gaussian Splatting (GS) has emerged as an efficient approach for high-quality novel view synthesis. While early GS variants struggled to accurately mode

PASR: Pose-Aware 3D Shape Retrieval from Occluded Single Views

TutorialsDGX agent

arXiv:2604.22658v1 Announce Type: new Abstract: Single-view 3D shape retrieval is a fundamental yet challenging task that is increasingly important with the growth of available 3D data. Existing appro

PoseFM: Relative Camera Pose Estimation Through Flow Matching

AgentsDGX agent

arXiv:2604.22350v1 Announce Type: new Abstract: Monocular visual odometry (VO) is a fundamental computer vision problem with applications in autonomous navigation, augmented reality and more. While de

Railway Artificial Intelligence Learning Benchmark (RAIL-BENCH): A Benchmark Suite for Perception in the Railway Domain

Model ReleasesDGX agent

arXiv:2604.22507v1 Announce Type: new Abstract: Automated train operation on existing railway infrastructure requires robust camera-based perception, yet the railway domain lacks public benchmark suit

Recent Advances in Multi-Agent Human Trajectory Prediction: A Comprehensive Review

Model ReleasesDGX agent

arXiv:2506.14831v3 Announce Type: replace Abstract: With the emergence of powerful data-driven methods in human trajectory prediction (HTP), gaining a finer understanding of multi-agent interactions l

Region Matters: Efficient and Reliable Region-Aware Visual Place Recognition

Model ReleasesDGX agent

arXiv:2604.22390v1 Announce Type: new Abstract: Visual Place Recognition (VPR) determines a query image's geographic location by matching it against geotagged databases. However, existing methods stru

ReLIC-SGG: Relation Lattice Completion for Open-Vocabulary Scene Graph Generation

ResearchDGX agent

arXiv:2604.22546v1 Announce Type: new Abstract: Open-vocabulary scene graph generation (SGG) aims to describe visual scenes with flexible relation phrases beyond a fixed predicate set. Existing method

Rethinking Token Pruning for Historical Screenshots in GUI Visual Agents: Semantic, Spatial, and Temporal Perspectives

ResearchDGX agent

arXiv:2603.26041v3 Announce Type: replace Abstract: In recent years, GUI visual agents built upon Multimodal Large Language Models (MLLMs) have demonstrated strong potential in navigation tasks. Howev

Revisiting Geometric Obfuscation with Dual Convergent Lines for Privacy-Preserving Image Queries in Visual Localization

ResearchDGX agent

arXiv:2604.22310v1 Announce Type: new Abstract: Privacy-Preserving Image Queries (PPIQ) are an emerging mechanism for cloud-based visual localization, enabling pose estimation from obfuscated features

Robust Camera-to-Mocap Calibration and Verification for Large-Scale Multi-Camera Data Capture

ApplicationsDGX agent

arXiv:2604.22118v1 Announce Type: new Abstract: Optical motion capture (mocap) systems are widely used for ground-truth capture in AR/VR, SLAM and robotics datasets. These datasets require extrinsic c

SAMIDARE: Advanced Tracking-by-Segmentation for Dense Scenarios

ResearchDGX agent

arXiv:2604.22162v1 Announce Type: new Abstract: Automated sports analysis demands robust multi-object tracking (MOT), yet segmentation-based methods often struggle with mask errors and ID switches in

Score-based Membership Inference on Diffusion Models

Model ReleasesDGX agent

arXiv:2509.25003v2 Announce Type: replace-cross Abstract: Membership inference attacks (MIAs) against Diffusion Models (DMs) raise pressing privacy concerns by revealing whether a sample was part of t

Segment Any-Quality Images with Generative Latent Space Enhancement

ApplicationsDGX agent

arXiv:2503.12507v3 Announce Type: replace Abstract: Despite their success, Segment Anything Models (SAMs) experience significant performance drops on severely degraded, low-quality images, limiting th

Segmentation of Gray Matters and White Matters from Brain MRI data

ResearchDGX agent

arXiv:2603.29171v3 Announce Type: replace Abstract: Accurate segmentation of brain tissues such as gray matter and white matter from magnetic resonance imaging is essential for studying brain anatomy,

Selective Depthwise Separable Convolution for Lightweight Joint Source-Channel Coding in Wireless Image Transmission

Model ReleasesDGX agent

arXiv:2604.22338v1 Announce Type: cross Abstract: Depthwise separable convolutional (DSConv) layers have been successfully applied to deep learning (DL)-based joint source-channel coding (JSCC) scheme

Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging

Model ReleasesDGX agent

arXiv:2510.05971v3 Announce Type: replace Abstract: The generalization of the Transformer architecture via MetaFormer has reshaped our understanding of its success in computer vision. By replacing sel

SIE3D: Single-Image Expressive 3D Avatar Generation via Semantic Embedding and Perceptual Expression Loss

HardwareDGX agent

arXiv:2509.24004v2 Announce Type: replace Abstract: Generating high-fidelity 3D head avatars from a single image is challenging, as current methods lack fine-grained, intuitive control over expression

Soft Anisotropic Diagrams for Differentiable Image Representation

HardwareDGX agent

arXiv:2604.21984v1 Announce Type: new Abstract: We introduce Soft Anisotropic Diagrams (SAD), an explicit and differentiable image representation parameterized by a set of adaptive sites in the image

SpaMEM: Benchmarking Dynamic Spatial Reasoning via Perception-Memory Integration in Embodied Environments

Model ReleasesDGX agent

arXiv:2604.22409v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have advanced static visual--spatial reasoning, yet they often fail to preserve long-horizon spatial coherence

SS3D: End2End Self-Supervised 3D from Web Videos

ResearchDGX agent

arXiv:2604.22686v1 Announce Type: new Abstract: We present SS3D, a web-scale SfM-based self-supervision pretraining pipeline for feed-forward 3D estimation from monocular video. Our model jointly pred

Structure-Guided Diffusion Model for EEG-Based Visual Cognition Reconstruction

TutorialsDGX agent

arXiv:2604.22649v1 Announce Type: cross Abstract: Objective: Decoding visual information from electroencephalography (EEG) is an important problem in neuroscience and brain-computer interface (BCI) re

Sum-of-Checks: Structured Reasoning for Surgical Safety with Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.22156v1 Announce Type: cross Abstract: Purpose: Accurate assessment of the Critical View of Safety (CVS) during laparoscopic cholecystectomy is essential to prevent bile duct injury, a comp

Thermal background reduction for mid-infrared imaging by low-rank background and sparse point-source modelling

SafetyDGX agent

arXiv:2604.22351v1 Announce Type: cross Abstract: Mid-infrared astronomy from the ground faces critical challenges in accurately detecting and quantifying sources due to the dominant spatially and tim

Towards Temporal Compositional Reasoning in Long-Form Sports Videos

Model ReleasesDGX agent

arXiv:2604.22226v1 Announce Type: new Abstract: Sports videos are a challenging domain for multimodal understanding because they involve complex and dynamic human activities. Despite rapid progress in

Transferable Physical-World Adversarial Patches Against Pedestrian Detection Models

SafetyDGX agent

arXiv:2604.22552v1 Announce Type: new Abstract: Physical adversarial patch attacks critically threaten pedestrian detection, causing surveillance and autonomous driving systems to miss pedestrians and

Uni-Encoder Meets Multi-Encoders: Representation Before Fusion for Brain Tumor Segmentation with Missing Modalities

ResearchDGX agent

arXiv:2604.22177v1 Announce Type: new Abstract: Multimodal MRI offers complementary information for brain tumor segmentation, but clinical scans often lack one or more modalities, which degrades segme

← Previous
1…172173174175176…209
Next →