AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
28 Apr 2026

ELSA: Exact Linear-Scan Attention for Fast and Memory-Light Vision Transformers

Model ReleasesDGX agent

arXiv:2604.23798v1 Announce Type: cross Abstract: Existing attention accelerators often trade exact softmax semantics, depend on fused Tensor Core kernels, or incur sequential depth that limits FP32 t

EMCompress: Video-LLMs with Endomorphic Multimodal Compression

Model ReleasesDGX agent

arXiv:2508.21094v3 Announce Type: replace Abstract: Video-LLMs face a fundamental tension in long-video reasoning: static, sparse frame sampling either dilutes evidence across task-irrelevant segments

Enhanced Privacy and Communication Efficiency in Non-IID Federated Learning with Adaptive Quantization and Differential Privacy

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.23426v1 Announce Type: new Abstract: Federated learning (FL) is a distributed machine learning method where multiple devices collaboratively train a model under the management of a central

Evaluating Remote Sensing Image Captions Beyond Metric Biases

ResearchDGX agent

arXiv:2604.22855v1 Announce Type: new Abstract: The core objective of image captioning is to achieve lossless semantic compression from visual signals into textual modalities. However, the reliance on

EX-FIQA: Leveraging Intermediate Early eXit Representations from Vision Transformers for Face Image Quality Assessment

Model ReleasesDGX agent

arXiv:2604.22842v1 Announce Type: new Abstract: Face Image Quality Assessment is crucial for reliable face recognition systems, yet existing Vision Transformer-based approaches rely exclusively on fin

EXACT: an explainable anomaly-aware vision foundation model for analysis of 3D chest CT

Local AiDGX agent

arXiv:2604.24146v1 Announce Type: new Abstract: Chest computed tomography (CT) is central to the detection and management of thoracic disease, yet the growing scale and complexity of volumetric imagin

Exploring Hierarchical Consistency and Unbiased Objectness for Open-Vocabulary Object Detection

Model ReleasesDGX agent

arXiv:2604.23344v1 Announce Type: new Abstract: Conventional object detectors typically operate under a closed-set assumption, limiting recognition to a predefined set of base classes seen during trai

FastAT Benchmark: A Comprehensive Framework for Fair Evaluation of Fast Adversarial Training Methods

Model ReleasesDGX agent

arXiv:2604.22853v1 Announce Type: new Abstract: Fast Adversarial Training (FastAT) seeks to achieve adversarial robustness at a fraction of the computational cost incurred by standard multi-step metho

FDIM: A Feature-distance-based Generic Video Quality Metric for Versatile Codecs

Model ReleasesDGX agent

arXiv:2604.24123v1 Announce Type: new Abstract: Video technology is advancing toward Ultra High Definition (UHD) and High Dynamic Range (HDR), which intensifies the need for higher compression efficie

FlashOverlap: Minimizing Tail Latency in Communication Overlap for Distributed LLM Training

ResearchDGX agent

arXiv:2604.24013v1 Announce Type: cross Abstract: The rapid growth in the size of large language models has necessitated the partitioning of computational workloads across accelerators such as GPUs, T

Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning

Model ReleasesDGX agent

arXiv:2507.14137v4 Announce Type: replace Abstract: We present Franca (pronounced Fran-ka): free one; the first fully open-source (data, code, weights) vision foundation model that matches and in many

From Edges to Depth: Probing the Spatial Hierarchy in Vision Transformers

ResearchDGX agent

arXiv:2604.23452v1 Announce Type: new Abstract: Vision Transformers trained only on image classification routinely transfer to tasks that demand spatial understanding, yet they receive no spatial supe

FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos

ResearchDGX agent

arXiv:2604.05621v2 Announce Type: replace Abstract: We present FunRec, a method for reconstructing functional 3D digital twins of indoor scenes directly from egocentric RGB-D interaction videos. Unlik

FUSER: Feed-Forward MUltiview 3D Registration Transformer and SE(3)^N Diffusion Refinement

ResearchDGX agent

arXiv:2512.09373v2 Announce Type: replace Abstract: Registration of multiview point clouds conventionally relies on extensive pairwise matching to build a pose graph for global synchronization, which

GA2-CLIP: Generic Attribute Anchor for Efficient Prompt Tuningin Video-Language Models

ResearchDGX agent

arXiv:2511.22125v2 Announce Type: replace Abstract: Visual and textual soft prompt tuning can effectively improve the adaptability of Vision-Language Models (VLMs) in downstream tasks. However, fine-t

GCP: Guarded Collaborative Perception with Spatial-Temporal Aware Malicious Agent Detection

Model ReleasesDGX agent

arXiv:2501.02450v2 Announce Type: replace Abstract: Collaborative perception significantly enhances autonomous driving safety by extending each vehicle's perception range through message sharing among

GenAssets: Generating in-the-wild 3D Assets in Latent Space

ResearchDGX agent

arXiv:2604.23010v1 Announce Type: new Abstract: High-quality 3D assets for traffic participants are critical for multi-sensor simulation, which is essential for the safe end-to-end development of auto

Generalising maximum mean discrepancy: kernelised functional Bregman divergences

Model ReleasesDGX agent

arXiv:2604.24047v1 Announce Type: cross Abstract: Bregman divergences play a pivotal role in statistics, machine learning and computational information geometry. Particularly in the context of machine

Generalizable CT-Free PET Attenuation and Scatter Correction for Pediatric Patients

ResearchDGX agent

arXiv:2604.22894v1 Announce Type: cross Abstract: Computed tomography (CT)-based attenuation and scatter correction improves quantitative PET but adds radiation exposure that is particularly undesirab

Geometric Analysis of Self-Supervised Vision Representations for Semantic Image Retrieval

ResearchDGX agent

arXiv:2604.24469v1 Announce Type: cross Abstract: Content-based image retrieval (CBIR) systems enable users to search images based on visual content instead of relying on metadata. The text domain has

Geometry-Conditioned Diffusion for Occlusion-Robust In-Bed Pose Estimation

ResearchDGX agent

arXiv:2604.23651v1 Announce Type: new Abstract: Robust in-bed human pose estimation under blanket occlusion remains challenging due to the scarcity of reliable labeled training data for heavily covere

GoClick: Lightweight Element Grounding Model for Autonomous GUI Interaction

Model ReleasesDGX agent

arXiv:2604.23941v1 Announce Type: new Abstract: Graphical User Interface (GUI) element grounding (precisely locating elements on screenshots based on natural language instructions) is fundamental for

Gradient-Guided Exploration of Generative Model's Latent Space for Controlled Iris Image Augmentations

ApplicationsDGX agent

arXiv:2511.09749v2 Announce Type: replace Abstract: Developing reliable iris recognition and presentation attack detection methods requires diverse datasets that capture realistic variations in iris f

Graph-augmented Segmentation of Complex Shapes in Laser Powder bed Fusion for Enhanced In Situ Inspection

Model ReleasesDGX agent

arXiv:2604.24234v1 Announce Type: new Abstract: The technological maturity of in situ inspection and monitoring methods in additive manufacturing is steadily increasing, enabling more efficient and pr

Group Orthogonal Low-Rank Adaptation for RGB-T Tracking

Model ReleasesDGX agent

arXiv:2512.05359v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning has emerged as a promising paradigm in RGB-T tracking, enabling downstream task adaptation by freezing pretrained pa

GS-DOT: Gaussian splatting-based image reconstruction for diffuse optical tomography

ResearchDGX agent

arXiv:2604.23675v1 Announce Type: cross Abstract: This work presents GS-DOT, a novel image reconstruction framework based on Gaussian Splatting (GS) for diffuse optical tomography (DOT). Inspired by G

H-SemiS: Hierarchical Fusion of Semi and Self-Supervised Learning for Knee Osteoarthritis Severity Grading

ResearchDGX agent

arXiv:2604.23335v1 Announce Type: new Abstract: Knee osteoarthritis (KOA) is a degenerative joint disease that can lead to chronic pain, reduced mobility, and long-term disability. Automated severity

HAC: Parameter-Efficient Hyperbolic Adaptation of CLIP for Zero-Shot VQA

Model ReleasesDGX agent

arXiv:2604.23665v1 Announce Type: new Abstract: Recent advances in representation learning have shown that hyperbolic geometry can offer a more expressive alternative to the Euclidean embeddings used

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation

HardwareDGX agent

arXiv:2604.23632v1 Announce Type: new Abstract: Real-time text-driven joint audio-video avatar generation requires jointly synthesizing portrait video and speech with high fidelity and precise synchro

Hierarchical Prototype-based Domain Priors for Multiple Instance Learning in Multimodal Histopathology Analysis

SafetyDGX agent

arXiv:2604.23982v1 Announce Type: new Abstract: Digital pathology has fundamentally altered diagnostic workflows by enabling the computational analysis of gigapixel Whole Slide Images (WSIs), yet effe

Hierarchical Spatio-Channel Clustering for Efficient Model Compression in Medical Image Analysis

Model ReleasesDGX agent

arXiv:2604.23375v1 Announce Type: new Abstract: Convolutional neural networks (CNNs) have become increasingly difficult to deploy in resource-constrained environments due to their large memory and com

Identity-Decoupled Anonymization for Visual Evidence in Multi-modal Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2604.23584v1 Announce Type: new Abstract: Multi-modal retrieval-augmented generation (MRAG) systems retrieve visual evidence from large image corpora to ground the responses of large multi-modal

ImmerIris: A Large-Scale Dataset and Benchmark for Off-Axis and Unconstrained Iris Recognition in Immersive Applications

Model ReleasesDGX agent

arXiv:2510.10113v3 Announce Type: replace Abstract: Recently, iris recognition is regaining prominence in immersive applications such as extended reality as a means of seamless user identification. Th

Improving Vision-language Models with Perception-centric Process Reward Models

Model ReleasesDGX agent

arXiv:2604.24583v1 Announce Type: new Abstract: Recent advancements in reinforcement learning with verifiable rewards (RLVR) have significantly improved the complex reasoning ability of vision-languag

Infrastructure-Guided Connectivity-Enhanced Road Crack Detection and Estimation

ResearchDGX agent

arXiv:2604.24616v1 Announce Type: new Abstract: In this paper, we report the world's first infrastructure-guided communication-enhanced road crack detection pipeline that is effective and implementabl

Insert In Style: A Zero-Shot Generative Framework for Harmonious Cross-Domain Object Composition

Model ReleasesDGX agent

arXiv:2511.15197v2 Announce Type: replace Abstract: Reference-based object composition involves integrating foreground reference image with background scene to produce harmonious fused image. This tas

INSIGHT: Indoor Scene Intelligence from Geometric-Semantic Hierarchy Transfer for Public~Safety

SafetyDGX agent

arXiv:2604.23095v1 Announce Type: new Abstract: Indoor environments lack the spatial intelligence infrastructure that GPS provides outdoors; first responders arriving at unfamiliar buildings typically

Instance Awareness of Multi-class Semantic Segmentation Loss Functions

ResearchDGX agent

arXiv:2604.24276v1 Announce Type: new Abstract: Instance-sensitive losses for semantic segmentation such as blob loss and CC loss were designed to address instance imbalance, ensuring small lesions ge

Instruction-Free Tuning of Large Vision Language Models for Medical Instruction Following

ResearchDGX agent

arXiv:2603.19482v2 Announce Type: replace Abstract: Large vision language models (LVLMs) have demonstrated impressive performance across a wide range of tasks. These capabilities largely stem from vis

IoT-Enhanced CNN-Based Labelled Crack Detection for Additive Manufacturing Image Annotation in Industry 4.0

ApplicationsDGX agent

arXiv:2604.22857v1 Announce Type: new Abstract: This paper presents an IoT-enhanced deep learning framework for automated crack detection in Additive Manufacturing (AM) surfaces using convolutional ne

ISExplore:Informative Segment Selection for Efficient Personalized 3D Talking Face Generation

ResearchDGX agent

arXiv:2511.07940v2 Announce Type: replace Abstract: Talking Face Generation (TFG) methods based on Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have recently achieved impressive prog

iWatchRoad: Scalable Detection and Geospatial Visualization of Potholes for Smart Cities

SafetyDGX agent

arXiv:2508.10945v2 Announce Type: replace Abstract: Potholes on the roads are a serious hazard and maintenance burden. This poses a significant threat to road safety and vehicle longevity, especially

JSSFF: A Joint Structural-Semantic Fusion Framework for Remote Sensing Image Captioning

SafetyDGX agent

arXiv:2604.24031v1 Announce Type: new Abstract: The encoder-decoder framework has become widely popular nowadays. In this model, the encoder extracts informative visual features from an input image, a

KAConvNet: Kolmogorov-Arnold Convolutional Networks for Vision Recognition

SafetyDGX agent

arXiv:2604.23320v1 Announce Type: new Abstract: The Convolutional Neural Networks (CNNs) have been the dominant and effective approach for general computer vision tasks. Recently, Kolmogorov-Arnold ne

Keypoint-based Dynamic Object 6-DoF Pose Tracking via Event Camera

ResearchDGX agent

arXiv:2604.23387v1 Announce Type: new Abstract: Accurate 6-DoF pose estimation of objects is critical for robots to perform precise manipulation tasks. However, for dynamic object pose estimation, con

Latent Inter-Frame Pruning: A Training-Free Method Bridging Traditional Video Compression and Modern Diffusion Transformers for Efficient Generation

HardwareDGX agent

arXiv:2604.23858v1 Announce Type: new Abstract: Video generation, while capable of generating realistic videos, is computationally expensive and slow, prohibiting real-time applications. In this paper

LatentBurst: A Fast and Efficient Multi Frame Super-Resolution for Hexadeca-Bayer Pattern CIS images

ResearchDGX agent

arXiv:2604.23268v1 Announce Type: new Abstract: This paper introduces a novel multi frame super-resolution network (MFSR) for burst hexadeca Bayer pattern Contact Image Sensor (CIS) images, which incl

LatentStealth: Unnoticeable and Efficient Adversarial Attacks on Expressive Human Pose and Shape Estimation

ApplicationsDGX agent

arXiv:2505.12009v2 Announce Type: replace Abstract: Expressive human pose and shape estimation (EHPS) plays a central role in digital human generation, particularly in live-streaming applications. How

LAVA: Layered Audio-Visual Anti-tampering Watermarking for Robust Deepfake Detection and Localization

SafetyDGX agent

arXiv:2604.23957v1 Announce Type: new Abstract: Proactive watermarking offers a promising approach for deepfake tamper detection and localization in short-form videos. However, existing methods often

Learning an Image Editing Model without Image Editing Pairs

ResearchDGX agent

arXiv:2510.14978v2 Announce Type: replace Abstract: Recent image editing models have achieved impressive results while following natural language editing instructions, but they rely on supervised fine

Learning Binary Sampling Patterns for Single-Pixel Imaging using Bilevel Optimisation

TutorialsDGX agent

arXiv:2508.19068v2 Announce Type: replace Abstract: Single-Pixel Imaging (SPI) enables the reconstruction of objects using a single detector through sequential illuminations with structured light patt

Learning from Imperfect Text Guidance: Robust Long-Tail Visual Recognition with High-Noise Label

SafetyDGX agent

arXiv:2604.23125v1 Announce Type: new Abstract: Real-world data often exhibit long-tailed distributions with numerous noisy labels, substantially degrading the performance of deep models. While prior

Learning from Noisy Prompts: Saliency-Guided Prompt Distillation for Robust Segmentation with SAM

Local AiDGX agent

arXiv:2604.23314v1 Announce Type: new Abstract: Segmentation is central to clinical diagnosis and monitoring, yet the reliability of modern foundation models in medical imaging still depends on the av

Learning Scene-Level Signed Directional Distance Function with Ellipsoidal Priors and Neural Residuals

TutorialsDGX agent

arXiv:2503.20066v2 Announce Type: replace-cross Abstract: Dense reconstruction and differentiable rendering are fundamental tightly connected operations in 3D vision and computer graphics. Recent neur

Learning to Decipher from Pixels -- A Case Study of Copiale

ApplicationsDGX agent

arXiv:2604.23683v1 Announce Type: new Abstract: Historical encrypted manuscripts require both paleographic interpretation of cipher symbols and cryptanalytic recovery of plaintext. Most existing compu

Learning to Identify Out-of-Distribution Objects for 3D LiDAR Anomaly Segmentation

AgentsDGX agent

arXiv:2604.23604v1 Announce Type: new Abstract: Understanding the surrounding environment is fundamental in autonomous driving and robotic perception. Distinguishing between known classes and previous

Learning Under Low Illumination: A Dataset and Algorithm for Traffic Sign Recognition

Model ReleasesDGX agent

arXiv:2511.17183v2 Announce Type: replace Abstract: Traffic signboards are vital for road safety and intelligent transportation systems, enabling navigation and autonomous driving. Yet, recognizing tr

LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models

SafetyDGX agent

arXiv:2604.23950v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have recently demonstrated remarkable capabilities in visual understanding and reasoning, but they also impose significant

Less is More in Semantic Space: Intrinsic Decoupling via Clifford-M for Fundus Image Classification

SafetyDGX agent

arXiv:2603.20806v2 Announce Type: replace Abstract: Multi-label fundus diagnosis requires features that capture both fine-grained lesions and large-scale retinal structure. Many multi-scale medical vi

Leveraging Spatial Transcriptomics as Alternative to Manual Annotations for Deep Learning-Based Nuclei Analysis

ResearchDGX agent

arXiv:2604.23481v1 Announce Type: new Abstract: Deep learning-based nuclei segmentation and classification in pathology images typically rely on large-scale pixel-level manual annotations, which are c

← Previous
1…169170171172173…209
Next →