AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
1 May 2026

InterPartAbility: Text-Guided Part Matching for Interpretable Person Re-Identification

ResearchDGX agent

arXiv:2604.27122v1 Announce Type: new Abstract: Text-to-image person re-identification (TI-ReID) relies on natural-language text description to retrieve top matching individuals from a large gallery o

Iterative Definition Refinement for Zero-Shot Classification via LLM-Based Semantic Prototype Optimization

Model ReleasesDGX agent

arXiv:2604.27335v1 Announce Type: new Abstract: Web filtering systems rely on accurate web content classification to block cyber threats, prevent data exfiltration, and ensure compliance. However, cla

JI-ADF: Joint-Individual Learning with Adaptive Decision Fusion for Multimodal Skin Lesion Classification

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2604.27343v1 Announce Type: new Abstract: Skin lesion classification is essential for early dermatological diagnosis, yet many existing computer-aided systems rely primarily on dermoscopic image

Judge, Then Drive: A Critic-Centric Vision Language Action Framework for Autonomous Driving

Model ReleasesDGX agent

arXiv:2604.27366v1 Announce Type: new Abstract: Recent advances in vision language action (VLA) models have shown remarkable potential for autonomous driving by directly mapping multimodal inputs to c

LA-Pose: Latent Action Pretraining Meets Pose Estimation

SafetyDGX agent

arXiv:2604.27448v1 Announce Type: new Abstract: This paper revisits camera pose estimation through the lens of self-supervised pretraining, focusing on inverse-dynamics pretraining as a scalable alter

LaST-R1: Reinforcing Action via Adaptive Physical Latent Reasoning for VLA Models

Model ReleasesDGX agent

arXiv:2604.28192v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have increasingly incorporated reasoning mechanisms for complex robotic manipulation. However, existing approaches

Leveraging Quantum-Based Architectures for Robust Diagnostics

ResearchDGX agent

arXiv:2511.12386v2 Announce Type: replace Abstract: Quantum machine learning has emerged as a promising approach for medical image analysis, particularly in settings where compact models and expressiv

Leveraging Verifier-Based Reinforcement Learning in Image Editing

ResearchDGX agent

arXiv:2604.27505v1 Announce Type: new Abstract: While Reinforcement Learning from Human Feedback (RLHF) has become a pivotal paradigm for text-to-image generation, its application to image editing rem

LM-CartSeg: Automated Segmentation of Lateral and Medial Cartilage and Subchondral Bone for Radiomics Analysis

ResearchDGX agent

arXiv:2512.03449v3 Announce Type: replace Abstract: Background and Objective: Radiomics of knee MRI requires robust, anatomically meaningful regions of interest (ROIs) that jointly capture cartilage a

Machine Unlearning for Class Removal through SISA-based Deep Neural Network Architectures

ApplicationsDGX agent

arXiv:2604.27804v1 Announce Type: new Abstract: The rapid proliferation of image generation models and other artificial intelligence (AI) systems has intensified concerns regarding data privacy and us

MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos

ResearchDGX agent

arXiv:2512.10881v2 Announce Type: replace Abstract: Motion capture now underpins content creation far beyond digital humans, yet most existing pipelines remain species- or template-specific. We formal

MoCapAnything V2: End-to-End Motion Capture for Arbitrary Skeletons

Local AiDGX agent

arXiv:2604.28130v1 Announce Type: new Abstract: Recent methods for arbitrary-skeleton motion capture from monocular video follow a factorized pipeline, where a Video-to-Pose network predicts joint pos

MSR:Hybrid Field Modeling for CT-MRI Rigid-Deformable Registration of the Cervical Spine with an Annotated Dataset

SafetyDGX agent

arXiv:2604.27654v1 Announce Type: new Abstract: Accurate CT-MRI registration of the cervical spine is essential for preoperative planning because this region is anatomically complex,highly variable,an

Noise2Map: End-to-End Diffusion Model for Semantic Segmentation and Change Detection

TutorialsDGX agent

arXiv:2604.27889v1 Announce Type: new Abstract: Semantic segmentation and change detection are two fundamental challenges in remote sensing, requiring models to capture either spatial semantics or tem

Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization

TutorialsDGX agent

arXiv:2512.10955v2 Announce Type: replace Abstract: Visual concept personalization aims to transfer only specific image attributes, such as identity, expression, lighting, and style, into unseen conte

OmniRobotHome: A Multi-Camera Platform for Real-Time Multiadic Human-Robot Interaction

SafetyDGX agent

arXiv:2604.28197v1 Announce Type: cross Abstract: Human-robot collaboration has been studied primarily in dyadic or sequential settings. However, real homes require multiadic collaboration, where mult

Parameter-Efficient Architectural Modifications for Translation-Invariant CNNs

Model ReleasesDGX agent

arXiv:2604.27870v1 Announce Type: new Abstract: Convolutional Neural Networks (CNNs) are widely assumed to be translation-invariant, yet standard architectures exhibit a startling fragility: even a si

PhotIQA: A photoacoustic image data set with image quality ratings

ResearchDGX agent

arXiv:2507.03478v2 Announce Type: replace-cross Abstract: Image quality assessment (IQA) is crucial in the evaluation stage of novel algorithms operating on images, including traditional and machine l

Physically-Informed Fuzzy Clustering of Vertical Sounding Ionograms

ResearchDGX agent

arXiv:2604.27721v1 Announce Type: cross Abstract: This paper presents a physically-informed fuzzy clustering of vertical sounding ionograms for automatically separating the ionogram into tracks suitab

PINN-Cast: Exploring the Role of Continuous-Depth NODE in Transformers and Physics Informed Loss as Soft Physical Constraints in Short-term Weather Forecasting

ResearchDGX agent

arXiv:2604.27313v1 Announce Type: cross Abstract: Operational weather prediction has long relied on physics-based numerical weather prediction (NWP), whose accuracy comes at the cost of substantial co

Primus: Enforcing Attention Usage for 3D Medical Image Segmentation

ResearchDGX agent

arXiv:2503.01835v2 Announce Type: replace Abstract: Transformers have achieved remarkable success across multiple fields, yet their impact on 3D medical image segmentation remains limited with convolu

PVeRA: Probabilistic Vector-Based Random Matrix Adaptation

Model ReleasesDGX agent

arXiv:2512.07703v2 Announce Type: replace Abstract: Large foundation models have emerged in the last years and are pushing performance boundaries for a variety of tasks. Training or even finetuning su

RayFormer: Modeling Inter- and Intra-Ray Similarity for NeRF-Based Video Snapshot Compressive Imaging

ApplicationsDGX agent

arXiv:2604.27702v1 Announce Type: new Abstract: Video snapshot compressive imaging (SCI) enables the reconstruction of dynamic scenes from a single snapshot measurement. Recently, NeRF-based methods h

Representation Frechet Loss for Visual Generation

ResearchDGX agent

arXiv:2604.28190v1 Announce Type: new Abstract: We show that Frechet Distance (FD), long considered impractical as a training objective, can in fact be effectively optimized in the representation spac

Representative Spectral Correlation Network for Multi-source Remote Sensing Image Classification

Model ReleasesDGX agent

arXiv:2604.27323v1 Announce Type: cross Abstract: Hyperspectral image (HSI) and SAR/LiDAR data offer complementary spectral and structural information for land-cover classification. However, their eff

Residual Gaussian Splatting for Ultra Sparse-View CBCT Reconstruction

SafetyDGX agent

arXiv:2604.27552v1 Announce Type: new Abstract: While 3D Gaussian splatting (3DGS) offers explicit and efficient scene representations for cone-beam computed tomography reconstruction, conventional ph

ResiHMR: Residual-Limb Aware Single-Image 3D Human Mesh Recovery for Individuals with Limb Loss

ApplicationsDGX agent

arXiv:2604.28025v1 Announce Type: new Abstract: Single-image human mesh recovery provides a compact 3D, person-centric representation that supports analysis, animation, AR and VR, rehabilitation, and

Rethinking Pulmonary Embolism Segmentation: A Study of Current Approaches and Challenges with an Open Weight Model

ResearchDGX agent

arXiv:2509.18308v3 Announce Type: replace Abstract: Pulmonary Embolism (PE) is a life-threatening condition for which accurate and timely detection is critical to patient care. However, our systematic

Revealing the Impact of Visual Text Style on Attribute-based Descriptions Produced by Large Visual Language Models

ResearchDGX agent

arXiv:2604.27553v1 Announce Type: new Abstract: When the visual style of text is considered, a wide variety can be observed in font, color, and size. However, when a word is read, its meaning is indep

REVIVE 3D: Refinement via Encoded Voluminous Inflated prior for Volume Enhancement

ResearchDGX agent

arXiv:2604.27504v1 Announce Type: new Abstract: Recent generative models have shown strong performance in generating diverse 3D assets from 2D images, a fundamental research topic in computer vision a

Robot Learning from Human Videos: A Survey

SafetyDGX agent

arXiv:2604.27621v1 Announce Type: cross Abstract: A critical bottleneck hindering further advancement in embodied AI and robotics is the challenge of scaling robot data. To address this, the field of

Sample-efficient evidence estimation of score based priors for model selection

SafetyDGX agent

arXiv:2602.20549v2 Announce Type: replace-cross Abstract: The choice of prior is central to solving ill-posed imaging inverse problems, making it essential to select one consistent with the measuremen

SECOS: Semantic Capture for Rigorous Classification in Open-World Semi-Supervised Learning

ResearchDGX agent

arXiv:2604.27596v1 Announce Type: new Abstract: In open-world semi-supervised learning (OWSSL), a model learns from labeled data and unlabeled data containing both known and novel classes. In practica

Self-Supervised Learning of Plant Image Representations

ResearchDGX agent

arXiv:2604.27538v1 Announce Type: new Abstract: Automated plant recognition plays a crucial role in biodiversity monitoring and conservation, yet current approaches rely heavily on supervised learning

Softmax-GS: Generalized Gaussians Learning When to Blend or Bound

Model ReleasesDGX agent

arXiv:2604.27437v1 Announce Type: new Abstract: 3D Gaussian Splatting (3D GS) is widely adopted for novel view synthesis due to its high training and rendering efficiency. However, its efficiency reli

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation

AgentsDGX agent

arXiv:2604.27620v1 Announce Type: new Abstract: Vision-and-Language Navigation (VLN) aims to enable an embodied agent to follow natural-language instructions and navigate to a target location in unsee

Sparse-View 3D Gaussian Splatting in the Wild

ApplicationsDGX agent

arXiv:2604.27422v1 Announce Type: new Abstract: We propose a 3D novel sparse-view synthesis framework for unconstrained real-world scenarios that contain distractors. Unlike existing methods that prim

Spectral Dynamic Attention Network for Hyperspectral Image Super-Resolution

Model ReleasesDGX agent

arXiv:2604.27326v1 Announce Type: cross Abstract: Hyperspectral image super-resolution is essential for enhancing the spatial fidelity of HSI data, yet existing deep learning methods often struggle wi

SQuadGen: Generating Simple Quad Layouts via Chart Distance Fields

ResearchDGX agent

arXiv:2604.27329v1 Announce Type: cross Abstract: 3D shapes from scanning, reconstruction, or AI-generated content often lack simple quad mesh layouts -- critical for efficient editing and modeling. E

Stop Holding Your Breath: CT-Informed Gaussian Splatting for Dynamic Bronchoscopy

ResearchDGX agent

arXiv:2604.28179v1 Announce Type: new Abstract: Bronchoscopic navigation relies on registering endoscopic video to a preoperative CT scan, but respiratory motion deforms the airway by 5-20 mm, creatin

Student Classroom Behavior Recognition Based on Improved YOLOv8s

ResearchDGX agent

arXiv:2604.27293v1 Announce Type: new Abstract: In classroom teaching, student behavior can reflect their learning state and classroom participation, which is of great significance for teaching qualit

TAFA-GSGC: Group-wise Scalable Point Cloud Geometry Compression with Progressive Residual Refinement

ResearchDGX agent

arXiv:2604.28045v1 Announce Type: new Abstract: Scalable compression is essential for bandwidth-adaptive transmission, yet most learned codecs are optimized for a fixed rate-distortion point, making r

Taming Noise-Induced Prototype Degradation for Privacy-Preserving Personalized Federated Fine-Tuning

ResearchDGX agent

arXiv:2604.27833v1 Announce Type: new Abstract: Prototype-based Personalized Federated Learning (ProtoPFL) enables efficient multi-domain adaptation by communicating compact class prototypes, but dire

TeD-Loc: Text Distillation for Weakly Supervised Object Localization

Local AiDGX agent

arXiv:2501.12632v2 Announce Type: replace Abstract: Weakly supervised object localization (WSOL) models are trained using only image-level class labels. They can predict both the object class and spat

Test-Time Distillation for Continual Model Adaptation

SafetyDGX agent

arXiv:2506.02671v3 Announce Type: replace Abstract: Deep neural networks often suffer performance degradation upon deployment due to distribution shifts. Continual Test-Time Adaptation (CTTA) aims to

Towards All-Day Perception for Off-Road Driving: A Large-Scale Multispectral Dataset and Comprehensive Benchmark

Model ReleasesDGX agent

arXiv:2604.27499v1 Announce Type: new Abstract: Off-road nighttime autonomous driving suffers from unreliable visible-light perception, making infrared modality crucial for accurate freespace detectio

Towards Generalizable Mapping of Hedges and Linear Woody Features from Earth Observation Data: a national Product for Germany

ResearchDGX agent

arXiv:2604.27247v1 Announce Type: new Abstract: Hedges and other linear woody features provide valuable ecosystem services, particularly within intensively managed agricultural landscapes. They are ke

TranSplat: Instant Object Relighting in Gaussian Splatting via Spherical Harmonic Radiance Transfer

ApplicationsDGX agent

arXiv:2503.22676v5 Announce Type: replace Abstract: We present TranSplat, a method for instant, accurate object relighting within the Gaussian Splatting (GS) framework. Rather than relying on costly i

TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On

Model ReleasesDGX agent

arXiv:2604.27958v1 Announce Type: new Abstract: Due to the scarcity of large-scale in-the-wild triplet data and the improper use of masks, the performance of video virtual try-on models remains limite

UHR-Net: An Uncertainty-Aware Hypergraph Refinement Network for Medical Image Segmentation

TutorialsDGX agent

arXiv:2604.28095v1 Announce Type: new Abstract: Accurate lesion segmentation is crucial for clinical diagnosis and treatment planning. However, lesions often resemble surrounding tissues and exhibit i

Uncertainty Quantification Framework for Aerial and UAV Photogrammetry through Error Propagation

ResearchDGX agent

arXiv:2507.13486v2 Announce Type: replace Abstract: Uncertainty quantification of the photogrammetry process is essential for providing per-point accuracy credentials of the point clouds. Unlike airbo

Understanding Adversarial Transferability in Vision-Language Models for Autonomous Driving: A Cross-Architecture Analysis

AgentsDGX agent

arXiv:2604.27414v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used in autonomous driving because they combine visual perception with language-based reasoning, supporti

Uni-HOI:A Unified framework for Learning the Joint distribution of Text and Human-Object Interaction

ResearchDGX agent

arXiv:2604.27491v1 Announce Type: new Abstract: Modeling 4D human-object interaction (HOI) is a compelling challenge in computer vision and an essential technology powering virtual and mixed-reality a

Unsupervised Machine Learning for Osteoporosis Diagnosis Using Singh Index Clustering on Hip Radiographs

ResearchDGX agent

arXiv:2411.15253v2 Announce Type: replace-cross Abstract: Osteoporosis, a prevalent condition among the aging population worldwide, is characterized by diminished bone mass and altered bone structure,

VeraRetouch: A Lightweight Fully Differentiable Framework for Multi-Task Reasoning Photo Retouching

Model ReleasesDGX agent

arXiv:2604.27375v1 Announce Type: new Abstract: Reasoning photo retouching has gained significant traction, requiring models to analyze image defects, give reasoning processes, and execute precise ret

VerteNet -- A Multi-Context Hybrid CNN Transformer for Accurate Vertebral Landmark Localization in Lateral Spine DXA Images

Local AiDGX agent

arXiv:2502.02097v3 Announce Type: replace Abstract: This aims to develop and validate a deep learning model that can accurately locate vertebral landmarks in lateral spine Dual energy X-ray Absorptiom

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

Model ReleasesDGX agent

arXiv:2604.28185v1 Announce Type: new Abstract: Recent visual generation models have made major progress in photorealism, typography, instruction following, and interactive editing, yet they still str

VTBench: A Multimodal Framework for Time-Series Classification with Chart-Based Representations

ResearchDGX agent

arXiv:2604.27259v1 Announce Type: new Abstract: Time-series classification (TSC) has advanced significantly with deep learning, yet most models rely solely on raw numerical inputs, overlooking alterna

World2Minecraft: Occupancy-Driven Simulated Scenes Construction

ApplicationsDGX agent

arXiv:2604.27578v1 Announce Type: new Abstract: Embodied intelligence requires high-fidelity simulation environments to support perception and decision-making, yet existing platforms often suffer from

YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal

ResearchDGX agent

arXiv:2604.27322v1 Announce Type: new Abstract: Recent advances in Diffusion Transformer (DiT)-based video generation technologies have shown impressive results for video object removal. However, thes

← Previous
1…163164165166167…209
Next →