AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
16 Apr 2026

Hybrid Approach for Enhancing Lesion Segmentation in Fundus Images

ResearchDGX agent

arXiv:2509.25549v2 Announce Type: replace Abstract: Choroidal nevi are common benign pigmented lesions in the eye, with a small risk of transforming into melanoma. Early detection is critical to impro

Hydra: Unifying Document Retrieval and Generation in a Single Vision-Language Model

HardwareDGX agent

arXiv:2603.28554v2 Announce Type: replace Abstract: Visual document understanding typically requires separate retrieval and generation models, doubling memory and system complexity. We present Hydra,

Learning Class Difficulty in Imbalanced Histopathology Segmentation via Dynamic Focal Attention

SafetyDGX agent

arXiv:2604.13479v1 Announce Type: cross Abstract: Semantic segmentation of histopathology images under class imbalance is typically addressed through frequency-based loss reweighting, which implicitly


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Learning Sewing Patterns via Latent Flow Matching of Implicit Fields

ResearchDGX agent

arXiv:2601.17740v2 Announce Type: replace Abstract: Sewing patterns define the structural foundation of garments and are essential for applications such as fashion design, fabrication, and physical si

Lite Any Stereo: Efficient Zero-Shot Stereo Matching

ApplicationsDGX agent

arXiv:2511.16555v3 Announce Type: replace Abstract: Recent advances in stereo matching have focused on accuracy, often at the cost of significantly increased model size. Traditionally, the community h

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding

Model ReleasesDGX agent

arXiv:2602.20913v2 Announce Type: replace Abstract: This paper addresses the critical and underexplored challenge of long video understanding with low computational budgets. We propose LongVideo-R1, a

MaMe & MaRe: Matrix-Based Token Merging and Restoration for Efficient Visual Perception and Synthesis

HardwareDGX agent

arXiv:2604.13432v1 Announce Type: new Abstract: Token compression is crucial for mitigating the quadratic complexity of self-attention mechanisms in Vision Transformers (ViTs), which often involve num

MApLe: Multi-instance Alignment of Diagnostic Reports and Large Medical Images

Local AiDGX agent

arXiv:2604.13970v1 Announce Type: new Abstract: In diagnostic reports, experts encode complex imaging data into clinically actionable information. They describe subtle pathological findings that are m

Med-CAM: Minimal Evidence for Explaining Medical Decision Making

SafetyDGX agent

arXiv:2604.13695v1 Announce Type: new Abstract: Reliable and interpretable decision-making is essential in medical imaging, where diagnostic outcomes directly influence patient care. Despite advances

MSGS: Multispectral 3D Gaussian Splatting

ApplicationsDGX agent

arXiv:2604.13340v1 Announce Type: new Abstract: We present a multispectral extension to 3D Gaussian Splatting (3DGS) for wavelength-aware view synthesis. Each Gaussian is augmented with spectral radia

Multi-Agent Object Detection Framework Based on Raspberry Pi YOLO Detector and Slack-Ollama Natural Language Interface

Local AiDGX agent

arXiv:2604.13345v1 Announce Type: new Abstract: The paper presents design and prototype implementation of an edge based object detection system within the new paradigm of AI agents orchestration. It g

Multi-Dimensional Knowledge Profiling with Large-Scale Literature Database and Hierarchical Retrieval

SafetyDGX agent

arXiv:2601.15170v2 Announce Type: replace Abstract: The rapid expansion of research across machine learning, vision, and language has produced a volume of publications that is increasingly difficult t

Multi-modal panoramic 3D outdoor datasets for place categorization

ResearchDGX agent

arXiv:2604.13142v1 Announce Type: cross Abstract: We present two multi-modal panoramic 3D outdoor (MPO) datasets for semantic place categorization with six categories: forest, coast, residential area,

Multitasking Embedding for Embryo Blastocyst Grading Prediction (MEmEBG)

TutorialsDGX agent

arXiv:2604.13217v1 Announce Type: new Abstract: Reliable evaluation of blastocyst quality is critical for the success of in vitro fertilization (IVF) treatments. Current embryo grading practices prima

MyoVision: A Mobile Research Tool and NEATBoost-Attention Ensemble Framework for Real Time Chicken Breast Myopathy Detection

ResearchDGX agent

arXiv:2604.13456v1 Announce Type: cross Abstract: Woody Breast (WB) and Spaghetti Meat (SM) myopathies significantly impact poultry meat quality, yet current detection methods rely either on subjectiv

Neural 3D Reconstruction of Planetary Surfaces from Descent-Phase Wide-Angle Imagery

ResearchDGX agent

arXiv:2604.13235v1 Announce Type: new Abstract: Digital elevation modeling of planetary surfaces is essential for studying past and ongoing geological processes. Wide-angle imagery acquired during spa

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding

Local AiDGX agent

arXiv:2604.14149v1 Announce Type: new Abstract: Long video understanding is inherently challenging for vision-language models (VLMs) because of the extensive number of frames. With each video frame ty

OneHOI: Unifying Human-Object Interaction Generation and Editing

ResearchDGX agent

arXiv:2604.14062v1 Announce Type: new Abstract: Human-Object Interaction (HOI) modelling captures how humans act upon and relate to objects, typically expressed as triplets. Existing approaches split

OPTED: Open Preprocessed Trachoma Eye Dataset Using Zero-Shot SAM 3 Segmentation

Model ReleasesDGX agent

arXiv:2603.06885v2 Announce Type: replace Abstract: Trachoma remains the leading infectious cause of blindness worldwide, with Sub-Saharan Africa bearing over 85% of the global burden and Ethiopia alo

PartNerFace: Part-based Neural Radiance Fields for Animatable Facial Avatar Reconstruction

TutorialsDGX agent

arXiv:2604.13918v1 Announce Type: new Abstract: We present PartNerFace, a part-based neural radiance fields approach, for reconstructing animatable facial avatar from monocular RGB videos. Existing so

PAT-VCM: Plug-and-Play Auxiliary Tokens for Video Coding for Machines

ResearchDGX agent

arXiv:2604.13294v1 Announce Type: new Abstract: Existing video coding for machines is often trained for a specific downstream task and model. As a result, the compressed representation becomes tightly

PatchPoison: Poisoning Multi-View Datasets to Degrade 3D Reconstruction

Model ReleasesDGX agent

arXiv:2604.13153v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has recently enabled highly photorealistic 3D reconstruction from casually captured multi-view images. However, this access

PBE-UNet: A light weight Progressive Boundary-Enhanced U-Net with Scale-Aware Aggregation for Ultrasound Image Segmentation

Model ReleasesDGX agent

arXiv:2604.13791v1 Announce Type: new Abstract: Accurate lesion segmentation in ultrasound images is essential for preventive screening and clinical diagnosis, yet remains challenging due to low contr

Person Re-Identification via Generalized Class Prototypes

ResearchDGX agent

arXiv:2510.17043v2 Announce Type: replace Abstract: Advanced feature extraction methods have significantly contributed to enhancing the task of person re-identification. In addition, modifications to

Physically-Guided Optical Inversion Enable Non-Contact Side-Channel Attack on Isolated Screens

ResearchDGX agent

arXiv:2604.13419v1 Announce Type: new Abstract: Noncontact exfiltration of electronic screen content poses a security challenge, with side-channel incursions as the principal vector. We introduce an o

POINTS-Seeker: Towards Training a Multimodal Agentic Search Model from Scratch

AgentsDGX agent

arXiv:2604.14029v1 Announce Type: new Abstract: While Large Multimodal Models (LMMs) demonstrate impressive visual perception, they remain epistemically constrained by their static parametric knowledg

PostureObjectstitch: Anomaly Image Generation Considering Assembly Relationships in Industrial Scenarios

ApplicationsDGX agent

arXiv:2604.13863v1 Announce Type: new Abstract: Image generation technology can synthesize condition-specific images to supplement real-world industrial anomaly data and enhance anomaly detection mode

Radar-Informed 3D Multi-Object Tracking under Adverse Conditions

ApplicationsDGX agent

arXiv:2604.13571v1 Announce Type: new Abstract: The challenge of 3D multi-object tracking (3D MOT) is achieving robustness in real-world applications, for example under adverse conditions and maintain

RadarSplat-RIO: Indoor Radar-Inertial Odometry with Gaussian Splatting-Based Radar Bundle Adjustment

ResearchDGX agent

arXiv:2604.13492v1 Announce Type: cross Abstract: Radar is more resilient to adverse weather and lighting conditions than visual and Lidar simultaneous localization and mapping (SLAM). However, most r

Reconstruction of a 3D wireframe from a single line drawing via generative depth estimation

ResearchDGX agent

arXiv:2604.13549v1 Announce Type: new Abstract: The conversion of 2D freehand sketches into 3D models remains a pivotal challenge in computer vision, bridging the gap between human creativity and digi

ReConText3D: Replay-based Continual Text-to-3D Generation

Model ReleasesDGX agent

arXiv:2604.13730v1 Announce Type: new Abstract: Continual learning enables models to acquire new knowledge over time while retaining previously learned capabilities. However, its application to text-t

Remote Sensing Image Super-Resolution for Imbalanced Textures: A Texture-Aware Diffusion Framework

Local AiDGX agent

arXiv:2604.13994v1 Announce Type: new Abstract: Generative diffusion priors have recently achieved state-of-the-art performance in natural image super-resolution, demonstrating a powerful capability t

Rethinking Image-to-3D Generation with Sparse Queries: Efficiency, Capacity, and Input-View Bias

Local AiDGX agent

arXiv:2604.13905v1 Announce Type: new Abstract: We present SparseGen, a novel framework for efficient image-to-3D generation, which exhibits low input-view bias while being significantly faster. Unlik

Rethinking Uncertainty in Segmentation: From Estimation to Decision

SafetyDGX agent

arXiv:2604.13262v1 Announce Type: new Abstract: In medical image segmentation, uncertainty estimates are often reported but rarely used to guide decisions. We study the missing step: how uncertainty m

Right Regions, Wrong Labels: Semantic Label Flips in Segmentation under Correlation Shift

ResearchDGX agent

arXiv:2604.13326v1 Announce Type: new Abstract: The robustness of machine learning models can be compromised by spurious correlations between non-causal features in the input data and target labels. A

RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph

SafetyDGX agent

arXiv:2511.07717v2 Announce Type: replace-cross Abstract: Estimating robot pose from a monocular RGB image is a challenge in robotics and computer vision. Existing methods typically build networks on

RobotPan: A 360^irc Surround-View Robotic Vision System for Embodied Perception

ResearchDGX agent

arXiv:2604.13476v1 Announce Type: cross Abstract: Surround-view perception is increasingly important for robotic navigation and loco-manipulation, especially in human-in-the-loop settings such as tele

ROSE: Retrieval-Oriented Segmentation Enhancement

Model ReleasesDGX agent

arXiv:2604.14147v1 Announce Type: new Abstract: Existing segmentation models based on multimodal large language models (MLLMs), such as LISA, often struggle with novel or emerging entities due to thei

SceneGlue: Scene-Aware Transformer for Feature Matching without Scene-Level Annotation

ResearchDGX agent

arXiv:2604.13941v1 Announce Type: new Abstract: Local feature matching plays a critical role in understanding the correspondence between cross-view images. However, traditional methods are constrained

SEDTalker: Emotion-Aware 3D Facial Animation Using Frame-Level Speech Emotion Diarization

ResearchDGX agent

arXiv:2604.13335v1 Announce Type: new Abstract: We introduce SEDTalker, an emotion-aware framework for speech-driven 3D facial animation that leverages frame-level speech emotion diarization to achiev

Seedance 2.0: Advancing Video Generation for World Complexity

Model ReleasesDGX agent

arXiv:2604.14148v1 Announce Type: new Abstract: Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecesso

Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios

Model ReleasesDGX agent

arXiv:2604.14041v1 Announce Type: new Abstract: Daily scenarios are characterized by visual richness, requiring Multimodal Large Language Models (MLLMs) to filter noise and identify decisive visual cl

See&Say: Vision Language Guided Safe Zone Detection for Autonomous Package Delivery Drones

SafetyDGX agent

arXiv:2604.13292v1 Announce Type: new Abstract: Autonomous drone delivery systems are rapidly advancing, but ensuring safe and reliable package drop-offs remains highly challenging in cluttered urban

SemAttNet: Towards Attention-based Semantic Aware Guided Depth Completion

Model ReleasesDGX agent

arXiv:2204.13635v2 Announce Type: replace Abstract: Depth completion involves recovering a dense depth map from a sparse map and an RGB image. Recent approaches focus on utilizing color images as guid

SemiFA: An Agentic Multi-Modal Framework for Autonomous Semiconductor Failure Analysis Report Generation

HardwareDGX agent

arXiv:2604.13236v1 Announce Type: new Abstract: Semiconductor failure analysis (FA) requires engineers to examine inspection images, correlate equipment telemetry, consult historical defect records, a

SiLVR: A Simple Language-based Video Reasoning Framework

ResearchDGX agent

arXiv:2505.24869v3 Announce Type: replace Abstract: Recent advances in test-time optimization have led to remarkable reasoning capabilities in Large Language Models (LLMs), enabling them to solve high

Simplicity Prevails: The Emergence of Generalizable AIGI Detection in Visual Foundation Models

Local AiDGX agent

arXiv:2602.01738v2 Announce Type: replace Abstract: While specialized detectors for AI-Generated Images (AIGI) achieve near-perfect accuracy on curated benchmarks, they suffer from a dramatic performa

SLQ: Bridging Modalities via Shared Latent Queries for Retrieval with Frozen MLLMs

Model ReleasesDGX agent

arXiv:2604.13710v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exhibit strong reasoning and world knowledge, yet adapting them for retrieval remains challenging. Existing app

SocialMirror: Reconstructing 3D Human Interaction Behaviors from Monocular Videos with Semantic and Geometric Guidance

Model ReleasesDGX agent

arXiv:2604.13581v1 Announce Type: new Abstract: Accurately reconstructing human behavior in close-interaction scenarios is crucial for enabling realistic virtual interactions in augmented reality, pre

SSD-GS: Scattering and Shadow Decomposition for Relightable 3D Gaussian Splatting

ResearchDGX agent

arXiv:2604.13333v1 Announce Type: new Abstract: We present SSD-GS, a physically-based relighting framework built upon 3D Gaussian Splatting (3DGS) that achieves high-quality reconstruction and photore

Target-Bench: Can Video World Models Achieve Mapless Path Planning with Semantic Targets?

Model ReleasesDGX agent

arXiv:2511.17792v2 Announce Type: replace Abstract: While recent video world models can generate highly realistic videos, their ability to perform semantic reasoning and planning remains unclear and u

Temporally Consistent Long-Term Memory for 3D Single Object Tracking

SafetyDGX agent

arXiv:2604.13789v1 Announce Type: new Abstract: 3D Single Object Tracking (3D-SOT) aims to localize a target object across a sequence of LiDAR point clouds, given its 3D bounding box in the first fram

The Gaussian Latent Machine: Efficient Prior and Posterior Sampling for Inverse Problems

ResearchDGX agent

arXiv:2505.12836v2 Announce Type: replace-cross Abstract: We consider the problem of sampling from a product-of-experts-type model that encompasses many standard prior and posterior distributions comm

The Spectrascapes Dataset: Street-view imagery beyond the visible captured using a mobile platform

ResearchDGX agent

arXiv:2604.13315v1 Announce Type: new Abstract: High-resolution data in spatial and temporal contexts is imperative for developing climate resilient cities. Current datasets for monitoring urban param

Tokenizing Semantic Segmentation with Run Length Encoding

TutorialsDGX agent

arXiv:2602.21627v3 Announce Type: replace Abstract: This paper presents a new unified approach to semantic segmentation in both images and videos by using language modeling to output the masks as sequ

Towards Generalizable Robotic Manipulation in Dynamic Environments

Model ReleasesDGX agent

arXiv:2603.15620v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models excel in static manipulation but struggle in dynamic environments with moving targets. This performance gap prim

Towards Multi-Object-Tracking with Radar on a Fast Moving Vehicle: On the Potential of Processing Radar in the Frequency Domain

AgentsDGX agent

arXiv:2604.14013v1 Announce Type: cross Abstract: We promote in this paper the processing of radar data in the frequency domain to achieve higher robustness against noise and structural errors, especi

Towards Patient-Specific Deformable Registration in Laparoscopic Surgery

ResearchDGX agent

arXiv:2604.13186v1 Announce Type: new Abstract: Unsafe surgical care is a critical health concern, often linked to limitations in surgeon experience, skills, and situational awareness. Integrating pat

Towards Successful Implementation of Automated Raveling Detection: Effects of Training Data Size, Illumination Difference, and Spatial Shift

Model ReleasesDGX agent

arXiv:2604.13322v1 Announce Type: new Abstract: Raveling, the loss of aggregates, is a major form of asphalt pavement surface distress, especially on highways. While research has shown that machine le

Towards Unconstrained Human-Object Interaction

ResearchDGX agent

arXiv:2604.14069v1 Announce Type: new Abstract: Human-Object Interaction (HOI) detection is a longstanding computer vision problem concerned with predicting the interaction between humans and objects.

← Previous
1…190191192193194…207
Next →