AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
25 May 2026

Efficient Learned Image Compression without Entropy Coding

ResearchDGX agent

arXiv:2605.23323v1 Announce Type: cross Abstract: Entropy coding is widely used in typical learned image compression (LIC) that converts latents into a compact bitstream. However, entropy coding is ty

Efficient One-Step Diffusion Restoration Model with Compact Token Compression and Linear Attention

Model ReleasesDGX agent

arXiv:2605.23451v1 Announce Type: new Abstract: Real-world image super-resolution aims to recover high-quality images from complex and unknown real-world degradations. However, existing generative Rea

Enhancing 3D Semantic Scene Completion with a Refinement Module

ResearchDGX agent

arXiv:2512.18363v2 Announce Type: replace Abstract: We propose ESSC-RM, a plug-and-play Enhancing framework for Semantic Scene Completion with a Refinement Module, which can be seamlessly integrated i


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Enhancing Blood Cells Classification using Hybrid Quantum Neural Networks

ResearchDGX agent

arXiv:2605.23324v1 Announce Type: new Abstract: Accurate classification of microscopic blood cells is still a critical task in medical image analysis, where subtle variations and limited data can chal

Exploring deep learning for Event-Based Saliency Prediction with a Transformer-based model

Model ReleasesDGX agent

arXiv:2605.23790v1 Announce Type: new Abstract: Saliency prediction has been extensively studied in RGB images and videos as a computational model of human visual attention. In contrast, predicting sa

ExpOS: Explainable Open-Surgery Skills Assessment Using 3D Hand Reconstruction

AgentsDGX agent

arXiv:2605.23653v1 Announce Type: new Abstract: Timely and transparent feedback is essential for effective surgical training, yet current assessment remains dependent on expert observation, limiting s

Extending Deep Event Visual Odometry with Sparse Point-Cloud Export

ResearchDGX agent

arXiv:2605.22890v1 Announce Type: cross Abstract: Event cameras are well suited for visual odometry under high-speed motion and challenging lighting conditions due to their low latency, high temporal

FAST-ME: Foundation-aware Adaptive Stopping for Motion Estimation for Efficient IoT Video Analysis

Model ReleasesDGX agent

arXiv:2605.23428v1 Announce Type: new Abstract: In modern multimedia systems, efficient video processing is critical, especially in resource-constrained environments such as IoT-based camera networks,

Flow Mismatching: Unsupervised Anomaly Detection via Velocity Discrepancies in Flow Matching Models

ResearchDGX agent

arXiv:2605.23070v1 Announce Type: new Abstract: We propose Flow Mismatching, an unsupervised anomaly detection method that deliberately avoids reconstruction-based paradigms. Instead, we treat flow ma

From Activation to Causality: Discovery of Causal Visual Representations in the Human Brain

Local AiDGX agent

arXiv:2605.23895v1 Announce Type: new Abstract: Identifying which brain regions represent a visual concept in the human brain is a central challenge in neuroscience. Existing approaches have localized

GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation

TutorialsDGX agent

arXiv:2506.14135v5 Announce Type: replace-cross Abstract: Accurate scene perception is critical for vision-based robotic manipulation. Existing approaches typically follow either a Vision-to-Action (V

GazeBehavior Annotation Toolkit (GBAT): AI-powered toolkit for automatic annotation of egocentric eye-tracking and video data of child-caregiver interaction

ResearchDGX agent

arXiv:2605.22962v1 Announce Type: new Abstract: Video recordings of child-caregiver interactions enable investigation of attentional dynamics during naturalistic behavior. Such multimodal recording al

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation

ApplicationsDGX agent

arXiv:2605.22882v1 Announce Type: new Abstract: Video world models can generate realistic futures from a single instruction, but they often fail to preserve consistent point-level motion over time. As

General Hazard Detection

SafetyDGX agent

arXiv:2605.23304v1 Announce Type: new Abstract: Hazard, as an abstract concept, is typically defined through cognitive-level logical reasoning rather than concrete examples. In contrast, existing haza

Generator-Refiner-Examiner: A Tri-Module Data Augmentation Framework for 3D Human Avatar Learning from Monocular Videos

ResearchDGX agent

arXiv:2605.23555v1 Announce Type: new Abstract: This paper addresses the challenge of reconstructing photorealistic and animatable 3D human avatars from monocular videos. While existing methods rely o

GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction

Local AiDGX agent

arXiv:2605.23888v1 Announce Type: new Abstract: We introduce a new approach to high-fidelity 3D scene reconstruction from multi-view RGB images that tightly couples reconstruction with a strong genera

Geo-Align: Video Generation Alignment via Metric Geometry Reward

SafetyDGX agent

arXiv:2605.23903v1 Announce Type: new Abstract: Camera-controlled video generation has achieved remarkable progress in recent years. However, existing video-to-video re-rendering methods primarily rel

GFSR: Geometric Fidelity and Spatial Refinement for Reliable Lane Detection

AgentsDGX agent

arXiv:2605.23327v1 Announce Type: new Abstract: Lane detection stands as a crucial perception task in autonomous driving and advanced driver assistance systems. However, existing methods still degrade

GlowGS: Generative Semantic Feature Learning for 3D Gaussian Splatting in Nighttime Glow Scenes

ResearchDGX agent

arXiv:2605.23602v1 Announce Type: new Abstract: Existing 3DGS methods effectively render high-quality novel views in clear-day scenes. However, they struggle with night scenes, particularly in glow re

GMENet: Generative Mixture of Experts Network for Multi-Center Glioma Diagnosis with Incomplete Imaging Sequences

TutorialsDGX agent

arXiv:2605.23183v1 Announce Type: cross Abstract: Contemporary glioma diagnosis integrates molecular features with histopathology to guide clinical decision-making. However, in clinical settings, dive

GT-SVJ: Generative-Transformer-Based Self-Supervised Video Judge For Efficient Video Reward Modeling

TutorialsDGX agent

arXiv:2602.05202v2 Announce Type: replace Abstract: Aligning video generative models with human preferences remains challenging: current approaches rely on Vision-Language Models (VLMs) for reward mod

HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction

ResearchDGX agent

arXiv:2605.23889v1 Announce Type: new Abstract: Online 3D reconstruction requires estimating camera pose and scene geometry under strict causal and bounded-memory constraints. Existing methods often s

Improved Vision-to-Chart Buoy Association with Learned World-to-Image Projection

TutorialsDGX agent

arXiv:2605.22942v1 Announce Type: new Abstract: This report presents a lightweight modification to the DETR-based fusion transformer baseline for the MaCVi 2026 Vision-to-Chart data association challe

Inconsistency-aware Multimodal Schrodinger Bridge for Deepfake Localization

ResearchDGX agent

arXiv:2605.23113v1 Announce Type: new Abstract: Audio-visual deepfake localization demands interval-level outputs that serve as temporal evidence. Despite recent progress, symmetric fusion under singl

IntentionNav: A Benchmark for Intent-Driven Object Navigation from Implicit Human Instruction

Model ReleasesDGX agent

arXiv:2605.23187v1 Announce Type: new Abstract: Existing object navigation benchmarks usually tell an embodied agent which object category to find, such as microwave or chair. Human-facing embodied AI

Joint Target-Less Intrinsic and Extrinsic Camera-LiDAR Calibration using Deep Point Correspondences

ResearchDGX agent

arXiv:2605.23397v1 Announce Type: new Abstract: Accurate camera-LiDAR calibration is a prerequisite for robust multi-modal perception in robotics. Recent target-less approaches based on deep point cor

LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation

ResearchDGX agent

arXiv:2605.23878v1 Announce Type: new Abstract: Modern video generators produce visually compelling clips but still struggle with physical and motion consistency, limiting their use as reliable world

LangFlash: Feed-forward 3D Language Gaussian Splatting from Sparse Unposed Images

ResearchDGX agent

arXiv:2605.23287v1 Announce Type: new Abstract: We present LangFlash, a feed-forward framework for 3D Language Gaussian Splatting that reconstructs 3D scenes parameterized by Gaussian primitives enric

Learning a Particle Dynamics Model with Real-world Videos

TutorialsDGX agent

arXiv:2605.23845v1 Announce Type: new Abstract: Data-driven learning approaches for physics simulation, sometimes referred to as world models, have emerged as promising alternatives to traditional phy

LQ-rPPG: A Label-Quantized Coarse-to-Fine Learning Framework for Remote Physiological Measurement

Model ReleasesDGX agent

arXiv:2605.23174v1 Announce Type: new Abstract: Remote photoplethysmography (rPPG) enables non-contact measurement of physiological signals from facial videos, offering strong potential for remote hea

Machine learning applied to emerald gemstone grading: framework proposal and creation of a public dataset

ResearchDGX agent

arXiv:2605.23777v1 Announce Type: new Abstract: The grading of gemstones is currently a manual procedure performed by gemologists. A popular approach uses reference stones, where those are visually in

MapGCLR: Geospatial Contrastive Learning of Representations for Online Vectorized HD Map Construction

AgentsDGX agent

arXiv:2603.10688v2 Announce Type: replace-cross Abstract: Autonomous vehicles rely on map information to understand the world around them. However, the creation and maintenance of offline high-definit

MDS-DETR: DETR with Masked Duplicate Suppressor

ResearchDGX agent

arXiv:2605.23507v1 Announce Type: new Abstract: The DEtection TRansformer (DETR) is a powerful end-to-end object detector, yet its one-to-one matching strategy suffers from slow convergence and low re

Millimeter-wave Imaging for Anthropometric Body Measurement

SafetyDGX agent

arXiv:2605.23064v1 Announce Type: new Abstract: Body shape and circumferences are clinically informative biomarkers for risk stratification, including measures such as waist to hip ratio, limb and tru

Mitigating Object Hallucinations via Sentence-Level Early Intervention

ResearchDGX agent

arXiv:2507.12455v3 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have revolutionized cross-modal understanding but continue to struggle with hallucinations - fabricated con

MuellerPT: Decomposition Driven Pretraining for Dense Learning in Mueller Polarimetry

ResearchDGX agent

arXiv:2605.23840v1 Announce Type: new Abstract: Mueller matrix imaging provides rich, physically meaningful contrast for biomedical tissue analysis, but supervised learning is hindered by scarce dense

NeuralBoneReg: An Instance-Specific Label-Free Point Cloud-Based Method for Multi-Modal Bone Surface Registration

SafetyDGX agent

arXiv:2511.14286v2 Announce Type: replace Abstract: In computer- and robot-assisted orthopedic surgery (CAOS), patient-specific surgical plans derived from preoperative imaging define target locations

NP-LoRA: Null Space Projection for Subject-Style LoRA Fusion

Model ReleasesDGX agent

arXiv:2511.11051v3 Announce Type: replace Abstract: Low-Rank Adaptation (LoRA) fusion enables the composition of subject and style representations for controllable generation without retraining. Howev

Occlusion-Aware Physics-Semantic Keyframe Selection for Robust Video Editing

ResearchDGX agent

arXiv:2605.23192v1 Announce Type: new Abstract: Video editing has recently achieved remarkable progress with diffusion-based generative models, enabling diverse object-level manipulations from natural

On the Provable Importance of Gradients for Language-Assisted Image Clustering

TutorialsDGX agent

arXiv:2510.16335v4 Announce Type: replace Abstract: This paper investigates the recently emerged problem of Language-assisted Image Clustering (LaIC), where textual semantics are leveraged to improve

PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion

HardwareDGX agent

arXiv:2605.23902v1 Announce Type: new Abstract: Most practical high-resolution text-to-image systems, including latent diffusion and autoregressive models, perform generation in a compact latent space

PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2503.06684v3 Announce Type: replace Abstract: Recent advances in diffusion-based text-to-image generation have demonstrated promising results through visual condition control. However, existing

PixIE: Prompted Pixel-Space Low-Light Image Enhancement

ResearchDGX agent

arXiv:2605.23531v1 Announce Type: new Abstract: Low-light images exhibit severe noise, contrast loss, and semantic ambiguity, making enhancement a joint problem of denoising and detail recovery. We pr

ProGIC: Progressive and Lightweight Generative Image Compression with Residual Vector Quantization

ResearchDGX agent

arXiv:2603.02897v2 Announce Type: replace Abstract: Recent advances in generative image compression (GIC) have delivered remarkable improvements in perceptual quality. However, many GICs rely on large

Recursive Block-Diagonal Coupling for Resource-Efficient Training of Vision Models

Model ReleasesDGX agent

arXiv:2605.23656v1 Announce Type: new Abstract: Training high-capacity vision models from scratch requires substantial computational resources. To improve training efficiency of a wide target model, e

Rethinking Transfer Learning for Industrial Inspection: DINOv3 vs. ImageNet Pretraining Across RGB and X-ray Tasks

ResearchDGX agent

arXiv:2605.23472v1 Announce Type: new Abstract: Vision foundation models pretrained on web-scale data have recently shown strong transfer capabilities on many downstream tasks, but their effectiveness

Revitalizing Dense Material Segmentation: Stabilized Vision Transformers and the Generalization Paradox

Model ReleasesDGX agent

arXiv:2605.23747v1 Announce Type: new Abstract: Material segmentation, the pixel-wise classification of physical surface properties, remains a challenging problem in computer vision, requiring physico

RiGS: Rigid-aware 4D Gaussian Splatting from a Single Monocular Video

TutorialsDGX agent

arXiv:2605.23672v1 Announce Type: new Abstract: Reconstructing dynamic 3D scenes from monocular videos is a fundamental yet highly challenging task, as real-world motions often involve both long-term

RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering

Model ReleasesDGX agent

arXiv:2605.23068v1 Announce Type: new Abstract: Reliable visual understanding in robot-assisted and minimally invasive surgery (RMIS/MIS) demands more than accurate masks: in clinical practice, clinic

RS2AD-LiDAR: End-to-End Autonomous Driving LiDAR Data Generation from Roadside Sensor Observations

AgentsDGX agent

arXiv:2605.23406v1 Announce Type: new Abstract: End-to-end autonomous driving solutions, which directly process multimodal sensory data and output fine-grained control commands, have gradually become

RT-NeRV: Rethinking Hybrid Neural Representations for Video via Residual Tokenization

ResearchDGX agent

arXiv:2403.12401v2 Announce Type: replace Abstract: Neural Representations for Videos(NeRV) have emerged as a promising paradigm for video compression by representing videos as compact neural networks

Sample-wise Targeted Adversarial Attacks on Test-time Adaptation

SafetyDGX agent

arXiv:2605.23411v1 Announce Type: cross Abstract: Test-time adaptation (TTA) effectively counters distribution shifts but exposes models to adversarial manipulation via the unlabeled test stream. Exis

Scene Reconstruction as Mapping Priors for 3D Detection

AgentsDGX agent

arXiv:2605.22997v1 Announce Type: new Abstract: In autonomous driving, mapping is critical for motion planning but remains an under-utilized resource for perception tasks such as 3D object detection.

SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models

Local AiDGX agent

arXiv:2605.23345v1 Announce Type: new Abstract: Interactive world models for first-person shooter (FPS) games must resolve high-frequency overlapping control signals at every frame without disrupting

Semantic-Aware Guided Drone Exploration for Language-Conditioned 3D Indoor Mapping

ApplicationsDGX agent

arXiv:2605.23160v1 Announce Type: cross Abstract: We present Semantic-Aware Guided Exploration, SAGE, a system for open-vocabulary exploration in unknown 3D indoor environments that preserves coverage

SLIP-RS: Structured-Attribute Language-Image Pre-Training for Remote Sensing Object Detection

ResearchDGX agent

arXiv:2605.23144v1 Announce Type: new Abstract: Existing language-image pre-training for remote sensing object detection is constrained by Monolithic Label Learning, which relies on exhaustively enume

Smart-Insertion-V: Photorealistic Video Insertion via a Closed-Loop Feedback Dual-Stream Framework

ResearchDGX agent

arXiv:2605.23891v1 Announce Type: new Abstract: Mask-free video object insertion has emerged as a challenging task, requiring harmonious integration of reference objects into source videos. However, e

Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition

SafetyDGX agent

arXiv:2605.23288v1 Announce Type: new Abstract: Recent Open-Vocabulary Action Recognition (OVAR) methods typically aggregate visual features into a global representation before computing text alignmen

STAMBRIDGE: Spectral-Temporal Amplitude-aware Mid-Feature Bridge for EEG Visual Decoding

Model ReleasesDGX agent

arXiv:2605.23137v1 Announce Type: cross Abstract: Electroencephalography (EEG) visual decoding remains challenging due to the modality gap between low-SNR neural signals and highly structured vision--

StereoGenBench: A Synthetic Multi-Camera Benchmark for Stereo Generation under Controlled Baseline Regimes

Model ReleasesDGX agent

arXiv:2605.23237v1 Announce Type: new Abstract: Stereo image and video generation, stereo geometry estimation, and condition-controlled view synthesis require paired data in which the variables that d

← Previous
1…115116117118119…211
Next →