AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
19 May 2026

Mining Forgery Traces from Reconstruction Error: A Weakly Supervised Framework for Multimodal Deepfake Temporal Localization

Local AiDGX agent

arXiv:2601.21458v2 Announce Type: replace Abstract: Modern deepfakes have evolved into localized and intermittent manipulations that require fine-grained temporal localization to mitigate severe digit

MIRAGE: Robust multi-modal architectures translate fMRI-to-image models from vision to mental imagery

Model ReleasesDGX agent

arXiv:2605.17198v1 Announce Type: cross Abstract: To be useful for downstream applications, vision decoding models that are trained to reconstruct seen images from human brain activity must be able to

Mitigating 3D Prostate Biparametric MRI Data Scarcity through Domain Adaptation using Locally-Trained Latent Diffusion Models for Prostate Cancer Detection


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research
DGX agent

arXiv:2507.06384v2 Announce Type: replace-cross Abstract: Objective: Latent diffusion models (LDMs) could mitigate data scarcity challenges affecting machine learning development for medical image int

MoASE++: Mixture of Activation Sparsity Experts with Domain-Adaptive On-policy Distillation for Continual Test Time Adaptation

SafetyDGX agent

arXiv:2605.17743v1 Announce Type: new Abstract: Continual test-time adaptation adapts a source-pretrained model to non-stationary, unlabeled target streams while retaining past competence, yet texture

MoCA3D: Monocular 3D Bounding Box Prediction in the Image Plane

ResearchDGX agent

arXiv:2603.19538v2 Announce Type: replace Abstract: Monocular 3D object understanding has largely been cast as a 2D RoI-to-3D box lifting problem. However, emerging downstream applications require ima

Mono-Hydra++: Real-Time Monocular Scene Graph Construction with Multi-Task Learning for 3D Indoor Mapping

Local AiDGX agent

arXiv:2605.17661v1 Announce Type: cross Abstract: Autonomous agile robots need more than metric geometry: they must understand objects, rooms, places, and spatial relations for search, inspection, exp

Monocular Depth Perception Enhancement Based on Joint Shading/Contrast Model and Motion Parallax (JSM)

ResearchDGX agent

arXiv:2605.17252v1 Announce Type: new Abstract: Stereoscopic 3D displays adopt a binocular depth cue to provide depth perception. However, users should be equipped with expensive special devices to ap

Monocular Open Vocabulary Occupancy Prediction for Indoor Scenes

Model ReleasesDGX agent

arXiv:2602.22667v2 Announce Type: replace Abstract: Open-vocabulary 3D occupancy is vital for embodied agents, which need to understand complex indoor environments where semantic categories are abunda

MorphSeek: Fine-grained Latent Representation-Level Policy Optimization for Deformable Image Registration

Model ReleasesDGX agent

arXiv:2511.17392v3 Announce Type: replace Abstract: Deformable image registration (DIR) remains a fundamental yet challenging problem in medical image analysis, largely due to the prohibitively high-d

Motion Cues from Image-based Point Tracking for LiDAR Scene Flow Estimation

AgentsDGX agent

arXiv:2605.16922v1 Announce Type: new Abstract: LiDAR scene flow estimation is essential for autonomous driving, as it provides 3D motion for each point. Self-supervised approaches use static-dynamic

MSIQ: Moment-based Scale-Invariant Quality Measure for Single Image Super-Resolution

SafetyDGX agent

arXiv:2605.17588v1 Announce Type: new Abstract: Assessing the quality of single image super-resolution (SISR) results remains an open methodological problem. Common full-reference metrics (PSNR, SSIM,

Multi-hop Relational Contrastive Learning: Extending Spatial Contrastive Pre-training Beyond Pairwise Relations

ResearchDGX agent

arXiv:2605.16456v1 Announce Type: new Abstract: Understanding how objects relate to each other in space is fundamental to scene understanding, yet most contrastive pre-training approaches only model p

Multi-Order Matching Network for Alignment-Free Depth Super-Resolution

SafetyDGX agent

arXiv:2511.16361v3 Announce Type: replace Abstract: Recent guided depth super-resolution methods are premised on the assumption of strict spatial alignment between depth and RGB, achieving high-qualit

NeRF-based Spacecraft Reconstruction from Close-Range Monocular Imagery Under Illumination Variability and Pose Uncertainty

AgentsDGX agent

arXiv:2605.18447v1 Announce Type: new Abstract: Autonomous rendezvous and proximity operations around uncooperative, unknown spacecraft are critical for active debris removal and on-orbit servicing mi

NERVE: A Neuromorphic Vision and Radar Ensemble for Multi-Sensor Fusion Research

ResearchDGX agent

arXiv:2605.16414v1 Announce Type: new Abstract: We present NERVE (Neuromorphic Vision and Radar Ensemble), a multi-sensor dataset comprising 257 minutes of synchronized recordings from five sensors: t

Network Knowledge Prior Guided Learning for Data-Efficient Surface Defect Detection

ApplicationsDGX agent

arXiv:2605.17780v1 Announce Type: new Abstract: Deep learning-based methods have become the de facto standard for industrial defect detection. However, their data-hungry nature and inherent 'black-box

NeuroLiDAR: Adaptive Frame Rate Depth Sensing via Neuromorphic Event-LiDAR Fusion

ResearchDGX agent

arXiv:2605.16805v1 Announce Type: new Abstract: LiDARs are widely used for 3D depth reconstruction, but their performance is often limited by inherent hardware constraints that impose trade-offs betwe

Neuroscience-inspired Staged Representation Learning with Disentangled Coarse- and Fine-Grained Semantics for EEG Visual Decoding

Model ReleasesDGX agent

arXiv:2605.16923v1 Announce Type: new Abstract: Decoding visual information from electroencephalography (EEG) signals remains a fundamental challenge in brain-computer interfaces and medical rehabilit

NEWTON: Agentic Planning for Physically Grounded Video Generation

SafetyDGX agent

arXiv:2605.18396v1 Announce Type: new Abstract: Video generation models produce visually compelling results but systematically violate physical commonsense -- on VideoPhy-2, the best model achieves on

Noise2Params: Unification and Parameter Determination from Noise via a Probabilistic Event Camera Model

Model ReleasesDGX agent

arXiv:2605.16317v1 Announce Type: new Abstract: Accurate, unified models for event cameras (ECs) remain elusive, hampering calibration and algorithm design. We develop a foundational probabilistic mod

Non-Colliding Biometric Identities for Digital Entities: Geometry, Capacity, and Million-Scale Virtual Identity Provisioning

ResearchDGX agent

arXiv:2605.18238v1 Announce Type: new Abstract: Digital entities such as AI agents and humanoid robots increasingly operate alongside real humans, yet their identity infrastructure is based on credent

Nonlinear Bipolar Compensation: Handling Outliers in Post-Training Quantization

ResearchDGX agent

arXiv:2605.16423v1 Announce Type: new Abstract: Network quantization has emerged as one of the most practical model compression techniques, which significantly reduces a model's memory and compute con

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation

ResearchDGX agent

arXiv:2605.17488v1 Announce Type: new Abstract: The landscape of joint audio and video generation has been fundamentally transformed by the advent of powerful foundation models. Despite these strides,

Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction

Model ReleasesDGX agent

arXiv:2605.17360v1 Announce Type: new Abstract: Real-time duplex interaction is essential for multimodal AI systems operating in real-world scenarios, where models must continuously process streaming

OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

Model ReleasesDGX agent

arXiv:2605.18577v1 Announce Type: new Abstract: Omni-proactive streaming video understanding, i.e., autonomously deciding when to speak and what to say from continuous audio-visual streams, is an emer

OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models

ResearchDGX agent

arXiv:2605.18041v1 Announce Type: new Abstract: Omnimodal large language models (OmniLLMs) have recently gained increasing attention for unified audio-video understanding. However, processing long mul

On Applicability of Synthetic Datasets for Facial Expression Recognition

ResearchDGX agent

arXiv:2605.17483v1 Announce Type: new Abstract: Facial Expression Recognition faces two core challenges. The first is class imbalance in public datasets, which skews the learning process and weakens g

Open Set Face Forgery Detection via Dual-Level Evidence Collection

ApplicationsDGX agent

arXiv:2512.04331v2 Announce Type: replace Abstract: The surge in face forgeries has increasingly undermined confidence in the authenticity of online content. As generation algorithms rapidly evolve, n

OpenGaFF: Open-Vocabulary Gaussian Feature Field with Codebook Attention

ResearchDGX agent

arXiv:2605.06088v2 Announce Type: replace Abstract: Understanding open-vocabulary 3D scenes with Gaussian-based representations remains challenging due to fragmented and spatially inconsistent semanti

OPTNet: Ordering Point Transformer Network for Post-disaster 3D Semantic Segmentation

ResearchDGX agent

arXiv:2605.17197v1 Announce Type: cross Abstract: Post-disaster damage assessment requires rapid and accurate semantic segmentation of 3D point clouds to identify critical infrastructure such as damag

P2GS: Physical Prior-guided Gaussian Splatting for Photometrically Consistent Urban Reconstruction

AgentsDGX agent

arXiv:2605.16925v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has recently emerged as a powerful explicit representation enabling fast, high-fidelity rendering, making it a promising fo

PanoWorld: A Generative Spatial World Model for Consistent Whole-House Panorama Synthesis

ResearchDGX agent

arXiv:2605.17916v1 Announce Type: new Abstract: Generating a consistent whole-house VR tour from a floorplan and style reference requires both photorealistic panoramas and cross-view spatial coherence

PartDiffuser: Part-wise 3D Mesh Generation via Discrete Diffusion

ApplicationsDGX agent

arXiv:2511.18801v3 Announce Type: replace Abstract: Existing autoregressive (AR) methods for generating artist-designed meshes struggle to balance global structural consistency with high-fidelity loca

Patch Ensembles for Robust Salmon Re-Identification with Weak Trajectory Labels

SafetyDGX agent

arXiv:2605.18038v1 Announce Type: new Abstract: Salmon re-identification in commercial net-pens is challenging due to large populations, which impose strict accuracy requirements and make large-scale

Patch-MoE Mamba: A Patch-Ordered Mixture-of-Experts State Space Architecture for Medical Image Segmentation

ResearchDGX agent

arXiv:2605.17719v1 Announce Type: new Abstract: CNN- and Transformer-based architectures have achieved strong performance in medical image segmentation, but CNNs are limited in modeling long-range dep

Patchwork: A compact representation for 3D polygonal shapes

ResearchDGX agent

arXiv:2605.16266v1 Announce Type: cross Abstract: We introduce Patchwork, a new general-purpose shape representation capable of modeling 2D and 3D geometry with a small number of parameters. Patchwork

PERL: Parameter Efficient Reasoning in CLIP Latent Space

Model ReleasesDGX agent

arXiv:2605.18464v1 Announce Type: new Abstract: Contrastively trained vision-language models such as CLIP provide strong zero-shot transfer by aligning images and text in a shared embedding space. How

PFlow-T: A Persistence-Driven Forward Process for Topology-Controlled Generation

ResearchDGX agent

arXiv:2605.17555v1 Announce Type: cross Abstract: Current topology aware diffusion models face an architectural mismatch by using Gaussian noise for corruption while recovering structural features thr

PhysSkin: Real-Time and Generalizable Physics-Based Animation via Self-Supervised Neural Skinning

TutorialsDGX agent

arXiv:2603.23194v2 Announce Type: replace-cross Abstract: Achieving real-time physics-based animation that generalizes across diverse 3D shapes and discretizations remains a fundamental challenge. We

PIXLRelight: Controllable Relighting via Intrinsic Conditioning

ResearchDGX agent

arXiv:2605.18735v1 Announce Type: new Abstract: We present PIXLRelight, a feed-forward approach for physically controllable single-image relighting. Existing methods either provide limited lighting co

PlantPose: Universal Plant Skeleton Estimation via Tree-constrained Graph Generation

ApplicationsDGX agent

arXiv:2605.17773v1 Announce Type: new Abstract: Accurate estimation of plant skeletal structures (e.g., branching structures) from images is essential for smart agriculture and plant science. Unlike h

Position: Age Estimation Models Do Not Process Biometric Data

SafetyDGX agent

arXiv:2605.17347v1 Announce Type: cross Abstract: When a neural network estimates someone's age from a photograph, does it process biometric data? The answer depends on whether identity-discriminative

Principal Component Analysis for Lunar Crater Detection

ResearchDGX agent

arXiv:2605.17125v1 Announce Type: new Abstract: Optical navigation is a critical component for lunar orbiter and lander missions. Image-based crater identification has emerged as a promising technolog

ProtoFlow: Mitigating Forgetting in Class-Incremental Remote Sensing Segmentation via Low-Curvature Prototype Flow

ResearchDGX agent

arXiv:2604.03212v2 Announce Type: replace Abstract: Remote sensing segmentation in real deployment is inherently continual: new semantic categories emerge, and acquisition conditions shift across seas

PySIFT: GPU-Resident Deterministic SIFT for Deep Learning Vision Pipelines

Local AiDGX agent

arXiv:2605.17869v1 Announce Type: new Abstract: A widespread assumption in local feature research holds that classical handcrafted descriptors are accuracy-limited relics best replaced by learned alte

QuadLink: Autoregressive Quad-Dominant Mesh Generation via Point-Relation Learning

ApplicationsDGX agent

arXiv:2605.16813v1 Announce Type: cross Abstract: The generation of production-ready quad-dominant meshes is a cornerstone of modern 3D content creation. Generating anisotropic quad-dominant meshes fr

Rad-VLSM: A Cross-Modal Framework with Semantics-Assisted Prompting for Medical Segmentation and Diagnosis

SafetyDGX agent

arXiv:2605.18130v1 Announce Type: new Abstract: Medical image segmentation is more clinically valuable when it supports diagnosis rather than merely producing lesion masks. However, diagnostically rel

RadGenome-Anatomy: A Large-Scale Anatomy-Labeled Chest Radiograph Dataset via Physically Grounded Volumetric Projection

ResearchDGX agent

arXiv:2605.17368v1 Announce Type: new Abstract: Anatomical structure labels for chest radiographs are essential for medical image segmentation and a broad range of downstream diagnostic tasks. However

Radial-Angular Geometry for Reliable Update Diagnosis in Noisy-Label Learning

Model ReleasesDGX agent

arXiv:2605.17429v1 Announce Type: cross Abstract: Noisy-label methods often estimate sample reliability from forward-space signals such as loss, confidence, or entropy. These signals indicate whether

RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture

TutorialsDGX agent

arXiv:2601.15891v2 Announce Type: replace Abstract: Recent advances in medical vision language models guide the learning of visual representations; however, this form of supervision is constrained by

RAVE: Re-Allocating Visual Attention in Large Multimodal Models

SafetyDGX agent

arXiv:2605.18359v1 Announce Type: new Abstract: Large multimodal models (LMMs) inherit the self-attention mechanism of pretrained language backbones, yet standard attention can exhibit suboptimal allo

Real-Time Neural Hair Denoising

ResearchDGX agent

arXiv:2605.17557v1 Announce Type: cross Abstract: We propose a lightweight real-time method for reconstructing strand-based hair G-Buffers from severely undersampled rasterized inputs. Our pipeline fi

ReBaR: Reference-Based Reasoning for Robust Pose Estimation from Monocular Images

Model ReleasesDGX agent

arXiv:2303.11675v3 Announce Type: replace Abstract: R}easoning for Robust Human Pose and Shape Estimation), designed to estimate human body shape and pose from single-view images. ReBaR effectively ad

REC-RL: Referring expression counting via Gaussian and range-based reward optimization

SafetyDGX agent

arXiv:2605.16460v1 Announce Type: new Abstract: Referring expression counting (REC) is an intention-driven task that requires context-aware visual reasoning. While recent vision-language models incorp

Resolving Representation Ambiguity in Feedforward Novel View Synthesis Transformer via Semantic-Spatial Decoupling

SafetyDGX agent

arXiv:2605.18599v1 Announce Type: new Abstract: Transformer-based models have advanced feedforward novel view synthesis (NVS). Current architectures such as GS-LRM and LVSM mix semantic information (e

Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?

ResearchDGX agent

arXiv:2511.08704v2 Announce Type: replace Abstract: This paper investigates the scaling properties of autoregressive next-pixel prediction, a simple, end-to-end yet under-explored framework for unifie

Rethinking Point Clouds as Sequences: A Causal Next-Token Predictive Learning Framework

Local AiDGX agent

arXiv:2605.17566v1 Announce Type: new Abstract: With the rapid progress of multimodal foundation models and predictive pre-training, an important open question is how to equip 3D point clouds with a p

Rethinking the State Update Gate for Long-Sequence Recurrent 3D Reconstruction

ResearchDGX agent

arXiv:2605.16981v1 Announce Type: new Abstract: Streaming 3D reconstruction under a strict constant-memory budget hinges on how the recurrent state is updated as the stream evolves. We profile TTT3R-s

RHINO: Reconstructing Human Interactions with Novel Objects from Monocular Videos

ApplicationsDGX agent

arXiv:2605.17014v1 Announce Type: new Abstract: Reconstructing people, objects, and their interactions in 3D is a long-standing goal for intelligent systems. Often the input is RGB video from a moving

Right Predictions, Misleading Explanations: On the Vulnerability of Vision-Language Model Explanations

SafetyDGX agent

arXiv:2605.16651v1 Announce Type: new Abstract: Explanation mechanisms are increasingly used to support transparency and trust in vision-language models (VLMs), particularly in settings where model de

← Previous
1…129130131132133…211
Next →