AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
28 Apr 2026

Light 'em Up: Enabling Few-Shot Low-Light 3D Gaussian Splatting with Multi-Scale Explicit Retinex Illumination Decoupling

ResearchDGX agent

arXiv:2604.24053v1 Announce Type: new Abstract: Full 360^irc novel view synthesis under low-light conditions remains challenging. Insufficient illumination, noise amplification, and view-dependent pho

LLMind: Bio-inspired Training-free Adaptive Visual Representations for Vision-Language Models

ResearchDGX agent

arXiv:2603.14882v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) typically assume a uniform spatial fidelity across the entire field of view of visual inputs, dedicating equal precisi

LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2603.03269v2 Announce Type: replace Abstract: Feedforward geometric foundation models achieve strong short-window reconstruction, yet scaling them to minutes-long videos is bottlenecked by quadr

Lost in the Vibrations: Vision Language Models Fail the Dynamic Gauges Test

Model ReleasesDGX agent

arXiv:2604.22829v1 Announce Type: new Abstract: The digital transformation of industrial manufacturing increasingly relies on the ability of autonomous robots to interact with legacy infrastructure, p

LunarDepthNet: Generation of Digital Elevation Models using Deep Learning and Monocular Satellite Images

TutorialsDGX agent

arXiv:2604.22848v1 Announce Type: new Abstract: Recent times have seen an increase in demand of high quality Digital Elevation Models (DEMs) for the lunar surface, because they are highly important fo

Majorization-Guided Test-Time Adaptation for Vision-Language Models under Modality-Specific Shift

Model ReleasesDGX agent

arXiv:2604.24602v1 Announce Type: new Abstract: Vision-language models transfer well in zero-shot settings, but at deployment the visual and textual branches often shift asymmetrically. Under this con

Mammographic Lesion Segmentation with Lightweight Models: A Comparative Study

ResearchDGX agent

arXiv:2604.23899v1 Announce Type: new Abstract: Breast cancer is a leading cause of cancer-related mortality among women worldwide, with mammography as the primary screening tool. While deep learning

MARRS: Masked Autoregressive Unit-based Reaction Synthesis

ResearchDGX agent

arXiv:2505.11334v4 Announce Type: replace Abstract: This work aims at a challenging task: human action-reaction synthesis, i.e., generating human reactions conditioned on the action sequence of anothe

MeshLAM: Feed-Forward One-Shot Animatable Textured Mesh Avatar Reconstruction

ResearchDGX agent

arXiv:2604.22865v1 Announce Type: new Abstract: We introduce MeshLAM, a feed-forward framework for one-shot animatable mesh head reconstruction that generates high-fidelity, animatable 3D head avatars

Micro-Expression-Aware Avatar Fingerprinting via Inter-Frame Feature Differencing

ResearchDGX agent

arXiv:2604.23247v1 Announce Type: new Abstract: Avatar fingerprinting, i.e., verifying who drives a synthetic talking-head video rather than whether it is real, is a critical safeguard for authorized

MIRAGE: A Micro-Interaction Relational Architecture for Grounded Exploration in Multi-Figure Artworks

SafetyDGX agent

arXiv:2604.23788v1 Announce Type: new Abstract: Appreciating multi-figure paintings requires understanding how characters relate through subtle cues like gaze alignment, gesture, and spatial arrangeme

Monocular Depth Estimation via Neural Network with Learnable Algebraic Group and Ring Structures

ResearchDGX agent

arXiv:2604.24328v1 Announce Type: new Abstract: Monocular depth estimation (MDE) has witnessed remarkable progress driven by Convolutional Neural Networks and transformer-based architectures. However,

Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation

TutorialsDGX agent

arXiv:2412.07160v3 Announce Type: replace Abstract: To equip artificial intelligence with a comprehensive understanding towards a temporal world, video and 4D panoptic scene graph generation abstracts

MotionHiFlow: Text-to-motion via hierarchical flow matching

SafetyDGX agent

arXiv:2604.23264v1 Announce Type: new Abstract: Text-to-motion generation aims to generate 3D human motions that are tightly aligned with the input text while remaining physically plausible and rich i

Multi-Scale Contrastive Learning for Video Temporal Grounding

ResearchDGX agent

arXiv:2412.07157v3 Announce Type: replace Abstract: Temporal grounding, which localizes video moments related to a natural language query, is a core problem of vision-language learning and video under

Multi-View Synergistic Learning with Vision-Language Adaption for Low-Resource Biomedical Image Classification

Model ReleasesDGX agent

arXiv:2604.23977v1 Announce Type: new Abstract: Accurate biomedical image classification under low-resource conditions remains challenging due to limited annotations, subtle inter-class visual differe

Multispectral airborne laser scanning dataset for tree species classification: MS-ALS-SPECIES

ResearchDGX agent

arXiv:2604.24370v1 Announce Type: new Abstract: The shift from stand-level to individual-tree-level forest assessments supports improved biodiversity mapping, particularly in boreal ecosystems where t

Multivariate Gaussian NeRF for Wide Field-of-View Ultrasound Reconstruction

ResearchDGX agent

arXiv:2604.24187v1 Announce Type: new Abstract: Wide Field-of-View (WFoV) reconstruction enhances 3D ultrasound imaging by providing valuable anatomical context for segmentation models and visualizati

MuSc-V2: Zero-Shot Multimodal Industrial Anomaly Classification and Segmentation with Mutual Scoring of Unlabeled Samples

ResearchDGX agent

arXiv:2511.10047v2 Announce Type: replace Abstract: Zero-shot anomaly classification (AC) and segmentation (AS) methods aim to identify and outline defects without using any labeled samples. In this p

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation

Model ReleasesDGX agent

arXiv:2604.23789v1 Announce Type: new Abstract: While video foundation models excel at single-shot generation, real-world cinematic storytelling inherently relies on complex multi-shot sequencing. Fur

NeuroClaw Technical Report

Model ReleasesDGX agent

arXiv:2604.24696v1 Announce Type: new Abstract: Agentic artificial intelligence systems promise to accelerate scientific workflows, but neuroimaging poses unique challenges: heterogeneous modalities (

Not All Directions Matter: Towards Structured and Task-Aware Low-Rank Model Adaptation

Model ReleasesDGX agent

arXiv:2603.14228v2 Announce Type: replace Abstract: Low-Rank Adaptation (LoRA) has become a cornerstone of parameter-efficient fine-tuning (PEFT). Yet, its efficacy is hampered by two fundamental limi

NVILA: Efficient Frontier Visual Language Models

ResearchDGX agent

arXiv:2412.04468v3 Announce Type: replace Abstract: Visual language models (VLMs) have made significant advances in accuracy in recent years. However, their efficiency has received much less attention

ODE-GS: Latent ODEs for Dynamic Scene Extrapolation with 3D Gaussian Splatting

Model ReleasesDGX agent

arXiv:2506.05480v4 Announce Type: replace-cross Abstract: We introduce ODE-GS, a novel approach that integrates 3D Gaussian Splatting with latent neural ordinary differential equations (ODEs) to enabl

Omni-o3: Deep Nested Omnimodal Deduction for Deliberative Audio-Visual Reasoning

SafetyDGX agent

arXiv:2604.24191v1 Announce Type: new Abstract: Omnimodal understanding entails a massive, highly redundant search space of cross-modal interactions, demanding focused and deliberative reasoning. Curr

OmniSch: A Multimodal PCB Schematic Benchmark For Structured Diagram Visual Reasoning

Model ReleasesDGX agent

arXiv:2604.00270v2 Announce Type: replace Abstract: Recent large multimodal models (LMMs) have made rapid progress in visual grounding, document understanding, and diagram reasoning tasks. However, th

OmniShotCut: Holistic Relational Shot Boundary Detection with Shot-Query Transformer

Model ReleasesDGX agent

arXiv:2604.24762v1 Announce Type: new Abstract: Shot Boundary Detection (SBD) aims to automatically identify shot changes and divide a video into coherent shots. While SBD was widely studied in the li

On-Device Vision Training, Deployment, and Inference on a Thumb-Sized Microcontroller

Model ReleasesDGX agent

arXiv:2604.23012v1 Announce Type: cross Abstract: This paper presents a complete, end-to-end on-device vision machine learning pipeline, comprising data acquisition, two-layer CNN training with Adam o

One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition

ResearchDGX agent

arXiv:2604.23173v1 Announce Type: new Abstract: Video Situation Recognition (VidSitu) addresses the challenging problem of 'who did what to whom, with what, how, and where' in a video. It tests thorou

Open-Vocabulary Semantic Segmentation Network Integrating Object-Level Label and Scene-Level Semantic Features for Multimodal Remote Sensing Images

ApplicationsDGX agent

arXiv:2604.24125v1 Announce Type: new Abstract: Semantic segmentation of multi-modal remote sensing imagery plays a pivotal role in land use/land cover (LULC) mapping, environmental monitoring, and pr

OpenVO: Open-World Visual Odometry with Temporal Dynamics Awareness

AgentsDGX agent

arXiv:2602.19035v2 Announce Type: replace Abstract: We introduce OpenVO, a novel framework for Open-world Visual Odometry (VO) with temporal awareness under limited input conditions. OpenVO effectivel

Oracle Noise: Faster Semantic Spherical Alignment for Interpretable Latent Optimization

SafetyDGX agent

arXiv:2604.23540v1 Announce Type: new Abstract: Text-to-image diffusion models have achieved remarkable generative capabilities, yet accurately aligning complex textual prompts with synthesized layout

Parameter-Efficient Multi-Task Learning via Progressive Task-Specific Adaptation

Model ReleasesDGX agent

arXiv:2509.19602v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning methods have emerged as a promising solution for adapting pre-trained models to various downstream tasks. While thes

PEPS: Positional Encoding Projected Sampling -- Extended

TutorialsDGX agent

arXiv:2604.24167v1 Announce Type: new Abstract: Implicit neural representations (INRs) are increasingly being used as tools to map coordinates to signals, encompassing applications from neural fields

Personalizing Causal Audio-Driven Facial Motion via Dynamic Multi-modal Retrieval

ResearchDGX agent

arXiv:2604.23692v1 Announce Type: cross Abstract: Audio-driven facial animation is essential for immersive digital interaction, yet existing frameworks fail to reconcile real-time streaming with high-

Phase-Separated Complex Hilbert PCA on Markerless 3D Pose Estimation Data: A Global Phase Network and Its Extension to a Continuous Field on the Body Surface

ResearchDGX agent

arXiv:2604.24415v1 Announce Type: cross Abstract: Quantitative analysis of the kinematic chain in sports motion is essential for performance evaluation and injury prevention. Conventional methods such

Physics-Informed Temporal U-Net for High-Fidelity Fluid Interpolation

ResearchDGX agent

arXiv:2604.23372v1 Announce Type: cross Abstract: Reconstructing high-fidelity fluid dynamics from sparse temporal observations is quite challenging, mainly due to the chaotic and non-linear nature of

PhysLayer: Language-Guided Layered Animation with Depth-Aware Physics

SafetyDGX agent

arXiv:2604.23574v1 Announce Type: new Abstract: Existing image-to-video generation methods often produce physically implausible motions and lack precise control over object dynamics. While prior appro

POCA: Pareto-Optimal Curriculum Alignment for Visual Text Generation

SafetyDGX agent

arXiv:2604.24171v1 Announce Type: new Abstract: Current visual text generation models struggle with the trade-off between text accuracy and overall image coherence. We find that achieving high text ac

Point Cloud Registration for Fusion between SPECT MPI and CTA Images

ResearchDGX agent

arXiv:2604.24524v1 Announce Type: new Abstract: Clinical fusion of Single Photon Emission Computed Tomography Myocardial Perfusion Imaging (SPECT MPI) and Computed Tomography Angiography (CTA) remains

Point-MF: One-step Point Cloud Generation from a Single Image via Mean Flows

TutorialsDGX agent

arXiv:2604.24586v1 Announce Type: new Abstract: Single-image point cloud reconstruction must infer complete 3D geometry, including occluded parts, from a single RGB image. While diffusion-based recons

PointTransformerX:Portable and Efficient 3D Point Cloud Processing without Sparse Algorithms

HardwareDGX agent

arXiv:2604.24169v1 Announce Type: new Abstract: 3D point cloud perception remains tightly coupled to custom CUDA operators for spatial operations, limiting portability and efficiency on non-NVIDIA, AM

POUR: A Provably Optimal Method for Unlearning Representations via Neural Collapse

ResearchDGX agent

arXiv:2511.19339v2 Announce Type: replace Abstract: In computer vision, machine unlearning aims to remove the influence of specific visual concepts or training images without retraining from scratch.

Probing CLIP's Comprehension of 360-Degree Textual and Visual Semantics

SafetyDGX agent

arXiv:2604.24642v1 Announce Type: new Abstract: The dream of instantly creating rich 360-degree panoramic worlds from text is rapidly becoming a reality, yet a crucial gap exists in our ability to rel

RACANet: Reliability-Aware Crowd Anchor Network for RGB-T Crowd Counting

Model ReleasesDGX agent

arXiv:2604.24543v1 Announce Type: new Abstract: RGB-Thermal (T) crowd counting aims to integrate visible-spectrum and thermal infrared information to improve the robustness of crowd density estimation

Radiomics- and Clinical Feature-Driven Prediction of Volumetric Response in Skull-Base Meningioma after CyberKnife Radiosurgery

ResearchDGX agent

arXiv:2604.24230v1 Announce Type: new Abstract: Skull-base meningiomas are often characterized by favorable long-term prognosis, yet their anatomical complexity and proximity to critical neurovascular

Reading in the Dark: Low-light Scene Text Recognition

Model ReleasesDGX agent

arXiv:2604.23685v1 Announce Type: new Abstract: Accurate text recognition in low-light environments is essential for intelligent systems in applications ranging from autonomous vehicles to smart surve

Resource-Constrained UAV-Based Weed Detection for Site-Specific Management on Edge Devices

Local AiDGX agent

arXiv:2604.23442v1 Announce Type: new Abstract: Weeds compete with crops for light, water, and nutrients, reducing yield and crop quality. Efficient weed detection is essential for site-specific weed

ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning

Model ReleasesDGX agent

arXiv:2604.24300v1 Announce Type: new Abstract: Current evaluations of spatial intelligence can be systematically invalid under modern vision-language model (VLM) settings. First, many benchmarks deri

Robust Deepfake Detection, NTIRE 2026 Challenge: Report

ApplicationsDGX agent

arXiv:2604.24163v1 Announce Type: new Abstract: Robustness is a long-overlooked problem in deepfake detection. However, detection performance is nearly worthless in the real world if it suffers under

Robust Grounding with MLLMs against Occlusion and Small Objects via Language-guided Semantic Cues

TutorialsDGX agent

arXiv:2604.24036v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) have enhanced grounding capabilities in general scenes, their robustness in crowded scenes remains undere

SATTC: Structure-Aware Label-Free Test-Time Calibration for Cross-Subject EEG-to-Image Retrieval

Local AiDGX agent

arXiv:2603.20738v2 Announce Type: replace Abstract: Cross-subject EEG-to-image retrieval for visual decoding is challenged by subject shift and hubness in the embedding space, which distort similarity

Seer: Language Instructed Video Prediction with Latent Diffusion Models

SafetyDGX agent

arXiv:2303.14897v4 Announce Type: replace Abstract: Imagining the future trajectory is the key for robots to make sound planning and successfully reach their goals. Therefore, text-conditioned video p

Self-Rewarding Vision-Language Model via Reasoning Decomposition

HardwareDGX agent

arXiv:2508.19652v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) often suffer from visual hallucinations: generating things that are not consistent with visual inputs and language sho

Self-Supervised Representation Learning via Hyperspherical Density Shaping

SafetyDGX agent

arXiv:2604.24498v1 Announce Type: new Abstract: Modern self-supervised representation learning methods often relies on empirical heuristics that are not theoretically grounded. In this study we propos

Semantic Segmentation for Histopathology using Learned Regularization based on Global Proportions

Model ReleasesDGX agent

arXiv:2604.24347v1 Announce Type: cross Abstract: In pathology, the spatial distribution and proportions of tissue types are key indicators of disease progression, and are more readily available than

SemiGDA: Generative Dual-distribution Alignment for Semi-Supervised Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2604.23274v1 Announce Type: new Abstract: Semi-supervised learning addresses label scarcity and high annotation costs in medical image segmentation by exploiting the latent information in unlabe

SemiSAM-O1: How far can we push the boundary of annotation-efficient medical image segmentation?

ResearchDGX agent

arXiv:2604.24109v1 Announce Type: new Abstract: Semi-supervised learning (SSL) has become a promising solution to alleviate the annotation burden of deep learning-based medical image segmentation mode

ServImage: An Image Generation and Editing Benchmark from Real-world Commercial Imaging Services

Model ReleasesDGX agent

arXiv:2604.24023v1 Announce Type: new Abstract: Recent image generation and editing models demonstrate robust adherence to instructions and high visual quality on academic benchmarks. However, their p

Shape: A Self-Supervised 3D Geometry Foundation Model for Industrial CAD Analysis

Model ReleasesDGX agent

arXiv:2604.22826v1 Announce Type: new Abstract: Industrial CAD workflows require robust, generalizable 3D geometric representations supporting accuracy and explainability. We introduce Shape, a self-s

← Previous
1…170171172173174…209
Next →