AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
6 May 2026

Enhanced 3D Brain Tumor Segmentation Using Assorted Precision Training

ResearchDGX agent

arXiv:2605.04008v1 Announce Type: new Abstract: A brain tumor is a medical disorder faced by individuals of all demographics. Medically, it is described as the spread of non-essential cells close to o

Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework

ResearchDGX agent

arXiv:2605.03390v1 Announce Type: new Abstract: Supervised talking head forgery detection faces severe generalization challenges due to the continuous evolution of generators. By reducing reliance on

Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation

TutorialsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.03790v1 Announce Type: new Abstract: With advances in multimodal research and deep learning, Multimodal Large Language Models (MLLMs) have emerged as a powerful paradigm for a wide range of

Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models

Model ReleasesDGX agent

arXiv:2605.03547v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs), trained on web-scale data, risk memorizing and regenerating copyrighted visual content such as characters and logo

FACTOR: Counterfactual Training-Free Test-Time Adaptation for Open-Vocabulary Object Detection

Model ReleasesDGX agent

arXiv:2605.03294v1 Announce Type: new Abstract: Open-vocabulary object detection often fails under distribution shifts, as it can be misled by spurious correlations between non-causal visual attribute

First Shape, Then Meaning: Efficient Geometry and Semantics Learning for Indoor Reconstruction

ApplicationsDGX agent

arXiv:2605.03463v1 Announce Type: new Abstract: Neural Surface Reconstruction has become a standard methodology for indoor 3D reconstruction, with Signed Distance Functions (SDFs) proving particularly

FluxFlow: Conservative Flow-Matching for Astronomical Image Super-Resolution

Model ReleasesDGX agent

arXiv:2605.03749v1 Announce Type: new Abstract: Ground-to-space astronomical super-resolution requires recovering space-quality images from ground-based observations that are simultaneously limited by

FreeTimeGS++: Secrets of Dynamic Gaussian Splatting and Their Principles

ResearchDGX agent

arXiv:2605.03337v1 Announce Type: new Abstract: The recent surge in 4D Gaussian Splatting (4DGS) has achieved impressive dynamic scene reconstruction. While these methods demonstrate remarkable perfor

From Code to Prediction: Fine-Tuning LLMs for Neural Network Performance Classification in NNGPT

Model ReleasesDGX agent

arXiv:2605.03686v1 Announce Type: cross Abstract: Automated Machine Learning (AutoML) frameworks increasingly leverage Large Language Models (LLMs) for tasks such as hyperparameter optimization and ne

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

SafetyDGX agent

arXiv:2507.07982v2 Announce Type: replace Abstract: Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models trained solely on raw v

GeoTopoDiff: Learning Geometry--Topology Graph Priors through Boundary-Constrained Mixed Diffusion for Sparse-Slice 3D Porous Reconstruction

ResearchDGX agent

arXiv:2605.03764v1 Announce Type: new Abstract: Diffusion-based voxel prior modelling is challenging for the reconstruction of large-scale 3D porous microstructures. Due to the demanding requirements

GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning

SafetyDGX agent

arXiv:2605.03403v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has recently shown strong performance in post-training large language models and vision-language models. It ra

Identity-Consistent Multi-Pose Generation of Contactless Fingerprints

Local AiDGX agent

arXiv:2605.03830v1 Announce Type: new Abstract: Contactless fingerprint recognition has gained increasing attention due to its advantages in hygiene and acquisition flexibility. However, the absence o

Illumination-Aware Contactless Fingerprint Spoof Detection via Paired Flash-Non-Flash Imaging

ResearchDGX agent

arXiv:2603.17679v2 Announce Type: replace Abstract: Contactless fingerprint recognition enables hygienic and convenient biometric authentication but poses new challenges for spoof detection due to the

IRIS: Intent Resolution via Inference-time Saccades for Open-Ended VQA in Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2602.16138v2 Announce Type: replace Abstract: We introduce IRIS (Intent Resolution via Inference-time Saccades), a novel training-free approach that uses eye-tracking data in real-time to resolv

Label-Efficient School Detection from Aerial Imagery via Weakly Supervised Pretraining and Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.03968v1 Announce Type: new Abstract: Accurate school detection is essential for supporting education initiatives, including infrastructure planning and expanding internet connectivity to un

LangPrecip: Language-Aware Multimodal Precipitation Nowcasting

ResearchDGX agent

arXiv:2512.22317v2 Announce Type: replace-cross Abstract: Short-term precipitation nowcasting is an inherently uncertain and under-constrained spatiotemporal forecasting problem, especially for rapidl

Large Language Models are Universal Reasoners for Visual Generation

SafetyDGX agent

arXiv:2605.04040v1 Announce Type: new Abstract: Text-to-image generation has advanced rapidly with diffusion models, progressing from CLIP and T5 conditioning to unified systems where a single LLM bac

Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures

ResearchDGX agent

arXiv:2605.04035v1 Announce Type: new Abstract: We propose HeadsUp, a scalable feed-forward method for reconstructing high-quality 3D Gaussian heads from large-scale multi-camera setups. Our method em

Learning Discriminative Signed Distance Functions from Multi-scale Level-of-detail Features for 3D Anomaly Detection

TutorialsDGX agent

arXiv:2605.03437v1 Announce Type: new Abstract: Detecting anomalies from 3D point clouds has received increasing attention in the field of computer vision, with some group-based or point-based methods

Learning to Segment using Summary Statistics and Weak Supervision

ResearchDGX agent

arXiv:2605.03059v1 Announce Type: new Abstract: Medical experts often manually segment images to obtain diagnostic statistics and discard the resulting annotations. We aim to train segmentation models

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute

HardwareDGX agent

arXiv:2504.17816v3 Announce Type: replace Abstract: Subject-driven video generation (SDV-Gen) aims to produce videos of a specific subject by adapting a pretrained video model, enabling personalized a

Mantis: Mamba-native Tuning is Efficient for 3D Point Cloud Foundation Models

Model ReleasesDGX agent

arXiv:2605.03438v1 Announce Type: new Abstract: Pre-trained 3D point cloud foundation models (PFMs) have demonstrated strong transferability across diverse downstream tasks. However, full fine-tuning

MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding

Local AiDGX agent

arXiv:2605.03398v1 Announce Type: new Abstract: Video Temporal Grounding (VTG) faces a cross-modal semantic gap that often leads to background features being incorrectly aligned with the query, while

MedSR-Vision: Deep Learning Framework for Multi-Domain Medical Image Super-Resolution

ResearchDGX agent

arXiv:2605.03343v1 Announce Type: new Abstract: Medical image super-resolution (MedSR) is essential for improving diagnostic precision across diverse imaging modalities such as MRI, CT, X-ray, Ultraso

Memorization In Stable Diffusion Is Unexpectedly Driven by CLIP Embeddings

SafetyDGX agent

arXiv:2605.02908v1 Announce Type: new Abstract: Understanding how textual embeddings contribute to memorization in text-to-image diffusion models is crucial for both interpretability and safety. This

Metadata, Wavelet, and Time Aware Diffusion Models for Satellite Image Super Resolution

ResearchDGX agent

arXiv:2506.23566v2 Announce Type: replace Abstract: The acquisition of high-resolution satellite imagery is often constrained by the spatial and temporal limitations of satellite sensors, as well as t

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models

Model ReleasesDGX agent

arXiv:2605.03485v1 Announce Type: new Abstract: Multidimensional human understanding is essential for real-world applications such as film analysis and virtual digital humans, yet current LVLM benchma

MILE: Mixture of Incremental LoRA Experts for Continual Semantic Segmentation across Domains and Modalities

Model ReleasesDGX agent

arXiv:2605.03555v1 Announce Type: new Abstract: Continual semantic segmentation requires models to adapt to new domains or modalities without sacrificing performance on previously learned tasks. Exper

Mix3R: Mixing Feed-forward Reconstruction and Generative 3D Priors for Joint Multi-view Aligned 3D Reconstruction and Pose Estimation

SafetyDGX agent

arXiv:2605.03359v1 Announce Type: new Abstract: Recent trends in sparse-view 3D reconstruction have taken two different paths: feed-forward reconstruction that predicts pixel-aligned point maps withou

MK-ResRecon: Multi-Kernel Residual Framework for Texture-Aware 3D MRI Refinement from Sparse 2D Slices

ResearchDGX agent

arXiv:2605.03432v1 Announce Type: new Abstract: Magnetic Resonance Imaging (MRI) acquisition remains a time-intensive and patient-straining process, as prolonged scan dura- tions increase the likeliho

Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration

Model ReleasesDGX agent

arXiv:2605.03820v1 Announce Type: new Abstract: Multimodal learning often grapples with the challenge of low-quality data, which predominantly manifests as two facets: modality imbalance and noisy cor

Normalized Matching Transformer

SafetyDGX agent

arXiv:2503.17715v3 Announce Type: replace Abstract: We introduce the Normalized Matching Transformer (NMT), a deep learning approach for efficient and accurate sparse semantic keypoint matching betwee

NucEval: A Robust Evaluation Framework for Nuclear Instance Segmentation

ResearchDGX agent

arXiv:2605.03144v1 Announce Type: new Abstract: In computational pathology, nuclear instance segmentation is a fundamental task with many downstream clinical applications. With the advent of deep lear

One Sequence to Segment Them All: Efficient Data Augmentation for CT and MRI Cross-Domain 3D Spine Segmentation

HardwareDGX agent

arXiv:2605.03098v1 Announce Type: new Abstract: Deep learning-based medical image segmentation is increasingly used to support clinical diagnosis and develop new treatment strategies. However, model p

Optimizing Grasping in Legged Robots: A Deep Learning Approach to Loco-Manipulation

ResearchDGX agent

arXiv:2508.17466v3 Announce Type: replace-cross Abstract: This paper presents a deep learning framework designed to enhance the grasping capabilities of quadrupeds equipped with arms, with a focus on

Orientation-Aware Unsupervised Domain Adaptation for Brain Tumor Classification Across Multi-Modal MRI

SafetyDGX agent

arXiv:2605.03490v1 Announce Type: new Abstract: The clinical integration of deep learning models for brain tumor diagnosis in neuro-oncology is severely constrained by limited expert-annotated MRI dat

Ortho-Hydra: Orthogonalized Experts for DiT LoRA

ResearchDGX agent

arXiv:2605.03252v1 Announce Type: cross Abstract: LoRA fine-tuning of diffusion transformers (DiT) on multi-style data suffers from style bleed: a single low-rank residual cannot represent several dis

Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback

Model ReleasesDGX agent

arXiv:2605.03848v1 Announce Type: new Abstract: Estimating how well a person performs an action, rather than which action is performed, is central to coaching, rehabilitation, and talent identificatio

Physically Guided Visual Mass Estimation from a Single RGB Image

TutorialsDGX agent

arXiv:2601.20303v2 Announce Type: replace Abstract: Estimating object mass from visual input is challenging because mass depends jointly on geometric volume and material-dependent density, neither of

PriorNet: Prior-Guided Engagement Estimation from Face Video

Model ReleasesDGX agent

arXiv:2605.03615v1 Announce Type: new Abstract: Engagement estimation from face video remains challenging because facial evidence is often incomplete, labeled data are limited, and engagement annotati

PROBE: Probabilistic Occupancy BEV Encoding with Analytical Translation Robustness for 3D Place Recognition

Model ReleasesDGX agent

arXiv:2603.05965v2 Announce Type: replace-cross Abstract: We present PROBE (PRobabilistic Occupancy BEV Encoding), a learning-free LiDAR place recognition descriptor that models each BEV cell's occupa

Quantifying the human visual exposome with vision language models

AgentsDGX agent

arXiv:2605.03863v1 Announce Type: cross Abstract: The visual environment is a fundamental yet unquantified determinant of mental health. While the concept of the environmental exposome is well establi

Quaternion Wavelet-Conditioned Diffusion Models for Image Super-Resolution

ResearchDGX agent

arXiv:2505.00334v3 Announce Type: replace Abstract: Image Super-Resolution is a fundamental problem in computer vision with broad applications spacing from medical imaging to satellite analysis. The a

Raising the Ceiling: Better Empirical Fixation Densities for Saliency Benchmarking

Model ReleasesDGX agent

arXiv:2605.03885v1 Announce Type: new Abstract: Empirical fixation densities, spatial distributions estimated from human eye-tracking data, are foundational to saliency benchmarking. They directly sha

RD-ViT: Recurrent-Depth Vision Transformer for Semantic Segmentation with Reduced Data Dependence Extending the Recurrent-Depth Transformer Architecture to Dense Prediction

Model ReleasesDGX agent

arXiv:2605.03999v1 Announce Type: new Abstract: Vision Transformers (ViTs) achieve state-of-the-art segmentation accuracy but require large training datasets because each layer has unique parameters t

Real Image Denoising with Knowledge Distillation for High-Performance Mobile NPUs

Model ReleasesDGX agent

arXiv:2605.03680v1 Announce Type: new Abstract: While deep-learning-based image restoration has achieved unprecedented fidelity, deployment on mobile Neural Processing Units (NPUs) remains bottlenecke

Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models

ResearchDGX agent

arXiv:2605.02912v1 Announce Type: new Abstract: Video Anomaly Detection (VAD) has traditionally been framed as binary classification or outlier detection, providing neither interpretable reasoning nor

ReLeaf: Benchmarking Leaf Segmentation across Domains and Species

Model ReleasesDGX agent

arXiv:2605.03784v1 Announce Type: new Abstract: Rising global food demand and growing climate pressure increase the need for sustainable, precise agricultural practices. Automated, individualized plan

Reservoir property image slices from the Groningen gas field for image translation and segmentation

ResearchDGX agent

arXiv:2605.03942v1 Announce Type: new Abstract: Reservoir characterization workflows increasingly rely on image-based and machine-learning/deep learning or even generative AI approaches, but openly av

Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence

ResearchDGX agent

arXiv:2605.03650v1 Announce Type: new Abstract: The de facto approach in video object-centric learning maintains temporal consistency through learned dynamics modules that predict future object repres

RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation

Model ReleasesDGX agent

arXiv:2507.00435v2 Announce Type: replace-cross Abstract: We introduce RoboEval, a structured evaluation framework and benchmark for robotic manipulation that augments binary success with principled b

Robustness and Transferability of Pix2Geomodel for Bidirectional Facies Property Translation in a Complex Reservoir

ApplicationsDGX agent

arXiv:2605.03919v1 Announce Type: cross Abstract: Reservoir geomodeling is central to subsurface characterization, but it remains challenging because conditioning data are sparse, geological heterogen

RPBA-Net: An Interpretable Residual Pyramid Bilateral Affine Network for RAW-Domain ISP Enhancement

Local AiDGX agent

arXiv:2605.03626v1 Announce Type: new Abstract: To address module fragmentation, uninterpretable mappings, and deployment constraints in RAW-domain demosaicing, color correction, and detail enhancemen

Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses

SafetyDGX agent

arXiv:2605.02900v1 Announce Type: cross Abstract: Embodied Artificial Intelligence (Embodied AI) integrates perception, cognition, planning, and interaction into agents that operate in open-world, saf

Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning

Model ReleasesDGX agent

arXiv:2605.03189v1 Announce Type: new Abstract: Image captioning has become an important task in computer vision, enabling models to generate natural language descriptions of visual content. While sev

Skeleton-Snippet Contrastive Learning with Multiscale Feature Fusion for Action Localization

ResearchDGX agent

arXiv:2512.16504v3 Announce Type: replace Abstract: The self-supervised pretraining paradigm has achieved great success in learning 3D action representations for skeleton-based action recognition usin

SoDa2: Single-Stage Open-Set Domain Adaptation via Decoupled Alignment for Cross-Scene Hyperspectral Image Classification

SafetyDGX agent

arXiv:2605.03371v1 Announce Type: new Abstract: Cross-scene hyperspectral image (HSI) classification stands as a fundamental research topic in remote sensing, with extensive applications spanning vari

Sparse Data Tree Canopy Segmentation: Fine-Tuning Leading Pretrained Models on Only 150 Images

ResearchDGX agent

arXiv:2601.10931v2 Announce Type: replace Abstract: Tree canopy detection from aerial imagery is an important task for environmental monitoring, urban planning, and ecosystem analysis. Simulating real

Speculative Coupled Decoding for Training-Free Lossless Acceleration of Autoregressive Visual Generation

ResearchDGX agent

arXiv:2510.24211v2 Announce Type: replace Abstract: Autoregressive (AR) modeling has recently emerged as a promising new paradigm in visual generation, but its practical adoption is severely constrain

← Previous
1…153154155156157…209
Next →