AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Safety

Large Language Models are Universal Reasoners for Visual Generation

DGX agent

arXiv:2605.04040v1 Announce Type: new Abstract: Text-to-image generation has advanced rapidly with diffusion models, progressing from CLIP and T5 conditioning to unified systems where a single LLM bac

safetyarxiv-cs-cv
6 May 2026
Research
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View Captures

DGX agent

arXiv:2605.04035v1 Announce Type: new Abstract: We propose HeadsUp, a scalable feed-forward method for reconstructing high-quality 3D Gaussian heads from large-scale multi-camera setups. Our method em

researcharxiv-cs-cv
6 May 2026
Tutorials

Learning Discriminative Signed Distance Functions from Multi-scale Level-of-detail Features for 3D Anomaly Detection

DGX agent

arXiv:2605.03437v1 Announce Type: new Abstract: Detecting anomalies from 3D point clouds has received increasing attention in the field of computer vision, with some group-based or point-based methods

tutorialsarxiv-cs-cv
6 May 2026
Research

Learning to Segment using Summary Statistics and Weak Supervision

DGX agent

arXiv:2605.03059v1 Announce Type: new Abstract: Medical experts often manually segment images to obtain diagnostic statistics and discard the resulting annotations. We aim to train segmentation models

researcharxiv-cs-cv
6 May 2026
Hardware

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute

DGX agent

arXiv:2504.17816v3 Announce Type: replace Abstract: Subject-driven video generation (SDV-Gen) aims to produce videos of a specific subject by adapting a pretrained video model, enabling personalized a

hardwarearxiv-cs-cv
6 May 2026
Model Releases

Mantis: Mamba-native Tuning is Efficient for 3D Point Cloud Foundation Models

DGX agent

arXiv:2605.03438v1 Announce Type: new Abstract: Pre-trained 3D point cloud foundation models (PFMs) have demonstrated strong transferability across diverse downstream tasks. However, full fine-tuning

model-releasesarxiv-cs-cv
6 May 2026
Local Ai

MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding

DGX agent

arXiv:2605.03398v1 Announce Type: new Abstract: Video Temporal Grounding (VTG) faces a cross-modal semantic gap that often leads to background features being incorrectly aligned with the query, while

local-aiarxiv-cs-cv
6 May 2026
Research

MedSR-Vision: Deep Learning Framework for Multi-Domain Medical Image Super-Resolution

DGX agent

arXiv:2605.03343v1 Announce Type: new Abstract: Medical image super-resolution (MedSR) is essential for improving diagnostic precision across diverse imaging modalities such as MRI, CT, X-ray, Ultraso

researcharxiv-cs-cv
6 May 2026
Safety

Memorization In Stable Diffusion Is Unexpectedly Driven by CLIP Embeddings

DGX agent

arXiv:2605.02908v1 Announce Type: new Abstract: Understanding how textual embeddings contribute to memorization in text-to-image diffusion models is crucial for both interpretability and safety. This

safetyarxiv-cs-cv
6 May 2026
Research

Metadata, Wavelet, and Time Aware Diffusion Models for Satellite Image Super Resolution

DGX agent

arXiv:2506.23566v2 Announce Type: replace Abstract: The acquisition of high-resolution satellite imagery is often constrained by the spatial and temporal limitations of satellite sensors, as well as t

researcharxiv-cs-cv
6 May 2026
Model Releases

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models

DGX agent

arXiv:2605.03485v1 Announce Type: new Abstract: Multidimensional human understanding is essential for real-world applications such as film analysis and virtual digital humans, yet current LVLM benchma

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

MILE: Mixture of Incremental LoRA Experts for Continual Semantic Segmentation across Domains and Modalities

DGX agent

arXiv:2605.03555v1 Announce Type: new Abstract: Continual semantic segmentation requires models to adapt to new domains or modalities without sacrificing performance on previously learned tasks. Exper

model-releasesarxiv-cs-cv
6 May 2026
Safety

Mix3R: Mixing Feed-forward Reconstruction and Generative 3D Priors for Joint Multi-view Aligned 3D Reconstruction and Pose Estimation

DGX agent

arXiv:2605.03359v1 Announce Type: new Abstract: Recent trends in sparse-view 3D reconstruction have taken two different paths: feed-forward reconstruction that predicts pixel-aligned point maps withou

safetyarxiv-cs-cv
6 May 2026
Research

MK-ResRecon: Multi-Kernel Residual Framework for Texture-Aware 3D MRI Refinement from Sparse 2D Slices

DGX agent

arXiv:2605.03432v1 Announce Type: new Abstract: Magnetic Resonance Imaging (MRI) acquisition remains a time-intensive and patient-straining process, as prolonged scan dura- tions increase the likeliho

researcharxiv-cs-cv
6 May 2026
Model Releases

Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration

DGX agent

arXiv:2605.03820v1 Announce Type: new Abstract: Multimodal learning often grapples with the challenge of low-quality data, which predominantly manifests as two facets: modality imbalance and noisy cor

model-releasesarxiv-cs-cv
6 May 2026
Safety

Normalized Matching Transformer

DGX agent

arXiv:2503.17715v3 Announce Type: replace Abstract: We introduce the Normalized Matching Transformer (NMT), a deep learning approach for efficient and accurate sparse semantic keypoint matching betwee

safetyarxiv-cs-cv
6 May 2026
Research

NucEval: A Robust Evaluation Framework for Nuclear Instance Segmentation

DGX agent

arXiv:2605.03144v1 Announce Type: new Abstract: In computational pathology, nuclear instance segmentation is a fundamental task with many downstream clinical applications. With the advent of deep lear

researcharxiv-cs-cv
6 May 2026
Hardware

One Sequence to Segment Them All: Efficient Data Augmentation for CT and MRI Cross-Domain 3D Spine Segmentation

DGX agent

arXiv:2605.03098v1 Announce Type: new Abstract: Deep learning-based medical image segmentation is increasingly used to support clinical diagnosis and develop new treatment strategies. However, model p

hardwarearxiv-cs-cv
6 May 2026
Research

Optimizing Grasping in Legged Robots: A Deep Learning Approach to Loco-Manipulation

DGX agent

arXiv:2508.17466v3 Announce Type: replace-cross Abstract: This paper presents a deep learning framework designed to enhance the grasping capabilities of quadrupeds equipped with arms, with a focus on

researcharxiv-cs-cv
6 May 2026
Safety

Orientation-Aware Unsupervised Domain Adaptation for Brain Tumor Classification Across Multi-Modal MRI

DGX agent

arXiv:2605.03490v1 Announce Type: new Abstract: The clinical integration of deep learning models for brain tumor diagnosis in neuro-oncology is severely constrained by limited expert-annotated MRI dat

safetyarxiv-cs-cv
6 May 2026
Research

Ortho-Hydra: Orthogonalized Experts for DiT LoRA

DGX agent

arXiv:2605.03252v1 Announce Type: cross Abstract: LoRA fine-tuning of diffusion transformers (DiT) on multi-style data suffers from style bleed: a single low-rank residual cannot represent several dis

researcharxiv-cs-cv
6 May 2026
Model Releases

Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback

DGX agent

arXiv:2605.03848v1 Announce Type: new Abstract: Estimating how well a person performs an action, rather than which action is performed, is central to coaching, rehabilitation, and talent identificatio

model-releasesarxiv-cs-cv
6 May 2026
Tutorials

Physically Guided Visual Mass Estimation from a Single RGB Image

DGX agent

arXiv:2601.20303v2 Announce Type: replace Abstract: Estimating object mass from visual input is challenging because mass depends jointly on geometric volume and material-dependent density, neither of

tutorialsarxiv-cs-cv
6 May 2026
Model Releases

PriorNet: Prior-Guided Engagement Estimation from Face Video

DGX agent

arXiv:2605.03615v1 Announce Type: new Abstract: Engagement estimation from face video remains challenging because facial evidence is often incomplete, labeled data are limited, and engagement annotati

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

PROBE: Probabilistic Occupancy BEV Encoding with Analytical Translation Robustness for 3D Place Recognition

DGX agent

arXiv:2603.05965v2 Announce Type: replace-cross Abstract: We present PROBE (PRobabilistic Occupancy BEV Encoding), a learning-free LiDAR place recognition descriptor that models each BEV cell's occupa

model-releasesarxiv-cs-cv
6 May 2026
Agents

Quantifying the human visual exposome with vision language models

DGX agent

arXiv:2605.03863v1 Announce Type: cross Abstract: The visual environment is a fundamental yet unquantified determinant of mental health. While the concept of the environmental exposome is well establi

agentsarxiv-cs-cv
6 May 2026
Research

Quaternion Wavelet-Conditioned Diffusion Models for Image Super-Resolution

DGX agent

arXiv:2505.00334v3 Announce Type: replace Abstract: Image Super-Resolution is a fundamental problem in computer vision with broad applications spacing from medical imaging to satellite analysis. The a

researcharxiv-cs-cv
6 May 2026
Model Releases

Raising the Ceiling: Better Empirical Fixation Densities for Saliency Benchmarking

DGX agent

arXiv:2605.03885v1 Announce Type: new Abstract: Empirical fixation densities, spatial distributions estimated from human eye-tracking data, are foundational to saliency benchmarking. They directly sha

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

RD-ViT: Recurrent-Depth Vision Transformer for Semantic Segmentation with Reduced Data Dependence Extending the Recurrent-Depth Transformer Architecture to Dense Prediction

DGX agent

arXiv:2605.03999v1 Announce Type: new Abstract: Vision Transformers (ViTs) achieve state-of-the-art segmentation accuracy but require large training datasets because each layer has unique parameters t

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

Real Image Denoising with Knowledge Distillation for High-Performance Mobile NPUs

DGX agent

arXiv:2605.03680v1 Announce Type: new Abstract: While deep-learning-based image restoration has achieved unprecedented fidelity, deployment on mobile Neural Processing Units (NPUs) remains bottlenecke

model-releasesarxiv-cs-cv
6 May 2026
Research

Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models

DGX agent

arXiv:2605.02912v1 Announce Type: new Abstract: Video Anomaly Detection (VAD) has traditionally been framed as binary classification or outlier detection, providing neither interpretable reasoning nor

researcharxiv-cs-cv
6 May 2026
Model Releases

ReLeaf: Benchmarking Leaf Segmentation across Domains and Species

DGX agent

arXiv:2605.03784v1 Announce Type: new Abstract: Rising global food demand and growing climate pressure increase the need for sustainable, precise agricultural practices. Automated, individualized plan

model-releasesarxiv-cs-cv
6 May 2026
Research

Reservoir property image slices from the Groningen gas field for image translation and segmentation

DGX agent

arXiv:2605.03942v1 Announce Type: new Abstract: Reservoir characterization workflows increasingly rely on image-based and machine-learning/deep learning or even generative AI approaches, but openly av

researcharxiv-cs-cv
6 May 2026
Research

Rethinking Temporal Consistency in Video Object-Centric Learning: From Prediction to Correspondence

DGX agent

arXiv:2605.03650v1 Announce Type: new Abstract: The de facto approach in video object-centric learning maintains temporal consistency through learned dynamics modules that predict future object repres

researcharxiv-cs-cv
6 May 2026
Model Releases

RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation

DGX agent

arXiv:2507.00435v2 Announce Type: replace-cross Abstract: We introduce RoboEval, a structured evaluation framework and benchmark for robotic manipulation that augments binary success with principled b

model-releasesarxiv-cs-cv
6 May 2026
Applications

Robustness and Transferability of Pix2Geomodel for Bidirectional Facies Property Translation in a Complex Reservoir

DGX agent

arXiv:2605.03919v1 Announce Type: cross Abstract: Reservoir geomodeling is central to subsurface characterization, but it remains challenging because conditioning data are sparse, geological heterogen

applicationsarxiv-cs-cv
6 May 2026
Local Ai

RPBA-Net: An Interpretable Residual Pyramid Bilateral Affine Network for RAW-Domain ISP Enhancement

DGX agent

arXiv:2605.03626v1 Announce Type: new Abstract: To address module fragmentation, uninterpretable mappings, and deployment constraints in RAW-domain demosaicing, color correction, and detail enhancemen

local-aiarxiv-cs-cv
6 May 2026
Safety

Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses

DGX agent

arXiv:2605.02900v1 Announce Type: cross Abstract: Embodied Artificial Intelligence (Embodied AI) integrates perception, cognition, planning, and interaction into agents that operate in open-world, saf

safetyarxiv-cs-cv
6 May 2026
Model Releases

Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning

DGX agent

arXiv:2605.03189v1 Announce Type: new Abstract: Image captioning has become an important task in computer vision, enabling models to generate natural language descriptions of visual content. While sev

model-releasesarxiv-cs-cv
6 May 2026
Research

Skeleton-Snippet Contrastive Learning with Multiscale Feature Fusion for Action Localization

DGX agent

arXiv:2512.16504v3 Announce Type: replace Abstract: The self-supervised pretraining paradigm has achieved great success in learning 3D action representations for skeleton-based action recognition usin

researcharxiv-cs-cv
6 May 2026
Safety

SoDa2: Single-Stage Open-Set Domain Adaptation via Decoupled Alignment for Cross-Scene Hyperspectral Image Classification

DGX agent

arXiv:2605.03371v1 Announce Type: new Abstract: Cross-scene hyperspectral image (HSI) classification stands as a fundamental research topic in remote sensing, with extensive applications spanning vari

safetyarxiv-cs-cv
6 May 2026
Research

Sparse Data Tree Canopy Segmentation: Fine-Tuning Leading Pretrained Models on Only 150 Images

DGX agent

arXiv:2601.10931v2 Announce Type: replace Abstract: Tree canopy detection from aerial imagery is an important task for environmental monitoring, urban planning, and ecosystem analysis. Simulating real

researcharxiv-cs-cv
6 May 2026
Research

Speculative Coupled Decoding for Training-Free Lossless Acceleration of Autoregressive Visual Generation

DGX agent

arXiv:2510.24211v2 Announce Type: replace Abstract: Autoregressive (AR) modeling has recently emerged as a promising new paradigm in visual generation, but its practical adoption is severely constrain

researcharxiv-cs-cv
6 May 2026
Model Releases

StateVLM: A State-Aware Vision-Language Model for Robotic Affordance Reasoning

DGX agent

arXiv:2605.03927v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown remarkable performance in various robotic tasks, as they can perceive visual information and understand natural

model-releasesarxiv-cs-cv
6 May 2026
Safety

Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation

DGX agent

arXiv:2605.03849v1 Announce Type: new Abstract: Distillation-based acceleration has become foundational for making autoregressive streaming video diffusion models practical, with distribution matching

safetyarxiv-cs-cv
6 May 2026
Applications

Synthetic Data Generation for Long-Tail Medical Image Classification: A Case Study in Skin Lesions

DGX agent

arXiv:2605.03221v1 Announce Type: new Abstract: Long-tailed class distributions are pervasive in multi-class medical datasets and pose significant challenges for deep learning models which typically u

applicationsarxiv-cs-cv
6 May 2026
Local Ai

TACO: Trajectory Aligning Cross-view Optimisation

DGX agent

arXiv:2605.03315v1 Announce Type: new Abstract: Cross-View Geo-localisation (CVGL) matches ground imagery against satellite tiles to give absolute position fixes, an alternative to GNSS where signals

local-aiarxiv-cs-cv
6 May 2026
Model Releases

Task-Aware Scanning Parameter Configuration for Robotic Inspection Using Vision Language Embeddings and Hyperdimensional Computing

DGX agent

arXiv:2605.03909v1 Announce Type: cross Abstract: Robotic laser profiling is widely used for dimensional verification and surface inspection, yet measurement fidelity is often dominated by sensor conf

model-releasesarxiv-cs-cv
6 May 2026
← Previous
1…194195196197198…263
Next →