AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Spectral Surgery: Class-Targeted Post-Hoc Rebalancing via Hessian Spike Perturbation

DGX agent

arXiv:2605.07790v1 Announce Type: cross Abstract: The Hessian spectrum of trained deep networks exhibits a characteristic structure: a continuous bulk of near-zero eigenvalues and a small number of la

researcharxiv-cs-cv
11 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

SphereVAD: Training-Free Video Anomaly Detection via Geodesic Inference on the Unit Hypersphere

DGX agent

arXiv:2605.08003v1 Announce Type: new Abstract: Video anomaly detection (VAD) aims to automatically identify events that deviate from normal patterns in untrimmed surveillance videos. Existing methods

researcharxiv-cs-cv
11 May 2026
Applications

SplatWeaver: Learning to Allocate Gaussian Primitives for Generalizable Novel View Synthesis

DGX agent

arXiv:2605.07287v1 Announce Type: new Abstract: Generalizable novel view synthesis aims to render unseen views from uncalibrated input images without requiring per-scene optimization. Recent feed-forw

applicationsarxiv-cs-cv
11 May 2026
Model Releases

SR^2-LoRA: Self-Rectifying Inter-layer Relations in Low-Rank Adaptation for Class-Incremental Learning

DGX agent

arXiv:2605.07420v1 Announce Type: cross Abstract: Pre-trained models with parameter-efficient fine-tuning (PEFT) have demonstrated promising potential for class-incremental learning (CIL), yet catastr

model-releasesarxiv-cs-cv
11 May 2026
Local Ai

ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation

DGX agent

arXiv:2605.07390v1 Announce Type: new Abstract: Generative models have achieved success in producing apparently coherent 2D videos, but remain challenging in the physical world due to lack of 4D spati

local-aiarxiv-cs-cv
11 May 2026
Research

STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation

DGX agent

arXiv:2605.08029v1 Announce Type: new Abstract: Deep generative models have advanced rapidly across text and vision, motivating unified multimodal systems that can understand, reason over, and generat

researcharxiv-cs-cv
11 May 2026
Research

Stochastic Transition-Map Distillation for Fast Probabilistic Inference

DGX agent

arXiv:2605.07661v1 Announce Type: cross Abstract: Diffusion models achieve strong generation quality, diversity, and distribution coverage, but their performance often comes with expensive inference.

researcharxiv-cs-cv
11 May 2026
Model Releases

Structure Over Scale: Learning Visual Reasoning from Pedagogical Video

DGX agent

arXiv:2601.23251v2 Announce Type: replace Abstract: State-of-the-art vision-language models (VLMs) score impressively on video benchmarks yet stumble on basic visual reasoning tasks involving spatial

model-releasesarxiv-cs-cv
11 May 2026
Research

Surgical Visual Understanding (SurgVU) Dataset

DGX agent

arXiv:2501.09209v2 Announce Type: replace Abstract: Owing to recent advances in machine learning and the ability to harvest large amounts of data during robotic-assisted surgeries, surgical data scien

researcharxiv-cs-cv
11 May 2026
Model Releases

SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes

DGX agent

arXiv:2602.04939v2 Announce Type: replace Abstract: Modern T2V/I2V generators synthesize people increasingly hard to distinguish from authentic footage, while current evaluation suites lag: legacy ben

model-releasesarxiv-cs-cv
11 May 2026
Tutorials

Tables Guide Vision: Learning to See the Heart through Tabular Data

DGX agent

arXiv:2503.14998v3 Announce Type: replace Abstract: Contrastive learning methods in computer vision typically rely on augmented views of the same image or multimodal pretraining strategies that align

tutorialsarxiv-cs-cv
11 May 2026
Model Releases

TAS-LoRA: Transformer Architecture Search with Mixture-of-LoRA Experts

DGX agent

arXiv:2605.07256v1 Announce Type: new Abstract: Transformer architecture search (TAS) discovers optimal vision transformer (ViT) architectures automatically, reducing human effort to manually design V

model-releasesarxiv-cs-cv
11 May 2026
Research

Task-Oriented Communication for Human Action Understanding via Edge-Cloud Co-Inference

DGX agent

arXiv:2605.07354v1 Announce Type: cross Abstract: The expanding application of smart sensing has created a growing demand for the accurate understanding of human action at the network edge. Traditiona

researcharxiv-cs-cv
11 May 2026
Tutorials

Task Relevance Is Not Local Replaceability: A Two-Axis View of Channel Information

DGX agent

arXiv:2605.07086v1 Announce Type: new Abstract: Channel importance in vision networks is usually summarized by a single score. That summary hides two different questions: how much a channel is related

tutorialsarxiv-cs-cv
11 May 2026
Research

Teacher-Feature Drifting: One-Step Diffusion Distillation with Pretrained Diffusion Representations

DGX agent

arXiv:2605.07327v1 Announce Type: new Abstract: Sampling from pretrained diffusion and flow-matching models typically requires many forward passes to generate diverse and high-fidelity images. Existin

researcharxiv-cs-cv
11 May 2026
Local Ai

Teaching Prompts to Coordinate: Hierarchical Layer-Grouped Prompt Tuning for Continual Learning

DGX agent

arXiv:2511.12090v3 Announce Type: replace Abstract: Prompt-based continual learning methods fine-tune only a small set of additional learnable parameters while keeping the pre-trained model's paramete

local-aiarxiv-cs-cv
11 May 2026
Research

Towards Explainable Industrial Anomaly Detection via Knowledge-Guided Latent Reasoning

DGX agent

arXiv:2602.09850v2 Announce Type: replace Abstract: Industrial anomaly detection demands precise reasoning over fine-grained defect patterns. However, existing multimodal large language models (MLLMs)

researcharxiv-cs-cv
11 May 2026
Safety

Towards Fairness under Label Bias in Image Segmentation: Impact, Measurement and Mitigation

DGX agent

arXiv:2605.06891v1 Announce Type: new Abstract: Labeled datasets reflect the biases of their annotation pipelines, which sometimes introduce label bias: group-conditional label errors that cause syste

safetyarxiv-cs-cv
11 May 2026
Research

Towards Highly-Constrained Human Motion Generation with Retrieval-Guided Diffusion Noise Optimization

DGX agent

arXiv:2605.08054v1 Announce Type: new Abstract: Generating human motion that satisfies customized zero-shot goal functions, enabling applications such as controllable character animation and behavior

researcharxiv-cs-cv
11 May 2026
Local Ai

Towards multi-modal forgery representation learning for AI-generated video detection and localization

DGX agent

arXiv:2605.07232v1 Announce Type: new Abstract: Recent advances in generative AI have democratized video creation at scale. AI-generated videos, including partially manipulated clips across visual and

local-aiarxiv-cs-cv
11 May 2026
Applications

Towards Photorealistic and Efficient Bokeh Rendering via Diffusion Framework

DGX agent

arXiv:2605.07429v1 Announce Type: new Abstract: Existing mobile devices are constrained by compact optical designs, such as small apertures, which make it difficult to produce natural, optically reali

applicationsarxiv-cs-cv
11 May 2026
Model Releases

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos

DGX agent

arXiv:2605.07593v1 Announce Type: new Abstract: Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory streams,

model-releasesarxiv-cs-cv
11 May 2026
Safety

TRAJGANR: Trajectory-Centric Urban Multimodal Learning via Geospatially Aligned Neural Representations

DGX agent

arXiv:2605.06990v1 Announce Type: new Abstract: Multimodal self-supervised learning (MSSL) has emerged as a key paradigm for pretraining geospatial foundation models. However, existing geospatial MSSL

safetyarxiv-cs-cv
11 May 2026
Research

TRAS: An Interactive Software for Tracing Tree Ring Cross Sections

DGX agent

arXiv:2605.08025v1 Announce Type: new Abstract: Tree ring marking remains a key step in dendrometry and dendrochronology, but it is often performed manually, making the process time-consuming, subject

researcharxiv-cs-cv
11 May 2026
Local Ai

TriDE: Triangle-Consistent Translation Directions for Global Camera Pose Estimation

DGX agent

arXiv:2605.06889v1 Announce Type: new Abstract: Pairwise translation directions are a key input to camera location estimation in global structure-from-motion. Existing estimators usually process each

local-aiarxiv-cs-cv
11 May 2026
Research

TriP: A Triangle Puzzle Approach to Robust Translation Averaging

DGX agent

arXiv:2605.07143v1 Announce Type: new Abstract: Translation averaging aims to recover camera locations from pairwise relative translation directions and is a fundamental component of global Structure-

researcharxiv-cs-cv
11 May 2026
Tutorials

TRUST: Test-Time Refinement using Uncertainty-Guided SSM Traverses

DGX agent

arXiv:2509.22813v2 Announce Type: replace Abstract: State Space Models (SSMs) have emerged as efficient alternatives to Vision Transformers (ViTs), with VMamba standing out as a pioneering architectur

tutorialsarxiv-cs-cv
11 May 2026
Research

Uncertainty Quantification for Cardiac Shape Reconstruction with Deep Signed Distance Functions via MCMC methods

DGX agent

arXiv:2605.07987v1 Announce Type: cross Abstract: Atlas-based approaches allow high-quality, patient-specific shape reconstructions of cardiac anatomy from sparse and/or noisy data such as point cloud

researcharxiv-cs-cv
11 May 2026
Applications

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models

DGX agent

arXiv:2605.07148v1 Announce Type: new Abstract: Decades of cognitive science establish that humans navigate environments by forming cognitive maps, defined as allocentric and topology-preserving repre

applicationsarxiv-cs-cv
11 May 2026
Safety

UniD-Shift: Towards Unified Semantic Segmentation via Interpretable Share-Private Multimodal Decomposition

DGX agent

arXiv:2605.07356v1 Announce Type: new Abstract: Semantic segmentation of large-scale 3D point clouds is crucial for applications such as autonomous driving and urban digital twins. However, the sparse

safetyarxiv-cs-cv
11 May 2026
Research

UniISP: A Unified ISP Framework for Both Human and Machine Vision

DGX agent

arXiv:2605.07359v1 Announce Type: new Abstract: Compared to RGB images, raw sensor data provides a richer representation of information, which is crucial for accurate recognition, particularly under c

researcharxiv-cs-cv
11 May 2026
Tutorials

UniV2D: Bridging Visual Restoration and Semantic Perception for Underwater Salient Object Detection

DGX agent

arXiv:2605.07146v1 Announce Type: new Abstract: Underwater salient object detection (USOD) plays a vital role in marine vision tasks but remains fundamentally challenging due to severe visual degradat

tutorialsarxiv-cs-cv
11 May 2026
Safety

VDEGaussian: Video Diffusion Enhanced 4D Gaussian Splatting for Dynamic Urban Scenes Modeling

DGX agent

arXiv:2508.02129v2 Announce Type: replace Abstract: Dynamic urban scene modeling is a rapidly evolving area with broad applications. While current approaches leveraging neural radiance fields or Gauss

safetyarxiv-cs-cv
11 May 2026
Research

Velocity-Space 3D Asset Editing

DGX agent

arXiv:2605.07385v1 Announce Type: cross Abstract: Editing a 3D asset locally, modifying a target region while preserving the rest, is a fundamental requirement of native 3D editing. Existing methods e

researcharxiv-cs-cv
11 May 2026
Tutorials

VesselRW: Weakly Supervised Subcutaneous Vessel Segmentation via Learned Random Walk Propagation

DGX agent

arXiv:2508.06819v3 Announce Type: replace Abstract: The task of parsing subcutaneous vessels in clinical images is often hindered by the high cost and limited availability of ground truth data, as wel

tutorialsarxiv-cs-cv
11 May 2026
Research

VIMCAN: Visual-Inertial 3D Human Pose Estimation with Hybrid Mamba-Cross-Attention Network

DGX agent

arXiv:2605.07552v1 Announce Type: new Abstract: The rapid advances in deep learning have significantly enhanced the accuracy of multimodal 3D human pose estimation (HPE). However, the state-of-the-art

researcharxiv-cs-cv
11 May 2026
Research

Weather-Robust Scene Semantics with Vision-Aligned 4D Radar

DGX agent

arXiv:2605.07367v1 Announce Type: cross Abstract: Cameras and LiDAR degrade in rain, fog, and snow, while millimeter-wave radar remains largely unaffected. We align a radar encoder to frozen SigLIP vi

researcharxiv-cs-cv
11 May 2026
Research

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion

DGX agent

arXiv:2605.07915v1 Announce Type: new Abstract: Tokenizers are a crucial component of latent diffusion models, as they define the latent space in which diffusion models operate. However, existing toke

researcharxiv-cs-cv
11 May 2026
Research

Zero-Shot Satellite Image Retrieval through Joint Embeddings: Application to Crisis Response

DGX agent

arXiv:2605.05405v2 Announce Type: replace Abstract: Semantic search of Earth observation archives remains challenging. Visual foundation models such as CLAY produce rich embeddings of satellite imager

researcharxiv-cs-cv
11 May 2026
Local Ai

3D Ultrasound-Derived Pseudo-CT Synthesis Using a Transformer-Augmented Residual Network for Real-Time Operator Guidance

DGX agent

arXiv:2605.04856v1 Announce Type: new Abstract: Computed tomography (CT) is indispensable for clinical diagnosis and image-guided interventions but exposes patients to ionizing radiation, motivating t

local-aiarxiv-cs-cv
7 May 2026
Research

A Bayesian Approach for Task-Specific Next-Best-View Selection with Uncertain Geometry

DGX agent

arXiv:2605.05095v1 Announce Type: cross Abstract: We develop a framework for task-specific active next-best-view selection in 3D reconstruction from point clouds, by casting the problem in the languag

researcharxiv-cs-cv
7 May 2026
Safety

A cross-modal network for facial expression recognition

DGX agent

arXiv:2605.04439v1 Announce Type: new Abstract: Deep neural networks enriched with structural information have been widely employed for facial expression recognition tasks. However, these methods ofte

safetyarxiv-cs-cv
7 May 2026
Research

A large-scale heterogeneous 3D magnetic resonance brain imaging dataset for self-supervised learning

DGX agent

arXiv:2506.14432v3 Announce Type: replace-cross Abstract: We present FOMO260K, a large-scale, heterogeneous dataset of 260,927 brain Magnetic Resonance Imaging (MRI) scans from 77,589 MRI sessions and

researcharxiv-cs-cv
7 May 2026
Research

A Reconstruction System for Industrial Pipeline Inner Walls Using Panoramic Image Stitching with Endoscopic Imaging

DGX agent

arXiv:2603.00714v2 Announce Type: replace Abstract: Visual analysis and reconstruction of pipeline inner walls remain challenging in industrial inspection scenarios. This paper presents a dedicated re

researcharxiv-cs-cv
7 May 2026
Model Releases

A unified Benchmark for Multi-Frame Image Restoration under Severe Refractive Warping

DGX agent

arXiv:2605.05079v1 Announce Type: new Abstract: Video sequence capturing through refractive dynamic media, such as a turbulent air or water surface, often suffer from severe geometric distortions and

model-releasesarxiv-cs-cv
7 May 2026
Research

Adapting Medical Vision Foundation Models for Volumetric Medical Image Segmentation via Active Learning and Selective Semi-supervised Fine-tuning

DGX agent

arXiv:2509.10784v3 Announce Type: replace-cross Abstract: Medical vision foundation models remain limited in downstream tasks, particularly volumetric medical image segmentation. While fine-tuning on

researcharxiv-cs-cv
7 May 2026
Research

Advancing Aesthetic Image Generation via Composition Transfer

DGX agent

arXiv:2605.04609v1 Announce Type: new Abstract: Composition is a cornerstone of visual aesthetics, influencing the appeal of an image. While its principles operate independently of specific content, i

researcharxiv-cs-cv
7 May 2026
Model Releases

Aes3D: Aesthetic Assessment in 3D Gaussian Splatting

DGX agent

arXiv:2605.05155v1 Announce Type: new Abstract: As 3D Gaussian Splatting (3DGS) gains attention in immersive media and digital content creation, assessing the aesthetics of 3D scenes becomes important

model-releasesarxiv-cs-cv
7 May 2026
← Previous
1…189190191192193…263
Next →