AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
11 May 2026

RedDiffuser: Auditing Multimodal Safety Failures in Vision-Language Models via Reinforced Diffusion

Model ReleasesDGX agent

arXiv:2503.06223v5 Announce Type: replace Abstract: Large Vision-Language Models (VLMs) are increasingly deployed in open-ended environments, where ensuring reliable safety under multimodal inputs is

Rethinking Dense Optical Flow without Test-Time Scaling

Model ReleasesDGX agent

arXiv:2605.08000v1 Announce Type: new Abstract: Recent progress in dense optical flow has been driven by increasingly complex architectures and multi-step refinement for test-time scaling. While these

RL-RIG: A Generative Spatial Reasoner via Intrinsic Reflection

ResearchDGX agent

arXiv:2602.19974v2 Announce Type: replace Abstract: Recent advancements in image generation have achieved impressive results in producing high-quality images. However, existing image generation models


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

S2M-Net: Spectral-Spatial Mixing for Medical Image Segmentation with Morphology-Aware Adaptive Loss

Model ReleasesDGX agent

arXiv:2601.01285v2 Announce Type: replace Abstract: Medical image segmentation requires balancing local precision for boundary-critical clinical applications, global context for anatomical coherence,

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models

SafetyDGX agent

arXiv:2605.07800v1 Announce Type: new Abstract: Recent video diffusion models (VDMs) synthesize visually convincing clips, yet still drop entities, mis-bind attributes, and weaken the interactions spe

Sat3R: Satellite DSM Reconstruction via RPC-Aware Depth Fine-tuning

Model ReleasesDGX agent

arXiv:2605.07264v1 Announce Type: new Abstract: Accurate Digital Surface Model (DSM) reconstruction from satellite imagery is critical for applications such as disaster response, urban planning, and l

SatSurfGS: Generalizable 2D Gaussian Splatting for Sparse-View Satellite Surface Reconstruction

Model ReleasesDGX agent

arXiv:2605.07181v1 Announce Type: new Abstract: Sparse-view satellite image surface reconstruction remains highly challenging, fundamentally because the reliability of multi-view matching under satell

Saving Foundation Flow-Matching Priors for Inverse Problems

ResearchDGX agent

arXiv:2511.16520v2 Announce Type: replace-cross Abstract: Foundation flow-matching (FM) models promise a universal prior for solving inverse problems (IPs), yet today they trail behind domain-specific

Scaling Continual Learning to 300+ Tasks with Bi-Level Routing Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2602.03473v2 Announce Type: replace-cross Abstract: Continual learning, especially class-incremental learning (CIL), on the basis of a pre-trained model (PTM) has garnered substantial research i

See Tomorrow, Act Today: Foresight-Driven Autonomous Driving

AgentsDGX agent

arXiv:2605.07195v1 Announce Type: new Abstract: Current end-to-end autonomous driving planners are fundamentally reactive: they condition on historical and present observations to predict future actio

Seeing Across Skies and Streets: Feedforward 3D Reconstruction from Satellite, Drone, and Ground Images

Local AiDGX agent

arXiv:2605.07978v1 Announce Type: new Abstract: Cross-view localization classically asks: where does this ground image lie on the satellite tile? Existing methods are typically limited to 3-DoF estima

SemanticDialect: Semantic-Aware Mixed-Format Quantization for Video Diffusion Transformers

HardwareDGX agent

arXiv:2603.02883v3 Announce Type: replace Abstract: Diffusion Transformers (DiTs) achieve state-of-the-art video generation quality, but their substantial memory and computational footprints hinder ed

Setting-Matched and Semantics-Scaled Benchmarking of One-Step Generative Models Against Multistep Diffusion and Flow Models

Model ReleasesDGX agent

arXiv:2603.14186v4 Announce Type: replace Abstract: State-of-the-art text-to-image models produce high-quality images, but inference remains expensive as generation requires several sequential ODE or

ShellfishNet: A Domain-Specific Benchmark for Visual Recognition of Marine Molluscs

Model ReleasesDGX agent

arXiv:2605.07338v1 Announce Type: new Abstract: The decline of global shellfish biodiversity poses a severe threat to coastal ecosystems. Although artificial intelligence (AI) technologies show potent

SIMI: Self-information Mining Network for Low-light Image Enhancement

ApplicationsDGX agent

arXiv:2605.07767v1 Announce Type: new Abstract: Poor lighting conditions significantly impact image quality, posing substantial challenges for image editing and visualization. Many existing enhancemen

SoftSAE: Dynamic Top-K Selection for Adaptive Sparse Autoencoders

Local AiDGX agent

arXiv:2605.06610v2 Announce Type: replace-cross Abstract: Sparse Autoencoders (SAEs) have become an important tool in mechanistic interpretability, helping to analyze internal representations in both

SoLAR: Error-Resilient Streamable Long-Horizon Free-Viewpoint Video Reconstruction with Anchor Activation and Latent Recalibration

ResearchDGX agent

arXiv:2605.07346v1 Announce Type: new Abstract: Free-Viewpoint Video (FVV) has emerged as a cornerstone of next-generation immersive media systems and attracted widespread attention. Previous methods

Spectral Surgery: Class-Targeted Post-Hoc Rebalancing via Hessian Spike Perturbation

ResearchDGX agent

arXiv:2605.07790v1 Announce Type: cross Abstract: The Hessian spectrum of trained deep networks exhibits a characteristic structure: a continuous bulk of near-zero eigenvalues and a small number of la

SphereVAD: Training-Free Video Anomaly Detection via Geodesic Inference on the Unit Hypersphere

ResearchDGX agent

arXiv:2605.08003v1 Announce Type: new Abstract: Video anomaly detection (VAD) aims to automatically identify events that deviate from normal patterns in untrimmed surveillance videos. Existing methods

SplatWeaver: Learning to Allocate Gaussian Primitives for Generalizable Novel View Synthesis

ApplicationsDGX agent

arXiv:2605.07287v1 Announce Type: new Abstract: Generalizable novel view synthesis aims to render unseen views from uncalibrated input images without requiring per-scene optimization. Recent feed-forw

SR^2-LoRA: Self-Rectifying Inter-layer Relations in Low-Rank Adaptation for Class-Incremental Learning

Model ReleasesDGX agent

arXiv:2605.07420v1 Announce Type: cross Abstract: Pre-trained models with parameter-efficient fine-tuning (PEFT) have demonstrated promising potential for class-incremental learning (CIL), yet catastr

ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation

Local AiDGX agent

arXiv:2605.07390v1 Announce Type: new Abstract: Generative models have achieved success in producing apparently coherent 2D videos, but remain challenging in the physical world due to lack of 4D spati

STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation

ResearchDGX agent

arXiv:2605.08029v1 Announce Type: new Abstract: Deep generative models have advanced rapidly across text and vision, motivating unified multimodal systems that can understand, reason over, and generat

Stochastic Transition-Map Distillation for Fast Probabilistic Inference

ResearchDGX agent

arXiv:2605.07661v1 Announce Type: cross Abstract: Diffusion models achieve strong generation quality, diversity, and distribution coverage, but their performance often comes with expensive inference.

Structure Over Scale: Learning Visual Reasoning from Pedagogical Video

Model ReleasesDGX agent

arXiv:2601.23251v2 Announce Type: replace Abstract: State-of-the-art vision-language models (VLMs) score impressively on video benchmarks yet stumble on basic visual reasoning tasks involving spatial

Surgical Visual Understanding (SurgVU) Dataset

ResearchDGX agent

arXiv:2501.09209v2 Announce Type: replace Abstract: Owing to recent advances in machine learning and the ability to harvest large amounts of data during robotic-assisted surgeries, surgical data scien

SynthForensics: Benchmarking and Evaluating People-Centric Synthetic Video Deepfakes

Model ReleasesDGX agent

arXiv:2602.04939v2 Announce Type: replace Abstract: Modern T2V/I2V generators synthesize people increasingly hard to distinguish from authentic footage, while current evaluation suites lag: legacy ben

Tables Guide Vision: Learning to See the Heart through Tabular Data

TutorialsDGX agent

arXiv:2503.14998v3 Announce Type: replace Abstract: Contrastive learning methods in computer vision typically rely on augmented views of the same image or multimodal pretraining strategies that align

TAS-LoRA: Transformer Architecture Search with Mixture-of-LoRA Experts

Model ReleasesDGX agent

arXiv:2605.07256v1 Announce Type: new Abstract: Transformer architecture search (TAS) discovers optimal vision transformer (ViT) architectures automatically, reducing human effort to manually design V

Task-Oriented Communication for Human Action Understanding via Edge-Cloud Co-Inference

ResearchDGX agent

arXiv:2605.07354v1 Announce Type: cross Abstract: The expanding application of smart sensing has created a growing demand for the accurate understanding of human action at the network edge. Traditiona

Task Relevance Is Not Local Replaceability: A Two-Axis View of Channel Information

TutorialsDGX agent

arXiv:2605.07086v1 Announce Type: new Abstract: Channel importance in vision networks is usually summarized by a single score. That summary hides two different questions: how much a channel is related

Teacher-Feature Drifting: One-Step Diffusion Distillation with Pretrained Diffusion Representations

ResearchDGX agent

arXiv:2605.07327v1 Announce Type: new Abstract: Sampling from pretrained diffusion and flow-matching models typically requires many forward passes to generate diverse and high-fidelity images. Existin

Teaching Prompts to Coordinate: Hierarchical Layer-Grouped Prompt Tuning for Continual Learning

Local AiDGX agent

arXiv:2511.12090v3 Announce Type: replace Abstract: Prompt-based continual learning methods fine-tune only a small set of additional learnable parameters while keeping the pre-trained model's paramete

Towards Explainable Industrial Anomaly Detection via Knowledge-Guided Latent Reasoning

ResearchDGX agent

arXiv:2602.09850v2 Announce Type: replace Abstract: Industrial anomaly detection demands precise reasoning over fine-grained defect patterns. However, existing multimodal large language models (MLLMs)

Towards Fairness under Label Bias in Image Segmentation: Impact, Measurement and Mitigation

SafetyDGX agent

arXiv:2605.06891v1 Announce Type: new Abstract: Labeled datasets reflect the biases of their annotation pipelines, which sometimes introduce label bias: group-conditional label errors that cause syste

Towards Highly-Constrained Human Motion Generation with Retrieval-Guided Diffusion Noise Optimization

ResearchDGX agent

arXiv:2605.08054v1 Announce Type: new Abstract: Generating human motion that satisfies customized zero-shot goal functions, enabling applications such as controllable character animation and behavior

Towards multi-modal forgery representation learning for AI-generated video detection and localization

Local AiDGX agent

arXiv:2605.07232v1 Announce Type: new Abstract: Recent advances in generative AI have democratized video creation at scale. AI-generated videos, including partially manipulated clips across visual and

Towards Photorealistic and Efficient Bokeh Rendering via Diffusion Framework

ApplicationsDGX agent

arXiv:2605.07429v1 Announce Type: new Abstract: Existing mobile devices are constrained by compact optical designs, such as small apertures, which make it difficult to produce natural, optically reali

TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos

Model ReleasesDGX agent

arXiv:2605.07593v1 Announce Type: new Abstract: Real-world audio-visual understanding requires chaining evidence that is sparse, temporally dispersed, and split across the visual and auditory streams,

TRAJGANR: Trajectory-Centric Urban Multimodal Learning via Geospatially Aligned Neural Representations

SafetyDGX agent

arXiv:2605.06990v1 Announce Type: new Abstract: Multimodal self-supervised learning (MSSL) has emerged as a key paradigm for pretraining geospatial foundation models. However, existing geospatial MSSL

TRAS: An Interactive Software for Tracing Tree Ring Cross Sections

ResearchDGX agent

arXiv:2605.08025v1 Announce Type: new Abstract: Tree ring marking remains a key step in dendrometry and dendrochronology, but it is often performed manually, making the process time-consuming, subject

TriDE: Triangle-Consistent Translation Directions for Global Camera Pose Estimation

Local AiDGX agent

arXiv:2605.06889v1 Announce Type: new Abstract: Pairwise translation directions are a key input to camera location estimation in global structure-from-motion. Existing estimators usually process each

TriP: A Triangle Puzzle Approach to Robust Translation Averaging

ResearchDGX agent

arXiv:2605.07143v1 Announce Type: new Abstract: Translation averaging aims to recover camera locations from pairwise relative translation directions and is a fundamental component of global Structure-

TRUST: Test-Time Refinement using Uncertainty-Guided SSM Traverses

TutorialsDGX agent

arXiv:2509.22813v2 Announce Type: replace Abstract: State Space Models (SSMs) have emerged as efficient alternatives to Vision Transformers (ViTs), with VMamba standing out as a pioneering architectur

Uncertainty Quantification for Cardiac Shape Reconstruction with Deep Signed Distance Functions via MCMC methods

ResearchDGX agent

arXiv:2605.07987v1 Announce Type: cross Abstract: Atlas-based approaches allow high-quality, patient-specific shape reconstructions of cardiac anatomy from sparse and/or noisy data such as point cloud

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models

ApplicationsDGX agent

arXiv:2605.07148v1 Announce Type: new Abstract: Decades of cognitive science establish that humans navigate environments by forming cognitive maps, defined as allocentric and topology-preserving repre

UniD-Shift: Towards Unified Semantic Segmentation via Interpretable Share-Private Multimodal Decomposition

SafetyDGX agent

arXiv:2605.07356v1 Announce Type: new Abstract: Semantic segmentation of large-scale 3D point clouds is crucial for applications such as autonomous driving and urban digital twins. However, the sparse

UniISP: A Unified ISP Framework for Both Human and Machine Vision

ResearchDGX agent

arXiv:2605.07359v1 Announce Type: new Abstract: Compared to RGB images, raw sensor data provides a richer representation of information, which is crucial for accurate recognition, particularly under c

UniV2D: Bridging Visual Restoration and Semantic Perception for Underwater Salient Object Detection

TutorialsDGX agent

arXiv:2605.07146v1 Announce Type: new Abstract: Underwater salient object detection (USOD) plays a vital role in marine vision tasks but remains fundamentally challenging due to severe visual degradat

VDEGaussian: Video Diffusion Enhanced 4D Gaussian Splatting for Dynamic Urban Scenes Modeling

SafetyDGX agent

arXiv:2508.02129v2 Announce Type: replace Abstract: Dynamic urban scene modeling is a rapidly evolving area with broad applications. While current approaches leveraging neural radiance fields or Gauss

Velocity-Space 3D Asset Editing

ResearchDGX agent

arXiv:2605.07385v1 Announce Type: cross Abstract: Editing a 3D asset locally, modifying a target region while preserving the rest, is a fundamental requirement of native 3D editing. Existing methods e

VesselRW: Weakly Supervised Subcutaneous Vessel Segmentation via Learned Random Walk Propagation

TutorialsDGX agent

arXiv:2508.06819v3 Announce Type: replace Abstract: The task of parsing subcutaneous vessels in clinical images is often hindered by the high cost and limited availability of ground truth data, as wel

VIMCAN: Visual-Inertial 3D Human Pose Estimation with Hybrid Mamba-Cross-Attention Network

ResearchDGX agent

arXiv:2605.07552v1 Announce Type: new Abstract: The rapid advances in deep learning have significantly enhanced the accuracy of multimodal 3D human pose estimation (HPE). However, the state-of-the-art

Weather-Robust Scene Semantics with Vision-Aligned 4D Radar

ResearchDGX agent

arXiv:2605.07367v1 Announce Type: cross Abstract: Cameras and LiDAR degrade in rain, fog, and snow, while millimeter-wave radar remains largely unaffected. We align a radar encoder to frozen SigLIP vi

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion

ResearchDGX agent

arXiv:2605.07915v1 Announce Type: new Abstract: Tokenizers are a crucial component of latent diffusion models, as they define the latent space in which diffusion models operate. However, existing toke

Zero-Shot Satellite Image Retrieval through Joint Embeddings: Application to Crisis Response

ResearchDGX agent

arXiv:2605.05405v2 Announce Type: replace Abstract: Semantic search of Earth observation archives remains challenging. Visual foundation models such as CLAY produce rich embeddings of satellite imager

7 May 2026

3D Ultrasound-Derived Pseudo-CT Synthesis Using a Transformer-Augmented Residual Network for Real-Time Operator Guidance

Local AiDGX agent

arXiv:2605.04856v1 Announce Type: new Abstract: Computed tomography (CT) is indispensable for clinical diagnosis and image-guided interventions but exposes patients to ionizing radiation, motivating t

A Bayesian Approach for Task-Specific Next-Best-View Selection with Uncertain Geometry

ResearchDGX agent

arXiv:2605.05095v1 Announce Type: cross Abstract: We develop a framework for task-specific active next-best-view selection in 3D reconstruction from point clouds, by casting the problem in the languag

A cross-modal network for facial expression recognition

SafetyDGX agent

arXiv:2605.04439v1 Announce Type: new Abstract: Deep neural networks enriched with structural information have been widely employed for facial expression recognition tasks. However, these methods ofte

A large-scale heterogeneous 3D magnetic resonance brain imaging dataset for self-supervised learning

ResearchDGX agent

arXiv:2506.14432v3 Announce Type: replace-cross Abstract: We present FOMO260K, a large-scale, heterogeneous dataset of 260,927 brain Magnetic Resonance Imaging (MRI) scans from 77,589 MRI sessions and

← Previous
1…149150151152153…209
Next →