AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
4 Aug 2026

SPARE: Structural Parameter-Free Affinity Regularization for Flow Matching

Model ReleasesDGX agent

arXiv:2608.01990v1 Announce Type: new Abstract: Denoising diffusion transformers achieve strong generation quality but converge slowly during training. Regularizing their internal representations has

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance

Local AiDGX agent

arXiv:2608.00502v1 Announce Type: new Abstract: Affordance grounding aims to localize the functional region for interaction, such as the handle to grasp or the button to press, rather than the whole o

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.01709v1 Announce Type: new Abstract: Vision-language models (VLMs) achieve strong semantic understanding but remain unreliable in metric spatial reasoning, particularly when queries require

SPECTRA: Band-Routed Embedding and Stage-Wise LoRA for Cross-Sensor Fine-Tuning of Geospatial Foundation Models

Model ReleasesDGX agent

arXiv:2608.01751v1 Announce Type: new Abstract: Geospatial foundation models (GeoFMs), pretrained on large-scale geospatial data such as Earth observation (EO), climate, and weather data, have shown p

SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection

Model ReleasesDGX agent

arXiv:2608.01334v1 Announce Type: new Abstract: AI-generated video (AIGV) detection aims to distinguish real videos from AI-generated ones. In practice, detectors trained on existing data often fail t

SpikeRestormer: Towards Energy-Efficient All-in-One Image Restoration via Unified Event Reasoning

ResearchDGX agent

arXiv:2608.02290v1 Announce Type: new Abstract: ANN-based All-in-One image restoration (AiOIR) unifies diverse degradation handling but incurs high computational costs, limiting its real-time deployme

SPIRIT: Spatio-temporal Pairwise Relational Modeling of Instrument-Tissue Interactions for Surgical Action Triplet Recognition

SafetyDGX agent

arXiv:2608.02188v1 Announce Type: new Abstract: Fine-grained understanding of surgical activity is essential for context-aware assistance in the operating room, including safety monitoring, adverse ev

SSR: Similarity-Shift Refinement for Training-Free Object-Centric Masks

ApplicationsDGX agent

arXiv:2608.01103v1 Announce Type: new Abstract: Object-centric models often produce fragmented masks, boundary leakage, and incorrect region merging. We introduce Similarity-Shift Refinement (SSR), a

ST-LoRA: Single Trajectory LoRA Ensemble for Uncertainty Aware Agricultural Segmentation

Model ReleasesDGX agent

arXiv:2608.01530v1 Announce Type: new Abstract: Reliable decision-support in digital agriculture requires accurate predictions and well-calibrated uncertainty estimates, particularly for dense predict

STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision

AgentsDGX agent

arXiv:2608.01535v1 Announce Type: new Abstract: Vision-language models (VLMs) are emerging as a key component of embodied intelligence, with growing applications in auto-labeling and end-to-end autono

STC-Net: Electroluminescence-Based Solar Cell Crack Segmentation for Power Loss Estimation

ResearchDGX agent

arXiv:2608.01714v1 Announce Type: new Abstract: Accurate crack assessment in electroluminescence (EL) images is important for photovoltaic (PV) reliability analysis, yet existing segmentation methods

STEAM:ASpatio-TEmporal Alignment Mixture-of-Experts Model with Hierarchical Pre-training for EEG Decoding

SafetyDGX agent

arXiv:2608.02070v1 Announce Type: new Abstract: Brain-computer interfaces (BCIs) have been widely used in motor rehabilitation, disease diagnosis, and other neural engineering scenarios. However, conv

Stipple: Real-Time Incremental Gaussian Splatting with Visual-Inertial Tracking

HardwareDGX agent

arXiv:2608.00931v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) provides efficient rendering of photo-realistic scenes, but its heavy preprocessing and training steps make it a poor fit

Stochastic Sequential Search in Very-High-Dimensional Feature Selection

ResearchDGX agent

arXiv:2608.01502v1 Announce Type: cross Abstract: Sequential subset search -- forward selection with floating backtracking and its descendants -- remains the quality reference in feature selection, bu

StreamSplat: Streaming Feed-Forward 3D Gaussian Splatting

ResearchDGX agent

arXiv:2608.01659v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting enables efficient novel-view synthesis without per-scene optimization, but most existing methods assume a fixed set o

StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring

ResearchDGX agent

arXiv:2608.01643v1 Announce Type: new Abstract: Real-time co-speech gesture generation must produce 3D motion clip by clip as speech arrives. Existing streaming methods are open-loop: each clip depend

Struct-GStream: Towards Efficient Free-Viewpoint Video Streaming at Low-Bitrates with Structured 3D Gaussians

Local AiDGX agent

arXiv:2608.01053v1 Announce Type: new Abstract: Constructing photorealistic Free-Viewpoint Videos (FVVs) of dynamic scenes from a set of posed 2D images has been an intriguing yet challenging task in

Structured Proxy Features for Multimodal NSCLC Survival Prediction from Pretreatment CT

Model ReleasesDGX agent

arXiv:2608.00446v1 Announce Type: new Abstract: Lung cancer results in roughly 1.8 million fatalities annually worldwide, with non-small cell lung cancer (NSCLC) comprising the majority of cases. Desp

StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field

Local AiDGX agent

arXiv:2608.01954v1 Announce Type: new Abstract: Fixed-layout indoor furniture styling requires selecting assets that form a coherent room without changing the prescribed furniture categories, position

SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation

Model ReleasesDGX agent

arXiv:2608.01977v1 Announce Type: new Abstract: Multimodal large models are increasingly used to generate scalable vector graphics (SVG), but reliable evaluation remains underexplored. Existing protoc

SVRepair: Structured Visual Reasoning for Automated Program Repair

Local AiDGX agent

arXiv:2602.06090v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently been applied to Automated Program Repair (APR), yet most existing approaches remain unimodal and fa

Swimm3R: Splatting with Medium-aware SfM for Underwater 3D Reconstruction

ResearchDGX agent

arXiv:2608.00950v1 Announce Type: new Abstract: We propose Swimm3R, a unified framework that combines medium-aware structure-from-motion (SfM) with Underwater Beta Splatting to address scattering- and

SWINSleepNet: A Hierarchical Context-Aware Framework for Sleep Staging (v2)

ResearchDGX agent

arXiv:2608.02183v1 Announce Type: new Abstract: Automatic sleep staging is a critical role in sleep disorder diagnosis, sleep quality assessment, and long-term health monitoring; however, existing app

T^2exture: Sparsely Perturbed Thermal-to-Texture Imaging

Model ReleasesDGX agent

arXiv:2608.02192v1 Announce Type: new Abstract: Thermal imaging remains effective under adverse illumination, yet passive long-wave infrared (LWIR) measurements often lack fine texture. Existing therm

TBSG-Net: Temporal Bipartite Scene Graph Network for Fine-Grained Video Moment Retrieval

Local AiDGX agent

arXiv:2608.02056v1 Announce Type: new Abstract: Recent advances in proposal-free Video Moment Retrieval (VMR) have highlighted the effectiveness of Static Scene Graphs (SSGs). By modeling objects and

Test-time Adaptation of Pelvic Bone Segmentation Models via Dynamic Reliability-Guided

ResearchDGX agent

arXiv:2608.00510v1 Announce Type: new Abstract: Reliable pelvic bone segmentation (PBS) from CT is essential for robot-assisted pelvic trauma surgery, yet deploying a source-trained model to a new hos

Test-Time Curriculum for Open-Set AIGC Detection

Model ReleasesDGX agent

arXiv:2608.00559v1 Announce Type: new Abstract: AI-generated image detectors deployed in open-world environments inevitably face distribution shifts as new and stronger generative models continue to e

The 1st AI Children Challenge

ApplicationsDGX agent

arXiv:2608.00356v1 Announce Type: new Abstract: The First AI Children Challenge aims to advance real-world applications of computer vision and AI in child healthcare, child education, and pediatrics.

The Push-Forward Transform for Continuous and Robust Comparison of Dynamic Shapes

Model ReleasesDGX agent

arXiv:2608.02306v1 Announce Type: new Abstract: We introduce a mathematical framework for shape comparison based on mapping functions from the shape domain to a common reference domain. This Push-Forw

Think in Sets for Streaming Video Token Compression

ResearchDGX agent

arXiv:2608.01169v1 Announce Type: new Abstract: Streaming VideoLLMs process frames causally while visual tokens grow continuously, making compression essential for controlling prefilling latency and m

Token Radius Attention for Efficient Video Generation

ResearchDGX agent

arXiv:2608.02504v1 Announce Type: new Abstract: Video Diffusion Transformers (VDiTs) enable high-fidelity generation but incur quadratic cost from dense 3D self-attention. Existing head- and block-lev

Towards Compact Unified Multimodal Tracking: Synergizing Knowledge Distillation with Structural Pruning

SafetyDGX agent

arXiv:2608.01488v1 Announce Type: new Abstract: Unified multimodal object tracking has achieved remarkable robustness by leveraging complementary sensor data (e.g., RGB, Thermal, Depth), yet the heavy

Training-Free Out-of-Distribution Detection for Pathology Whole-Slide Images

ResearchDGX agent

arXiv:2608.01407v1 Announce Type: new Abstract: Safe deployment of AI methods in medicine requires robust guardrails that detect when input data deviate from the training distribution to ensure that m

Transformer Geometry Observatory TGO-III: Semantic Geometry Observatory

ResearchDGX agent

arXiv:2608.01876v1 Announce Type: new Abstract: With the widespread adoption of Vision Transformers in modern AI, the need to analyze their inherent representational behavior has become increasingly i

TravKAN: Fast and Interpretable Nonlinear Traversability Analysis with Kolmogorov-Arnold Networks

SafetyDGX agent

arXiv:2608.02320v1 Announce Type: cross Abstract: Traversability analysis is a fundamental capability for autonomous mobile robots operating in unstructured environments. While modern machine learning

TurboClear: One-Step Object-Effect Removal via Region-Calibrated Distribution Matching and Fusion

ResearchDGX agent

arXiv:2608.01288v1 Announce Type: new Abstract: Recently, diffusion-based removal methods have achieved promising visual quality in removing both target objects and their associated effects. However,

Two Sides of the Same Coin: Co-Evolving Search for Cross-Task Attacks on Vision-Language Models

ResearchDGX agent

arXiv:2608.02137v1 Announce Type: new Abstract: Vision-language models (VLMs) exhibit strong generalization across multimodal tasks but remain vulnerable to adversarial perturbations. Existing attacks

UAV-Based Environmental Monitoring of Rip-Current Indicators Using Wavelet-Derived Texture Features

SafetyDGX agent

arXiv:2608.02448v1 Announce Type: new Abstract: Rip currents are recurrent coastal natural hazards that threaten beachgoers and create operational challenges for lifeguards and coastal managers. Relia

UCBound-Net: Uncertainty-Guided Boundary-Aware Continual Learning for Domain-Incremental Ultrasound Segmentation

Model ReleasesDGX agent

arXiv:2608.01518v1 Announce Type: new Abstract: Continual learning in clinical imaging faces a dual challenge: a model must assimilate knowledge from new anatomical domains while retaining representat

UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction

SafetyDGX agent

arXiv:2608.01298v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have emerged as a core architecture in generative modeling due to their scalability and adaptability to multimodal tasks.

Uncertainty Quantification for Visual Object Pose Estimation: S-Lemma Ellipsoidal Bounds

ApplicationsDGX agent

arXiv:2511.21666v2 Announce Type: replace-cross Abstract: Quantifying the uncertainty of an object's pose estimate is essential for robust control and planning. Although pose estimation is a well-stud

Understanding and Overcoming Cross-modal Fusion Bias in Multimodal Anomaly Detection From A Fisher Information Perspective

SafetyDGX agent

arXiv:2608.00986v1 Announce Type: new Abstract: Current advancements in Multimodal Anomaly Detection (MAD) are largely driven by enhancing multimodal fusion, particularly through the integration of RG

Understanding Synergistic Interactions among Pathology Foundation Models via Adaptive Fusion

ResearchDGX agent

arXiv:2608.01370v1 Announce Type: new Abstract: Pathology foundation models (PFMs) provide strong tile-level representations via self-supervised pre-training on large-scale pathology images. Yet, PFMs

UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation

ResearchDGX agent

arXiv:2608.01944v1 Announce Type: new Abstract: Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes w

UniqueSplat: View-conditioned 3D Gaussian Splatting for Generalizable 3D Reconstruction

TutorialsDGX agent

arXiv:2608.02145v1 Announce Type: new Abstract: In this paper, we propose UniqueSplat, a view-conditioned feed-forward 3D Gaussian Splatting model to reconstruct customized 3D radiance fields for each

UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization

ResearchDGX agent

arXiv:2608.01706v1 Announce Type: new Abstract: Recent geometric foundation models enable feed-forward inference for SLAM, but their predictions are strongly dependent on the input view set, which lea

Unleashing the Power of Text: Text-Guided Flow Matching for Image Fusion under Complex Degradations

TutorialsDGX agent

arXiv:2608.00530v1 Announce Type: new Abstract: Infrared-visible image fusion under realistic degradation scenarios is a challenging task, as degradations not only cause a loss of reliable modality-sp

USP-Mamba: Unmixing-Derived Spectral and Structural Prompting for Hyperspectral Image Super-Resolution

SafetyDGX agent

arXiv:2608.02401v1 Announce Type: new Abstract: Hyperspectral image super-resolution aims to reconstruct high-resolution imagery while preserving dense spectral information. Recently, Mamba-based mode

VARPose: Flexible 2D Pose Densification via Visual Autoregressive Modeling for Enhanced 3D Lifting

ResearchDGX agent

arXiv:2608.02214v1 Announce Type: new Abstract: Visual AutoRegressive Modeling (VAR) has excelled in natural image generation via next-scale prediction, but its use on topology-structured data like hu

VaRS-Doc: Interpretation-Aware Variant Representations via Latent Self-Probing for Visual Document Retrieval

ApplicationsDGX agent

arXiv:2608.01211v1 Announce Type: new Abstract: Visual document retrieval has recently become increasingly important in applications such as enterprise search, scientific literature discovery, and ret

VC-Tooler: Learning Compositional and Adaptive Visual Tool Use

AgentsDGX agent

arXiv:2608.02217v1 Announce Type: new Abstract: Agentic multimodal reasoning extends passive image understanding by allowing VLMs to actively acquire and refine visual evidence through visual tool int

VGER: Voxel-Guided Global Event Ranking for Event Cloud Attribution

ResearchDGX agent

arXiv:2608.01470v1 Announce Type: new Abstract: Event cameras produce sparse and asynchronous event streams that provide rich spatio-temporal information for efficient perception. Recent advances in e

Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh

Model ReleasesDGX agent

arXiv:2608.00094v1 Announce Type: new Abstract: Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a

Volcanic Clouds Detection through QCNN and Geostationary Satellite Multispectral Imagery

SafetyDGX agent

arXiv:2608.00072v1 Announce Type: new Abstract: Recent advances in quantum computing are opening new possibilities for Earth Observation (EO) data analysis. Quantum machine learning (QML) approaches o

VR3D: View-Robust 3D Representation Learning for Aerial-Ground Person Re-Identification

Model ReleasesDGX agent

arXiv:2608.02598v1 Announce Type: new Abstract: Aerial-ground person re-identification is a challenging task due to cross-platform viewpoint variations, which cause severe occlusion and geometric defo

WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA

SafetyDGX agent

arXiv:2608.01035v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deployment is sev

What Carries the Signal in Pathology Foundation-Model Atlases? A Patient-Level Controlled Benchmark in Breast Cancer

Model ReleasesDGX agent

arXiv:2608.00105v1 Announce Type: new Abstract: Pathology foundation models are reported to encode molecular programmes in tissue morphology, but the evidence is usually a cohort-wide ranked gene list

When Extreme Darkness Meets Motion Blur: MeanFlow for Unified RAW Restoration

ResearchDGX agent

arXiv:2608.01720v1 Announce Type: new Abstract: Extremely low-light RAW enhancement aims to recover severely attenuated sensor signals, yet existing methods often focus on illumination and noise while

When Measurement Conventions Masquerade as Calibration Gains in Cardiac Digital Twins

Model ReleasesDGX agent

arXiv:2608.01602v1 Announce Type: new Abstract: Cardiac digital twins convert clinical images into physiological measurements through observation operators, yet calibration studies often assume a fixe

Where Does Generative Difficulty Reside? An Empirical Study of Target Representations

Local AiDGX agent

arXiv:2608.00626v1 Announce Type: new Abstract: The target representation defines the distribution an image generator must learn, yet it is often treated as an interchangeable interface. This assumpti

← Previous
1…1819202122…207
Next →