AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Applications

SSR: Similarity-Shift Refinement for Training-Free Object-Centric Masks

DGX agent

arXiv:2608.01103v1 Announce Type: new Abstract: Object-centric models often produce fragmented masks, boundary leakage, and incorrect region merging. We introduce Similarity-Shift Refinement (SSR), a

applicationsarxiv-cs-cv
4 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

ST-LoRA: Single Trajectory LoRA Ensemble for Uncertainty Aware Agricultural Segmentation

DGX agent

arXiv:2608.01530v1 Announce Type: new Abstract: Reliable decision-support in digital agriculture requires accurate predictions and well-calibrated uncertainty estimates, particularly for dense predict

model-releasesarxiv-cs-cv
4 Aug 2026
Agents

STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision

DGX agent

arXiv:2608.01535v1 Announce Type: new Abstract: Vision-language models (VLMs) are emerging as a key component of embodied intelligence, with growing applications in auto-labeling and end-to-end autono

agentsarxiv-cs-cv
4 Aug 2026
Research

STC-Net: Electroluminescence-Based Solar Cell Crack Segmentation for Power Loss Estimation

DGX agent

arXiv:2608.01714v1 Announce Type: new Abstract: Accurate crack assessment in electroluminescence (EL) images is important for photovoltaic (PV) reliability analysis, yet existing segmentation methods

researcharxiv-cs-cv
4 Aug 2026
Safety

STEAM:ASpatio-TEmporal Alignment Mixture-of-Experts Model with Hierarchical Pre-training for EEG Decoding

DGX agent

arXiv:2608.02070v1 Announce Type: new Abstract: Brain-computer interfaces (BCIs) have been widely used in motor rehabilitation, disease diagnosis, and other neural engineering scenarios. However, conv

safetyarxiv-cs-cv
4 Aug 2026
Hardware

Stipple: Real-Time Incremental Gaussian Splatting with Visual-Inertial Tracking

DGX agent

arXiv:2608.00931v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) provides efficient rendering of photo-realistic scenes, but its heavy preprocessing and training steps make it a poor fit

hardwarearxiv-cs-cv
4 Aug 2026
Research

Stochastic Sequential Search in Very-High-Dimensional Feature Selection

DGX agent

arXiv:2608.01502v1 Announce Type: cross Abstract: Sequential subset search -- forward selection with floating backtracking and its descendants -- remains the quality reference in feature selection, bu

researcharxiv-cs-cv
4 Aug 2026
Research

StreamSplat: Streaming Feed-Forward 3D Gaussian Splatting

DGX agent

arXiv:2608.01659v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting enables efficient novel-view synthesis without per-scene optimization, but most existing methods assume a fixed set o

researcharxiv-cs-cv
4 Aug 2026
Research

StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring

DGX agent

arXiv:2608.01643v1 Announce Type: new Abstract: Real-time co-speech gesture generation must produce 3D motion clip by clip as speech arrives. Existing streaming methods are open-loop: each clip depend

researcharxiv-cs-cv
4 Aug 2026
Local Ai

Struct-GStream: Towards Efficient Free-Viewpoint Video Streaming at Low-Bitrates with Structured 3D Gaussians

DGX agent

arXiv:2608.01053v1 Announce Type: new Abstract: Constructing photorealistic Free-Viewpoint Videos (FVVs) of dynamic scenes from a set of posed 2D images has been an intriguing yet challenging task in

local-aiarxiv-cs-cv
4 Aug 2026
Model Releases

Structured Proxy Features for Multimodal NSCLC Survival Prediction from Pretreatment CT

DGX agent

arXiv:2608.00446v1 Announce Type: new Abstract: Lung cancer results in roughly 1.8 million fatalities annually worldwide, with non-small cell lung cancer (NSCLC) comprising the majority of cases. Desp

model-releasesarxiv-cs-cv
4 Aug 2026
Local Ai

StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field

DGX agent

arXiv:2608.01954v1 Announce Type: new Abstract: Fixed-layout indoor furniture styling requires selecting assets that form a coherent room without changing the prescribed furniture categories, position

local-aiarxiv-cs-cv
4 Aug 2026
Model Releases

SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation

DGX agent

arXiv:2608.01977v1 Announce Type: new Abstract: Multimodal large models are increasingly used to generate scalable vector graphics (SVG), but reliable evaluation remains underexplored. Existing protoc

model-releasesarxiv-cs-cv
4 Aug 2026
Local Ai

SVRepair: Structured Visual Reasoning for Automated Program Repair

DGX agent

arXiv:2602.06090v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently been applied to Automated Program Repair (APR), yet most existing approaches remain unimodal and fa

local-aiarxiv-cs-cv
4 Aug 2026
Research

Swimm3R: Splatting with Medium-aware SfM for Underwater 3D Reconstruction

DGX agent

arXiv:2608.00950v1 Announce Type: new Abstract: We propose Swimm3R, a unified framework that combines medium-aware structure-from-motion (SfM) with Underwater Beta Splatting to address scattering- and

researcharxiv-cs-cv
4 Aug 2026
Research

SWINSleepNet: A Hierarchical Context-Aware Framework for Sleep Staging (v2)

DGX agent

arXiv:2608.02183v1 Announce Type: new Abstract: Automatic sleep staging is a critical role in sleep disorder diagnosis, sleep quality assessment, and long-term health monitoring; however, existing app

researcharxiv-cs-cv
4 Aug 2026
Model Releases

T^2exture: Sparsely Perturbed Thermal-to-Texture Imaging

DGX agent

arXiv:2608.02192v1 Announce Type: new Abstract: Thermal imaging remains effective under adverse illumination, yet passive long-wave infrared (LWIR) measurements often lack fine texture. Existing therm

model-releasesarxiv-cs-cv
4 Aug 2026
Local Ai

TBSG-Net: Temporal Bipartite Scene Graph Network for Fine-Grained Video Moment Retrieval

DGX agent

arXiv:2608.02056v1 Announce Type: new Abstract: Recent advances in proposal-free Video Moment Retrieval (VMR) have highlighted the effectiveness of Static Scene Graphs (SSGs). By modeling objects and

local-aiarxiv-cs-cv
4 Aug 2026
Research

Test-time Adaptation of Pelvic Bone Segmentation Models via Dynamic Reliability-Guided

DGX agent

arXiv:2608.00510v1 Announce Type: new Abstract: Reliable pelvic bone segmentation (PBS) from CT is essential for robot-assisted pelvic trauma surgery, yet deploying a source-trained model to a new hos

researcharxiv-cs-cv
4 Aug 2026
Model Releases

Test-Time Curriculum for Open-Set AIGC Detection

DGX agent

arXiv:2608.00559v1 Announce Type: new Abstract: AI-generated image detectors deployed in open-world environments inevitably face distribution shifts as new and stronger generative models continue to e

model-releasesarxiv-cs-cv
4 Aug 2026
Applications

The 1st AI Children Challenge

DGX agent

arXiv:2608.00356v1 Announce Type: new Abstract: The First AI Children Challenge aims to advance real-world applications of computer vision and AI in child healthcare, child education, and pediatrics.

applicationsarxiv-cs-cv
4 Aug 2026
Model Releases

The Push-Forward Transform for Continuous and Robust Comparison of Dynamic Shapes

DGX agent

arXiv:2608.02306v1 Announce Type: new Abstract: We introduce a mathematical framework for shape comparison based on mapping functions from the shape domain to a common reference domain. This Push-Forw

model-releasesarxiv-cs-cv
4 Aug 2026
Research

Think in Sets for Streaming Video Token Compression

DGX agent

arXiv:2608.01169v1 Announce Type: new Abstract: Streaming VideoLLMs process frames causally while visual tokens grow continuously, making compression essential for controlling prefilling latency and m

researcharxiv-cs-cv
4 Aug 2026
Research

Token Radius Attention for Efficient Video Generation

DGX agent

arXiv:2608.02504v1 Announce Type: new Abstract: Video Diffusion Transformers (VDiTs) enable high-fidelity generation but incur quadratic cost from dense 3D self-attention. Existing head- and block-lev

researcharxiv-cs-cv
4 Aug 2026
Safety

Towards Compact Unified Multimodal Tracking: Synergizing Knowledge Distillation with Structural Pruning

DGX agent

arXiv:2608.01488v1 Announce Type: new Abstract: Unified multimodal object tracking has achieved remarkable robustness by leveraging complementary sensor data (e.g., RGB, Thermal, Depth), yet the heavy

safetyarxiv-cs-cv
4 Aug 2026
Research

Training-Free Out-of-Distribution Detection for Pathology Whole-Slide Images

DGX agent

arXiv:2608.01407v1 Announce Type: new Abstract: Safe deployment of AI methods in medicine requires robust guardrails that detect when input data deviate from the training distribution to ensure that m

researcharxiv-cs-cv
4 Aug 2026
Research

Transformer Geometry Observatory TGO-III: Semantic Geometry Observatory

DGX agent

arXiv:2608.01876v1 Announce Type: new Abstract: With the widespread adoption of Vision Transformers in modern AI, the need to analyze their inherent representational behavior has become increasingly i

researcharxiv-cs-cv
4 Aug 2026
Safety

TravKAN: Fast and Interpretable Nonlinear Traversability Analysis with Kolmogorov-Arnold Networks

DGX agent

arXiv:2608.02320v1 Announce Type: cross Abstract: Traversability analysis is a fundamental capability for autonomous mobile robots operating in unstructured environments. While modern machine learning

safetyarxiv-cs-cv
4 Aug 2026
Research

TurboClear: One-Step Object-Effect Removal via Region-Calibrated Distribution Matching and Fusion

DGX agent

arXiv:2608.01288v1 Announce Type: new Abstract: Recently, diffusion-based removal methods have achieved promising visual quality in removing both target objects and their associated effects. However,

researcharxiv-cs-cv
4 Aug 2026
Research

Two Sides of the Same Coin: Co-Evolving Search for Cross-Task Attacks on Vision-Language Models

DGX agent

arXiv:2608.02137v1 Announce Type: new Abstract: Vision-language models (VLMs) exhibit strong generalization across multimodal tasks but remain vulnerable to adversarial perturbations. Existing attacks

researcharxiv-cs-cv
4 Aug 2026
Safety

UAV-Based Environmental Monitoring of Rip-Current Indicators Using Wavelet-Derived Texture Features

DGX agent

arXiv:2608.02448v1 Announce Type: new Abstract: Rip currents are recurrent coastal natural hazards that threaten beachgoers and create operational challenges for lifeguards and coastal managers. Relia

safetyarxiv-cs-cv
4 Aug 2026
Model Releases

UCBound-Net: Uncertainty-Guided Boundary-Aware Continual Learning for Domain-Incremental Ultrasound Segmentation

DGX agent

arXiv:2608.01518v1 Announce Type: new Abstract: Continual learning in clinical imaging faces a dual challenge: a model must assimilate knowledge from new anatomical domains while retaining representat

model-releasesarxiv-cs-cv
4 Aug 2026
Safety

UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction

DGX agent

arXiv:2608.01298v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have emerged as a core architecture in generative modeling due to their scalability and adaptability to multimodal tasks.

safetyarxiv-cs-cv
4 Aug 2026
Applications

Uncertainty Quantification for Visual Object Pose Estimation: S-Lemma Ellipsoidal Bounds

DGX agent

arXiv:2511.21666v2 Announce Type: replace-cross Abstract: Quantifying the uncertainty of an object's pose estimate is essential for robust control and planning. Although pose estimation is a well-stud

applicationsarxiv-cs-cv
4 Aug 2026
Safety

Understanding and Overcoming Cross-modal Fusion Bias in Multimodal Anomaly Detection From A Fisher Information Perspective

DGX agent

arXiv:2608.00986v1 Announce Type: new Abstract: Current advancements in Multimodal Anomaly Detection (MAD) are largely driven by enhancing multimodal fusion, particularly through the integration of RG

safetyarxiv-cs-cv
4 Aug 2026
Research

Understanding Synergistic Interactions among Pathology Foundation Models via Adaptive Fusion

DGX agent

arXiv:2608.01370v1 Announce Type: new Abstract: Pathology foundation models (PFMs) provide strong tile-level representations via self-supervised pre-training on large-scale pathology images. Yet, PFMs

researcharxiv-cs-cv
4 Aug 2026
Research

UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation

DGX agent

arXiv:2608.01944v1 Announce Type: new Abstract: Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes w

researcharxiv-cs-cv
4 Aug 2026
Tutorials

UniqueSplat: View-conditioned 3D Gaussian Splatting for Generalizable 3D Reconstruction

DGX agent

arXiv:2608.02145v1 Announce Type: new Abstract: In this paper, we propose UniqueSplat, a view-conditioned feed-forward 3D Gaussian Splatting model to reconstruct customized 3D radiance fields for each

tutorialsarxiv-cs-cv
4 Aug 2026
Research

UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization

DGX agent

arXiv:2608.01706v1 Announce Type: new Abstract: Recent geometric foundation models enable feed-forward inference for SLAM, but their predictions are strongly dependent on the input view set, which lea

researcharxiv-cs-cv
4 Aug 2026
Tutorials

Unleashing the Power of Text: Text-Guided Flow Matching for Image Fusion under Complex Degradations

DGX agent

arXiv:2608.00530v1 Announce Type: new Abstract: Infrared-visible image fusion under realistic degradation scenarios is a challenging task, as degradations not only cause a loss of reliable modality-sp

tutorialsarxiv-cs-cv
4 Aug 2026
Safety

USP-Mamba: Unmixing-Derived Spectral and Structural Prompting for Hyperspectral Image Super-Resolution

DGX agent

arXiv:2608.02401v1 Announce Type: new Abstract: Hyperspectral image super-resolution aims to reconstruct high-resolution imagery while preserving dense spectral information. Recently, Mamba-based mode

safetyarxiv-cs-cv
4 Aug 2026
Research

VARPose: Flexible 2D Pose Densification via Visual Autoregressive Modeling for Enhanced 3D Lifting

DGX agent

arXiv:2608.02214v1 Announce Type: new Abstract: Visual AutoRegressive Modeling (VAR) has excelled in natural image generation via next-scale prediction, but its use on topology-structured data like hu

researcharxiv-cs-cv
4 Aug 2026
Applications

VaRS-Doc: Interpretation-Aware Variant Representations via Latent Self-Probing for Visual Document Retrieval

DGX agent

arXiv:2608.01211v1 Announce Type: new Abstract: Visual document retrieval has recently become increasingly important in applications such as enterprise search, scientific literature discovery, and ret

applicationsarxiv-cs-cv
4 Aug 2026
Agents

VC-Tooler: Learning Compositional and Adaptive Visual Tool Use

DGX agent

arXiv:2608.02217v1 Announce Type: new Abstract: Agentic multimodal reasoning extends passive image understanding by allowing VLMs to actively acquire and refine visual evidence through visual tool int

agentsarxiv-cs-cv
4 Aug 2026
Research

VGER: Voxel-Guided Global Event Ranking for Event Cloud Attribution

DGX agent

arXiv:2608.01470v1 Announce Type: new Abstract: Event cameras produce sparse and asynchronous event streams that provide rich spatio-temporal information for efficient perception. Recent advances in e

researcharxiv-cs-cv
4 Aug 2026
Model Releases

Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh

DGX agent

arXiv:2608.00094v1 Announce Type: new Abstract: Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a

model-releasesarxiv-cs-cv
4 Aug 2026
Safety

Volcanic Clouds Detection through QCNN and Geostationary Satellite Multispectral Imagery

DGX agent

arXiv:2608.00072v1 Announce Type: new Abstract: Recent advances in quantum computing are opening new possibilities for Earth Observation (EO) data analysis. Quantum machine learning (QML) approaches o

safetyarxiv-cs-cv
4 Aug 2026
Model Releases

VR3D: View-Robust 3D Representation Learning for Aerial-Ground Person Re-Identification

DGX agent

arXiv:2608.02598v1 Announce Type: new Abstract: Aerial-ground person re-identification is a challenging task due to cross-platform viewpoint variations, which cause severe occlusion and geometric defo

model-releasesarxiv-cs-cv
4 Aug 2026
← Previous
1…2526272829…261
Next →