AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Agents

Semantic Browsing: Controllable Diversity for Image Generation

DGX agent

arXiv:2606.23679v1 Announce Type: new Abstract: Modern text-to-image models excel in visual fidelity and prompt adherence. However, this strict adherence comes at the cost of diversity: generated samp

agentsarxiv-cs-cv
23 Jun 2026
Safety

Semi-Supervised Vision-Language-Action Model

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2606.21493v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models enable robots to predict actions directly from visual observations and language instructions, but adapting them to n

safetyarxiv-cs-cv
23 Jun 2026
Research

SenseExpo: Spatial Exploration and Navigation via Scene Estimation from Expeditious Predictive Operators

DGX agent

arXiv:2503.16000v2 Announce Type: replace Abstract: We present extbf{SenseExpo}, a lightweight single-robot exploration framework that integrates a compact map prediction network into a frontier-based

researcharxiv-cs-cv
23 Jun 2026
Research

Shear-Free Viewport Magnification for 360-Degree via Spherical Mobius Boosts

DGX agent

arXiv:2606.20684v1 Announce Type: new Abstract: Viewport-adaptive 360-degree imaging seeks to allocate a fixed sampling budget to the region a viewer is likely to observe. Existing view-biased project

researcharxiv-cs-cv
23 Jun 2026
Research

ShuffleFlow: Scalable Posterior Inference for Bayesian Inverse Imaging

DGX agent

arXiv:2606.21099v1 Announce Type: new Abstract: Variational inference (VI) is a powerful method for principled posterior inference for scientific inverse imaging. VI learns the posterior distribution,

researcharxiv-cs-cv
23 Jun 2026
Research

SimAC: A Simple Anti-Customization Method for Protecting Face Privacy against Text-to-Image Synthesis of Diffusion Models

DGX agent

arXiv:2312.07865v4 Announce Type: replace Abstract: Despite the success of diffusion-based customization methods on visual content creation, increasing concerns have been raised about such techniques

researcharxiv-cs-cv
23 Jun 2026
Agents

SIMSplat: Language-Aligned 4D Gaussian Splatting for Driving Scenario Generation

DGX agent

arXiv:2510.02469v2 Announce Type: replace-cross Abstract: Driving scene manipulation using real-world sensor data has emerged as a promising alternative to traditional driving simulators. Despite adva

agentsarxiv-cs-cv
23 Jun 2026
Model Releases

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning

DGX agent

arXiv:2606.22873v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications. This broad deployment expands the

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Skeleton-to-Image Encoding: Enabling Skeleton Representation Learning via Vision-Pretrained Models

DGX agent

arXiv:2603.05963v2 Announce Type: replace Abstract: Recent advances in large-scale pretrained vision models have demonstrated impressive capabilities across a wide range of downstream tasks, including

researcharxiv-cs-cv
23 Jun 2026
Safety

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models

DGX agent

arXiv:2606.23041v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in visual understanding but remain constrained in visual generation due to the

safetyarxiv-cs-cv
23 Jun 2026
Agents

SPARC: A Multi-Agent System for Electrical Circuit Question Answering

DGX agent

arXiv:2606.20643v1 Announce Type: cross Abstract: Electrical circuit diagram QA tasks require complex mathematical reasoning, which remains challenging for multimodal LLMs. We present SPARC, a multi-a

agentsarxiv-cs-cv
23 Jun 2026
Tutorials

Sparse Point-Guided Fusion of Supervised and Self-Supervised Learning Model for Seaweed Segmentation

DGX agent

arXiv:2606.21026v1 Announce Type: new Abstract: The ocean plays a critical role in sustainable development, particularly in climate change mitigation. Among marine ecosystems, blue carbon ecosystems a

tutorialsarxiv-cs-cv
23 Jun 2026
Model Releases

SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation

DGX agent

arXiv:2605.24354v2 Announce Type: replace Abstract: Recently, world models have made significant progress in enhancing end-to-end driving systems through both future situation forecasting and improved

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Spatially Grounded Concept-Based Image Classification

DGX agent

arXiv:2510.04180v2 Announce Type: replace Abstract: Deep neural networks can achieve high accuracy while relying on evidence that is hard to inspect or misaligned with the intended task. Concept Bottl

researcharxiv-cs-cv
23 Jun 2026
Research

Spatio-Temporal Wildfire Spread Prediction in Canada using a Video Swin-Hybrid-U-Net and Satellite Imagery

DGX agent

arXiv:2606.20693v1 Announce Type: new Abstract: Background: Wildfires in Canada present increasing threats to ecosystems, communities, and infrastructure, demanding accurate forecasting tools to aid m

researcharxiv-cs-cv
23 Jun 2026
Research

Specificity- and Calibration-Aware Breast Ultrasound Segmentation via Entropy-Guided Boundary Supervision

DGX agent

arXiv:2606.22308v1 Announce Type: cross Abstract: Lesion segmentation in breast ultrasound involves two related challenges. In images with lesions, speckle noise, low tissue contrast, and posterior ac

researcharxiv-cs-cv
23 Jun 2026
Safety

Spectral Gating via Damped Oscillations for Adaptive Implicit Neural Representations

DGX agent

arXiv:2606.23129v1 Announce Type: new Abstract: Implicit Neural Representations (INRs) have been proven successful in encoding continuous signals through coordinate-based networks, yet facing a spectr

safetyarxiv-cs-cv
23 Jun 2026
Tutorials

Spectral GS-SLAM: Observability-Aware, Degeneracy-Robust Tracking for Real-Time 3D Gaussian Splatting SLAM

DGX agent

arXiv:2606.21258v1 Announce Type: cross Abstract: Recent 3DGS-SLAM systems enable real-time operation by leveraging conventional feature matching or ICP-based tracking, thereby avoiding the heavy dens

tutorialsarxiv-cs-cv
23 Jun 2026
Local Ai

SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs

DGX agent

arXiv:2606.20244v2 Announce Type: replace Abstract: Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and easy to over

local-aiarxiv-cs-cv
23 Jun 2026
Safety

Stabilizing Consistency Training: A Flow Map Analysis and Self-Distillation

DGX agent

arXiv:2601.22679v2 Announce Type: replace-cross Abstract: Consistency models have been proposed for fast generative modeling, achieving results competitive with diffusion and flow models. However, the

safetyarxiv-cs-cv
23 Jun 2026
Local Ai

SteerVTE: Seamless Video Text Editing with Style and Glyph Control

DGX agent

arXiv:2606.23254v1 Announce Type: new Abstract: Visual text editing aims to precisely modify text in images and videos while preserving stylistic consistency and visual realism. Despite significant ad

local-aiarxiv-cs-cv
23 Jun 2026
Research

Stochastic Signed Distance Processes

DGX agent

arXiv:2606.20856v1 Announce Type: new Abstract: Multi-view surface reconstruction is a core problem in computer vision. One prominent line of work represents the surface implicitly as a signed distanc

researcharxiv-cs-cv
23 Jun 2026
Model Releases

Streaming Dense Voxel Representations for 3D Occupancy Prediction

DGX agent

arXiv:2503.22087v3 Announce Type: replace Abstract: In this paper, we explore dense voxel streaming for accurate and efficient 3D occupancy prediction. While dense voxel representations offer fine-gra

model-releasesarxiv-cs-cv
23 Jun 2026
Research

StreamPPG: Low-Latency rPPG Estimation via Consistent Privileged Learning

DGX agent

arXiv:2606.23186v1 Announce Type: new Abstract: Remote photoplethysmography (rPPG) estimates the blood volume pulse (BVP) signal from facial videos, enabling contact-free health monitoring. Convention

researcharxiv-cs-cv
23 Jun 2026
Research

StructSAM: Structure- and Spectrum-Preserving Token Merging for Segment Anything Models

DGX agent

arXiv:2603.07307v2 Announce Type: replace Abstract: Recent token merging techniques for Vision Transformers (ViTs) provide substantial speedups by reducing the number of tokens processed by self-atten

researcharxiv-cs-cv
23 Jun 2026
Tutorials

Structural Assessment for Understanding and Guiding Dataset Distillation in Discrete Token Space

DGX agent

arXiv:2606.21705v1 Announce Type: new Abstract: Dataset distillation (DD) has proven to reduce training cost while preserving accuracy. While promising, the factors that make one distilled dataset mor

tutorialsarxiv-cs-cv
23 Jun 2026
Model Releases

Structured Hyperedge Adaptation for Parameter-Efficient Fine-Tuning of Vision Transformers

DGX agent

arXiv:2606.22383v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) has become a practical solution for adapting large pretrained vision transformers (ViTs) to downstream tasks whil

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Structured Spectral Graph Representation Learning for Multi-label Abnormality Analysis from 3D CT Scans

DGX agent

arXiv:2510.10779v5 Announce Type: replace Abstract: With the growing volume of CT examinations, there is an increasing demand for automated tools such as organ segmentation, abnormality detection, and

researcharxiv-cs-cv
23 Jun 2026
Model Releases

Subject-Level Unknown-Identity Identification from Leap Motion Controller 2 Hand Landmarks

DGX agent

arXiv:2606.22986v1 Announce Type: new Abstract: This work studies subject recognition from Leap Motion Controller 2 (LMC2) hand landmark data under a subject-level unknown-identity identification prot

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Surgical Anatomy Recognition with Context Learning using Foundation Representations

DGX agent

arXiv:2606.22124v1 Announce Type: new Abstract: Accurate recognition of anatomical structures is essential for safe and effective minimally invasive surgery (MIS), yet it remains underexplored in surg

researcharxiv-cs-cv
23 Jun 2026
Safety

Synergistic Dual-Branch Adaptation for Multi-modal Generalized Category Discovery

DGX agent

arXiv:2606.21446v1 Announce Type: new Abstract: Generalized Category Discovery (GCD) aims to classify old categories and discover new ones from unlabeled data. Recent multi-modal approaches introduce

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

T-IMPACT: A Severity-Aware Benchmark for Contextual Image-Text Manipulation

DGX agent

arXiv:2606.22339v1 Announce Type: new Abstract: Recent advances in vision-language models and generative editing systems have made it increasingly easy to produce persuasive multimodal misinformation

model-releasesarxiv-cs-cv
23 Jun 2026
Research

T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition

DGX agent

arXiv:2606.21607v1 Announce Type: new Abstract: Vision-language models such as CLIP have recently achieved strong performance on a wide range of visual understanding tasks. However, most existing mode

researcharxiv-cs-cv
23 Jun 2026
Research

T-VSS: Test-Time Visual Subspace Steering for Adversarial Robustness of Vision-Language Models

DGX agent

arXiv:2606.23132v1 Announce Type: new Abstract: Vision-language models (VLMs) achieve strong zero-shot recognition, but they remain highly vulnerable to adversarial perturbations. Recent test-time ada

researcharxiv-cs-cv
23 Jun 2026
Research

Technical Report for ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge: Exploring Query-Based Segmentation and Increased Spatial Context for Outdoor Scene Understanding

DGX agent

arXiv:2606.21456v1 Announce Type: new Abstract: In this report, we present our submission to the GOOSE 2D Fine-Grained Semantic Segmentation Challenge, organized as part of the Workshop on Field Robot

researcharxiv-cs-cv
23 Jun 2026
Model Releases

Technical Report for the ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge: Pretraining-Diverse Ensemble of Foundation Vision Encoders for Robust Outdoor Scene Understanding

DGX agent

arXiv:2606.23113v1 Announce Type: new Abstract: This report presents our solution for the ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge, which requires parsing unstructured outdoor s

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

TeleStyle V2: Beyond Content-Preserving Style Transfer with Self-Distillation and Distribution-Matching-Distillation

DGX agent

arXiv:2606.20709v1 Announce Type: new Abstract: Given a content reference and a style reference, content-preserving style transfer requires the model to generate stylized outputs with content and styl

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Temporally Aware Densification for Dynamic 3D Gaussian Splatting

DGX agent

arXiv:2606.23212v1 Announce Type: new Abstract: Despite modeling temporal motion, dynamic 3D Gaussian Splatting (3DGS) methods still inherit a static densification strategy that is ill-suited for dyna

model-releasesarxiv-cs-cv
23 Jun 2026
Safety

Test-Time Alignment of Text-to-Image Diffusion Models via Null-Text Embedding Optimisation

DGX agent

arXiv:2511.20889v2 Announce Type: replace Abstract: Test-time alignment (TTA) aims to adapt models to specific rewards during inference. However, existing methods tend to either under-optimise or over

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

Thalia: A Global, Multi-Modal Dataset for Volcanic Activity Monitoring

DGX agent

arXiv:2505.17782v4 Announce Type: replace Abstract: Monitoring volcanic activity is of paramount importance to safeguarding lives, infrastructure, and ecosystems. However, only a small fraction of kno

model-releasesarxiv-cs-cv
23 Jun 2026
Research

The First Assessment of PhiSat-2 Imagery for Monocular Building Height Estimation

DGX agent

arXiv:2603.29245v3 Announce Type: replace Abstract: Monocular building height estimation from optical imagery is important for characterizing urban vertical structure, yet remains challenging due to t

researcharxiv-cs-cv
23 Jun 2026
Applications

The Impact of VAE Design on Latent Pose Representations for Diffusion-based Sign Language Production

DGX agent

arXiv:2606.22959v1 Announce Type: cross Abstract: Latent diffusion approaches to sign language production (SLP) rely on an initial stage that learns an encoding of sign pose sequences, enabling genera

applicationsarxiv-cs-cv
23 Jun 2026
Model Releases

The MAMA-MIA Challenge: Advancing Generalizability and Fairness in Breast MRI Tumor Segmentation and Treatment Response Prediction

DGX agent

arXiv:2603.01250v3 Announce Type: replace Abstract: Breast cancer is the most frequently diagnosed malignancy among women worldwide and a leading cause of cancer-related mortality. Dynamic contrast-en

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

The Power of Light: Improving Synthetic-to-Real Domain Adaptation through Physically-Based Indirect Illumination

DGX agent

arXiv:2606.22574v1 Announce Type: new Abstract: While synthetic data generation resolves the manual labeling bottleneck in computer vision, minimizing the syn-to-real domain gap requires optimizing re

model-releasesarxiv-cs-cv
23 Jun 2026
Safety

The Scissors Effect: When Resize-Based Input Diversity Helps or Hurts Transfer Attacks

DGX agent

arXiv:2606.22516v1 Announce Type: cross Abstract: Input Diversity (DI), which applies random resizing and padding at each attack iteration, is a near-default ingredient of transfer-based adversarial a

safetyarxiv-cs-cv
23 Jun 2026
Research

The Unreasonable Effectiveness of VLMs for Zero-shot Procedural Mistake Detection

DGX agent

arXiv:2606.21579v1 Announce Type: new Abstract: Procedural mistake detection is important for quality control and user assistance across many disciplines. Recent work in this field has achieved signif

researcharxiv-cs-cv
23 Jun 2026
Agents

Three-Step Hierarchical Transformer for Multi-Pedestrian Trajectory Prediction

DGX agent

arXiv:2606.23058v1 Announce Type: new Abstract: Pedestrian trajectory prediction requires modeling temporal dynamics, multimodal cues, and social interactions in crowded environments. Existing methods

agentsarxiv-cs-cv
23 Jun 2026
Research

TIDY: Thermal Infrared Image Denoising via Wavelet Domain Entropy and Directional Stripe Index

DGX agent

arXiv:2606.19813v1 Announce Type: cross Abstract: Thermal infrared (TIR) imaging has been a popular choice for field robotics due to its robust perception capability under low light visual degradation

researcharxiv-cs-cv
23 Jun 2026
← Previous
1…102103104105106…263
Next →