AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
23 Jun 2026

RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation

Model ReleasesDGX agent

arXiv:2606.23221v1 Announce Type: new Abstract: Recent years have witnessed remarkable progress in image generation and editing, particularly regarding instruction following and visual fidelity. Howev

RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild

Model ReleasesDGX agent

arXiv:2606.23344v1 Announce Type: new Abstract: Accurate document layout analysis remains a critical bottleneck for document parsing systems, due to the intricate coupling among heterogeneous document

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2511.20651v2 Announce Type: replace Abstract: Reinforcement learning (RL) has recently emerged as a promising approach for aligning text-to-image generative models with human preferences. A key

S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix

ResearchDGX agent

arXiv:2508.08048v2 Announce Type: replace Abstract: While video generation models excel at producing high-quality monocular videos, generating 3D stereoscopic and spatial videos for immersive applicat

Safe Few-Step Generation via Velocity Editing

SafetyDGX agent

arXiv:2606.23267v1 Announce Type: new Abstract: Flow matching has recently emerged as a strong paradigm for state-of-the-art text-to-image (T2I) generation, enabling high-quality generation with a sma

SAGE: An Expert-Annotated South Asian GI Endoscopy Dataset for Multimodal Learning and Hallucination Analysis

SafetyDGX agent

arXiv:2606.22144v1 Announce Type: new Abstract: Gastrointestinal cancers represent a growing health burden in the South Asian region, driven largely by rapid changes in socio-economic conditions & lif

SARIF: Segment Anything for Robust Image Forensics

Local AiDGX agent

arXiv:2606.21108v1 Announce Type: new Abstract: Image forgery localization remains challenging due to diverse manipulation techniques and distribution shifts. Existing forgery localization models achi

SATURN: Symbolic Spatial Reasoning for Multi-Perspective Grounding

Model ReleasesDGX agent

arXiv:2606.22694v1 Announce Type: new Abstract: Vision-Language Models (VLMs) remain unreliable when spatial reasoning requires composing relations whose meanings depend on frames of reference. Existi

Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology

TutorialsDGX agent

arXiv:2606.20477v2 Announce Type: replace Abstract: We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations. We introduce RefRad2D, a lar

ScalePredictor: Instance-aware Scale Learning for Accurate Quantization of Vision Transformers

ResearchDGX agent

arXiv:2606.21947v1 Announce Type: new Abstract: Vision Transformers have achieved remarkable success in many fields, yet their deployment on edge devices remains challenging due to their substantial c

Scaling Diverse Language Generation for 3D Visual Grounding

Local AiDGX agent

arXiv:2606.20946v1 Announce Type: cross Abstract: Developing robust models for 3D visual grounding (3DVG), the localization of entities in a 3D scene described in natural language, is important for en

Scaling Self-Play for End-to-End Driving

SafetyDGX agent

arXiv:2606.19641v2 Announce Type: replace-cross Abstract: End-to-end autonomous driving models are typically trained on offline human-demonstration datasets that provide limited state coverage and oft

Scaling State-Space Models from Lines to Paragraphs: An Ablation of Mamba-based OCR

ResearchDGX agent

arXiv:2606.23524v1 Announce Type: new Abstract: End-to-end OCR increasingly relies on autoregressive sequence models, where the quadratic cost of Transformer attention limits efficient transcription o

Scaling up fine-grained intracranial vessel annotations in computed tomography angiography

ResearchDGX agent

arXiv:2606.21756v1 Announce Type: cross Abstract: In this work, we present SemanticVessel, a dataset for fine-grained brain vessel segmentation in computed tomography angiography scans. Based on the d

ScalingAttention: Discovering Intrinsic Sparse Attention Topology for Video Diffusion Transformers

SafetyDGX agent

arXiv:2606.23019v1 Announce Type: new Abstract: While Diffusion Transformers (DiTs) have revolutionized high-fidelity video generation, their reliance on 3D full attention creates a quadratic computat

Scene-agnostic ALS boresight self-calibration

Model ReleasesDGX agent

arXiv:2606.23101v1 Announce Type: new Abstract: ALS boresight calibration has relied for two decades on dedicated flight patterns over structured scenes containing planar surfaces of varied aspect and

Scene-Level Heterogeneous Physics Simulation with 3D Gaussian Splats

ApplicationsDGX agent

arXiv:2606.21753v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) has achieved state-of-the-art photorealistic rendering, but the representation gap prevents these assets from being physi

SCOPE: Scale-Consistent One-Pass Estimation of 3D Geometry

SafetyDGX agent

arXiv:2606.21300v1 Announce Type: new Abstract: We present SCOPE (Scale-Consistent One-Pass Estimation of 3D Geometry), a novel approach for estimating 3D geometry from extended monocular video sequen

SCRUB-FL: Sanitizing and Cleansing Representations via Unlearning of Backdoors

Local AiDGX agent

arXiv:2606.22700v1 Announce Type: cross Abstract: Federated Learning (FL) enables collaborative model training without sharing raw data, making it a promising paradigm for privacy-sensitive applicatio

SEED: Simple ViT and Evolving Harness for Explainable Text Forgery Detection

Local AiDGX agent

arXiv:2606.21138v1 Announce Type: new Abstract: AI-assisted image editing threatens trust in financial, legal, and identity records. The GenText-Forensics Challenge at ACM MM 2026 addresses this by re

Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation

SafetyDGX agent

arXiv:2606.19120v2 Announce Type: replace-cross Abstract: On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-level targets conditioned

SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion

Model ReleasesDGX agent

arXiv:2606.22568v1 Announce Type: new Abstract: Training image generation foundation models consumes substantial resources. Previous methods have attempted to leverage semantic guidance to accelerate

SegTME-UNI2: A Foundation Model-Based Framework for Generalisable Multiclass Cell Segmentation and LLM-Driven Tumour Microenvironment Characterisation in Histopathology

Model ReleasesDGX agent

arXiv:2606.17702v2 Announce Type: replace Abstract: Characterising the tumour microenvironment (TME) from routine H&E-stained histology images requires simultaneous cell segmentation, feature extracti

Self-Supervised Dual-Frequency Phase Decomposition for Single-Shot Composite Fringe Projection Profilometry

ResearchDGX agent

arXiv:2606.21027v1 Announce Type: new Abstract: Single-shot fringe projection profilometry (FPP) has been actively studied for real-time measurement, dynamic object reconstruction, and motion-sensitiv

Semantic Browsing: Controllable Diversity for Image Generation

AgentsDGX agent

arXiv:2606.23679v1 Announce Type: new Abstract: Modern text-to-image models excel in visual fidelity and prompt adherence. However, this strict adherence comes at the cost of diversity: generated samp

Semi-Supervised Vision-Language-Action Model

SafetyDGX agent

arXiv:2606.21493v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models enable robots to predict actions directly from visual observations and language instructions, but adapting them to n

SenseExpo: Spatial Exploration and Navigation via Scene Estimation from Expeditious Predictive Operators

ResearchDGX agent

arXiv:2503.16000v2 Announce Type: replace Abstract: We present extbf{SenseExpo}, a lightweight single-robot exploration framework that integrates a compact map prediction network into a frontier-based

Shear-Free Viewport Magnification for 360-Degree via Spherical Mobius Boosts

ResearchDGX agent

arXiv:2606.20684v1 Announce Type: new Abstract: Viewport-adaptive 360-degree imaging seeks to allocate a fixed sampling budget to the region a viewer is likely to observe. Existing view-biased project

ShuffleFlow: Scalable Posterior Inference for Bayesian Inverse Imaging

ResearchDGX agent

arXiv:2606.21099v1 Announce Type: new Abstract: Variational inference (VI) is a powerful method for principled posterior inference for scientific inverse imaging. VI learns the posterior distribution,

SimAC: A Simple Anti-Customization Method for Protecting Face Privacy against Text-to-Image Synthesis of Diffusion Models

ResearchDGX agent

arXiv:2312.07865v4 Announce Type: replace Abstract: Despite the success of diffusion-based customization methods on visual content creation, increasing concerns have been raised about such techniques

SIMSplat: Language-Aligned 4D Gaussian Splatting for Driving Scenario Generation

AgentsDGX agent

arXiv:2510.02469v2 Announce Type: replace-cross Abstract: Driving scene manipulation using real-world sensor data has emerged as a promising alternative to traditional driving simulators. Despite adva

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning

Model ReleasesDGX agent

arXiv:2606.22873v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications. This broad deployment expands the

Skeleton-to-Image Encoding: Enabling Skeleton Representation Learning via Vision-Pretrained Models

ResearchDGX agent

arXiv:2603.05963v2 Announce Type: replace Abstract: Recent advances in large-scale pretrained vision models have demonstrated impressive capabilities across a wide range of downstream tasks, including

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models

SafetyDGX agent

arXiv:2606.23041v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success in visual understanding but remain constrained in visual generation due to the

SPARC: A Multi-Agent System for Electrical Circuit Question Answering

AgentsDGX agent

arXiv:2606.20643v1 Announce Type: cross Abstract: Electrical circuit diagram QA tasks require complex mathematical reasoning, which remains challenging for multimodal LLMs. We present SPARC, a multi-a

Sparse Point-Guided Fusion of Supervised and Self-Supervised Learning Model for Seaweed Segmentation

TutorialsDGX agent

arXiv:2606.21026v1 Announce Type: new Abstract: The ocean plays a critical role in sustainable development, particularly in climate change mitigation. Among marine ecosystems, blue carbon ecosystems a

SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation

Model ReleasesDGX agent

arXiv:2605.24354v2 Announce Type: replace Abstract: Recently, world models have made significant progress in enhancing end-to-end driving systems through both future situation forecasting and improved

Spatially Grounded Concept-Based Image Classification

ResearchDGX agent

arXiv:2510.04180v2 Announce Type: replace Abstract: Deep neural networks can achieve high accuracy while relying on evidence that is hard to inspect or misaligned with the intended task. Concept Bottl

Spatio-Temporal Wildfire Spread Prediction in Canada using a Video Swin-Hybrid-U-Net and Satellite Imagery

ResearchDGX agent

arXiv:2606.20693v1 Announce Type: new Abstract: Background: Wildfires in Canada present increasing threats to ecosystems, communities, and infrastructure, demanding accurate forecasting tools to aid m

Specificity- and Calibration-Aware Breast Ultrasound Segmentation via Entropy-Guided Boundary Supervision

ResearchDGX agent

arXiv:2606.22308v1 Announce Type: cross Abstract: Lesion segmentation in breast ultrasound involves two related challenges. In images with lesions, speckle noise, low tissue contrast, and posterior ac

Spectral Gating via Damped Oscillations for Adaptive Implicit Neural Representations

SafetyDGX agent

arXiv:2606.23129v1 Announce Type: new Abstract: Implicit Neural Representations (INRs) have been proven successful in encoding continuous signals through coordinate-based networks, yet facing a spectr

Spectral GS-SLAM: Observability-Aware, Degeneracy-Robust Tracking for Real-Time 3D Gaussian Splatting SLAM

TutorialsDGX agent

arXiv:2606.21258v1 Announce Type: cross Abstract: Recent 3DGS-SLAM systems enable real-time operation by leveraging conventional feature matching or ICP-based tracking, thereby avoiding the heavy dens

SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs

Local AiDGX agent

arXiv:2606.20244v2 Announce Type: replace Abstract: Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and easy to over

Stabilizing Consistency Training: A Flow Map Analysis and Self-Distillation

SafetyDGX agent

arXiv:2601.22679v2 Announce Type: replace-cross Abstract: Consistency models have been proposed for fast generative modeling, achieving results competitive with diffusion and flow models. However, the

SteerVTE: Seamless Video Text Editing with Style and Glyph Control

Local AiDGX agent

arXiv:2606.23254v1 Announce Type: new Abstract: Visual text editing aims to precisely modify text in images and videos while preserving stylistic consistency and visual realism. Despite significant ad

Stochastic Signed Distance Processes

ResearchDGX agent

arXiv:2606.20856v1 Announce Type: new Abstract: Multi-view surface reconstruction is a core problem in computer vision. One prominent line of work represents the surface implicitly as a signed distanc

Streaming Dense Voxel Representations for 3D Occupancy Prediction

Model ReleasesDGX agent

arXiv:2503.22087v3 Announce Type: replace Abstract: In this paper, we explore dense voxel streaming for accurate and efficient 3D occupancy prediction. While dense voxel representations offer fine-gra

StreamPPG: Low-Latency rPPG Estimation via Consistent Privileged Learning

ResearchDGX agent

arXiv:2606.23186v1 Announce Type: new Abstract: Remote photoplethysmography (rPPG) estimates the blood volume pulse (BVP) signal from facial videos, enabling contact-free health monitoring. Convention

StructSAM: Structure- and Spectrum-Preserving Token Merging for Segment Anything Models

ResearchDGX agent

arXiv:2603.07307v2 Announce Type: replace Abstract: Recent token merging techniques for Vision Transformers (ViTs) provide substantial speedups by reducing the number of tokens processed by self-atten

Structural Assessment for Understanding and Guiding Dataset Distillation in Discrete Token Space

TutorialsDGX agent

arXiv:2606.21705v1 Announce Type: new Abstract: Dataset distillation (DD) has proven to reduce training cost while preserving accuracy. While promising, the factors that make one distilled dataset mor

Structured Hyperedge Adaptation for Parameter-Efficient Fine-Tuning of Vision Transformers

Model ReleasesDGX agent

arXiv:2606.22383v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) has become a practical solution for adapting large pretrained vision transformers (ViTs) to downstream tasks whil

Structured Spectral Graph Representation Learning for Multi-label Abnormality Analysis from 3D CT Scans

ResearchDGX agent

arXiv:2510.10779v5 Announce Type: replace Abstract: With the growing volume of CT examinations, there is an increasing demand for automated tools such as organ segmentation, abnormality detection, and

Subject-Level Unknown-Identity Identification from Leap Motion Controller 2 Hand Landmarks

Model ReleasesDGX agent

arXiv:2606.22986v1 Announce Type: new Abstract: This work studies subject recognition from Leap Motion Controller 2 (LMC2) hand landmark data under a subject-level unknown-identity identification prot

Surgical Anatomy Recognition with Context Learning using Foundation Representations

ResearchDGX agent

arXiv:2606.22124v1 Announce Type: new Abstract: Accurate recognition of anatomical structures is essential for safe and effective minimally invasive surgery (MIS), yet it remains underexplored in surg

Synergistic Dual-Branch Adaptation for Multi-modal Generalized Category Discovery

SafetyDGX agent

arXiv:2606.21446v1 Announce Type: new Abstract: Generalized Category Discovery (GCD) aims to classify old categories and discover new ones from unlabeled data. Recent multi-modal approaches introduce

T-IMPACT: A Severity-Aware Benchmark for Contextual Image-Text Manipulation

Model ReleasesDGX agent

arXiv:2606.22339v1 Announce Type: new Abstract: Recent advances in vision-language models and generative editing systems have made it increasingly easy to produce persuasive multimodal misinformation

T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition

ResearchDGX agent

arXiv:2606.21607v1 Announce Type: new Abstract: Vision-language models such as CLIP have recently achieved strong performance on a wide range of visual understanding tasks. However, most existing mode

T-VSS: Test-Time Visual Subspace Steering for Adversarial Robustness of Vision-Language Models

ResearchDGX agent

arXiv:2606.23132v1 Announce Type: new Abstract: Vision-language models (VLMs) achieve strong zero-shot recognition, but they remain highly vulnerable to adversarial perturbations. Recent test-time ada

Technical Report for ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge: Exploring Query-Based Segmentation and Increased Spatial Context for Outdoor Scene Understanding

ResearchDGX agent

arXiv:2606.21456v1 Announce Type: new Abstract: In this report, we present our submission to the GOOSE 2D Fine-Grained Semantic Segmentation Challenge, organized as part of the Workshop on Field Robot

Technical Report for the ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge: Pretraining-Diverse Ensemble of Foundation Vision Encoders for Robust Outdoor Scene Understanding

Model ReleasesDGX agent

arXiv:2606.23113v1 Announce Type: new Abstract: This report presents our solution for the ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge, which requires parsing unstructured outdoor s

← Previous
1…8182838485…211
Next →