AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
19 May 2026

Robo-Cortex: A Self-Evolving Embodied Agent via Dual-Grain Cognitive Memory and Autonomous Knowledge Induction

Local AiDGX agent

arXiv:2605.18729v1 Announce Type: cross Abstract: The ability to navigate and interact with complex environments is central to real-world embodied agents, yet navigation in unseen environments remains

ROVR-Open-Dataset: A Large-Scale Depth Dataset for Autonomous Driving

Model ReleasesDGX agent

arXiv:2508.13977v3 Announce Type: replace Abstract: Depth estimation is a fundamental component of spatial perception for autonomous driving and other unmanned systems operating in open urban environm

RSEdit: Text-Guided Image Editing for Remote Sensing

ResearchDGX agent

arXiv:2603.13708v2 Announce Type: replace Abstract: In this paper, we explore text-guided image editing in the remote sensing domain using generative modeling. We propose rsedit, a collection of model


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

RT-Splatting: Joint Reflection-Transmission Modeling with Gaussian Splatting

ResearchDGX agent

arXiv:2605.18263v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) enables real-time novel view synthesis with high visual quality. However, existing methods struggle with semi-transparent s

SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training

SafetyDGX agent

arXiv:2605.18719v1 Announce Type: new Abstract: Diffusion models have been widely studied for removing unsafe content learned during pre-training. Existing methods require expensive supervised data, e

SAM 2++: Tracking Anything at Any Granularity

Model ReleasesDGX agent

arXiv:2510.18822v4 Announce Type: replace Abstract: Due to the varying granularity of target states across different tasks, most existing trackers are tailored to a single task, which specificity limi

SAMRI: Segment Any MRI

Local AiDGX agent

arXiv:2510.26635v3 Announce Type: replace-cross Abstract: Summary: SAMRI is an MRI-specialized adaptation of the Segment Anything Model achieving superior whole-body MRI segmentation, particularly for

SCAR: Self-Supervised Continuous Action Representation Learning

ResearchDGX agent

arXiv:2605.16412v1 Announce Type: cross Abstract: Despite the central role of action in embodied intelligence, learning transferable action representations from visual transitions remains a fundamenta

SCARED-C: Corrected Camera Poses for Endoscopic Depth Estimation

Model ReleasesDGX agent

arXiv:2605.16628v1 Announce Type: new Abstract: The SCARED dataset is a widely used benchmark for endoscopic depth estimation, offering ground-truth 3D reconstructions captured with a structured light

SeamCam: Quantifying Seamless Camouflage via Multi-Cue Visual Detectability

Local AiDGX agent

arXiv:2605.16515v1 Announce Type: new Abstract: Animals are described as effectively camouflaged when they blend seamlessly with their surrounding, yet no standardized quantitative measure of this sea

See Silhouettes in Motion with Neuromorphic Vision

ResearchDGX agent

arXiv:2605.17984v1 Announce Type: cross Abstract: Quasi-bimodal objects, such as text, road signs, and barcodes, play a basic yet vital role in daily visual communication. By boiling these down to cle

Seeing Together:Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2605.18431v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have made substantial progress in egocentric video understanding, but their ability to reason cooperatively fro

Seeking the Unfamiliar but Memorable: Conceptual Creativity as Meta-Learning

ApplicationsDGX agent

arXiv:2605.16477v1 Announce Type: cross Abstract: What does it mean to create a new concept, rather than retrieve a familiar one? Repeatedly sampling a generative model at the same prompt produces var

SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation

ResearchDGX agent

arXiv:2605.17630v1 Announce Type: new Abstract: Here's a trimmed version under 1920 characters: Open-vocabulary segmentation models such as SAM3 achieve strong performance through concept-level text p

Semantics Disentanglement and Composition for Universal Image Coding with Efficiently LLM Reasoning and Generative Diffusion

ResearchDGX agent

arXiv:2412.18158v2 Announce Type: replace Abstract: Learned image compression methods have shown impressive performance but are often highly specialized for either human perception or specific machine

Semi-LAR: Semi-supervised Contrastive Learning with Linear Attention for Removal of Nighttime Flares

ResearchDGX agent

arXiv:2605.18156v1 Announce Type: new Abstract: Lens flare removal is challenging due to the large spatial extent of flare artifacts and their entanglement with scene structures, while existing method

Setting the Stage: Text-Driven Scene-Consistent Image Generation

Model ReleasesDGX agent

arXiv:2512.12598v3 Announce Type: replace Abstract: We focus on the foundational task of Scene Staging: given a reference scene image and a text condition specifying an actor category to be generated

SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft Signals

SafetyDGX agent

arXiv:2605.18039v1 Announce Type: new Abstract: Learning dense correspondences across deformable 3D shapes remains a long-standing challenge due to structural variability, non-isometric deformation, a

Shallow Deep Learning Can Still Excel in Fine-Grained Few-Shot Learning

ResearchDGX agent

arXiv:2507.22041v2 Announce Type: replace Abstract: Deep learning has witnessed the extensive utilization across a wide spectrum of domains, including fine-grained few-shot learning (FGFSL) which heav

SHED: Style-Homogenized Embedding Alignment for Domain Generalization

SafetyDGX agent

arXiv:2605.16973v1 Announce Type: new Abstract: Domain generalization aims to enhance model robustness against unseen domains with embedding distribution shifts. While large-scale vision-language mode

Simple Approximation and Derivative Free Inference-Time Scaling for Diffusion Models via Sequential Monte Carlo on Path Measures

SafetyDGX agent

arXiv:2605.17850v1 Announce Type: cross Abstract: iffusion-based generative models increasingly rely on inference-time guidance, adding a drift term or reweighting mixture of experts, to improve sampl

Single Image Reflection Removal with Patch Reflectance Prior

TutorialsDGX agent

arXiv:2312.03798v2 Announce Type: replace Abstract: Single Image Reflection Removal (SIRR) in real-world images is a challenging task due to diverse image degradations occurring on the glass surface d

SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning

Model ReleasesDGX agent

arXiv:2605.17949v1 Announce Type: new Abstract: Remote sensing vision-language models commonly rely on pretrained visual encoders to convert images into semantic features before language-model reasoni

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration

SafetyDGX agent

arXiv:2605.17423v1 Announce Type: new Abstract: We study series-level cinematic remaking, a long-horizon video-to-video generation problem that localizes full episodes or films via stylization or acto

Sparse Autoencoders are Topic Models

TutorialsDGX agent

arXiv:2511.16309v2 Announce Type: replace Abstract: Sparse autoencoders (SAEs) are used to analyze embeddings, but their role and practical value are debated. We propose a new perspective on SAEs by d

SpecSem-Net: Integrating Spectral and Semantic Features for Robust AI-generated Video Detection

Model ReleasesDGX agent

arXiv:2605.17311v1 Announce Type: new Abstract: The remarkable visual fidelity of recent commercial video generative models, such as Sora and Veo, renders robust AI-generated video detection increasin

Spectral Progressive Diffusion for Efficient Image and Video Generation

ResearchDGX agent

arXiv:2605.18736v1 Announce Type: new Abstract: Diffusion models have been shown to implicitly generate visual content autoregressively in the frequency domain, where low-frequency components are gene

Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI

Local AiDGX agent

arXiv:2605.18466v1 Announce Type: new Abstract: Segmenting vocal tract articulators in real-time MRI (rtMRI) is a challenging dynamic image segmentation problem characterized by low contrast, rapid mo

SPIKE: An Adaptive Dual Controller Framework for Cost-Efficient Long-Horizon Game Agents

ResearchDGX agent

arXiv:2605.18636v1 Announce Type: new Abstract: Long-horizon multimodal agents in open-world games must stay goal-directed across many low-level interactions under tight token and latency budgets. Exi

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation

TutorialsDGX agent

arXiv:2605.18267v1 Announce Type: new Abstract: Normalizing flows (NFs) provide exact likelihoods and deterministic invertible sampling, but have historically lagged behind diffusion models for large-

Stabilizing, Scaling & Enhancing MeanFlow for Large-scale Diffusion Distillation

Model ReleasesDGX agent

arXiv:2605.17834v1 Announce Type: new Abstract: Diffusion models exhibit remarkable generative capability, but their high latency limits practical deployment. Many studies have attempted to reduce sam

Stable and Near-Reversible Diffusion ODE Solvers for Image Editing

SafetyDGX agent

arXiv:2605.16399v1 Announce Type: new Abstract: The inversion of diffusion models plays a central role in image editing. Algebraically reversible ODE solvers provide an appealing approach to diffusion

Stable Routing for Mixture-of-Experts in Class-Incremental Learning

SafetyDGX agent

arXiv:2605.17571v1 Announce Type: new Abstract: Class-incremental learning (CIL) requires models to learn new classes sequentially while preserving prior knowledge. Recently, approaches that combine p

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data

Model ReleasesDGX agent

arXiv:2605.18287v1 Announce Type: new Abstract: It is infeasible to encompass all possible disturbances within the training dataset. This raises a critical question regarding the robustness of Vision-

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth

TutorialsDGX agent

arXiv:2605.18603v1 Announce Type: new Abstract: Vision-Language Models (VLMs) deployed as situated agents in high-resolution visual environments require active perception -- the ability to dynamically

Statistical Hand Shape Modeling from Clinical CT Scans Using Deep Learning and Implicit Skinning

ResearchDGX agent

arXiv:2605.16980v1 Announce Type: new Abstract: Accurate segmentation and statistical shape modeling of hand anatomy have significant implications for medical diagnostics, ergonomics, and biomechanics

SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation

Model ReleasesDGX agent

arXiv:2511.19320v2 Announce Type: replace Abstract: Preserving first-frame identity while ensuring precise motion control is a fundamental challenge in human image animation. The Image-to-Motion Bindi

StreamingEffect: Real-Time Human-Centric Video Effect Generation

HardwareDGX agent

arXiv:2605.17019v1 Announce Type: new Abstract: Streaming video effect generation is highly desirable for live human-centric applications such as e-commerce streaming, entertainment, and vlogging, yet

StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion Model

ResearchDGX agent

arXiv:2511.14223v3 Announce Type: replace Abstract: This paper focuses on the task of speech-driven 3D facial animation, which aims to generate realistic and synchronized facial motions driven by spee

Stroke of Surprise: Progressive Semantic Illusions in Vector Sketching

ResearchDGX agent

arXiv:2602.12280v2 Announce Type: replace Abstract: Visual illusions traditionally rely on spatial manipulations such as multi-view consistency. In this work, we introduce Progressive Semantic Illusio

Supervise Less, See More: Training-free Nuclear Instance Segmentation with Prototype-Guided Prompting

Model ReleasesDGX agent

arXiv:2511.19953v2 Announce Type: replace Abstract: Accurate nuclear instance segmentation is a pivotal task in computational pathology, supporting data-driven clinical insights and facilitating downs

Supervised contrastive learning for cell stage classification of animal embryos

ResearchDGX agent

arXiv:2502.07360v3 Announce Type: replace-cross Abstract: Videomicroscopy, when combined with machine learning, offers a promising approach for studying the early development of in vitro produced (IVP

SurgLQA: Scalable Long-Horizon Surgical Video Question Answering

Model ReleasesDGX agent

arXiv:2605.17915v1 Announce Type: new Abstract: Surgical Video Question Answering (VideoQA) provides a promising paradigm for dynamic intraoperative interpretation, enabling real-time decision support

SVL: Spike-based Vision-language Pretraining for Efficient 3D Open-world Understanding

SafetyDGX agent

arXiv:2505.17674v2 Announce Type: replace Abstract: Spiking Neural Networks (SNNs) provide an energy-efficient way to extract 3D spatio-temporal features. However, existing SNNs still exhibit a signif

SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation

AgentsDGX agent

arXiv:2605.16530v1 Announce Type: new Abstract: Realistic surgical simulation plays a crucial role in training novice surgeons and in the development of autonomous agents. World models can scale such

Symmetry Matters: Auditing and Symmetrizing 3D Generative Models

ResearchDGX agent

arXiv:2512.18953v2 Announce Type: replace Abstract: Symmetry is a strong prior present in many object categories, yet standard benchmarks for 3D generative models rarely report whether this prior is p

Synthetic Aperture Radar Image Change Detection Based on Global Dynamic Context-Aware Network

Local AiDGX agent

arXiv:2605.16764v1 Announce Type: new Abstract: Convolutional neural networks (CNNs) have been extensively and successfully applied to the task of synthetic aperture radar (SAR) image change detection

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms

Model ReleasesDGX agent

arXiv:2605.17336v1 Announce Type: cross Abstract: Tactile sensing is a fundamental modality for embodied intelligence, offering unique and direct feedback on contact geometry, material properties, and

TAME: Test-Time Adversarial Prompt Tuning via Mixture-of-Experts for Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.17577v1 Announce Type: new Abstract: Large-scale pre-trained Vision-Language models (VLMs), such as CLIP, exhibit strong zero-shot generalization, yet remain highly vulnerable to impercepti

Test-Time Hinting for Black-Box Vision-Language Models

ResearchDGX agent

arXiv:2605.16410v1 Announce Type: new Abstract: Test-time scaling (TTS) methods have proven highly effective for LLMs, yet their application to vision-language models (VLMs) remains relatively underex

The Learnability Gap in Medical Latent Diffusion

TutorialsDGX agent

arXiv:2605.17087v1 Announce Type: new Abstract: Generative data augmentation with latent diffusion models is a promising strategy for addressing class imbalance in medical imaging, yet current approac

The MixCount Dataset: Bridging the Data Gap for Open-Vocabulary Object Counting

Model ReleasesDGX agent

arXiv:2605.18063v1 Announce Type: new Abstract: Object counting is a foundational vision task with over a decade of dedicated research, yet state-of-the-art models still fail systematically in the mix

The Silent Brush: Evaluating Artistic Style Leakage in AI Art Generation

TutorialsDGX agent

arXiv:2605.17500v1 Announce Type: cross Abstract: Generative text-to-image models are typically trained on large-scale web-scraped datasets that include diverse visual content such as copyrighted and

Thermal-Only Crowd Counting with Deployment-Time Privacy Protection

ApplicationsDGX agent

arXiv:2605.17042v1 Announce Type: new Abstract: While RGB-Thermal crowd counting has shown promise, the paradigm faces critical limitations: RGB data raises privacy concerns in public surveillance, an

Threats to Arabic Handwriting Recognition: Investigating Black-Box Adversarial Attacks on embedded ConvNet models

Model ReleasesDGX agent

arXiv:2605.18058v1 Announce Type: new Abstract: Arabic handwriting recognition (AHR) has made significant progress with deep learning models. AHR research has largely focused on performance, with secu

TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval

Model ReleasesDGX agent

arXiv:2605.18434v1 Announce Type: cross Abstract: E-commerce image search often takes a cropped image as the query, while each candidate is represented by full item images and structured text. This im

Token-Space Mask Prediction for Efficient Vision Transformer Segmentation

HardwareDGX agent

arXiv:2605.18177v1 Announce Type: new Abstract: Query-based Vision Transformer segmentation models typically reconstruct dense spatial feature maps to predict masks, inheriting design patterns from co

Topo-GS: Continuous Volumetric Embedding of High-Dimensional Data via Topological Gaussian Splatting

ResearchDGX agent

arXiv:2605.17011v1 Announce Type: cross Abstract: Dimensionality reduction algorithms map high-dimensional data into visualizable 2D or 3D spaces, but traditionally rely on a discrete point-cloud para

TouchMap-OR: Multi-View 3D Mapping of Hand-Surface Contacts

ResearchDGX agent

arXiv:2605.17638v1 Announce Type: new Abstract: Hand-surface interactions between clinicians, patients, and medical equipment play a central role in pathogen transmission during medical procedures. Ho

Towards Generalized Image Manipulation Localization via Score-based Model

Local AiDGX agent

arXiv:2605.16879v1 Announce Type: new Abstract: With the rapid evolution of synthetic media, Image Manipulation Localization (IML) has emerged as a critical component in multimedia forensics for ensur

← Previous
1…130131132133134…211
Next →