AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Applications

Seeking the Unfamiliar but Memorable: Conceptual Creativity as Meta-Learning

DGX agent

arXiv:2605.16477v1 Announce Type: cross Abstract: What does it mean to create a new concept, rather than retrieve a familiar one? Repeatedly sampling a generative model at the same prompt produces var

applicationsarxiv-cs-cv
19 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation

DGX agent

arXiv:2605.17630v1 Announce Type: new Abstract: Here's a trimmed version under 1920 characters: Open-vocabulary segmentation models such as SAM3 achieve strong performance through concept-level text p

researcharxiv-cs-cv
19 May 2026
Research

Semantics Disentanglement and Composition for Universal Image Coding with Efficiently LLM Reasoning and Generative Diffusion

DGX agent

arXiv:2412.18158v2 Announce Type: replace Abstract: Learned image compression methods have shown impressive performance but are often highly specialized for either human perception or specific machine

researcharxiv-cs-cv
19 May 2026
Research

Semi-LAR: Semi-supervised Contrastive Learning with Linear Attention for Removal of Nighttime Flares

DGX agent

arXiv:2605.18156v1 Announce Type: new Abstract: Lens flare removal is challenging due to the large spatial extent of flare artifacts and their entanglement with scene structures, while existing method

researcharxiv-cs-cv
19 May 2026
Model Releases

Setting the Stage: Text-Driven Scene-Consistent Image Generation

DGX agent

arXiv:2512.12598v3 Announce Type: replace Abstract: We focus on the foundational task of Scene Staging: given a reference scene image and a text condition specifying an actor category to be generated

model-releasesarxiv-cs-cv
19 May 2026
Safety

SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft Signals

DGX agent

arXiv:2605.18039v1 Announce Type: new Abstract: Learning dense correspondences across deformable 3D shapes remains a long-standing challenge due to structural variability, non-isometric deformation, a

safetyarxiv-cs-cv
19 May 2026
Research

Shallow Deep Learning Can Still Excel in Fine-Grained Few-Shot Learning

DGX agent

arXiv:2507.22041v2 Announce Type: replace Abstract: Deep learning has witnessed the extensive utilization across a wide spectrum of domains, including fine-grained few-shot learning (FGFSL) which heav

researcharxiv-cs-cv
19 May 2026
Safety

SHED: Style-Homogenized Embedding Alignment for Domain Generalization

DGX agent

arXiv:2605.16973v1 Announce Type: new Abstract: Domain generalization aims to enhance model robustness against unseen domains with embedding distribution shifts. While large-scale vision-language mode

safetyarxiv-cs-cv
19 May 2026
Safety

Simple Approximation and Derivative Free Inference-Time Scaling for Diffusion Models via Sequential Monte Carlo on Path Measures

DGX agent

arXiv:2605.17850v1 Announce Type: cross Abstract: iffusion-based generative models increasingly rely on inference-time guidance, adding a drift term or reweighting mixture of experts, to improve sampl

safetyarxiv-cs-cv
19 May 2026
Tutorials

Single Image Reflection Removal with Patch Reflectance Prior

DGX agent

arXiv:2312.03798v2 Announce Type: replace Abstract: Single Image Reflection Removal (SIRR) in real-world images is a challenging task due to diverse image degradations occurring on the glass surface d

tutorialsarxiv-cs-cv
19 May 2026
Model Releases

SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning

DGX agent

arXiv:2605.17949v1 Announce Type: new Abstract: Remote sensing vision-language models commonly rely on pretrained visual encoders to convert images into semantic features before language-model reasoni

model-releasesarxiv-cs-cv
19 May 2026
Safety

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration

DGX agent

arXiv:2605.17423v1 Announce Type: new Abstract: We study series-level cinematic remaking, a long-horizon video-to-video generation problem that localizes full episodes or films via stylization or acto

safetyarxiv-cs-cv
19 May 2026
Tutorials

Sparse Autoencoders are Topic Models

DGX agent

arXiv:2511.16309v2 Announce Type: replace Abstract: Sparse autoencoders (SAEs) are used to analyze embeddings, but their role and practical value are debated. We propose a new perspective on SAEs by d

tutorialsarxiv-cs-cv
19 May 2026
Model Releases

SpecSem-Net: Integrating Spectral and Semantic Features for Robust AI-generated Video Detection

DGX agent

arXiv:2605.17311v1 Announce Type: new Abstract: The remarkable visual fidelity of recent commercial video generative models, such as Sora and Veo, renders robust AI-generated video detection increasin

model-releasesarxiv-cs-cv
19 May 2026
Research

Spectral Progressive Diffusion for Efficient Image and Video Generation

DGX agent

arXiv:2605.18736v1 Announce Type: new Abstract: Diffusion models have been shown to implicitly generate visual content autoregressively in the frequency domain, where low-frequency components are gene

researcharxiv-cs-cv
19 May 2026
Local Ai

Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI

DGX agent

arXiv:2605.18466v1 Announce Type: new Abstract: Segmenting vocal tract articulators in real-time MRI (rtMRI) is a challenging dynamic image segmentation problem characterized by low contrast, rapid mo

local-aiarxiv-cs-cv
19 May 2026
Research

SPIKE: An Adaptive Dual Controller Framework for Cost-Efficient Long-Horizon Game Agents

DGX agent

arXiv:2605.18636v1 Announce Type: new Abstract: Long-horizon multimodal agents in open-world games must stay goal-directed across many low-level interactions under tight token and latency budgets. Exi

researcharxiv-cs-cv
19 May 2026
Tutorials

SRC-Flow: Compact Semantic Representations Enable Normalizing Flows for Image Generation

DGX agent

arXiv:2605.18267v1 Announce Type: new Abstract: Normalizing flows (NFs) provide exact likelihoods and deterministic invertible sampling, but have historically lagged behind diffusion models for large-

tutorialsarxiv-cs-cv
19 May 2026
Model Releases

Stabilizing, Scaling & Enhancing MeanFlow for Large-scale Diffusion Distillation

DGX agent

arXiv:2605.17834v1 Announce Type: new Abstract: Diffusion models exhibit remarkable generative capability, but their high latency limits practical deployment. Many studies have attempted to reduce sam

model-releasesarxiv-cs-cv
19 May 2026
Safety

Stable and Near-Reversible Diffusion ODE Solvers for Image Editing

DGX agent

arXiv:2605.16399v1 Announce Type: new Abstract: The inversion of diffusion models plays a central role in image editing. Algebraically reversible ODE solvers provide an appealing approach to diffusion

safetyarxiv-cs-cv
19 May 2026
Safety

Stable Routing for Mixture-of-Experts in Class-Incremental Learning

DGX agent

arXiv:2605.17571v1 Announce Type: new Abstract: Class-incremental learning (CIL) requires models to learn new classes sequentially while preserving prior knowledge. Recently, approaches that combine p

safetyarxiv-cs-cv
19 May 2026
Model Releases

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data

DGX agent

arXiv:2605.18287v1 Announce Type: new Abstract: It is infeasible to encompass all possible disturbances within the training dataset. This raises a critical question regarding the robustness of Vision-

model-releasesarxiv-cs-cv
19 May 2026
Tutorials

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth

DGX agent

arXiv:2605.18603v1 Announce Type: new Abstract: Vision-Language Models (VLMs) deployed as situated agents in high-resolution visual environments require active perception -- the ability to dynamically

tutorialsarxiv-cs-cv
19 May 2026
Research

Statistical Hand Shape Modeling from Clinical CT Scans Using Deep Learning and Implicit Skinning

DGX agent

arXiv:2605.16980v1 Announce Type: new Abstract: Accurate segmentation and statistical shape modeling of hand anatomy have significant implications for medical diagnostics, ergonomics, and biomechanics

researcharxiv-cs-cv
19 May 2026
Model Releases

SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation

DGX agent

arXiv:2511.19320v2 Announce Type: replace Abstract: Preserving first-frame identity while ensuring precise motion control is a fundamental challenge in human image animation. The Image-to-Motion Bindi

model-releasesarxiv-cs-cv
19 May 2026
Hardware

StreamingEffect: Real-Time Human-Centric Video Effect Generation

DGX agent

arXiv:2605.17019v1 Announce Type: new Abstract: Streaming video effect generation is highly desirable for live human-centric applications such as e-commerce streaming, entertainment, and vlogging, yet

hardwarearxiv-cs-cv
19 May 2026
Research

StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion Model

DGX agent

arXiv:2511.14223v3 Announce Type: replace Abstract: This paper focuses on the task of speech-driven 3D facial animation, which aims to generate realistic and synchronized facial motions driven by spee

researcharxiv-cs-cv
19 May 2026
Research

Stroke of Surprise: Progressive Semantic Illusions in Vector Sketching

DGX agent

arXiv:2602.12280v2 Announce Type: replace Abstract: Visual illusions traditionally rely on spatial manipulations such as multi-view consistency. In this work, we introduce Progressive Semantic Illusio

researcharxiv-cs-cv
19 May 2026
Model Releases

Supervise Less, See More: Training-free Nuclear Instance Segmentation with Prototype-Guided Prompting

DGX agent

arXiv:2511.19953v2 Announce Type: replace Abstract: Accurate nuclear instance segmentation is a pivotal task in computational pathology, supporting data-driven clinical insights and facilitating downs

model-releasesarxiv-cs-cv
19 May 2026
Research

Supervised contrastive learning for cell stage classification of animal embryos

DGX agent

arXiv:2502.07360v3 Announce Type: replace-cross Abstract: Videomicroscopy, when combined with machine learning, offers a promising approach for studying the early development of in vitro produced (IVP

researcharxiv-cs-cv
19 May 2026
Model Releases

SurgLQA: Scalable Long-Horizon Surgical Video Question Answering

DGX agent

arXiv:2605.17915v1 Announce Type: new Abstract: Surgical Video Question Answering (VideoQA) provides a promising paradigm for dynamic intraoperative interpretation, enabling real-time decision support

model-releasesarxiv-cs-cv
19 May 2026
Safety

SVL: Spike-based Vision-language Pretraining for Efficient 3D Open-world Understanding

DGX agent

arXiv:2505.17674v2 Announce Type: replace Abstract: Spiking Neural Networks (SNNs) provide an energy-efficient way to extract 3D spatio-temporal features. However, existing SNNs still exhibit a signif

safetyarxiv-cs-cv
19 May 2026
Agents

SWoMo: Neuro-Symbolic World Model for Cataract Surgery Simulation

DGX agent

arXiv:2605.16530v1 Announce Type: new Abstract: Realistic surgical simulation plays a crucial role in training novice surgeons and in the development of autonomous agents. World models can scale such

agentsarxiv-cs-cv
19 May 2026
Research

Symmetry Matters: Auditing and Symmetrizing 3D Generative Models

DGX agent

arXiv:2512.18953v2 Announce Type: replace Abstract: Symmetry is a strong prior present in many object categories, yet standard benchmarks for 3D generative models rarely report whether this prior is p

researcharxiv-cs-cv
19 May 2026
Local Ai

Synthetic Aperture Radar Image Change Detection Based on Global Dynamic Context-Aware Network

DGX agent

arXiv:2605.16764v1 Announce Type: new Abstract: Convolutional neural networks (CNNs) have been extensively and successfully applied to the task of synthetic aperture radar (SAR) image change detection

local-aiarxiv-cs-cv
19 May 2026
Model Releases

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms

DGX agent

arXiv:2605.17336v1 Announce Type: cross Abstract: Tactile sensing is a fundamental modality for embodied intelligence, offering unique and direct feedback on contact geometry, material properties, and

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

TAME: Test-Time Adversarial Prompt Tuning via Mixture-of-Experts for Vision-Language Models

DGX agent

arXiv:2605.17577v1 Announce Type: new Abstract: Large-scale pre-trained Vision-Language models (VLMs), such as CLIP, exhibit strong zero-shot generalization, yet remain highly vulnerable to impercepti

model-releasesarxiv-cs-cv
19 May 2026
Research

Test-Time Hinting for Black-Box Vision-Language Models

DGX agent

arXiv:2605.16410v1 Announce Type: new Abstract: Test-time scaling (TTS) methods have proven highly effective for LLMs, yet their application to vision-language models (VLMs) remains relatively underex

researcharxiv-cs-cv
19 May 2026
Tutorials

The Learnability Gap in Medical Latent Diffusion

DGX agent

arXiv:2605.17087v1 Announce Type: new Abstract: Generative data augmentation with latent diffusion models is a promising strategy for addressing class imbalance in medical imaging, yet current approac

tutorialsarxiv-cs-cv
19 May 2026
Model Releases

The MixCount Dataset: Bridging the Data Gap for Open-Vocabulary Object Counting

DGX agent

arXiv:2605.18063v1 Announce Type: new Abstract: Object counting is a foundational vision task with over a decade of dedicated research, yet state-of-the-art models still fail systematically in the mix

model-releasesarxiv-cs-cv
19 May 2026
Tutorials

The Silent Brush: Evaluating Artistic Style Leakage in AI Art Generation

DGX agent

arXiv:2605.17500v1 Announce Type: cross Abstract: Generative text-to-image models are typically trained on large-scale web-scraped datasets that include diverse visual content such as copyrighted and

tutorialsarxiv-cs-cv
19 May 2026
Applications

Thermal-Only Crowd Counting with Deployment-Time Privacy Protection

DGX agent

arXiv:2605.17042v1 Announce Type: new Abstract: While RGB-Thermal crowd counting has shown promise, the paradigm faces critical limitations: RGB data raises privacy concerns in public surveillance, an

applicationsarxiv-cs-cv
19 May 2026
Model Releases

Threats to Arabic Handwriting Recognition: Investigating Black-Box Adversarial Attacks on embedded ConvNet models

DGX agent

arXiv:2605.18058v1 Announce Type: new Abstract: Arabic handwriting recognition (AHR) has made significant progress with deep learning models. AHR research has largely focused on performance, with secu

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval

DGX agent

arXiv:2605.18434v1 Announce Type: cross Abstract: E-commerce image search often takes a cropped image as the query, while each candidate is represented by full item images and structured text. This im

model-releasesarxiv-cs-cv
19 May 2026
Hardware

Token-Space Mask Prediction for Efficient Vision Transformer Segmentation

DGX agent

arXiv:2605.18177v1 Announce Type: new Abstract: Query-based Vision Transformer segmentation models typically reconstruct dense spatial feature maps to predict masks, inheriting design patterns from co

hardwarearxiv-cs-cv
19 May 2026
Research

Topo-GS: Continuous Volumetric Embedding of High-Dimensional Data via Topological Gaussian Splatting

DGX agent

arXiv:2605.17011v1 Announce Type: cross Abstract: Dimensionality reduction algorithms map high-dimensional data into visualizable 2D or 3D spaces, but traditionally rely on a discrete point-cloud para

researcharxiv-cs-cv
19 May 2026
Research

TouchMap-OR: Multi-View 3D Mapping of Hand-Surface Contacts

DGX agent

arXiv:2605.17638v1 Announce Type: new Abstract: Hand-surface interactions between clinicians, patients, and medical equipment play a central role in pathogen transmission during medical procedures. Ho

researcharxiv-cs-cv
19 May 2026
Local Ai

Towards Generalized Image Manipulation Localization via Score-based Model

DGX agent

arXiv:2605.16879v1 Announce Type: new Abstract: With the rapid evolution of synthetic media, Image Manipulation Localization (IML) has emerged as a critical component in multimedia forensics for ensur

local-aiarxiv-cs-cv
19 May 2026
← Previous
1…163164165166167…263
Next →