AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection

DGX agent

arXiv:2605.30544v1 Announce Type: new Abstract: Visual monitoring systems that rely on cloud-based AI inference expose raw image data to external services, creating fundamental tensions with the data-

model-releasesarxiv-cs-cv
1 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

On the Illusion of Gender Bias in Face Recognition: Explaining the Fairness Issue Through Non-demographic Attributes

DGX agent

arXiv:2501.12020v2 Announce Type: replace Abstract: Face recognition systems (FRS) exhibit significant accuracy differences based on the user's gender. Since such a gender gap reduces the trustworthin

safetyarxiv-cs-cv
1 Jun 2026
Safety

Optimizing Rank for High-Fidelity Implicit Neural Representations

DGX agent

arXiv:2512.14366v2 Announce Type: replace Abstract: Implicit Neural Representations (INRs) based on vanilla Multi-Layer Perceptrons (MLPs) are widely believed to be incapable of representing high-freq

safetyarxiv-cs-cv
1 Jun 2026
Safety

Parallel Tempering Initial Sampling in Inference-Time Reward Alignment

DGX agent

arXiv:2605.30991v1 Announce Type: cross Abstract: Inference-time reward alignment steers pretrained diffusion and flow-based generative models to satisfy user-specified rewards without retraining. Rec

safetyarxiv-cs-cv
1 Jun 2026
Research

PEEK: Picking Essential frames via Efficient Knowledge distillation

DGX agent

arXiv:2605.31029v1 Announce Type: new Abstract: Video-language models can process only a limited number of frames, making frame selection a key bottleneck for efficient video captioning. Most captioni

researcharxiv-cs-cv
1 Jun 2026
Tutorials

Personalize Your Large Vision-language Models With In-context Prompt Tuning

DGX agent

arXiv:2605.31513v1 Announce Type: new Abstract: Large vision-language models (LVLMs) have demonstrated strong general multimodal capability and are increasingly deployed in downstream systems. This tr

tutorialsarxiv-cs-cv
1 Jun 2026
Research

PolSAR Image Classification using a Hybrid Complex-Valued Network (HybridCVNet)

DGX agent

arXiv:2605.31137v1 Announce Type: new Abstract: Recently, convolutional neural networks (CNNs) have become popular for image classification due to their effectiveness in computer vision tasks. Now, re

researcharxiv-cs-cv
1 Jun 2026
Research

Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic Conditioning

DGX agent

arXiv:2605.31115v1 Announce Type: new Abstract: Dual-hand action segmentation, densely predicting actions for both hands from untrimmed videos, is essential for understanding complex bimanual activiti

researcharxiv-cs-cv
1 Jun 2026
Research

Position-Blind Ptychography: Viability of image reconstruction via data-driven variational inference

DGX agent

arXiv:2509.25269v3 Announce Type: replace-cross Abstract: In this work, we present and investigate the novel blind inverse problem of position-blind ptychography, i.e., ptychographic phase retrieval w

researcharxiv-cs-cv
1 Jun 2026
Safety

PostCam: Camera-Controllable Novel-View Video Generation with Query-Shared Cross-Attention

DGX agent

arXiv:2511.17185v2 Announce Type: replace Abstract: We propose PostCam, a streamlined framework for novel-view video generation that achieves superior detail preservation and precise camera trajectory

safetyarxiv-cs-cv
1 Jun 2026
Model Releases

PRISM: Progressive Reasoning through Iterative Slot Memory for Vision

DGX agent

arXiv:2605.30942v1 Announce Type: new Abstract: Modern vision models process images in a single feed-forward pass, which limits their ability to recover missing evidence or refine uncertain representa

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Probabilistic Precipitation Nowcasting with Rectified Flow Transformers

DGX agent

arXiv:2605.31204v1 Announce Type: new Abstract: Accurate weather forecasts are essential across various domains and are safety-critical in extreme weather conditions. Compared to simulation-based fore

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

QVGGT: Post-Training Quantized Visual Geometry Grounded Transformer

DGX agent

arXiv:2605.31124v1 Announce Type: new Abstract: Estimating 3D attributes directly from images has advanced rapidly with the Visual Geometry Grounded Transformer (VGGT), which predicts camera parameter

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Recognizing Co-Speech Gestures in-the-Wild

DGX agent

arXiv:2605.31589v1 Announce Type: new Abstract: While humans naturally gesture during speech, only a sparse subset of these movements are visually depictive and semantically linked to specific spoken

model-releasesarxiv-cs-cv
1 Jun 2026
Research

Rectified flow-based prediction of post-treatment brain MRI from pre-radiotherapy priors for patients with glioma

DGX agent

arXiv:2603.08385v2 Announce Type: replace-cross Abstract: Brain tumors result in 20 years of lost life on average. Standard therapies induce complex structural changes in the brain that are monitored

researcharxiv-cs-cv
1 Jun 2026
Applications

ReGuLaR: Relation-Grounded Latent Reasoning for Large Vision-Language Models

DGX agent

arXiv:2605.30587v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning has significantly improved the reasoning ability of large vision-language models (LVLMs) by verbalizing intermediate re

applicationsarxiv-cs-cv
1 Jun 2026
Tutorials

Remembering by Reconstructing: Domain Incremental Learning With Test-Time Training on Video Streams

DGX agent

arXiv:2605.31108v1 Announce Type: new Abstract: In this work we introduce a novel approach to domain incremental learning, adapting models over time to evolving, non-stationary data. In contrast to ot

tutorialsarxiv-cs-cv
1 Jun 2026
Tutorials

Representation Forcing for Bottleneck-Free Unified Multimodal Models

DGX agent

arXiv:2605.31604v1 Announce Type: new Abstract: Unified multimodal models (UMMs) aim to handle perception and generation in a single model. Yet existing UMMs still rely on a frozen, separately pretrai

tutorialsarxiv-cs-cv
1 Jun 2026
Local Ai

Rethinking Efficient Crack Segmentation with Task-Aligned Structural-Directional Modeling

DGX agent

arXiv:2605.31048v1 Announce Type: new Abstract: Recent crack segmentation methods often follow generic semantic segmentation designs, using stronger backbones, hybrid CNN-Transformer-Mamba encoders, a

local-aiarxiv-cs-cv
1 Jun 2026
Tutorials

Robust Dreamer: Deviation-Aware Latent Gaussian Memory for Action-Controlled AR Video Generation

DGX agent

arXiv:2605.30855v1 Announce Type: new Abstract: Frame-wise action-controlled image-to-video generation is a promising paradigm for interactive world simulation, where each control signal should elicit

tutorialsarxiv-cs-cv
1 Jun 2026
Safety

Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise Optimization

DGX agent

arXiv:2412.03876v2 Announce Type: replace Abstract: Text-to-Image (T2I) diffusion models are widely recognized for their ability to generate high-quality and diverse images based on text prompts. Howe

safetyarxiv-cs-cv
1 Jun 2026
Model Releases

SAW-Bench: Learning Situated Awareness in the Real World

DGX agent

arXiv:2602.16682v2 Announce Type: replace Abstract: A core aspect of human perception is situated awareness, the ability to relate ourselves to the surrounding physical environment and reason over pos

model-releasesarxiv-cs-cv
1 Jun 2026
Local Ai

SDF2Net: Shallow to Deep Feature Fusion Network for PolSAR Image Classification

DGX agent

arXiv:2402.17672v2 Announce Type: replace Abstract: Polarimetric synthetic aperture radar (PolSAR) images encompass valuable information that can facilitate extensive land cover interpretation and gen

local-aiarxiv-cs-cv
1 Jun 2026
Model Releases

Self-Tuning Regularization for Image Scanning Microscopy

DGX agent

arXiv:2605.31426v1 Announce Type: cross Abstract: Image Scanning Microscopy (ISM) is a fluorescence imaging technique that combines detector-array acquisition and computational reconstruction to achie

model-releasesarxiv-cs-cv
1 Jun 2026
Research

Sinkhorn Normalization of Diffusion Kernels

DGX agent

arXiv:2507.06161v2 Announce Type: replace Abstract: Smoothing a signal based on local neighborhoods is a core operation in machine learning and geometry processing. On well-structured domains such as

researcharxiv-cs-cv
1 Jun 2026
Research

Skin Lesion Classification Based on ResNet-50 Enhanced With Adaptive Spatial Feature Fusion

DGX agent

arXiv:2510.03876v2 Announce Type: replace Abstract: Skin cancer classification is challenging due to high inter-class similarity, intra-class variability, and artifacts in dermoscopic images. To addre

researcharxiv-cs-cv
1 Jun 2026
Research

SLAP: The Semantic Least Action Principle for Variational Video-Language Modeling

DGX agent

arXiv:2605.30750v1 Announce Type: new Abstract: In the era of Large Video-Language Models (LVLMs), the computational necessity of sparse frame sampling creates a fundamental ``temporal gap'', renderin

researcharxiv-cs-cv
1 Jun 2026
Research

SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation

DGX agent

arXiv:2605.31033v1 Announce Type: new Abstract: Streaming video generation models typically rely on temporal-centric memory, which organizes historical context as raw frames, chunk segments, or unclus

researcharxiv-cs-cv
1 Jun 2026
Research

SMART: SMPLest-X Mesh Adaptation and RAFT Tracking for Soccer Pose Estimation

DGX agent

arXiv:2605.31551v1 Announce Type: new Abstract: We present our approach to the FIFA Skeletal Tracking Challenge 2026, which requires estimating 3D world-space poses of soccer players from broadcast vi

researcharxiv-cs-cv
1 Jun 2026
Model Releases

SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models

DGX agent

arXiv:2605.31597v1 Announce Type: new Abstract: Measuring structured object understanding in vision foundation models remains challenging due to inconsistent evaluation protocols and limited part-leve

model-releasesarxiv-cs-cv
1 Jun 2026
Safety

SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing

DGX agent

arXiv:2605.25193v2 Announce Type: replace Abstract: Visual and acoustic events in the physical world are inherently coupled, yet existing video editing methods typically adopt decoupled pipelines, lac

safetyarxiv-cs-cv
1 Jun 2026
Research

SteerFace: Debiasing Synthetic Face Generation via Adaptive Residue Perturbation

DGX agent

arXiv:2605.30894v1 Announce Type: new Abstract: The shortage of legally compliant data for face recognition training has sparked growing interest in using synthetic data as an alternative. While recen

researcharxiv-cs-cv
1 Jun 2026
Research

Student Capacity Moderates Knowledge Distillation Effectiveness: A Systematic Study Across ResNet Teacher-Student Pairs on CIFAR-10

DGX agent

arXiv:2605.31191v1 Announce Type: cross Abstract: We investigate how teacher-student capacity relationships modulate knowledge distillation (KD) effectiveness in ResNet-based image classification on C

researcharxiv-cs-cv
1 Jun 2026
Local Ai

SurGe: Improved Surface Geometry in Point Maps

DGX agent

arXiv:2605.31577v1 Announce Type: new Abstract: Recent feedforward 3D reconstruction methods predict point maps and estimate global 3D geometry remarkably well. However, their predictions still exhibi

local-aiarxiv-cs-cv
1 Jun 2026
Model Releases

SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence

DGX agent

arXiv:2605.31529v1 Announce Type: new Abstract: True video intelligence demands more than recognizing what is visible: it requires reasoning about why events unfold, predicting what would change under

model-releasesarxiv-cs-cv
1 Jun 2026
Safety

TALON: Token-Aligned Lightweight Adapters for 6-DoF Spacecraft Pose Estimation

DGX agent

arXiv:2605.31217v1 Announce Type: new Abstract: Monocular 6-DoF spacecraft pose estimation methods predominantly process individual frames, discarding the temporal information present in an image sequ

safetyarxiv-cs-cv
1 Jun 2026
Safety

Task-Focused Memorization for Multimodal Agents

DGX agent

arXiv:2605.31075v1 Announce Type: new Abstract: Long-term memory is essential for multimodal agents to build coherent experience, accumulate world knowledge, and achieve continual learning. However, c

safetyarxiv-cs-cv
1 Jun 2026
Tutorials

Text-guided Feature Disentanglement for Cross-modal Gait Recognition

DGX agent

arXiv:2605.30784v1 Announce Type: new Abstract: Gait recognition is a biometric technique that identifies individuals based on their walking patterns, offering advantages in long-range, non-intrusive

tutorialsarxiv-cs-cv
1 Jun 2026
Model Releases

The Regularizing Power of Language-Training Deepfake Detectors

DGX agent

arXiv:2605.31192v1 Announce Type: new Abstract: Recently, thanks to the advent of Multimodal-LLMs, deepfake detectors are striving not only to be generalizable but also interpretable. We propose that

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces

DGX agent

arXiv:2602.07864v2 Announce Type: replace Abstract: Spatial intelligence is crucial for vision--language models (VLMs), yet many scene-centric benchmarks evaluate unconstrained environments where a si

model-releasesarxiv-cs-cv
1 Jun 2026
Research

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens

DGX agent

arXiv:2605.31294v1 Announce Type: new Abstract: Recent advances in Audio-LLMs like GPT-4o have ushered in an era of conversational interaction with language models. Conversational avatars however, sti

researcharxiv-cs-cv
1 Jun 2026
Local Ai

Topologically Consistent Multi-view 3D Head Reconstruction via Coarse-Guided Layered Surface Sampling

DGX agent

arXiv:2605.31283v1 Announce Type: new Abstract: We present SHELLS (Semantic Head Estimation via Layered Local Sampling), an efficient feed-forward framework for 3D head reconstruction in dense semanti

local-aiarxiv-cs-cv
1 Jun 2026
Research

Triangle Splatting SLAM

DGX agent

arXiv:2605.31419v1 Announce Type: new Abstract: We present a dense RGB-D SLAM system using differentiable triangles as the 3D map representation. While 3D Gaussian Splatting has emerged as the leading

researcharxiv-cs-cv
1 Jun 2026
Research

TTE-CAM: Self-Explainable Class Activation Maps for Pretrained Black-Box CNNs

DGX agent

arXiv:2603.26885v2 Announce Type: replace Abstract: Convolutional neural networks (CNNs) achieve state-of-the-art performance in medical image analysis yet remain opaque, limiting adoption in high-sta

researcharxiv-cs-cv
1 Jun 2026
Safety

Unfolding Generative Flows with Koopman Operators: Trajectory-Preserving Linearization

DGX agent

arXiv:2506.22304v3 Announce Type: replace-cross Abstract: Continuous Normalizing Flows (CNFs) enable elegant generative modeling but remain bottlenecked by their iterative nature requiring costly samp

safetyarxiv-cs-cv
1 Jun 2026
Research

UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis

DGX agent

arXiv:2510.15710v3 Announce Type: replace Abstract: Medical workflows routinely combine reading images with producing visual and textual outputs, making both image understanding and generation central

researcharxiv-cs-cv
1 Jun 2026
Safety

Unsupervised Defect Detection for Surgical Instruments

DGX agent

arXiv:2509.21561v2 Announce Type: replace Abstract: Ensuring the safety of surgical instruments requires reliable detection of visual defects. However, manual inspection is prone to error, and existin

safetyarxiv-cs-cv
1 Jun 2026
Research

VAD-GS: Visibility-Aware Densification for 3D Gaussian Splatting in Dynamic Urban Scenes

DGX agent

arXiv:2510.09364v2 Announce Type: replace Abstract: 3D Gaussian splatting (3DGS) has demonstrated impressive performance in synthesizing high-fidelity novel views. Nonetheless, its effectiveness criti

researcharxiv-cs-cv
1 Jun 2026
← Previous
1…133134135136137…263
Next →