AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Model Releases

SlimVLM: Sensitivity-aware Dynamic Structured Pruning with Adaptive Visual Token Selection for Efficient Vision-Language Models

DGX agent

arXiv:2608.03580v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have demonstrated remarkable performance in processing and understanding both text and images, their large parameter

model-releasesarxiv-cs-cv
5 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference

DGX agent

arXiv:2608.03335v1 Announce Type: new Abstract: Video diffusion transformers (vDiTs) generate high quality but pay quadratic self-attention cost, making inference prohibitive at video-token scales. Th

safetyarxiv-cs-cv
5 Aug 2026
Research

SpreadMark: Robust Image Watermarking via Spread-Spectrum Embedding

DGX agent

arXiv:2608.03165v1 Announce Type: cross Abstract: Invisible image watermarks are increasingly used for deepfake detection and provenance tracking, where they must survive not only incidental distortio

researcharxiv-cs-cv
5 Aug 2026
Research

SRAP: SVD-Refined Adversarial Perturbations for Imperceptible Face-Swap Defense

DGX agent

arXiv:2608.03395v1 Announce Type: new Abstract: Deepfake technologies pose increasing threats to facial privacy and identity security, motivating proactive defenses that protect facial images before m

researcharxiv-cs-cv
5 Aug 2026
Model Releases

StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision

DGX agent

arXiv:2603.29368v2 Announce Type: replace Abstract: Driven by the advancement of 3D devices, stereo vision tasks including stereo matching and stereo conversion have emerged as a critical research fro

model-releasesarxiv-cs-cv
5 Aug 2026
Applications

Stop Replacing Noise with Noise: Two-Source Reliability Assessment for Label Correction and Sample Reweighting in Label-Noise Learning

DGX agent

arXiv:2608.03432v1 Announce Type: cross Abstract: Refurbishment-based noisy-label learning mixes an observed label with a model-derived pseudo target, typically using one sample-wise cleanliness score

applicationsarxiv-cs-cv
5 Aug 2026
Safety

StreamDAM: Presence-Aware Memory for Real-Time Streaming Video Object Segmentation

DGX agent

arXiv:2608.03912v1 Announce Type: new Abstract: Quality-tier video object segmentation (VOS) trackers such as DAM4SAM top accuracy leaderboards, but they are measured offline, one frame at a time with

safetyarxiv-cs-cv
5 Aug 2026
Tutorials

Style-Aware Gloss Control for Generative Non-Photorealistic Rendering

DGX agent

arXiv:2602.16611v3 Announce Type: replace-cross Abstract: Humans can infer material characteristics of objects from their visual appearance, and this ability extends to artistic depictions, where simi

tutorialsarxiv-cs-cv
5 Aug 2026
Research

Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation

DGX agent

arXiv:2608.03158v1 Announce Type: new Abstract: Daily activities require humans to coordinate whole-body motion with the motion of surrounding objects. Despite recent progress in human-object interact

researcharxiv-cs-cv
5 Aug 2026
Model Releases

SUV: Future Scene Understanding as Video Generation for End-to-End Driving

DGX agent

arXiv:2608.03084v1 Announce Type: new Abstract: End-to-end driving requires a coherent understanding of future scenes, yet existing methods model these scenes using task-specific heads and output form

model-releasesarxiv-cs-cv
5 Aug 2026
Safety

T2VAttack: Adversarial Attack on Text-to-Video Diffusion Models

DGX agent

arXiv:2512.23953v2 Announce Type: replace Abstract: The rapid evolution of Text-to-Video (T2V) diffusion models has driven remarkable advancements in generating high-quality, temporally coherent video

safetyarxiv-cs-cv
5 Aug 2026
Research

TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models

DGX agent

arXiv:2608.03057v1 Announce Type: new Abstract: Static quantization assigns one weight precision to every denoising step. To preserve quality, that precision must accommodate the most quantization-sen

researcharxiv-cs-cv
5 Aug 2026
Local Ai

TDVR: Joint Text Disambiguation and Viewpoint Reasoning for Zero-Shot 3D Visual Grounding

DGX agent

arXiv:2608.03763v1 Announce Type: new Abstract: Zero-shot 3D visual grounding aims to localize specific objects based on textual descriptions and 3D visual input. However, the effectiveness of existin

local-aiarxiv-cs-cv
5 Aug 2026
Safety

Test Time Adaptation Methods for Point Cloud Registration in Laparoscopic Surgery

DGX agent

arXiv:2608.02883v1 Announce Type: new Abstract: 3D point cloud registration in laparoscopic surgery estimates the transformation between an intraoperative organ reconstructed from video and its preope

safetyarxiv-cs-cv
5 Aug 2026
Model Releases

Test-Time Augmentation for Tabular-to-Image Classifiers under Distribution Shifts

DGX agent

arXiv:2608.03557v1 Announce Type: new Abstract: Tabular-to-image methods that convert tabular data into visual representations have emerged as a novel paradigm for leveraging the high performance of d

model-releasesarxiv-cs-cv
5 Aug 2026
Research

Toward Visual Grounding: A Survey

DGX agent

arXiv:2412.20206v4 Announce Type: replace Abstract: Visual Grounding, also known as Referring Expression Comprehension and Phrase Grounding, aims to ground the specific region(s) within the image(s) b

researcharxiv-cs-cv
5 Aug 2026
Research

Towards Reliable and Reproducible Fetal Brain Biometry: A Deep Learning Approach Using MRI

DGX agent

arXiv:2608.03724v1 Announce Type: new Abstract: Fetal brain biometry is essential for quantitative assessment of brain development, supporting gestational age estimation, developmental monitoring, and

researcharxiv-cs-cv
5 Aug 2026
Research

Two-Way Garment Transfer: Unified Diffusion Framework for Dressing and Undressing Synthesis

DGX agent

arXiv:2508.04551v2 Announce Type: replace Abstract: While recent advances in virtual try-on (VTON) have achieved realistic garment transfer to human subjects, its inverse task, virtual try-off (VTOFF)

researcharxiv-cs-cv
5 Aug 2026
Research

UniEvo-RS: Omni-Prompt Unified Remote Sensing Segmentation with Representative Exemplar-Driven Prototype Evolution

DGX agent

arXiv:2608.03911v1 Announce Type: new Abstract: Prompt-driven vision-language models (VLMs) hold immense promise for accelerating dense remote sensing (RS) annotation, but static models suffer from se

researcharxiv-cs-cv
5 Aug 2026
Model Releases

UniWorld-Design: From Pixel Generation to Layer-Native Design

DGX agent

arXiv:2608.03971v1 Announce Type: new Abstract: We introduce UniWorld-Design, a framework that redefines image generation from flat pixel synthesis to structured visual composition, with semantic RGBA

model-releasesarxiv-cs-cv
5 Aug 2026
Research

Unsupervised Adversarial Domain Adaptation for Uterine layer Segmentation: From Labeled Cine to Unlabeled Dynamic EPI MRI

DGX agent

arXiv:2608.03762v1 Announce Type: cross Abstract: Uterine peristalsis is a key physiological phenomenon responsible for various functions across the menstrual cycle, intimately linked to uterine wall

researcharxiv-cs-cv
5 Aug 2026
Agents

VLC Fusion: Vision-Language Conditioned Sensor Fusion for Robust Object Detection

DGX agent

arXiv:2505.12715v2 Announce Type: replace Abstract: Although fusing multiple sensor modalities can enhance object detection performance, existing fusion approaches often overlook subtle variations in

agentsarxiv-cs-cv
5 Aug 2026
Safety

What is the Right Embedding Space for Contrastive Learning in REC?

DGX agent

arXiv:2505.22850v2 Announce Type: replace Abstract: Referring Expression Counting (REC) requires distinguishing visually similar objects described by fine-grained text cues. Existing methods tackle th

safetyarxiv-cs-cv
5 Aug 2026
Model Releases

When Classes Evolve: A Benchmark and Framework for Stage-Aware Class-Incremental Learning

DGX agent

arXiv:2602.00573v2 Announce Type: replace-cross Abstract: Class-Incremental Learning (CIL) aims to sequentially learn new classes while mitigating catastrophic forgetting of previously learned knowled

model-releasesarxiv-cs-cv
5 Aug 2026
Model Releases

When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware

DGX agent

arXiv:2608.03649v1 Announce Type: new Abstract: Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead,

model-releasesarxiv-cs-cv
5 Aug 2026
Applications

XiDepth: a Lightweight and Efficient Network for Self-supervised Monocular Depth Estimation

DGX agent

arXiv:2608.03666v1 Announce Type: new Abstract: Self-supervised monocular depth estimation has emerged as an appealing solution to design lightweight and effective models for deployment on computation

applicationsarxiv-cs-cv
5 Aug 2026
Research

3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering

DGX agent

arXiv:2608.01185v1 Announce Type: new Abstract: Recent 3D vision-language models (3D VLMs) construct geometry aware tokens by projecting 2D visual features into world coordinates, enabling spatial rea

researcharxiv-cs-cv
4 Aug 2026
Model Releases

A Benchmark Dataset for MLLM-Generated Image Detection: GPT Image2 & Nano Banana2

DGX agent

arXiv:2608.01258v1 Announce Type: new Abstract: The realism of images generated by multimodal large language models (MLLMs), such as GPT Image2 and Nano Banana2, has improved rapidly in recent years.

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

A General-Purpose VLM Can Teach an Astronomy Foundation Model to Better Recognize Galaxy Morphology

DGX agent

arXiv:2608.02300v1 Announce Type: new Abstract: Existing astronomy foundation models provide strong galaxy representations, but adapting them to new survey conditions and survey-specific morphology re

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis Generation

DGX agent

arXiv:2608.02505v1 Announce Type: cross Abstract: Can scientific abduction occur without continuous sensorimotor embodiment? Recent arguments in AI and philosophy of science hold that genuine hypothes

model-releasesarxiv-cs-cv
4 Aug 2026
Research

ABRA: Teleporting Fine-Tuned Knowledge Across Domains for Open-Vocabulary Object Detection

DGX agent

arXiv:2603.12409v2 Announce Type: replace Abstract: Although recent Open-Vocabulary Object Detection architectures, such as Grounding DINO, demonstrate strong zero-shot capabilities, their performance

researcharxiv-cs-cv
4 Aug 2026
Model Releases

Accuracy Does Not Guarantee Human-Likeness: Cross-Domain Human-Centered Benchmark in Monocular Depth Estimation

DGX agent

arXiv:2512.08163v2 Announce Type: replace Abstract: Deep neural networks (DNNs) are increasingly used as functional models of human vision, yet standard monocular depth estimation (MDE) benchmarks lar

model-releasesarxiv-cs-cv
4 Aug 2026
Applications

Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery

DGX agent

arXiv:2608.02471v1 Announce Type: new Abstract: Computational attention models could help surgeons manage the visual demands of laparoscopy, but they require dense spatial labels that are difficult to

applicationsarxiv-cs-cv
4 Aug 2026
Tutorials

AdaForensics: Learning A Characteristic-aware Adaptive Deepfake Detector

DGX agent

arXiv:2608.02160v1 Announce Type: new Abstract: In this paper, we propose a characteristic-aware adaptive network named AdaForensics for deepfake detection. Most existing methods learn a fixed network

tutorialsarxiv-cs-cv
4 Aug 2026
Safety

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning

DGX agent

arXiv:2608.01980v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning can improve performance on difficult video questions but often wastes decoding tokens on simple ones. We study whether

safetyarxiv-cs-cv
4 Aug 2026
Model Releases

AeroLLE: Constrained Pseudo-Supervision for Nighttime Aerial Image Enhancement with the AeroNight-1.5K Benchmark

DGX agent

arXiv:2608.00702v1 Announce Type: new Abstract: Nighttime aerial image enhancement is challenged by spatially nonuniform exposure, mixed illumination, and weak structural evidence, while registered no

model-releasesarxiv-cs-cv
4 Aug 2026
Agents

AIMold: An Autonomous AI-based Pipeline for Complex Mold Design

DGX agent

arXiv:2608.00800v1 Announce Type: new Abstract: Injection molding is the cornerstone of mass-producing plastic components. While current algorithms can automate mold design for basic geometries using

agentsarxiv-cs-cv
4 Aug 2026
Local Ai

AirSplat: Alignment and Rating for Robust Feed-Forward 3D Gaussian Splatting

DGX agent

arXiv:2603.25129v2 Announce Type: replace Abstract: While 3D Vision Foundation Models (3DVFMs) have demonstrated remarkable zero-shot capabilities in visual geometry estimation, their direct applicati

local-aiarxiv-cs-cv
4 Aug 2026
Model Releases

An Accessible Solution for Deformable Image Registration Compared with Learning-Based Approaches

DGX agent

arXiv:2608.02248v1 Announce Type: cross Abstract: Deformable image registration (DIR) is a core problem in medical image analysis; but, unlike labeling decision problems such as classification and seg

model-releasesarxiv-cs-cv
4 Aug 2026
Applications

ArcAD: Anomaly-Rectified Calibration for Cold-Start Supervised Anomaly Detection

DGX agent

arXiv:2607.02252v2 Announce Type: replace Abstract: The deployment of Industrial Anomaly Detection (IAD) in real-world manufacturing frequently encounters a challenging cold-start bottleneck, in which

applicationsarxiv-cs-cv
4 Aug 2026
Research

ARM: Detector-Agnostic Changepoint Attribution with Finite-Sample Error Control

DGX agent

arXiv:2608.01691v1 Announce Type: cross Abstract: Detecting a change in a multivariate series answers only the first of two questions; the operational question is which coordinates changed. Existing a

researcharxiv-cs-cv
4 Aug 2026
Research

Artificial Intelligence for the Characterization of Particles and Fibers by Optical Microscopy

DGX agent

arXiv:2608.00361v1 Announce Type: new Abstract: Optical microscopy of particle and fiber dispersions involves interpreting subtle visual cues influenced by specimen morphology, chemical composition, m

researcharxiv-cs-cv
4 Aug 2026
Model Releases

Asleep at the Wheel: JEPA's Limitations in Evaluating Novel Driving Data

DGX agent

arXiv:2608.01336v1 Announce Type: new Abstract: Modern autonomous-driving fleets record far more video than human reviewers can inspect. This motivates the need for an automatic clip triage mechanism,

model-releasesarxiv-cs-cv
4 Aug 2026
Applications

Assessing the Benefits of Combining Advanced Deep Learning Techniques for Post-Disaster Building Damage Assessment from UAV Imagery

DGX agent

arXiv:2608.01906v1 Announce Type: new Abstract: Rapid and accurate post-disaster building damage assessment is essential, yet remains a challenging task. Unmanned Aerial Vehicle (UAV) imagery offers a

applicationsarxiv-cs-cv
4 Aug 2026
Safety

ASTRA: Asynchronous Spatio-Temporal Reconstruction via Trajectory Alignment

DGX agent

arXiv:2608.02006v1 Announce Type: new Abstract: Dynamic 3D scene reconstruction has achieved remarkable success under the assumption of strictly synchronized multi-camera inputs. However, in real-worl

safetyarxiv-cs-cv
4 Aug 2026
Tutorials

Astrolabe: Spherical-Map Guidance Across Diffusion Pipelines for Full-Body Capture from Unconstrained Images

DGX agent

arXiv:2608.01276v1 Announce Type: new Abstract: Full-body capture from unconstrained photographs requires global correspondence across arbitrary views, poses, crops, and occlusions. Yet pose, geometry

tutorialsarxiv-cs-cv
4 Aug 2026
Research

Attention-Steered Vision-Language Models for Sign Language Translation

DGX agent

arXiv:2608.00235v1 Announce Type: new Abstract: Vision-language models (VLMs) have emerged as a powerful framework for multimodal video understanding. However, they remain limited in the sign language

researcharxiv-cs-cv
4 Aug 2026
Model Releases

Automatic LV Localization and Short-Axis Plane Estimation from Arbitrary CMR Slice

DGX agent

arXiv:2608.00145v1 Announce Type: cross Abstract: Accurate estimation of left ventricular (LV) orientation is essential for cardiac magnetic resonance (CMR) imaging and downstream analysis. Existing m

model-releasesarxiv-cs-cv
4 Aug 2026
← Previous
1…1617181920…259
Next →