AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlog
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Model Releases

Staying True to the Origin: Continuous Image Stylization with Smooth Transitions

DGX agent

arXiv:2608.08125v1 Announce Type: new Abstract: Recent advances in generative models have achieved remarkable performance in text- and image-conditioned editing. However, preserving the content of a g

model-releasesarxiv-cs-cv
11 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

SUMI: Scalable Unified Model for 3D Point Cloud Inference

DGX agent

arXiv:2608.08115v1 Announce Type: new Abstract: Point cloud completion commonly follows a coarse-to-fine paradigm, where a low-density coarse shape is first predicted and then upsampled to the target

researcharxiv-cs-cv
11 Aug 2026
Model Releases

SwissCrop25: A National Multi-Year Benchmark for Operational Crop Mapping

DGX agent

arXiv:2608.09497v1 Announce Type: new Abstract: Operational crop mapping requires models that generalise across years, resolve fine-grained crop taxonomies, and distinguish cropland from surrounding l

model-releasesarxiv-cs-cv
11 Aug 2026
Safety

SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model

DGX agent

arXiv:2608.07948v1 Announce Type: new Abstract: VAR has gained widespread popularity due to its next-scale prediction paradigm. However, it faces substantial performance bottlenecks when handling comp

safetyarxiv-cs-cv
11 Aug 2026
Research

Task-Adaptive 3D Cross-Field MRI Translation via Field-Conditioned Content-Style Pretraining

DGX agent

arXiv:2608.09264v1 Announce Type: new Abstract: Magnetic field strength is a major source of domain shift in magnetic resonance imaging (MRI), affecting signal-to-noise ratio, tissue contrast, spatial

researcharxiv-cs-cv
11 Aug 2026
Research

TeaMatch: Teachable Cross-Modal Representation Learning for 2D-3D Matching

DGX agent

arXiv:2608.09590v1 Announce Type: new Abstract: Learning reliable correspondences between images and point clouds is fundamental for 2D-3D matching. Despite recent progress in detection-free methods,

researcharxiv-cs-cv
11 Aug 2026
Model Releases

Test-Time Prototype Adaptation for Open-Vocabulary Semantic Segmentation

DGX agent

arXiv:2608.08290v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation (OVSS) repurposes a pretrained CLIP encoder for dense prediction without additional labeled supervision. Existing

model-releasesarxiv-cs-cv
11 Aug 2026
Agents

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning

DGX agent

arXiv:2608.09682v1 Announce Type: new Abstract: Tool-augmented vision-language models increasingly 'think with images': they call crop, zoom, or code tools and reason over the returned pixels. However

agentsarxiv-cs-cv
11 Aug 2026
Research

Tokenizer Generator Coupling in Medical Image Generation

DGX agent

arXiv:2608.07713v1 Announce Type: new Abstract: Latent medical image generators usually treat the tokenizer as fixed preprocessing. We test whether this separation is valid in a controlled ChestMNIST

researcharxiv-cs-cv
11 Aug 2026
Model Releases

Topology-Aware Global-Local Mamba Networks for Palm Vein Biometrics

DGX agent

arXiv:2608.08951v1 Announce Type: new Abstract: Palm-vein recognition is a fine-grained biometric task in which both local vascular texture and the global layout of the vessel tree carry discriminativ

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

Toward Mask Annotation-Free Surgical Instrument Segmentation from Endoscopic Images Using Text-Prompted Segment Anything Model 3 (SAM3)

DGX agent

arXiv:2608.08844v1 Announce Type: new Abstract: Surgical instrument segmentation is a fundamental task for computer-assisted interventions, yet most existing methods rely on pixel-level annotations or

model-releasesarxiv-cs-cv
11 Aug 2026
Applications

Towards Adaptive Super-Resolution and Quality Assessment via Test-Time Adaptation

DGX agent

arXiv:2608.08508v1 Announce Type: new Abstract: This paper presents doctoral research on adaptive video super-resolution and perceptual quality modeling under real-world conditions. Existing video sup

applicationsarxiv-cs-cv
11 Aug 2026
Agents

Towards Collaborative Joint Perception and Prediction: Framework, Baseline Evaluation, and Deployment Perspectives

DGX agent

arXiv:2608.09541v1 Announce Type: new Abstract: Connected Autonomous Vehicles (CAVs) increasingly exploit Vehicle-to-Everything (V2X) communication to exchange multi-source sensor information, enablin

agentsarxiv-cs-cv
11 Aug 2026
Safety

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework

DGX agent

arXiv:2608.09529v1 Announce Type: new Abstract: As an important subfield of cross-modal generation, synthesizing static visual content in the form of images from audio, namely audio-to-image (A2I) gen

safetyarxiv-cs-cv
11 Aug 2026
Research

TriView-YOLO: Early Multi-View Fusion for Ground Penetrating Radar Cavity Detection in Soft, High-Water-Content Soils

DGX agent

arXiv:2608.09522v1 Announce Type: new Abstract: Automated detection of subsurface cavities from Ground Penetrating Radar (GPR) is most difficult in soft, high-water-content ground, where conductive, w

researcharxiv-cs-cv
11 Aug 2026
Research

Tropical Cyclone Forecasting via Latent Rectified Flow using Satellite Imagery and Atmospheric Fields

DGX agent

arXiv:2608.08354v1 Announce Type: new Abstract: Tropical cyclones are growing more destructive in a changing climate, and efficient forecasting of their structure and track has become a necessity. Dee

researcharxiv-cs-cv
11 Aug 2026
Research

Uncertainty-Aware 4D Gaussian Splatting for Monocular Occluded Human Rendering

DGX agent

arXiv:2602.06343v3 Announce Type: replace Abstract: High-fidelity rendering of dynamic humans from monocular videos typically degrades catastrophically under occlusions. Existing solutions incorporate

researcharxiv-cs-cv
11 Aug 2026
Local Ai

UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation

DGX agent

arXiv:2608.09143v1 Announce Type: new Abstract: Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodifi

local-aiarxiv-cs-cv
11 Aug 2026
Tutorials

UniScale: Arbitrary-Scale Industrial Anomaly Generation

DGX agent

arXiv:2608.07864v1 Announce Type: new Abstract: Industrial anomaly inspection faces a major challenge due to the lack of real-world anomaly samples. While generative models are used to create anomaly

tutorialsarxiv-cs-cv
11 Aug 2026
Research

Unsupervised Domain Adaptation for Multitask Image Analysis in Realistic Context with Extreme Label Shift; Application to the CTAO first Large Sized Telescope

DGX agent

arXiv:2608.09630v1 Announce Type: cross Abstract: Unsupervised domain adaptation is a widespread set of methods that leverages the knowledge of a labeled source domain to train a model to perform well

researcharxiv-cs-cv
11 Aug 2026
Model Releases

Unsupervised Point Cloud Registration with Self-Distillation

DGX agent

arXiv:2409.07558v2 Announce Type: replace Abstract: Rigid point cloud registration is a fundamental problem and highly relevant in robotics and autonomous driving. Nowadays deep learning methods can b

model-releasesarxiv-cs-cv
11 Aug 2026
Research

Unveiling the Secret of AdaLN-Zero in Diffusion Transformer

DGX agent

arXiv:2608.09438v1 Announce Type: new Abstract: Diffusion transformer (DiT), a rapidly emerging architecture for image generation, has gained much attention. However, despite ongoing efforts to improv

researcharxiv-cs-cv
11 Aug 2026
Research

UPolarSQ: Polar Representation Learning for Optic Disc and Peripapillary Atrophy Segmentation and Quantification in Fundus Photographs

DGX agent

arXiv:2608.08771v1 Announce Type: new Abstract: Myopia-induced posterior-pole remodeling is frequently accompanied by Optic Disc (OD) deformation and Peripapillary Atrophy (PPA), both of which provide

researcharxiv-cs-cv
11 Aug 2026
Safety

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models

DGX agent

arXiv:2608.08622v1 Announce Type: new Abstract: Large vision-language models (LVLMs) have demonstrated strong performance in open-ended video understanding, yet they remain prone to fluent responses u

safetyarxiv-cs-cv
11 Aug 2026
Safety

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction

DGX agent

arXiv:2608.09448v1 Announce Type: cross Abstract: Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains d

safetyarxiv-cs-cv
11 Aug 2026
Model Releases

VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation

DGX agent

arXiv:2608.09573v1 Announce Type: new Abstract: Natural-language-driven 'vibe coding' enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of thei

model-releasesarxiv-cs-cv
11 Aug 2026
Applications

View-Adaptive Renderer for View-Consistent 2D-to-3D Generation

DGX agent

arXiv:2608.09110v1 Announce Type: new Abstract: Reconstructing 3D shapes from a single image remains a fundamental yet challenging problem in computer vision. Traditional monocular 3D generation pipel

applicationsarxiv-cs-cv
11 Aug 2026
Model Releases

VIGIL: Tackling Hallucination Detection in Image Recontextualization

DGX agent

arXiv:2602.14633v2 Announce Type: replace Abstract: We introduce VIGIL (Visual Inconsistency & Generative In-context Lucidity), a benchmark dataset and framework that provides a fine-grained categoriz

model-releasesarxiv-cs-cv
11 Aug 2026
Research

Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties

DGX agent

arXiv:2608.07726v1 Announce Type: new Abstract: Estimating volumetric mechanical properties, including Young's modulus, Poisson's ratio, and density at each voxel, is intrinsically ambiguous from visi

researcharxiv-cs-cv
11 Aug 2026
Research

VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs

DGX agent

arXiv:2510.16598v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) encounter significant computational and memory bottlenecks from the massive number of visual tokens generat

researcharxiv-cs-cv
11 Aug 2026
Local Ai

Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding

DGX agent

arXiv:2608.08832v1 Announce Type: new Abstract: Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nod

local-aiarxiv-cs-cv
11 Aug 2026
Model Releases

VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling

DGX agent

arXiv:2608.08630v1 Announce Type: new Abstract: Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-atte

model-releasesarxiv-cs-cv
11 Aug 2026
Research

VOICE: A Vision-Omics Foundation Model Integrating Direct and Retrieval-Based Prediction of In-situ Single-Cell Gene Expression

DGX agent

arXiv:2608.08366v1 Announce Type: new Abstract: Spatial transcriptomics can resolve gene expression at single-cell resolution, but it is costly, limited to targeted panels of a few hundred to a few th

researcharxiv-cs-cv
11 Aug 2026
Local Ai

Warp-free Cross-view Geo-localization via Feature-space Consensus Mining

DGX agent

arXiv:2608.09321v1 Announce Type: new Abstract: Cross-view geo-localization is challenging due to drastic viewpoint changes and large appearance discrepancies between street-level and satellite imager

local-aiarxiv-cs-cv
11 Aug 2026
Hardware

What Irregularity Costs: CUDA C++, Rust, and Triton on a Hash-Blocked GPU Workload

DGX agent

arXiv:2608.08287v1 Announce Type: new Abstract: GPU language comparisons are almost always run on tiled dense linear algebra, where every toolchain is good and the differences are small. We implement

hardwarearxiv-cs-cv
11 Aug 2026
Model Releases

When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery

DGX agent

arXiv:2608.08132v1 Announce Type: new Abstract: Reconstruction of 3D objects from a single image is a challenging research problem in computer vision. The key challenge is the lack of critical informa

model-releasesarxiv-cs-cv
11 Aug 2026
Research

Where Is the Bee? Detecting Tiny Pollinators with a Single Collaborative-Head Transformer

DGX agent

arXiv:2608.08580v1 Announce Type: new Abstract: The CVPPA@ECCV 2026 BuzzSpot Challenge asks us to detect bees, bumblebees, hoverflies, and moths in 1920x1080 field keyframes. Its annotations carry 2 d

researcharxiv-cs-cv
11 Aug 2026
Model Releases

Wiener Representation Filtering for VLM Hallucination Suppression

DGX agent

arXiv:2608.08167v1 Announce Type: new Abstract: Vision-language models (VLMs) excel at open-ended captioning and visual QA but often describe objects, attributes, or relations absent from the image, a

model-releasesarxiv-cs-cv
11 Aug 2026
Research

World Simulator: Queer Erotica and the Absurdity of AI Video Models That Promise the World

DGX agent

arXiv:2608.07510v1 Announce Type: cross Abstract: Increasingly, AI video models are marketed as 'world simulators,' suggesting their ability to model infinite realities. Despite such claims, these mod

researcharxiv-cs-cv
11 Aug 2026
Safety

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

DGX agent

arXiv:2608.09730v1 Announce Type: new Abstract: Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicit

safetyarxiv-cs-cv
11 Aug 2026
Local Ai

XClipGS: Exact Half-Space Clipping for Medical Volume Gaussian Splatting

DGX agent

arXiv:2608.07760v1 Announce Type: new Abstract: Gaussian-splatting proxies enable interactive rendering of volumetric medical scans, but a clipping plane exposes anatomy not constrained by external-vi

local-aiarxiv-cs-cv
11 Aug 2026
Research

XEns-CKD: An Explainable Ensemble-Based Approach for Chronic Kidney Disease Stage Detection

DGX agent

arXiv:2608.07561v1 Announce Type: new Abstract: Chronic kidney disease (CKD) is a silent disease. Its progression may not significantly hamper a person's daily routine. Human kidney function can be cl

researcharxiv-cs-cv
11 Aug 2026
Model Releases

XFeat Revisited: Reproducibility and Evaluation of a Lightweight Image Matcher

DGX agent

arXiv:2608.09519v1 Announce Type: new Abstract: We present a reproducibility study of XFeat, a lightweight local feature extractor and matcher designed to identify corresponding points across images e

model-releasesarxiv-cs-cv
11 Aug 2026
Research

You Only Flow Once: Calibrated and Real-Time Radar Pose Estimation with Multi-Hypothesis Normalizing Flows

DGX agent

arXiv:2608.09579v1 Announce Type: new Abstract: Sparse and noisy millimeter-wave radar point cloud observations often correspond to multiple plausible human poses, making deterministic pose estimation

researcharxiv-cs-cv
11 Aug 2026
Model Releases

Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No

DGX agent

arXiv:2608.08315v1 Announce Type: new Abstract: Multimodal LLMs that recognise events reliably still fail to say when they happen. Prompted for timestamps, strong VLMs reach as little as 3.8% R@0.5 on

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

Zero-shot 2D Grounding with Novel Affordance Types

DGX agent

arXiv:2608.08929v1 Announce Type: new Abstract: 2D affordance grounding aims to locate the region of an object that a human can interact with. Existing research focuses on recognizing affordance types

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

Zero-Shot Traffic Accident Detection via a Coarse-to-Fine VLM-Tracking Pipeline

DGX agent

arXiv:2608.08867v1 Announce Type: new Abstract: Traffic surveillance cameras capture accidents continuously, yet converting raw CCTV footage into structured event records that pinpoint when, where, an

model-releasesarxiv-cs-cv
11 Aug 2026
Research

ZOMP: Zeroth-Order Multi-Modal Prompt Tuning for Vision-Language Models

DGX agent

arXiv:2608.08060v1 Announce Type: new Abstract: Fine-tuning vision-language models such as CLIP typically requires backpropagation (BP) through the full model, which is infeasible when only forward-pa

researcharxiv-cs-cv
11 Aug 2026
← Previous
1…678910…259
Next →