AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Safety

LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models

DGX agent

arXiv:2608.05706v1 Announce Type: new Abstract: World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied int

safetyarxiv-cs-cv
7 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

DGX agent

arXiv:2608.06060v1 Announce Type: new Abstract: Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-

researcharxiv-cs-cv
7 Aug 2026
Safety

Learning visual representations for compositional analysis of artworks and photographs

DGX agent

arXiv:2608.06142v1 Announce Type: new Abstract: Composition, the deliberate arrangement of visual elements, is central to how meaning, emotion, and aesthetic quality are conveyed in artwork, yet it re

safetyarxiv-cs-cv
7 Aug 2026
Applications

LiteKD-Net: Lightweight Knowledge-Distilled Network for Mobile Image Denoising

DGX agent

arXiv:2608.05739v1 Announce Type: new Abstract: Mobile image denoising requires both good restoration quality and low computational cost. In addition, it's annoying to collect large-scale LQ-GT clean

applicationsarxiv-cs-cv
7 Aug 2026
Model Releases

LoDA: A Level of Detection Aware Method and a Multimodal Sensing Benchmark for Object Level Change Detection

DGX agent

arXiv:2608.05356v1 Announce Type: new Abstract: High-definition 3D LiDAR maps are important for autonomous driving and smart-city services, which require reliable detection of object-level changes in

model-releasesarxiv-cs-cv
7 Aug 2026
Model Releases

MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding

DGX agent

arXiv:2607.16284v2 Announce Type: replace Abstract: Micro-Actions (MAs) are subtle and spontaneous human behaviors that provide important non-verbal cues in social interaction and affective communicat

model-releasesarxiv-cs-cv
7 Aug 2026
Research

Mapping Armenian Paris: Extracting and Geocoding Commercial Advertisements from the 20th-Century Diaspora Press

DGX agent

arXiv:2608.05911v1 Announce Type: new Abstract: This paper presents an end-to-end, IIIF-based pipeline that turns the digitised Armenian press of France into an interactive map of the 20th-century Par

researcharxiv-cs-cv
7 Aug 2026
Safety

MapTCL: Temporal Consistency Learning via Bidirectional Alignment for Vectorized HD Map Construction

DGX agent

arXiv:2608.05209v1 Announce Type: new Abstract: Constructing reliable online HD maps remains challenging in dynamic urban environments due to moving objects and occlusions. While recent works employ f

safetyarxiv-cs-cv
7 Aug 2026
Model Releases

MASS: Multiplayer World Models with Authoritative Shared State

DGX agent

arXiv:2608.06257v1 Announce Type: new Abstract: Current video world models struggle in multiplayer environments because they entangle world state with view-dependent visual latents, leading to redunda

model-releasesarxiv-cs-cv
7 Aug 2026
Model Releases

MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers

DGX agent

arXiv:2608.05878v1 Announce Type: new Abstract: Text-to-image diffusion transformers learn about objects and scenes by learning to generate them, making them strong candidates for training-free zero-s

model-releasesarxiv-cs-cv
7 Aug 2026
Research

MirrorNet: Can Medical Image Anonymization Really Protect Patient Identity?

DGX agent

arXiv:2608.05938v1 Announce Type: new Abstract: Medical images are routinely de-identified---names, dates, and other metadata removed---and then shared for research, teaching, and public benchmarks un

researcharxiv-cs-cv
7 Aug 2026
Research

MOSAIK: Multi-Patch Content-Aware Spatial Allocation of Image Tokens for Efficient Generation

DGX agent

arXiv:2608.05450v1 Announce Type: new Abstract: Pixel-space diffusion models avoid the reconstruction ceiling of latent diffusion models by generating directly in image space. However, their substanti

researcharxiv-cs-cv
7 Aug 2026
Model Releases

Multi-Representation Geometric Hierarchy Fusion: An Implicit-Submap Driven Framework for Resilient 3D Place Recognition

DGX agent

arXiv:2506.14243v4 Announce Type: replace Abstract: LiDAR-based place recognition is critical for long-term autonomous driving without GPS. Existing handcrafted feature methods face dual limitations.

model-releasesarxiv-cs-cv
7 Aug 2026
Research

Multi-Year Geospatial Reasoning using Interannually-Consistent Historical Predictions as a Free Input Modality

DGX agent

arXiv:2608.05979v1 Announce Type: new Abstract: Machine learning, and deep networks in particular, are increasingly used to derive higher-level Earth observation (EO) products such as annual land-cove

researcharxiv-cs-cv
7 Aug 2026
Research

NeuroAdaptTrainer: A Fiji/ImageJ Plugin for YOLO-Based Neuron Segmentation, InteractiveCorrection and Transfer Learning

DGX agent

arXiv:2608.05226v1 Announce Type: new Abstract: Neuron counting and segmentation in microscopy images of neuronal cultures is a routine and time-consuming task in neuroscience research, traditionally

researcharxiv-cs-cv
7 Aug 2026
Applications

nnMIL: A generalizable multiple instance learning framework for computational pathology

DGX agent

arXiv:2511.14907v2 Announce Type: replace Abstract: Computational pathology holds substantial promise for improving diagnosis and guiding treatment decisions. Recent pathology foundation models enable

applicationsarxiv-cs-cv
7 Aug 2026
Model Releases

OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction

DGX agent

arXiv:2608.05539v1 Announce Type: new Abstract: Recent vision-language models (VLMs) can generate executable CAD programs from images, but existing methods mainly target coarse, general-purpose 3D obj

model-releasesarxiv-cs-cv
7 Aug 2026
Research

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding

DGX agent

arXiv:2608.05707v1 Announce Type: new Abstract: Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Sinc

researcharxiv-cs-cv
7 Aug 2026
Safety

Ordered Diffusion for 3D Human Registration

DGX agent

arXiv:2608.05804v1 Announce Type: new Abstract: 3D human registration has historically been treated as a regression task, assuming a unique ground-truth alignment exists between the template and an in

safetyarxiv-cs-cv
7 Aug 2026
Research

OTLesMix: Wasserstein Barycenter and Optimal Transport Map for Synthetic Lesion Generation with Diverse Shapes and Locations

DGX agent

arXiv:2608.06264v1 Announce Type: new Abstract: The development of deep learning over the past decade has revolutionized medical imaging segmentation, allowing the extraction of precise descriptors fr

researcharxiv-cs-cv
7 Aug 2026
Safety

Overcoming Attention Drift: Homogeneity-Heterogeneity Guided Feature Aggregation for Low-Light Remote Sensing Image Enhancement

DGX agent

arXiv:2608.05843v1 Announce Type: new Abstract: Restoring high-fidelity remote sensing imagery from extreme low-light degradation is indispensable for reliable Earth observation and downstream machine

safetyarxiv-cs-cv
7 Aug 2026
Research

PaCoNet: Deep Data Extraction for Parallel Coordinates

DGX agent

arXiv:2608.06030v1 Announce Type: new Abstract: Extracting data from visualizations has long challenged computer vision, with current research focused on bar, line, and pie charts, among other low-dim

researcharxiv-cs-cv
7 Aug 2026
Model Releases

Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection

DGX agent

arXiv:2604.04444v2 Announce Type: replace Abstract: Open-vocabulary object detection (OVOD) enables models to detect any object category, including unseen ones. Benefiting from large-scale pre-trainin

model-releasesarxiv-cs-cv
7 Aug 2026
Research

Patient Pose Assessment Using a CT-Based Framework for Synthetic Data Generation

DGX agent

arXiv:2608.06126v1 Announce Type: new Abstract: An adequate diagnostic quality of radiographs is essential for reliable diagnoses and treatment planning. The patient's pose during radiography is one o

researcharxiv-cs-cv
7 Aug 2026
Safety

PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models

DGX agent

arXiv:2608.05720v1 Announce Type: new Abstract: We propose PhyLatent, a dynamics-relevant training objective for JointEmbedding Predictive Architecture (JEPA) world models. Our key observation is that

safetyarxiv-cs-cv
7 Aug 2026
Applications

Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation

DGX agent

arXiv:2608.05341v1 Announce Type: new Abstract: Vision-Language Models (VLMs) for radiology report generation are typically trained on retrospective clinical reports, which suffer from omission noise:

applicationsarxiv-cs-cv
7 Aug 2026
Local Ai

Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured Environments

DGX agent

arXiv:2608.06170v1 Announce Type: cross Abstract: Hierarchical 3D scene graphs are a promising representation for high-level spatial reasoning in autonomous mobile platforms. However, existing extract

local-aiarxiv-cs-cv
7 Aug 2026
Research

PromptForSegCXR: Prompt-Driven Multi-Organ and Multi-Disease Segmentation in Chest X-rays using a Multi-stage Fusion Mechanism

DGX agent

arXiv:2507.00673v2 Announce Type: replace-cross Abstract: Image segmentation is central to automated medical image analysis, enabling precise identification of anatomical structures and pathological r

researcharxiv-cs-cv
7 Aug 2026
Research

Respect Your Zero-Shot Uncertainty: Conservative Calibration for Test-Time-Adapted Vision-Language Models

DGX agent

arXiv:2608.05945v1 Announce Type: new Abstract: Test-time adaptation (TTA) can improve the recognition accuracy of vision-language models under distribution shift, but often degrades calibration, maki

researcharxiv-cs-cv
7 Aug 2026
Tutorials

Reversible Unlearnable Examples: Towards the Copyright Protection in Deep Learning Era

DGX agent

arXiv:2608.06211v1 Announce Type: cross Abstract: Significant advancements in deep learning have been made possible by the utilization of large datasets, underscoring the critical importance of copyri

tutorialsarxiv-cs-cv
7 Aug 2026
Safety

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

DGX agent

arXiv:2608.05903v1 Announce Type: new Abstract: Mainstream World-Action Models (WAMs) adapt pretrained video generation models (VGMs) for robot control, transferring their learned dynamics prior for a

safetyarxiv-cs-cv
7 Aug 2026
Safety

Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training

DGX agent

arXiv:2608.06125v1 Announce Type: new Abstract: Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human prefer

safetyarxiv-cs-cv
7 Aug 2026
Safety

SCI-CLIP: Segment-Centric Inference with Reference Memory for Training-Free Open-Vocabulary Segmentation

DGX agent

arXiv:2608.05627v1 Announce Type: new Abstract: Training-free open-vocabulary segmentation remains limited by a missing inference abstraction. Frozen vision-language features are produced at patch lev

safetyarxiv-cs-cv
7 Aug 2026
Research

SciQNet: Two-Stage Multimodal Adaptation for Scientific Image Quality Assessment

DGX agent

arXiv:2608.05691v1 Announce Type: new Abstract: Scientific images are essential for communicating experimental observations, quantitative evidence and conceptual knowledge. Unlike natural images, thei

researcharxiv-cs-cv
7 Aug 2026
Research

Shape-Aware Oriented Bounding Box (OBB) to Horizontal Bounding Box (HBB) Conversion

DGX agent

arXiv:2608.05858v1 Announce Type: new Abstract: Accurate object detection in aerial and satellite imagery is dependent upon the bounding box representation. This is especially true for spatially orien

researcharxiv-cs-cv
7 Aug 2026
Applications

Sparse Mixture-of-Experts for Non-Uniform Noise Reduction in MRI Images

DGX agent

arXiv:2501.14198v3 Announce Type: replace-cross Abstract: Magnetic Resonance Imaging (MRI) is an essential diagnostic tool in clinical settings, but its utility is often hindered by noise artifacts in

applicationsarxiv-cs-cv
7 Aug 2026
Tutorials

SR-JEPA: Learning Predictive Latent State in 3D Scenes

DGX agent

arXiv:2608.05774v1 Announce Type: new Abstract: Joint-embedding predictive architectures learn by predicting latent representations of missing observations, yet many masked JEPAs are evaluated primari

tutorialsarxiv-cs-cv
7 Aug 2026
Research

STAIL: Semantic Text-Anchored Incremental Learning for Medical Imaging via Large Language Models

DGX agent

arXiv:2608.05808v1 Announce Type: new Abstract: Deep learning models applied to medical image analysis suffer from severe catastrophic forgetting when continually adapting to new clinical tasks in dyn

researcharxiv-cs-cv
7 Aug 2026
Model Releases

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

DGX agent

arXiv:2608.05703v1 Announce Type: new Abstract: Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-s

model-releasesarxiv-cs-cv
7 Aug 2026
Research

StyleComposer: Training-Free Multi-Reference Style Composition

DGX agent

arXiv:2608.05213v1 Announce Type: new Abstract: The style of a painting is not monolithic: color, texture, and structure may come from different sources. Existing reference-guided methods transfer the

researcharxiv-cs-cv
7 Aug 2026
Research

Support Operation Factorization: Compositional Readout of Frozen Vision Encoders under Controlled Interventions

DGX agent

arXiv:2608.06174v1 Announce Type: new Abstract: Compositional analysis of frozen vision encoders should determine both what changed and where it changed. Standard factor probes score these axes separa

researcharxiv-cs-cv
7 Aug 2026
Model Releases

TAU-Bench: From Anomaly Instance Tracking to Fine-Grained Video Anomaly Understanding

DGX agent

arXiv:2608.05699v1 Announce Type: new Abstract: Humans understand anomalous events through a coherent perceptual process in which they identify the focal instance, follow its behavior as the event unf

model-releasesarxiv-cs-cv
7 Aug 2026
Model Releases

Text-Guided Refinement of Multi-sequence Glioma Subregion Segmentation with a Vision-Language Foundation Model

DGX agent

arXiv:2608.05389v1 Announce Type: new Abstract: Background: Accurate glioma subregion delineation is important for radiotherapy planning and longitudinal monitoring, but manual contour correction is t

model-releasesarxiv-cs-cv
7 Aug 2026
Model Releases

The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents

DGX agent

arXiv:2608.06065v1 Announce Type: new Abstract: GUI agents are commonly trained offline from successful interaction trajectories. Standard training decomposes each trajectory into prefix-action pairs:

model-releasesarxiv-cs-cv
7 Aug 2026
Local Ai

TLNM: Externally Validated Tooth Detection, Numbering and Segmentation from Smartphone Photographs Using Mask R-CNN

DGX agent

arXiv:2608.06275v1 Announce Type: new Abstract: Oral health issues affect billions globally, but the cost and limited access to professional dental care hinder preventive oral healthcare. Research rel

local-aiarxiv-cs-cv
7 Aug 2026
Research

To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation

DGX agent

arXiv:2608.05879v1 Announce Type: new Abstract: Text-driven 3D generation has advanced rapidly in creating large-scale outdoor environments and detailed indoor scenes, but these domains are usually sy

researcharxiv-cs-cv
7 Aug 2026
Research

Topology-Aware Neighborhood Learning for Source-Free Cross-Scene Hyperspectral Image Classification

DGX agent

arXiv:2608.05964v1 Announce Type: new Abstract: Domain adaptation has advanced cross-scene hyperspectral image classification, significantly improving discriminative capability in complex scenarios. H

researcharxiv-cs-cv
7 Aug 2026
Model Releases

Tree-NET: Enhancing 2D Medical Image Segmentation Through Efficient Low-Level Feature Training

DGX agent

arXiv:2501.02140v2 Announce Type: replace-cross Abstract: This paper introduces Tree-NET, a novel framework for medical image segmentation that leverages bottleneck supervision to enhance both segment

model-releasesarxiv-cs-cv
7 Aug 2026
← Previous
1…1011121314…259
Next →