AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
7 Aug 2026

Multi-Representation Geometric Hierarchy Fusion: An Implicit-Submap Driven Framework for Resilient 3D Place Recognition

Model ReleasesDGX agent

arXiv:2506.14243v4 Announce Type: replace Abstract: LiDAR-based place recognition is critical for long-term autonomous driving without GPS. Existing handcrafted feature methods face dual limitations.

Multi-Year Geospatial Reasoning using Interannually-Consistent Historical Predictions as a Free Input Modality

ResearchDGX agent

arXiv:2608.05979v1 Announce Type: new Abstract: Machine learning, and deep networks in particular, are increasingly used to derive higher-level Earth observation (EO) products such as annual land-cove

NeuroAdaptTrainer: A Fiji/ImageJ Plugin for YOLO-Based Neuron Segmentation, InteractiveCorrection and Transfer Learning


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research
DGX agent

arXiv:2608.05226v1 Announce Type: new Abstract: Neuron counting and segmentation in microscopy images of neuronal cultures is a routine and time-consuming task in neuroscience research, traditionally

nnMIL: A generalizable multiple instance learning framework for computational pathology

ApplicationsDGX agent

arXiv:2511.14907v2 Announce Type: replace Abstract: Computational pathology holds substantial promise for improving diagnosis and guiding treatment decisions. Recent pathology foundation models enable

OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction

Model ReleasesDGX agent

arXiv:2608.05539v1 Announce Type: new Abstract: Recent vision-language models (VLMs) can generate executable CAD programs from images, but existing methods mainly target coarse, general-purpose 3D obj

One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding

ResearchDGX agent

arXiv:2608.05707v1 Announce Type: new Abstract: Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Sinc

Ordered Diffusion for 3D Human Registration

SafetyDGX agent

arXiv:2608.05804v1 Announce Type: new Abstract: 3D human registration has historically been treated as a regression task, assuming a unique ground-truth alignment exists between the template and an in

OTLesMix: Wasserstein Barycenter and Optimal Transport Map for Synthetic Lesion Generation with Diverse Shapes and Locations

ResearchDGX agent

arXiv:2608.06264v1 Announce Type: new Abstract: The development of deep learning over the past decade has revolutionized medical imaging segmentation, allowing the extraction of precise descriptors fr

Overcoming Attention Drift: Homogeneity-Heterogeneity Guided Feature Aggregation for Low-Light Remote Sensing Image Enhancement

SafetyDGX agent

arXiv:2608.05843v1 Announce Type: new Abstract: Restoring high-fidelity remote sensing imagery from extreme low-light degradation is indispensable for reliable Earth observation and downstream machine

PaCoNet: Deep Data Extraction for Parallel Coordinates

ResearchDGX agent

arXiv:2608.06030v1 Announce Type: new Abstract: Extracting data from visualizations has long challenged computer vision, with current research focused on bar, line, and pie charts, among other low-dim

Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection

Model ReleasesDGX agent

arXiv:2604.04444v2 Announce Type: replace Abstract: Open-vocabulary object detection (OVOD) enables models to detect any object category, including unseen ones. Benefiting from large-scale pre-trainin

Patient Pose Assessment Using a CT-Based Framework for Synthetic Data Generation

ResearchDGX agent

arXiv:2608.06126v1 Announce Type: new Abstract: An adequate diagnostic quality of radiographs is essential for reliable diagnoses and treatment planning. The patient's pose during radiography is one o

PhyLatent: Learning Dynamics-Relevant Representations for JEPA World Models

SafetyDGX agent

arXiv:2608.05720v1 Announce Type: new Abstract: We propose PhyLatent, a dynamics-relevant training objective for JointEmbedding Predictive Architecture (JEPA) world models. Our key observation is that

Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation

ApplicationsDGX agent

arXiv:2608.05341v1 Announce Type: new Abstract: Vision-Language Models (VLMs) for radiology report generation are typically trained on retrospective clinical reports, which suffer from omission noise:

Prior-SG: Task and Prior Driven Region Segmentation for Scene Graphs in Arbitrarily-Structured Environments

Local AiDGX agent

arXiv:2608.06170v1 Announce Type: cross Abstract: Hierarchical 3D scene graphs are a promising representation for high-level spatial reasoning in autonomous mobile platforms. However, existing extract

PromptForSegCXR: Prompt-Driven Multi-Organ and Multi-Disease Segmentation in Chest X-rays using a Multi-stage Fusion Mechanism

ResearchDGX agent

arXiv:2507.00673v2 Announce Type: replace-cross Abstract: Image segmentation is central to automated medical image analysis, enabling precise identification of anatomical structures and pathological r

Respect Your Zero-Shot Uncertainty: Conservative Calibration for Test-Time-Adapted Vision-Language Models

ResearchDGX agent

arXiv:2608.05945v1 Announce Type: new Abstract: Test-time adaptation (TTA) can improve the recognition accuracy of vision-language models under distribution shift, but often degrades calibration, maki

Reversible Unlearnable Examples: Towards the Copyright Protection in Deep Learning Era

TutorialsDGX agent

arXiv:2608.06211v1 Announce Type: cross Abstract: Significant advancements in deep learning have been made possible by the utilization of large datasets, underscoring the critical importance of copyri

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

SafetyDGX agent

arXiv:2608.05903v1 Announce Type: new Abstract: Mainstream World-Action Models (WAMs) adapt pretrained video generation models (VGMs) for robot control, transferring their learned dynamics prior for a

Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training

SafetyDGX agent

arXiv:2608.06125v1 Announce Type: new Abstract: Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human prefer

SCI-CLIP: Segment-Centric Inference with Reference Memory for Training-Free Open-Vocabulary Segmentation

SafetyDGX agent

arXiv:2608.05627v1 Announce Type: new Abstract: Training-free open-vocabulary segmentation remains limited by a missing inference abstraction. Frozen vision-language features are produced at patch lev

SciQNet: Two-Stage Multimodal Adaptation for Scientific Image Quality Assessment

ResearchDGX agent

arXiv:2608.05691v1 Announce Type: new Abstract: Scientific images are essential for communicating experimental observations, quantitative evidence and conceptual knowledge. Unlike natural images, thei

Shape-Aware Oriented Bounding Box (OBB) to Horizontal Bounding Box (HBB) Conversion

ResearchDGX agent

arXiv:2608.05858v1 Announce Type: new Abstract: Accurate object detection in aerial and satellite imagery is dependent upon the bounding box representation. This is especially true for spatially orien

Sparse Mixture-of-Experts for Non-Uniform Noise Reduction in MRI Images

ApplicationsDGX agent

arXiv:2501.14198v3 Announce Type: replace-cross Abstract: Magnetic Resonance Imaging (MRI) is an essential diagnostic tool in clinical settings, but its utility is often hindered by noise artifacts in

SR-JEPA: Learning Predictive Latent State in 3D Scenes

TutorialsDGX agent

arXiv:2608.05774v1 Announce Type: new Abstract: Joint-embedding predictive architectures learn by predicting latent representations of missing observations, yet many masked JEPAs are evaluated primari

STAIL: Semantic Text-Anchored Incremental Learning for Medical Imaging via Large Language Models

ResearchDGX agent

arXiv:2608.05808v1 Announce Type: new Abstract: Deep learning models applied to medical image analysis suffer from severe catastrophic forgetting when continually adapting to new clinical tasks in dyn

StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding

Model ReleasesDGX agent

arXiv:2608.05703v1 Announce Type: new Abstract: Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-s

StyleComposer: Training-Free Multi-Reference Style Composition

ResearchDGX agent

arXiv:2608.05213v1 Announce Type: new Abstract: The style of a painting is not monolithic: color, texture, and structure may come from different sources. Existing reference-guided methods transfer the

Support Operation Factorization: Compositional Readout of Frozen Vision Encoders under Controlled Interventions

ResearchDGX agent

arXiv:2608.06174v1 Announce Type: new Abstract: Compositional analysis of frozen vision encoders should determine both what changed and where it changed. Standard factor probes score these axes separa

TAU-Bench: From Anomaly Instance Tracking to Fine-Grained Video Anomaly Understanding

Model ReleasesDGX agent

arXiv:2608.05699v1 Announce Type: new Abstract: Humans understand anomalous events through a coherent perceptual process in which they identify the focal instance, follow its behavior as the event unf

Text-Guided Refinement of Multi-sequence Glioma Subregion Segmentation with a Vision-Language Foundation Model

Model ReleasesDGX agent

arXiv:2608.05389v1 Announce Type: new Abstract: Background: Accurate glioma subregion delineation is important for radiotherapy planning and longitudinal monitoring, but manual contour correction is t

The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents

Model ReleasesDGX agent

arXiv:2608.06065v1 Announce Type: new Abstract: GUI agents are commonly trained offline from successful interaction trajectories. Standard training decomposes each trajectory into prefix-action pairs:

TLNM: Externally Validated Tooth Detection, Numbering and Segmentation from Smartphone Photographs Using Mask R-CNN

Local AiDGX agent

arXiv:2608.06275v1 Announce Type: new Abstract: Oral health issues affect billions globally, but the cost and limited access to professional dental care hinder preventive oral healthcare. Research rel

To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation

ResearchDGX agent

arXiv:2608.05879v1 Announce Type: new Abstract: Text-driven 3D generation has advanced rapidly in creating large-scale outdoor environments and detailed indoor scenes, but these domains are usually sy

Topology-Aware Neighborhood Learning for Source-Free Cross-Scene Hyperspectral Image Classification

ResearchDGX agent

arXiv:2608.05964v1 Announce Type: new Abstract: Domain adaptation has advanced cross-scene hyperspectral image classification, significantly improving discriminative capability in complex scenarios. H

Tree-NET: Enhancing 2D Medical Image Segmentation Through Efficient Low-Level Feature Training

Model ReleasesDGX agent

arXiv:2501.02140v2 Announce Type: replace-cross Abstract: This paper introduces Tree-NET, a novel framework for medical image segmentation that leverages bottleneck supervision to enhance both segment

TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs

ResearchDGX agent

arXiv:2608.05616v1 Announce Type: new Abstract: Despite the remarkable progress of large vision language models (LVLMs), object hallucination remains a fundamental challenge that hinders their trustwo

Uncertainty-Aware World Model for Aerial Image-Goal Navigation

ApplicationsDGX agent

arXiv:2608.05597v1 Announce Type: new Abstract: Aerial image-goal navigation requires an unmanned aerial vehicle (UAV) to reach a target location specified by a goal image. Existing world-model-based

Universal Concept Disruption for SAM3 Image Segmentation

ResearchDGX agent

arXiv:2608.05983v1 Announce Type: new Abstract: SAM3 extends promptable segmentation from geometry-driven mask prediction to open-vocabulary concept segmentation, where a text-conditioned grounding mo

UQ-Loc: Uncertainty-Aware LiDAR Scene Coordinate Regression

Local AiDGX agent

arXiv:2608.06307v1 Announce Type: new Abstract: LiDAR-based Scene Coordinate Regression (SCR) maps point clouds directly to 3D scene coordinates, enabling precise 6-DoF localisation without explicit m

URNet: A Unified Reparameterized Network for Efficient RGB-D Semantic Segmentation

ResearchDGX agent

arXiv:2608.05671v1 Announce Type: new Abstract: Previous RGB-D semantic segmentation methods commonly employ dual encoders to separately process RGB and depth inputs, followed by dedicated modules for

Versatile Video Representation via Feed-Forward 2D Gaussian Splatting Tokenization

ResearchDGX agent

arXiv:2508.11183v2 Announce Type: replace Abstract: Recent video representation methods that rely on fixed-grid, patch-wise tokenization often exhibit limited versatility.Spatially, uniformly allocati

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing

Model ReleasesDGX agent

arXiv:2608.05485v1 Announce Type: new Abstract: Evaluating generated videos remains challenging because existing benchmarks rely on fixed evaluation content, cover only a subset of generation and edit

Visual Intention Grounding for Egocentric Assistants

Model ReleasesDGX agent

arXiv:2504.13621v2 Announce Type: replace Abstract: Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object qu

VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances

TutorialsDGX agent

arXiv:2608.05215v1 Announce Type: cross Abstract: Learning manipulation skills from human videos is promising for scalable robot learning. However, the embodiment mismatch between humans and robots ma

Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification

Model ReleasesDGX agent

arXiv:2608.05776v1 Announce Type: new Abstract: Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditi

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation

ResearchDGX agent

arXiv:2608.05648v1 Announce Type: new Abstract: Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

ResearchDGX agent

arXiv:2608.05803v1 Announce Type: new Abstract: Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming

Local AiDGX agent

arXiv:2608.05663v1 Announce Type: new Abstract: Real-time long-form avatar audio--video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consis

VSMP-IMU: Video-Grounded Semantic Motion Programs for Sensor-Aware Synthetic IMU Generation

ResearchDGX agent

arXiv:2608.05782v1 Announce Type: new Abstract: Wearable human activity recognition (HAR) is often limited by the scarcity of labeled sensor data, especially in low-resource, class-imbalanced, and sub

Wan-Animate-2: Pushing the Application Boundaries of Character Animation

ResearchDGX agent

arXiv:2608.06009v1 Announce Type: new Abstract: Character image animation remains a foundational yet challenging task in computer vision. Existing approaches can be broadly categorized into three para

What Drives Test-Time Adaptation for CLIP? A Controlled Empirical Study from an Update Perspective

Model ReleasesDGX agent

arXiv:2606.14299v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) such as CLIP have become a standard backbone for open-vocabulary recognition, yet their zero-shot predictions remain v

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

Local AiDGX agent

arXiv:2608.05369v1 Announce Type: cross Abstract: Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in r

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

SafetyDGX agent

arXiv:2608.05799v1 Announce Type: cross Abstract: Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to

6 Aug 2026

A GitOps-Driven Annotation Catalog for Fully Automatic Railway Operations

ApplicationsDGX agent

arXiv:2608.04724v1 Announce Type: cross Abstract: Automatic train operation (ATO) at grade of automation 3 and above (GoA3-GoA4) requires robust AI-based perception systems capable of reliably detecti

A Multi-Sensor Dataset for Monitoring the Operational Environment of Rail Vehicles

ResearchDGX agent

arXiv:2608.04704v1 Announce Type: new Abstract: Reliable environment monitoring is essential for the safe and efficient operation of automated railway systems, covering all Grades of Automation (GoA),

ACA-GS: Adaptive-Capacity Anchored Gaussian Splatting for Compact Dynamic Radiance Fields

Model ReleasesDGX agent

arXiv:2608.04581v1 Announce Type: new Abstract: Recent advances in 4D Gaussian Splatting (4DGS) enable high-fidelity, real-time spatiotemporal rendering, but expose a fundamental trade-off between mot

Advancing Utility Pole and Sign Detection Through Deep Learning

Model ReleasesDGX agent

arXiv:2608.04061v1 Announce Type: new Abstract: Utility poles are an essential part of the infrastructure used to support power distribution systems and other critical public services. Their regular i

Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle

Model ReleasesDGX agent

arXiv:2608.04314v1 Announce Type: cross Abstract: Once visual content enters an AI pipeline, its owner often retains little technical control over how it is used. Legal and regulatory remedies can add

Aligning Fetal Anatomy with Kinematic Tree Log-Euclidean PolyRigid Transforms

ResearchDGX agent

arXiv:2603.02371v2 Announce Type: replace Abstract: Automated analysis of articulated bodies is crucial in medical imaging. Existing surface-based models often ignore internal volumetric structures an

← Previous
1…89101112…207
Next →