AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
15 Jul 2026

Together, Then Apart: Balancing Alignment and Distinctiveness for Multimodal Survival Analysis

SafetyDGX agent

arXiv:2511.18089v2 Announce Type: replace Abstract: Multimodal survival analysis aims to improve cancer prognosis using heterogeneous biomedical data, such as histopathology images and genomic profile

Towards Vision-Free CIR: Attribute-Augmented Scoring and LLM-Based Reranking for Zero-Shot Composed Image Retrieval

ResearchDGX agent

arXiv:2607.12621v1 Announce Type: new Abstract: Recent work has shown that 'Vision-Free'' approaches (representing images as text) can be effective for standard image retrieval tasks. However, it rema

Traj-VLN: Learning Pixel-Space Interaction via Autoregressive Trajectory Generation

AgentsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.10744v2 Announce Type: replace Abstract: Benefiting from the powerful priors embedded in large-scale pre-training data and the emerging commonsense reasoning ability, large language models

TSCA-Net: Temporal-Spatial Clique Attention for Interpretable Multimodal Pedestrian Trajectory Prediction

Local AiDGX agent

arXiv:2607.11939v1 Announce Type: new Abstract: Accurate pedestrian trajectory prediction in crowded environments remains challenging due to the multimodal uncertainty of human motion and the variable

UMSS: Towards Unsupervised Multi-modal Semantic Segmentation

TutorialsDGX agent

arXiv:2607.12372v1 Announce Type: new Abstract: Multimodal semantic segmentation (MSS) is essential for robust perception in complex environments, yet its potential remains largely untapped because of

Uncertainty-Aware Multi-Source Retinal Fluid Segmentation in OCT

ResearchDGX agent

arXiv:2607.12212v1 Announce Type: cross Abstract: Measuring retinal fluid from optical coherence tomography (OCT) drives treatment decisions in macular disease, but manual annotation is slow and segme

UniMedSeg: Unified In-Context Learning for Multi-Paradigm 2D/3D Medical Image Segmentation

ResearchDGX agent

arXiv:2607.12896v1 Announce Type: new Abstract: Medical image segmentation foundation models are expected to generalize across diverse clinical scenarios, yet existing universal methods remain fragmen

UniVR: Thinking in Visual Space for Unified Visual Reasoning

Model ReleasesDGX agent

arXiv:2607.12800v1 Announce Type: new Abstract: Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation in

VanillaBench: The Hidden Accuracy Cost of Adversarial Robustness

Model ReleasesDGX agent

arXiv:2607.12545v1 Announce Type: cross Abstract: Adversarial robustness research has produced hundreds of defended models over the past decade, yet the literature almost universally reports robustnes

ViCo3D: Empowering LiDAR-based Collaborative 3D Object Detection with Vision Foundation Models

AgentsDGX agent

arXiv:2607.12959v1 Announce Type: new Abstract: LiDAR-based collaborative 3D perception in Vehicle-to-Everything (V2X) systems typically relies on fusing bird's-eye-view (BEV) features across agents.

Virtual Chromoendscopy with Tunable Visibility Enhancement

ResearchDGX agent

arXiv:2607.12416v1 Announce Type: new Abstract: Chromoendoscopy (CE) is a common clinical practice that sprays indigo carmine blue dye onto the gastric surface to improve the visibility of gastric les

VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models

AgentsDGX agent

arXiv:2606.13460v2 Announce Type: replace Abstract: Semantic 3D occupancy provides a voxelized world state for autonomous driving and robot decision making, but object and rare-class errors can affect

VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression

Model ReleasesDGX agent

arXiv:2607.12756v1 Announce Type: new Abstract: Vision-language models (VLMs) process large numbers of visual tokens, resulting in substantial inference latency and memory overhead. This has motivated

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors

ResearchDGX agent

arXiv:2512.15748v2 Announce Type: replace-cross Abstract: Visual Species Recognition (VSR) is a fundamental task in scientific disciplines that require species-level identification, including ecology,

WanToFight: Real-Time Generative Game Engine for Multi-Player Combat Interaction

Local AiDGX agent

arXiv:2607.12592v1 Announce Type: new Abstract: We present WanToFight, a generative game engine that simulates real-time, two-player The King of Fighters '97 (KOF~'97) gameplay from keyboard input. Pr

What Does a Temporal Benchmark Score Measure? Decomposing Channel Use in Video VLM Evaluation

Model ReleasesDGX agent

arXiv:2607.12304v1 Announce Type: new Abstract: A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions. 1.

X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

Model ReleasesDGX agent

arXiv:2607.12993v1 Announce Type: new Abstract: We present X-lens, a compact feed-forward model for metric depth estimation from a variable number of calibrated fisheye and pinhole views. To support r

10 Jul 2026

3D Reconstruction of deciduous Trees using low-cost UAV- and Crane-based Photogrammetry for Monitoring Shoot Elongation across entire Canopies

ApplicationsDGX agent

arXiv:2607.07905v1 Announce Type: new Abstract: Tree growth determines how much CO2 is sequestered from the atmosphere and temporarily stored in woody biomass. At the same time tree growth is affected

A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding

Local AiDGX agent

arXiv:2512.21414v2 Announce Type: replace Abstract: Recent tool-use frameworks powered by vision-language models (VLMs) improve image understanding by grounding model predictions with specialized tool

Anatomically Guided Latent Diffusion for Brain MRI Progression Modeling

ResearchDGX agent

arXiv:2601.14584v2 Announce Type: replace Abstract: Accurately modeling longitudinal brain MRI progression is crucial for understanding neurodegenerative diseases and predicting individualized structu

ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation

Model ReleasesDGX agent

arXiv:2607.08741v1 Announce Type: cross Abstract: Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics. While rece

Are Current Continual Learning Methods Truly Agnostic? Introducing OPRE, a Step Toward Agnostic Continual Learning

ResearchDGX agent

arXiv:2511.08226v2 Announce Type: replace-cross Abstract: In order to achieve Continual Learning (CL), the problem of catastrophic forgetting, one that has plagued neural networks since their inceptio

ARGUS: Accelerated, Robust, General, and Unsupervised Cell Tracking Solutions

HardwareDGX agent

arXiv:2607.08297v1 Announce Type: new Abstract: Background and Objective: Quantitative analysis of cell dynamics is central to modern biological research, providing critical insights into immune cell

Asynchronous Federated Continual Segmentation with Evolving Clients and Label Spaces

Local AiDGX agent

arXiv:2503.15414v3 Announce Type: replace-cross Abstract: Federated learning seeks to foster collaboration among distributed clients while preserving the privacy of their local data. Traditional feder

Attention-Based Segmentation of WMHs and Differentiation of Vascular vs. Demyelinating Lesions

ResearchDGX agent

arXiv:2607.08171v1 Announce Type: new Abstract: White Matter Hyperintensities (WMHs) are commonly observed in brain Magnetic Resonance Imaging (MRI) scans. They are associated with various neurologica

Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation

Model ReleasesDGX agent

arXiv:2607.08397v1 Announce Type: new Abstract: Referring Image Segmentation (RIS) aims to segment image regions specified by natural language, enabling fine-grained and controllable visual understand

Benchmark Evaluation of Feredated Learning on Multi-organ Images

Model ReleasesDGX agent

arXiv:2607.08219v1 Announce Type: new Abstract: The privacy requirements of medical data and its substantial variations across organs and modalities hinder the clinical implementation of medical AI. F

Beyond wheelchairs and blindfolds: Investigating disability stereotypes in T2I models with INCLUDE-BENCH

Model ReleasesDGX agent

arXiv:2607.08515v1 Announce Type: new Abstract: Text-to-image (T2I) models have been shown to exhibit social biases. Prior work has mainly focused on gender, skin tone, and cultural representation wit

BiasBench: A reproducible benchmark for tuning the biases of event cameras

Model ReleasesDGX agent

arXiv:2504.18235v2 Announce Type: replace Abstract: Event-based cameras are bio-inspired sensors that detect light changes asynchronously for each pixel. They are increasingly used in fields like comp

Borrowing from anything: A generalizable framework for reference-guided instance editing

SafetyDGX agent

arXiv:2512.15138v2 Announce Type: replace Abstract: Reference-guided instance editing is fundamentally limited by semantic entanglement, where a reference's intrinsic appearance is intertwined with it

Classical versus Deep Mirror-Symmetry Scoring: A Benchmark of Thirteen Methods

Model ReleasesDGX agent

arXiv:2607.08379v1 Announce Type: new Abstract: Quantifying how mirror-symmetric an image is about a given axis (symmetry scoring) underpins applications from visual aesthetics to medical imaging, yet

Closing the Null Space: Guidance-Aware Quantization for Classifier-Free Diffusion

Model ReleasesDGX agent

arXiv:2607.08241v1 Announce Type: new Abstract: Deploying classifier-free guidance (CFG) diffusion models under real-world compute budgets requires quantization, yet existing post-training quantizatio

Computation, Condensation, and the Incompleteness Between Them: A Coupled Foundation of Intelligence

AgentsDGX agent

arXiv:2303.04203v4 Announce Type: replace-cross Abstract: The theory of computation was built to answer Turing's question: what is effectively calculable by an unbounded, immortal, disembodied agent f

ConRad: Efficient Conformal Prediction for Radiomics

ResearchDGX agent

arXiv:2607.08084v1 Announce Type: cross Abstract: Radiomic features derived from medical images and segmentation masks are used to support decision making in clinical imaging pipelines. In practice, t

Continual Test-Time Adaptation in Computer Vision: Methods, Benchmarks, and Future Directions

Model ReleasesDGX agent

arXiv:2607.08164v1 Announce Type: new Abstract: Deep neural nets achieve remarkable performance when training and test data share the same distribution, but this assumption frequently breaks in real-w

CT-CLIP Representations for Multimodal Lung Cancer Survival Prediction

ResearchDGX agent

arXiv:2607.08503v1 Announce Type: new Abstract: Accurate prognosis prediction is important for treatment planning in lung cancer, but deep learning-driven survival modelling is often limited by the sc

Data Alchemy: Mitigating Cross-Site Model Variability Through Test Time Data Calibration

ResearchDGX agent

arXiv:2407.13632v2 Announce Type: replace Abstract: Deploying deep learning-based imaging tools across various clinical sites poses significant challenges due to inherent domain shifts and regulatory

DeltaDeno: Zero-Shot Anomaly Generation via Delta-Denoising Attribution

SafetyDGX agent

arXiv:2511.16920v2 Announce Type: replace Abstract: Anomaly generation is often framed as few-shot fine-tuning with anomalous samples, which contradicts the scarcity that motivates generation and tend

DeltaV: Thinking with Visual State Updates in Unified Large Multimodal Models

Model ReleasesDGX agent

arXiv:2607.08434v1 Announce Type: new Abstract: Current Unified Large Multimodal Models (ULMMs) support interleaved multimodal reasoning through textual reasoning and intermediate visual states, but t

Diagnosing Corruption-Induced Reliability Failures in Vision-Language Models

SafetyDGX agent

arXiv:2511.19032v2 Announce Type: replace Abstract: Visual corruptions can change vision--language model (VLM) behavior in ways that top-1 accuracy does not capture. A model may keep the same answer w

Dive Into the Implicit Biases of Low-rank Vision-language Alignment

Model ReleasesDGX agent

arXiv:2607.08194v1 Announce Type: new Abstract: Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requi

Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?

ResearchDGX agent

arXiv:2607.08514v1 Announce Type: new Abstract: Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models

Do Transformations Reveal the Truth? Generative Residual Learning for Generalized AI-Generated Image Detection

Model ReleasesDGX agent

arXiv:2607.08674v1 Announce Type: new Abstract: The rapid advancement of generative AI has enabled the creation of highly realistic deepfake media, posing significant threats, including misinformation

Dual-Correlation Hypergraph Network for Unaligned RGBT Video Object Detection and A Large-scale Benchmark

Model ReleasesDGX agent

arXiv:2607.08191v1 Announce Type: new Abstract: RGB-Thermal (RGBT) Video Object Detection (VOD) has gained significant traction due to its ability to overcome the limitations of conventional RGB-based

Effective Gaussian Management for High-fidelity Scene Reconstruction

ResearchDGX agent

arXiv:2509.12742v4 Announce Type: replace Abstract: This paper proposes an effective Gaussian management framework for high-fidelity scene reconstruction of both appearance and geometry. Unlike recent

Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding

ApplicationsDGX agent

arXiv:2512.14236v2 Announce Type: replace Abstract: The growing demand for immersive 3D content calls for automated monocular-to-stereo video conversion. We present Elastic3D, a controllable, direct e

Enhancing In-context Panoramic Generation via Geometric-aware Pretraining

TutorialsDGX agent

arXiv:2607.08765v1 Announce Type: new Abstract: In this work, we present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with downstream t

Enhancing the KidSat Model: Integrating Geographical Encoding and Data Quality Assessment for Childhood Poverty Prediction

ResearchDGX agent

arXiv:2607.08281v1 Announce Type: new Abstract: Accurate poverty mapping using satellite imagery is often hindered by (i) noisy and sparse survey-derived supervision, (ii) image quality issues such as

Equivariant Quantum Clustering with Differential Privacy: Parameter-Efficient Privacy-Preserving Analysis Across Heterogeneous Sensitive Datasets

Model ReleasesDGX agent

arXiv:2607.08092v1 Announce Type: cross Abstract: Privacy-preserving clustering is critical for analyzing sensitive data in healthcare, cybersecurity, and enterprise applications, where maintaining da

EVIS: A Physics-Grounded Event Camera Plugin for NVIDIA Isaac Sim

HardwareDGX agent

arXiv:2607.08098v1 Announce Type: new Abstract: Event cameras offer microsecond temporal resolution, low latency, and high dynamic range, making them attractive for robotics. However, labeled event-ca

False Confidence: Automated Labels Confound Fairness Audits in Cervical Spine Segmentation

Model ReleasesDGX agent

arXiv:2607.07852v1 Announce Type: cross Abstract: Automated segmentation of cervical-spine MRI is increasingly used in clinical workflows, yet no fairness audit exists for this anatomy. We show that a

FedTR: Federated Learning Framework with Transfer Learning for Industrial Visual Inspection

Local AiDGX agent

arXiv:2607.08014v1 Announce Type: new Abstract: Federated learning (FL) is a collaborative learning scheme to train deep learning models, where collaborating parties can consolidate their models witho

FunHOI: Annotation-Free 3D Hand-Object Interaction Generation via Functional Text Guidance

ResearchDGX agent

arXiv:2502.20805v3 Announce Type: replace-cross Abstract: Hand-object interaction(HOI) is the fundamental link between human and environment, yet its dexterous and complex pose significantly challenge

Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos

Model ReleasesDGX agent

arXiv:2512.01803v3 Announce Type: replace Abstract: Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elu

Geometry and Gradient-based Partitioning for Panoramic Outdoor Reconstruction

Model ReleasesDGX agent

arXiv:2607.08769v1 Announce Type: new Abstract: Scaling 3D Gaussian Splatting (3DGS) to large outdoor scenes is costly in both data acquisition and computation. Adopting panoramic images with equirect

GERD: Geometric event response data generation

ApplicationsDGX agent

arXiv:2412.03259v3 Announce Type: replace Abstract: Event-based vision sensors offer high temporal resolution, high dynamic range, and low power consumption, yet event-based vision models lag behind c

GIRAF: Towards Generalizable Human Interactions with Articulated Objects

ResearchDGX agent

arXiv:2607.07880v1 Announce Type: new Abstract: Synthesizing realistic full-body human interactions with articulated objects is a fundamental challenge for embodied AI and graphics, with applications

GRE-Diff: Gaussian Room Embeddings for Structured Layout Diffusion

ResearchDGX agent

arXiv:2607.08086v1 Announce Type: new Abstract: Designing functional and aesthetically coherent floor plans requires exploring a vast space of possible room arrangements, a task that quickly becomes o

GSurf: Learning Signed Distance Fields from Splatting Opaque Gaussians for High-quality 3D Reconstruction

ResearchDGX agent

arXiv:2411.15723v4 Announce Type: replace Abstract: High-fidelity surface reconstruction from multi-view images is a core problem in 3D computer vision. While neural implicit surfaces like SDFs offer

HairWeaver: Few-Shot Photorealistic Hair Motion Synthesis with Sim-to-Real Guided Video Diffusion

SafetyDGX agent

arXiv:2602.11117v2 Announce Type: replace Abstract: We present HairWeaver, a diffusion-based pipeline that animates a single human image with realistic and expressive hair dynamics. While existing met

← Previous
1…4243444546…209
Next →