AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
9 Jul 2026

Attention in Geometry: Scalable Spatial Modeling via Adaptive Density Fields and FAISS-Accelerated Kernels

ResearchDGX agent

arXiv:2601.06135v3 Announce Type: replace-cross Abstract: Spatial computation in geographic systems increasingly requires query-conditioned, local, interpretable aggregation under metric constraints.

Automatic Echocardiography Segmentation via Transition Probability Correlation for Stable Semantic Extraction

SafetyDGX agent

arXiv:2607.07580v1 Announce Type: new Abstract: While echocardiography is essential for cardiovascular diagnosis, inherent speckle noise and low signal-to-noise ratio often lead to ambiguous semantic

Bi-PT: Bidirectional Cross-Attention Point Transformers for Four-Chamber Heart Reconstruction from Sparse Cardiac MRI Data

Research

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2607.06923v1 Announce Type: new Abstract: We propose Bi-PT, a pipeline for reconstructing 3D four-chamber human heart meshes from clinical sparsely sampled cardiac magnetic resonance imaging (CM

BUS: Brain-Inspired Unsupervised Self-Reflection for Advanced Multimodal Reasoning

ResearchDGX agent

arXiv:2607.07361v1 Announce Type: new Abstract: Current Vision-Language Models (VLMs) often struggle to handle complex visual tasks that require consistent and fine-grained reasoning. Recent methods a

Cardiac MRI Through-Plane Super-Resolution Guided by Reference and Memory

ResearchDGX agent

arXiv:2607.07581v1 Announce Type: new Abstract: Clinical cardiac MRI is commonly acquired with high in-plane resolution but coarse through-plane resolution to reduce scan time and accommodate breath-h

CEVAR: Centerline Embedding Extraction for Endovascular Aneurysm Repair

ResearchDGX agent

arXiv:2606.15667v2 Announce Type: replace Abstract: Long-term mortality rates after endovascular aneurysm repair (EVAR) remain elevated due to post-EVAR rupture caused by loss of seal in stent graft s

CoFINN: Conservation Flux Informed Neural Networks for Physics Problems Governed by Conservation Laws

ResearchDGX agent

arXiv:2607.06587v1 Announce Type: new Abstract: We present CoFINN (Conservation Flux Informed Neural Networks), a physics-informed deep learning framework for predicting compressible flow fields gover

ColorFM: An Optimization-to-Learning Framework for Color Transfer via Flow Matching

ResearchDGX agent

arXiv:2607.07119v1 Announce Type: new Abstract: Color transfer aims to align the color distribution of a source image with that of a reference image while preserving structural and semantic consistenc

CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views

AgentsDGX agent

arXiv:2607.06691v1 Announce Type: new Abstract: Human-human collaboration is a fundamental aspect of everyday life, essential to success in a wide range of goal-directed activities from household task

Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering

Model ReleasesDGX agent

arXiv:2607.07179v1 Announce Type: new Abstract: Document Visual Question Answering (DocVQA) presents a complex multimodal challenge, requiring models to exploit visual, textual, and layout information

Compass: Prostate Cancer Detection Needs Multi-View Context

ResearchDGX agent

arXiv:2607.06919v1 Announce Type: new Abstract: Artificial intelligence (AI) analysis of micro-ultrasound (muUS) has shown promise for prostate cancer (PCa) detection. However, most existing AI method

Context-Aware Slum Mapping in Sub-Saharan Africa Using Sentinel-1 Texture and Local Climate Zones

Local AiDGX agent

arXiv:2607.07532v1 Announce Type: new Abstract: Accurate mapping of informal settlements remains a major challenge in Sub-Saharan African (SSA) cities because optical imagery often fails to distinguis

CRIS: Cross-Plane Self-Supervised Isotropic Restoration for Anisotropic Volumetric Imaging Across Modalities

ResearchDGX agent

arXiv:2606.15967v2 Announce Type: replace Abstract: Anisotropic volumetric acquisitions are common in clinical MRI and volume electron microscopy (vEM), where sparse through-plane sampling creates thi

DiffCVE: Diffusion-based Compressed Video Enhancement

ResearchDGX agent

arXiv:2607.07195v1 Announce Type: new Abstract: Perceptual quality enhancement of severely compressed videos remains challenging due to complex artifact patterns and substantial information loss. Rece

Discovering Geometric Biases in 3D Face Reconstruction: A Curvature-Aware Spectral Framework for Fairness Evaluation

SafetyDGX agent

arXiv:2607.07486v1 Announce Type: new Abstract: 3D Morphable Models (3DMMs) remain the standard parametric shape priors for many state-of-the-art 3D face reconstruction algorithms. However, as these m

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

SafetyDGX agent

arXiv:2607.07608v1 Announce Type: cross Abstract: Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling wi

DYNA-PRUNER: Input-Adaptive Data-Model Co-Pruning for Efficient and Scalable Spatio-Temporal Media Prediction

HardwareDGX agent

arXiv:2606.15346v2 Announce Type: replace Abstract: Spatio-temporal prediction supports radar/satellite nowcasting and city-scale traffic monitoring, but modern models are often too expensive for real

Dynamic Object Detection and Tracking in Construction: A Fisheye Camera and LiDAR Sensor Fusion Model

ApplicationsDGX agent

arXiv:2607.06896v1 Announce Type: cross Abstract: Robust dynamic object detection and tracking are essential for enabling robots to operate safely and effectively alongside humans in complex environme

ECHO: Ego-Centric modeling of Human-Object interactions

ResearchDGX agent

arXiv:2508.21556v3 Announce Type: replace Abstract: Modeling human-object interactions (HOI) from an egocentric perspective is a critical yet challenging task, particularly when relying on sparse sign

EdgeCompress: Coupling Multidimensional Model Compression and Dynamic Inference for EdgeAI

ResearchDGX agent

arXiv:2607.06982v1 Announce Type: new Abstract: Convolutional neural networks (CNNs) have demonstrated encouraging results in image classification tasks. However, the prohibitive computational cost of

EditVerse3D: High-Quality 3D Object Editing with Region-Aware Learning

Local AiDGX agent

arXiv:2607.07187v1 Announce Type: new Abstract: Local editing of 3D objects remains a long-standing challenge. When interacting with 3D content, humans naturally tend to specify a coarse region of int

Ego-Human Motion Prediction with 3D-Aware LLM

Model ReleasesDGX agent

arXiv:2607.07001v1 Announce Type: new Abstract: Anticipating human motion from an egocentric perspective is fundamental for proactive assistance in AR/VR, human-robot collaboration, and embodied AI. W

EmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AI

SafetyDGX agent

arXiv:2607.07459v1 Announce Type: cross Abstract: We present EmbodiedGen V2, a generative 3D world engine for building executable sim-ready environments for embodied intelligence. Sim-ready 3D asset g

Ensemble Deep Learning Approaches for AI-Altered Video Detection

TutorialsDGX agent

arXiv:2607.06872v1 Announce Type: new Abstract: The increasing accessibility of artificial intelligence has led to a rapid rise in AI-generated videos, making it more difficult to distinguish between

EventVGGT: Exploring Cross-Modal Distillation for Consistent Event-based Depth Estimation

ResearchDGX agent

arXiv:2603.09385v2 Announce Type: replace Abstract: Event cameras offer superior sensitivity to high-speed motion and extreme lighting, making event-based monocular depth estimation a promising approa

Face-trace: Open-Set Attribution and Progressive Discovery of Synthetic Face Generators

ApplicationsDGX agent

arXiv:2607.07545v1 Announce Type: new Abstract: Recent advances in generative Artificial Intelligence have made synthetic face images increasingly realistic, creating new challenges for multimedia for

FMMC: Harnessing the Power of Foundation Models for Accurate Material Classification

Model ReleasesDGX agent

arXiv:2603.17390v2 Announce Type: replace Abstract: Material classification has emerged as a critical task in computer vision and graphics, supporting the assignment of accurate material properties to

Format-Controlled Multi-Scale JPEG Compression Response Analysis for Image-Level Forgery Screening

Local AiDGX agent

arXiv:2607.06615v1 Announce Type: cross Abstract: Image forgery detection is a critical task in digital forensics, yet many deep-learning localization approaches are typically GPU-accelerated and comp

From Data Completeness to Data Sufficiency: A Task-Driven Imaging Framework for Intraoperative CBCT under Quality-Time-Dose Trade-offs

ResearchDGX agent

arXiv:2607.07039v1 Announce Type: cross Abstract: Mobile C-arm cone-beam computed tomography (CBCT) has been widely used for real-time intraoperative 3D imaging. However, current practice often mechan

From My View to Yours: Learning Egocentric Cues from Exocentric Video using Privileged Egocentric Supervision

Model ReleasesDGX agent

arXiv:2501.05711v4 Announce Type: replace Abstract: Vision Language Models (VLMs) have achieved strong performance across a wide range of video understanding tasks. However, their viewpoint-invariant

G-PROBE: Cross-FOV Place Recognition and Certainty-Coupled Localization for 3D Point Clouds

ResearchDGX agent

arXiv:2607.06782v1 Announce Type: cross Abstract: Global localization from 3D point clouds remains challenging under limited or asymmetric fields of view (FOV), which fail to provide the dense, symmet

G-ZAP: A Generalizable Zero-Shot Framework for Arbitrary-Scale Pansharpening

ApplicationsDGX agent

arXiv:2603.14412v2 Announce Type: replace Abstract: Pansharpening aims to fuse a high-resolution panchromatic (PAN) image and a low-resolution multispectral (LRMS) image to produce a high-resolution m

Gen4U: Unifying Video Generation and Understanding via Diffusion

SafetyDGX agent

arXiv:2607.06856v1 Announce Type: new Abstract: Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-a

General Incomplete Multimodal Learning via Dynamic Quality Perception

ApplicationsDGX agent

arXiv:2607.06943v1 Announce Type: new Abstract: Multimodal learning robust to missing modalities is essential for real-world applications. Existing methods mainly focus on inter-modality missing, wher

Geometric Collapse: When Vision Models Fail to Verify Physical Causality

ResearchDGX agent

arXiv:2607.06871v1 Announce Type: new Abstract: Recent progress in large-scale self-supervised learning has improved dense geometric prediction, but it remains unclear whether such scaling yields infe

Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation

ResearchDGX agent

arXiv:2512.05044v2 Announce Type: replace Abstract: Generating interactive and dynamic 4D scenes from a single static image remains a core challenge. Most existing generate-then-reconstruct and recons

GP-4DGS: Probabilistic 4D Gaussian Splatting from Monocular Video via Variational Gaussian Processes

ResearchDGX agent

arXiv:2604.02915v2 Announce Type: replace Abstract: We present GP-4DGS, a novel framework that integrates Gaussian Processes (GPs) into 4D Gaussian Splatting (4DGS) for principled probabilistic modeli

Hardware-aware Graph Neural Networks prunning for embedded event-based vision

ResearchDGX agent

arXiv:2607.06739v1 Announce Type: new Abstract: Event-based cameras are gaining popularity as the sensor of choice for mobile robotics, due to their high performance in dynamic environments. However,

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework

SafetyDGX agent

arXiv:2602.23615v3 Announce Type: replace Abstract: Current Large Multimodal Models (LMMs) struggle with high-resolution visual inputs during the reasoning process, as the number of image tokens incre

HPR-SAM: Hierarchical Probabilistic Representation Learning for Prompt-free SAM-based Medical Image Segmentation

ResearchDGX agent

arXiv:2607.06972v1 Announce Type: new Abstract: Prompt-free adaptation of the Segment Anything Model (SAM) has emerged as a promising paradigm for automatic medical image segmentation. Existing method

HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation

Model ReleasesDGX agent

arXiv:2506.08797v2 Announce Type: replace Abstract: To address key limitations in human-object interaction (HOI) video generation -- specifically the reliance on curated motion data, limited generaliz

Infinite Worlds with Versatile Interactions

HardwareDGX agent

arXiv:2607.07534v1 Announce Type: new Abstract: We present LingBot-World 2.0 (also known as LingBot-World-Infinity), an advanced iteration of LingBot-World featuring four distinct upgrades. (1) Our mo

InfraQR: Edge-Placed QR-Inspired Structured Patch Attacks on Infrared Vision-Language Models

Model ReleasesDGX agent

arXiv:2607.07288v1 Announce Type: new Abstract: Infrared vision-language models are increasingly used for perception under low-light and adverse visual conditions, yet their robustness to localized st

Latency-Constrained DNN Architecture Learning for Edge Systems using Zerorized Batch Normalization

HardwareDGX agent

arXiv:2607.06922v1 Announce Type: cross Abstract: Deep learning applications have been widely adopted on edge devices, to mitigate the privacy and latency issues of accessing cloud servers. Deciding t

Learning to Unify Deformable Shape and Texture Representations for Cardiac Video Classification

ResearchDGX agent

arXiv:2607.07518v1 Announce Type: new Abstract: Deformable shape representations have proven to be robust complements to texture features in cardiac image classification, offering geometric priors tha

LEMUR 2: Unlocking Neural Network Diversity for AI

ResearchDGX agent

arXiv:2607.06839v1 Announce Type: cross Abstract: Existing NAS benchmarks (e.g., NAS-Bench, NATS-Bench) cover only narrow, task-specific regions of the architectural design space and lack cross-domain

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

Model ReleasesDGX agent

arXiv:2607.07673v1 Announce Type: new Abstract: Medicine is inherently multimodal, requiring clinicians to synthesize information across diverse data streams. Yet the development of multimodal foundat

MegaFlow: Zero-Shot Large Displacement Optical Flow

Local AiDGX agent

arXiv:2603.25739v2 Announce Type: replace Abstract: Accurate estimation of large displacement optical flow remains a critical challenge. Existing methods typically rely on iterative local search or/an

MMAgent-R^2: Learning to Rerank and Reject for Agentic mRAG

AgentsDGX agent

arXiv:2607.07383v1 Announce Type: new Abstract: Knowledge-based Visual Question Answering (KB-VQA) requires models to retrieve visual entities matching the query image from large-scale encyclopedic kn

MMDiff: Extending Diffusion Transformers for Multi-Modal Generation

ResearchDGX agent

arXiv:2606.16673v2 Announce Type: replace Abstract: Diffusion transformers have demonstrated remarkable generative capabilities, yet the rich perceptual representations computed across their denoising

Naming the Concepts Classifiers Rely On: Language-Anchored Decomposition for Faithful Explanation

Local AiDGX agent

arXiv:2607.07264v1 Announce Type: new Abstract: Deep neural networks are widely deployed in high-stakes visual applications where interpretability is critical, yet existing explanations face a trade-o

NoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction

ResearchDGX agent

arXiv:2607.07168v1 Announce Type: new Abstract: Pose-Free Feed-forward 3D Gaussian Splatting (3DGS) has recently emerged as a powerful paradigm for fast scene reconstruction. However, its performance

Physically Grounded Monocular Depth via Nanophotonic Wavefront Encoding

ResearchDGX agent

arXiv:2503.15770v3 Announce Type: replace-cross Abstract: Depth foundation models (DFMs) offer strong learned priors for 3D perception from single RGB images but lack physical depth cues, leading to a

Pixel-Precise Explainable Stress Indexing: A Semantic Segmentation Framework for Disease Severity Quantification in Field Crops

ResearchDGX agent

arXiv:2607.06585v1 Announce Type: new Abstract: Plant diseases, resulting from both biotic and abiotic stresses, cause an estimated 20-40% loss in global agricultural yield annually, resulting in econ

Prior-matched evaluation of operational Earth-observation classifiers: a three-number reporting method demonstrated on Sentinel-1 internal-wave detection

ResearchDGX agent

arXiv:2607.07146v1 Announce Type: cross Abstract: The Internal Waves Service screens the Sentinel-1 Wave-mode archive for internal solitary waves, routing detections to experts whose adjudication time

Prototype-Anchored Generalized Manifold Regression for Unknown-Domain Object Detection

TutorialsDGX agent

arXiv:2607.07192v1 Announce Type: new Abstract: In this paper, we study Single-Domain Generalized Object Detection (Single-DGOD), which aims to transfer a detector trained on a single source domain to

PUF: Plug-and-Play Uncertainty-Aware Fusion for Online 3D Scene Graph Generation

ResearchDGX agent

arXiv:2607.07170v1 Announce Type: new Abstract: Online 3D scene graph generation builds a persistent, structured representation of a scene by incrementally fusing 2D observations into a global 3D grap

Rail Track Extraction from Rasterized Classified Point Clouds Using a Full-Resolution, Fully Convolutional Recurrent Neural Network

ResearchDGX agent

arXiv:2607.06829v1 Announce Type: new Abstract: Rail track extraction is essential for effective railway asset management and maintenance, especially in automated inspection and mapping workflows. Thi

Retrieving and Refining Winning Noise Tickets for Diffusion-Based Motion Generation

Model ReleasesDGX agent

arXiv:2607.06843v1 Announce Type: new Abstract: Diffusion-based text-to-motion models synthesize realistic human motions but often exhibit semantic drift from the input text. Motion is inherently temp

ROAD-Waymo: A Large-Scale Action Awareness Dataset for Autonomous Driving

Model ReleasesDGX agent

arXiv:2411.01683v3 Announce Type: replace Abstract: Autonomous Vehicle (AV) perception systems require more than simply seeing, via e.g., object detection or scene segmentation. They need a holistic u

← Previous
1…4445464748…209
Next →