AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
7 Jul 2026

MAGE: View-guided Point Cloud Completion with Efficient Modality Alignment and Adaptive Geometry Enhancement

SafetyDGX agent

arXiv:2607.02568v1 Announce Type: new Abstract: View-based point cloud completion aims to recover a complete 3D shape from a partial point cloud, guided by a single-view image. However, existing appro

MambaRefine-CD: MambaVision with Region-Boundary Temporal Refinement

ResearchDGX agent

arXiv:2607.04403v1 Announce Type: cross Abstract: Binary change detection in remote sensing requires both complete changed-region localization and accurate boundary delineation. We present MambaRefine

Measuring 3D Spatial Geometric Consistency in Dynamic Video Generation

ResearchDGX agent

arXiv:2603.19048v2 Announce Type: replace Abstract: Recent generative models can produce high-fidelity videos, yet they often exhibit 3D spatial geometric inconsistencies. Existing evaluation methods


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

MedMambaLite: Hardware-Aware Mamba for Medical Image Classification

Local AiDGX agent

arXiv:2508.05049v1 Announce Type: cross Abstract: AI-powered medical devices have driven the need for real-time, on-device inference such as biomedical image classification. Deployment of deep learnin

MergeSurv: Merging-Based Continual Learning for Survival Analysis on Whole-Slide Images

ResearchDGX agent

arXiv:2607.04747v1 Announce Type: new Abstract: Survival analysis on Whole Slide Images (WSIs) is important in computational pathology for prognosis estimation and treatment planning. However, existin

MetaMax: Improved Open-Set Deep Neural Networks via Weibull Calibration

ResearchDGX agent

arXiv:2211.10872v2 Announce Type: replace Abstract: Open-set recognition refers to the problem in which classes that were not seen during training appear at inference time. This requires the ability t

MGCA-Net: Multi-Grained Category-Aware Network for Open-Vocabulary Temporal Action Localization

ResearchDGX agent

arXiv:2511.13039v2 Announce Type: replace Abstract: Open-Vocabulary Temporal Action Localization (OV-TAL) aims to recognize and localize instances of any desired action categories in videos without ex

Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World Models

SafetyDGX agent

arXiv:2409.16663v5 Announce Type: replace-cross Abstract: We propose the use of latent space generative world models to address the covariate shift problem in autonomous driving. A world model is a ne

Mixture-of-Gaussians-Guided Schedule Design for Brownian Bridge Diffusion Models

SafetyDGX agent

arXiv:2607.03517v1 Announce Type: cross Abstract: Brownian Bridge Diffusion Models (BBDM) offer an appealing framework for image restoration and inverse problems by constructing a stochastic bridge fr

MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training

Model ReleasesDGX agent

arXiv:2602.06285v2 Announce Type: replace Abstract: Recent research in geospatial machine learning has demonstrated that models pretrained with self-supervised learning on Earth observation data can p

Model Confidence-Guided Multi-Image Fusion of Fundus Images for Diabetic Retinopathy Diagnosis

ResearchDGX agent

arXiv:2607.03643v1 Announce Type: cross Abstract: Purpose: Early screening for eye diseases is critical in low- and middle-income countries where access to care is limited. We investigate whether a co

Motion Estimation Techniques for Volumetric Video Attribute Compression

ResearchDGX agent

arXiv:2607.03576v1 Announce Type: cross Abstract: Point cloud compression relies on techniques to compress both geometry and attributes. Motion-based approaches for dynamic solid point cloud geometry

MUSON: A Reasoning-oriented Multimodal Dataset for Socially Compliant Navigation in Urban Environments

Model ReleasesDGX agent

arXiv:2512.22867v2 Announce Type: replace Abstract: Socially compliant navigation requires structured reasoning about dynamic pedestrians and physical constraints to ensure safe and interpretable deci

MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

SafetyDGX agent

arXiv:2607.05376v1 Announce Type: new Abstract: Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis

NABLA: Neighborhood Adaptive Block-Level Attention

ResearchDGX agent

arXiv:2507.13546v2 Announce Type: replace Abstract: Recent progress in transformer-based architectures has demonstrated remarkable success in video generation tasks. However, the quadratic complexity

Natural Language Camera Movement Understanding

Model ReleasesDGX agent

arXiv:2607.03043v1 Announce Type: new Abstract: Understanding camera movement in natural language is critical for training and evaluating video generation models, among other applications. However, we

NavEYE: Vision-Centered Multi-Sensor Fusion-Based Situational Awareness System for Intelligent Surface Vehicles

SafetyDGX agent

arXiv:2607.03915v1 Announce Type: new Abstract: With the rapid development of sensor and artificial intelligence (AI) technologies, intelligent surface vehicles (ISVs) have gained increasing attention

Observable- and Positional-Encoding-Dependent Symmetry Readout from Neural Network Weights

ResearchDGX agent

arXiv:2607.03108v1 Announce Type: cross Abstract: Post-hoc analysis of trained neural network weights often seeks to recover geometric structure directly from the parameters. We show that, for positio

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion

ResearchDGX agent

arXiv:2603.06577v2 Announce Type: replace Abstract: While recent multimodal large language models (MLLMs) have made impressive strides, they predominantly employ a conventional autoregressive architec

OmniDS: Dual-Stream Context Fusion for Omnidirectional Depth from Fisheye Cameras

SafetyDGX agent

arXiv:2607.03038v1 Announce Type: new Abstract: Omnidirectional depth estimation from multi-fisheye camera rigs is complicated by visibility conflicts: wide baselines cause different cameras to observ

OmniLayout: A Schematic-Coupled Multimodal Benchmark for Constraint-Aware Geometric Reasoning in PCB Layout

Model ReleasesDGX agent

arXiv:2607.03261v1 Announce Type: new Abstract: Recent large language models (LLMs) have demonstrated remarkable progress in 3D spatial reasoning, spatial grounding, and fine-grained geometric underst

Open-Attribute Person Retrieval: Finding People Through Distinctive and Novel Attributes

SafetyDGX agent

arXiv:2508.01389v3 Announce Type: replace Abstract: Person retrieval in surveillance videos often depends on attributes described by witnesses or operators. However, the most useful cues in practice a

Overloading Large Vision-Language Models for Jailbreaking

SafetyDGX agent

arXiv:2607.02961v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) exhibit remarkable vision-language capabilities and are increasingly deployed in real-world applications such as pe

PAGE: Towards Practical Human-level Gaze Target Estimation

ResearchDGX agent

arXiv:2607.04860v1 Announce Type: new Abstract: Gaze target estimation, the task of predicting where a person is looking in a scene, is crucial to understanding human attention and intent. It is a cha

Paired Uterine Whole-Slide Images and Pathology Reports for Multimodal Computational Pathology

ResearchDGX agent

arXiv:2607.04020v1 Announce Type: new Abstract: Uterine diseases represent an important category of gynecologic pathology and require accurate histopathological assessment for diagnosis and treatment

Perceiving Better Moments: Cover Frame Reselection and Enhancement for Live Photos with the Live2K Dataset

Model ReleasesDGX agent

arXiv:2607.04151v1 Announce Type: new Abstract: Modern smartphones capture Live Photos, short video bursts surrounding a still image, offering a dynamic and engaging photographic experience. However,

Perceptual Flow Matching for Few-Step Generative Modeling

ResearchDGX agent

arXiv:2607.03524v1 Announce Type: new Abstract: We propose Perceptual Flow Matching (PFM), a simple yet effective framework for few-step generation in flow-matching models. Rather than performing velo

Phi-SegNet: Phase-Integrated Supervision for Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2601.16064v2 Announce Type: replace-cross Abstract: Deep learning has substantially advanced medical image segmentation, yet achieving robust generalization across diverse imaging modalities and

PhysMirror: Physics-Aware Mirror Object Generation

Model ReleasesDGX agent

arXiv:2607.03470v1 Announce Type: new Abstract: Synthesizing physically accurate mirror reflections remains a fundamental challenge for modern text-to-image diffusion models, which are increasingly cr

Piecewise Dynamic Diffusion Regularization for Reconstruction of Cardiac Cine MRI

ResearchDGX agent

arXiv:2607.03299v1 Announce Type: cross Abstract: Real-time cardiac cine MRI enables visualization of the beating heart during free breathing, but severe undersampling and motion make reconstruction h

PixelPilot: Scalable Vision-Language-Action Models for End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2607.04637v1 Announce Type: new Abstract: Vision-Language-Action Models (VLAs), which leverage the advanced reasoning capabilities of Vision-Language Models (VLMs), show promising generalization

PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space

ResearchDGX agent

arXiv:2607.05373v1 Announce Type: new Abstract: 3D reconstruction and generation are commonly tackled by separate paradigms: pixel-based regression for reconstruction, and latent diffusion for generat

Present but Not Remembered: Auditing How Frozen VLAs Encode, Deploy, and Steer Visual History

ApplicationsDGX agent

arXiv:2607.03372v1 Announce Type: new Abstract: A frozen vision-language-action model (VLA) receives recent observations at every decision step, yet prior work has focused on adding memory rather than

PreSIST: Vision-Language-Informed Object Persistence Prediction in Open-World Scenes

ResearchDGX agent

arXiv:2607.04057v1 Announce Type: new Abstract: Robots deployed over long periods must reason about environments that change over time. Existing long-term perception systems often address object chang

Pretreatment MRI reveals a latent, molecular-subtype-independent structural phenotype that organizes treatment trajectories and recurrence risk

ResearchDGX agent

arXiv:2607.02768v1 Announce Type: cross Abstract: Pathologic complete response and tumor shrinkage measure whether breast cancer responds to neoadjuvant therapy, but not whether that response was stru

PRIMA: Pre-training with Risk-integrated Image-Metadata Alignment for Medical Diagnosis via LLM

SafetyDGX agent

arXiv:2602.23297v2 Announce Type: replace Abstract: Medical diagnosis requires the effective synthesis of visual manifestations and clinical metadata. However, existing methods often treat metadata as

Prior Bias in Vision Language Models on UML Diagram Interpretation

Model ReleasesDGX agent

arXiv:2607.02853v1 Announce Type: new Abstract: Vision Language Models (VLMs) are increasingly applied to software engineering artifacts, especially UML class diagrams whose meaning depends on visual

PRISM3D: Probabilistic Refinement and Robust Initialization for Physically Consistent Scene Modeling under Extreme Motion Blur

Model ReleasesDGX agent

arXiv:2607.03855v1 Announce Type: new Abstract: We address the inverse problem of blind 3D scene reconstruction from extremely motion-blurred images, a scenario where traditional Structure-from-Motion

Privacy-Preserving Industrial Ergonomics: mmWave-Based Automated REBA Scoring and Pose Estimation

ResearchDGX agent

arXiv:2607.02611v1 Announce Type: new Abstract: Work-related Musculoskeletal Disorders (WMSDs) require continuous ergonomic assessments. While Rapid Entire Body Assessment (REBA) is a gold-standard ob

Probabilistic Robustness in Medical Image Classification

SafetyDGX agent

arXiv:2607.03797v1 Announce Type: new Abstract: Deep learning (DL) has shown strong performance in medical image classification, but its trustworthy deployment remains challenging in safety-critical c

Probe-EM: Targeted Neuron Tracing via Training-Free Semantic Verification

ResearchDGX agent

arXiv:2607.04696v1 Announce Type: new Abstract: Establishing large-scale, high-resolution neural connectivity maps is fundamental to elucidating the structural basis of brain function. However, when p

Probing Geospatial SSL Representations with Environmental Signals

Model ReleasesDGX agent

arXiv:2607.05207v1 Announce Type: new Abstract: Self-supervised learning (SSL) is designed to learn generic, transferable representations rather than representations optimized for a single task. Most

Probing Identity-Specific Motion Signatures: A Controlled Diagnostic Study

Model ReleasesDGX agent

arXiv:2607.03633v1 Announce Type: new Abstract: Identity recognition (e.g., person, animal re-identification) has traditionally relied heavily on static appearance cues. Yet motion--consistent, indivi

ProCon: Projection-Consistency Memory for Training-Free Anomaly Detection

ResearchDGX agent

arXiv:2607.04894v1 Announce Type: new Abstract: Memory-based anomaly detection is attractive because it localizes defects from normal images without training a decoder or synthesizing pseudo anomalies

Property-Constrained 3D Porous Media Reconstruction from 2D Images via Conditional Generative Adversarial Networks

ResearchDGX agent

arXiv:2607.02693v1 Announce Type: new Abstract: This study presents a conditional Generative Adversarial Network (cGAN) framework for generating 3D porous media volumes with controlled porosity, train

Provable Pruning for Efficient 3D Gaussian Splatting via Coresets

ResearchDGX agent

arXiv:2607.02721v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) enables high-quality real-time novel-view synthesis, but practical scenes often contain millions of Gaussians, making compr

ProxyUp: Training-Free Proxy-Conditioned Video Generation for Controllable Dynamics

SafetyDGX agent

arXiv:2607.03732v1 Announce Type: new Abstract: Precise control over complex dynamics remains challenging for modern video generative models, as text prompts alone often cannot specify physically plau

Purify then Guide: Rethinking Domain Generalization for Multimodal Face Anti-Spoofing

Model ReleasesDGX agent

arXiv:2505.09484v2 Announce Type: replace Abstract: Face Anti-Spoofing (FAS) is essential for the security of facial recognition systems in diverse scenarios such as payment processing and surveillanc

Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control

ResearchDGX agent

arXiv:2607.04978v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) underpin a growing family of latent world models for control from raw pixels, but every existing JEPA

QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding

SafetyDGX agent

arXiv:2607.04559v1 Announce Type: new Abstract: The performance of vision-language models (VLMs) in video understanding declines with increasing video duration, as video moments unrelated to the query

RADIANCE: Relative Adaptive Denoising with IP-Adapter for Novel Concept Enhancement

SafetyDGX agent

arXiv:2607.05088v1 Announce Type: new Abstract: Text-to-image (T2I) diffusion models have achieved striking progress but still struggle to synthesize rare concepts involving unusual attribute-object p

RAF: Reliability-Aware Fusion of Camera, LiDAR, and 4D RADAR for Robust 3D Object Detection in Adverse Weather

TutorialsDGX agent

arXiv:2607.04587v1 Announce Type: new Abstract: Robust 3D object detection in adverse weather conditions is challenging due to sensor limitations. Although combining complementary modalities such as L

RayTun3R: Online Camera Adaptation in 3D Foundation Models

Model ReleasesDGX agent

arXiv:2607.02711v1 Announce Type: new Abstract: Recent 3D foundation models, such as DUSt3R, MASt3R, VGGT, pi^3, and Depth Anything 3, provide strong feed-forward depth and pose estimates on pinhole i

REAL-OW: Rehearsal-free Open World Object Detection with Low-Rank Adaptation and Dual-Stage Objectness Modeling

ApplicationsDGX agent

arXiv:2607.03004v1 Announce Type: new Abstract: Open-World Object Detection (OWOD) requires detectors to identify previously unseen objects as unknown and incrementally incorporate them into the set o

Real-Time LiDAR Gaussian Splatting SLAM

ResearchDGX agent

arXiv:2607.04127v1 Announce Type: new Abstract: We present a real-time LiDAR-based framework for Gaussian Splatting SLAM that tightly couples fast G-ICP registration with spherical rasterization-based

ReCal3R: Reliability-Calibrated Learning Rates for Streaming 3D Reconstruction

SafetyDGX agent

arXiv:2607.05356v1 Announce Type: new Abstract: Streaming 3D reconstruction relies on a compact recurrent scene state to process long image streams in linear time and bounded memory. However, repeated

Reconstructing Rational Functions on Finite Abelian Groups with Higher Autocorrelations

ResearchDGX agent

arXiv:2503.21022v2 Announce Type: replace Abstract: The higher-order autocorrelations of integer-valued or rational-valued functions on finite Abelian groups appear naturally in X-ray crystallography,

Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation

Model ReleasesDGX agent

arXiv:2601.14788v2 Announce Type: replace Abstract: Diffusion models have seen widespread adoption for text-driven human motion generation and related tasks due to their impressive generative capabili

Reference-Induced Consensus for Selective Posed-Reference Visual Localization

SafetyDGX agent

arXiv:2607.04722v1 Announce Type: new Abstract: We present RIC-Loc (Reference-Induced Consensus localization), a scene-training-free posed-reference localizer that is SfM-point-map-free in its main es

Reliability-Aware CT-MRI Registration: A Quality Engineering Framework with Stability Analysis and Risk Classification

SafetyDGX agent

arXiv:2607.02585v1 Announce Type: new Abstract: Multimodal CT-MRI registration is central to image-guided radiotherapy, surgical navigation, and diagnostic workflows, but most pipelines report only ag

← Previous
1…5152535455…209
Next →