AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
28 Jul 2026

Layering Virtual Try-On

Model ReleasesDGX agent

arXiv:2607.22924v1 Announce Type: new Abstract: In the real world, fashion is about layering: adding a jacket over a shirt, or a sequence of adding and removing layers, rather than just a single-layer

LCMamNet: A Lightweight Cross-scale Mamba Network for Infrared Small Target Detection

Local AiDGX agent

arXiv:2607.24184v1 Announce Type: new Abstract: Infrared small target detection (IRSTD) is important for low-altitude perception, unmanned-system warning, and security monitoring. However, weak target

Learning-based Hierarchical Tracheal Anatomy Understanding from Sparse Surgical Demonstration Annotations for Ultrasound Robots

Local AiDGX agent

arXiv:2607.22789v1 Announce Type: cross Abstract: Tracheostomy requires precise localization of the tracheal incision site; however, conventional manual palpation is subjective and often unreliable, w


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Learning Dense 2D-3D Correspondence for X-ray-to-CT Registration of Knee Bones

TutorialsDGX agent

arXiv:2607.22803v1 Announce Type: cross Abstract: Recovering the 6-DoF pose of the knee bones from a plain radiograph, given the patient's segmented pre-operative CT, turns a routine low-dose image in

Learning Sampling Parameters for Diffusion Models

Model ReleasesDGX agent

arXiv:2607.23488v1 Announce Type: cross Abstract: Text-to-image diffusion models expose many inference-time sampling parameters, including prompts, negative prompts, classifier-free guidance scales, a

LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models

Local AiDGX agent

arXiv:2606.16586v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) remain unreliable on fine-grained visual perception, even when high-resolution inputs preserve the necessar

Long-Tailed Medical Image Classification

SafetyDGX agent

arXiv:2607.23883v1 Announce Type: new Abstract: In this paper, we examine the difficulties of using standard techniques for medical image classification due to long-tailed distributions (wherein rarer

LoTA-N2N: Local Trace Adaptation for Zero-Shot Self-Supervised Image Denoising

Model ReleasesDGX agent

arXiv:2607.24135v1 Announce Type: new Abstract: Single-image self-supervised denoising replaces unavailable clean targets with surrogate targets constructed from noisy observations. Its effectiveness

Low-light Image Enhancement via Multi-scale Attention combined with Fourier Transform

ApplicationsDGX agent

arXiv:2607.24002v1 Announce Type: new Abstract: Low-light image enhancement (LLIE) aims to improve image quality and clarity in diverse and demanding low-illumination environments. However, existing d

LowAux-RDNet: Low-Pass Residual Supervision with Scene-Balanced Real-World Training for Single-Image Reflection Removal

Model ReleasesDGX agent

arXiv:2607.22707v1 Announce Type: new Abstract: Single-image reflection removal aims to recover a clean transmission layer from one image captured through glass. We study an explicit decomposition pip

Mamba-CL: Optimizing Selective State Space Model in Null Space for Continual Learning

TutorialsDGX agent

arXiv:2411.15469v3 Announce Type: replace Abstract: Continual Learning (CL) aims to equip AI models with the ability to learn a sequence of tasks over time, without forgetting previously learned knowl

Manifold-Constrained Noise Optimization for Diverse Diffusion Sampling

ResearchDGX agent

arXiv:2607.23937v1 Announce Type: new Abstract: Few-step distilled diffusion models generate high-quality images quickly, but often lose per-prompt diversity, producing near-identical samples across r

Markerless Motion Capture in Routine Clinical Upper Limb Assessments: Validity and Insights Beyond Ordinal Scoring

ResearchDGX agent

arXiv:2607.23608v1 Announce Type: new Abstract: The Action Research Arm Test (ARAT) is a widely-used upper limb outcome measure in neurorehabilitation, but its ordinal scoring is subjective and suffer

MATS: A novel multi-modality multi-task learning framework for 3D perception in autonomous driving

Model ReleasesDGX agent

arXiv:2607.24224v1 Announce Type: new Abstract: Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autono

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

Model ReleasesDGX agent

arXiv:2607.24424v1 Announce Type: new Abstract: Vision-language models commonly project all tokens produced by a pretrained vision encoder into a large language model. However, final-layer features ca

MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation

ResearchDGX agent

arXiv:2607.23504v1 Announce Type: new Abstract: Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to maintain long-horizon visual history for trajectory consistency wh

Meshless Domain Randomization via Explicit Parameter Perturbation of 3D Gaussian Splatting

Model ReleasesDGX agent

arXiv:2607.22890v1 Announce Type: cross Abstract: Domain Randomization (DR) is a standard technique for closing the Sim-to-Real gap, yet traditional DR pipelines rely on classical computer graphics re

Metric Surface Reconstruction of Neurosurgical Scenes from Monocular Operating Microscope Images and Microscope Pose

ResearchDGX agent

arXiv:2607.22773v1 Announce Type: cross Abstract: Objective: We evaluated whether metric 3D geometry of neurosurgical operative exposure can be recovered from standard monocular operating-microscope i

MicroZoom: Structure-Preserving Detail Synthesis at Extreme Scale

TutorialsDGX agent

arXiv:2607.24729v1 Announce Type: new Abstract: We introduce MicroZoom, a generative framework for gigapixel image synthesis at the microscopic scale. Given a standard photograph and a sparse set of c

MIME: Multimodal Interactive Motion Encoder

SafetyDGX agent

arXiv:2607.22702v1 Announce Type: new Abstract: Text-motion representation learning has advanced rapidly, with growing interest in multi person interactions for animation, AR/VR, and embodied AI. Thes

Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding

Model ReleasesDGX agent

arXiv:2607.24407v1 Announce Type: new Abstract: Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and comp

MMOE: Modernizing Diffusion Transformers with Efficient Expert Design

Model ReleasesDGX agent

arXiv:2607.24665v1 Announce Type: new Abstract: Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capa

mmSimPrior: Learning Simulation Priors for Data-Efficient Real-World Generalizable Radar-Based Human Motion Reconstruction

ApplicationsDGX agent

arXiv:2607.22973v1 Announce Type: new Abstract: Millimeter-wave (mmWave) radar offers privacy-preserving and lighting-robust sensing for human motion reconstruction, but learning models that generaliz

MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving

AgentsDGX agent

arXiv:2607.23511v1 Announce Type: new Abstract: End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a

MSG-Loc: Multi-Label Likelihood-based Semantic Graph Matching for Object-Level Global Localization

ApplicationsDGX agent

arXiv:2512.03522v3 Announce Type: replace-cross Abstract: Robots are often required to localize in environments with unknown object classes and semantic ambiguity. However, when performing global loca

MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction

ResearchDGX agent

arXiv:2607.24436v1 Announce Type: new Abstract: High-fidelity 3D generative modeling increasingly relies on the latent diffusion paradigm, where the reconstruction quality of the underlying 3D VAE bec

Multi-Modal Object Re-Identification with Prompt-S6 and Semantic-Aware Knowledge Guidance

SafetyDGX agent

arXiv:2607.23451v1 Announce Type: new Abstract: Multi-modal object Re-Identification (ReID) aims to retrieve specific objects by integrating complementary information from multiple modalities. However

Multiview Multi-Person Human Mesh Recovery Under Large Scenes with Occlusions

Model ReleasesDGX agent

arXiv:2607.24302v1 Announce Type: new Abstract: Human mesh recovery (HMR) aims to recover 3D human meshes from images. Most existing HMR benchmarks and methods focus on either multi-person reconstruct

Mutual Modality Trust with Lightweight Reconstruction Regularization for Fine-grained Tire Pattern Recognition

SafetyDGX agent

arXiv:2607.23979v1 Announce Type: new Abstract: Visual tire recognition serves as a core supporting technique for vehicle safety monitoring, autonomous driving perception and automated automotive main

Neuromorphic Object Detection: An In-Depth Study and Future Directions

Model ReleasesDGX agent

arXiv:2607.23576v1 Announce Type: new Abstract: Conventional frame-based cameras face significant challenges in detecting objects under high-speed motion blur or in low-light environments. Neuromorphi

Nova3D: Code-Native Generation of Programmable 3D Assets

Model ReleasesDGX agent

arXiv:2607.22738v1 Announce Type: cross Abstract: Current 3D generative models mostly produce a final surface: a visually strong but largely opaque mesh. Interactive 3D worlds need more than a surface

NSL-SLAM: High-Fidelity Neural Structured-Light Depth for Practical SLAM and Reconstruction

Model ReleasesDGX agent

arXiv:2607.24495v1 Announce Type: new Abstract: Structured-light (SL) cameras power depth sensing in millions of devices, and recent neural SL decoding methods have substantially improved their depth

Occlusion-Point Reuse for Ray-Traced Ambient Occlusion and Shadow

ResearchDGX agent

arXiv:2607.23122v1 Announce Type: cross Abstract: Ambient occlusion (AO) and soft shadows are critical visibility cues for spatial perception in real-time rendering. Hardware ray tracing provides a di

OmniCache: Multidimensional Hierarchical Feature Caching For Diffusion Models

ResearchDGX agent

arXiv:2607.23844v1 Announce Type: new Abstract: High-resolution image and video diffusion models, including SD3, FLUX, and recent video diffusion transformers, have substantially improved generative q

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars

ResearchDGX agent

arXiv:2607.23023v1 Announce Type: new Abstract: Recent advances in diffusion-based generative models have enabled real-time audio-driven avatar generation and unified audio-visual synthesis, providing

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

HardwareDGX agent

arXiv:2607.23193v1 Announce Type: new Abstract: Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

SafetyDGX agent

arXiv:2607.23855v1 Announce Type: cross Abstract: Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Des

Operator learning for models of tear film breakup

ResearchDGX agent

arXiv:2601.08001v2 Announce Type: replace-cross Abstract: Tear film (TF) breakup is a key driver of understanding dry eye disease, yet estimating TF thickness and osmolarity from fluorescence (FL) ima

ORGAN: Object-Centric Representation Learning using Cycle Consistent Generative Adversarial Networks

ApplicationsDGX agent

arXiv:2603.02063v2 Announce Type: replace Abstract: Although data generation is often straightforward, extracting information from data is more difficult. Object-centric representation learning can ex

Out-of-Length Scene Text Recognition: A Two-Axis Diagnosis and a Training-Free Fix

Model ReleasesDGX agent

arXiv:2607.23194v1 Announce Type: new Abstract: Scene Text Recognition (STR) models are trained almost exclusively on word crops of at most 25 characters, yet real deployments (signage, product labels

Panda: Unsupervised Pelvic Anomaly Detection for Real-Time MR Imaging

ResearchDGX agent

arXiv:2607.24703v1 Announce Type: new Abstract: Female pelvic diseases remain an under researched area characterized by often delayed diagnosis. While pelvic MRI offers superior soft-tissue contrast f

Parallel Swin Transformer-Enhanced 3D MRI-to-CT Synthesis for MRI-Only Radiotherapy Planning

ResearchDGX agent

arXiv:2602.05387v2 Announce Type: replace Abstract: MRI provides superior soft tissue contrast without ionizing radiation; however, the absence of electron density information limits its direct use fo

Parameter-Efficient Adaptation of SAM3 for Prompt-Driven Surgical Concept Segmentation

Model ReleasesDGX agent

arXiv:2607.23694v1 Announce Type: new Abstract: Efficient surgical segmentation empowers clinical diagnosis, intraoperative monitoring, and downstream robotic pipelines for reconstruction and simulati

PathSelect: Sequential Token Selection for Whole Slide Pathology

TutorialsDGX agent

arXiv:2607.23631v1 Announce Type: new Abstract: Gigapixel Whole-Slide Images (WSIs) present a fundamental computational bottleneck for vision-language models (VLMs) due to extreme sequence lengths. Ex

PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models

ResearchDGX agent

arXiv:2607.22726v1 Announce Type: new Abstract: Despite advances in Video Large Language Models (VLLMs) that have displayed promising outcomes in video understanding, the redundancy in the long-durati

Perturbation-Aware Diffusion-Guided Hybrid Segmentation for Robust and Annotation-Efficient Plant Stress Phenotyping

SafetyDGX agent

arXiv:2607.23680v1 Announce Type: new Abstract: Semantic segmentation in agricultural imagery is often evaluated under in-domain protocols, yet practical deployment requires robustness to appearance p

Phenology-based learning framework for yield estimation and harvest forecasting of raspberry fruits

Model ReleasesDGX agent

arXiv:2411.00967v2 Announce Type: replace Abstract: The future of agriculture is intertwined with automation. Accurate fruit detection, yield estimation, and harvest time prediction are crucial for ef

PointCHR: Point Cloud Analysis via Curvature-Aware Hyperbolic Rectification

ResearchDGX agent

arXiv:2607.24052v1 Announce Type: new Abstract: High-curvature regions in 3D point clouds encapsulate critical fine-grained geometric semantics yet exhibit a distinct long-tail sparsity in their spati

PriSAR: 3D Geometric-Prior-Guided Diffusion for Parameter-Controlled SAR Image Generation

Model ReleasesDGX agent

arXiv:2607.22963v1 Announce Type: cross Abstract: Synthetic aperture radar (SAR) image generation can mitigate data scarcity, but controllablegeneration under sparse observation angles remains difficu

PRISM: Prompt Refinement via Image-grounded Self-rewarding Mechanism for Text-to-Image Generation

SafetyDGX agent

arXiv:2607.24353v1 Announce Type: new Abstract: Text-to-image generation models can synthesize high-quality images from natural language descriptions, but their performance remains highly sensitive to

pyALDIC: A Python Implementation of Augmented Lagrangian Digital Image Correlation with a GUI, Adaptive Meshing, and Mask-Aware Subset Splitting

ResearchDGX agent

arXiv:2607.22755v1 Announce Type: cross Abstract: pyALDIC is an open-source Python implementation of augmented Lagrangian digital image correlation (AL-DIC) for full-field displacement and strain meas

QueenVIS: Rethinking Image-Only Training for Video Instance Segmentation via Query Enrichment

Local AiDGX agent

arXiv:2607.24598v1 Announce Type: new Abstract: Video instance segmentation (VIS) requires models to detect, segment, and track object identities across frames, and most methods enforce temporal consi

Real-Time Human-Centric World Modeling for Upper-Body Human-Object Interaction

Local AiDGX agent

arXiv:2607.23517v1 Announce Type: new Abstract: We present a real-time human-centric world model for upper-body interactive generation, aiming to synthesize coherent local world dynamics centered on a

Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding

SafetyDGX agent

arXiv:2607.24199v1 Announce Type: new Abstract: Understanding and complying with traffic regulations is a safety-critical requirement for autonomous driving, yet remains challenging due to the diversi

Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention

Local AiDGX agent

arXiv:2511.12940v2 Announce Type: replace Abstract: Recent advancements in video generation has shifted from bidirectional models for short videos to autoregressive ones for ultra long video generatio

ReflexTrack: A Feedback-Driven Agent for Training-Free Referring Video Object Segmentation

Local AiDGX agent

arXiv:2607.24098v1 Announce Type: new Abstract: Referring video object segmentation (RVOS) requires segmenting a target specified by natural language throughout a video. Recent agentic approaches comb

Rethinking Expert Training for Model Merging with Prompt Learning

Model ReleasesDGX agent

arXiv:2607.24465v1 Announce Type: new Abstract: Model merging aims to combine multiple domain-specialized experts trained from a shared foundation model into a single multi-task model. Existing approa

RoadVGGT: Road-Structure-Aware Feed-Forward Road Surface Reconstruction

AgentsDGX agent

arXiv:2607.23758v1 Announce Type: new Abstract: Large-scale road surface reconstruction supports high-definition mapping, autonomous-driving perception, annotation, and simulation. Existing road-speci

Robust 6-DoF Object Pose Tracking with Built-In Recovery under Occlusions and Rapid Object Motions

SafetyDGX agent

arXiv:2607.23468v1 Announce Type: new Abstract: Real-time 6-DoF object pose tracking is essential for many robotics applications, and several approaches exist. Yet even today's approaches remain unrel

RODR: Riemannian Orthogonally Decoupled Regularization for Disentangled Manifold Representation

ResearchDGX agent

arXiv:2607.23958v1 Announce Type: new Abstract: Point cloud denoising is essentially a geometric recovery task that aims to reconstruct the intrinsic structure of a smooth 2D Riemannian manifold embed

← Previous
1…3031323334…209
Next →