AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
23 Jun 2026

Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG

Local AiDGX agent

arXiv:2606.21968v1 Announce Type: new Abstract: Vision-Language Models (VLMs) struggle as query-relevant objects become smaller. To address this, recent training-free approaches dynamically retrieve a

Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do

TutorialsDGX agent

arXiv:2606.22565v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) has become a standard method for improving reasoning capabilities in large language models (LLMs) by eliciting step-by-step thi

LUMINA-26: Low-Light Understanding for Modeling and Interpreting Night-time Actions

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.23118v1 Announce Type: new Abstract: Low-light human action recognition remains a challenging problem due to poor illumination, amplified noise, motion ambiguity, and diverse real-world sce

LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2509.23729v3 Announce Type: replace Abstract: Large Language Models (LLMs) with multimodal capabilities have revolutionized vision-language tasks, but their deployment often requires huge memory

LVQAC: Lattice Vector Quantization Coupled with Spatially Adaptive Companding for Efficient Learned Image Compression

ResearchDGX agent

arXiv:2304.12319v2 Announce Type: replace-cross Abstract: Recently, numerous end-to-end optimized image compression neural networks have been developed and proved themselves as leaders in rate-distort

Maintain Plasticity in Long-timescale Continual Test-time Adaptation

SafetyDGX agent

arXiv:2412.20034v2 Announce Type: replace Abstract: Continual test-time domain adaptation (CTTA) aims to adjust pre-trained source models to perform well over time across non-stationary target environ

MambaADv2: Evolving Duality-enhanced State Space Model for Unsupervised Anomaly Detection

Local AiDGX agent

arXiv:2606.23126v1 Announce Type: new Abstract: While recent advancements in anomaly detection have demonstrated the efficacy of CNN- and Transformer-based approaches, these architectures face inheren

MammoExpert: Benchmarking Chain-of-Thought Reasoning in Mammography Diagnosis

Model ReleasesDGX agent

arXiv:2606.21119v1 Announce Type: new Abstract: Mammography is an essential tool for breast cancer detection, with millions of examinations conducted annually. However, publicly available high-quality

MapReason-OSM: Can Vision-Language Models Make Graph-Verifiable Mobility Decisions from Street Maps ?

Model ReleasesDGX agent

arXiv:2606.22597v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used to read maps for logistics, delivery, and accessible navigation, where the output is an actionable d

MAPS: Multi-Anchor Projection Similarity for Joint Vision-Language Geo-Localization

SafetyDGX agent

arXiv:2606.22543v1 Announce Type: new Abstract: Humans localize places by integrating perceptual cues from vision with semantic reasoning from language, forming a scene understanding that is both intu

MaRS: Robust Out-of-Distribution Detection via Mahalanobis Residual Scoring

ResearchDGX agent

arXiv:2606.22649v1 Announce Type: new Abstract: Foundation models provide highly descriptive representations for medical images, yet their reliability degrades under distribution shifts arising from c

MEDLAYXPLAIN: Benchmarking the Expert-Lay Gap in Medical Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.21194v1 Announce Type: new Abstract: Medical Vision-Language Models (Med-VLMs) achieve strong expert-level performance, yet their ability to generate patient-accessible descriptions remains

MeGAS: Thermomechanical Dynamic Gaussian Splatting for Thermophysical Scene Editing

ResearchDGX agent

arXiv:2606.23455v1 Announce Type: new Abstract: Recent advances integrate physically grounded Newtonian dynamics with neural rendering frameworks, narrowing the gap between photorealistic scene recons

MemoryVAM: Integrating Memory into Video Action Model for Robot Manipulation

SafetyDGX agent

arXiv:2606.20679v1 Announce Type: cross Abstract: Video-world-model policies learn action-relevant representations by predicting future observations. However, they condition on only a short observatio

Mesh2GS: White-Box 3DGS Construction via Plenoptic Sampling

Model ReleasesDGX agent

arXiv:2606.21898v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a promising method for high-quality, real-time 3D reconstruction. To associate 3DGS with mesh representati

MeshFlow: Mesh Generation with Equivariant Flow Matching

ResearchDGX agent

arXiv:2606.23489v1 Announce Type: cross Abstract: Meshes are among the most common 3D scene representations, but directly generating meshes is challenging because the representation contains important

MILE: A Mechanically Isomorphic Exoskeleton Data Collection System with Fingertip Visuotactile Sensing for Dexterous Manipulation

ResearchDGX agent

arXiv:2512.00324v3 Announce Type: replace-cross Abstract: Imitation learning provides a promising approach to dexterous hand manipulation, but its effectiveness is limited by the lack of large-scale,

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding

ResearchDGX agent

arXiv:2601.07298v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) excel at single-image understanding, they exhibit significantly degraded performance in multi-image r

Mirage: a Clean-Label Backdoor against LiDAR 3D Object Detection

Model ReleasesDGX agent

arXiv:2606.20752v1 Announce Type: new Abstract: Deep neural network-based LiDAR 3D object detection serves as a critical perception component in safety-critical autonomous systems. However, recent stu

MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents

Local AiDGX agent

arXiv:2606.20717v1 Announce Type: new Abstract: Multimodal Large Language Model (MLLM)-based web agents provide practical, high-precision solutions for visual browser automation; however, they inheren

MIRCaps: A Large-Scale Mixed-Domain Dataset with Image-Level and Region-Level Captions for Fine-Grained Vision-Language Learning

TutorialsDGX agent

arXiv:2606.21419v1 Announce Type: new Abstract: Despite recent progress in Vision-Language Models (VLMs), mixed-domain image-caption datasets for both general-purpose and CCTV-based video surveillance

Mitigating Cross-Image Information Leakage in Multi-Image Understanding with Large Vision-Language Models

ResearchDGX agent

arXiv:2508.13744v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) exhibit strong performance on single-image tasks. However, their performance degrades significantly when handli

Mitigating Measurement-Induced Training Instability in Hybrid Quantum Neural Networks for Protein Classification

Model ReleasesDGX agent

arXiv:2606.22551v1 Announce Type: cross Abstract: Hybrid Quantum Neural Network (QNN) classifiers produce logits as expectation values of quantum measurement operators. For standard Pauli measurements

MMGist: A Comprehensive Multimodal Benchmark for 2027

Model ReleasesDGX agent

arXiv:2606.22437v1 Announce Type: new Abstract: We conduct a systematic study of 18 widely used vision-language benchmarks and identify three major issues: 1) many items do not rely on visual cues and

MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos

Model ReleasesDGX agent

arXiv:2603.14145v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have shown strong performance in visual and audio understanding when evaluated in isolation. However,

Modular Diffusion Models for Structured Visual Recognition

Local AiDGX agent

arXiv:2606.22702v1 Announce Type: new Abstract: Traditional supervised methods for structured visual recognition tasks -- such as object detection, segmentation, and scene graph generation -- often pr

MoECodec: Image Compression for joint human and machine perception via Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2606.21033v1 Announce Type: cross Abstract: Image compression for machines calls for a unified codec that serves multiple downstream vision tasks. Existing approaches either adopt task-specific

MoFusion: A Framework for Denoising-Diffusion-based Motion Synthesis

ResearchDGX agent

arXiv:2212.04495v3 Announce Type: replace Abstract: Conventional methods for human motion synthesis are either deterministic or struggle with the trade-off between motion diversity and motion quality.

MOOZY: A Patient-First Foundation Model for Computational Pathology

Model ReleasesDGX agent

arXiv:2603.27048v3 Announce Type: replace Abstract: Computational pathology needs whole-slide image (WSI) foundation models that transfer across diverse clinical tasks, yet current approaches remain l

Morphology-Aware Multimodal Representation Learning for Insect Phylogenetic Reconstruction

Model ReleasesDGX agent

arXiv:2606.22077v1 Announce Type: new Abstract: Morphological traits provide important evidence for phylogenetic reconstruction and evolutionary relationship analysis. Recent image-based approaches ha

Motion-Aware Reinforcement Learning For Object Localization

SafetyDGX agent

arXiv:2606.21764v1 Announce Type: new Abstract: We present MARLNet (Motion-Aware Reinforcement Learning Network), a PPO-based bounding-box refinement agent that incorporates a constant-velocity motion

MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning

Model ReleasesDGX agent

arXiv:2606.23061v1 Announce Type: new Abstract: Motion instruction generation in cross-video comparison aims to produce corrective feedback that describes the differences between a query and a referen

MotionMAR: Multi-scale Auto-Regressive Human Motion Reconstruction from Sparse Observations

ResearchDGX agent

arXiv:2606.23000v1 Announce Type: new Abstract: Human motion follows a temporal hierarchical structure, transitioning from low-frequency global trajectories to high-frequency details. Inspired by the

MotionPyramid: Hierarchical Motion Representation and Residual Interfaces

SafetyDGX agent

arXiv:2606.20705v1 Announce Type: new Abstract: We ask whether the representational hierarchy seen in perception, from local primitives such as edges to higher level structures such as parts and objec

MS-rPPG: Multi-spectral State Space Model for Remote Photoplethysmography in Driver Monitoring Systems

ApplicationsDGX agent

arXiv:2606.21115v1 Announce Type: new Abstract: Remote photoplethysmography (rPPG) is a camera-based technique for measuring physiological signals, particularly cardiac activity. From the remotely mea

Multi-cancer detection using a computationally efficient CNN with transfer learning

HardwareDGX agent

arXiv:2606.22400v1 Announce Type: new Abstract: This study introduces a computationally efficient convolutional neural network (CNN) architecture enhanced with transfer learning for multi-cancer detec

Multi-Depth Concept Extraction for Post-Hoc Vision Encoder Explanation

ResearchDGX agent

arXiv:2411.19700v5 Announce Type: replace Abstract: Explainable AI methods for vision models aim to identify the parts of the input that are important for the final prediction and subsequently relate

Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation

ResearchDGX agent

arXiv:2606.22197v1 Announce Type: new Abstract: Dynamic 3D Gaussian splatting faces a fundamental tension between motion consistency and visual fidelity. Deformation-based approaches preserve temporal

MultiMem: Measuring and Mitigating Memorization in Multi-Modal Contrastive Learninga

ResearchDGX agent

arXiv:2606.22220v1 Announce Type: new Abstract: Memorization in machine learning models enables high performance on rare in-distribution samples by capturing their atypical patterns. However, it also

Multimodal Image Colorization: Quantifying the Impact of Text-Conditioned Guidance on Grayscale-to-Color Translation

ResearchDGX agent

arXiv:2606.20722v1 Announce Type: cross Abstract: Grayscale images are commonly found in historical photography restoration, medical imaging, and artistic media. However, automatically applying color

muMatch: Foundation Models for Semi-supervised Learning and Domain Adaptation in EM

ResearchDGX agent

arXiv:2606.21605v1 Announce Type: new Abstract: Vision foundation models have substantially advanced computer vision, enabling state-of-the-art performance in zero- and few-shot settings. They have be

MythraGen: Two-Stage Retrieval Augmented Art Generation Framework

ResearchDGX agent

arXiv:2606.22924v1 Announce Type: new Abstract: Text-to-image generation has seen rapid advancements, especially with the development of generative models. However, challenges remain in achieving high

Native space based pipelines outperform template space based pipeline in subcortical segmentation

TutorialsDGX agent

arXiv:2606.21463v1 Announce Type: new Abstract: Accurate segmentation of subcortical regions is critical for neurosurgical planning and functional research. Most automated methods rely on template spa

NegAS: Negative Label Guided Attention and Scoring for Out-of-Distribution Object Detection with Vision-Language Models

SafetyDGX agent

arXiv:2606.22537v1 Announce Type: new Abstract: Out-of-Distribution (OOD) detection is essential for ensuring the robustness and reliability of object detection systems deployed in safety-critical app

NeoJaundice-AI: Smartphone-Based Neonatal Jaundice Detection Using Dual-Input Deep Learning and Synthetic Augmentation

ResearchDGX agent

arXiv:2606.20689v1 Announce Type: new Abstract: Neonatal jaundice (hyperbilirubinemia) is one of the most common conditions affecting newborns worldwide, with India alone recording roughly 15 million

NeoLoc-68: End-to-end 68-point neonatal facial landmark localisation in neonatal clinical environments

Local AiDGX agent

arXiv:2606.20823v1 Announce Type: new Abstract: Facial landmark localisation is a prerequisite for developing automated, non-contact neonatal pain assessment methods. Clinicians use pain scales to jud

Neural Architecture Distributions: A New Paradigm for Stochastic Segmentation

SafetyDGX agent

arXiv:2606.21061v1 Announce Type: new Abstract: Stochastic segmentation seeks to represent multiple plausible masks for a single image, which is essential in safety- and quality-critical applications

NeuroShield: A Device-Agnostic Foundation Model for EEG Authentication

ResearchDGX agent

arXiv:2606.20673v1 Announce Type: cross Abstract: A central challenge in EEG authentication is that models are typically tied to the acquisition settings in which they are trained. In particular, vari

NGPS: Structure-Preserving Self-Supervised Denoising via Neighbor-Guided Patch Sampling

TutorialsDGX agent

arXiv:2606.23200v1 Announce Type: cross Abstract: Neighboring-slice self-supervised denoising is attractive for volumetric medical imaging, yet inter-slice misalignment breaks anatomical correspondenc

NoduLoCC2026: Lung Nodule Localization and Classification Contest from Chest X-Ray Images

ResearchDGX agent

arXiv:2606.21290v1 Announce Type: new Abstract: We propose NoduLoCC2026, a challenge on lung nodule detection and localization in chest X-ray images. We have provided a dataset for both tasks and rece

Non-line-of-sight imaging with arbitrary relay surface geometries via 3D Gaussian Transient Rendering

AgentsDGX agent

arXiv:2606.21270v1 Announce Type: cross Abstract: Imaging objects hidden outside the direct line of sight expands the effective field of view and is critical for applications such as autonomous drivin

Null-Space Diffusion Distillation Unlocks Speed, Fidelity and Realism in Lensless Imaging

ResearchDGX agent

arXiv:2511.12024v3 Announce Type: replace Abstract: Lensless imaging reconstructs scenes from highly multiplexed measurements, resulting in a severely ill-posed inverse problem. In this work, we ident

NullFlow: One-Step Generative Reconstruction

TutorialsDGX agent

arXiv:2606.22696v1 Announce Type: new Abstract: We propose NullFlow, a principled framework for one-step generative image reconstruction. Our key idea is to confine the generative flow to a measuremen

Object-Centric Dataset Resources for Constrained-Data Image Generation and Augmentation

ResearchDGX agent

arXiv:2606.21113v1 Announce Type: new Abstract: Object-centric image generation is important in settings with few labeled examples, including pedestrian analysis in smart-city scenes, traffic-sign ins

Ocean4D: Generative Underwater 4D Reconstruction via Medium-Aware Video Diffusion

ResearchDGX agent

arXiv:2606.23298v1 Announce Type: new Abstract: Underwater 4D reconstruction remains challenging due to the coupling between degraded light transport in participating media and dynamic water variation

Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion

ResearchDGX agent

arXiv:2606.21135v1 Announce Type: new Abstract: Human motion generation has been widely studied across diverse input modalities, text, music, and video, and recent efforts have unified these into sing

OmniNWM: Omniscient Driving Navigation World Models

SafetyDGX agent

arXiv:2510.18313v5 Announce Type: replace Abstract: Autonomous driving world models are expected to work effectively across three core dimensions: state, action, and reward. However, existing methods

OmniSpace: Efficient Geometry Awareness for Autonomous Vehicles MLLMs

AgentsDGX agent

arXiv:2606.22617v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable performance on 2D visual tasks, yet enhancing their spatial intelligence for real-worl

On-Manifold Variational Learning with Heat-Kernel Priors

ResearchDGX agent

arXiv:2606.18658v2 Announce Type: replace Abstract: Learning unsupervised representations of medical imaging cohorts can reveal clinically meaningful prototypes without expert labels, which are often

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception

SafetyDGX agent

arXiv:2606.20764v1 Announce Type: new Abstract: Reliable spatial decision automation, such as autonomous driving and maritime surveillance, critically depends on robust visual perception. However, rea

← Previous
1…7980818283…211
Next →