AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Local Ai

Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG

DGX agent

arXiv:2606.21968v1 Announce Type: new Abstract: Vision-Language Models (VLMs) struggle as query-relevant objects become smaller. To address this, recent training-free approaches dynamically retrieve a

local-aiarxiv-cs-cv
23 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Tutorials

Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do

DGX agent

arXiv:2606.22565v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) has become a standard method for improving reasoning capabilities in large language models (LLMs) by eliciting step-by-step thi

tutorialsarxiv-cs-cv
23 Jun 2026
Model Releases

LUMINA-26: Low-Light Understanding for Modeling and Interpreting Night-time Actions

DGX agent

arXiv:2606.23118v1 Announce Type: new Abstract: Low-light human action recognition remains a challenging problem due to poor illumination, amplified noise, motion ambiguity, and diverse real-world sce

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models

DGX agent

arXiv:2509.23729v3 Announce Type: replace Abstract: Large Language Models (LLMs) with multimodal capabilities have revolutionized vision-language tasks, but their deployment often requires huge memory

model-releasesarxiv-cs-cv
23 Jun 2026
Research

LVQAC: Lattice Vector Quantization Coupled with Spatially Adaptive Companding for Efficient Learned Image Compression

DGX agent

arXiv:2304.12319v2 Announce Type: replace-cross Abstract: Recently, numerous end-to-end optimized image compression neural networks have been developed and proved themselves as leaders in rate-distort

researcharxiv-cs-cv
23 Jun 2026
Safety

Maintain Plasticity in Long-timescale Continual Test-time Adaptation

DGX agent

arXiv:2412.20034v2 Announce Type: replace Abstract: Continual test-time domain adaptation (CTTA) aims to adjust pre-trained source models to perform well over time across non-stationary target environ

safetyarxiv-cs-cv
23 Jun 2026
Local Ai

MambaADv2: Evolving Duality-enhanced State Space Model for Unsupervised Anomaly Detection

DGX agent

arXiv:2606.23126v1 Announce Type: new Abstract: While recent advancements in anomaly detection have demonstrated the efficacy of CNN- and Transformer-based approaches, these architectures face inheren

local-aiarxiv-cs-cv
23 Jun 2026
Model Releases

MammoExpert: Benchmarking Chain-of-Thought Reasoning in Mammography Diagnosis

DGX agent

arXiv:2606.21119v1 Announce Type: new Abstract: Mammography is an essential tool for breast cancer detection, with millions of examinations conducted annually. However, publicly available high-quality

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

MapReason-OSM: Can Vision-Language Models Make Graph-Verifiable Mobility Decisions from Street Maps ?

DGX agent

arXiv:2606.22597v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used to read maps for logistics, delivery, and accessible navigation, where the output is an actionable d

model-releasesarxiv-cs-cv
23 Jun 2026
Safety

MAPS: Multi-Anchor Projection Similarity for Joint Vision-Language Geo-Localization

DGX agent

arXiv:2606.22543v1 Announce Type: new Abstract: Humans localize places by integrating perceptual cues from vision with semantic reasoning from language, forming a scene understanding that is both intu

safetyarxiv-cs-cv
23 Jun 2026
Research

MaRS: Robust Out-of-Distribution Detection via Mahalanobis Residual Scoring

DGX agent

arXiv:2606.22649v1 Announce Type: new Abstract: Foundation models provide highly descriptive representations for medical images, yet their reliability degrades under distribution shifts arising from c

researcharxiv-cs-cv
23 Jun 2026
Model Releases

MEDLAYXPLAIN: Benchmarking the Expert-Lay Gap in Medical Vision-Language Models

DGX agent

arXiv:2606.21194v1 Announce Type: new Abstract: Medical Vision-Language Models (Med-VLMs) achieve strong expert-level performance, yet their ability to generate patient-accessible descriptions remains

model-releasesarxiv-cs-cv
23 Jun 2026
Research

MeGAS: Thermomechanical Dynamic Gaussian Splatting for Thermophysical Scene Editing

DGX agent

arXiv:2606.23455v1 Announce Type: new Abstract: Recent advances integrate physically grounded Newtonian dynamics with neural rendering frameworks, narrowing the gap between photorealistic scene recons

researcharxiv-cs-cv
23 Jun 2026
Safety

MemoryVAM: Integrating Memory into Video Action Model for Robot Manipulation

DGX agent

arXiv:2606.20679v1 Announce Type: cross Abstract: Video-world-model policies learn action-relevant representations by predicting future observations. However, they condition on only a short observatio

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

Mesh2GS: White-Box 3DGS Construction via Plenoptic Sampling

DGX agent

arXiv:2606.21898v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a promising method for high-quality, real-time 3D reconstruction. To associate 3DGS with mesh representati

model-releasesarxiv-cs-cv
23 Jun 2026
Research

MeshFlow: Mesh Generation with Equivariant Flow Matching

DGX agent

arXiv:2606.23489v1 Announce Type: cross Abstract: Meshes are among the most common 3D scene representations, but directly generating meshes is challenging because the representation contains important

researcharxiv-cs-cv
23 Jun 2026
Research

MILE: A Mechanically Isomorphic Exoskeleton Data Collection System with Fingertip Visuotactile Sensing for Dexterous Manipulation

DGX agent

arXiv:2512.00324v3 Announce Type: replace-cross Abstract: Imitation learning provides a promising approach to dexterous hand manipulation, but its effectiveness is limited by the lack of large-scale,

researcharxiv-cs-cv
23 Jun 2026
Research

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding

DGX agent

arXiv:2601.07298v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) excel at single-image understanding, they exhibit significantly degraded performance in multi-image r

researcharxiv-cs-cv
23 Jun 2026
Model Releases

Mirage: a Clean-Label Backdoor against LiDAR 3D Object Detection

DGX agent

arXiv:2606.20752v1 Announce Type: new Abstract: Deep neural network-based LiDAR 3D object detection serves as a critical perception component in safety-critical autonomous systems. However, recent stu

model-releasesarxiv-cs-cv
23 Jun 2026
Local Ai

MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents

DGX agent

arXiv:2606.20717v1 Announce Type: new Abstract: Multimodal Large Language Model (MLLM)-based web agents provide practical, high-precision solutions for visual browser automation; however, they inheren

local-aiarxiv-cs-cv
23 Jun 2026
Tutorials

MIRCaps: A Large-Scale Mixed-Domain Dataset with Image-Level and Region-Level Captions for Fine-Grained Vision-Language Learning

DGX agent

arXiv:2606.21419v1 Announce Type: new Abstract: Despite recent progress in Vision-Language Models (VLMs), mixed-domain image-caption datasets for both general-purpose and CCTV-based video surveillance

tutorialsarxiv-cs-cv
23 Jun 2026
Research

Mitigating Cross-Image Information Leakage in Multi-Image Understanding with Large Vision-Language Models

DGX agent

arXiv:2508.13744v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) exhibit strong performance on single-image tasks. However, their performance degrades significantly when handli

researcharxiv-cs-cv
23 Jun 2026
Model Releases

Mitigating Measurement-Induced Training Instability in Hybrid Quantum Neural Networks for Protein Classification

DGX agent

arXiv:2606.22551v1 Announce Type: cross Abstract: Hybrid Quantum Neural Network (QNN) classifiers produce logits as expectation values of quantum measurement operators. For standard Pauli measurements

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

MMGist: A Comprehensive Multimodal Benchmark for 2027

DGX agent

arXiv:2606.22437v1 Announce Type: new Abstract: We conduct a systematic study of 18 widely used vision-language benchmarks and identify three major issues: 1) many items do not rely on visual cues and

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos

DGX agent

arXiv:2603.14145v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have shown strong performance in visual and audio understanding when evaluated in isolation. However,

model-releasesarxiv-cs-cv
23 Jun 2026
Local Ai

Modular Diffusion Models for Structured Visual Recognition

DGX agent

arXiv:2606.22702v1 Announce Type: new Abstract: Traditional supervised methods for structured visual recognition tasks -- such as object detection, segmentation, and scene graph generation -- often pr

local-aiarxiv-cs-cv
23 Jun 2026
Model Releases

MoECodec: Image Compression for joint human and machine perception via Mixture-of-Experts

DGX agent

arXiv:2606.21033v1 Announce Type: cross Abstract: Image compression for machines calls for a unified codec that serves multiple downstream vision tasks. Existing approaches either adopt task-specific

model-releasesarxiv-cs-cv
23 Jun 2026
Research

MoFusion: A Framework for Denoising-Diffusion-based Motion Synthesis

DGX agent

arXiv:2212.04495v3 Announce Type: replace Abstract: Conventional methods for human motion synthesis are either deterministic or struggle with the trade-off between motion diversity and motion quality.

researcharxiv-cs-cv
23 Jun 2026
Model Releases

MOOZY: A Patient-First Foundation Model for Computational Pathology

DGX agent

arXiv:2603.27048v3 Announce Type: replace Abstract: Computational pathology needs whole-slide image (WSI) foundation models that transfer across diverse clinical tasks, yet current approaches remain l

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Morphology-Aware Multimodal Representation Learning for Insect Phylogenetic Reconstruction

DGX agent

arXiv:2606.22077v1 Announce Type: new Abstract: Morphological traits provide important evidence for phylogenetic reconstruction and evolutionary relationship analysis. Recent image-based approaches ha

model-releasesarxiv-cs-cv
23 Jun 2026
Safety

Motion-Aware Reinforcement Learning For Object Localization

DGX agent

arXiv:2606.21764v1 Announce Type: new Abstract: We present MARLNet (Motion-Aware Reinforcement Learning Network), a PPO-based bounding-box refinement agent that incorporates a constant-velocity motion

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning

DGX agent

arXiv:2606.23061v1 Announce Type: new Abstract: Motion instruction generation in cross-video comparison aims to produce corrective feedback that describes the differences between a query and a referen

model-releasesarxiv-cs-cv
23 Jun 2026
Research

MotionMAR: Multi-scale Auto-Regressive Human Motion Reconstruction from Sparse Observations

DGX agent

arXiv:2606.23000v1 Announce Type: new Abstract: Human motion follows a temporal hierarchical structure, transitioning from low-frequency global trajectories to high-frequency details. Inspired by the

researcharxiv-cs-cv
23 Jun 2026
Safety

MotionPyramid: Hierarchical Motion Representation and Residual Interfaces

DGX agent

arXiv:2606.20705v1 Announce Type: new Abstract: We ask whether the representational hierarchy seen in perception, from local primitives such as edges to higher level structures such as parts and objec

safetyarxiv-cs-cv
23 Jun 2026
Applications

MS-rPPG: Multi-spectral State Space Model for Remote Photoplethysmography in Driver Monitoring Systems

DGX agent

arXiv:2606.21115v1 Announce Type: new Abstract: Remote photoplethysmography (rPPG) is a camera-based technique for measuring physiological signals, particularly cardiac activity. From the remotely mea

applicationsarxiv-cs-cv
23 Jun 2026
Hardware

Multi-cancer detection using a computationally efficient CNN with transfer learning

DGX agent

arXiv:2606.22400v1 Announce Type: new Abstract: This study introduces a computationally efficient convolutional neural network (CNN) architecture enhanced with transfer learning for multi-cancer detec

hardwarearxiv-cs-cv
23 Jun 2026
Research

Multi-Depth Concept Extraction for Post-Hoc Vision Encoder Explanation

DGX agent

arXiv:2411.19700v5 Announce Type: replace Abstract: Explainable AI methods for vision models aim to identify the parts of the input that are important for the final prediction and subsequently relate

researcharxiv-cs-cv
23 Jun 2026
Research

Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation

DGX agent

arXiv:2606.22197v1 Announce Type: new Abstract: Dynamic 3D Gaussian splatting faces a fundamental tension between motion consistency and visual fidelity. Deformation-based approaches preserve temporal

researcharxiv-cs-cv
23 Jun 2026
Research

MultiMem: Measuring and Mitigating Memorization in Multi-Modal Contrastive Learninga

DGX agent

arXiv:2606.22220v1 Announce Type: new Abstract: Memorization in machine learning models enables high performance on rare in-distribution samples by capturing their atypical patterns. However, it also

researcharxiv-cs-cv
23 Jun 2026
Research

Multimodal Image Colorization: Quantifying the Impact of Text-Conditioned Guidance on Grayscale-to-Color Translation

DGX agent

arXiv:2606.20722v1 Announce Type: cross Abstract: Grayscale images are commonly found in historical photography restoration, medical imaging, and artistic media. However, automatically applying color

researcharxiv-cs-cv
23 Jun 2026
Research

muMatch: Foundation Models for Semi-supervised Learning and Domain Adaptation in EM

DGX agent

arXiv:2606.21605v1 Announce Type: new Abstract: Vision foundation models have substantially advanced computer vision, enabling state-of-the-art performance in zero- and few-shot settings. They have be

researcharxiv-cs-cv
23 Jun 2026
Research

MythraGen: Two-Stage Retrieval Augmented Art Generation Framework

DGX agent

arXiv:2606.22924v1 Announce Type: new Abstract: Text-to-image generation has seen rapid advancements, especially with the development of generative models. However, challenges remain in achieving high

researcharxiv-cs-cv
23 Jun 2026
Tutorials

Native space based pipelines outperform template space based pipeline in subcortical segmentation

DGX agent

arXiv:2606.21463v1 Announce Type: new Abstract: Accurate segmentation of subcortical regions is critical for neurosurgical planning and functional research. Most automated methods rely on template spa

tutorialsarxiv-cs-cv
23 Jun 2026
Safety

NegAS: Negative Label Guided Attention and Scoring for Out-of-Distribution Object Detection with Vision-Language Models

DGX agent

arXiv:2606.22537v1 Announce Type: new Abstract: Out-of-Distribution (OOD) detection is essential for ensuring the robustness and reliability of object detection systems deployed in safety-critical app

safetyarxiv-cs-cv
23 Jun 2026
Research

NeoJaundice-AI: Smartphone-Based Neonatal Jaundice Detection Using Dual-Input Deep Learning and Synthetic Augmentation

DGX agent

arXiv:2606.20689v1 Announce Type: new Abstract: Neonatal jaundice (hyperbilirubinemia) is one of the most common conditions affecting newborns worldwide, with India alone recording roughly 15 million

researcharxiv-cs-cv
23 Jun 2026
Local Ai

NeoLoc-68: End-to-end 68-point neonatal facial landmark localisation in neonatal clinical environments

DGX agent

arXiv:2606.20823v1 Announce Type: new Abstract: Facial landmark localisation is a prerequisite for developing automated, non-contact neonatal pain assessment methods. Clinicians use pain scales to jud

local-aiarxiv-cs-cv
23 Jun 2026
Safety

Neural Architecture Distributions: A New Paradigm for Stochastic Segmentation

DGX agent

arXiv:2606.21061v1 Announce Type: new Abstract: Stochastic segmentation seeks to represent multiple plausible masks for a single image, which is essential in safety- and quality-critical applications

safetyarxiv-cs-cv
23 Jun 2026
Research

NeuroShield: A Device-Agnostic Foundation Model for EEG Authentication

DGX agent

arXiv:2606.20673v1 Announce Type: cross Abstract: A central challenge in EEG authentication is that models are typically tied to the acquisition settings in which they are trained. In particular, vari

researcharxiv-cs-cv
23 Jun 2026
← Previous
1…99100101102103…263
Next →