AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Safety

Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation

DGX agent

arXiv:2606.01565v1 Announce Type: cross Abstract: Vision-Language Navigation in Continuous Environments (VLN-CE) poses a formidable challenge for autonomous agents, requiring seamless integration of n

safetyarxiv-cs-cv
2 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Hierarchically Decoupled Mixture-of-Experts for Robust Traffic Sign Recognition in Complex Driving Scenarios

DGX agent

arXiv:2606.01822v1 Announce Type: new Abstract: Traffic sign detection is a fundamental component of environmental perception in autonomous driving and intelligent transportation systems. However, mos

model-releasesarxiv-cs-cv
2 Jun 2026
Research

HiGS: A Hierarchical Rendering Architecture for Real-Time 3D Gaussian Splatting

DGX agent

arXiv:2606.00352v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has become the standard for real-time novel view synthesis on commodity GPUs. Its pipeline ties spatial partitioning and ra

researcharxiv-cs-cv
2 Jun 2026
Local Ai

Hist2Style: Histogram-Guided Stylization with Bilateral Grids

DGX agent

arXiv:2606.01819v1 Announce Type: new Abstract: Photorealistic style transfer aims to match the color and tone of an input image to that of a style target while preserving the content and details of t

local-aiarxiv-cs-cv
2 Jun 2026
Safety

HiTokSR: A Coarse-to-Fine Tokenizer with Hierarchical Codebooks for High-Fidelity Real-World Image Super-Resolution

DGX agent

arXiv:2606.01157v1 Announce Type: new Abstract: Vector-quantized (VQ) generative models have shown promising results in real-world image super-resolution (Real-ISR). However, existing methods typicall

safetyarxiv-cs-cv
2 Jun 2026
Safety

HOLA: Holistic Multi-Modal Alignment for Open-Set 3D Recognition

DGX agent

arXiv:2606.01334v1 Announce Type: new Abstract: Open-set 3D recognition requires models that generalize to rare or unseen categories. Recent approaches address this by distilling language-vision knowl

safetyarxiv-cs-cv
2 Jun 2026
Research

Honey, I Shrunk the Arc de Triomphe!

DGX agent

arXiv:2606.02379v1 Announce Type: new Abstract: Metric scale monocular geometry estimation has seen significant progress through large-scale data aggregation, yet current foundation models suffer from

researcharxiv-cs-cv
2 Jun 2026
Research

HumanNOVA: Photorealistic, Universal and Rapid 3D Human Avatar Modeling from a Single Image

DGX agent

arXiv:2606.02573v1 Announce Type: new Abstract: In this paper, we present HumanNOVA, a photorealistic, universal, and rapid model for generating 3D human avatars from a single RGB image. Achieving bot

researcharxiv-cs-cv
2 Jun 2026
Agents

HyperDet: 3D Object Detection with Hyper 4D Radar Point Clouds

DGX agent

arXiv:2602.11554v3 Announce Type: replace-cross Abstract: How far can 3D object detection go using 4D radar alone? Despite offering weather-robust and velocity-aware sensing for autonomous perception,

agentsarxiv-cs-cv
2 Jun 2026
Research

HyperVQ: Enabling Hyperprior Entropy Modeling for VQ-Based Generative Image Compression

DGX agent

arXiv:2512.07192v2 Announce Type: replace Abstract: Vector Quantization (VQ) based generative image compression has achieved remarkable perceptual quality. However, existing VQ codecs suffer from two

researcharxiv-cs-cv
2 Jun 2026
Model Releases

hZACH-ViT: Curved Latent Geometry for Compact Vision Transformers in Low-Data Medical Imaging

DGX agent

arXiv:2606.00906v1 Announce Type: new Abstract: Compact Vision Transformers are attractive for medical imaging in low-data and resource-constrained settings, but most existing variants assume that Euc

model-releasesarxiv-cs-cv
2 Jun 2026
Research

iLRM: An Iterative Large 3D Reconstruction Model

DGX agent

arXiv:2507.23277v3 Announce Type: replace Abstract: Feed-forward 3D modeling has emerged as a promising approach for rapid and high-quality 3D reconstruction. In particular, directly generating explic

researcharxiv-cs-cv
2 Jun 2026
Research

IMA++: ISIC Archive Multi-Annotator Dermoscopic Skin Lesion Segmentation Dataset

DGX agent

arXiv:2512.21472v2 Announce Type: replace Abstract: Multi-annotator medical image segmentation is an important research problem, but requires annotated datasets that are expensive to collect. Dermosco

researcharxiv-cs-cv
2 Jun 2026
Research

Images as Tables: In-Context Learning with TabPFN for Low-Data Detection of AI-Generated Images

DGX agent

arXiv:2606.00872v1 Announce Type: new Abstract: AI-generated image detection is a moving-target problem: detectors trained on one generator often fail when a new generator appears, and only a few labe

researcharxiv-cs-cv
2 Jun 2026
Research

Improving Combined Detection and Classification of TEM Defects via Mask-Conditioned Latent Diffusion Augmentation

DGX agent

arXiv:2606.02532v1 Announce Type: new Abstract: Analyzing microstructural defects in transmission electron microscopy (TEM) images, particularly in irradiated metal alloys, is often limited by the ava

researcharxiv-cs-cv
2 Jun 2026
Research

Improving Visual Grounding in Remote Sensing via Cluster-Guided Refinement and Model Ensemble Voting

DGX agent

arXiv:2606.00556v1 Announce Type: new Abstract: Visual grounding aims to locate image regions that correspond to natural language descriptions and is a key component of interpretable vision systems. I

researcharxiv-cs-cv
2 Jun 2026
Research

Improving Visual Token Reduction via Rectifying Distortions for Efficient Multimodal LLM Inference

DGX agent

arXiv:2606.01711v1 Announce Type: new Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks, yet the quadratic computation

researcharxiv-cs-cv
2 Jun 2026
Model Releases

InsightVQA: High-Dimensional Emotion-Cognitive Visual Question Answering Benchmark

DGX agent

arXiv:2606.02171v1 Announce Type: new Abstract: Visual emotion understanding requires models not only to recognize emotional states, but also to why they arise and perform higher-level cognitive reaso

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

InstructSAM: Segment Any Instance with Any Instructions

DGX agent

arXiv:2605.26102v2 Announce Type: replace Abstract: In this paper, we introduce InstructSAM, a unified and streamlined framework designed for multi-instance segmentation under arbitrary instructions.

model-releasesarxiv-cs-cv
2 Jun 2026
Safety

Interpretable Modeling of Driver Attention Shifts with a Vision--Language Model

DGX agent

arXiv:2508.05852v2 Announce Type: replace Abstract: Driver gaze is commonly modeled as a spatial heatmap, but heatmaps alone are difficult for humans to interpret because they do not explain which roa

safetyarxiv-cs-cv
2 Jun 2026
Safety

IntraStyler: Intra-Domain Style Synthesis for Cross-Modality MRI Domain Adaptation

DGX agent

arXiv:2601.00212v2 Announce Type: replace Abstract: Segmentation of vestibular schwannoma and cochlea from T2 MRI is clinically important yet annotation-intensive. Domain adaptation (DA) has been wide

safetyarxiv-cs-cv
2 Jun 2026
Safety

KG-FairDiff: Knowledge Graph-Guided Prompt Refinement for Demographically Fair Text-to-Image Generation

DGX agent

arXiv:2606.01282v1 Announce Type: new Abstract: Text-to-Image (TTI) systems are now everyday infrastructure for journalism, education, advertising, and public communication, and the demographic and cu

safetyarxiv-cs-cv
2 Jun 2026
Research

LagerNVS: Latent Geometry for Fully Neural Real-time Novel View Synthesis

DGX agent

arXiv:2603.20176v3 Announce Type: replace Abstract: Recent work has shown that neural networks can perform 3D tasks such as Novel View Synthesis (NVS) without explicit 3D reconstruction. Even so, we a

researcharxiv-cs-cv
2 Jun 2026
Research

LastAct: Trajectory-Guided Latest-Activity Localization for Real-Time Smart-Home Activity Recognition

DGX agent

arXiv:2606.00260v1 Announce Type: new Abstract: Human Activity Recognition (HAR) from ambient sensors enables smart-home applications such as health monitoring and assisted living. In realistic deploy

researcharxiv-cs-cv
2 Jun 2026
Research

Learning Action-Conditional and Object-Centric Gaussian Splatting World Models for Rigid Objects

DGX agent

arXiv:2606.01950v1 Announce Type: cross Abstract: World models enable intelligent agents to predict the consequences of their actions on the environment. In this paper, we propose Multi Rigid Object G

researcharxiv-cs-cv
2 Jun 2026
Research

Learning Label-Efficient Interpretable Medical Image Diagnosis via Semi-supervised Hypergraph Concept Bottleneck Model

DGX agent

arXiv:2606.01698v1 Announce Type: new Abstract: Deep learning has revolutionized medical image analysis, delivering exceptional diagnostic accuracy across diverse applications. Yet, the lack of interp

researcharxiv-cs-cv
2 Jun 2026
Research

Learning Neural Deformation Representation for 4D Dynamic Shape Generation

DGX agent

arXiv:2606.01021v1 Announce Type: new Abstract: Recent developments in 3D shape representation opened new possibilities for generating detailed 3D shapes. Despite these advances, there are few studies

researcharxiv-cs-cv
2 Jun 2026
Research

Learning to Trim: End-to-End Causal Graph Pruning with Dynamic Anatomical Feature Banks for Medical VQA

DGX agent

arXiv:2603.26028v2 Announce Type: replace Abstract: Medical Visual Question Answering (MedVQA) models often exhibit limited generalization due to reliance on dataset-specific correlations, such as rec

researcharxiv-cs-cv
2 Jun 2026
Safety

LFA: Layer Feature Attention for Run-Time Introspection of 2D Object Detectors in Automated Driving

DGX agent

arXiv:2606.00372v1 Announce Type: new Abstract: Reliable object detection is critical for automated driving, yet even state-of-the-art detectors inevitably make errors that can compromise safety. Intr

safetyarxiv-cs-cv
2 Jun 2026
Model Releases

LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models

DGX agent

arXiv:2606.02535v1 Announce Type: new Abstract: Large-scale generative models have demonstrated remarkable capabilities across image generation and editing tasks. However, their performance in low-lev

model-releasesarxiv-cs-cv
2 Jun 2026
Research

LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation

DGX agent

arXiv:2606.02553v1 Announce Type: new Abstract: Autoregressive (AR) video diffusion enables variable-length synthesis, but long-horizon generation often suffers from accumulated errors and identity dr

researcharxiv-cs-cv
2 Jun 2026
Safety

Markerless Augmented Reality Registration for Surgical Guidance: A Multi-Anatomy Clinical Accuracy Study

DGX agent

arXiv:2511.02086v2 Announce Type: replace Abstract: Purpose: In this paper, we develop and clinically evaluate a depth-only, markerless augmented reality (AR) registration pipeline on a head-mounted d

safetyarxiv-cs-cv
2 Jun 2026
Model Releases

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

DGX agent

arXiv:2606.00793v1 Announce Type: new Abstract: Recent advancements in video-based world models have demonstrated an unprecedented ability to synthesize high-fidelity visual sequences. However, a fund

model-releasesarxiv-cs-cv
2 Jun 2026
Applications

MCPDepth: Omnidirectional Depth Estimation via Stereo Matching from Multi-Cylindrical Panoramas

DGX agent

arXiv:2408.01653v4 Announce Type: replace Abstract: Omnidirectional depth estimation presents a significant challenge due to the inherent distortions in panoramic images. Despite notable advancements,

applicationsarxiv-cs-cv
2 Jun 2026
Safety

Measurement Geometry and Design for Trustworthy Generative Inverse Problems

DGX agent

arXiv:2606.02309v1 Announce Type: cross Abstract: Generative models are increasingly used as priors for inverse problems, but their ability to produce realistic images creates a basic trust problem: a

safetyarxiv-cs-cv
2 Jun 2026
Model Releases

Med-URWKV{ag}: Toward Enhanced Pretrained Pure VRWKV Models for Medical Image Segmentation

DGX agent

arXiv:2506.10858v2 Announce Type: replace-cross Abstract: Medical image segmentation is a fundamental task in computer-aided diagnosis and treatment. Existing approaches based on CNNs, ViTs, Mamba, an

model-releasesarxiv-cs-cv
2 Jun 2026
Research

MipSLAM: Alias-Free Gaussian Splatting SLAM

DGX agent

arXiv:2603.06989v3 Announce Type: replace Abstract: This paper introduces MipSLAM, a frequency-aware 3D Gaussian Splatting (3DGS) SLAM framework capable of high-fidelity anti-aliased novel view synthe

researcharxiv-cs-cv
2 Jun 2026
Model Releases

MixerSENet: A Lightweight Framework for Efficient Hyperspectral Image Classification

DGX agent

arXiv:2606.01700v1 Announce Type: new Abstract: In this paper, a novel framework, MixerSENet, is introduced for hyperspectral image (HSI) classification, designed to address the challenges of computat

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue

DGX agent

arXiv:2606.00622v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) demonstrate remarkable visual understanding, yet their reliability in interactive settings is severely undermin

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

MMDG-Bench: A Benchmark for Multimodal Domain Generalization

DGX agent

arXiv:2606.00891v1 Announce Type: new Abstract: Multi-modal Domain Generalization (MMDG) seeks to leverage complementary modalities to enhance model robustness on unseen domains. Despite extensive pro

model-releasesarxiv-cs-cv
2 Jun 2026
Research

MMTalker: Multiresolution 3D Talking Head Synthesis with Multimodal Feature Fusion

DGX agent

arXiv:2604.02941v2 Announce Type: replace Abstract: Speech-driven three-dimensional (3D) facial animation synthesis aims to build a mapping from one-dimensional (1D) speech signals to time-varying 3D

researcharxiv-cs-cv
2 Jun 2026
Research

Modeling Robotics Dataset Construction as an Artifact-Based Build Process

DGX agent

arXiv:2606.00162v1 Announce Type: cross Abstract: Robotic systems generate large volumes of multimodal sensor data, but converting ROS bag recordings into machine learning datasets is often handled by

researcharxiv-cs-cv
2 Jun 2026
Research

MORPHOS: Autoregressive 4D Generation with Temporal Structured Latents

DGX agent

arXiv:2606.02491v1 Announce Type: new Abstract: We present MORPHOS, a novel autoregressive framework that generates dynamic 3D assets from videos across diverse representations, including meshes, 3D G

researcharxiv-cs-cv
2 Jun 2026
Model Releases

Motion-aware Event Suppression for Event Cameras

DGX agent

arXiv:2602.23204v3 Announce Type: replace Abstract: In this work, we introduce the first framework for Motion-aware Event Suppression, which learns to filter events triggered by IMOs and ego-motion in

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

MotionDreamer: Universal Skeletal Motion Generation for 3D Rigged Shapes

DGX agent

arXiv:2606.01518v1 Announce Type: new Abstract: Motion generation for rigged shapes is vital for scalable 4D asset production. However, template-based methods are limited by specific topologies and fa

model-releasesarxiv-cs-cv
2 Jun 2026
Research

MPMWorlds: Material-Point-Method Simulations for Inferring and Extrapolating Physical Dynamics

DGX agent

arXiv:2606.01538v1 Announce Type: cross Abstract: To study the ability to infer physical dynamics from videos and extrapolate them forward in time, we assemble a dataset of 2D Material Point Method (M

researcharxiv-cs-cv
2 Jun 2026
Model Releases

MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching

DGX agent

arXiv:2606.01985v1 Announce Type: new Abstract: Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing de

model-releasesarxiv-cs-cv
2 Jun 2026
Safety

Multi-modal Video Representation Alignment for Robust Self-supervised Driver Distraction Detection

DGX agent

arXiv:2606.02352v1 Announce Type: new Abstract: Robust self-supervised learning of multi-modal video representations is critical for real-world applications such as driver distraction detection, where

safetyarxiv-cs-cv
2 Jun 2026
← Previous
1…127128129130131…263
Next →