AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Model Releases

iDocV2: Leveraging Self-Supervision and Open-Set Detection for Improving Pattern Spotting in Historical Documents

DGX agent

arXiv:2604.16726v1 Announce Type: new Abstract: Considering the imminent massification of digital books, it has become critical to facilitate searching collections through graphical patterns. Current

model-releasesarxiv-cs-cv
21 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

IMA-MoE: An Interpretable Modality-Aware Mixture-of-Experts Framework for Characterizing the Neurobiological Signatures of Binge Eating Disorder

DGX agent

arXiv:2604.17028v1 Announce Type: new Abstract: Binge eating disorder (BED) is the most prevalent eating disorder. However, current diagnostic frameworks remain largely grounded in symptom-based crite

researcharxiv-cs-cv
21 Apr 2026
Safety

Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback

DGX agent

arXiv:2412.02617v2 Announce Type: replace-cross Abstract: Large text-to-video models hold immense potential for a wide range of downstream applications. However, they struggle to accurately depict dyn

safetyarxiv-cs-cv
21 Apr 2026
Research

Improving Radio Interferometry Imaging by Explicitly Modeling Cross-Domain Consistency in Reconstruction

DGX agent

arXiv:2604.16794v1 Announce Type: new Abstract: Radio astronomy plays a crucial role in understanding the universe, particularly within the realm of non-thermal astrophysics. Images of celestial objec

researcharxiv-cs-cv
21 Apr 2026
Research

IncepDeHazeGAN: Novel Satellite Image Dehazing

DGX agent

arXiv:2604.16609v1 Announce Type: new Abstract: Dehazing is a technique in computer vision for enhancing the visual quality of images captured in cloudy or foggy conditions. Dehazing helps to recover

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Incoherent Deformation, Not Capacity: Diagnosing and Mitigating Overfitting in Dynamic Gaussian Splatting

DGX agent

arXiv:2604.16747v1 Announce Type: new Abstract: Dynamic 3D Gaussian Splatting methods achieve strong training-view PSNR on monocular video but generalize poorly: on the D-NeRF benchmark we measure an

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

IncreFA: Breaking the Static Wall of Generative Model Attribution

DGX agent

arXiv:2604.17736v1 Announce Type: new Abstract: As AI generative models evolve at unprecedented speed, image attribution has become a moving target. New diffusion, adversarial and autoregressive gener

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Inductive Convolution Nuclear Norm Minimization for Tensor Completion with Arbitrary Sampling

DGX agent

arXiv:2604.17001v1 Announce Type: new Abstract: The recently established Convolution Nuclear Norm Minimization (CNNM) addresses the problem of extit{tensor completion with arbitrary sampling} (TCAS),

researcharxiv-cs-cv
21 Apr 2026
Research

Inference-Time Temporal Probability Smoothing for Stable Video Segmentation with SAM2 under Weak Prompts

DGX agent

arXiv:2604.17115v1 Announce Type: new Abstract: Interactive video segmentation models such as SAM2 have demonstrated strong generalization across diverse visual domains. However, under weak user super

researcharxiv-cs-cv
21 Apr 2026
Safety

Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception

DGX agent

arXiv:2604.17651v1 Announce Type: new Abstract: World models, generative AI systems that simulate how environments evolve, are transforming autonomous driving, yet all existing approaches adopt an ego

safetyarxiv-cs-cv
21 Apr 2026
Research

Instant Colorization of Gaussian Splats

DGX agent

arXiv:2604.17155v1 Announce Type: new Abstract: Gaussian Splatting has recently become one of the most popular frameworks for photorealistic 3D scene reconstruction and rendering. While current raster

researcharxiv-cs-cv
21 Apr 2026
Safety

Instinct vs. Reflection: Unifying Token and Verbalized Confidence in Multimodal Large Models

DGX agent

arXiv:2604.17274v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in various perception and reasoning tasks. Despite this success, ens

safetyarxiv-cs-cv
21 Apr 2026
Agents

Instruction-as-State: Environment-Guided and State-Conditioned Semantic Understanding for Embodied Navigation

DGX agent

arXiv:2604.18223v1 Announce Type: new Abstract: Vision-and-Language Navigation requires agents to follow natural-language instructions in visually changing environments. A central challenge is the dyn

agentsarxiv-cs-cv
21 Apr 2026
Applications

Integrating Feature Selection and Machine Learning for Nitrogen Assessment in Grapevine Leaves using In-Field Hyperspectral Imaging

DGX agent

arXiv:2507.17869v3 Announce Type: replace-cross Abstract: Nitrogen (N) is one of the most critical nutrients in winegrape production, influencing vine vigor, fruit composition, and wine quality. Becau

applicationsarxiv-cs-cv
21 Apr 2026
Model Releases

INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval

DGX agent

arXiv:2604.18051v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) is a challenging image retrieval paradigm that enables to retrieve target images based on multimodal queries consisting o

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts

DGX agent

arXiv:2509.10813v3 Announce Type: replace Abstract: The advancement of Embodied AI heavily relies on large-scale, simulatable 3D scene datasets characterized by scene diversity and realistic layouts.

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Is SAM3 ready for pathology segmentation?

DGX agent

arXiv:2604.18225v1 Announce Type: new Abstract: Is Segment Anything Model 3 (SAM3) capable in segmenting Any Pathology Images? Digital pathology segmentation spans tissue-level and nuclei-level scales

researcharxiv-cs-cv
21 Apr 2026
Research

Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models

DGX agent

arXiv:2512.02636v3 Announce Type: replace-cross Abstract: Log-likelihood evaluation enables important capabilities in generative models, including model comparison, certain fine-tuning objectives, and

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription

DGX agent

arXiv:2502.20295v2 Announce Type: replace-cross Abstract: Handwriting text recognition (HTR) remains a challenging task. Existing approaches require fine-tuning on labeled data, which is impractical t

model-releasesarxiv-cs-cv
21 Apr 2026
Research

KaLDeX: Kalman Filter based Linear Deformable Cross Attention for Retina Vessel Segmentation

DGX agent

arXiv:2410.21160v2 Announce Type: replace-cross Abstract: Background and Objective: In the realm of ophthalmic imaging, accurate vascular segmentation is paramount for diagnosing and managing various

researcharxiv-cs-cv
21 Apr 2026
Model Releases

KIRA: Knowledge-Intensive Image Retrieval and Reasoning Architecture for Specialized Visual Domains

DGX agent

arXiv:2604.16915v1 Announce Type: new Abstract: Retrieval augmented generation (RAG) has transformed text based question answering, yet its extension to visual domains remains hindered by fundamental

model-releasesarxiv-cs-cv
21 Apr 2026
Applications

LAGS: Low-Altitude Gaussian Splatting with Groupwise Heterogeneous Graph Learning

DGX agent

arXiv:2604.16910v1 Announce Type: new Abstract: Low-altitude Gaussian splatting (LAGS) facilitates 3D scene reconstruction by aggregating aerial images from distributed drones. However, as LAGS priori

applicationsarxiv-cs-cv
21 Apr 2026
Research

Latent-Compressed Variational Autoencoder for Video Diffusion Models

DGX agent

arXiv:2604.16479v1 Announce Type: new Abstract: Video variational autoencoders (VAEs) used in latent diffusion models typically require a sufficiently large number of latent channels to ensure high-qu

researcharxiv-cs-cv
21 Apr 2026
Model Releases

LayerCache: Exploiting Layer-wise Velocity Heterogeneity for Efficient Flow Matching Inference

DGX agent

arXiv:2604.16492v1 Announce Type: new Abstract: Flow Matching models achieve state-of-the-art image generation quality but incur substantial inference cost due to iterative denoising through large Tra

model-releasesarxiv-cs-cv
21 Apr 2026
Research

LBFTI: Layer-Based Facial Template Inversion for Identity-Preserving Fine-Grained Face Reconstruction

DGX agent

arXiv:2604.18358v1 Announce Type: new Abstract: In face recognition systems, facial templates are widely adopted for identity authentication due to their compliance with the data minimization principl

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Learned Nonlocal Feature Matching and Filtering for RAW Image Denoising

DGX agent

arXiv:2604.17453v1 Announce Type: cross Abstract: Being one of the oldest and most basic problems in image processing, image denoising has seen a resurgence spurred by rapid advances in deep learning.

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

LiquidTAD: An Efficient Method for Temporal Action Detection via Liquid Neural Dynamics

DGX agent

arXiv:2604.18274v1 Announce Type: new Abstract: Temporal Action Detection (TAD) in untrimmed videos is currently dominated by Transformer-based architectures. While high-performing, their quadratic co

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length

DGX agent

arXiv:2512.04677v5 Announce Type: replace Abstract: Audio-driven avatar interaction demands real-time, streaming, and infinite-length generation -- capabilities fundamentally at odds with the sequenti

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing

DGX agent

arXiv:2604.17021v1 Announce Type: new Abstract: Video editing aims to modify input videos according to user intent. Recently, end-to-end training methods have garnered widespread attention, constructi

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

LLaMA-XR: A Novel Framework for Radiology Report Generation using LLaMA and QLoRA Fine Tuning

DGX agent

arXiv:2506.03178v2 Announce Type: replace-cross Abstract: Automated radiology report generation holds significant potential to reduce radiologists' workload and enhance diagnostic accuracy. However, g

model-releasesarxiv-cs-cv
21 Apr 2026
Research

LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding

DGX agent

arXiv:2501.05067v3 Announce Type: replace Abstract: In this paper, we introduce LLaVA-Octopus, a novel video multimodal large language model. LLaVA-Octopus adaptively weights features from different v

researcharxiv-cs-cv
21 Apr 2026
Agents

LLM as a Tool, Not an Agent: Code-Mined Tree Transformations for Neural Architecture Search

DGX agent

arXiv:2604.16555v1 Announce Type: cross Abstract: Neural Architecture Search (NAS) aims to automatically discover high-performing deep neural network (DNN) architectures. However, conventional algorit

agentsarxiv-cs-cv
21 Apr 2026
Local Ai

LOD-Net: Locality-Aware 3D Object Detection Using Multi-Scale Transformer Network

DGX agent

arXiv:2604.16696v1 Announce Type: new Abstract: 3D object detection in point cloud data remains a challenging task due to the sparsity and lack of global structure inherent in the input. In this work,

local-aiarxiv-cs-cv
21 Apr 2026
Model Releases

Long-CODE: Isolating Pure Long-Context as an Orthogonal Dimension in Video Evaluation

DGX agent

arXiv:2604.17428v1 Announce Type: new Abstract: As video generation models achieve unprecedented capabilities, the demand for robust video evaluation metrics becomes increasingly critical. Traditional

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Long-Text-to-Image Generation via Compositional Prompt Decomposition

DGX agent

arXiv:2604.18258v1 Announce Type: new Abstract: While modern text-to-image (T2I) models excel at generating images from intricate prompts, they struggle to capture the key details when the inputs are

model-releasesarxiv-cs-cv
21 Apr 2026
Agents

LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation

DGX agent

arXiv:2604.17190v1 Announce Type: new Abstract: Aerial Vision-and-Language Navigation (Aerial VLN) enables unmanned aerial vehicles (UAVs) to follow natural language instructions and navigate complex

agentsarxiv-cs-cv
21 Apr 2026
Tutorials

Lorentz Framework for Semantic Segmentation

DGX agent

arXiv:2604.16836v1 Announce Type: new Abstract: Semantic segmentation in hyperbolic space enables compact modeling of hierarchical structure while providing inherent uncertainty quantification. Prior

tutorialsarxiv-cs-cv
21 Apr 2026
Research

Low Light Image Enhancement Challenge at NTIRE 2026

DGX agent

arXiv:2604.17669v1 Announce Type: new Abstract: This paper presents a comprehensive review of the NTIRE 2026 Low Light Image Enhancement Challenge, highlighting the proposed solutions and final result

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Lumos3D: A Single-Forward Framework for Low-Light 3D Scene Restoration

DGX agent

arXiv:2511.09818v2 Announce Type: replace Abstract: Restoring 3D scenes with low-light conditions is challenging, and most existing methods depend on precomputed camera poses and scene-specific optimi

model-releasesarxiv-cs-cv
21 Apr 2026
Applications

MambaKick: Early Penalty Direction Prediction from HAR Embeddings

DGX agent

arXiv:2604.16588v1 Announce Type: new Abstract: Penalty kicks in soccer are decided under extreme time constraints, where goalkeepers benefit from anticipating shot direction from the kickers motion b

applicationsarxiv-cs-cv
21 Apr 2026
Safety

Mammo-FM: Breast-specific foundational model for Integrated Mammographic Diagnosis, Prognosis, and Reporting

DGX agent

arXiv:2512.00198v2 Announce Type: replace Abstract: Breast cancer is one of the leading causes of death among women worldwide. We introduce Mammo-FM, the first foundation model specifically for mammog

safetyarxiv-cs-cv
21 Apr 2026
Model Releases

MARCO: Navigating the Unseen Space of Semantic Correspondence

DGX agent

arXiv:2604.18267v1 Announce Type: new Abstract: Recent advances in semantic correspondence rely on dual-encoder architectures, combining DINOv2 with diffusion backbones. While accurate, these billion-

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Marrying Text-to-Motion Generation with Skeleton-Based Action Recognition

DGX agent

arXiv:2604.17090v1 Announce Type: new Abstract: Human action recognition and motion generation are two active research problems in human-centric computer vision, both aiming to align motion with textu

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems

DGX agent

arXiv:2503.16549v2 Announce Type: replace Abstract: Despite strong results on many tasks, multimodal large language models (MLLMs) still underperform on visual mathematical problem solving, especially

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis

DGX agent

arXiv:2411.17690v3 Announce Type: replace-cross Abstract: Unified decoder-only transformers have shown promise for multimodal generation, yet the mechanisms by which they synchronize modalities with h

safetyarxiv-cs-cv
21 Apr 2026
Research

Medial Axis Aware Learning of Signed Distance Functions

DGX agent

arXiv:2604.16512v1 Announce Type: new Abstract: We propose a novel variational method to compute a highly accurate global signed distance function (SDF) to a given point cloud. To this end, the jump s

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Medical Image Understanding Improves Survival Prediction via Visual Instruction Tuning

DGX agent

arXiv:2604.18250v1 Announce Type: new Abstract: Accurate prognostication and risk estimation are essential for guiding clinical decision-making and optimizing patient management. While radiologist-ass

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

MEDN: Motion-Emotion Feature Decoupling Network for Micro-Expression Recognition

DGX agent

arXiv:2604.17899v1 Announce Type: new Abstract: Unlike macro-expression, micro-expression does not follow a strictly consistent mapping rule between emotions and Action Units (AUs). As a result, some

model-releasesarxiv-cs-cv
21 Apr 2026
← Previous
1…228229230231232…261
Next →