AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Model Releases

Hero-Mamba: Mamba-based Dual Domain Learning for Underwater Image Enhancement

DGX agent

arXiv:2604.16266v1 Announce Type: new Abstract: Underwater images often suffer from severe degradation, such as color distortion, low contrast, and blurred details, due to light absorption and scatter

model-releasesarxiv-cs-cv
20 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Hierarchical Codec Diffusion for Video-to-Speech Generation

DGX agent

arXiv:2604.15923v1 Announce Type: cross Abstract: Video-to-Speech (VTS) generation aims to synthesize speech from a silent video without auditory signals. However, existing VTS methods disregard the h

safetyarxiv-cs-cv
20 Apr 2026
Tutorials

HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images

DGX agent

arXiv:2603.02210v3 Announce Type: replace Abstract: Human-product images, which showcase the integration of humans and products, play a vital role in advertising, e-commerce, and digital marketing. Th

tutorialsarxiv-cs-cv
20 Apr 2026
Model Releases

HyCal: A Training-Free Prototype Calibration Method for Cross-Discipline Few-Shot Class-Incremental Learning

DGX agent

arXiv:2604.15678v1 Announce Type: new Abstract: Pretrained Vision-Language Models (VLMs) like CLIP show promise in continual learning, but existing Few-Shot Class-Incremental Learning (FSCIL) methods

model-releasesarxiv-cs-cv
20 Apr 2026
Model Releases

IA-CLAHE: Image-Adaptive Clip Limit Estimation for CLAHE

DGX agent

arXiv:2604.16010v1 Announce Type: new Abstract: This paper proposes image-adaptive contrast limited adaptive histogram equalization (IA-CLAHE). Conventional CLAHE is widely used to boost the performan

model-releasesarxiv-cs-cv
20 Apr 2026
Applications

Information Router for Mitigating Modality Dominance in Vision-Language Models

DGX agent

arXiv:2604.16264v1 Announce Type: new Abstract: Vision Language models (VLMs) have demonstrated strong performance across a wide range of benchmarks, yet they often suffer from modality dominance, whe

applicationsarxiv-cs-cv
20 Apr 2026
Model Releases

InstructTable: Improving Table Structure Recognition Through Instructions

DGX agent

arXiv:2604.02880v2 Announce Type: replace Abstract: Table structure recognition (TSR) holds widespread practical importance by parsing tabular images into structured representations, yet encounters si

model-releasesarxiv-cs-cv
20 Apr 2026
Tutorials

Learning Affine-Equivariant Proximal Operators

DGX agent

arXiv:2604.15556v1 Announce Type: cross Abstract: Proximal operators are fundamental across many applications in signal processing and machine learning, including solving ill-posed inverse problems. R

tutorialsarxiv-cs-cv
20 Apr 2026
Safety

Learning to Look before Learning to Like: Incorporating Human Visual Cognition into Aesthetic Quality Assessment

DGX agent

arXiv:2604.15853v1 Announce Type: new Abstract: Automated Aesthetic Quality Assessment (AQA) treats images primarily as static pixel vectors, aligning predictions with human-rating scores largely thro

safetyarxiv-cs-cv
20 Apr 2026
Research

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens

DGX agent

arXiv:2602.12370v2 Announce Type: replace Abstract: Recent progress in large models has led to significant advances in unified multimodal generation and understanding. However, the development of mode

researcharxiv-cs-cv
20 Apr 2026
Research

LP^{2}DH: A Locality-Preserving Pixel-Difference Hashing Framework for Dynamic Texture Recognition

DGX agent

arXiv:2604.15707v1 Announce Type: new Abstract: Spatiotemporal Local Binary Pattern (STLBP) is a widely used dynamic texture descriptor, but it suffers from extremely high dimensionality. To tackle th

researcharxiv-cs-cv
20 Apr 2026
Safety

M3R: Localized Rainfall Nowcasting with Meteorology-Informed MultiModal Attention

DGX agent

arXiv:2604.15377v1 Announce Type: cross Abstract: Accurate and timely rainfall nowcasting is crucial for disaster mitigation and water resource management. Despite recent advances in deep learning, pr

safetyarxiv-cs-cv
20 Apr 2026
Model Releases

Making Image Editing Easier via Adaptive Task Reformulation with Agentic Executions

DGX agent

arXiv:2604.15917v1 Announce Type: new Abstract: Instruction guided image editing has advanced substantially with recent generative models, yet it still fails to produce reliable results across many se

model-releasesarxiv-cs-cv
20 Apr 2026
Research

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation

DGX agent

arXiv:2510.09065v2 Announce Type: replace-cross Abstract: We introduce MMAudioSep, a generative model for video/text-queried sound separation that is founded on a pretrained video-to-audio model. By l

researcharxiv-cs-cv
20 Apr 2026
Model Releases

MMGait: Towards Multi-Modal Gait Recognition

DGX agent

arXiv:2604.15979v1 Announce Type: new Abstract: Gait recognition has emerged as a powerful biometric technique for identifying individuals at a distance without requiring user cooperation. Most existi

model-releasesarxiv-cs-cv
20 Apr 2026
Applications

Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions

DGX agent

arXiv:2604.16135v1 Announce Type: new Abstract: Recent advances in generative motion synthesis have enabled the production of realistic human motions from diverse input modalities. However, synthesizi

applicationsarxiv-cs-cv
20 Apr 2026
Research

Neural Gabor Splatting: Enhanced Gaussian Splatting with Neural Gabor for High-frequency Surface Reconstruction

DGX agent

arXiv:2604.15941v1 Announce Type: new Abstract: Recent years have witnessed the rapid emergence of 3D Gaussian splatting (3DGS) as a powerful approach for 3D reconstruction and novel view synthesis. I

researcharxiv-cs-cv
20 Apr 2026
Model Releases

neuralCAD-Edit: An Expert Benchmark for Multimodal-Instructed 3D CAD Model Editing

DGX agent

arXiv:2604.16170v1 Announce Type: new Abstract: We introduce neuralCAD-Edit, the first benchmark for editing 3D CAD models collected from expert CAD engineers. Instead of text conditioning as in prior

model-releasesarxiv-cs-cv
20 Apr 2026
Model Releases

P3T: Prototypical Point-level Prompt Tuning with Enhanced Generalization for 3D Vision-Language Models

DGX agent

arXiv:2604.15703v1 Announce Type: new Abstract: With the rise of pre-trained models in the 3D point cloud domain for a wide range of real-world applications, adapting them to downstream tasks has beco

model-releasesarxiv-cs-cv
20 Apr 2026
Model Releases

PILOT: A Promptable Interleaved Layout-aware OCR Transformer

DGX agent

arXiv:2504.03621v2 Announce Type: replace Abstract: Classical OCR pipelines decompose document reading into detection, segmentation, and recognition stages, which makes them sensitive to localization

model-releasesarxiv-cs-cv
20 Apr 2026
Model Releases

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation

DGX agent

arXiv:2604.15670v1 Announce Type: new Abstract: Reasoning segmentation has recently expanded from ground-level scenes to remote-sensing imagery, yet UAV data poses distinct challenges, including obliq

model-releasesarxiv-cs-cv
20 Apr 2026
Safety

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding

DGX agent

arXiv:2604.15770v1 Announce Type: new Abstract: Accurate open-vocabulary 3D scene understanding requires semantic representations that are both language-aligned and spatially precise at the pixel leve

safetyarxiv-cs-cv
20 Apr 2026
Local Ai

PolarMAE: Efficient Fetal Ultrasound Pre-training via Semantic Screening and Polar-Guided Masking

DGX agent

arXiv:2604.15893v1 Announce Type: new Abstract: Intelligent fetal ultrasound (US) interpretation is crucial for prenatal diagnosis, but high annotation costs and operator-induced variance make unsuper

local-aiarxiv-cs-cv
20 Apr 2026
Applications

Polyglot: Multilingual Style Preserving Speech-Driven Facial Animation

DGX agent

arXiv:2604.16108v1 Announce Type: new Abstract: Speech-Driven Facial Animation (SDFA) has gained significant attention due to its applications in movies, video games, and virtual reality. However, mos

applicationsarxiv-cs-cv
20 Apr 2026
Research

Proper Body Landmark Subset Enables More Accurate and 5X Faster Recognition of Isolated Signs in LIBRAS

DGX agent

arXiv:2510.24887v4 Announce Type: replace Abstract: This paper examines the feasibility of utilizing lightweight body landmark detection for recognizing isolated signs in Brazilian Sign Language (LIBR

researcharxiv-cs-cv
20 Apr 2026
Applications

ProtoTTA: Prototype-Guided Test-Time Adaptation

DGX agent

arXiv:2604.15494v1 Announce Type: cross Abstract: Deep networks that rely on prototypes-interpretable representations that can be related to the model input-have gained significant attention for balan

applicationsarxiv-cs-cv
20 Apr 2026
Research

Ranking XAI Methods for Head and Neck Cancer Outcome Prediction

DGX agent

arXiv:2604.16034v1 Announce Type: new Abstract: For head and neck cancer (HNC) patients, prognostic outcome prediction can support personalized treatment strategy selection. Improving prediction perfo

researcharxiv-cs-cv
20 Apr 2026
Model Releases

Reasoning over Video: Evaluating How MLLMs Extract, Integrate, and Reconstruct Spatiotemporal Evidence

DGX agent

arXiv:2603.13091v2 Announce Type: replace Abstract: The growing interest in embodied agents increases the demand for spatiotemporal video understanding, yet existing benchmarks largely emphasize extra

model-releasesarxiv-cs-cv
20 Apr 2026
Model Releases

Repurposing 3D Generative Model for Autoregressive Layout Generation

DGX agent

arXiv:2604.16299v1 Announce Type: new Abstract: We introduce LaviGen, a framework that repurposes 3D generative models for 3D layout generation. Unlike previous methods that infer object layouts from

model-releasesarxiv-cs-cv
20 Apr 2026
Applications

Saturation-Aware Space-Variant Blind Image Deblurring

DGX agent

arXiv:2604.16200v1 Announce Type: new Abstract: This paper presents a novel saturation aware space variant blind image deblurring framework designed to address challenges posed by saturated pixels in

applicationsarxiv-cs-cv
20 Apr 2026
Research

Scalable spatial point process models for forensic footwear analysis

DGX agent

arXiv:2602.07006v2 Announce Type: replace Abstract: Shoe print evidence recovered from crime scenes plays a key role in forensic investigations. By examining shoe prints, investigators can determine d

researcharxiv-cs-cv
20 Apr 2026
Safety

Scalable Unseen Objects 6-DoF Absolute Pose Estimation with Robotic Integration

DGX agent

arXiv:2503.05578v4 Announce Type: replace Abstract: Pose estimation-guided unseen object 6-DoF robotic manipulation is a key task in robotics. However, the scalability of current pose estimation metho

safetyarxiv-cs-cv
20 Apr 2026
Research

Self-Supervised Angular Deblurring in Photoacoustic Reconstruction via Noisier2Inverse

DGX agent

arXiv:2604.15681v1 Announce Type: new Abstract: Photoacoustic tomography (PAT) is an emerging imaging modality that combines the complementary strengths of optical contrast and ultrasonic resolution.

researcharxiv-cs-cv
20 Apr 2026
Agents

SENSE: Stereo OpEN Vocabulary SEmantic Segmentation

DGX agent

arXiv:2604.15946v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation enables models to segment objects or image regions beyond fixed class sets, offering flexibility in dynamic enviro

agentsarxiv-cs-cv
20 Apr 2026
Research

Splats in Splats++: Robust and Generalizable 3D Gaussian Splatting Steganography

DGX agent

arXiv:2604.15862v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has recently redefined the paradigm of 3D reconstruction, striking an unprecedented balance between visual fidelity and com

researcharxiv-cs-cv
20 Apr 2026
Research

SPLIT: Self-supervised Partitioning for Learned Inversion in Nonlinear Tomography

DGX agent

arXiv:2604.15651v1 Announce Type: new Abstract: Machine learning has achieved impressive performance in tomographic reconstruction, but supervised training requires paired measurements and ground-trut

researcharxiv-cs-cv
20 Apr 2026
Model Releases

SSFT: A Lightweight Spectral-Spatial Fusion Transformer for Generic Hyperspectral Classification

DGX agent

arXiv:2604.15828v1 Announce Type: new Abstract: Hyperspectral imaging enables fine-grained recognition of materials by capturing rich spectral signatures, but learning robust classifiers is challengin

model-releasesarxiv-cs-cv
20 Apr 2026
Research

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

DGX agent

arXiv:2602.05638v3 Announce Type: replace Abstract: While foundation models have advanced surgical video analysis, current approaches rely predominantly on pixel-level reconstruction objectives that w

researcharxiv-cs-cv
20 Apr 2026
Local Ai

TableSeq: Unified Generation of Structure, Content, and Layout

DGX agent

arXiv:2604.16070v1 Announce Type: new Abstract: We present TableSeq, an image-only, end-to-end framework for joint table structure recognition, content recognition, and cell localization. The model fo

local-aiarxiv-cs-cv
20 Apr 2026
Research

The Amazing Stability of Flow Matching

DGX agent

arXiv:2604.16079v1 Announce Type: new Abstract: The success of deep generative models in generating high-quality and diverse samples is often attributed to particular architectures and large training

researcharxiv-cs-cv
20 Apr 2026
Research

Topology-Driven Fusion of nnU-Net and MedNeXt for Accurate Brain Tumor Segmentation on Sub-Saharan Africa Dataset

DGX agent

arXiv:2604.15964v1 Announce Type: cross Abstract: Accurate automatic brain tumor segmentation in Low and Middle-Income (LMIC) countries is challenging due to the lack of defined national imaging proto

researcharxiv-cs-cv
20 Apr 2026
Research

Towards In-Context Tone Style Transfer with A Large-Scale Triplet Dataset

DGX agent

arXiv:2604.16114v1 Announce Type: new Abstract: Tone style transfer for photo retouching aims to adapt the stylistic tone of the reference image to a given content image. However, the lack of high-qua

researcharxiv-cs-cv
20 Apr 2026
Model Releases

Towards Realistic Open-Vocabulary Remote Sensing Segmentation: Benchmark and Baseline

DGX agent

arXiv:2604.15652v1 Announce Type: new Abstract: Open-vocabulary remote sensing image segmentation (OVRSIS) remains underexplored due to fragmented datasets, limited training diversity, and the lack of

model-releasesarxiv-cs-cv
20 Apr 2026
Research

Training Flow Matching: The Role of Weighting and Parameterization

DGX agent

arXiv:2603.06454v2 Announce Type: replace Abstract: We study the training objectives of denoising-based generative models, with a particular focus on loss weighting and output parameterization, includ

researcharxiv-cs-cv
20 Apr 2026
Local Ai

Two-Stage Framework for Efficient UAV-Based Wildfire Video Analysis with Adaptive Compression and Fire Source Detection

DGX agent

arXiv:2508.16739v2 Announce Type: replace Abstract: Unmanned Aerial Vehicles (UAVs) have become increasingly important in disaster emergency response by facilitating aerial video analysis. Due to the

local-aiarxiv-cs-cv
20 Apr 2026
Model Releases

TwoHamsters: Benchmarking Multi-Concept Compositional Unsafety in Text-to-Image Models

DGX agent

arXiv:2604.15967v1 Announce Type: cross Abstract: Despite the remarkable synthesis capabilities of text-to-image (T2I) models, safeguarding them against content violations remains a persistent challen

model-releasesarxiv-cs-cv
20 Apr 2026
Research

UA-Net: Uncertainty-Aware Network for TRISO Image Semantic Segmentation

DGX agent

arXiv:2604.15542v1 Announce Type: new Abstract: Tristructural isotropic (TRISO)-coated particle fuels undergo dimensional changes and chemical reactions during high-temperature neutron irradiation. Po

researcharxiv-cs-cv
20 Apr 2026
Local Ai

VeRVE: Versatile Retrieval for Videos via Unified Embeddings

DGX agent

arXiv:2601.12193v3 Announce Type: replace Abstract: Modern video retrieval systems are expected to handle diverse tasks ranging from corpus-level retrieval, fine-grained moment localization to flexibl

local-aiarxiv-cs-cv
20 Apr 2026
← Previous
1…234235236237238…261
Next →