AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
28 May 2026

Inpainting-Style Conditional Diffusion for Multivariable Time Series Forecasting

Model ReleasesDGX agent

arXiv:2605.28324v1 Announce Type: new Abstract: In this paper, we propose a novel conditional diffusion-based framework for multivariable time-series solar power forecasting. The proposed method refor

Internally Referenced Low-Light Enhancement

Model ReleasesDGX agent

arXiv:2605.28605v1 Announce Type: new Abstract: Self-supervised low-light image enhancement (LLIE) is highly appealing as it eliminates the reliance on external paired data. However, the lack of exter

Intra-YOLO: A Small Object Detection Model for Caries and Molar-Incisor Hypomineralization in Intraoral Photography Based on Transfer Learning with Reinforcement Learning

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.28157v1 Announce Type: new Abstract: This study developed a computer-aided diagnosis (CAD) system for detecting caries and molar-incisor hypomineralization (MIH) in intraoral photographs. T

IRPO: Boosting Image Restoration via Post-training GRPO

ResearchDGX agent

arXiv:2512.00814v3 Announce Type: replace Abstract: Post-training has become effective for high-level generation, but its role in low-level vision remains underexplored. Existing image restoration met

Janus-LoRA: A Balanced Low-Rank Adaptation for Continual Learning

Model ReleasesDGX agent

arXiv:2605.28495v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has emerged as a promising paradigm for Continual Learning. It independently updates its low-rank factors (A and B), creating

JECA^2: Judgment-Explanation Consistent Adversarial Attack against Forensic Vision-Language Models

SafetyDGX agent

arXiv:2605.28609v1 Announce Type: new Abstract: Forensic vision-language models (VLMs) have recently been developed to detect image tampering and provide natural-language explanations. However, their

Learning to Label: A Reinforced Self-Evolving Framework for Semi-supervised Referring Expression Segmentation

ResearchDGX agent

arXiv:2605.28239v1 Announce Type: new Abstract: Semi-supervised referring expression segmentation (SS-RES) aims to achieve precise pixel-level language grounding under limited annotation, yet suffers

LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts

ResearchDGX agent

arXiv:2602.11564v2 Announce Type: replace Abstract: Recent advances in video diffusion models have significantly improved visual quality, yet ultra-high-resolution (UHR) video generation remains a for

LV-OSD: Language-Vision-Complementary Open-Set Object Detection

Model ReleasesDGX agent

arXiv:2605.28271v1 Announce Type: new Abstract: Object detection is an important task in computer vision, which aims to detect the objects of interest. through the given category list or query images.

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning

Model ReleasesDGX agent

arXiv:2605.27960v1 Announce Type: new Abstract: Despite their popularity and success, Multimodal Large Language Models (MLLMs) often struggle to interpret images accurately, which limits their reasoni

MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation

Model ReleasesDGX agent

arXiv:2605.28173v1 Announce Type: new Abstract: End-to-end manga generation is a structured visual storytelling task that requires story decomposition, recurring character and scene grounding, page la

MeniOmni: A Structured Multimodal Benchmark for Holistic Meniscus Injury Assessment

Model ReleasesDGX agent

arXiv:2605.28161v1 Announce Type: new Abstract: Clinical diagnosis of meniscus injuries requires radiologists to integrate volumetric MRI evidence with patient context (e.g., sex, age, BMI) and to pro

MMRad-22K: A Structured Multimodal Evidence Dataset for Chest X-ray Report Generation

Local AiDGX agent

arXiv:2602.12843v2 Announce Type: replace Abstract: Chest X-ray (CXR) reporting follows a region-based clinical workflow in which radiologists inspect anatomical regions and integrate localized findin

MORI-Seg: Learning Morphological Geometry for Instance Segmentation without Instance Annotations

ResearchDGX agent

arXiv:2605.28261v1 Announce Type: new Abstract: Instance-level quantification of kidney functional units is essential for morphometric analysis, yet most publicly available pathology datasets provide

NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval

Model ReleasesDGX agent

arXiv:2603.12824v2 Announce Type: replace-cross Abstract: Vision-Language Model (VLM) based retrievers have advanced visual document retrieval (VDR) to impressive quality. They require the same multi-

Neural Image Space Tessellation efect

TutorialsDGX agent

arXiv:2602.23754v2 Announce Type: replace-cross Abstract: We present Neural Image Space Tessellation effect (NIST), a lightweight screen-space post-processing approach for reducing the faceted silhoue

Next-Scale Autoregressive Models for Text-to-Motion Generation

ResearchDGX agent

arXiv:2604.03799v2 Announce Type: replace Abstract: Autoregressive (AR) models offer stable and efficient training, but standard next-token prediction is not well aligned with the temporal structure r

NL-MambaXCT: Self-Supervised Nested-Learning Mamba for Nomex Honeycomb X-ray CT Defect Classification

Model ReleasesDGX agent

arXiv:2605.27454v1 Announce Type: cross Abstract: X-ray computed tomography (XCT) is widely used for non-destructive testing of Nomex honeycomb structures in aerospace manufacturing, but industrial in

No Safe Dose: How Training Data Drives Unsafe Image Generation

SafetyDGX agent

arXiv:2605.28137v1 Announce Type: new Abstract: Text-to-image models trained on large-scale data often inevitably ingest unsafe content. While some people observe input-output amplifications, it remai

ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency

ApplicationsDGX agent

arXiv:2508.18271v2 Announce Type: replace Abstract: 3D object inpainting is commonly achieved via multi-view 2D image completion, yet independently inpainted views often suffer from cross-view inconsi

{Omega}-QVLA: Robust Quantization for Vision-Language-Action Models via Composite Rotation and Per-step Scaling

Model ReleasesDGX agent

arXiv:2605.28803v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models unify perception, reasoning, and control within a single policy, yet their multi-billion-parameter backbones and dif

OmniEgo-R^2: A Routed Reasoning Framework for the 1st Cross-Domain EgoCross Challenge at CVPR 2026

ResearchDGX agent

arXiv:2605.24481v2 Announce Type: replace Abstract: The 1st Cross-Domain EgoCross Challenge at EgoVis, CVPR 2026 evaluates whether multimodal large language models can reason over egocentric videos ac

On the Equivariant Learning of the Q-tensor Order Parameter

Model ReleasesDGX agent

arXiv:2605.27679v1 Announce Type: cross Abstract: We construct and evaluate group-equivariant neural networks for the prediction of the two-dimensional Q-tensor order parameter of nematic liquid cryst

OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning

Local AiDGX agent

arXiv:2605.28691v1 Announce Type: new Abstract: Diffusion Transformers achieve strong video generation quality, but the quadratic cost of full attention limits efficiency. We introduce OSP-Next, an ef

Pattern Recognition Tasks with Personalized Federated Learning

ResearchDGX agent

arXiv:2605.27816v1 Announce Type: new Abstract: Personalized Federated Learning (PFL) constitutes a novel paradigm that tailors Machine Learning (ML) models to individual clients, thereby furnishing p

PocketGS: On-Device Training of 3D Gaussian Splatting for High Perceptual Modeling

Local AiDGX agent

arXiv:2601.17354v5 Announce Type: replace Abstract: While 3D Gaussian Splatting (3DGS) enables real-time rendering, its training demands workstation-level compute and memory, making mobile deployment

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation

Model ReleasesDGX agent

arXiv:2605.28237v1 Announce Type: cross Abstract: Real-world navigation is fundamentally driven by Points of Interest (POIs), yet reaching a precise POI remains a critical 'final-meters' challenge. Ex

PointQ-Bench: Benchmarking Diagnostic and Interpretable Point Cloud Quality Assessment

Model ReleasesDGX agent

arXiv:2605.28241v1 Announce Type: new Abstract: Point cloud quality plays a critical role in 3D acquisition, reconstruction, rendering, and perception, yet existing point cloud quality assessment (PCQ

Privacy Protection Against Personalized Text-to-Image Synthesis via Cross-image Consistency Constraints

ResearchDGX agent

arXiv:2504.12747v2 Announce Type: replace Abstract: The rapid advancement of diffusion models and personalization techniques has made it possible to recreate individual portraits from just a few publi

Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation

ResearchDGX agent

arXiv:2605.28230v1 Announce Type: new Abstract: Modern video generative models produce visually impressive results, yet frequently violate basic physical principles. We propose Proprio, a training-fre

Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation

Model ReleasesDGX agent

arXiv:2605.28091v1 Announce Type: new Abstract: Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple

RASR: Retrieval-Augmented Super Resolution for Practical Reference-based Image Restoration

Model ReleasesDGX agent

arXiv:2508.09449v2 Announce Type: replace Abstract: Reference-based Super Resolution (RefSR) improves upon Single Image Super Resolution (SISR) by leveraging high-quality reference images to enhance t

Reflective Dialogue between Teacher and Solver Agents for Video Question Answering

Model ReleasesDGX agent

arXiv:2605.27885v1 Announce Type: new Abstract: Various approaches have been proposed to adapt Vision-Language Models (VLMs) to specialized domains for Video Question Answering, including fine-tuning

Representation-Conditioned Diffusion Models for Guided Training Data Generation

ApplicationsDGX agent

arXiv:2605.27495v1 Announce Type: new Abstract: Data availability remains a critical bottleneck in many deep learning applications. Large-scale datasets are often expensive to collect, curate and anno

Resolution-free neural surrogates for geometric parameterization and mapping with spatially varying fields

Model ReleasesDGX agent

arXiv:2605.28551v1 Announce Type: new Abstract: Many imaging problems require computing spatial transformations induced by spatially varying intensity, feature, or density fields. Canonical examples i

Rethinking Video-Language Model from the Language Input Perspective

ApplicationsDGX agent

arXiv:2605.27920v1 Announce Type: new Abstract: Driven by the wave of large language models, Video-Language Models (VLMs) have become a significant yet challenging technology to bridge the gap between

REVEAL: Reference-Grounded Reasoning for Multimodal Manipulation Detection

ResearchDGX agent

arXiv:2605.28459v1 Announce Type: new Abstract: Multimodal manipulation detection aims to simultaneously identify forged image--text pairs and localize tampered regions, yet existing methods typically

Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification

Model ReleasesDGX agent

arXiv:2512.12887v3 Announce Type: replace Abstract: 3D medical image classification is essential for modern clinical workflows. Medical foundation models (FMs) have emerged as a promising approach for

SA4Depth: Consistent Pose-Depth Scale Alignment for Self-Supervised Monocular Depth Estimation

SafetyDGX agent

arXiv:2605.28477v1 Announce Type: new Abstract: Self-supervised depth estimation from monocular sequences relies on the joint learning of a depth and a pose network. Despite abundant research done to

SAFE-Diff: Scale-Aware Attention and Feature-Dispersive Diffusion with Uncertainty Estimation for Contrast-Enhanced Breast MRI Synthesis

ResearchDGX agent

arXiv:2605.25767v2 Announce Type: replace Abstract: Synthesizing high fidelity contrast enhanced MRI is clinically valuable for safer and more efficient breast cancer screening, yet remains challengin

SAM-Enhanced Segmentation on Road Datasets: Balancing Critical Classes in Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.28136v1 Announce Type: new Abstract: Dense semantic segmentation is essential for autonomous driving, yet many multi-modal datasets lack pixel-level annotations. The Zenseact Open Dataset (

SeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined Grouping

Model ReleasesDGX agent

arXiv:2605.28735v1 Announce Type: new Abstract: Transparent objects are common in daily life, and it is important to understand their multilayer depth, including the transparent surface and the object

Segment to Focus: Guiding Latent Action Models in the Presence of Distractors

AgentsDGX agent

arXiv:2602.02259v2 Announce Type: replace-cross Abstract: Latent action models (LAMs) offer a promising path to pre-training embodied agents on large amounts of action-free video. They infer latent ac

Self-Prophetic Decoding to Unlock Visual Search in LVLMs

ResearchDGX agent

arXiv:2605.28741v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) are rapidly evolving toward true multimodal reasoning, with visual search representing a concrete instantiation of

Self-Supervised Online Robot-Agnostic Traversability Estimation for Open-World Environments

Model ReleasesDGX agent

arXiv:2605.28442v1 Announce Type: cross Abstract: Self-supervised online traversability estimation enables robots to continuously learn from unlabeled open-world experiences and adapt their navigation

SEMAGIC: Learning Semantically Consistent Deformable 3D Representations from In-the-Wild Images

SafetyDGX agent

arXiv:2605.27938v1 Announce Type: new Abstract: Learning deformable 3D object models from single-view in-the-wild images has enabled impressive 3D shape reconstruction without supervision. However, it

SIGMA: Bridging Structural and Distributional Gaps for Vision Foundation Model Adaptation

Model ReleasesDGX agent

arXiv:2605.27893v1 Announce Type: new Abstract: Vision Foundation Models (VFMs) have demonstrated impressive representational capabilities. However, adapting them to downstream tasks via full fine-tun

SIGMA: Semantic-Difference Instruction-Grounding Mask Annotator for Text-Driven Image Manipulation Localization

Local AiDGX agent

arXiv:2605.27924v1 Announce Type: new Abstract: Text-driven image editing has advanced rapidly, but reliably localizing these manipulations requires image manipulation localization (IML) models traine

Sketch2Motion: Text-driven 2D Sketch to 3D Animation via Diffusion-guided Skeleton Optimization

TutorialsDGX agent

arXiv:2605.28394v1 Announce Type: new Abstract: Animation of 2D hand-drawn sketches provides an effective medium for visual communication. However, these sketches pose challenges, particularly in hand

ST-ColoNet: Spatio-Temporal Colon Segment Recognition via Hybrid Attention and Edge-Guided Feature Learning

ResearchDGX agent

arXiv:2605.28119v1 Announce Type: new Abstract: Colo-segment recognition in colonoscopy videos is a key requirement for many downstream tasks, but existing automatic recognition methods only use colon

Stay Fair! Ensuring Group Fairness in Diffusion Models Across Guidance Scales

Model ReleasesDGX agent

arXiv:2605.28036v1 Announce Type: new Abstract: Diffusion models steer conditional generation with a tunable guidance scale to trade off prompt alignment and diversity. However, existing debiasing tec

Structure-Guided Visual Perturbation Neutralization for LVLMs

SafetyDGX agent

arXiv:2605.27927v1 Announce Type: new Abstract: Image inputs enable Large Vision Language Models (LVLMs) to perceive fine-grained visual information, but also introduce a pixel-level attack surface th

Structure over Pixels: Learning Variable-Length Visual Programs

ResearchDGX agent

arXiv:2605.27696v1 Announce Type: new Abstract: Discrete visual tokenizers translate images into ordered sequences of codes, providing a natural representation for structural description of scenes. Ye

Super-Resolved Canopy Height Mapping from Sentinel-2 Time Series Using Airborne LiDAR HD Reference Data across Metropolitan France

ResearchDGX agent

arXiv:2512.11524v3 Announce Type: replace Abstract: Fine-scale forest monitoring is essential for understanding canopy structure and its dynamics, which are key indicators of carbon stocks, biodiversi

Toward Semantic-Agnostic and Shape-Aware Vision-Language Segmentation Models

ApplicationsDGX agent

arXiv:2605.28348v1 Announce Type: new Abstract: Vision-language segmentation models have recently achieved strong performance by leveraging high-level semantic object categories expressed in natural l

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs

SafetyDGX agent

arXiv:2605.27894v1 Announce Type: new Abstract: Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these

Transfer learning RGB models to hyperspectral images with trainable tensor decompositions

ResearchDGX agent

arXiv:2605.28331v1 Announce Type: new Abstract: Transfer learning makes it possible to use large vision networks on a variety of domains, by specializing their models' general filters to new tasks. Ho

Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation

HardwareDGX agent

arXiv:2605.27582v1 Announce Type: cross Abstract: Embodied navigation requires an agent to map language and visual observations to a stream of spatial actions that drive a real robot through environme

VEOcc: Voxel-Centric Online Semantic Occupancy Prediction For Embodied Scene Understanding

AgentsDGX agent

arXiv:2605.25059v2 Announce Type: replace Abstract: Crucial for autonomous exploration, online 3D occupancy prediction and mapping incrementally constructs dense spatial representations on the fly. Ho

VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning

Model ReleasesDGX agent

arXiv:2510.08555v2 Announce Type: replace Abstract: Existing controllable video generation methods are typically designed for rigid, task-specific settings, such as first-frame image-to-video, inpaint

← Previous
1…111112113114115…211
Next →