AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

Inpainting-Style Conditional Diffusion for Multivariable Time Series Forecasting

DGX agent

arXiv:2605.28324v1 Announce Type: new Abstract: In this paper, we propose a novel conditional diffusion-based framework for multivariable time-series solar power forecasting. The proposed method refor

model-releasesarxiv-cs-cv
28 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Internally Referenced Low-Light Enhancement

DGX agent

arXiv:2605.28605v1 Announce Type: new Abstract: Self-supervised low-light image enhancement (LLIE) is highly appealing as it eliminates the reliance on external paired data. However, the lack of exter

model-releasesarxiv-cs-cv
28 May 2026
Research

Intra-YOLO: A Small Object Detection Model for Caries and Molar-Incisor Hypomineralization in Intraoral Photography Based on Transfer Learning with Reinforcement Learning

DGX agent

arXiv:2605.28157v1 Announce Type: new Abstract: This study developed a computer-aided diagnosis (CAD) system for detecting caries and molar-incisor hypomineralization (MIH) in intraoral photographs. T

researcharxiv-cs-cv
28 May 2026
Research

IRPO: Boosting Image Restoration via Post-training GRPO

DGX agent

arXiv:2512.00814v3 Announce Type: replace Abstract: Post-training has become effective for high-level generation, but its role in low-level vision remains underexplored. Existing image restoration met

researcharxiv-cs-cv
28 May 2026
Model Releases

Janus-LoRA: A Balanced Low-Rank Adaptation for Continual Learning

DGX agent

arXiv:2605.28495v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has emerged as a promising paradigm for Continual Learning. It independently updates its low-rank factors (A and B), creating

model-releasesarxiv-cs-cv
28 May 2026
Safety

JECA^2: Judgment-Explanation Consistent Adversarial Attack against Forensic Vision-Language Models

DGX agent

arXiv:2605.28609v1 Announce Type: new Abstract: Forensic vision-language models (VLMs) have recently been developed to detect image tampering and provide natural-language explanations. However, their

safetyarxiv-cs-cv
28 May 2026
Research

Learning to Label: A Reinforced Self-Evolving Framework for Semi-supervised Referring Expression Segmentation

DGX agent

arXiv:2605.28239v1 Announce Type: new Abstract: Semi-supervised referring expression segmentation (SS-RES) aims to achieve precise pixel-level language grounding under limited annotation, yet suffers

researcharxiv-cs-cv
28 May 2026
Research

LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts

DGX agent

arXiv:2602.11564v2 Announce Type: replace Abstract: Recent advances in video diffusion models have significantly improved visual quality, yet ultra-high-resolution (UHR) video generation remains a for

researcharxiv-cs-cv
28 May 2026
Model Releases

LV-OSD: Language-Vision-Complementary Open-Set Object Detection

DGX agent

arXiv:2605.28271v1 Announce Type: new Abstract: Object detection is an important task in computer vision, which aims to detect the objects of interest. through the given category list or query images.

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning

DGX agent

arXiv:2605.27960v1 Announce Type: new Abstract: Despite their popularity and success, Multimodal Large Language Models (MLLMs) often struggle to interpret images accurately, which limits their reasoni

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

MangaFlow: An End-to-End Agentic Framework for Controllable Story to Manga Generation

DGX agent

arXiv:2605.28173v1 Announce Type: new Abstract: End-to-end manga generation is a structured visual storytelling task that requires story decomposition, recurring character and scene grounding, page la

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

MeniOmni: A Structured Multimodal Benchmark for Holistic Meniscus Injury Assessment

DGX agent

arXiv:2605.28161v1 Announce Type: new Abstract: Clinical diagnosis of meniscus injuries requires radiologists to integrate volumetric MRI evidence with patient context (e.g., sex, age, BMI) and to pro

model-releasesarxiv-cs-cv
28 May 2026
Local Ai

MMRad-22K: A Structured Multimodal Evidence Dataset for Chest X-ray Report Generation

DGX agent

arXiv:2602.12843v2 Announce Type: replace Abstract: Chest X-ray (CXR) reporting follows a region-based clinical workflow in which radiologists inspect anatomical regions and integrate localized findin

local-aiarxiv-cs-cv
28 May 2026
Research

MORI-Seg: Learning Morphological Geometry for Instance Segmentation without Instance Annotations

DGX agent

arXiv:2605.28261v1 Announce Type: new Abstract: Instance-level quantification of kidney functional units is essential for morphometric analysis, yet most publicly available pathology datasets provide

researcharxiv-cs-cv
28 May 2026
Model Releases

NanoVDR: Distilling a 2B Vision-Language Retriever into a 70M Text-Only Encoder for Visual Document Retrieval

DGX agent

arXiv:2603.12824v2 Announce Type: replace-cross Abstract: Vision-Language Model (VLM) based retrievers have advanced visual document retrieval (VDR) to impressive quality. They require the same multi-

model-releasesarxiv-cs-cv
28 May 2026
Tutorials

Neural Image Space Tessellation efect

DGX agent

arXiv:2602.23754v2 Announce Type: replace-cross Abstract: We present Neural Image Space Tessellation effect (NIST), a lightweight screen-space post-processing approach for reducing the faceted silhoue

tutorialsarxiv-cs-cv
28 May 2026
Research

Next-Scale Autoregressive Models for Text-to-Motion Generation

DGX agent

arXiv:2604.03799v2 Announce Type: replace Abstract: Autoregressive (AR) models offer stable and efficient training, but standard next-token prediction is not well aligned with the temporal structure r

researcharxiv-cs-cv
28 May 2026
Model Releases

NL-MambaXCT: Self-Supervised Nested-Learning Mamba for Nomex Honeycomb X-ray CT Defect Classification

DGX agent

arXiv:2605.27454v1 Announce Type: cross Abstract: X-ray computed tomography (XCT) is widely used for non-destructive testing of Nomex honeycomb structures in aerospace manufacturing, but industrial in

model-releasesarxiv-cs-cv
28 May 2026
Safety

No Safe Dose: How Training Data Drives Unsafe Image Generation

DGX agent

arXiv:2605.28137v1 Announce Type: new Abstract: Text-to-image models trained on large-scale data often inevitably ingest unsafe content. While some people observe input-output amplifications, it remai

safetyarxiv-cs-cv
28 May 2026
Applications

ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency

DGX agent

arXiv:2508.18271v2 Announce Type: replace Abstract: 3D object inpainting is commonly achieved via multi-view 2D image completion, yet independently inpainted views often suffer from cross-view inconsi

applicationsarxiv-cs-cv
28 May 2026
Model Releases

{Omega}-QVLA: Robust Quantization for Vision-Language-Action Models via Composite Rotation and Per-step Scaling

DGX agent

arXiv:2605.28803v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models unify perception, reasoning, and control within a single policy, yet their multi-billion-parameter backbones and dif

model-releasesarxiv-cs-cv
28 May 2026
Research

OmniEgo-R^2: A Routed Reasoning Framework for the 1st Cross-Domain EgoCross Challenge at CVPR 2026

DGX agent

arXiv:2605.24481v2 Announce Type: replace Abstract: The 1st Cross-Domain EgoCross Challenge at EgoVis, CVPR 2026 evaluates whether multimodal large language models can reason over egocentric videos ac

researcharxiv-cs-cv
28 May 2026
Model Releases

On the Equivariant Learning of the Q-tensor Order Parameter

DGX agent

arXiv:2605.27679v1 Announce Type: cross Abstract: We construct and evaluate group-equivariant neural networks for the prediction of the two-dimensional Q-tensor order parameter of nematic liquid cryst

model-releasesarxiv-cs-cv
28 May 2026
Local Ai

OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning

DGX agent

arXiv:2605.28691v1 Announce Type: new Abstract: Diffusion Transformers achieve strong video generation quality, but the quadratic cost of full attention limits efficiency. We introduce OSP-Next, an ef

local-aiarxiv-cs-cv
28 May 2026
Research

Pattern Recognition Tasks with Personalized Federated Learning

DGX agent

arXiv:2605.27816v1 Announce Type: new Abstract: Personalized Federated Learning (PFL) constitutes a novel paradigm that tailors Machine Learning (ML) models to individual clients, thereby furnishing p

researcharxiv-cs-cv
28 May 2026
Local Ai

PocketGS: On-Device Training of 3D Gaussian Splatting for High Perceptual Modeling

DGX agent

arXiv:2601.17354v5 Announce Type: replace Abstract: While 3D Gaussian Splatting (3DGS) enables real-time rendering, its training demands workstation-level compute and memory, making mobile deployment

local-aiarxiv-cs-cv
28 May 2026
Model Releases

POINav: Benchmarking and Enhancing Final-Meters Arrival in Real-World Vision-Language Navigation

DGX agent

arXiv:2605.28237v1 Announce Type: cross Abstract: Real-world navigation is fundamentally driven by Points of Interest (POIs), yet reaching a precise POI remains a critical 'final-meters' challenge. Ex

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

PointQ-Bench: Benchmarking Diagnostic and Interpretable Point Cloud Quality Assessment

DGX agent

arXiv:2605.28241v1 Announce Type: new Abstract: Point cloud quality plays a critical role in 3D acquisition, reconstruction, rendering, and perception, yet existing point cloud quality assessment (PCQ

model-releasesarxiv-cs-cv
28 May 2026
Research

Privacy Protection Against Personalized Text-to-Image Synthesis via Cross-image Consistency Constraints

DGX agent

arXiv:2504.12747v2 Announce Type: replace Abstract: The rapid advancement of diffusion models and personalization techniques has made it possible to recreate individual portraits from just a few publi

researcharxiv-cs-cv
28 May 2026
Research

Proprio: Latent Self-Scoring and Inference-Time Refinement for Physically Plausible Video Generation

DGX agent

arXiv:2605.28230v1 Announce Type: new Abstract: Modern video generative models produce visually impressive results, yet frequently violate basic physical principles. We propose Proprio, a training-fre

researcharxiv-cs-cv
28 May 2026
Model Releases

Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation

DGX agent

arXiv:2605.28091v1 Announce Type: new Abstract: Text-to-Image generation has evolved from basic image synthesis into a frequently used core capability in professional creative workflows, where simple

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

RASR: Retrieval-Augmented Super Resolution for Practical Reference-based Image Restoration

DGX agent

arXiv:2508.09449v2 Announce Type: replace Abstract: Reference-based Super Resolution (RefSR) improves upon Single Image Super Resolution (SISR) by leveraging high-quality reference images to enhance t

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

Reflective Dialogue between Teacher and Solver Agents for Video Question Answering

DGX agent

arXiv:2605.27885v1 Announce Type: new Abstract: Various approaches have been proposed to adapt Vision-Language Models (VLMs) to specialized domains for Video Question Answering, including fine-tuning

model-releasesarxiv-cs-cv
28 May 2026
Applications

Representation-Conditioned Diffusion Models for Guided Training Data Generation

DGX agent

arXiv:2605.27495v1 Announce Type: new Abstract: Data availability remains a critical bottleneck in many deep learning applications. Large-scale datasets are often expensive to collect, curate and anno

applicationsarxiv-cs-cv
28 May 2026
Model Releases

Resolution-free neural surrogates for geometric parameterization and mapping with spatially varying fields

DGX agent

arXiv:2605.28551v1 Announce Type: new Abstract: Many imaging problems require computing spatial transformations induced by spatially varying intensity, feature, or density fields. Canonical examples i

model-releasesarxiv-cs-cv
28 May 2026
Applications

Rethinking Video-Language Model from the Language Input Perspective

DGX agent

arXiv:2605.27920v1 Announce Type: new Abstract: Driven by the wave of large language models, Video-Language Models (VLMs) have become a significant yet challenging technology to bridge the gap between

applicationsarxiv-cs-cv
28 May 2026
Research

REVEAL: Reference-Grounded Reasoning for Multimodal Manipulation Detection

DGX agent

arXiv:2605.28459v1 Announce Type: new Abstract: Multimodal manipulation detection aims to simultaneously identify forged image--text pairs and localize tampered regions, yet existing methods typically

researcharxiv-cs-cv
28 May 2026
Model Releases

Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification

DGX agent

arXiv:2512.12887v3 Announce Type: replace Abstract: 3D medical image classification is essential for modern clinical workflows. Medical foundation models (FMs) have emerged as a promising approach for

model-releasesarxiv-cs-cv
28 May 2026
Safety

SA4Depth: Consistent Pose-Depth Scale Alignment for Self-Supervised Monocular Depth Estimation

DGX agent

arXiv:2605.28477v1 Announce Type: new Abstract: Self-supervised depth estimation from monocular sequences relies on the joint learning of a depth and a pose network. Despite abundant research done to

safetyarxiv-cs-cv
28 May 2026
Research

SAFE-Diff: Scale-Aware Attention and Feature-Dispersive Diffusion with Uncertainty Estimation for Contrast-Enhanced Breast MRI Synthesis

DGX agent

arXiv:2605.25767v2 Announce Type: replace Abstract: Synthesizing high fidelity contrast enhanced MRI is clinically valuable for safer and more efficient breast cancer screening, yet remains challengin

researcharxiv-cs-cv
28 May 2026
Model Releases

SAM-Enhanced Segmentation on Road Datasets: Balancing Critical Classes in Autonomous Driving

DGX agent

arXiv:2605.28136v1 Announce Type: new Abstract: Dense semantic segmentation is essential for autonomous driving, yet many multi-modal datasets lack pixel-level annotations. The Zenseact Open Dataset (

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

SeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined Grouping

DGX agent

arXiv:2605.28735v1 Announce Type: new Abstract: Transparent objects are common in daily life, and it is important to understand their multilayer depth, including the transparent surface and the object

model-releasesarxiv-cs-cv
28 May 2026
Agents

Segment to Focus: Guiding Latent Action Models in the Presence of Distractors

DGX agent

arXiv:2602.02259v2 Announce Type: replace-cross Abstract: Latent action models (LAMs) offer a promising path to pre-training embodied agents on large amounts of action-free video. They infer latent ac

agentsarxiv-cs-cv
28 May 2026
Research

Self-Prophetic Decoding to Unlock Visual Search in LVLMs

DGX agent

arXiv:2605.28741v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) are rapidly evolving toward true multimodal reasoning, with visual search representing a concrete instantiation of

researcharxiv-cs-cv
28 May 2026
Model Releases

Self-Supervised Online Robot-Agnostic Traversability Estimation for Open-World Environments

DGX agent

arXiv:2605.28442v1 Announce Type: cross Abstract: Self-supervised online traversability estimation enables robots to continuously learn from unlabeled open-world experiences and adapt their navigation

model-releasesarxiv-cs-cv
28 May 2026
Safety

SEMAGIC: Learning Semantically Consistent Deformable 3D Representations from In-the-Wild Images

DGX agent

arXiv:2605.27938v1 Announce Type: new Abstract: Learning deformable 3D object models from single-view in-the-wild images has enabled impressive 3D shape reconstruction without supervision. However, it

safetyarxiv-cs-cv
28 May 2026
Model Releases

SIGMA: Bridging Structural and Distributional Gaps for Vision Foundation Model Adaptation

DGX agent

arXiv:2605.27893v1 Announce Type: new Abstract: Vision Foundation Models (VFMs) have demonstrated impressive representational capabilities. However, adapting them to downstream tasks via full fine-tun

model-releasesarxiv-cs-cv
28 May 2026
Local Ai

SIGMA: Semantic-Difference Instruction-Grounding Mask Annotator for Text-Driven Image Manipulation Localization

DGX agent

arXiv:2605.27924v1 Announce Type: new Abstract: Text-driven image editing has advanced rapidly, but reliably localizing these manipulations requires image manipulation localization (IML) models traine

local-aiarxiv-cs-cv
28 May 2026
← Previous
1…139140141142143…263
Next →