AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
20 May 2026

Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era

SafetyDGX agent

arXiv:2605.18903v1 Announce Type: cross Abstract: Vision-Language Models in Continual Learning (VLM-CL) aim to continuously adapt to new multimodal tasks while retaining prior knowledge. The emerging

RECIPE: Procedural Planning via Grounding in Instructional Video

Model ReleasesDGX agent

arXiv:2605.19976v1 Announce Type: new Abstract: Visual planning asks a model to generate the remaining steps of a procedure in natural language given a partial video context and a goal. Progress on th

Replacement Learning: Training Neural Networks with Fewer Parameters

Model ReleasesDGX agent

arXiv:2605.19533v1 Announce Type: new Abstract: End-to-end training with full-depth backpropagation remains the dominant paradigm for optimizing deep neural networks, but its efficiency deteriorates a


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Return of Frustratingly Easy Unsupervised Video Domain Adaptation

ResearchDGX agent

arXiv:2605.19510v1 Announce Type: new Abstract: Unsupervised video domain adaptation (UVDA) is a practical but under-explored problem. In this paper, we propose a frustratingly easy UVDA method, calle

Robust Mitigation of Age-Dependent Confounding Effects via Sample-Difficulty Decorrelation

ResearchDGX agent

arXiv:2605.19230v1 Announce Type: new Abstract: Age dependent performance disparities in medical image classification often arise because age acts as a confounder, linking imaging morphology with dise

RoomPilot: Controllable Indoor Scene Synthesis via Multimodal Semantic Parsing

ResearchDGX agent

arXiv:2512.11234v2 Announce Type: replace Abstract: Generating controllable indoor scenes is fundamental to applications in game development, architectural visualization, and embodied AI. However, exi

SafeAlign-VLA: A Negative-Enhanced Safe Alignment Framework for Risk-Aware Autonomous Driving

SafetyDGX agent

arXiv:2605.19524v1 Announce Type: cross Abstract: End-to-end autonomous driving systems excel in common scenarios but struggle with safety-critical long-tail cases. Vision-Language-Action (VLA) models

Scalable, Energy-Efficient Optical-Neural Architecture for Multiplexed Deepfake Video Detection

ApplicationsDGX agent

arXiv:2605.19360v1 Announce Type: new Abstract: The rapid proliferation of AI-generated visual media has created an urgent need for efficient, trustworthy deepfake detection systems. However, existing

Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling

SafetyDGX agent

arXiv:2503.06310v4 Announce Type: replace Abstract: Generating coherent long-form video sequences from discrete text prompts remains challenging due to difficulties in maintaining temporal coherence,

SEAL: Semantic Aware Image Watermarking

ResearchDGX agent

arXiv:2503.12172v4 Announce Type: replace-cross Abstract: Generative models have rapidly evolved to generate realistic outputs. However, their synthetic outputs increasingly challenge the clear distin

Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation

ResearchDGX agent

arXiv:2605.19340v1 Announce Type: new Abstract: Vision foundation models (VFMs) have achieved strong performance across various vision tasks. However, it still remains challenging to apply VFMs for cr

Self-Creative Text-to-Object Generation using Semantic-Aware Spatial Weighting

SafetyDGX agent

arXiv:2605.19554v1 Announce Type: new Abstract: Instilling creativity in text-to-image (T2I) generation presents a significant challenge, as it requires synthesized images to exhibit not only visual n

Semantic-Enriched Latent Visual Reasoning

Model ReleasesDGX agent

arXiv:2605.19342v1 Announce Type: new Abstract: Multimodal latent-space reasoning aims to replace explicit thinking with images by performing visual reasoning directly in a compact latent space. Howev

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction

ResearchDGX agent

arXiv:2605.20110v1 Announce Type: new Abstract: Referring segmentation grounds natural-language queries to pixel-level masks, but extending it to complex scenarios with multiple instances, cross-categ

Smartphone-based Circular Plot Sampling for Forest Inventory

TutorialsDGX agent

arXiv:2605.19213v1 Announce Type: new Abstract: Circular sample plots are a cornerstone of forest inventory, yet accurate measurement of tree diameter at breast height (DBH) and spatial location withi

Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction

ResearchDGX agent

arXiv:2605.06270v2 Announce Type: replace Abstract: Feed-forward 3D reconstruction models based on Vision Transformers can directly estimate scene geometry and camera poses from a small set of input i

Sparse Mixture-of-Experts Routing in Visual Diffusion Transformers:Diagnosis, Boundary Calibration and Evolutionary Roadmap from Routing Collapse to Selective Deadlock

ResearchDGX agent

arXiv:2605.19378v1 Announce Type: new Abstract: This paper systematically diagnoses the training failure modes of Token-Choice sparse Mixture-of-Experts (MoE) on video Diffusion Transformers. Starting

Spatially Prompted Visual Trajectory Prediction for Egocentric Manipulation

SafetyDGX agent

arXiv:2605.20085v1 Announce Type: new Abstract: Robotic manipulation is often specified through language instructions or task identifiers, yet cluttered environments with similar objects are better ha

Spectral Gradient Surgery for Domain-Generalizable Dataset Distillation

ResearchDGX agent

arXiv:2605.18836v1 Announce Type: cross Abstract: Dataset Distillation (DD) synthesizes a compact synthetic dataset that preserves the training utility of a full dataset. However, its standard formula

SpecX: A Large-Scale Benchmark for Multi-Modal Spectroscopy and Cross-Paradigm Evaluation

Model ReleasesDGX agent

arXiv:2605.18791v1 Announce Type: cross Abstract: Existing spectral benchmarks are limited in scale, modality alignment, and evaluation scope, and typically focus on either specialized models or multi

SphericalDreamer: Generating Navigable Immersive 3D Worlds with Panorama Fusion

ResearchDGX agent

arXiv:2605.19974v1 Announce Type: new Abstract: The generation of immersive and navigable 3D environments is increasingly prevalent with the growing adoption of virtual reality and 3D content. However

Stage-adaptive Token Selection for Efficient Omni-modal LLMs

ResearchDGX agent

arXiv:2605.20035v1 Announce Type: new Abstract: Omni-modal large language models (om-LLMs) achieve unified audio-visual understanding by encoding video and audio into temporally aligned token sequence

Structural Energy Guidance for View-Consistent Text-to-3D Generation

SafetyDGX agent

arXiv:2605.19876v1 Announce Type: new Abstract: Text-to-3D generation based on diffusion models often suffers from the Janus problem, leading to inconsistent geometry across viewpoints. This work iden

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding

Model ReleasesDGX agent

arXiv:2605.19866v1 Announce Type: new Abstract: Vision-Language Models (VLMs) parse documents end-to-end but frequently break down on layouts unlike those seen in training. We attribute this to a two-

Structuring Open-Ended NAS: Semi-Automated Design Knowledge Structuring with LLMs for Efficient Neural Architecture Search

TutorialsDGX agent

arXiv:2605.19247v1 Announce Type: new Abstract: Current neural architecture search (NAS) methods are often limited by their predefined, restrictive search spaces. While recent large language model (LL

SVG360: Editable Multiview Vector Graphics from a Single SVG

ResearchDGX agent

arXiv:2511.16766v3 Announce Type: replace Abstract: Scalable Vector Graphics are a standard representation for editable visual design, yet they are usually authored as single view two dimensional illu

SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution

ResearchDGX agent

arXiv:2605.19319v1 Announce Type: new Abstract: Visual prediction has emerged as a promising paradigm for embodied control, where future observations are generated and then translated into actions. Ho

Taming Real-World Space-Time Video Super-Resolution with One-Step Diffusion

ApplicationsDGX agent

arXiv:2601.20308v2 Announce Type: replace Abstract: Diffusion models have demonstrated exceptional success in video super-resolution (VSR), exhibiting powerful capabilities for generating fine-grained

Tango3D: Towards Alignment for Global and Local 2D-3D Correspondence

Local AiDGX agent

arXiv:2605.19727v1 Announce Type: new Abstract: Existing 3D foundation models typically align point clouds to frozen vision-language spaces like CLIP, which achieve strong cross-modal retrieval by com

TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards

Model ReleasesDGX agent

arXiv:2605.19320v1 Announce Type: new Abstract: Faithful text rendering remains a persistent weakness of large text-to-image generative models, as it requires both semantic instruction following and f

TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation

ResearchDGX agent

arXiv:2409.08248v2 Announce Type: replace Abstract: In this paper, we introduce TextBoost, an efficient one-shot personalization approach for text-to-image diffusion models. Traditional personalizatio

Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning

Local AiDGX agent

arXiv:2605.19491v1 Announce Type: new Abstract: Traditional whole slide image (WSI) analysis methods typically rely on the multiple instance learning (MIL) paradigm, which extracts patch-level feature

TideGS: Scalable Training of Over One Billion 3D Gaussian Splatting Primitives via Out-of-Core Optimization

Model ReleasesDGX agent

arXiv:2605.20150v1 Announce Type: new Abstract: Training 3D Gaussian Splatting (3DGS) at billion-primitive scale is fundamentally memory-bound: each Gaussian primitive carries a large attribute vector

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs

Model ReleasesDGX agent

arXiv:2605.19528v1 Announce Type: new Abstract: 3D localization in Multimodal Large Language Models (MLLMs), including 3D object detection and 3D visual grounding, is fundamentally limited by camera i

Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models

ResearchDGX agent

arXiv:2605.19137v1 Announce Type: new Abstract: Video foundation models achieve strong performance across many video understanding tasks, but typically require large-scale pre-training on massive vide

Towards Fine-Grained Robustness: Attention-Guided Test-Time Prompt Tuning for Vision-Language Models

TutorialsDGX agent

arXiv:2605.19956v1 Announce Type: new Abstract: Vision-Language Models (VLMs), such as CLIP, have achieved significant zero-shot performance on downstream tasks with various fine-tuning adaptation met

TrajectoryMover: Generative Movement of Object Trajectories in Videos

ResearchDGX agent

arXiv:2603.29092v3 Announce Type: replace Abstract: Generative video editing has enabled several intuitive editing operations for short video clips that would previously have been difficult to achieve

Trust It or Not: Evidential Uncertainty for Feed-Forward 3D Reconstruction with Trust3R

ResearchDGX agent

arXiv:2605.19539v1 Announce Type: new Abstract: Geometric foundation models hold promise for unconstrained dense geometry prediction from uncalibrated images. However, in current feed-forward designs,

UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register

ResearchDGX agent

arXiv:2605.19622v1 Announce Type: new Abstract: Representation learning with Vision Transformers (ViTs) has advanced rapidly, yet the utility of large-scale models in spatially sensitive tasks is hind

Universal Skeleton Understanding via Differentiable Rendering and MLLMs

SafetyDGX agent

arXiv:2603.18003v4 Announce Type: replace Abstract: Multimodal large language models (MLLMs) exhibit strong visual-language reasoning, yet cannot process structured, non-visual data such as human skel

Unsupervised Unfolded rPCA (U2-rPCA): Deep Interpretable Clutter Filtering for Ultrasound Microvascular Imaging

ResearchDGX agent

arXiv:2510.00660v2 Announce Type: replace Abstract: High-sensitivity clutter filtering is a fundamental step in ultrasound microvascular imaging. Singular value decomposition (SVD) and robust principa

Vision Harnessing Agent for Open Ad-hoc Segmentation

Model ReleasesDGX agent

arXiv:2605.19410v1 Announce Type: new Abstract: Segmentation has become easy when the concept is known, requiring retrieval of a learned visual grounding from text. It remains hard for open ad-hoc con

WBCAtt+: Fine-Grained Pixel-Level Morphological Annotations for White Blood Cell Images

ResearchDGX agent

arXiv:2605.19692v1 Announce Type: new Abstract: The microscopic examination of white blood cells (WBCs) plays a fundamental role in pathology and is essential for diagnosing blood disorders such as le

What Makes Synthetic Data Effective in Image Segmentation

Model ReleasesDGX agent

arXiv:2605.19289v1 Announce Type: new Abstract: Driven by rapid advances in large-scale generative models, synthetic data has emerged as a promising solution for visual understanding. While modern dif

When Preference Labels Fall Short: Aligning Diffusion Models from Real Data

SafetyDGX agent

arXiv:2605.19839v1 Announce Type: new Abstract: Preference alignment aims to guide generative models by learning from comparisons between preferred and non-preferred samples. In practice, most existin

Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making

SafetyDGX agent

arXiv:2602.07008v2 Announce Type: replace Abstract: Reliable models should not only predict correctly, but also justify decisions with acceptable evidence. Yet conventional supervised learning typical

White-Balance First, Adjust Later: Cross-Camera Color Constancy via Vision-Language Evaluation

ResearchDGX agent

arXiv:2605.19613v1 Announce Type: new Abstract: Color constancy aims to keep object colors consistent under varying illumination. Cross-camera generalization in color constancy remains challenging bec

Worst-Group Equalized Odds Regularization for Multi-Attribute Fair Medical Image Classification

SafetyDGX agent

arXiv:2605.19214v1 Announce Type: cross Abstract: Diagnostic performance in medical AI varies systematically across demographic groups, yet subgroup AUC can mask clinically important disparities. At a

WoundFormer: Multi-Scale Spatial Feature Fusion for Multi-Class Wound Tissue Segmentation

Model ReleasesDGX agent

arXiv:2605.19868v1 Announce Type: new Abstract: Chronic wounds such as diabetic foot ulcers and pressure injuries require accurate tissue-level assessment to guide treatment planning and monitor heali

X-Ray cardiac angiographic vessel segmentation based on pixel classification using machine learning and region growing

ResearchDGX agent

arXiv:2605.20073v1 Announce Type: new Abstract: This work proposes a pixel-classification approach for vessel segmentation in x-ray angiograms. The proposal uses textural features such as anisotropic

XFlowMap: Cross-Scale Generalization and Mapping of Massive Origin-Destination Data

ResearchDGX agent

arXiv:2605.18777v1 Announce Type: cross Abstract: Mapping large origin-destination (OD) datasets remains challenging because flow maps become cluttered, meaningful patterns occur at multiple spatial s

19 May 2026

3D Densification for Multi-Map Monocular VSLAM in Endoscopy

ResearchDGX agent

arXiv:2503.14346v3 Announce Type: replace Abstract: Multi-map Sparse Monocular visual Simultaneous Localization and Mapping applied to monocular endoscopic sequences has proven efficient to robustly r

3D Skew Gaussian Splatting with Any Camera Trajectory Visualization Engine

HardwareDGX agent

arXiv:2605.18334v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) has revolutionized real-time photorealistic view synthesis, its fundamental reliance on symmetric Gaussian distributi

A Comprehensive Survey of Action Quality Assessment: Method and Benchmark

Model ReleasesDGX agent

arXiv:2412.11149v2 Announce Type: replace Abstract: Action Quality Assessment (AQA) aims to automatically evaluate how well human actions are performed and has been widely applied in sports analysis,

A Conditional U-Net Pipeline with Pre- and Post-Processing for Aerial RGB-to-Thermal Image Translation

ResearchDGX agent

arXiv:2605.17564v1 Announce Type: new Abstract: Paired RGB-thermal data has shown significant utility across a range of applications, including image fusion, object tracking, and anomaly detection; ho

A Dataset for the Recognition of Historical and Handwritten Music Scores in Western Notation

ResearchDGX agent

arXiv:2605.18436v1 Announce Type: new Abstract: A large amount of musical heritage has been digitised by memory institutions: libraries, museums, and archives. Nevertheless, the field of Optical Music

A Large-Scale Study on the Accuracy vs Cost Trade-offs of Training and Evaluation Settings in Fine-Grained Image Recognition

ResearchDGX agent

arXiv:2605.18700v1 Announce Type: new Abstract: Prior work on fine-grained image recognition (FGIR) has established the importance of the backbone selection, but has neglected the accuracy-vs-cost tra

A Retrieval-Augmented Generation Approach to Extracting Algorithmic Logic from Neural Networks

ResearchDGX agent

arXiv:2512.04329v2 Announce Type: replace Abstract: Reusing existing neural-network components is central to research efficiency, yet discovering, extracting, and validating such modules across thousa

A simple approach for biometrics: Finger-knuckle prints recognition based on a Sobel filter and similarity measures

ResearchDGX agent

arXiv:2605.17673v1 Announce Type: new Abstract: The objective of this work is to propose a novel methodology for the finger knuckle print recognition, which is essentially a digital photo of the finge

A Systematic Analysis of Out-of-Distribution Detection Under Representation and Training Paradigm Shifts

Model ReleasesDGX agent

arXiv:2511.11934v3 Announce Type: replace-cross Abstract: We present a systematic benchmark of out-of-distribution (OOD) detection CSFs through a representation-centric lens. Our study spans CNN and V

← Previous
1…125126127128129…211
Next →