AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Tutorials

Structuring Open-Ended NAS: Semi-Automated Design Knowledge Structuring with LLMs for Efficient Neural Architecture Search

DGX agent

arXiv:2605.19247v1 Announce Type: new Abstract: Current neural architecture search (NAS) methods are often limited by their predefined, restrictive search spaces. While recent large language model (LL

tutorialsarxiv-cs-cv
20 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

SVG360: Editable Multiview Vector Graphics from a Single SVG

DGX agent

arXiv:2511.16766v3 Announce Type: replace Abstract: Scalable Vector Graphics are a standard representation for editable visual design, yet they are usually authored as single view two dimensional illu

researcharxiv-cs-cv
20 May 2026
Research

SWEET: Sparse World Modeling with Image Editing for Embodied Task Execution

DGX agent

arXiv:2605.19319v1 Announce Type: new Abstract: Visual prediction has emerged as a promising paradigm for embodied control, where future observations are generated and then translated into actions. Ho

researcharxiv-cs-cv
20 May 2026
Applications

Taming Real-World Space-Time Video Super-Resolution with One-Step Diffusion

DGX agent

arXiv:2601.20308v2 Announce Type: replace Abstract: Diffusion models have demonstrated exceptional success in video super-resolution (VSR), exhibiting powerful capabilities for generating fine-grained

applicationsarxiv-cs-cv
20 May 2026
Local Ai

Tango3D: Towards Alignment for Global and Local 2D-3D Correspondence

DGX agent

arXiv:2605.19727v1 Announce Type: new Abstract: Existing 3D foundation models typically align point clouds to frozen vision-language spaces like CLIP, which achieve strong cross-modal retrieval by com

local-aiarxiv-cs-cv
20 May 2026
Model Releases

TextAlign: Preference Alignment for Text Rendering with Hierarchical Rewards

DGX agent

arXiv:2605.19320v1 Announce Type: new Abstract: Faithful text rendering remains a persistent weakness of large text-to-image generative models, as it requires both semantic instruction following and f

model-releasesarxiv-cs-cv
20 May 2026
Research

TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation

DGX agent

arXiv:2409.08248v2 Announce Type: replace Abstract: In this paper, we introduce TextBoost, an efficient one-shot personalization approach for text-to-image diffusion models. Traditional personalizatio

researcharxiv-cs-cv
20 May 2026
Local Ai

Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning

DGX agent

arXiv:2605.19491v1 Announce Type: new Abstract: Traditional whole slide image (WSI) analysis methods typically rely on the multiple instance learning (MIL) paradigm, which extracts patch-level feature

local-aiarxiv-cs-cv
20 May 2026
Model Releases

TideGS: Scalable Training of Over One Billion 3D Gaussian Splatting Primitives via Out-of-Core Optimization

DGX agent

arXiv:2605.20150v1 Announce Type: new Abstract: Training 3D Gaussian Splatting (3DGS) at billion-primitive scale is fundamentally memory-bound: each Gaussian primitive carries a large attribute vector

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs

DGX agent

arXiv:2605.19528v1 Announce Type: new Abstract: 3D localization in Multimodal Large Language Models (MLLMs), including 3D object detection and 3D visual grounding, is fundamentally limited by camera i

model-releasesarxiv-cs-cv
20 May 2026
Research

Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models

DGX agent

arXiv:2605.19137v1 Announce Type: new Abstract: Video foundation models achieve strong performance across many video understanding tasks, but typically require large-scale pre-training on massive vide

researcharxiv-cs-cv
20 May 2026
Tutorials

Towards Fine-Grained Robustness: Attention-Guided Test-Time Prompt Tuning for Vision-Language Models

DGX agent

arXiv:2605.19956v1 Announce Type: new Abstract: Vision-Language Models (VLMs), such as CLIP, have achieved significant zero-shot performance on downstream tasks with various fine-tuning adaptation met

tutorialsarxiv-cs-cv
20 May 2026
Research

TrajectoryMover: Generative Movement of Object Trajectories in Videos

DGX agent

arXiv:2603.29092v3 Announce Type: replace Abstract: Generative video editing has enabled several intuitive editing operations for short video clips that would previously have been difficult to achieve

researcharxiv-cs-cv
20 May 2026
Research

Trust It or Not: Evidential Uncertainty for Feed-Forward 3D Reconstruction with Trust3R

DGX agent

arXiv:2605.19539v1 Announce Type: new Abstract: Geometric foundation models hold promise for unconstrained dense geometry prediction from uncalibrated images. However, in current feed-forward designs,

researcharxiv-cs-cv
20 May 2026
Research

UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register

DGX agent

arXiv:2605.19622v1 Announce Type: new Abstract: Representation learning with Vision Transformers (ViTs) has advanced rapidly, yet the utility of large-scale models in spatially sensitive tasks is hind

researcharxiv-cs-cv
20 May 2026
Safety

Universal Skeleton Understanding via Differentiable Rendering and MLLMs

DGX agent

arXiv:2603.18003v4 Announce Type: replace Abstract: Multimodal large language models (MLLMs) exhibit strong visual-language reasoning, yet cannot process structured, non-visual data such as human skel

safetyarxiv-cs-cv
20 May 2026
Research

Unsupervised Unfolded rPCA (U2-rPCA): Deep Interpretable Clutter Filtering for Ultrasound Microvascular Imaging

DGX agent

arXiv:2510.00660v2 Announce Type: replace Abstract: High-sensitivity clutter filtering is a fundamental step in ultrasound microvascular imaging. Singular value decomposition (SVD) and robust principa

researcharxiv-cs-cv
20 May 2026
Model Releases

Vision Harnessing Agent for Open Ad-hoc Segmentation

DGX agent

arXiv:2605.19410v1 Announce Type: new Abstract: Segmentation has become easy when the concept is known, requiring retrieval of a learned visual grounding from text. It remains hard for open ad-hoc con

model-releasesarxiv-cs-cv
20 May 2026
Research

WBCAtt+: Fine-Grained Pixel-Level Morphological Annotations for White Blood Cell Images

DGX agent

arXiv:2605.19692v1 Announce Type: new Abstract: The microscopic examination of white blood cells (WBCs) plays a fundamental role in pathology and is essential for diagnosing blood disorders such as le

researcharxiv-cs-cv
20 May 2026
Model Releases

What Makes Synthetic Data Effective in Image Segmentation

DGX agent

arXiv:2605.19289v1 Announce Type: new Abstract: Driven by rapid advances in large-scale generative models, synthetic data has emerged as a promising solution for visual understanding. While modern dif

model-releasesarxiv-cs-cv
20 May 2026
Safety

When Preference Labels Fall Short: Aligning Diffusion Models from Real Data

DGX agent

arXiv:2605.19839v1 Announce Type: new Abstract: Preference alignment aims to guide generative models by learning from comparisons between preferred and non-preferred samples. In practice, most existin

safetyarxiv-cs-cv
20 May 2026
Safety

Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making

DGX agent

arXiv:2602.07008v2 Announce Type: replace Abstract: Reliable models should not only predict correctly, but also justify decisions with acceptable evidence. Yet conventional supervised learning typical

safetyarxiv-cs-cv
20 May 2026
Research

White-Balance First, Adjust Later: Cross-Camera Color Constancy via Vision-Language Evaluation

DGX agent

arXiv:2605.19613v1 Announce Type: new Abstract: Color constancy aims to keep object colors consistent under varying illumination. Cross-camera generalization in color constancy remains challenging bec

researcharxiv-cs-cv
20 May 2026
Safety

Worst-Group Equalized Odds Regularization for Multi-Attribute Fair Medical Image Classification

DGX agent

arXiv:2605.19214v1 Announce Type: cross Abstract: Diagnostic performance in medical AI varies systematically across demographic groups, yet subgroup AUC can mask clinically important disparities. At a

safetyarxiv-cs-cv
20 May 2026
Model Releases

WoundFormer: Multi-Scale Spatial Feature Fusion for Multi-Class Wound Tissue Segmentation

DGX agent

arXiv:2605.19868v1 Announce Type: new Abstract: Chronic wounds such as diabetic foot ulcers and pressure injuries require accurate tissue-level assessment to guide treatment planning and monitor heali

model-releasesarxiv-cs-cv
20 May 2026
Research

X-Ray cardiac angiographic vessel segmentation based on pixel classification using machine learning and region growing

DGX agent

arXiv:2605.20073v1 Announce Type: new Abstract: This work proposes a pixel-classification approach for vessel segmentation in x-ray angiograms. The proposal uses textural features such as anisotropic

researcharxiv-cs-cv
20 May 2026
Research

XFlowMap: Cross-Scale Generalization and Mapping of Massive Origin-Destination Data

DGX agent

arXiv:2605.18777v1 Announce Type: cross Abstract: Mapping large origin-destination (OD) datasets remains challenging because flow maps become cluttered, meaningful patterns occur at multiple spatial s

researcharxiv-cs-cv
20 May 2026
Research

3D Densification for Multi-Map Monocular VSLAM in Endoscopy

DGX agent

arXiv:2503.14346v3 Announce Type: replace Abstract: Multi-map Sparse Monocular visual Simultaneous Localization and Mapping applied to monocular endoscopic sequences has proven efficient to robustly r

researcharxiv-cs-cv
19 May 2026
Hardware

3D Skew Gaussian Splatting with Any Camera Trajectory Visualization Engine

DGX agent

arXiv:2605.18334v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) has revolutionized real-time photorealistic view synthesis, its fundamental reliance on symmetric Gaussian distributi

hardwarearxiv-cs-cv
19 May 2026
Model Releases

A Comprehensive Survey of Action Quality Assessment: Method and Benchmark

DGX agent

arXiv:2412.11149v2 Announce Type: replace Abstract: Action Quality Assessment (AQA) aims to automatically evaluate how well human actions are performed and has been widely applied in sports analysis,

model-releasesarxiv-cs-cv
19 May 2026
Research

A Conditional U-Net Pipeline with Pre- and Post-Processing for Aerial RGB-to-Thermal Image Translation

DGX agent

arXiv:2605.17564v1 Announce Type: new Abstract: Paired RGB-thermal data has shown significant utility across a range of applications, including image fusion, object tracking, and anomaly detection; ho

researcharxiv-cs-cv
19 May 2026
Research

A Dataset for the Recognition of Historical and Handwritten Music Scores in Western Notation

DGX agent

arXiv:2605.18436v1 Announce Type: new Abstract: A large amount of musical heritage has been digitised by memory institutions: libraries, museums, and archives. Nevertheless, the field of Optical Music

researcharxiv-cs-cv
19 May 2026
Research

A Large-Scale Study on the Accuracy vs Cost Trade-offs of Training and Evaluation Settings in Fine-Grained Image Recognition

DGX agent

arXiv:2605.18700v1 Announce Type: new Abstract: Prior work on fine-grained image recognition (FGIR) has established the importance of the backbone selection, but has neglected the accuracy-vs-cost tra

researcharxiv-cs-cv
19 May 2026
Research

A Retrieval-Augmented Generation Approach to Extracting Algorithmic Logic from Neural Networks

DGX agent

arXiv:2512.04329v2 Announce Type: replace Abstract: Reusing existing neural-network components is central to research efficiency, yet discovering, extracting, and validating such modules across thousa

researcharxiv-cs-cv
19 May 2026
Research

A simple approach for biometrics: Finger-knuckle prints recognition based on a Sobel filter and similarity measures

DGX agent

arXiv:2605.17673v1 Announce Type: new Abstract: The objective of this work is to propose a novel methodology for the finger knuckle print recognition, which is essentially a digital photo of the finge

researcharxiv-cs-cv
19 May 2026
Model Releases

A Systematic Analysis of Out-of-Distribution Detection Under Representation and Training Paradigm Shifts

DGX agent

arXiv:2511.11934v3 Announce Type: replace-cross Abstract: We present a systematic benchmark of out-of-distribution (OOD) detection CSFs through a representation-centric lens. Our study spans CNN and V

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

Accelerating Rectified Flow Models via Trajectory-Aware Caching

DGX agent

arXiv:2605.16789v1 Announce Type: new Abstract: Diffusion and rectified flow (RF) models generate high-fidelity images and videos, but their iterative velocity-field evaluations are computationally ex

model-releasesarxiv-cs-cv
19 May 2026
Research

Adaptive double-phase Rudin--Osher--Fatemi denoising model

DGX agent

arXiv:2510.04382v2 Announce Type: replace-cross Abstract: Even though more than 30 years have passed since the seminal Rudin--Osher--Fatemi (ROF) paper on total variation (TV) denoising, it remains re

researcharxiv-cs-cv
19 May 2026
Model Releases

Adaptive Fused Prior Transfer for Controllable Generative Image Compression

DGX agent

arXiv:2605.16817v1 Announce Type: cross Abstract: Learned image compression has achieved competitive rate-distortion performance, but very-low-bitrate reconstruction remains difficult because the tran

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

Advancing Narrative Long Video Generation via Training-Free Identity-Aware Memory

DGX agent

arXiv:2605.18733v1 Announce Type: new Abstract: Autoregressive video generation has improved rapidly in visual fidelity and interactivity, but it still suffers from long-term inconsistency and memory

model-releasesarxiv-cs-cv
19 May 2026
Model Releases

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech

DGX agent

arXiv:2605.17583v1 Announce Type: new Abstract: While existing text-to-speech (TTS) models exhibit high expressiveness, fine-grained control over composite instructions remains challenging due to the

model-releasesarxiv-cs-cv
19 May 2026
Safety

AIM: Adversarial Information Masking for Faithfulness Evaluation of Saliency Maps

DGX agent

arXiv:2605.16905v1 Announce Type: cross Abstract: Post-hoc saliency methods are widely used to interpret deep neural networks, but their faithfulness is difficult to evaluate reliably. Existing evalua

safetyarxiv-cs-cv
19 May 2026
Safety

Alignment and Safety of Diffusion Models via Reinforcement Learning and Reward Modeling: A Survey

DGX agent

arXiv:2505.17352v2 Announce Type: replace Abstract: Diffusion models have become a central paradigm for image and multimodal generation, yet their deployment raises persistent questions about alignmen

safetyarxiv-cs-cv
19 May 2026
Safety

An Efficient Streaming Video Understanding Framework with Agentic Control

DGX agent

arXiv:2605.17921v1 Announce Type: new Abstract: Streaming video requires handling dynamic information density under strict latency budgets. Yet, existing methods typically employ static strategies, su

safetyarxiv-cs-cv
19 May 2026
Applications

Articulation in Prime: Primitive-Based Articulated Object Understanding from a Single Casual Video

DGX agent

arXiv:2605.18645v1 Announce Type: new Abstract: Retrieving the 3D kinematics of articulated objects from monocular video is a fundamental challenge in computer vision. Existing methods rely on complex

applicationsarxiv-cs-cv
19 May 2026
Model Releases

ArtMesh: Part-Aware Articulated Mesh Fields with Motion-Consistent Dynamics

DGX agent

arXiv:2605.16582v1 Announce Type: new Abstract: We present ArtMesh, a mesh-native method for reconstructing articulated objects explicitly as connected triangle meshes with per-part rigid motion from

model-releasesarxiv-cs-cv
19 May 2026
Research

AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents

DGX agent

arXiv:2605.17933v1 Announce Type: new Abstract: Vision-language model (VLM) agents increasingly rely on memory-augmented reinforcement learning to reuse experience across long-horizon tasks, yet most

researcharxiv-cs-cv
19 May 2026
Local Ai

AtlasVid: Efficient Ultra-High-Resolution Long Video Generation via Decoupled Global-Local Modeling

DGX agent

arXiv:2605.16649v1 Announce Type: new Abstract: Recent diffusion-based video generators have achieved remarkable visual fidelity and prompt controllability, yet scaling them to ultra-high-resolution (

local-aiarxiv-cs-cv
19 May 2026
← Previous
1…157158159160161…263
Next →