AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

DGX agent

arXiv:2605.20183v1 Announce Type: new Abstract: Video generation is rapidly evolving from single-shot synthesis to complex multi-shot audio-video (MSAV) narratives to meet real-world demands. However,

model-releasesarxiv-cs-cv
20 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Multi-axis Analysis of Image Manipulation Localization

DGX agent

arXiv:2605.20174v1 Announce Type: new Abstract: Advanced image editing software enables easy creation of highly convincing image manipulations, which has been made even more accessible in recent years

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

DGX agent

arXiv:2511.14159v2 Announce Type: replace Abstract: Evaluating the robustness of Large Vision-Language Models (LVLMs) is essential for their continued development and responsible deployment in real-wo

model-releasesarxiv-cs-cv
20 May 2026
Research

Navigating the Emotion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition

DGX agent

arXiv:2605.18884v1 Announce Type: cross Abstract: Multimodal emotion recognition aims to integrate text, audio, and video sources to understand human affective states. Although multimodal large langua

researcharxiv-cs-cv
20 May 2026
Safety

Neuron Incidence Redistribution for Fairness in Medical Image Classification

DGX agent

arXiv:2605.19393v1 Announce Type: new Abstract: Deep learning models for medical image classification are susceptible to subgroup performance disparities across demographic attributes such as age, gen

safetyarxiv-cs-cv
20 May 2026
Model Releases

Next-Acceleration-Scale Prediction for Autoregressive MRI Reconstruction

DGX agent

arXiv:2605.19354v1 Announce Type: cross Abstract: MRI reconstruction is an inherently ill-posed inverse problem, since incomplete measurements admit many plausible solutions. This ambiguity becomes mo

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

NGL: Natural Garment Language for Training-Free Sewing Pattern Estimation

DGX agent

arXiv:2602.20700v2 Announce Type: replace Abstract: Estimating sewing patterns from images is a practical approach for creating high-quality 3D garments, but it remains challenging due to the scarcity

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

No Hard Negatives Required: Concept Centric Learning Leads to Compositionality without Degrading Zero-shot Capabilities of Contrastive Models

DGX agent

arXiv:2603.25722v2 Announce Type: replace Abstract: Contrastive vision-language (V&L) models remain a popular choice for various applications. However, several limitations have emerged, most notably t

model-releasesarxiv-cs-cv
20 May 2026
Safety

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer

DGX agent

arXiv:2511.22940v3 Announce Type: replace Abstract: Recent advances in diffusion models have greatly improved pose-driven character animation. However, existing methods are limited to spatially aligne

safetyarxiv-cs-cv
20 May 2026
Tutorials

OP2GS: Object-Aware 3D Gaussian Splatting with Dual-Opacity Primitives

DGX agent

arXiv:2605.20044v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) provides an explicit and efficient scene representation, but its primitives lack inherent object-level identity, hindering

tutorialsarxiv-cs-cv
20 May 2026
Model Releases

PEPL: Precision-Enhanced Pseudo-Labeling for Fine-Grained Image Classification in Semi-Supervised Learning

DGX agent

arXiv:2409.03192v2 Announce Type: replace Abstract: Fine-grained image classification has witnessed significant advancements with the advent of deep learning and computer vision technologies. However,

model-releasesarxiv-cs-cv
20 May 2026
Local Ai

Perceptual misalignment of texture representations in convolutional neural networks

DGX agent

arXiv:2604.01341v2 Announce Type: replace Abstract: Mathematical modeling of visual textures traces back to Julesz's intuition that texture perception in humans is based on local correlations between

local-aiarxiv-cs-cv
20 May 2026
Research

Personalized Face Privacy Protection From a Single Image

DGX agent

arXiv:2605.19032v1 Announce Type: new Abstract: Photos of faces uploaded online are vulnerable to malicious actors who can scrape facial images from online sources and intrude on personal privacy via

researcharxiv-cs-cv
20 May 2026
Model Releases

Physics-in-the-Loop: A Hybrid Agentic Architecture for Validated CAD Engineering Design

DGX agent

arXiv:2605.19717v1 Announce Type: new Abstract: Large Language Models (LLMs) can generate Computer-Aided Design (CAD), yet lack physical comprehension required for reliable engineering design. Instead

model-releasesarxiv-cs-cv
20 May 2026
Safety

Physics-informed simulation framework for realistic sonar image generation and statistical validation

DGX agent

arXiv:2605.19712v1 Announce Type: new Abstract: Synthetic sonar datasets offer a scalable alternative to costly real-world acquisition, yet their utility remains limited by the absence of rigorous qua

safetyarxiv-cs-cv
20 May 2026
Research

PiG-Avatar: Hierarchical Neural-Field-Guided Gaussian Avatars

DGX agent

arXiv:2605.20185v1 Announce Type: cross Abstract: Existing Gaussian avatar methods typically parameterize geometry on a body-template surface, which entangles the avatar's representation space with th

researcharxiv-cs-cv
20 May 2026
Model Releases

PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset

DGX agent

arXiv:2605.20147v1 Announce Type: new Abstract: Text-to-Image (T2I) models have recently seen notable progress around 1K and 2K resolution. With the extreme desire for better visual experience and the

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

PrAda: Few-Shot Visual Adaptation for Text-Prompted Segmentation

DGX agent

arXiv:2605.19623v1 Announce Type: new Abstract: Segmenting images is critical for visual understanding but demands extensive pixel-level annotations. Foundational models have enabled new paradigms for

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

Preferences Order, Ratings Anchor: From Fused Expert Aesthetic Ground Truth to Self-Distillation

DGX agent

arXiv:2605.19776v1 Announce Type: new Abstract: Pairwise preferences and pointwise ratings are the two dominant annotation protocols in image aesthetic assessment (IAA), yet existing benchmarks adopt

model-releasesarxiv-cs-cv
20 May 2026
Applications

Prognostic Value of Lung Ultrasound Biomarkers for Readmission Risk in Congestive Heart Failure: A Pilot Data-Driven Analysis

DGX agent

arXiv:2605.18878v1 Announce Type: cross Abstract: Hospital readmission within 30 days of discharge is a leading driver of morbidity, mortality, and avoidable healthcare expenditure in congestive heart

applicationsarxiv-cs-cv
20 May 2026
Model Releases

ProJo4D: Progressive Joint Optimization for Sparse-View Inverse Physics Estimation

DGX agent

arXiv:2506.05317v3 Announce Type: replace Abstract: Neural rendering has advanced significantly in 3D reconstruction and novel view synthesis, and integrating physics into these frameworks opens new a

model-releasesarxiv-cs-cv
20 May 2026
Research

PureCC: Pure Learning for Text-to-Image Concept Customization

DGX agent

arXiv:2603.07561v2 Announce Type: replace Abstract: Existing concept customization methods have achieved remarkable outcomes in high-fidelity and multi-concept customization. However, they often negle

researcharxiv-cs-cv
20 May 2026
Safety

Rapid patient-specific neural networks for intraoperative X-ray to volume registration

DGX agent

arXiv:2503.16309v2 Announce Type: replace-cross Abstract: Advanced navigation techniques in image-guided interventions and surgical robotics require the rapid and precise alignment of 3D preoperative

safetyarxiv-cs-cv
20 May 2026
Model Releases

Real-World On-Vehicle Evaluation of Embedding-Based Anomaly Detection

DGX agent

arXiv:2605.19744v1 Announce Type: new Abstract: Detecting anomalies in traffic scenes is crucial for ensuring safety in autonomous driving, yet collecting representative anomalous data remains challen

model-releasesarxiv-cs-cv
20 May 2026
Safety

Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era

DGX agent

arXiv:2605.18903v1 Announce Type: cross Abstract: Vision-Language Models in Continual Learning (VLM-CL) aim to continuously adapt to new multimodal tasks while retaining prior knowledge. The emerging

safetyarxiv-cs-cv
20 May 2026
Model Releases

RECIPE: Procedural Planning via Grounding in Instructional Video

DGX agent

arXiv:2605.19976v1 Announce Type: new Abstract: Visual planning asks a model to generate the remaining steps of a procedure in natural language given a partial video context and a goal. Progress on th

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

Replacement Learning: Training Neural Networks with Fewer Parameters

DGX agent

arXiv:2605.19533v1 Announce Type: new Abstract: End-to-end training with full-depth backpropagation remains the dominant paradigm for optimizing deep neural networks, but its efficiency deteriorates a

model-releasesarxiv-cs-cv
20 May 2026
Research

Return of Frustratingly Easy Unsupervised Video Domain Adaptation

DGX agent

arXiv:2605.19510v1 Announce Type: new Abstract: Unsupervised video domain adaptation (UVDA) is a practical but under-explored problem. In this paper, we propose a frustratingly easy UVDA method, calle

researcharxiv-cs-cv
20 May 2026
Research

Robust Mitigation of Age-Dependent Confounding Effects via Sample-Difficulty Decorrelation

DGX agent

arXiv:2605.19230v1 Announce Type: new Abstract: Age dependent performance disparities in medical image classification often arise because age acts as a confounder, linking imaging morphology with dise

researcharxiv-cs-cv
20 May 2026
Research

RoomPilot: Controllable Indoor Scene Synthesis via Multimodal Semantic Parsing

DGX agent

arXiv:2512.11234v2 Announce Type: replace Abstract: Generating controllable indoor scenes is fundamental to applications in game development, architectural visualization, and embodied AI. However, exi

researcharxiv-cs-cv
20 May 2026
Safety

SafeAlign-VLA: A Negative-Enhanced Safe Alignment Framework for Risk-Aware Autonomous Driving

DGX agent

arXiv:2605.19524v1 Announce Type: cross Abstract: End-to-end autonomous driving systems excel in common scenarios but struggle with safety-critical long-tail cases. Vision-Language-Action (VLA) models

safetyarxiv-cs-cv
20 May 2026
Applications

Scalable, Energy-Efficient Optical-Neural Architecture for Multiplexed Deepfake Video Detection

DGX agent

arXiv:2605.19360v1 Announce Type: new Abstract: The rapid proliferation of AI-generated visual media has created an urgent need for efficient, trustworthy deepfake detection systems. However, existing

applicationsarxiv-cs-cv
20 May 2026
Safety

Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling

DGX agent

arXiv:2503.06310v4 Announce Type: replace Abstract: Generating coherent long-form video sequences from discrete text prompts remains challenging due to difficulties in maintaining temporal coherence,

safetyarxiv-cs-cv
20 May 2026
Research

SEAL: Semantic Aware Image Watermarking

DGX agent

arXiv:2503.12172v4 Announce Type: replace-cross Abstract: Generative models have rapidly evolved to generate realistic outputs. However, their synthetic outputs increasingly challenge the clear distin

researcharxiv-cs-cv
20 May 2026
Research

Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation

DGX agent

arXiv:2605.19340v1 Announce Type: new Abstract: Vision foundation models (VFMs) have achieved strong performance across various vision tasks. However, it still remains challenging to apply VFMs for cr

researcharxiv-cs-cv
20 May 2026
Safety

Self-Creative Text-to-Object Generation using Semantic-Aware Spatial Weighting

DGX agent

arXiv:2605.19554v1 Announce Type: new Abstract: Instilling creativity in text-to-image (T2I) generation presents a significant challenge, as it requires synthesized images to exhibit not only visual n

safetyarxiv-cs-cv
20 May 2026
Model Releases

Semantic-Enriched Latent Visual Reasoning

DGX agent

arXiv:2605.19342v1 Announce Type: new Abstract: Multimodal latent-space reasoning aims to replace explicit thinking with images by performing visual reasoning directly in a compact latent space. Howev

model-releasesarxiv-cs-cv
20 May 2026
Research

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction

DGX agent

arXiv:2605.20110v1 Announce Type: new Abstract: Referring segmentation grounds natural-language queries to pixel-level masks, but extending it to complex scenarios with multiple instances, cross-categ

researcharxiv-cs-cv
20 May 2026
Tutorials

Smartphone-based Circular Plot Sampling for Forest Inventory

DGX agent

arXiv:2605.19213v1 Announce Type: new Abstract: Circular sample plots are a cornerstone of forest inventory, yet accurate measurement of tree diameter at breast height (DBH) and spatial location withi

tutorialsarxiv-cs-cv
20 May 2026
Research

Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction

DGX agent

arXiv:2605.06270v2 Announce Type: replace Abstract: Feed-forward 3D reconstruction models based on Vision Transformers can directly estimate scene geometry and camera poses from a small set of input i

researcharxiv-cs-cv
20 May 2026
Research

Sparse Mixture-of-Experts Routing in Visual Diffusion Transformers:Diagnosis, Boundary Calibration and Evolutionary Roadmap from Routing Collapse to Selective Deadlock

DGX agent

arXiv:2605.19378v1 Announce Type: new Abstract: This paper systematically diagnoses the training failure modes of Token-Choice sparse Mixture-of-Experts (MoE) on video Diffusion Transformers. Starting

researcharxiv-cs-cv
20 May 2026
Safety

Spatially Prompted Visual Trajectory Prediction for Egocentric Manipulation

DGX agent

arXiv:2605.20085v1 Announce Type: new Abstract: Robotic manipulation is often specified through language instructions or task identifiers, yet cluttered environments with similar objects are better ha

safetyarxiv-cs-cv
20 May 2026
Research

Spectral Gradient Surgery for Domain-Generalizable Dataset Distillation

DGX agent

arXiv:2605.18836v1 Announce Type: cross Abstract: Dataset Distillation (DD) synthesizes a compact synthetic dataset that preserves the training utility of a full dataset. However, its standard formula

researcharxiv-cs-cv
20 May 2026
Model Releases

SpecX: A Large-Scale Benchmark for Multi-Modal Spectroscopy and Cross-Paradigm Evaluation

DGX agent

arXiv:2605.18791v1 Announce Type: cross Abstract: Existing spectral benchmarks are limited in scale, modality alignment, and evaluation scope, and typically focus on either specialized models or multi

model-releasesarxiv-cs-cv
20 May 2026
Research

SphericalDreamer: Generating Navigable Immersive 3D Worlds with Panorama Fusion

DGX agent

arXiv:2605.19974v1 Announce Type: new Abstract: The generation of immersive and navigable 3D environments is increasingly prevalent with the growing adoption of virtual reality and 3D content. However

researcharxiv-cs-cv
20 May 2026
Research

Stage-adaptive Token Selection for Efficient Omni-modal LLMs

DGX agent

arXiv:2605.20035v1 Announce Type: new Abstract: Omni-modal large language models (om-LLMs) achieve unified audio-visual understanding by encoding video and audio into temporally aligned token sequence

researcharxiv-cs-cv
20 May 2026
Safety

Structural Energy Guidance for View-Consistent Text-to-3D Generation

DGX agent

arXiv:2605.19876v1 Announce Type: new Abstract: Text-to-3D generation based on diffusion models often suffers from the Janus problem, leading to inconsistent geometry across viewpoints. This work iden

safetyarxiv-cs-cv
20 May 2026
Model Releases

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding

DGX agent

arXiv:2605.19866v1 Announce Type: new Abstract: Vision-Language Models (VLMs) parse documents end-to-end but frequently break down on layouts unlike those seen in training. We attribute this to a two-

model-releasesarxiv-cs-cv
20 May 2026
← Previous
1…156157158159160…263
Next →