AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
20 May 2026

GeoMamba: A Geometry-driven MambaVision Framework and Dataset for Fine-grained Optical-SAR Object Retrieval

ResearchDGX agent

arXiv:2605.19734v1 Announce Type: new Abstract: Multi-source remote sensing enables complementary observation of ground objects, while cross-modal fine-grained object retrieval remains challenging, es

GLUT: 3D Gaussian Lookup Table for Continuous Color Transformation

ResearchDGX agent

arXiv:2605.19889v1 Announce Type: cross Abstract: 3D Lookup Tables (3D LUTs) are widely used for color mapping, but their grid-based representation requires discretizing the RGB space, leading to a ca

GoTTA be Diverse: Rethinking Memory Policies for Test-Time Adaptation

Model ReleasesDGX agent

arXiv:2605.19890v1 Announce Type: new Abstract: Test-time adaptation (TTA) enables a pre-trained model to adapt online to an unlabeled test stream under distribution shift. While most TTA research foc


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

GRLoc: Geometric Representation Regression for Visual Localization

Local AiDGX agent

arXiv:2511.13864v2 Announce Type: replace Abstract: Absolute Pose Regression (APR) has emerged as a compelling paradigm for visual localization. However, APR models typically operate as black boxes, d

Hard-Label Black-Box Attacks on 3D Point Clouds

SafetyDGX agent

arXiv:2412.00404v2 Announce Type: replace Abstract: With the maturity of depth sensors in various 3D safety-critical applications, 3D point cloud models have been shown to be vulnerable to adversarial

HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding

Model ReleasesDGX agent

arXiv:2605.19223v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) exhibit strong performance on standard video tasks, their ability to faithfully summarize and reason over

HEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models

AgentsDGX agent

arXiv:2605.19631v1 Announce Type: cross Abstract: End-to-end autonomous driving has emerged as a compelling alternative to traditional modular pipelines by directly mapping raw sensor data to driving

Hierarchical Schedule Optimization for Fast and Robust Diffusion Model Sampling

Local AiDGX agent

arXiv:2511.11688v3 Announce Type: replace-cross Abstract: Diffusion probabilistic models have set a new standard for generative fidelity but are hindered by a slow iterative sampling process. A powerf

HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance

SafetyDGX agent

arXiv:2506.07209v2 Announce Type: replace-cross Abstract: We present HOI-PAGE, a new approach that prioritizes part-level affordance reasoning to generate high-fidelity 4D human-object interactions (H

iDiff: Interpretable Difference-aware Framework for Pairwise Image Quality Assessment

Local AiDGX agent

arXiv:2605.19522v1 Announce Type: new Abstract: Pairwise image quality assessment (IQA) in professional photography requires a model not only to identify the preferred image between two candidates, bu

iGSP:Implicit Gradient Subspace Projection for Efficient Continual Learning of Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.19301v1 Announce Type: new Abstract: Vision-Language Models require efficient adaptation to continually emerging downstream tasks. While Parameter-Efficient Fine-Tuning mitigates catastroph

Improved visual-information-driven model for crowd simulation and its modular application

SafetyDGX agent

arXiv:2504.03758v4 Announce Type: replace-cross Abstract: Crowd movement simulation is crucial for pedestrian safety management and facility design. Data-driven models offer the potential to improve r

INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference

TutorialsDGX agent

arXiv:2605.18853v1 Announce Type: cross Abstract: Edge deployment of Vision-Language Models (VLMs) faces a tradeoff between latency and accuracy: cloud execution provides high-quality predictions but

InterLight: Leveraging Intrinsic Illumination Priors for Low-Light Image Enhancement

ResearchDGX agent

arXiv:2605.19982v1 Announce Type: new Abstract: Low-Light Image Enhancement (LLIE) has long been a challenging problem in low-level vision, as insufficient illumination often leads to low contrast, de

Interpretable Computer Vision for Defect Detection in X-ray Tomography of Aerospace SiC/SiC Composites

ResearchDGX agent

arXiv:2605.20159v1 Announce Type: new Abstract: Non-destructive testing of aerospace SiC/SiC composites via X-ray computed tomography (XCT) relies on expert visual assessment, with current workflows o

Inverse Design of Metasurface based Absorbers using Physics Guided Conditional Diffusion Models

SafetyDGX agent

arXiv:2605.19611v1 Announce Type: new Abstract: Inverse design of metasurfaces for specific electromagnetic responses requires generating geometries that satisfy stringent spectral constraints while m

LaCoVL-FER: Landmark-Guided Contrastive Learning Network with Vision-Language Enhancement for Facial Expression Recognition

ApplicationsDGX agent

arXiv:2605.19821v1 Announce Type: new Abstract: Facial Expression Recognition (FER) in the wild is still challenging due to uncontrolled variations in pose, occlusion, and illumination. Most existing

Landscape-Awareness for Geometric View Diffusion Model

Local AiDGX agent

arXiv:2605.19865v1 Announce Type: new Abstract: Accurate camera viewpoint estimation under sparse-view conditions remains challenging, particularly in two-view scenarios. Recent approaches leverage di

Landslide Detection and Mapping Using Deep Learning Across Multi-Source Satellite Data and Geographic Regions

ResearchDGX agent

arXiv:2507.01123v2 Announce Type: replace Abstract: Landslides pose severe threats to infrastructure, economies, and human lives, necessitating accurate detection and predictive mapping across diverse

Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection

ResearchDGX agent

arXiv:2504.00470v2 Announce Type: replace-cross Abstract: To develop a trustworthy AI system, which aim to identify the input regions that most influence the models decisions. The primary task of exis

LMM-Track4D: Eliciting 4D Dynamic Reasoning in LMMs via Trajectory-Grounded Dialogue

Model ReleasesDGX agent

arXiv:2605.19390v1 Announce Type: new Abstract: Recent large multimodal models (LMMs) have become increasingly capable on image and video understanding, yet still struggle to sustain 4D continuous spa

Locate-then-Sparsify: Attribution Guided Sparse Strategy for Visual Hallucination Mitigation

ResearchDGX agent

arXiv:2603.16284v2 Announce Type: replace Abstract: Despite the significant advancements in Large Vision-Language Models (LVLMs), their tendency to generate hallucinations undermines reliability and r

Low-Compute Watermark Removal via Dual-Domain Natural Projection

SafetyDGX agent

arXiv:2510.07538v2 Announce Type: replace Abstract: Effective removal of semantic watermarks requires balancing three competing objectives: high removal success, low perceptual distortion, and low com

MAM-CLIP: Vision-Language Pretraining on Mammography Atlases for BI-RADS Classification

Model ReleasesDGX agent

arXiv:2605.19359v1 Announce Type: new Abstract: Deep learning methods have demonstrated promising results in predicting BI-RADS scores from mammography images. However, the interpretation of these ima

MapAnything: Evaluating Monocular Metric Depth Models for 3D Urban Asset Localization

ResearchDGX agent

arXiv:2509.14839v2 Announce Type: replace Abstract: City administrations increasingly rely on comprehensive databases and urban digital twins of city assets, such as traffic signs and trees, as well a

Matern Noise for Triangulation-Agnostic Flow Matching on Meshes

ResearchDGX agent

arXiv:2605.19305v1 Announce Type: cross Abstract: This paper tackles the task of learning to generate signals over triangle meshes in a triangulation-agnostic manner, meaning the trained model can be

MatPhys: Learning Material-Aware Physics Parameters for Deformable Object Simulation from Videos

ResearchDGX agent

arXiv:2605.19386v1 Announce Type: new Abstract: Reconstructing simulation-ready deformable objects is important for vision, graphics, and robotics. Existing physics-driven methods can recover physical

Mechanisms of Object Localization in Vision-Language Models

Local AiDGX agent

arXiv:2605.19792v1 Announce Type: new Abstract: Visually-grounded language models (VLMs) are highly effective in linking visual and textual information, yet they often struggle with basic classificati

MedFM-Robust: Benchmarking Robustness of Medical Foundation Models

Model ReleasesDGX agent

arXiv:2605.19027v1 Announce Type: new Abstract: Medical foundation models (MedFMs) have emerged as transformative tools in healthcare, demonstrating capabilities across diverse clinical applications.

MetaEarth-MM: Unified Multimodal Remote Sensing Image Generation with Scene-centered Joint Modeling

ResearchDGX agent

arXiv:2605.20090v1 Announce Type: new Abstract: Multi-modal remote sensing images are vital for Earth observation, yet complete paired observations are often scarce in practice. Existing generative me

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems

Model ReleasesDGX agent

arXiv:2605.19307v1 Announce Type: new Abstract: Visual Question Answering (VQA), as the representative multimodal task, serves as a key benchmark for evaluating the reasoning capabilities of Multimoda

Minimalist Visual Inertial Odometry

ApplicationsDGX agent

arXiv:2605.19990v1 Announce Type: cross Abstract: Visual-Inertial Odometry(VIO), which is critical to mobile robot navigation, uses cameras with a large number of pixels. Capturing and processing came

MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency

Model ReleasesDGX agent

arXiv:2510.25897v2 Announce Type: replace Abstract: The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one rew

MMGS: 10imes Compressed 3DGS through Optimal Transport Aggregation based on Multi-view Ranking

ResearchDGX agent

arXiv:2605.19304v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) has revolutionized 3D reconstruction, it suffers from significant overhead due to massive redundant primitives. Exist

Motion-2-To-3: Leveraging 2D Motion Data for 3D Motion Generations

ResearchDGX agent

arXiv:2412.13111v2 Announce Type: replace Abstract: Text-driven human motion synthesis has showcased its potential for revolutionizing motion design in the movie and game industry. Existing methods of

MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation

Model ReleasesDGX agent

arXiv:2605.18956v1 Announce Type: new Abstract: Recent motion-language models unify tasks like comprehension and generation but operate at a coarse granularity, lacking fine-grained understanding and

MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation

Model ReleasesDGX agent

arXiv:2605.20183v1 Announce Type: new Abstract: Video generation is rapidly evolving from single-shot synthesis to complex multi-shot audio-video (MSAV) narratives to meet real-world demands. However,

Multi-axis Analysis of Image Manipulation Localization

Model ReleasesDGX agent

arXiv:2605.20174v1 Announce Type: new Abstract: Advanced image editing software enables easy creation of highly convincing image manipulations, which has been made even more accessible in recent years

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs

Model ReleasesDGX agent

arXiv:2511.14159v2 Announce Type: replace Abstract: Evaluating the robustness of Large Vision-Language Models (LVLMs) is essential for their continued development and responsible deployment in real-wo

Navigating the Emotion Tree: Hierarchical Hyperbolic RAG for Multimodal Emotion Recognition

ResearchDGX agent

arXiv:2605.18884v1 Announce Type: cross Abstract: Multimodal emotion recognition aims to integrate text, audio, and video sources to understand human affective states. Although multimodal large langua

Neuron Incidence Redistribution for Fairness in Medical Image Classification

SafetyDGX agent

arXiv:2605.19393v1 Announce Type: new Abstract: Deep learning models for medical image classification are susceptible to subgroup performance disparities across demographic attributes such as age, gen

Next-Acceleration-Scale Prediction for Autoregressive MRI Reconstruction

Model ReleasesDGX agent

arXiv:2605.19354v1 Announce Type: cross Abstract: MRI reconstruction is an inherently ill-posed inverse problem, since incomplete measurements admit many plausible solutions. This ambiguity becomes mo

NGL: Natural Garment Language for Training-Free Sewing Pattern Estimation

Model ReleasesDGX agent

arXiv:2602.20700v2 Announce Type: replace Abstract: Estimating sewing patterns from images is a practical approach for creating high-quality 3D garments, but it remains challenging due to the scarcity

No Hard Negatives Required: Concept Centric Learning Leads to Compositionality without Degrading Zero-shot Capabilities of Contrastive Models

Model ReleasesDGX agent

arXiv:2603.25722v2 Announce Type: replace Abstract: Contrastive vision-language (V&L) models remain a popular choice for various applications. However, several limitations have emerged, most notably t

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer

SafetyDGX agent

arXiv:2511.22940v3 Announce Type: replace Abstract: Recent advances in diffusion models have greatly improved pose-driven character animation. However, existing methods are limited to spatially aligne

OP2GS: Object-Aware 3D Gaussian Splatting with Dual-Opacity Primitives

TutorialsDGX agent

arXiv:2605.20044v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) provides an explicit and efficient scene representation, but its primitives lack inherent object-level identity, hindering

PEPL: Precision-Enhanced Pseudo-Labeling for Fine-Grained Image Classification in Semi-Supervised Learning

Model ReleasesDGX agent

arXiv:2409.03192v2 Announce Type: replace Abstract: Fine-grained image classification has witnessed significant advancements with the advent of deep learning and computer vision technologies. However,

Perceptual misalignment of texture representations in convolutional neural networks

Local AiDGX agent

arXiv:2604.01341v2 Announce Type: replace Abstract: Mathematical modeling of visual textures traces back to Julesz's intuition that texture perception in humans is based on local correlations between

Personalized Face Privacy Protection From a Single Image

ResearchDGX agent

arXiv:2605.19032v1 Announce Type: new Abstract: Photos of faces uploaded online are vulnerable to malicious actors who can scrape facial images from online sources and intrude on personal privacy via

Physics-in-the-Loop: A Hybrid Agentic Architecture for Validated CAD Engineering Design

Model ReleasesDGX agent

arXiv:2605.19717v1 Announce Type: new Abstract: Large Language Models (LLMs) can generate Computer-Aided Design (CAD), yet lack physical comprehension required for reliable engineering design. Instead

Physics-informed simulation framework for realistic sonar image generation and statistical validation

SafetyDGX agent

arXiv:2605.19712v1 Announce Type: new Abstract: Synthetic sonar datasets offer a scalable alternative to costly real-world acquisition, yet their utility remains limited by the absence of rigorous qua

PiG-Avatar: Hierarchical Neural-Field-Guided Gaussian Avatars

ResearchDGX agent

arXiv:2605.20185v1 Announce Type: cross Abstract: Existing Gaussian avatar methods typically parameterize geometry on a body-template surface, which entangles the avatar's representation space with th

PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset

Model ReleasesDGX agent

arXiv:2605.20147v1 Announce Type: new Abstract: Text-to-Image (T2I) models have recently seen notable progress around 1K and 2K resolution. With the extreme desire for better visual experience and the

PrAda: Few-Shot Visual Adaptation for Text-Prompted Segmentation

Model ReleasesDGX agent

arXiv:2605.19623v1 Announce Type: new Abstract: Segmenting images is critical for visual understanding but demands extensive pixel-level annotations. Foundational models have enabled new paradigms for

Preferences Order, Ratings Anchor: From Fused Expert Aesthetic Ground Truth to Self-Distillation

Model ReleasesDGX agent

arXiv:2605.19776v1 Announce Type: new Abstract: Pairwise preferences and pointwise ratings are the two dominant annotation protocols in image aesthetic assessment (IAA), yet existing benchmarks adopt

Prognostic Value of Lung Ultrasound Biomarkers for Readmission Risk in Congestive Heart Failure: A Pilot Data-Driven Analysis

ApplicationsDGX agent

arXiv:2605.18878v1 Announce Type: cross Abstract: Hospital readmission within 30 days of discharge is a leading driver of morbidity, mortality, and avoidable healthcare expenditure in congestive heart

ProJo4D: Progressive Joint Optimization for Sparse-View Inverse Physics Estimation

Model ReleasesDGX agent

arXiv:2506.05317v3 Announce Type: replace Abstract: Neural rendering has advanced significantly in 3D reconstruction and novel view synthesis, and integrating physics into these frameworks opens new a

PureCC: Pure Learning for Text-to-Image Concept Customization

ResearchDGX agent

arXiv:2603.07561v2 Announce Type: replace Abstract: Existing concept customization methods have achieved remarkable outcomes in high-fidelity and multi-concept customization. However, they often negle

Rapid patient-specific neural networks for intraoperative X-ray to volume registration

SafetyDGX agent

arXiv:2503.16309v2 Announce Type: replace-cross Abstract: Advanced navigation techniques in image-guided interventions and surgical robotics require the rapid and precise alignment of 3D preoperative

Real-World On-Vehicle Evaluation of Embedding-Based Anomaly Detection

Model ReleasesDGX agent

arXiv:2605.19744v1 Announce Type: new Abstract: Detecting anomalies in traffic scenes is crucial for ensuring safety in autonomous driving, yet collecting representative anomalous data remains challen

← Previous
1…124125126127128…211
Next →