AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
7 May 2026

ScriptHOI: Learning Scripted State Transitions for Open-Vocabulary Human-Object Interaction Detection

ResearchDGX agent

arXiv:2605.05057v1 Announce Type: new Abstract: Open-vocabulary human-object interaction (HOI) detection requires recognizing interaction phrases that may not appear as annotated categories during tra

Segmenting proto-halos with vision transformers

ResearchDGX agent

arXiv:2508.00049v2 Announce Type: cross Abstract: The formation of dark-matter halos from small cosmological perturbations generated in the early universe is a highly non-linear process typically mode

Shape2Animal: Creative Animal Generation from Natural Silhouettes

ApplicationsDGX agent

arXiv:2506.20616v3 Announce Type: replace Abstract: Humans possess a unique ability to perceive meaningful patterns in ambiguous stimuli, a cognitive phenomenon known as pareidolia. This paper introdu


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation

Model ReleasesDGX agent

arXiv:2511.06754v3 Announce Type: replace-cross Abstract: Inspired by how humans reason over discrete objects and their relationships, we explore whether compact object-centric and object-relation rep

StableI2I: Spotting Unintended Changes in Image-to-Image Transition

Model ReleasesDGX agent

arXiv:2605.04453v1 Announce Type: new Abstract: In most real-world image-to-image (I2I) scenarios, existing evaluations primarily focus on instruction following and the perceptual quality or aesthetic

Stream-T1: Test-Time Scaling for Streaming Video Generation

TutorialsDGX agent

arXiv:2605.04461v1 Announce Type: new Abstract: While Test-Time Scaling (TTS) offers a promising direction to enhance video generation without the surging costs of training, current test-time video ge

Structured 3D Latents Are Surprisingly Powerful: Unleashing Generalizable Style with 2D Diffusion

ResearchDGX agent

arXiv:2605.04412v1 Announce Type: new Abstract: 3D asset generation plays a pivotal role in fields such as gaming and virtual reality, enabling the rapid synthesis of high-fidelity 3D objects from a s

SV-GS: Sparse View 4D Reconstruction with Skeleton-Driven Gaussian Splatting

TutorialsDGX agent

arXiv:2601.00285v2 Announce Type: replace Abstract: Reconstructing a dynamic target moving over a large area is challenging. Standard approaches for dynamic object reconstruction require dense coverag

Syn4D: A Multiview Synthetic 4D Dataset

ResearchDGX agent

arXiv:2605.05207v1 Announce Type: new Abstract: Dense 3D reconstruction and tracking of dynamic scenes from monocular video remains an important open challenge in computer vision. Progress in this are

Taming Outlier Tokens in Diffusion Transformers

ResearchDGX agent

arXiv:2605.05206v1 Announce Type: new Abstract: We study outlier tokens in Diffusion Transformers (DiTs) for image generation. Prior work has shown that Vision Transformers (ViTs) can produce a small

Temporal Structure Matters for Efficient Test-Time Adaptation in Wearable Human Activity Recognition

SafetyDGX agent

arXiv:2605.04617v1 Announce Type: new Abstract: Wearable human activity recognition (WHAR) models often suffer from performance degradation under real-world cross-user distribution shifts. Test-time a

Topology-Constrained Quantized nnUNet for Efficient and Anatomically Accurate 3D Tooth Segmentation

ApplicationsDGX agent

arXiv:2605.04201v1 Announce Type: new Abstract: We propose a topology-constrained quantized nnUNet framework for efficient and anatomically accurate 3D tooth segmentation, addressing the challenges of

Topology-Preserving Data Augmentation for Ring-Type Polygon Annotations

ResearchDGX agent

arXiv:2603.14764v3 Announce Type: replace Abstract: Geometric data augmentation is widely used in segmentation workflows, but polygon annotations are often assumed to remain valid after transformation

Towards General Preference Alignment: Diffusion Models at Nash Equilibrium

SafetyDGX agent

arXiv:2605.04494v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has been popular for aligning text-to-image (T2I) diffusion models with human preferences. As a main

Towards Generative Location Awareness for Disaster Response: A Probabilistic Cross-view Geolocalization Approach

Local AiDGX agent

arXiv:2512.20056v2 Announce Type: replace-cross Abstract: As Earth's climate changes, it is impacting disasters and extreme weather events across the planet. Record-breaking heat waves, drenching rain

UAV as Urban Construction Change Monitor: A New Benchmark and Change Captioning Model

Model ReleasesDGX agent

arXiv:2605.04409v1 Announce Type: new Abstract: Remote Sensing Image Change Captioning (RSICC) aims to generate spatially grounded natural language descriptions of scene evolution from bi-temporal ima

UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning

SafetyDGX agent

arXiv:2508.11196v2 Announce Type: replace Abstract: Recent advances in vision-language models (VLMs) have demonstrated strong generalization in natural image tasks. However, their performance often de

UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization

SafetyDGX agent

arXiv:2511.08195v3 Announce Type: replace Abstract: UI-to-code aims to translate UI screenshots into executable front-end code. Despite progress with vision-language models (VLMs), most existing metho

ULF-Loc: Unbiased Landmark Feature for Robust Visual Localization with 3D Gaussian Splatting

SafetyDGX agent

arXiv:2605.04730v1 Announce Type: new Abstract: Visual localization is a core technology for augmented reality and autonomous navigation. Recent methods combine the efficient rendering of 3D Gaussian

UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings

SafetyDGX agent

arXiv:2505.11815v2 Announce Type: replace Abstract: Current vision-language models have been explored for multi-modal embedding tasks like information retrieval. However, they face significant challen

UniPCB: A Generation-Assisted Detection Framework for PCB Defect Inspection

ResearchDGX agent

arXiv:2605.04635v1 Announce Type: new Abstract: Printed Circuit Board (PCB) defect inspection faces two compounding challenges: scarce and imbalanced defect samples that limit model training, and insu

VC-FeS: Viewpoint-Conditioned Feature Selection for Vehicle Re-identification in Thermal Vision

ResearchDGX agent

arXiv:2605.04750v1 Announce Type: new Abstract: Identification of less-articulated objects using single-channel images, such as thermal images, is important in many applications, such as surveillance.

Velox: Learning Representations of 4D Geometry and Appearance

TutorialsDGX agent

arXiv:2605.04527v1 Announce Type: new Abstract: We introduce a framework for learning latent representations of 4D objects which are descriptive, faithfully capturing object geometry and appearance; c

Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning

Model ReleasesDGX agent

arXiv:2506.06856v3 Announce Type: replace Abstract: Visual reasoning is crucial for understanding complex multimodal data and advancing Artificial General Intelligence. Existing methods enhance the re

VL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking

Model ReleasesDGX agent

arXiv:2605.04574v1 Announce Type: new Abstract: UAV-ground visual tracking (UGVT) aims to simultaneously track the same object from both the UAV and the ground view. However, existing two-stream metho

VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA

Model ReleasesDGX agent

arXiv:2605.04870v1 Announce Type: new Abstract: Video text-based visual question answering (Video TextVQA) aims to answer questions by reasoning over visual textual content appearing in videos. Despit

Wasserstein-Aligned Localisation for VLM-Based Distributional OOD Detection in Medical Imaging

Model ReleasesDGX agent

arXiv:2605.05161v1 Announce Type: new Abstract: Zero-shot anomaly localisation via vision-language models (VLMs) offers a compelling approach for rare pathology detection, yet its performance is funda

What Matters in Practical Learned Image Compression

Local AiDGX agent

arXiv:2605.05148v1 Announce Type: new Abstract: One of the major differentiators unlocked by learned codecs relative to their hard-coded traditional counterparts is their ability to be optimized direc

6 May 2026

3D Human Face Reconstruction with 3DMM face model from RGB image

ResearchDGX agent

arXiv:2605.03996v1 Announce Type: new Abstract: Nowadays as convolution neural networks demonstrate its powerful problem-solving ability in the area of image processing, efforts have been made to reco

4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere

ResearchDGX agent

arXiv:2602.10094v2 Announce Type: replace Abstract: We present 4RC, a unified feed-forward framework for 4D reconstruction from monocular videos. Unlike existing approaches that typically decouple mot

A Benchmark for Interactive World Models with a Unified Action Generation Framework

Model ReleasesDGX agent

arXiv:2605.03941v1 Announce Type: new Abstract: Achieving Artificial General Intelligence (AGI) requires agents that learn and interact adaptively, with interactive world models providing scalable env

A Deeper Dive into the Irreversibility of PolyProtect: Making Protected Face Templates Harder to Invert

Model ReleasesDGX agent

arXiv:2605.03857v1 Announce Type: new Abstract: This work presents a deeper analysis of the 'irreversibility' property of PolyProtect, a biometric template protection method initially proposed for sec

A Framework for Exploring and Disentangling Intersectional Bias: A Case Study in Fetal Ultrasound

Model ReleasesDGX agent

arXiv:2605.02942v1 Announce Type: cross Abstract: Bias in medical AI is often framed as a problem of representation. However, in image-based tasks such as fetal ultrasound, performance disparities can

A Partition-Based Generating Function for Row-Convex Polyominoes

ResearchDGX agent

arXiv:2605.03203v1 Announce Type: cross Abstract: An alternative generating function is proposed to enumerate row-convex polyominoes without internal holes on a discrete grid. The approach is based on

A Robust Unsupervised Domain Adaptation Framework for Medical Image Classification Using RKHS-MMD

SafetyDGX agent

arXiv:2605.03787v1 Announce Type: new Abstract: Labeling medical images is a major bottleneck in the field of medical imaging, as it requires domain-specific expertise, and it gets further complicated

A User-Centric Analysis of Explainability in AI-Based Medical Image Diagnosis

ResearchDGX agent

arXiv:2605.02903v1 Announce Type: cross Abstract: In recent years, AI systems in the medical domain have advanced significantly. However, despite outperforming humans, they are rarely used in practice

A Vision-Based Shared-Control Teleoperation Scheme for Controlling the Robotic Arm of a Four-Legged Robot

SafetyDGX agent

arXiv:2508.14994v3 Announce Type: replace-cross Abstract: In hazardous and remote environments, robotic systems perform critical tasks demanding improved safety and efficiency. Among these, quadruped

AHPA: Adaptive Hierarchical Prior Alignment for Diffusion Transformers

Local AiDGX agent

arXiv:2605.03317v1 Announce Type: new Abstract: Representation alignment has recently emerged as an effective paradigm for accelerating Diffusion Transformer training. Despite their success, existing

AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock

ApplicationsDGX agent

arXiv:2507.22101v2 Announce Type: replace Abstract: Crops, fisheries and livestock form the backbone of global food production, essential to feed the ever-growing global population. However, these sec

AniMatrix: An Anime Video Generation Model that Thinks in Art, Not Physics

ApplicationsDGX agent

arXiv:2605.03652v1 Announce Type: new Abstract: Video generation models internalize physical realism as their prior. Anime deliberately violates physics: smears, impact frames, chibi shifts; and its t

Approaching human parity in the quality of automated organoid image segmentation

TutorialsDGX agent

arXiv:2605.03053v1 Announce Type: new Abstract: Organoids are complex, three dimensional, self-organizing cell cultures which manifest organ-like features and represent a powerful platform for studyin

Audio-Visual Intelligence in Large Foundation Models

SafetyDGX agent

arXiv:2605.04045v1 Announce Type: new Abstract: Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines

Before Forgetting, Learn to Remember: Revisiting Foundational Learning Failures in LVLM Unlearning Benchmarks

Model ReleasesDGX agent

arXiv:2605.03759v1 Announce Type: new Abstract: While Large Vision-Language Models (LVLMs) offer powerful capabilities, they pose privacy risks by unintentionally memorizing sensitive personal informa

Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs

Model ReleasesDGX agent

arXiv:2512.09874v2 Announce Type: replace Abstract: Correctly parsing mathematical formulas from PDFs is critical for training large language models and building scientific knowledge bases from academ

BEVCALIB: LiDAR-Camera Calibration via Geometry-Guided Bird's-Eye View Representations

AgentsDGX agent

arXiv:2506.02587v2 Announce Type: replace Abstract: Accurate LiDAR-camera calibration is fundamental to fusing multi-modal perception in autonomous driving and robotic systems. Traditional calibration

BFORE: Butterfly-Firefly Optimized Retinex Enhancement for Low-Light Image Quality Improvement

Model ReleasesDGX agent

arXiv:2605.03509v1 Announce Type: new Abstract: Low-light image enhancement is a fundamental challenge in computer vision and multimedia applications, as images captured under insufficient illuminatio

Boundary-Aware Uncertainty Quantification for Wildfire Spread Prediction

ResearchDGX agent

arXiv:2605.03148v1 Announce Type: new Abstract: Reliable wildfire spread prediction is vital for risk-aware emergency planning, yet most deep learning models lack principled uncertainty quantification

Can Multimodal Large Language Models Understand Pathologic Movements? A Pilot Study on Seizure Semiology

ResearchDGX agent

arXiv:2605.03352v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated robust capabilities in recognizing everyday human activities, yet their potential for analyzi

Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding

TutorialsDGX agent

arXiv:2512.17817v3 Announce Type: replace Abstract: While 3DGS has emerged as a high-fidelity scene representation, encoding rich, general-purpose features directly from its primitives remains under-e

Conditions for well-posed color recovery in scattering media

ResearchDGX agent

arXiv:2605.03837v1 Announce Type: new Abstract: Recovering scene color from images captured in scattering media is a fundamental inverse problem in optical imaging. Yet the problem is intrinsically il

Context- and Pixel-aware Large Language Model for Video Quality Assessment

ResearchDGX agent

arXiv:2505.16025v3 Announce Type: replace Abstract: Video quality assessment (VQA) is a challenging research topic with broad applications. Traditional hand-crafted and discriminative learning-based V

CropVLM: A Domain-Adapted Vision-Language Model for Open-Set Crop Analysis

Model ReleasesDGX agent

arXiv:2605.03259v1 Announce Type: new Abstract: High-throughput plant phenotyping, the quantitative measurement of observable plant traits, is critical for modern breeding but remains constrained by a

DALPHIN: Benchmarking Digital Pathology AI Copilots Against Pathologists on an Open Multicentric Dataset

Model ReleasesDGX agent

arXiv:2605.03544v1 Announce Type: new Abstract: Foundation models with visual question answering capabilities for digital pathology are emerging. Such unprecedented technology requires independent ben

deSEO: Physics-Aware Dataset Creation for High-Resolution Satellite Image Shadow Removal

ResearchDGX agent

arXiv:2605.03610v1 Announce Type: new Abstract: Shadows cast by terrain and tall structures remain a major obstacle for high-resolution satellite image analysis, degrading classification, detection, a

Diffusion Masked Pretraining for Dynamic Point Cloud

ResearchDGX agent

arXiv:2605.03639v1 Announce Type: new Abstract: Dynamic point cloud pretraining is still dominated by masked reconstruction objectives. However, these objectives inherit two key limitations. Existing

DINO Soars: DINOv3 for Open-Vocabulary Semantic Segmentation of Remote Sensing Imagery

Model ReleasesDGX agent

arXiv:2605.03175v1 Announce Type: new Abstract: The remote sensing (RS) domain suffers from a lack of densely labeled datasets, which are costly to obtain. Thus, models that can segment RS imagery wel

DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models

SafetyDGX agent

arXiv:2605.03877v1 Announce Type: new Abstract: Dataset distillation enables efficient training by distilling the information of large-scale datasets into significantly smaller synthetic datasets. Dif

Dual-Foundation Models for Unsupervised Domain Adaptation

AgentsDGX agent

arXiv:2605.03365v1 Announce Type: new Abstract: Semantic segmentation provides pixel-level scene understanding essential for autonomous driving and fine-grained perception tasks. However, training seg

Dynamic Distillation and Gradient Consistency for Robust Long-Tailed Incremental Learning

ResearchDGX agent

arXiv:2605.03364v1 Announce Type: new Abstract: The task of Long-tailed Class Incremental Learning (LT-CIL) addresses the sequential learning of new classes from datasets with imbalanced class distrib

EMOVIS: Emotion-Optimized Image Processing

ResearchDGX agent

arXiv:2605.03131v1 Announce Type: cross Abstract: In cinematography, visual attributes such as color grading, contrast, and brightness are manipulated to reinforce the emotional narrative of a scene.

← Previous
1…152153154155156…209
Next →