AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere

DGX agent

arXiv:2602.10094v2 Announce Type: replace Abstract: We present 4RC, a unified feed-forward framework for 4D reconstruction from monocular videos. Unlike existing approaches that typically decouple mot

researcharxiv-cs-cv
6 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

A Benchmark for Interactive World Models with a Unified Action Generation Framework

DGX agent

arXiv:2605.03941v1 Announce Type: new Abstract: Achieving Artificial General Intelligence (AGI) requires agents that learn and interact adaptively, with interactive world models providing scalable env

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

A Deeper Dive into the Irreversibility of PolyProtect: Making Protected Face Templates Harder to Invert

DGX agent

arXiv:2605.03857v1 Announce Type: new Abstract: This work presents a deeper analysis of the 'irreversibility' property of PolyProtect, a biometric template protection method initially proposed for sec

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

A Framework for Exploring and Disentangling Intersectional Bias: A Case Study in Fetal Ultrasound

DGX agent

arXiv:2605.02942v1 Announce Type: cross Abstract: Bias in medical AI is often framed as a problem of representation. However, in image-based tasks such as fetal ultrasound, performance disparities can

model-releasesarxiv-cs-cv
6 May 2026
Research

A Partition-Based Generating Function for Row-Convex Polyominoes

DGX agent

arXiv:2605.03203v1 Announce Type: cross Abstract: An alternative generating function is proposed to enumerate row-convex polyominoes without internal holes on a discrete grid. The approach is based on

researcharxiv-cs-cv
6 May 2026
Safety

A Robust Unsupervised Domain Adaptation Framework for Medical Image Classification Using RKHS-MMD

DGX agent

arXiv:2605.03787v1 Announce Type: new Abstract: Labeling medical images is a major bottleneck in the field of medical imaging, as it requires domain-specific expertise, and it gets further complicated

safetyarxiv-cs-cv
6 May 2026
Research

A User-Centric Analysis of Explainability in AI-Based Medical Image Diagnosis

DGX agent

arXiv:2605.02903v1 Announce Type: cross Abstract: In recent years, AI systems in the medical domain have advanced significantly. However, despite outperforming humans, they are rarely used in practice

researcharxiv-cs-cv
6 May 2026
Safety

A Vision-Based Shared-Control Teleoperation Scheme for Controlling the Robotic Arm of a Four-Legged Robot

DGX agent

arXiv:2508.14994v3 Announce Type: replace-cross Abstract: In hazardous and remote environments, robotic systems perform critical tasks demanding improved safety and efficiency. Among these, quadruped

safetyarxiv-cs-cv
6 May 2026
Local Ai

AHPA: Adaptive Hierarchical Prior Alignment for Diffusion Transformers

DGX agent

arXiv:2605.03317v1 Announce Type: new Abstract: Representation alignment has recently emerged as an effective paradigm for accelerating Diffusion Transformer training. Despite their success, existing

local-aiarxiv-cs-cv
6 May 2026
Applications

AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock

DGX agent

arXiv:2507.22101v2 Announce Type: replace Abstract: Crops, fisheries and livestock form the backbone of global food production, essential to feed the ever-growing global population. However, these sec

applicationsarxiv-cs-cv
6 May 2026
Applications

AniMatrix: An Anime Video Generation Model that Thinks in Art, Not Physics

DGX agent

arXiv:2605.03652v1 Announce Type: new Abstract: Video generation models internalize physical realism as their prior. Anime deliberately violates physics: smears, impact frames, chibi shifts; and its t

applicationsarxiv-cs-cv
6 May 2026
Tutorials

Approaching human parity in the quality of automated organoid image segmentation

DGX agent

arXiv:2605.03053v1 Announce Type: new Abstract: Organoids are complex, three dimensional, self-organizing cell cultures which manifest organ-like features and represent a powerful platform for studyin

tutorialsarxiv-cs-cv
6 May 2026
Safety

Audio-Visual Intelligence in Large Foundation Models

DGX agent

arXiv:2605.04045v1 Announce Type: new Abstract: Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines

safetyarxiv-cs-cv
6 May 2026
Model Releases

Before Forgetting, Learn to Remember: Revisiting Foundational Learning Failures in LVLM Unlearning Benchmarks

DGX agent

arXiv:2605.03759v1 Announce Type: new Abstract: While Large Vision-Language Models (LVLMs) offer powerful capabilities, they pose privacy risks by unintentionally memorizing sensitive personal informa

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs

DGX agent

arXiv:2512.09874v2 Announce Type: replace Abstract: Correctly parsing mathematical formulas from PDFs is critical for training large language models and building scientific knowledge bases from academ

model-releasesarxiv-cs-cv
6 May 2026
Agents

BEVCALIB: LiDAR-Camera Calibration via Geometry-Guided Bird's-Eye View Representations

DGX agent

arXiv:2506.02587v2 Announce Type: replace Abstract: Accurate LiDAR-camera calibration is fundamental to fusing multi-modal perception in autonomous driving and robotic systems. Traditional calibration

agentsarxiv-cs-cv
6 May 2026
Model Releases

BFORE: Butterfly-Firefly Optimized Retinex Enhancement for Low-Light Image Quality Improvement

DGX agent

arXiv:2605.03509v1 Announce Type: new Abstract: Low-light image enhancement is a fundamental challenge in computer vision and multimedia applications, as images captured under insufficient illuminatio

model-releasesarxiv-cs-cv
6 May 2026
Research

Boundary-Aware Uncertainty Quantification for Wildfire Spread Prediction

DGX agent

arXiv:2605.03148v1 Announce Type: new Abstract: Reliable wildfire spread prediction is vital for risk-aware emergency planning, yet most deep learning models lack principled uncertainty quantification

researcharxiv-cs-cv
6 May 2026
Research

Can Multimodal Large Language Models Understand Pathologic Movements? A Pilot Study on Seizure Semiology

DGX agent

arXiv:2605.03352v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated robust capabilities in recognizing everyday human activities, yet their potential for analyzi

researcharxiv-cs-cv
6 May 2026
Tutorials

Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding

DGX agent

arXiv:2512.17817v3 Announce Type: replace Abstract: While 3DGS has emerged as a high-fidelity scene representation, encoding rich, general-purpose features directly from its primitives remains under-e

tutorialsarxiv-cs-cv
6 May 2026
Research

Conditions for well-posed color recovery in scattering media

DGX agent

arXiv:2605.03837v1 Announce Type: new Abstract: Recovering scene color from images captured in scattering media is a fundamental inverse problem in optical imaging. Yet the problem is intrinsically il

researcharxiv-cs-cv
6 May 2026
Research

Context- and Pixel-aware Large Language Model for Video Quality Assessment

DGX agent

arXiv:2505.16025v3 Announce Type: replace Abstract: Video quality assessment (VQA) is a challenging research topic with broad applications. Traditional hand-crafted and discriminative learning-based V

researcharxiv-cs-cv
6 May 2026
Model Releases

CropVLM: A Domain-Adapted Vision-Language Model for Open-Set Crop Analysis

DGX agent

arXiv:2605.03259v1 Announce Type: new Abstract: High-throughput plant phenotyping, the quantitative measurement of observable plant traits, is critical for modern breeding but remains constrained by a

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

DALPHIN: Benchmarking Digital Pathology AI Copilots Against Pathologists on an Open Multicentric Dataset

DGX agent

arXiv:2605.03544v1 Announce Type: new Abstract: Foundation models with visual question answering capabilities for digital pathology are emerging. Such unprecedented technology requires independent ben

model-releasesarxiv-cs-cv
6 May 2026
Research

deSEO: Physics-Aware Dataset Creation for High-Resolution Satellite Image Shadow Removal

DGX agent

arXiv:2605.03610v1 Announce Type: new Abstract: Shadows cast by terrain and tall structures remain a major obstacle for high-resolution satellite image analysis, degrading classification, detection, a

researcharxiv-cs-cv
6 May 2026
Research

Diffusion Masked Pretraining for Dynamic Point Cloud

DGX agent

arXiv:2605.03639v1 Announce Type: new Abstract: Dynamic point cloud pretraining is still dominated by masked reconstruction objectives. However, these objectives inherit two key limitations. Existing

researcharxiv-cs-cv
6 May 2026
Model Releases

DINO Soars: DINOv3 for Open-Vocabulary Semantic Segmentation of Remote Sensing Imagery

DGX agent

arXiv:2605.03175v1 Announce Type: new Abstract: The remote sensing (RS) domain suffers from a lack of densely labeled datasets, which are costly to obtain. Thus, models that can segment RS imagery wel

model-releasesarxiv-cs-cv
6 May 2026
Safety

DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models

DGX agent

arXiv:2605.03877v1 Announce Type: new Abstract: Dataset distillation enables efficient training by distilling the information of large-scale datasets into significantly smaller synthetic datasets. Dif

safetyarxiv-cs-cv
6 May 2026
Agents

Dual-Foundation Models for Unsupervised Domain Adaptation

DGX agent

arXiv:2605.03365v1 Announce Type: new Abstract: Semantic segmentation provides pixel-level scene understanding essential for autonomous driving and fine-grained perception tasks. However, training seg

agentsarxiv-cs-cv
6 May 2026
Research

Dynamic Distillation and Gradient Consistency for Robust Long-Tailed Incremental Learning

DGX agent

arXiv:2605.03364v1 Announce Type: new Abstract: The task of Long-tailed Class Incremental Learning (LT-CIL) addresses the sequential learning of new classes from datasets with imbalanced class distrib

researcharxiv-cs-cv
6 May 2026
Research

EMOVIS: Emotion-Optimized Image Processing

DGX agent

arXiv:2605.03131v1 Announce Type: cross Abstract: In cinematography, visual attributes such as color grading, contrast, and brightness are manipulated to reinforce the emotional narrative of a scene.

researcharxiv-cs-cv
6 May 2026
Research

Enhanced 3D Brain Tumor Segmentation Using Assorted Precision Training

DGX agent

arXiv:2605.04008v1 Announce Type: new Abstract: A brain tumor is a medical disorder faced by individuals of all demographics. Medically, it is described as the spread of non-essential cells close to o

researcharxiv-cs-cv
6 May 2026
Research

Enhancing Self-Supervised Talking Head Forgery Detection via a Training-Free Dual-System Framework

DGX agent

arXiv:2605.03390v1 Announce Type: new Abstract: Supervised talking head forgery detection faces severe generalization challenges due to the continuous evolution of generators. By reducing reliance on

researcharxiv-cs-cv
6 May 2026
Tutorials

Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation

DGX agent

arXiv:2605.03790v1 Announce Type: new Abstract: With advances in multimodal research and deep learning, Multimodal Large Language Models (MLLMs) have emerged as a powerful paradigm for a wide range of

tutorialsarxiv-cs-cv
6 May 2026
Model Releases

Erase Persona, Forget Lore: Benchmarking Multimodal Copyright Unlearning in Large Vision Language Models

DGX agent

arXiv:2605.03547v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs), trained on web-scale data, risk memorizing and regenerating copyrighted visual content such as characters and logo

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

FACTOR: Counterfactual Training-Free Test-Time Adaptation for Open-Vocabulary Object Detection

DGX agent

arXiv:2605.03294v1 Announce Type: new Abstract: Open-vocabulary object detection often fails under distribution shifts, as it can be misled by spurious correlations between non-causal visual attribute

model-releasesarxiv-cs-cv
6 May 2026
Applications

First Shape, Then Meaning: Efficient Geometry and Semantics Learning for Indoor Reconstruction

DGX agent

arXiv:2605.03463v1 Announce Type: new Abstract: Neural Surface Reconstruction has become a standard methodology for indoor 3D reconstruction, with Signed Distance Functions (SDFs) proving particularly

applicationsarxiv-cs-cv
6 May 2026
Model Releases

FluxFlow: Conservative Flow-Matching for Astronomical Image Super-Resolution

DGX agent

arXiv:2605.03749v1 Announce Type: new Abstract: Ground-to-space astronomical super-resolution requires recovering space-quality images from ground-based observations that are simultaneously limited by

model-releasesarxiv-cs-cv
6 May 2026
Research

FreeTimeGS++: Secrets of Dynamic Gaussian Splatting and Their Principles

DGX agent

arXiv:2605.03337v1 Announce Type: new Abstract: The recent surge in 4D Gaussian Splatting (4DGS) has achieved impressive dynamic scene reconstruction. While these methods demonstrate remarkable perfor

researcharxiv-cs-cv
6 May 2026
Model Releases

From Code to Prediction: Fine-Tuning LLMs for Neural Network Performance Classification in NNGPT

DGX agent

arXiv:2605.03686v1 Announce Type: cross Abstract: Automated Machine Learning (AutoML) frameworks increasingly leverage Large Language Models (LLMs) for tasks such as hyperparameter optimization and ne

model-releasesarxiv-cs-cv
6 May 2026
Safety

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

DGX agent

arXiv:2507.07982v2 Announce Type: replace Abstract: Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models trained solely on raw v

safetyarxiv-cs-cv
6 May 2026
Research

GeoTopoDiff: Learning Geometry--Topology Graph Priors through Boundary-Constrained Mixed Diffusion for Sparse-Slice 3D Porous Reconstruction

DGX agent

arXiv:2605.03764v1 Announce Type: new Abstract: Diffusion-based voxel prior modelling is challenging for the reconstruction of large-scale 3D porous microstructures. Due to the demanding requirements

researcharxiv-cs-cv
6 May 2026
Safety

GRPO-TTA: Test-Time Visual Tuning for Vision-Language Models via GRPO-Driven Reinforcement Learning

DGX agent

arXiv:2605.03403v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has recently shown strong performance in post-training large language models and vision-language models. It ra

safetyarxiv-cs-cv
6 May 2026
Local Ai

Identity-Consistent Multi-Pose Generation of Contactless Fingerprints

DGX agent

arXiv:2605.03830v1 Announce Type: new Abstract: Contactless fingerprint recognition has gained increasing attention due to its advantages in hygiene and acquisition flexibility. However, the absence o

local-aiarxiv-cs-cv
6 May 2026
Research

Illumination-Aware Contactless Fingerprint Spoof Detection via Paired Flash-Non-Flash Imaging

DGX agent

arXiv:2603.17679v2 Announce Type: replace Abstract: Contactless fingerprint recognition enables hygienic and convenient biometric authentication but poses new challenges for spoof detection due to the

researcharxiv-cs-cv
6 May 2026
Model Releases

IRIS: Intent Resolution via Inference-time Saccades for Open-Ended VQA in Large Vision-Language Models

DGX agent

arXiv:2602.16138v2 Announce Type: replace Abstract: We introduce IRIS (Intent Resolution via Inference-time Saccades), a novel training-free approach that uses eye-tracking data in real-time to resolv

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

Label-Efficient School Detection from Aerial Imagery via Weakly Supervised Pretraining and Fine-Tuning

DGX agent

arXiv:2605.03968v1 Announce Type: new Abstract: Accurate school detection is essential for supporting education initiatives, including infrastructure planning and expanding internet connectivity to un

model-releasesarxiv-cs-cv
6 May 2026
Research

LangPrecip: Language-Aware Multimodal Precipitation Nowcasting

DGX agent

arXiv:2512.22317v2 Announce Type: replace-cross Abstract: Short-term precipitation nowcasting is an inherently uncertain and under-constrained spatiotemporal forecasting problem, especially for rapidl

researcharxiv-cs-cv
6 May 2026
← Previous
1…193194195196197…263
Next →