AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
24 Jun 2026

VisChronos: Revolutionizing Image Captioning Through Real-Life Events

ApplicationsDGX agent

arXiv:2606.24058v1 Announce Type: new Abstract: This paper aims to bridge the semantic gap between visual content and natural language understanding by leveraging historical events in the real world a

VisCritic: Visual State Comparison as Process Reward for GUI Agents

Model ReleasesDGX agent

arXiv:2606.24525v1 Announce Type: new Abstract: GUI agents powered by vision-language models show strong potential for automating digital tasks, yet frequently fail in long-horizon scenarios due to th

VistaRef: Boosting Visual Spatial Orientation Awareness for Pointing-to-Object Detection

Local AiDGX agent

arXiv:2606.24498v1 Announce Type: new Abstract: Grounding deictic gestures in natural images is fundamental to AR and human-robot collaboration, providing a basis for seamless spatial interaction. Whi


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering

ApplicationsDGX agent

arXiv:2606.24602v1 Announce Type: new Abstract: Despite remarkable progress in multimodal understanding, current MLLMs still exhibit limitations in video text understanding, particularly when semantic

VSANet: View-aware Sparse Attention Network for Light Field Image Denoising

ResearchDGX agent

arXiv:2606.24737v1 Announce Type: new Abstract: Light field (LF) image denoising is challenging due to the high-dimensional structure of LF data. While noise is independent across sub-aperture images,

What Do Flow-Based Inverse Solvers Approximate? A Posterior-Transport View

SafetyDGX agent

arXiv:2606.24516v1 Announce Type: new Abstract: A growing family of training-free solvers -- FlowDPS, FLOWER, PnP-Flow and their diffusion ancestors (DPS, DAPS) -- repurpose a pretrained flow-matching

23 Jun 2026

2D Versus 3D Diffusion for In Silico Training of Interventional X-ray AI Models

ResearchDGX agent

arXiv:2606.21414v1 Announce Type: cross Abstract: The ability to synthesize realistic X-ray images has catalyzed the development of AI models for X-ray image-guided procedures, which otherwise suffer

360Anything: Geometry-Free Lifting of Images and Videos to 360{eg}

SafetyDGX agent

arXiv:2601.16192v2 Announce Type: replace Abstract: Lifting perspective images and videos to 360{eg} panoramas enables immersive 3D world generation. Existing approaches often rely on explicit geometr

3D Vessel Reconstruction from Sparse-View Dynamic DSA Images via Vessel Probability Guided Attenuation Learning

AgentsDGX agent

arXiv:2405.10705v3 Announce Type: replace-cross Abstract: Digital Subtraction Angiography (DSA) is one of the gold standards for vascular disease diagnosis. With the help of a contrast agent, time-res

4DVLT: Dynamic Scene Understanding with Worldline-Centered Vision-Language Tracking

Model ReleasesDGX agent

arXiv:2606.22631v1 Announce Type: new Abstract: 4D dynamic scene understanding requires grounding language to a persistent worldline that binds identity, metric 3D motion, and synchronized multi-view

A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models

ResearchDGX agent

arXiv:2511.15098v2 Announce Type: replace Abstract: Discrete diffusion-based multimodal large language models (dMLLMs) have emerged as a promising alternative to autoregressive MLLMs thanks to their a

A Controlled Study of CLIP-Based Body-Scene Fusion for Emotion Recognition in Context

SafetyDGX agent

arXiv:2606.22072v1 Announce Type: new Abstract: Apparent emotion in natural images is often not visible from the face alone. The face may be small, hidden, or neutral, while posture and scene context

A DVDrive Approach for doScenes Instructed Driving Challenge

Local AiDGX agent

arXiv:2606.21623v1 Announce Type: new Abstract: Instruction-conditioned trajectory prediction is an emerging problem in autonomous driving, where a model predicts the future ego trajectory not only fr

A Latent Representation Learning Framework for Hyperspectral Image Emulation in Remote Sensing

Model ReleasesDGX agent

arXiv:2603.21911v2 Announce Type: replace Abstract: Synthetic hyperspectral image (HSI) generation is essential for large-scale simulation, algorithm development, and mission design, yet traditional r

A Linear Fractional Transformation Model and Calibration Method for Light Field Camera

Model ReleasesDGX agent

arXiv:2511.03962v2 Announce Type: replace Abstract: Accurate intrinsic calibration is a crucial yet challenging prerequisite for 3D reconstruction using light field cameras. Existing calibration model

A Neurosymbolic Framework for Interpretable Skeleton-Based Seizure Detection via Concept-Driven Logical Reasoning

ResearchDGX agent

arXiv:2606.21252v1 Announce Type: new Abstract: Video-based seizure detection is essential for the management of epilepsy patients, offering a non-invasive complement to electroencephalography. While

A Physics-Informed, Behavior-Aware Digital Twin for Robust Multimodal Forecasting of Core Body Temperature in Precision Livestock Farming

ApplicationsDGX agent

arXiv:2604.04098v2 Announce Type: replace Abstract: Precision livestock farming requires accurate and timely heat stress prediction to ensure animal welfare and optimize farm management. This study pr

A Projection-Based Surrogate Gradient Interpretation for Neural Codec Wrappers

ResearchDGX agent

arXiv:2606.20671v1 Announce Type: new Abstract: Neural wrappers are learned pre-and postprocessing networks designed to enhance the performance of conventional video codecs. Although these approaches

A Skin-Tone-Aware Dual-Representation Remote Photoplethysmography Framework for Contactless Respiratory Rate Estimation

Model ReleasesDGX agent

arXiv:2606.21511v1 Announce Type: cross Abstract: Respiratory rate is a vital indicator of pulmonary and cardiovascular health, yet conventional methods for estimating respiratory rate are often intru

A Smart Classroom Behavior Analysis Framework with a New Highly Congested Classroom Dataset

Model ReleasesDGX agent

arXiv:2606.21568v1 Announce Type: new Abstract: Student behavior detection is important for intelligent classroom analysis but remains challenging in large-class scenarios due to dense instance co-occ

A Test-time Actor-Critic Approach to News Images Generation

ResearchDGX agent

arXiv:2606.21304v1 Announce Type: new Abstract: This paper introduces the CERTH-ITI solution for the MediaEval NewsImages 2026 challenge, which focuses on generating images related to news headlines.

A UAV-Based Multi-Modal Vision System for Automated Sideslope Deformation Monitoring and Hazard Detection

SafetyDGX agent

arXiv:2606.20681v1 Announce Type: new Abstract: Slope hazards constitute a major safety threat to expressway infrastructure, and their evolution is typically manifested as slow surface deformation. Co

A Viscosity Semigroup Framework for Stable Image Reconstruction

ResearchDGX agent

arXiv:2606.20620v1 Announce Type: new Abstract: Starting from the axiomatic formulation of scale-space theory, we develop a viscosity-solution framework for multiscale image representations arising fr

Accurate identification and measurement of the precipitate area by two-stage deep neural networks in novel chromium-based alloys

ResearchDGX agent

arXiv:2606.22112v1 Announce Type: new Abstract: The performance of advanced materials for extreme environments is underpinned by their microstructure, including the size and distribution of reinforcin

ACE-GS: Acing the Trade-off with Accurate, Compact and Efficient 3D Gaussian Splatting

Model ReleasesDGX agent

arXiv:2606.21244v1 Announce Type: new Abstract: 3D Gaussian Splatting achieves exceptional real-time rendering, but its substantial computational and storage demands hinder widespread deployment. Exis

Adaptive Beam Selection for Efficient Scanning Probe Tomography

ResearchDGX agent

arXiv:2606.21713v1 Announce Type: cross Abstract: In X-ray tomography, reconstruction quality generally improves with larger numbers of projections. However, more projections increase experiment costs

Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization

TutorialsDGX agent

arXiv:2606.21736v1 Announce Type: new Abstract: Single domain generalization (SDG) aims to learn a robust model, which could perform well on many unseen domains while there is only one single domain a

AEF-Econ: Toward Plug-and-Play Socioeconomic Foundation Embeddings from AlphaEarth for Urban Remote Sensing

Model ReleasesDGX agent

arXiv:2606.20697v1 Announce Type: new Abstract: AlphaEarth Foundations (AEF) unify global remote sensing foundation embeddings through multimodal self-supervised learning, but their pretraining focuse

AgroSense 2.0: Cross-Modal Transformer Fusion with Geospatial Raster Integration and Interpretable Multi-Task Learning for Precision Crop Recommendation

ResearchDGX agent

arXiv:2606.21892v1 Announce Type: cross Abstract: Crop recommendation systems in precision agriculture have long suffered from a fundamental modality gap: visual soil characterization and chemical nut

AI-Augmented Thyroid Scintigraphy for Robust Classification of Disease

ResearchDGX agent

arXiv:2503.00366v3 Announce Type: replace-cross Abstract: Thyroid scintigraphy is vital for diagnosing thyroid disorders, yet deep learning (DL) models in this domain often struggle with limited, imba

AIR: Adaptive Interleaved Reasoning with Code in MLLMs

AgentsDGX agent

arXiv:2606.23678v1 Announce Type: new Abstract: Following the paradigm shift initiated by OpenAI o3, interleaved reasoning with code to enhance multimodal large language models (MLLMs) has become a pi

An approach with Visual and Tabular Mamba to multimodal medical data using Mixed Fusion

ResearchDGX agent

arXiv:2606.20738v1 Announce Type: new Abstract: This article presents a complementary approach for integrating multimodal medical data in cancer classification, based on state space models represented

Anatomically Consistent TMJ Disc Segmentation via Semantic Anchoring and Clinical Priors

Local AiDGX agent

arXiv:2606.21177v1 Announce Type: cross Abstract: Segmenting the temporomandibular joint (TMJ) disc from MRI is essential for accurate diagnosis of internal derangement, yet it remains unreliable in p

Arbor: Explicit Geometric Conditioning for Controllable 3D Asset Generation

ResearchDGX agent

arXiv:2606.23514v1 Announce Type: new Abstract: Text and image conditioned 3D models now generate convincing assets, but they still offer little direct control over the space an object should occupy o

Arc-Length Parameterized Interpolating Splines

ResearchDGX agent

arXiv:2606.21209v1 Announce Type: cross Abstract: We present an iterative algorithm to compute an arc-length parameterized spline interpolating a set of points. This differs from other methods where t

ARGUSTRACK: A Multi-View Annotation System for Multi-Object Tracking

SafetyDGX agent

arXiv:2606.20687v1 Announce Type: new Abstract: Multi-Camera Multi-Target (MCMT) tracking has emerged as a critical capability for applications ranging from autonomous driving to animal behavior monit

Artic-O: End-to-End Articulated Object Reconstruction via Latent Geometry Learning

ResearchDGX agent

arXiv:2606.21938v1 Announce Type: new Abstract: Reconstructing articulated objects from sparse images requires recovering complete geometry, movable parts, and motion parameters. Recent methods typica

ASCII Art Turns LLMs into VLA Controllers

Model ReleasesDGX agent

arXiv:2606.21470v1 Announce Type: cross Abstract: Vision--Language--Action (VLA) controllers are often built by extending vision--language models (VLMs) with action supervision, relying on multimodal

Attention at Rest Stays at Rest: Breaking Visual Inertia for Cognitive Hallucination Mitigation

ResearchDGX agent

arXiv:2604.01989v3 Announce Type: replace Abstract: Like a body at rest that stays at rest, we find that visual attention in multimodal large language models (MLLMs) exhibits pronounced inertia, remai

Attention-Spectrum Regularization for Replay-Free Continual Multimodal LLMs

ResearchDGX agent

arXiv:2606.23063v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly required to adapt to non-stationary streams of visual domains, question types, and user instru

Automated sign detection across the Electronic Babylonian Library: A large-scale dataset and end-to-end cuneiform OCR pipeline

ResearchDGX agent

arXiv:2606.22608v1 Announce Type: new Abstract: Learning to read cuneiform tablets is an extremely demanding task; consequently, of the roughly half million excavated tablets, only a small fraction ha

Autonomous and Self-Adapting System for Synthetic Media Detection and Attribution

AgentsDGX agent

arXiv:2504.03615v2 Announce Type: replace Abstract: Rapid advances in generative AI have enabled the creation of highly realistic synthetic images, which, while beneficial in many domains, also pose s

Autonomous Subsea Cable Search and Tracking with Graph-Optimised Priors and Visual Tracking

AgentsDGX agent

arXiv:2606.23606v1 Announce Type: cross Abstract: Global communications rely on subsea cable infrastructure that remains vulnerable to damage from natural hazards and human activity. Autonomous underw

AwakeForest: An Interactive Geospatial Platform for Large-Scale Forest Imagery

ResearchDGX agent

arXiv:2606.23542v1 Announce Type: new Abstract: Forest imagery analysis often involves multiple tightly coupled vision tasks, which must be performed under substantial variation in geographic regions,

BAC-JEPA: Label-Efficient Breast Arterial Calcification Segmentation via Synthetic Mammography-Guided Supervision

HardwareDGX agent

arXiv:2606.22089v1 Announce Type: new Abstract: Breast arterial calcification (BAC) on screening mammograms is an emerging cardiovascular risk biomarker, but quantitative use requires reproducible seg

BadDreamer: Transferable Backdoor Attacks against Video World Models for Autonomous Driving

SafetyDGX agent

arXiv:2606.21172v1 Announce Type: new Abstract: Video world models are increasingly used in autonomous driving to forecast future scene evolution and provide future-aware spatio-temporal representatio

BELDE: Building a Large-scale Earth-observation Land-cover Dataset for Europe

Model ReleasesDGX agent

arXiv:2606.20909v1 Announce Type: new Abstract: Earth observation imagery plays a critical role in environmental monitoring, urban planning, disaster assessment, and climate analysis. While multi-spec

Benchmarking Vision-Language Models for Microscopic Plant Image Understanding

Model ReleasesDGX agent

arXiv:2606.22497v1 Announce Type: new Abstract: Microscopic imaging provides essential visual evidence for studying plant biology and pathology at the cellular and subcellular levels. However, existin

BEV-Denoise: Learning Intrinsic Noise for Accurate Bird's-Eye-View Semantic Segmentation

ApplicationsDGX agent

arXiv:2606.22931v1 Announce Type: new Abstract: In this paper, we present a framework dubbed extbf{BEV-Denoise} that estimates and removes intrinsic noise from learned Bird's-Eye-View (BEV) features t

Beyond a Single Light: A Large-Scale Aerial Dataset for Urban Scene Reconstruction Under Varying Illumination

ApplicationsDGX agent

arXiv:2512.14200v2 Announce Type: replace Abstract: Recent advances in Neural Radiance Fields and 3D Gaussian Splatting have demonstrated strong potential for large-scale UAV-based 3D reconstruction t

Beyond Damage Assessment: Recyclable Material Detection in Aerial Disaster Imagery Using a Lightweight Patch-Based Framework

ResearchDGX agent

arXiv:2606.21279v1 Announce Type: new Abstract: Nowadays, more and more disasters of different natures are appearing. Several disaster assessment approaches have been developed in order to identify da

Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification

ResearchDGX agent

arXiv:2606.21838v1 Announce Type: new Abstract: Multimodal contrastive learning has enabled zero-shot visual classification by aligning images with textual categories. However, in hierarchically struc

Beyond ROC-AUC: Operating-Point Performance Reporting for Biometric Verification

ResearchDGX agent

arXiv:2606.20680v1 Announce Type: new Abstract: A biometric verifier is often deployed with a strict false match budget, so only a narrow, low false match rate (FMR) slice of the score range is used.

Beyond Templates: Revisiting Zero-Shot Remote Sensing through Meta-Prompting

TutorialsDGX agent

arXiv:2606.20702v1 Announce Type: new Abstract: Vision-language models (VLMs) have sparked growing interest in zero-shot Earth Observation (EO) downstream tasks, with further gains enabled by remote-s

Beyond the LUMIR challenge: The pathway to foundational registration models

Model ReleasesDGX agent

arXiv:2505.24160v3 Announce Type: replace-cross Abstract: Medical image challenges have played a transformative role in advancing the field, catalyzing innovation and establishing new performance benc

BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation

Model ReleasesDGX agent

arXiv:2511.22973v2 Announce Type: replace Abstract: Long video generation is a critical step toward building realistic world models, requiring both high visual fidelity and long-range interaction cons

Biological Sex Determination in Cadavers Using Deep Learning Algorithms from Computed Tomography Images of Pelvis and Skull

ResearchDGX agent

arXiv:2606.22515v1 Announce Type: new Abstract: Sexual identification of decomposed cadavers challenges traditional methods dependent on visual anthropological analysis. This study evaluates state-of-

Black-Box Continual Learning for Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.22999v1 Announce Type: new Abstract: The rapid deployment of Vision-Language Models (VLMs) in dynamic environments necessitates the ability to learn continuously without forgetting. However

Boosting Neural Video Codec via Scale-Driven Online Flow Refinement

ResearchDGX agent

arXiv:2606.23023v1 Announce Type: new Abstract: Although state-of-the-art neural video codecs (NVCs) have achieved remarkable performance, they suffer from limited generalization when encountering com

Boundary-by-Mask: Few-Shot Instance Segmentation with Mask-Conditioned Boundary Learning for Texture-Poor Industrial Parts

ResearchDGX agent

arXiv:2606.21594v1 Announce Type: new Abstract: Recent advances in large pre-trained models have led to remarkable progress in instance segmentation on general images. However, industrial scenarios re

← Previous
1…7576777879…211
Next →