AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Applications

VisChronos: Revolutionizing Image Captioning Through Real-Life Events

DGX agent

arXiv:2606.24058v1 Announce Type: new Abstract: This paper aims to bridge the semantic gap between visual content and natural language understanding by leveraging historical events in the real world a

applicationsarxiv-cs-cv
24 Jun 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

VisCritic: Visual State Comparison as Process Reward for GUI Agents

DGX agent

arXiv:2606.24525v1 Announce Type: new Abstract: GUI agents powered by vision-language models show strong potential for automating digital tasks, yet frequently fail in long-horizon scenarios due to th

model-releasesarxiv-cs-cv
24 Jun 2026
Local Ai

VistaRef: Boosting Visual Spatial Orientation Awareness for Pointing-to-Object Detection

DGX agent

arXiv:2606.24498v1 Announce Type: new Abstract: Grounding deictic gestures in natural images is fundamental to AR and human-robot collaboration, providing a basis for seamless spatial interaction. Whi

local-aiarxiv-cs-cv
24 Jun 2026
Applications

ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering

DGX agent

arXiv:2606.24602v1 Announce Type: new Abstract: Despite remarkable progress in multimodal understanding, current MLLMs still exhibit limitations in video text understanding, particularly when semantic

applicationsarxiv-cs-cv
24 Jun 2026
Research

VSANet: View-aware Sparse Attention Network for Light Field Image Denoising

DGX agent

arXiv:2606.24737v1 Announce Type: new Abstract: Light field (LF) image denoising is challenging due to the high-dimensional structure of LF data. While noise is independent across sub-aperture images,

researcharxiv-cs-cv
24 Jun 2026
Safety

What Do Flow-Based Inverse Solvers Approximate? A Posterior-Transport View

DGX agent

arXiv:2606.24516v1 Announce Type: new Abstract: A growing family of training-free solvers -- FlowDPS, FLOWER, PnP-Flow and their diffusion ancestors (DPS, DAPS) -- repurpose a pretrained flow-matching

safetyarxiv-cs-cv
24 Jun 2026
Research

2D Versus 3D Diffusion for In Silico Training of Interventional X-ray AI Models

DGX agent

arXiv:2606.21414v1 Announce Type: cross Abstract: The ability to synthesize realistic X-ray images has catalyzed the development of AI models for X-ray image-guided procedures, which otherwise suffer

researcharxiv-cs-cv
23 Jun 2026
Safety

360Anything: Geometry-Free Lifting of Images and Videos to 360{eg}

DGX agent

arXiv:2601.16192v2 Announce Type: replace Abstract: Lifting perspective images and videos to 360{eg} panoramas enables immersive 3D world generation. Existing approaches often rely on explicit geometr

safetyarxiv-cs-cv
23 Jun 2026
Agents

3D Vessel Reconstruction from Sparse-View Dynamic DSA Images via Vessel Probability Guided Attenuation Learning

DGX agent

arXiv:2405.10705v3 Announce Type: replace-cross Abstract: Digital Subtraction Angiography (DSA) is one of the gold standards for vascular disease diagnosis. With the help of a contrast agent, time-res

agentsarxiv-cs-cv
23 Jun 2026
Model Releases

4DVLT: Dynamic Scene Understanding with Worldline-Centered Vision-Language Tracking

DGX agent

arXiv:2606.22631v1 Announce Type: new Abstract: 4D dynamic scene understanding requires grounding language to a persistent worldline that binds identity, metric 3D motion, and synchronized multi-view

model-releasesarxiv-cs-cv
23 Jun 2026
Research

A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models

DGX agent

arXiv:2511.15098v2 Announce Type: replace Abstract: Discrete diffusion-based multimodal large language models (dMLLMs) have emerged as a promising alternative to autoregressive MLLMs thanks to their a

researcharxiv-cs-cv
23 Jun 2026
Safety

A Controlled Study of CLIP-Based Body-Scene Fusion for Emotion Recognition in Context

DGX agent

arXiv:2606.22072v1 Announce Type: new Abstract: Apparent emotion in natural images is often not visible from the face alone. The face may be small, hidden, or neutral, while posture and scene context

safetyarxiv-cs-cv
23 Jun 2026
Local Ai

A DVDrive Approach for doScenes Instructed Driving Challenge

DGX agent

arXiv:2606.21623v1 Announce Type: new Abstract: Instruction-conditioned trajectory prediction is an emerging problem in autonomous driving, where a model predicts the future ego trajectory not only fr

local-aiarxiv-cs-cv
23 Jun 2026
Model Releases

A Latent Representation Learning Framework for Hyperspectral Image Emulation in Remote Sensing

DGX agent

arXiv:2603.21911v2 Announce Type: replace Abstract: Synthetic hyperspectral image (HSI) generation is essential for large-scale simulation, algorithm development, and mission design, yet traditional r

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

A Linear Fractional Transformation Model and Calibration Method for Light Field Camera

DGX agent

arXiv:2511.03962v2 Announce Type: replace Abstract: Accurate intrinsic calibration is a crucial yet challenging prerequisite for 3D reconstruction using light field cameras. Existing calibration model

model-releasesarxiv-cs-cv
23 Jun 2026
Research

A Neurosymbolic Framework for Interpretable Skeleton-Based Seizure Detection via Concept-Driven Logical Reasoning

DGX agent

arXiv:2606.21252v1 Announce Type: new Abstract: Video-based seizure detection is essential for the management of epilepsy patients, offering a non-invasive complement to electroencephalography. While

researcharxiv-cs-cv
23 Jun 2026
Applications

A Physics-Informed, Behavior-Aware Digital Twin for Robust Multimodal Forecasting of Core Body Temperature in Precision Livestock Farming

DGX agent

arXiv:2604.04098v2 Announce Type: replace Abstract: Precision livestock farming requires accurate and timely heat stress prediction to ensure animal welfare and optimize farm management. This study pr

applicationsarxiv-cs-cv
23 Jun 2026
Research

A Projection-Based Surrogate Gradient Interpretation for Neural Codec Wrappers

DGX agent

arXiv:2606.20671v1 Announce Type: new Abstract: Neural wrappers are learned pre-and postprocessing networks designed to enhance the performance of conventional video codecs. Although these approaches

researcharxiv-cs-cv
23 Jun 2026
Model Releases

A Skin-Tone-Aware Dual-Representation Remote Photoplethysmography Framework for Contactless Respiratory Rate Estimation

DGX agent

arXiv:2606.21511v1 Announce Type: cross Abstract: Respiratory rate is a vital indicator of pulmonary and cardiovascular health, yet conventional methods for estimating respiratory rate are often intru

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

A Smart Classroom Behavior Analysis Framework with a New Highly Congested Classroom Dataset

DGX agent

arXiv:2606.21568v1 Announce Type: new Abstract: Student behavior detection is important for intelligent classroom analysis but remains challenging in large-class scenarios due to dense instance co-occ

model-releasesarxiv-cs-cv
23 Jun 2026
Research

A Test-time Actor-Critic Approach to News Images Generation

DGX agent

arXiv:2606.21304v1 Announce Type: new Abstract: This paper introduces the CERTH-ITI solution for the MediaEval NewsImages 2026 challenge, which focuses on generating images related to news headlines.

researcharxiv-cs-cv
23 Jun 2026
Safety

A UAV-Based Multi-Modal Vision System for Automated Sideslope Deformation Monitoring and Hazard Detection

DGX agent

arXiv:2606.20681v1 Announce Type: new Abstract: Slope hazards constitute a major safety threat to expressway infrastructure, and their evolution is typically manifested as slow surface deformation. Co

safetyarxiv-cs-cv
23 Jun 2026
Research

A Viscosity Semigroup Framework for Stable Image Reconstruction

DGX agent

arXiv:2606.20620v1 Announce Type: new Abstract: Starting from the axiomatic formulation of scale-space theory, we develop a viscosity-solution framework for multiscale image representations arising fr

researcharxiv-cs-cv
23 Jun 2026
Research

Accurate identification and measurement of the precipitate area by two-stage deep neural networks in novel chromium-based alloys

DGX agent

arXiv:2606.22112v1 Announce Type: new Abstract: The performance of advanced materials for extreme environments is underpinned by their microstructure, including the size and distribution of reinforcin

researcharxiv-cs-cv
23 Jun 2026
Model Releases

ACE-GS: Acing the Trade-off with Accurate, Compact and Efficient 3D Gaussian Splatting

DGX agent

arXiv:2606.21244v1 Announce Type: new Abstract: 3D Gaussian Splatting achieves exceptional real-time rendering, but its substantial computational and storage demands hinder widespread deployment. Exis

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Adaptive Beam Selection for Efficient Scanning Probe Tomography

DGX agent

arXiv:2606.21713v1 Announce Type: cross Abstract: In X-ray tomography, reconstruction quality generally improves with larger numbers of projections. However, more projections increase experiment costs

researcharxiv-cs-cv
23 Jun 2026
Tutorials

Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization

DGX agent

arXiv:2606.21736v1 Announce Type: new Abstract: Single domain generalization (SDG) aims to learn a robust model, which could perform well on many unseen domains while there is only one single domain a

tutorialsarxiv-cs-cv
23 Jun 2026
Model Releases

AEF-Econ: Toward Plug-and-Play Socioeconomic Foundation Embeddings from AlphaEarth for Urban Remote Sensing

DGX agent

arXiv:2606.20697v1 Announce Type: new Abstract: AlphaEarth Foundations (AEF) unify global remote sensing foundation embeddings through multimodal self-supervised learning, but their pretraining focuse

model-releasesarxiv-cs-cv
23 Jun 2026
Research

AgroSense 2.0: Cross-Modal Transformer Fusion with Geospatial Raster Integration and Interpretable Multi-Task Learning for Precision Crop Recommendation

DGX agent

arXiv:2606.21892v1 Announce Type: cross Abstract: Crop recommendation systems in precision agriculture have long suffered from a fundamental modality gap: visual soil characterization and chemical nut

researcharxiv-cs-cv
23 Jun 2026
Research

AI-Augmented Thyroid Scintigraphy for Robust Classification of Disease

DGX agent

arXiv:2503.00366v3 Announce Type: replace-cross Abstract: Thyroid scintigraphy is vital for diagnosing thyroid disorders, yet deep learning (DL) models in this domain often struggle with limited, imba

researcharxiv-cs-cv
23 Jun 2026
Agents

AIR: Adaptive Interleaved Reasoning with Code in MLLMs

DGX agent

arXiv:2606.23678v1 Announce Type: new Abstract: Following the paradigm shift initiated by OpenAI o3, interleaved reasoning with code to enhance multimodal large language models (MLLMs) has become a pi

agentsarxiv-cs-cv
23 Jun 2026
Research

An approach with Visual and Tabular Mamba to multimodal medical data using Mixed Fusion

DGX agent

arXiv:2606.20738v1 Announce Type: new Abstract: This article presents a complementary approach for integrating multimodal medical data in cancer classification, based on state space models represented

researcharxiv-cs-cv
23 Jun 2026
Local Ai

Anatomically Consistent TMJ Disc Segmentation via Semantic Anchoring and Clinical Priors

DGX agent

arXiv:2606.21177v1 Announce Type: cross Abstract: Segmenting the temporomandibular joint (TMJ) disc from MRI is essential for accurate diagnosis of internal derangement, yet it remains unreliable in p

local-aiarxiv-cs-cv
23 Jun 2026
Research

Arbor: Explicit Geometric Conditioning for Controllable 3D Asset Generation

DGX agent

arXiv:2606.23514v1 Announce Type: new Abstract: Text and image conditioned 3D models now generate convincing assets, but they still offer little direct control over the space an object should occupy o

researcharxiv-cs-cv
23 Jun 2026
Research

Arc-Length Parameterized Interpolating Splines

DGX agent

arXiv:2606.21209v1 Announce Type: cross Abstract: We present an iterative algorithm to compute an arc-length parameterized spline interpolating a set of points. This differs from other methods where t

researcharxiv-cs-cv
23 Jun 2026
Safety

ARGUSTRACK: A Multi-View Annotation System for Multi-Object Tracking

DGX agent

arXiv:2606.20687v1 Announce Type: new Abstract: Multi-Camera Multi-Target (MCMT) tracking has emerged as a critical capability for applications ranging from autonomous driving to animal behavior monit

safetyarxiv-cs-cv
23 Jun 2026
Research

Artic-O: End-to-End Articulated Object Reconstruction via Latent Geometry Learning

DGX agent

arXiv:2606.21938v1 Announce Type: new Abstract: Reconstructing articulated objects from sparse images requires recovering complete geometry, movable parts, and motion parameters. Recent methods typica

researcharxiv-cs-cv
23 Jun 2026
Model Releases

ASCII Art Turns LLMs into VLA Controllers

DGX agent

arXiv:2606.21470v1 Announce Type: cross Abstract: Vision--Language--Action (VLA) controllers are often built by extending vision--language models (VLMs) with action supervision, relying on multimodal

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Attention at Rest Stays at Rest: Breaking Visual Inertia for Cognitive Hallucination Mitigation

DGX agent

arXiv:2604.01989v3 Announce Type: replace Abstract: Like a body at rest that stays at rest, we find that visual attention in multimodal large language models (MLLMs) exhibits pronounced inertia, remai

researcharxiv-cs-cv
23 Jun 2026
Research

Attention-Spectrum Regularization for Replay-Free Continual Multimodal LLMs

DGX agent

arXiv:2606.23063v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly required to adapt to non-stationary streams of visual domains, question types, and user instru

researcharxiv-cs-cv
23 Jun 2026
Research

Automated sign detection across the Electronic Babylonian Library: A large-scale dataset and end-to-end cuneiform OCR pipeline

DGX agent

arXiv:2606.22608v1 Announce Type: new Abstract: Learning to read cuneiform tablets is an extremely demanding task; consequently, of the roughly half million excavated tablets, only a small fraction ha

researcharxiv-cs-cv
23 Jun 2026
Agents

Autonomous and Self-Adapting System for Synthetic Media Detection and Attribution

DGX agent

arXiv:2504.03615v2 Announce Type: replace Abstract: Rapid advances in generative AI have enabled the creation of highly realistic synthetic images, which, while beneficial in many domains, also pose s

agentsarxiv-cs-cv
23 Jun 2026
Agents

Autonomous Subsea Cable Search and Tracking with Graph-Optimised Priors and Visual Tracking

DGX agent

arXiv:2606.23606v1 Announce Type: cross Abstract: Global communications rely on subsea cable infrastructure that remains vulnerable to damage from natural hazards and human activity. Autonomous underw

agentsarxiv-cs-cv
23 Jun 2026
Research

AwakeForest: An Interactive Geospatial Platform for Large-Scale Forest Imagery

DGX agent

arXiv:2606.23542v1 Announce Type: new Abstract: Forest imagery analysis often involves multiple tightly coupled vision tasks, which must be performed under substantial variation in geographic regions,

researcharxiv-cs-cv
23 Jun 2026
Hardware

BAC-JEPA: Label-Efficient Breast Arterial Calcification Segmentation via Synthetic Mammography-Guided Supervision

DGX agent

arXiv:2606.22089v1 Announce Type: new Abstract: Breast arterial calcification (BAC) on screening mammograms is an emerging cardiovascular risk biomarker, but quantitative use requires reproducible seg

hardwarearxiv-cs-cv
23 Jun 2026
Safety

BadDreamer: Transferable Backdoor Attacks against Video World Models for Autonomous Driving

DGX agent

arXiv:2606.21172v1 Announce Type: new Abstract: Video world models are increasingly used in autonomous driving to forecast future scene evolution and provide future-aware spatio-temporal representatio

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

BELDE: Building a Large-scale Earth-observation Land-cover Dataset for Europe

DGX agent

arXiv:2606.20909v1 Announce Type: new Abstract: Earth observation imagery plays a critical role in environmental monitoring, urban planning, disaster assessment, and climate analysis. While multi-spec

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Benchmarking Vision-Language Models for Microscopic Plant Image Understanding

DGX agent

arXiv:2606.22497v1 Announce Type: new Abstract: Microscopic imaging provides essential visual evidence for studying plant biology and pathology at the cellular and subcellular levels. However, existin

model-releasesarxiv-cs-cv
23 Jun 2026
← Previous
1…9495969798…263
Next →