AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
29 Jul 2026

Open-Ended CT Volume Segmentation with Weak Supervision from Language

ResearchDGX agent

arXiv:2607.25860v1 Announce Type: new Abstract: We introduce a method for training a text-conditioned segmentation model for CT scans, which combines voxel-level supervision with coarse but scalable s

OpenPVMapper: A Multi-source, Nationwide Database of Rooftop Photovoltaic Systems in France

Model ReleasesDGX agent

arXiv:2607.25153v1 Announce Type: new Abstract: Rooftop photovoltaic (PV) systems account for the vast majority of PV grid connections, yet no open, comprehensive, installation-level dataset of these

OrthKD: Extracting Generalized Clinical Knowledge from Heterogeneous Teachers for Lightweight Deployment

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.25545v1 Announce Type: cross Abstract: Deploying diabetic retinopathy (DR) screening models in primary care requires edge-efficient systems that remain accurate, safe, and reliable under do

PanoLess: Environment Reconstruction from Partial Reflective Views

Model ReleasesDGX agent

arXiv:2607.25362v1 Announce Type: new Abstract: Reflections from shiny objects and glass facades naturally extend the field of view of a camera, capturing the surrounding environment without the need

Parallel Decoding Distillation for Fast Image and Video Generation

Model ReleasesDGX agent

arXiv:2607.26004v1 Announce Type: new Abstract: Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2607.24957v1 Announce Type: new Abstract: We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Model

Quasi-SVD: Learning a Lie-constrained matrix factorisation for real-time imaging

HardwareDGX agent

arXiv:2607.25967v1 Announce Type: new Abstract: Singular Value Decomposition (SVD) underlies matrix factorisation tasks across computational imaging, with medical applications increasingly demanding r

RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection

Model ReleasesDGX agent

arXiv:2607.25392v1 Announce Type: new Abstract: We introduce RDVSv2, a large-scale benchmark for RGB-D video salient object detection (RGB-D VSOD) with dense frame-level annotations. Existing datasets

Reading Legends on Ancient Coins: An Object Detection Approach for Character Recognition on a Novel Roman Republican Dataset

ResearchDGX agent

arXiv:2607.25455v1 Announce Type: new Abstract: When it comes to the proper classification of ancient coins with respect to their time and issuer, the textual inscriptions on these coins, also known a

ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition

Model ReleasesDGX agent

arXiv:2607.25565v1 Announce Type: new Abstract: Recovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, yet remains challenging since edita

Safety-Aware Cascaded Inference for Crop Damage Assessment with Controlled Error Trade-offs

SafetyDGX agent

arXiv:2607.25468v1 Announce Type: new Abstract: In picture-based agricultural insurance for smallholder farmers, missed damage detections carry substantially higher cost than false alarms: a farmer wh

SAM-MI: A Mask-Injected Framework for Enhancing Open-Vocabulary Semantic Segmentation with SAM

Model ReleasesDGX agent

arXiv:2511.20027v2 Announce Type: replace Abstract: Open-vocabulary semantic segmentation (OVSS) aims to segment and recognize objects universally. Trained on extensive high-quality segmentation data,

Schrodinger's Cat: Probabilistic Representation and Prediction of Potential Scene Kinematics

ResearchDGX agent

arXiv:2607.25984v1 Announce Type: new Abstract: Predicting how a scene may evolve from partial observations requires reasoning about multiple possible futures rather than committing to a single trajec

ScoreShield: Differentially Private Release of Similarity Scores

Model ReleasesDGX agent

arXiv:2607.25041v1 Announce Type: cross Abstract: A growing number of applications, such as biometrics and retrieval-augmented generation (RAG), rely on cosine similarity scores computed between vecto

Sense it with your eyes: Sensation Generation and Understanding for Advertisements

Model ReleasesDGX agent

arXiv:2607.25314v1 Announce Type: new Abstract: Sensory advertising evokes human senses through visual cues, enabling audiences to mentally simulate experiences and increasing persuasive impact. Despi

SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models

ResearchDGX agent

arXiv:2607.25818v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs), such as Qwen2.5-VL and InternVL3, generate large numbers of vision tokens for high-resolution inputs, l

SurgSLOT: Segment Anything in Surgical Videos via Semantic Long-term Tracking

Model ReleasesDGX agent

arXiv:2511.16618v2 Announce Type: replace Abstract: Surgical scene understanding demands temporally consistent tracking of instruments and tissues. For clinical use, such tracking should generalize to

TIGA: Trajectory-Injected Generative Attack against Black-box AIGC Detectors

ResearchDGX agent

arXiv:2607.25894v1 Announce Type: new Abstract: Recent diffusion models have achieved remarkable realism in facial image synthesis, posing growing challenges to artificial intelligence-generated conte

Towards Faithful Sentimental Image Captioning via Evidence-Aware Multi-Agent Reasoning

Model ReleasesDGX agent

arXiv:2607.25789v1 Announce Type: new Abstract: Sentimental Image Captioning (SIC) requires balancing emotional expression with visual fidelity. Existing methods often struggle with this trade-off, le

Towards Reliable Stain Transfer: An Iterative Data-Model Co-Optimization Framework Based on Multimodal Expert-Guided Assessment

ResearchDGX agent

arXiv:2607.25393v1 Announce Type: new Abstract: Histopathological examination primarily relies on hematoxylin and eosin (H&E) and immunohistochemistry (IHC) staining. Although IHC provides critical mo

Track-Leakage-Free Hold-Out Self-Validation for Photogrammetric Reconstruction: Protocol, Sensitivity, and Limits

Model ReleasesDGX agent

arXiv:2607.24852v1 Announce Type: new Abstract: Automated photogrammetric inspection emits metric measurements from a 3D reconstruction whose own correctness is normally unknown without an external su

Unifying Active Learning and Semi-Supervised Learning for Medical Image Segmentation

TutorialsDGX agent

arXiv:2607.25014v1 Announce Type: new Abstract: In practical settings, medical image segmentation models are often developed with limited annotated data rather than fully labeled datasets. Training fr

Universal Pansharpening Model

Model ReleasesDGX agent

arXiv:2603.03831v2 Announce Type: replace Abstract: Pansharpening generates the high-resolution multi-spectral (MS) image by integrating spatial details from a texture-rich panchromatic (PAN) image an

VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening

SafetyDGX agent

arXiv:2607.26042v1 Announce Type: new Abstract: We present VetClaw, an edge-cloud multimodal agentic system for early veterinary disease screening. VetClaw uses a camera module as an edge sensing devi

WHTMix: Efficient Stereo Depth Estimation via Walsh-Hadamard Token Mixing

ResearchDGX agent

arXiv:2607.25234v1 Announce Type: new Abstract: Stereo depth estimation for driving, robotics and augmented reality must run at high resolution under tight latency budgets, yet in transformer-based ma

Wonder: Video World Model Done Better

ResearchDGX agent

arXiv:2607.26037v1 Announce Type: new Abstract: We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wond

28 Jul 2026

A Controlled Visual-Backbone Benchmark for Multimodal Short-Term Solar Irradiance Forecasting

Model ReleasesDGX agent

arXiv:2607.23633v1 Announce Type: cross Abstract: Sky-image irradiance studies often compare forecasting systems in which the image encoder, temporal model, fusion block, target definition, and traini

A Diagnostic Gap Framework for Evaluating Reconstruction Fidelity in Weakly Supervised Mammography

ResearchDGX agent

arXiv:2607.22740v1 Announce Type: new Abstract: Weakly supervised pipelines for medical imaging have become increasingly popular over the years. These systems often include multiple stages and compone

A Modern ConvNet for Solar Filament Detection

ResearchDGX agent

arXiv:2607.24525v1 Announce Type: cross Abstract: Automated solar filament detection using deep learning faces several challenges. Semantic segmentation of solar filaments is a complicated multiscale

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions

ResearchDGX agent

arXiv:2607.23235v1 Announce Type: new Abstract: Image captioning is a primary task in vision--language research, yet assessing how faithfully a caption preserves image semantics without relying on ref

A Reference-Free Framework for Evaluating Single-Frame ISP Pipelines

ResearchDGX agent

arXiv:2607.23321v1 Announce Type: cross Abstract: Evaluating camera image signal processing (ISP) pipelines requires measuring low-level artifacts introduced by operations such as denoising, demosaici

A Scale-adaptive Vision Model Links C. elegans Neuronal Morphology to Behavior for Neurotoxicity Assessment

Model ReleasesDGX agent

arXiv:2607.23183v1 Announce Type: cross Abstract: Neurological disorders are a leading cause of global disability and are increasingly linked to environmental chemical exposures. Yet neurotoxicity ass

A Unified Stereo Geometry Estimation Framework for Disparity and Surface Normal

ResearchDGX agent

arXiv:2607.24024v1 Announce Type: new Abstract: Stereo matching and surface normal estimation are fundamental tasks in 3D vision. However, existing feed-forward stereo methods still struggle to produc

Accuracy potential of visual localization exploiting high-end street-level imagery

AgentsDGX agent

arXiv:2607.24409v1 Announce Type: new Abstract: Accurate and reliable pose information with respect to a reference frame is increasingly demanded across applications such as autonomous navigation, sur

Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models

ResearchDGX agent

arXiv:2603.05147v2 Announce Type: replace Abstract: Current research on Vision-Language-Action (VLA) models predominantly focuses on enhancing generalization through reasoning techniques. While effect

AdaKAN: A dual-branch adaptive Kolmogorov-Arnold network for medical image segmentation

ResearchDGX agent

arXiv:2607.22891v1 Announce Type: new Abstract: Medical image segmentation is a fundamental task in computer-aided diagnosis, yet it remains challenging due to the complexity of anatomical structures

AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognition

Model ReleasesDGX agent

arXiv:2607.02271v2 Announce Type: replace Abstract: Vein recognition is a secure biometric technology often constrained by limited annotated data and imaging variations. While data augmentation mitiga

Ambient pressure compensation and robust position control of oil-filled electric joint systems for underwater manipulators

ResearchDGX agent

arXiv:2607.24384v1 Announce Type: new Abstract: Electric joint systems are significant elements of an underwater manipulator for its actuation, drive, and control. Working in an underwater environment

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars

Model ReleasesDGX agent

arXiv:2607.24013v1 Announce Type: new Abstract: Production-ready audio-driven avatar generation requires efficient inference without sacrificing fidelity or motion expressiveness. However, existing ac

ATCNet-CIAM for Multi-Session Motor Imagery EEG Signal Classification

ResearchDGX agent

arXiv:2607.23522v1 Announce Type: new Abstract: Motor imagery (MI)-based electroencephalography is widely used in non-invasive brain--computer interfaces (BCIs), but robust decoding remains challengin

BATON: A Multimodal Benchmark for Bidirectional Automation Transition Observation in Naturalistic Driving

Model ReleasesDGX agent

arXiv:2604.07263v2 Announce Type: replace-cross Abstract: Existing driving automation (DA) systems on production vehicles rely on human drivers to decide when to engage DA while requiring them to rema

Benchmarking the Domain Gap: Model Selection Instability Under Domain Shift in Video Capsule Endoscopy

ResearchDGX agent

arXiv:2607.22736v1 Announce Type: new Abstract: Video capsule endoscopy (VCE) classification is typically evaluated within a single dataset, yet clinical deployment demands robustness across acquisiti

Beyond Appearance: A Multi-cue Framework and Large-scale Benchmark for Pedestrian Association and Tracking on Mobile Aerial-Ground Platforms

Model ReleasesDGX agent

arXiv:2607.23803v1 Announce Type: new Abstract: Multi-view Multi-object Association and Tracking (MvMoAT) associates objects across camera views and tracks them over time, supporting identity persiste

Beyond Error-vs-Discard Characteristic: Toward Stable and Reliable Evaluation for Face Image Quality Assessment

ResearchDGX agent

arXiv:2607.22752v1 Announce Type: new Abstract: Face Image Quality Assessment (FIQA) aims to estimate the utility of facial images for reliable recognition. The evaluation of FIQA methods is predomina

BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion

TutorialsDGX agent

arXiv:2607.24110v1 Announce Type: new Abstract: Mobile infrared-visible imaging typically pairs a compact infrared sensor with a high-resolution visible camera for complementary perception. While cros

Breaking the Synthetic-Real Domain Shortcut for Training-Free Generative Replay-based Class Incremental Learning

SafetyDGX agent

arXiv:2607.22994v1 Announce Type: new Abstract: Class-incremental learning (CIL) requires models to continuously acquire new knowledge while avoiding catastrophic forgetting. While exemplar replay is

Calibration-Free 3D Multi-Camera People Tracking for Indoor Environment

ResearchDGX agent

arXiv:2607.22731v1 Announce Type: new Abstract: Multi-Camera People Tracking (MCPT) traditionally relies on precise intrinsic and extrinsic camera calibration to project 2D detections into a unified 3

CameraAnything: Refilming Videos with Arbitrary Camera Control

Model ReleasesDGX agent

arXiv:2607.24591v1 Announce Type: new Abstract: We introduce CameraAnything, the first unified framework for camera controlled video editing that enables joint control of both intrinsic and extrinsic

Cascade Forgery Mining Network for Fingerprint Presentation Attack Detection

ResearchDGX agent

arXiv:2607.24090v1 Announce Type: new Abstract: Fingerprint Presentation Attack Detection (PAD) is a critical component of fingerprint identification systems, serving as a protective measure against u

Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework

Model ReleasesDGX agent

arXiv:2607.22715v1 Announce Type: new Abstract: The rapid growth of Artificial Intelligence-generated content (AIGC) is reshaping video production and circulation, exposing children to an increasing v

Codebook Capacity Governs Perceptual Quality Across Resolutions in Hierarchical Discrete Video Compression

ResearchDGX agent

arXiv:2607.23366v1 Announce Type: cross Abstract: Learned video codecs based on continuous latent representations typically require resolution-specific retraining or rate-distortion (RD) recalibration

Color Fundus Photography Analysis: Co-evolution of Data, Preprocessing, and Modeling toward Multimodal AI

ResearchDGX agent

arXiv:2607.23972v1 Announce Type: new Abstract: Color Fundus Photography (CFP) is a primary non-invasive imaging modality for large-scale screening of ophthalmic and systemic diseases. Existing survey

ConFusion: Continuous Fusion Space Learning for Fine-Grained Controllable Infrared and Visible Image Fusion

SafetyDGX agent

arXiv:2607.23600v1 Announce Type: new Abstract: Controllable infrared-visible image fusion aims to integrate complementary thermal and structural information with flexible region-aware modulation, pro

Consistent Evidence, Robust Recognition: Faithful Attribution Regularization under Geometric Transformations

Model ReleasesDGX agent

arXiv:2607.23835v1 Announce Type: new Abstract: Attribution methods are widely used to characterize the evidence underlying model predictions, yet their potential to improve model behavior remains und

Contrastive Parameter Disentanglement for Multi-modal Remote Sensing Image Generation

Model ReleasesDGX agent

arXiv:2607.23673v1 Announce Type: new Abstract: Existing remote sensing image generation methods are largely confined to single-modality synthesis and therefore fail to exploit the complementary infor

Controllable Diversity in Normalization-Based Implicit Ensembles via Softmax-Temperature Modulation

Model ReleasesDGX agent

arXiv:2607.23860v1 Announce Type: cross Abstract: Deep ensembles provide the most reliable uncertainty estimates in deep learning, but their cost grows linearly with the number of members. Implicit en

Counterfactual Motion Reliability Learning for Robust UAV Tracking

TutorialsDGX agent

arXiv:2607.23209v1 Announce Type: new Abstract: Infrared unmanned aerial vehicle (UAV) tracking is challenging because the target is often small, low-contrast, and easily confused with thermal distrac

CrossSpine: Multi-scale Cross-sequence Attention with Anatomical Priors for Automated Pfirrmann Grading

TutorialsDGX agent

arXiv:2607.22728v1 Announce Type: new Abstract: Automated grading of Lumbar Disc Degeneration is essential for the objective quantification of structural changes associated with low back pain. Observi

DailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative Models

Model ReleasesDGX agent

arXiv:2607.24016v1 Announce Type: new Abstract: Recent advances in generative models have shifted AI-generated image detection from identifying easily distinguishable, fully synthetic images to identi

DAP-Pose: Deep Temporal Alignment and Physics-aware Cross-modal Sensor Fusion for Robust Pose Estimation

Model ReleasesDGX agent

arXiv:2607.23755v1 Announce Type: new Abstract: Robust and accurate pose estimation with multi-modal sensors is fundamental for autonomous vehicles and mobile robotic systems in complex environments.

← Previous
1…2829303132…209
Next →