AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
23 Jun 2026

One-Shot Data Selection for Medical Image Classification via Graph Coverage

Local AiDGX agent

arXiv:2606.22002v1 Announce Type: new Abstract: Training medical image classifiers on entire datasets is wasteful when annotation budgets are limited: not all samples contribute equally, yet acquiring

Open Annotations and Synthetic Data for Field Localisation in Indian Bank Cheques

Model ReleasesDGX agent

arXiv:2606.20682v1 Announce Type: new Abstract: Automated cheque processing requires localising key fields (date, legal amount, IFSC code, account number, signature, and payee name) before any recogni

OphthaDT: Generative Digital Twins for Forecasting Visual Acuity Trajectories in Ophthalmology

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.22101v1 Announce Type: cross Abstract: Precision medicine in ophthalmology requires accurate longitudinal predictions, but the fragmented nature of multimodal clinical data remains a barrie

Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback

Model ReleasesDGX agent

arXiv:2510.02561v2 Announce Type: replace Abstract: Recent advances in large video-language models (VLMs) rely on extensive fine-tuning techniques that strengthen alignment between textual and visual

ORBIT: Training-Free Multi-Attribute Behavioral Steering via Orthogonal Subspace Rotation

Model ReleasesDGX agent

arXiv:2606.22357v1 Announce Type: cross Abstract: Language models are widely used in assistant settings, where controlling behavioral attributes is often essential. Activation steering modifies hidden

OrthoMotion:Disentangling Camera and Subject Motion via Geometry Semantics Orthogonal Attention

ResearchDGX agent

arXiv:2606.22835v1 Announce Type: new Abstract: Controllable video generation demands independent command of the camera and the subject, yet 2D conditioning entangles them: camera- and object-induced

OSOG: A Differentiable, Physics-Informed Synthetic Data Engine for Micro-Optical Environments

ApplicationsDGX agent

arXiv:2606.21381v1 Announce Type: new Abstract: Deep learning in computational microscopy is severely constrained by the scarcity of densely annotated datasets. While synthetic data generation has bri

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture

Local AiDGX agent

arXiv:2606.23256v1 Announce Type: new Abstract: The increasing maturity of embodied AI platforms has driven a growing interest in procedural video representation learning to support intelligent assist

PaaF: Raising the perceived quality of INR-Based Image Compression

ResearchDGX agent

arXiv:2606.21655v1 Announce Type: cross Abstract: Implicit Neural Representations (INRs) have recently emerged as a promising paradigm for image compression, offering a fundamentally different approac

PG-MAP: Joint MAP Optimization for Inference-Time Alignment of Diffusion and Flow-Matching Models

SafetyDGX agent

arXiv:2606.22958v1 Announce Type: cross Abstract: Inference-time alignment of pretrained text-to-image models is typically performed along a single control axis, such as classifier-free guidance, atte

PHAST-Net: Attention-Guided, Physics-Informed Network for Unified Estimation of Ideal Time-Frequency Representations

ResearchDGX agent

arXiv:2606.23665v1 Announce Type: cross Abstract: We introduce PHAST-Net, an attention-guided, physics-informed network for unified estimation of Ideal Time-Frequency Representations (ITFRs), spanning

phi-Scene: Physically Grounded Image-to-3D Scene Reconstruction

SafetyDGX agent

arXiv:2606.21596v1 Announce Type: new Abstract: Reconstructing compositional 3D scenes from a single image is a fundamental challenge in 3D world modeling. Recent methods can recover high-fidelity, co

PHOEBI: An Open-World Benchmark for Bacterial Identification in Phase-Contrast Microscopy

Model ReleasesDGX agent

arXiv:2606.22890v1 Announce Type: new Abstract: Optical microscopy enables rapid, label-free imaging of live bacteria and is the standard instrument for species identification across clinical, environ

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation

TutorialsDGX agent

arXiv:2512.24551v4 Announce Type: replace Abstract: Recent advances in text-to-video (T2V) generation have achieved good visual quality, yet synthesizing videos that faithfully follow physical laws re

PhysFlow: Frequency Decoupled with Dual-Field Rectified Flow for Remote Photoplethysmography

Model ReleasesDGX agent

arXiv:2606.23226v1 Announce Type: new Abstract: Remote Photoplethysmography (rPPG) enables contactless pulse estimation from facial videos, serving as a vital tool for health monitoring. However, curr

Physically-guided Image Generation for Multi-Projection Mapping

SafetyDGX agent

arXiv:2606.22477v1 Announce Type: new Abstract: Projection Mapping (PM) enables seamless superimposition of digital content onto real-world 3D objects, serving as a fundamental technique for immersive

Physics-Guided Spatiotemporal State Space Modeling for Lookahead Molten Pool Segmentation in Laser Wire-Feed Welding

ResearchDGX agent

arXiv:2606.23028v1 Announce Type: new Abstract: Real-time weld-pool perception is critical for closed-loop control in laser wire-feed welding, where sensing, computation, and actuator response introdu

PIAvatar: Physically Interactive Avatars via Deformation Gradient Decoupling

ResearchDGX agent

arXiv:2606.21162v1 Announce Type: cross Abstract: 3D human avatars have shown impressive visual fidelity driven by pose-conditioned models, yet they still lack the physical ability required for intera

PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards

SafetyDGX agent

arXiv:2602.01624v2 Announce Type: replace Abstract: Text-to-video (T2V) generation aims to synthesize videos with high visual quality and temporal consistency that are semantically aligned with input

Poisson2Gaussian: Noise Gaussianization to Enhance Image Denoising

TutorialsDGX agent

arXiv:2606.23098v1 Announce Type: new Abstract: The quantum nature of light determines the inherent Poisson stochasticity of photon detection, which is ubiquitous in photography, microscopy, and astro

Policy-as-Data: Learning Generalizable HOI Diffusion Models from Simulated Physics

SafetyDGX agent

arXiv:2606.22806v1 Announce Type: new Abstract: Synthesizing realistic Human-Object Interactions (HOI) is critical for creating embodied avatars and functional virtual environments. However, current d

PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models

SafetyDGX agent

arXiv:2606.22540v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models provide a unified paradigm for robotic manipulation, yet their real-world deployment is often bottlenecked by execut

Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking

Model ReleasesDGX agent

arXiv:2606.23604v1 Announce Type: new Abstract: The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion estimation. How

Polynomial Dice Loss for Medical Image Segmentation

ResearchDGX agent

arXiv:2606.23373v1 Announce Type: new Abstract: Medical image segmentation is a fundamental task for medical image processing and computer-assisted intervention, yet data imbalance and small lesion de

Pose Anything Anywhere:Model-free Object Poses from Arbitrary References

SafetyDGX agent

arXiv:2606.23634v1 Announce Type: new Abstract: Estimating the 6D pose of unseen objects is a fundamental yet challenging problem for open-world robotics and embodied perception. Model-based methods a

Precision Recall Controllable Radiology Report Generation via Hybrid Natural Language and Clinical Reward Learning

Model ReleasesDGX agent

arXiv:2606.21447v1 Announce Type: cross Abstract: Automated radiology report generation (RRG) has gained increasing attention because it can reduce the heavy workload of clinical report writing. Howev

Predicting Immune Biomarkers with MultiModal Mixture-of-Expert Pathology Foundation Models Empowers Precision Oncology

ResearchDGX agent

arXiv:2606.18123v2 Announce Type: replace Abstract: Predicting immune biomarkers associated with the tumor immune microenvironment (TIME) is critical for advancing precision oncology, yet existing app

Privacy-Preserving Person Re-Identification from Temporal Sequences with Transformer and Hungarian Optimization

ResearchDGX agent

arXiv:2606.23230v1 Announce Type: new Abstract: Person re-identification (Re-ID) is a crucial task in surveillance and human behavior analysis, often used in public spaces such as transport hubs. Trad

Projection-Volume Fidelity Divergence: Diagnosing and Controlling Optimization Drift in Sparse-View 3D Gaussian Tomography

ResearchDGX agent

arXiv:2606.22525v1 Announce Type: new Abstract: Sparse-view computed tomography is a severely ill-posed inverse problem, where recent 3D Gaussian Splatting methods offer an efficient explicit represen

Prompt-Calibrated SAM 3 for Open-Vocabulary Remote Sensing Semantic Segmentation

ResearchDGX agent

arXiv:2606.21863v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation (OVSS) in remote sensing images aims to segment categories beyond a fixed label space. Recent SAM 3-based methods

Prompting Diffusion Models for Zero-Shot Instance Segmentation

ResearchDGX agent

arXiv:2606.22660v1 Announce Type: new Abstract: Several disruptive research directions have recently emerged in computer vision, including foundation models achieving previously unseen zero-shot perfo

PROTON: Prototype-Based Test-Time Online OOD Detection for Medical VLMs

Model ReleasesDGX agent

arXiv:2606.20913v1 Announce Type: new Abstract: Medical vision-language models (VLMs) enable zero-shot clinical image classification, yet reliably detecting out-of-distribution (OOD) inputs at deploym

Quantile Adaptive Temperature Scaling for Confidence Calibration

ResearchDGX agent

arXiv:2606.21749v1 Announce Type: new Abstract: Deep neural networks often produce poorly calibrated confidence estimates, overstating their certainty even when predictions are incorrect. Temperature

Quantum Visual Fields with Neural Amplitude Encoding

ResearchDGX agent

arXiv:2508.10900v2 Announce Type: replace Abstract: Quantum Implicit Neural Representations (QINRs) have emerged as a promising paradigm that leverages parametrised quantum circuits to encode and proc

Radial Basis Function Networks as Projection Heads in Self-Supervised Learning

ResearchDGX agent

arXiv:2606.21590v1 Announce Type: new Abstract: Self-supervised learning (SSL) typically relies on a backbone encoder followed by a small multilayer perceptron (MLP) projection head, which is conventi

RAPID: A Reproducible Multi-Agent Pipeline for Interpretable Disaster Damage Assessment from Satellite and Street-View Imagery

AgentsDGX agent

arXiv:2606.21819v1 Announce Type: new Abstract: Due to the increasing frequency and intensity of extreme climate events, there is a clear demand for intelligent, scalable, and autonomous approaches to

Rapid Quantification of Outdoor Object Visibility in Urban Setting Using Connected-Vehicle Fields of View

ResearchDGX agent

arXiv:2506.03365v3 Announce Type: replace-cross Abstract: Identifying locations that offer maximum visual exposure to passing vehicular traffic is a core problem in urban analytics, with applications

RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation

ResearchDGX agent

arXiv:2606.22749v1 Announce Type: new Abstract: Pre-trained Vision Foundation Models (VFMs) have become central to modern computer vision due to their powerful semantic representations and strong gene

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations

Model ReleasesDGX agent

arXiv:2606.22766v1 Announce Type: new Abstract: Audio Description aims to generate concise narrations of essential visual content in audio-visual media for blind and low-vision audiences. Existing met

Real-Time Multimodal Activity-Aware Error Detection in Robot-Assisted Surgery

SafetyDGX agent

arXiv:2606.23593v1 Announce Type: cross Abstract: Robot-assisted minimally invasive surgery improves surgical precision but introduces complexity, making technical error detection essential for ensuri

Real-time pedestrian attribute recognition with YOLOv8 and ResNet18

HardwareDGX agent

arXiv:2606.21200v1 Announce Type: new Abstract: Pedestrian attribute recognition (PAR) assigns semantic labels to detected pedestrians and is useful in surveillance, video retrieval, and human-centere

Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild

Model ReleasesDGX agent

arXiv:2603.04205v2 Announce Type: replace Abstract: While Vision-Language Models (VLMs) achieve near-perfect scores on digital document benchmarks like OmniDocBench, their performance in the unpredict

ReconMIL: Synergizing Latent Space Reconstruction with Bi-Stream Mamba for Whole Slide Image Analysis

ResearchDGX agent

arXiv:2603.19925v2 Announce Type: replace-cross Abstract: Whole slide image (WSI) analysis heavily relies on multiple instance learning (MIL). While recent methods benefit from large-scale foundation

Region-Specific Calibration Achieves Excellent Inter-Device Reliability for Smartphone Dermatology: A Multi-Device Benchmark on Korean Facial Skin

Model ReleasesDGX agent

arXiv:2512.21988v3 Announce Type: replace-cross Abstract: Background: Smartphone-based dermatology requires inter-device colorimetric reliability that holds across calibration regimes, yet quantitativ

REKEY: Metadata-Grounded Visual-Key Regeneration for Contamination-Resilient VQA Evaluation

Model ReleasesDGX agent

arXiv:2606.20736v1 Announce Type: new Abstract: Static visual question answering (VQA) benchmarks age quickly: Once the items leak into training corpora, scores can reflect memorization rather than ge

Reliability-Guided Adaptive Ensembling for Robust Test-Time Adaptation

ResearchDGX agent

arXiv:2606.22351v1 Announce Type: cross Abstract: Test-time adaptation (TTA) can mitigate domain shift without source data, but it is highly brittle under adversarially contaminated test streams, wher

RelightAnyone: A Generalized Relightable 3D Gaussian Head Model

SafetyDGX agent

arXiv:2601.03357v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has become a standard approach to reconstruct and render photorealistic 3D head avatars. A major challenge is to religh

Render-FM: Feedforward Model for Real-time Photorealistic Volumetric Rendering

ResearchDGX agent

arXiv:2505.17338v2 Announce Type: replace Abstract: Photorealistic volumetric rendering of CT scans greatly benefits clinical workflows, yet neural approaches such as Neural Radiance Fields (NeRF) and

Resolving Multi-Target Association in OFDM-based ISAC via Vision-aided Multi-Modal Learning

ResearchDGX agent

arXiv:2606.22195v1 Announce Type: new Abstract: Orthogonal frequency division multiplexing (OFDM)-based integrated sensing and communication (ISAC) systems commonly extract target parameters by peak-s

Rethinking Object-Centric Representations for Video Dynamics Modeling

SafetyDGX agent

arXiv:2606.23436v1 Announce Type: new Abstract: Unsupervised video object tracking aims to decompose dynamic scenes into persistent, object-centric entities without manual annotations. Many recent app

Rethinking Prototype-based Similarity Learning for Few-Shot Object Detection

TutorialsDGX agent

arXiv:2606.23069v1 Announce Type: new Abstract: Few-shot object detection aims to detect novel object categories from only a few labeled examples, avoiding costly large-scale annotation. Recent protot

Rethinking the Adaptation of Vision Foundation Models for Efficient Cell Segmentation

ResearchDGX agent

arXiv:2606.21913v1 Announce Type: new Abstract: Cell segmentation is critical for computational pathology and biomedical discovery. While recent Vision Foundation Models (VFMs) have demonstrated remar

Retrieval-Augmented Anatomical Guidance for Text-to-CT Generation

ResearchDGX agent

arXiv:2603.08305v2 Announce Type: replace Abstract: Text-conditioned generative models for volumetric medical imaging provide semantic control but lack explicit anatomical guidance, often resulting in

Robot Self-Improvement via Human-Video Dynamics Models

SafetyDGX agent

arXiv:2606.21406v1 Announce Type: cross Abstract: A central question in robot learning is how to acquire skills from the kinds of data that humans learn from: passive observation, embodied practice, a

Robust 3DGS-based SLAM via Adaptive Kernel Smoothing

Model ReleasesDGX agent

arXiv:2511.23221v2 Announce Type: replace Abstract: In this paper, we challenge the conventional notion in 3DGS-SLAM that rendering quality is the primary determinant of tracking accuracy. We argue th

Robust Image-Driven Phenotyping of Ovarian Tumor Cells using Optimized Dynamic Features in Hyperbolic Channels

ResearchDGX agent

arXiv:2606.20703v1 Announce Type: new Abstract: Label-free, image-based cellular mechanophenotyping in microfluidic devices provides a high-throughput method for single-cell profiling. However, while

Robust Representation Learning in Masked Autoencoders

SafetyDGX agent

arXiv:2602.03531v2 Announce Type: replace-cross Abstract: Masked Autoencoders (MAEs) achieve impressive performance in image classification tasks, yet the internal representations they learn remain le

Robust Zero-Shot Generalization for Open-Vocabulary Action Recognition via Task Arithmetic

ApplicationsDGX agent

arXiv:2606.20734v1 Announce Type: new Abstract: Open Vocabulary Action Recognition (OVAR) enables the recognition of novel actions by leveraging vision-language representations, overcoming the limitat

Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City

AgentsDGX agent

arXiv:2606.20980v1 Announce Type: new Abstract: As Self-Driving Cars continue to expand internationally and use multi-modal systems such as VLMs as a cognitive backbone for their Action models; how we

Rotation-Aware Point-Cloud Embeddings for Vision-Based In-Hand Reorientation

SafetyDGX agent

arXiv:2606.21788v1 Announce Type: cross Abstract: Point-cloud goals provide a direct way to specify dexterous in-hand reorientation: instead of defining an object-specific pose frame or estimating 6D

← Previous
1…8081828384…211
Next →