AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Image class translation: visual inspection of class-specific hypotheticals and classification based on translation distance

DGX agent

arXiv:2408.08973v3 Announce Type: replace Abstract: Purpose: A major barrier to the implementation of artificial intelligence for medical applications is automated CNNs' lack of explainability and hig

researcharxiv-cs-cv
8 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

Implicit Data Synthesis for Contrastive Unsupervised Data Augmentation

DGX agent

arXiv:2606.07498v1 Announce Type: new Abstract: Scientific observations generate large quantities of unlabeled data which is laborious to hand-label, making unsupervised learning techniques valuable f

researcharxiv-cs-cv
8 Jun 2026
Research

JA-SIREN: Deterministic Initialization for Sinusoidal Networks via Spectral Matching

DGX agent

arXiv:2606.06671v1 Announce Type: new Abstract: Existing implicit neural representation (INR) approaches suffer from stochastic initialization that does not guarantee consistent or high-quality perfor

researcharxiv-cs-cv
8 Jun 2026
Safety

LARA: Latent Action Representation Alignment for Vision-Language-Action Models

DGX agent

arXiv:2606.07100v1 Announce Type: new Abstract: Visual-language action (VLA) models enable robots to predict actions directly from observations and language instructions, but their performance depends

safetyarxiv-cs-cv
8 Jun 2026
Tutorials

Lighting-Aware Representation Learning under Controllable Lighting Variation

DGX agent

arXiv:2606.06899v1 Announce Type: new Abstract: Variations in illumination remain a major challenge for visual representation learning, as they induce substantial appearance changes both across and wi

tutorialsarxiv-cs-cv
8 Jun 2026
Applications

LRMIL: Efficient Low-Resolution Multiple Instance Learning via High-Resolution Knowledge Distillation for Whole Slide Image Classification

DGX agent

arXiv:2606.06864v1 Announce Type: new Abstract: Multiple instance learning (MIL) has become a standard paradigm for whole slide image (WSI) analysis in digital pathology, as it enables slide-level pre

applicationsarxiv-cs-cv
8 Jun 2026
Applications

LUCID: Learning Unified Control for Image Deflaring and Exposure Mastery in Nighttime Photography

DGX agent

arXiv:2606.06901v1 Announce Type: new Abstract: Photography is the art of painting with light, yet nighttime scenes are shaped by competing degradations: intense flares obscure scene structure, while

applicationsarxiv-cs-cv
8 Jun 2026
Research

MedSIGHT: Towards Grounded Visual Comprehension in Medical Large Vision-Language Models

DGX agent

arXiv:2606.06760v1 Announce Type: new Abstract: Medical large vision-language models (Med-LVLMs) have recently achieved remarkable progress in vision-language comprehension and medical image segmentat

researcharxiv-cs-cv
8 Jun 2026
Research

Mind the Gap: Disentangling Performance Bottlenecks in Video Instance Segmentation

DGX agent

arXiv:2606.07394v1 Announce Type: new Abstract: In Video Instance Segmentation (VIS), classification, segmentation, and tracking objectives are jointly evaluated, but their individual contributions to

researcharxiv-cs-cv
8 Jun 2026
Applications

Multi-FRuGaL: Multimodal Flexible Redundancy-aware Decomposed Gated Learning for Cancer Diagnosis and Prognosis

DGX agent

arXiv:2606.06867v1 Announce Type: new Abstract: Modern medicine relies on heterogeneous data sources spanning radiology, pathology, text reports, and structured clinical information. However, real-wor

applicationsarxiv-cs-cv
8 Jun 2026
Hardware

MVSegNet: A Lightweight Boundary-Aware Network for Fetal Lateral Ventricle Segmentation and Atrial Width Estimation in Prenatal Ultrasound

DGX agent

arXiv:2606.06958v1 Announce Type: new Abstract: Fetal ventriculomegaly is assessed by measuring the atrial width of the lateral ventricle in prenatal ultrasound. Accurate segmentation is essential for

hardwarearxiv-cs-cv
8 Jun 2026
Model Releases

OpenGlass: Open-Source Smart Glasses for On-Device Event-Based Gesture Recognition

DGX agent

arXiv:2606.07431v1 Announce Type: new Abstract: Smart eyewear enables unobtrusive, context-aware interaction through multimodal sensors and on-device intelligence, but is severely limited by power, me

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

OPTIMUS-Prime: Minimal and Sufficient Concept Explanations for Deep Vision Models

DGX agent

arXiv:2606.07180v1 Announce Type: new Abstract: The growing demand for transparency in automated decision-making has propelled eXplainable Artificial Intelligence (XAI) to the forefront of machine lea

model-releasesarxiv-cs-cv
8 Jun 2026
Research

PARSE: Part-Aware Relational Spatial Modeling

DGX agent

arXiv:2603.07704v2 Announce Type: replace Abstract: Inter-object relations underpin spatial intelligence, yet existing representations -- linguistic prepositions or object-level scene graphs -- are to

researcharxiv-cs-cv
8 Jun 2026
Model Releases

Physics-Driven Semantic Scattering Structure Understanding of Aircraft Target in SAR Images

DGX agent

arXiv:2606.06847v1 Announce Type: cross Abstract: Synthetic aperture radar (SAR) has become indispensable for target interpretation owing to its all-day and all-weather observation capability. In SAR

model-releasesarxiv-cs-cv
8 Jun 2026
Research

polyDAG: Polynomial Acyclicity Constraints for Efficient Continuous Causal Discovery in Visual Semantic Graphs

DGX agent

arXiv:2606.06908v1 Announce Type: new Abstract: Modern image-analysis pipelines often convert images into structured semantic variables, such as facial attributes, object concepts, and scene descripto

researcharxiv-cs-cv
8 Jun 2026
Research

Real-Time AttentionBender: Granular Interactive Network Bending of Video Diffusion Transformers

DGX agent

arXiv:2606.06497v1 Announce Type: cross Abstract: Generative video models have achieved remarkable visual fidelity, yet their prompt-only interface offers thin creative agency and obscures the model's

researcharxiv-cs-cv
8 Jun 2026
Model Releases

RealDocBench: A Benchmark for Field-Level QA and Layout Understanding on Real-World Regulated Documents

DGX agent

arXiv:2606.07401v1 Announce Type: new Abstract: Document parsing systems are increasingly deployed in high-stakes, regulated workflows such as mortgage underwriting, financial reporting, supply-chain

model-releasesarxiv-cs-cv
8 Jun 2026
Research

Reconstructing Multi-Decadal Forest Disturbances: A Spatio-Temporal Transformer Approach

DGX agent

arXiv:2606.07249v1 Announce Type: new Abstract: Accurate monitoring of forest disturbances is essential for understanding carbon dynamics and land management, yet traditional approaches typically rely

researcharxiv-cs-cv
8 Jun 2026
Research

RigPAPR: Rig-Based Animation of Static Neural Point Clouds from a Fixed-Viewpoint Video

DGX agent

arXiv:2606.06685v1 Announce Type: new Abstract: Static neural point reconstructions capture a subject at high fidelity from posed images. Given such a reconstruction, we aim to animate it to follow a

researcharxiv-cs-cv
8 Jun 2026
Model Releases

RISE: Single Static Radar-based Indoor Scene Understanding

DGX agent

arXiv:2511.14019v3 Announce Type: replace Abstract: Robust and privacy-preserving indoor scene understanding remains a fundamental open problem. While optical sensors such as RGB and LiDAR offer high

model-releasesarxiv-cs-cv
8 Jun 2026
Safety

Robotic Policy Adaptation via Weight-Space Meta-Learning

DGX agent

arXiv:2606.07217v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are emerging as a promising paradigm for robotic manipulation, enabling general-purpose policies trained from larg

safetyarxiv-cs-cv
8 Jun 2026
Model Releases

RPC-GS: Gaussian Splatting with native RPC Rendering for Satellite Imagery

DGX agent

arXiv:2606.06690v1 Announce Type: new Abstract: We present RPC-GS, the first Gaussian Splatting framework for satellite imagery that operates natively with Rational Polynomial Camera (RPC) models. The

model-releasesarxiv-cs-cv
8 Jun 2026
Research

S23DR 2026 Winning Solution

DGX agent

arXiv:2606.06695v1 Announce Type: new Abstract: This text presents the winning solution to the S23DR 2026 challenge for structured 3D wireframe reconstruction from sparse SfM, fitted depth, and semant

researcharxiv-cs-cv
8 Jun 2026
Model Releases

Seeing Without Exposing: Adaptive Privacy Control for Open-World, Context-Hungry MLLMs

DGX agent

arXiv:2606.07175v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have raised new privacy challenges. On the data side, user-provided inputs often include unpredictable sensitiv

model-releasesarxiv-cs-cv
8 Jun 2026
Research

SegMoTE: Token-Level Mixture of Experts for Medical Image Segmentation

DGX agent

arXiv:2602.19213v2 Announce Type: replace Abstract: Medical image segmentation is vital for clinical diagnosis and quantitative analysis, yet remains challenging due to the heterogeneity of imaging mo

researcharxiv-cs-cv
8 Jun 2026
Safety

Semantic-Structural Alignment for Generative Pictorial Charts

DGX agent

arXiv:2606.06498v1 Announce Type: cross Abstract: Traditional statistical graphics are precise but often lack the visual appeal, memorability, and engagement of pictorial charts. We present a generati

safetyarxiv-cs-cv
8 Jun 2026
Model Releases

Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning

DGX agent

arXiv:2606.07436v1 Announce Type: new Abstract: This paper explores agentic 3D spatial understanding, i.e., MLLM agents performing 3D reasoning through tool use. Existing methods often misuse tools an

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

Spatial-Temporal Decoupled Adapter for Micro-gesture Online Recognition

DGX agent

arXiv:2606.07355v1 Announce Type: new Abstract: Micro-gesture online recognition aims to temporally localize and classify subtle gestures in untrimmed videos. Owing to their extremely short duration,

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors

DGX agent

arXiv:2606.06891v1 Announce Type: new Abstract: Despite advances in 3D scene understanding, existing 3D Large Multimodal Models operate in offline settings, requiring complete scene observations or pr

model-releasesarxiv-cs-cv
8 Jun 2026
Local Ai

Streaming Video Generation with Streaming Force Control

DGX agent

arXiv:2606.07508v1 Announce Type: new Abstract: We introduce StreamForce, a streaming video generation framework that enables physically grounded control through continuous force inputs. Unlike prior

local-aiarxiv-cs-cv
8 Jun 2026
Model Releases

SVHighlights: Towards Extremely Long Sport Video Highlight Detection

DGX agent

arXiv:2606.06926v1 Announce Type: new Abstract: While highlight detection for long-form videos is of great practical importance, most existing methods remain limited to short-form content, largely due

model-releasesarxiv-cs-cv
8 Jun 2026
Research

T2LM: Long-Term 3D Human Motion Generation from Multiple Sentences

DGX agent

arXiv:2406.00636v2 Announce Type: replace Abstract: In this paper, we address the challenging problem of long-term 3D human motion generation. Specifically, we aim to generate a long sequence of smoot

researcharxiv-cs-cv
8 Jun 2026
Research

The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models

DGX agent

arXiv:2603.22278v2 Announce Type: replace Abstract: Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their prop

researcharxiv-cs-cv
8 Jun 2026
Research

TraRA: Trajectory-level Recognition Aggregation for Video Text Spotting in Urban Surveillance

DGX agent

arXiv:2606.07161v1 Announce Type: new Abstract: Video Text Spotting (VTS) is essential for urban surveillance and intelligent transportation systems, enabling automated reading of street signs, vehicl

researcharxiv-cs-cv
8 Jun 2026
Safety

TrioPose: Native Triple-Stream Diffusion Transformers for Pose-Guided Text-to-Image Generation

DGX agent

arXiv:2606.07053v1 Announce Type: new Abstract: Pose-guided text-to-image generation often suffers from limb distortions and feature crosstalk in complex multi-person scenarios. While existing UNet-ba

safetyarxiv-cs-cv
8 Jun 2026
Model Releases

Twin: Tuning Learning Rate and Weight Decay of Deep Homogeneous Classifiers without Validation

DGX agent

arXiv:2403.05532v2 Announce Type: replace-cross Abstract: We introduce Tune without Validation (Twin), a simple and effective pipeline for tuning learning rate and weight decay of homogeneous classifi

model-releasesarxiv-cs-cv
8 Jun 2026
Safety

Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information Flows

DGX agent

arXiv:2606.06875v1 Announce Type: new Abstract: Diffusion transformers (DiTs) equipped with multimodal attention (MM-Attn) have become a dominant paradigm for image generation. However, preventing the

safetyarxiv-cs-cv
8 Jun 2026
Model Releases

UniSHARP: Universal Sharp Monocular View Synthesis

DGX agent

arXiv:2606.07514v1 Announce Type: new Abstract: In this work, we focus on extending SHARP, the popular photorealistic view synthesis method, for universal monocular rendering across a continuum of cam

model-releasesarxiv-cs-cv
8 Jun 2026
Research

Unmixing ATR-{mu}FTIR spectroscopic images of cross-sections of historical oil paintings

DGX agent

arXiv:2603.06673v2 Announce Type: replace Abstract: Spectroscopic imaging (SI) has become central to heritage science because it enables non-invasive, spatially resolved characterisation of materials

researcharxiv-cs-cv
8 Jun 2026
Research

Unregistered Spectral Image Fusion: Unmixing, Adversarial Learning, and Recoverability

DGX agent

arXiv:2603.21510v3 Announce Type: replace-cross Abstract: This paper addresses the fusion of a pair of spatially unregistered hyperspectral image (HSI) and multispectral image (MSI) covering roughly o

researcharxiv-cs-cv
8 Jun 2026
Research

Unsupervised Learning Based Focal Stack Camera Depth Estimation

DGX agent

arXiv:2203.07904v3 Announce Type: replace-cross Abstract: We propose an unsupervised deep learning based method to estimate depth from focal stack camera images. On the NYU-v2 dataset, our method achi

researcharxiv-cs-cv
8 Jun 2026
Applications

USU-Corn-WeedDB: A UAV RGB Image Dataset for Multi-Species Weed Detection in Forage Corn

DGX agent

arXiv:2606.06709v1 Announce Type: new Abstract: Weed pressure in forage corn production causes yield losses of up to 31.5%, yet site-specific weed management (SSWM) systems built on UAV imagery and de

applicationsarxiv-cs-cv
8 Jun 2026
Research

Varifold Moment Invariants for Sustainable and Explainable Contour Feature Extraction

DGX agent

arXiv:2606.07333v1 Announce Type: new Abstract: We introduce Varifold Moments Invariants (VMI) as a unifying framework for many previously introduced Moment Invariants. These invariants are deeply rel

researcharxiv-cs-cv
8 Jun 2026
Model Releases

VeriDrive: Verifiable Counterfactual Supervision for Cost-Efficient Vision-Language Planning

DGX agent

arXiv:2606.07338v1 Announce Type: new Abstract: Vision-language driving models increasingly use reasoning supervision to bridge perception, prediction, and planning, but existing driving rationales ar

model-releasesarxiv-cs-cv
8 Jun 2026
Model Releases

VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object Segmentation

DGX agent

arXiv:2606.06819v1 Announce Type: new Abstract: Reasoning Video Object Segmentation (RVOS) demands a sophisticated integration of temporal dynamics, spatial details, and linguistic reasoning to achiev

model-releasesarxiv-cs-cv
8 Jun 2026
Hardware

VIRTUS-FPP: Virtual Sensor Modeling for Fringe Projection Profilometry in NVIDIA Isaac Sim

DGX agent

arXiv:2509.22685v2 Announce Type: replace-cross Abstract: Fringe projection profilometry (FPP) is a high-precision structured-light sensing technique for 3D surface reconstruction, yet its practical d

hardwarearxiv-cs-cv
8 Jun 2026
Research

When CLIP Sees More, It Fights Back Harder: Multi-View Guided Adaptive Counterattacks for Test-Time Adversarial Robustness

DGX agent

arXiv:2606.06938v1 Announce Type: new Abstract: Vision-language models such as CLIP have achieved remarkable zero-shot recognition capabilities, yet their robustness against adversarial perturbations

researcharxiv-cs-cv
8 Jun 2026
← Previous
1…115116117118119…263
Next →