AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlog
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Rheos: Modelling Continuous Motion Dynamics in Hierarchical 3D Scene Graphs

DGX agent

arXiv:2603.20239v2 Announce Type: replace-cross Abstract: 3D Scene Graphs (3DSGs) provide hierarchical, multi-resolution abstractions that encode the geometric and semantic structure of an environment

researcharxiv-cs-cv
29 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

RPM-Distill: Physiology-guided Adaptive Cross-modal Distillation for Robust Remote Physiological Measurement

DGX agent

arXiv:2606.28089v1 Announce Type: new Abstract: Video-based remote physiological measurement (RPM) is highly accessible but remains fragile under varying illumination, skin tones, and motion. Radio fr

safetyarxiv-cs-cv
29 Jun 2026
Model Releases

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning

DGX agent

arXiv:2606.28266v1 Announce Type: new Abstract: Remote Sensing Image Change Captioning (RSICC) aims to describe changes between bi-temporal remote sensing images and holds significant research and app

model-releasesarxiv-cs-cv
29 Jun 2026
Safety

Scalable and Differentiable Point-Cloud Registration Using Maximum Mean Discrepancy

DGX agent

arXiv:2606.27818v1 Announce Type: new Abstract: We present MMD-Reg, a novel correspondence-free approach to point-cloud registration that is differentiable and has linear computational complexity in t

safetyarxiv-cs-cv
29 Jun 2026
Safety

ScaLe-INR: Scale and Learn Implicit Neural Representations

DGX agent

arXiv:2606.27862v1 Announce Type: new Abstract: Implicit Neural Representations (INRs) parameterized by multilayer perceptrons excel at modeling continuous signals. However, a key challenge persists a

safetyarxiv-cs-cv
29 Jun 2026
Local Ai

Scene and Human in One World: Reconstruction in a Feedforward Pass

DGX agent

arXiv:2606.27720v1 Announce Type: new Abstract: Reconstructing humans in dynamic scenes from moving monocular cameras remains challenging due to scale ambiguity, human-scene misalignment, and occlusio

local-aiarxiv-cs-cv
29 Jun 2026
Safety

Scene Generation at Absolute Scale: Utilizing Semantic and Geometric Guidance From Text for Accurate and Interpretable 3D Indoor Scene Generation

DGX agent

arXiv:2603.13910v2 Announce Type: replace Abstract: We present GuidedSceneGen, a text-to-3D generation framework that produces metrically accurate, globally consistent, and semantically interpretable

safetyarxiv-cs-cv
29 Jun 2026
Local Ai

SelectAnyTree: A Promptable Instance Segmentation Model for 3D Forest LiDAR Point Clouds

DGX agent

arXiv:2606.27491v1 Announce Type: new Abstract: Automated instance segmentation of forest LiDAR point clouds is increasingly critical as forest monitoring moves toward scalable, detailed, 3D measureme

local-aiarxiv-cs-cv
29 Jun 2026
Model Releases

SemCityLoc: Aerial 6DoF Localization Using Semantic 3D City Models

DGX agent

arXiv:2606.27444v1 Announce Type: new Abstract: Aerial 6DoF localization typically relies on precise GNSS signals or radiometrically rich 3D reconstructions, limiting scalability and on-board deployme

model-releasesarxiv-cs-cv
29 Jun 2026
Safety

SHIFT: Motion Alignment in Video Diffusion Models with Adversarial Hybrid Fine-Tuning

DGX agent

arXiv:2603.17426v2 Announce Type: replace Abstract: Image-conditioned video diffusion models achieve impressive visual realism but often suffer from weakened motion fidelity, e.g., reduced motion dyna

safetyarxiv-cs-cv
29 Jun 2026
Safety

SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models

DGX agent

arXiv:2606.27741v1 Announce Type: new Abstract: Recent advances in video diffusion models have greatly improved visual fidelity, yet their generated motions often violate physical plausibility. We obs

safetyarxiv-cs-cv
29 Jun 2026
Research

Spectral Subsurface Scattering from RGB via Biophysical Skin Inversion

DGX agent

arXiv:2606.27604v1 Announce Type: cross Abstract: In this paper we present a spectral optical inversion for skin for path tracing-based rendering of subsurface scattering. Skin is a complex multilayer

researcharxiv-cs-cv
29 Jun 2026
Research

StableMotion: One-Step Motion Estimation with Diffusion Prior

DGX agent

arXiv:2505.06668v2 Announce Type: replace Abstract: We present StableMotion, a novel framework that leverages geometric and content priors from pretrained large-scale image diffusion models for motion

researcharxiv-cs-cv
29 Jun 2026
Safety

StructSplat: Generalizable 3D Gaussian Splatting from Uncalibrated Sparse Views

DGX agent

arXiv:2606.28321v1 Announce Type: new Abstract: We present StructSplat, a feed-forward and generalizable 3D Gaussian reconstruction framework that operates directly on uncalibrated images without requ

safetyarxiv-cs-cv
29 Jun 2026
Model Releases

Structured-Li-GS: Structured 3D Gaussians Splatting with LiDAR Incorporation and Spatial Constraints

DGX agent

arXiv:2606.27509v1 Announce Type: new Abstract: In this study, we develop a Structured framework for Gaussian Splatting (3DGS) with LiDAR integration (Structured-Li-GS). It is a lightweight Gaussian S

model-releasesarxiv-cs-cv
29 Jun 2026
Model Releases

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery

DGX agent

arXiv:2505.10764v4 Announce Type: replace Abstract: Innovations in digital intelligence are transforming robotic surgery with more informed decision-making. Real-time awareness of surgical instrument

model-releasesarxiv-cs-cv
29 Jun 2026
Safety

SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation

DGX agent

arXiv:2508.06115v3 Announce Type: replace Abstract: Semantic segmentation in open-vocabulary scenarios presents significant challenges due to the wide range and granularity of semantic categories. Exi

safetyarxiv-cs-cv
29 Jun 2026
Model Releases

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

DGX agent

arXiv:2510.03117v2 Announce Type: replace Abstract: This study focuses on a challenging yet promising task, Text-to-Sounding-Video (T2SV) generation, which aims to generate a video with synchronized a

model-releasesarxiv-cs-cv
29 Jun 2026
Safety

TempAct: Advancing Temporal Plausibility in Autoregressive Video Generation via Planner-Executor RL

DGX agent

arXiv:2606.28016v1 Announce Type: new Abstract: Autoregressive (AR) video diffusion models enable low-latency streaming generation by synthesizing videos chunk by chunk with cached visual context, but

safetyarxiv-cs-cv
29 Jun 2026
Local Ai

Temporal-Emerged Prompting for Segment Anything in Multiframe Infrared Small Target Detection

DGX agent

arXiv:2606.27655v1 Announce Type: new Abstract: Accurately localizing and segmenting small targets in low signal-to-noise ratio (SNR) infrared sequences remains a challenging task. Since targets are o

local-aiarxiv-cs-cv
29 Jun 2026
Local Ai

Tessellating The Earth

DGX agent

arXiv:2606.27514v1 Announce Type: new Abstract: Geolocation encoders, which map geographic coordinates to learned representations, are emerging as an effective means of capturing visual and non-visual

local-aiarxiv-cs-cv
29 Jun 2026
Safety

Text as Illumination: Spatial Contrastive Retinex Learning for Language-guided Medical Image Segmentation

DGX agent

arXiv:2606.27794v1 Announce Type: new Abstract: Language-guided Medical Image Segmentation (LMIS) has shown great potential to improve the delineation of anatomical structures and lesions by integrati

safetyarxiv-cs-cv
29 Jun 2026
Model Releases

TextDS: Parameter-Efficient Representation Alignment for Scene Text Detection under Distribution Shifts

DGX agent

arXiv:2606.28077v1 Announce Type: new Abstract: In real-world deployments, scene text detectors inevitably face distribution shifts beyond the training distribution. Prior work often depends on large-

model-releasesarxiv-cs-cv
29 Jun 2026
Model Releases

There and Back Again: A Flexible-Frame Transformer for Multi-Exposure Fusion

DGX agent

arXiv:2606.27905v1 Announce Type: new Abstract: Multi-exposure fusion (MEF) brings the dynamic range of conventional cameras closer to that of human vision, producing images with rich scene content. G

model-releasesarxiv-cs-cv
29 Jun 2026
Research

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs

DGX agent

arXiv:2510.00705v3 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) often struggle with fine-grained perception, such as identifying small objects in high-resolution images or

researcharxiv-cs-cv
29 Jun 2026
Tutorials

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots

DGX agent

arXiv:2606.28133v1 Announce Type: cross Abstract: We study whether we can learn novel manipulation skills from human actions to a bi-manual robot with parallel grippers. Human action data is cheap, ab

tutorialsarxiv-cs-cv
29 Jun 2026
Local Ai

TruEye: Fine-Grained Detection of AI-Generated Human Subjects in Images

DGX agent

arXiv:2606.27505v1 Announce Type: new Abstract: AI generated images are proliferating across the Internet. While some are used for entertainment, others are weaponized for fraud and social engineering

local-aiarxiv-cs-cv
29 Jun 2026
Model Releases

TRUST: Efficient Abdominal Trauma Recognition via Image-to-Ultrasound-Video Transfer Learning

DGX agent

arXiv:2606.27777v1 Announce Type: new Abstract: Abdominal ultrasound is indispensable for rapid, noninvasive trauma triage. However, interpreting the subtle dynamic cues embedded in continuous scannin

model-releasesarxiv-cs-cv
29 Jun 2026
Safety

Two-Stage Cross-Domain Cervical Abnormality Screening with Cytopathological Image Synthesis and Knowledge Distillation

DGX agent

arXiv:2606.27678v1 Announce Type: new Abstract: Cross-domain diagnosis remains a major challenge in cervical cell pathology due to pronounced domain shifts across institutions and the subtle visual di

safetyarxiv-cs-cv
29 Jun 2026
Research

Unbiased Diffusion Variational Inversion via Principled Posterior Matching

DGX agent

arXiv:2605.25042v2 Announce Type: replace Abstract: Existing score-based methods for inverse problems often resort to approximate minimization of the KL divergence between the inversion distribution a

researcharxiv-cs-cv
29 Jun 2026
Model Releases

Understanding Cross-Rig Generalization in Automotive Perception: a Multi-Rig Benchmark and Rig Variation Metrics

DGX agent

arXiv:2606.27554v1 Announce Type: new Abstract: Camera-based perception systems for autonomous driving are typically developed and evaluated using fixed sensor rigs, while real-world vehicle fleets ex

model-releasesarxiv-cs-cv
29 Jun 2026
Research

Understanding How MLLMs Describe Artworks Using Token Activation Maps

DGX agent

arXiv:2606.27947v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) describe artworks with remarkable fluency, yet the visual reasoning behind their outputs remains opaque. When a

researcharxiv-cs-cv
29 Jun 2026
Model Releases

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning

DGX agent

arXiv:2606.27828v1 Announce Type: new Abstract: Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic visual evidence rather than merely r

model-releasesarxiv-cs-cv
29 Jun 2026
Model Releases

VLM-Aware Meta-Optic Front-End Design for Frozen Vision-Language Models

DGX agent

arXiv:2606.27646v1 Announce Type: new Abstract: Conventional machine-vision pipelines typically rely on high-quality optics that produce clean, human-interpretable images, and optical design has there

model-releasesarxiv-cs-cv
29 Jun 2026
Agents

VLM-Guided Visual Place Recognition for Planet-Scale Geo-Localization

DGX agent

arXiv:2507.17455v2 Announce Type: replace Abstract: Geo-localization from a single image at planet scale (essentially an advanced or extreme version of the kidnapped robot problem) is a fundamental an

agentsarxiv-cs-cv
29 Jun 2026
Tutorials

Web2Grasp: Learning Functional Grasps from Web Images of Hand-Object Interactions

DGX agent

arXiv:2505.05517v3 Announce Type: replace Abstract: Functional grasping is essential for enabling dexterous multi-finger robot hands to manipulate objects effectively. Prior work largely focuses on po

tutorialsarxiv-cs-cv
29 Jun 2026
Model Releases

ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion Retrieval

DGX agent

arXiv:2606.27708v1 Announce Type: new Abstract: Adapting a foundation vision-language encoder to a specialized retrieval task creates a fundamental tradeoff: gains on the target distribution come at t

model-releasesarxiv-cs-cv
29 Jun 2026
Model Releases

6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models

DGX agent

arXiv:2512.04238v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly integrated into clinical workflows. However, existing benchmarks primarily assess performance on comm

model-releasesarxiv-cs-cv
26 Jun 2026
Research

A UAV-Based Multispectral and RGB Dataset for Multi-Stage Paddy Crop Monitoring in Indian Agricultural Fields

DGX agent

arXiv:2601.01084v2 Announce Type: replace Abstract: We present a large-scale unmanned aerial vehicle (UAV)-based RGB and multispectral image dataset collected over paddy fields in the Vijayawada regio

researcharxiv-cs-cv
26 Jun 2026
Applications

Adversarial Robustness of AI-Generated Image Detectors in the Real World

DGX agent

arXiv:2410.01574v4 Announce Type: replace Abstract: The rapid advancement of Generative Artificial Intelligence (GenAI) capabilities is accompanied by a concerning rise in its misuse. In particular th

applicationsarxiv-cs-cv
26 Jun 2026
Applications

Appearance-Preserving Refinement of Generated 3D Assets for Monochromatic Fabrication

DGX agent

arXiv:2606.26850v1 Announce Type: cross Abstract: Recent advances in 3D mesh generation have enabled the creation of visually realistic assets. However, much of their visual fidelity is encoded in tex

applicationsarxiv-cs-cv
26 Jun 2026
Model Releases

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

DGX agent

arXiv:2606.27376v1 Announce Type: new Abstract: Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision,

model-releasesarxiv-cs-cv
26 Jun 2026
Research

Attention-Based Prototype Calibration for Multi-Rater Few-Shot Medical Image Segmentation

DGX agent

arXiv:2606.16325v2 Announce Type: replace Abstract: Few-shot medical image segmentation methods typically assume a single ground-truth annotation, overlooking systematic variability across expert rate

researcharxiv-cs-cv
26 Jun 2026
Applications

Beyond Aesthetics: Quantifying Information Loss in Turbid Scenes

DGX agent

arXiv:2606.26295v1 Announce Type: new Abstract: Visibility in underwater environments degrades rapidly under turbid conditions, yet the effects on computer-vision models remain unclear. This issue is

applicationsarxiv-cs-cv
26 Jun 2026
Research

Beyond Single-Source Cognitive Taskonomy:Multi-Source Task Relations through fMRI Transfer Learning

DGX agent

arXiv:2606.26279v1 Announce Type: new Abstract: Cognitive tasks are organized by shared and specialized neural processes. Masked fMRI reconstruction provides a common self-supervised objective for qua

researcharxiv-cs-cv
26 Jun 2026
Research

Budget-Aware Keyboardless Interaction

DGX agent

arXiv:2606.26508v1 Announce Type: cross Abstract: Interacting with computers typically relies on traditional input devices such as keyboards, mice, and monitors, which can be cumbersome for users seek

researcharxiv-cs-cv
26 Jun 2026
Safety

Calibrated Harmonic Overlaid Implicit Neural Representations for Multi-Dimensional Data

DGX agent

arXiv:2606.26763v1 Announce Type: new Abstract: Implicit neural representation (INR) has emerged as a powerful prior for multi-dimensional data (e.g., multispectral images and videos). However, most I

safetyarxiv-cs-cv
26 Jun 2026
Research

Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting

DGX agent

arXiv:2606.26754v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) provides an efficient and explicit representation for novel view synthesis, enforcing stylistic coherence across view

researcharxiv-cs-cv
26 Jun 2026
← Previous
1…8586878889…263
Next →