AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
29 Jun 2026

Radar Guided Camera Verification for Automatic Emergency Braking Rethinking Object Detection in Radar Camera Fusion

Local AiDGX agent

arXiv:2606.27556v1 Announce Type: new Abstract: Radar camera fusion is widely used in Automatic Emergency Braking AEB systems because radar provides reliable range and velocity measurements while came

RAE-NWM: Navigation World Model in Dense Visual Representation Space

TutorialsDGX agent

arXiv:2603.09241v2 Announce Type: replace Abstract: Visual navigation requires agents to reach goals in complex environments through perception and planning. World models address this task by simulati

RANSAC Scoring Done Right

Model ReleasesDGX agent

arXiv:2606.27385v1 Announce Type: cross Abstract: The most widely used RANSAC variants score candidate models by counting inliers or summing per-point scores that saturate beyond a residual threshold.


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

ReScene: Structured Indoor Scene Reconstruction from Multi-View Captures

Model ReleasesDGX agent

arXiv:2606.28060v1 Announce Type: new Abstract: Constructing simulation-ready 3D scenes from multi-view captures is a key bottleneck for Embodied Artificial Intelligence, as downstream tasks require o

ReWorld: Learning Better Representations for World Action Models

SafetyDGX agent

arXiv:2606.27504v1 Announce Type: new Abstract: World Action Models (WAMs) model future environment evolution under action conditioning, offering a scalable paradigm for autonomous driving. However, e

Rheos: Modelling Continuous Motion Dynamics in Hierarchical 3D Scene Graphs

ResearchDGX agent

arXiv:2603.20239v2 Announce Type: replace-cross Abstract: 3D Scene Graphs (3DSGs) provide hierarchical, multi-resolution abstractions that encode the geometric and semantic structure of an environment

RPM-Distill: Physiology-guided Adaptive Cross-modal Distillation for Robust Remote Physiological Measurement

SafetyDGX agent

arXiv:2606.28089v1 Announce Type: new Abstract: Video-based remote physiological measurement (RPM) is highly accessible but remains fragile under varying illumination, skin tones, and motion. Radio fr

RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning

Model ReleasesDGX agent

arXiv:2606.28266v1 Announce Type: new Abstract: Remote Sensing Image Change Captioning (RSICC) aims to describe changes between bi-temporal remote sensing images and holds significant research and app

Scalable and Differentiable Point-Cloud Registration Using Maximum Mean Discrepancy

SafetyDGX agent

arXiv:2606.27818v1 Announce Type: new Abstract: We present MMD-Reg, a novel correspondence-free approach to point-cloud registration that is differentiable and has linear computational complexity in t

ScaLe-INR: Scale and Learn Implicit Neural Representations

SafetyDGX agent

arXiv:2606.27862v1 Announce Type: new Abstract: Implicit Neural Representations (INRs) parameterized by multilayer perceptrons excel at modeling continuous signals. However, a key challenge persists a

Scene and Human in One World: Reconstruction in a Feedforward Pass

Local AiDGX agent

arXiv:2606.27720v1 Announce Type: new Abstract: Reconstructing humans in dynamic scenes from moving monocular cameras remains challenging due to scale ambiguity, human-scene misalignment, and occlusio

Scene Generation at Absolute Scale: Utilizing Semantic and Geometric Guidance From Text for Accurate and Interpretable 3D Indoor Scene Generation

SafetyDGX agent

arXiv:2603.13910v2 Announce Type: replace Abstract: We present GuidedSceneGen, a text-to-3D generation framework that produces metrically accurate, globally consistent, and semantically interpretable

SelectAnyTree: A Promptable Instance Segmentation Model for 3D Forest LiDAR Point Clouds

Local AiDGX agent

arXiv:2606.27491v1 Announce Type: new Abstract: Automated instance segmentation of forest LiDAR point clouds is increasingly critical as forest monitoring moves toward scalable, detailed, 3D measureme

SemCityLoc: Aerial 6DoF Localization Using Semantic 3D City Models

Model ReleasesDGX agent

arXiv:2606.27444v1 Announce Type: new Abstract: Aerial 6DoF localization typically relies on precise GNSS signals or radiometrically rich 3D reconstructions, limiting scalability and on-board deployme

SHIFT: Motion Alignment in Video Diffusion Models with Adversarial Hybrid Fine-Tuning

SafetyDGX agent

arXiv:2603.17426v2 Announce Type: replace Abstract: Image-conditioned video diffusion models achieve impressive visual realism but often suffer from weakened motion fidelity, e.g., reduced motion dyna

SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models

SafetyDGX agent

arXiv:2606.27741v1 Announce Type: new Abstract: Recent advances in video diffusion models have greatly improved visual fidelity, yet their generated motions often violate physical plausibility. We obs

Spectral Subsurface Scattering from RGB via Biophysical Skin Inversion

ResearchDGX agent

arXiv:2606.27604v1 Announce Type: cross Abstract: In this paper we present a spectral optical inversion for skin for path tracing-based rendering of subsurface scattering. Skin is a complex multilayer

StableMotion: One-Step Motion Estimation with Diffusion Prior

ResearchDGX agent

arXiv:2505.06668v2 Announce Type: replace Abstract: We present StableMotion, a novel framework that leverages geometric and content priors from pretrained large-scale image diffusion models for motion

StructSplat: Generalizable 3D Gaussian Splatting from Uncalibrated Sparse Views

SafetyDGX agent

arXiv:2606.28321v1 Announce Type: new Abstract: We present StructSplat, a feed-forward and generalizable 3D Gaussian reconstruction framework that operates directly on uncalibrated images without requ

Structured-Li-GS: Structured 3D Gaussians Splatting with LiDAR Incorporation and Spatial Constraints

Model ReleasesDGX agent

arXiv:2606.27509v1 Announce Type: new Abstract: In this study, we develop a Structured framework for Gaussian Splatting (3DGS) with LiDAR integration (Structured-Li-GS). It is a lightweight Gaussian S

SurgXBench: Explainable Vision-Language Model Benchmark for Surgery

Model ReleasesDGX agent

arXiv:2505.10764v4 Announce Type: replace Abstract: Innovations in digital intelligence are transforming robotic surgery with more informed decision-making. Real-time awareness of surgical instrument

SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation

SafetyDGX agent

arXiv:2508.06115v3 Announce Type: replace Abstract: Semantic segmentation in open-vocabulary scenarios presents significant challenges due to the wide range and granularity of semantic categories. Exi

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

Model ReleasesDGX agent

arXiv:2510.03117v2 Announce Type: replace Abstract: This study focuses on a challenging yet promising task, Text-to-Sounding-Video (T2SV) generation, which aims to generate a video with synchronized a

TempAct: Advancing Temporal Plausibility in Autoregressive Video Generation via Planner-Executor RL

SafetyDGX agent

arXiv:2606.28016v1 Announce Type: new Abstract: Autoregressive (AR) video diffusion models enable low-latency streaming generation by synthesizing videos chunk by chunk with cached visual context, but

Temporal-Emerged Prompting for Segment Anything in Multiframe Infrared Small Target Detection

Local AiDGX agent

arXiv:2606.27655v1 Announce Type: new Abstract: Accurately localizing and segmenting small targets in low signal-to-noise ratio (SNR) infrared sequences remains a challenging task. Since targets are o

Tessellating The Earth

Local AiDGX agent

arXiv:2606.27514v1 Announce Type: new Abstract: Geolocation encoders, which map geographic coordinates to learned representations, are emerging as an effective means of capturing visual and non-visual

Text as Illumination: Spatial Contrastive Retinex Learning for Language-guided Medical Image Segmentation

SafetyDGX agent

arXiv:2606.27794v1 Announce Type: new Abstract: Language-guided Medical Image Segmentation (LMIS) has shown great potential to improve the delineation of anatomical structures and lesions by integrati

TextDS: Parameter-Efficient Representation Alignment for Scene Text Detection under Distribution Shifts

Model ReleasesDGX agent

arXiv:2606.28077v1 Announce Type: new Abstract: In real-world deployments, scene text detectors inevitably face distribution shifts beyond the training distribution. Prior work often depends on large-

There and Back Again: A Flexible-Frame Transformer for Multi-Exposure Fusion

Model ReleasesDGX agent

arXiv:2606.27905v1 Announce Type: new Abstract: Multi-exposure fusion (MEF) brings the dynamic range of conventional cameras closer to that of human vision, producing images with rich scene content. G

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs

ResearchDGX agent

arXiv:2510.00705v3 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) often struggle with fine-grained perception, such as identifying small objects in high-resolution images or

Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robots

TutorialsDGX agent

arXiv:2606.28133v1 Announce Type: cross Abstract: We study whether we can learn novel manipulation skills from human actions to a bi-manual robot with parallel grippers. Human action data is cheap, ab

TruEye: Fine-Grained Detection of AI-Generated Human Subjects in Images

Local AiDGX agent

arXiv:2606.27505v1 Announce Type: new Abstract: AI generated images are proliferating across the Internet. While some are used for entertainment, others are weaponized for fraud and social engineering

TRUST: Efficient Abdominal Trauma Recognition via Image-to-Ultrasound-Video Transfer Learning

Model ReleasesDGX agent

arXiv:2606.27777v1 Announce Type: new Abstract: Abdominal ultrasound is indispensable for rapid, noninvasive trauma triage. However, interpreting the subtle dynamic cues embedded in continuous scannin

Two-Stage Cross-Domain Cervical Abnormality Screening with Cytopathological Image Synthesis and Knowledge Distillation

SafetyDGX agent

arXiv:2606.27678v1 Announce Type: new Abstract: Cross-domain diagnosis remains a major challenge in cervical cell pathology due to pronounced domain shifts across institutions and the subtle visual di

Unbiased Diffusion Variational Inversion via Principled Posterior Matching

ResearchDGX agent

arXiv:2605.25042v2 Announce Type: replace Abstract: Existing score-based methods for inverse problems often resort to approximate minimization of the KL divergence between the inversion distribution a

Understanding Cross-Rig Generalization in Automotive Perception: a Multi-Rig Benchmark and Rig Variation Metrics

Model ReleasesDGX agent

arXiv:2606.27554v1 Announce Type: new Abstract: Camera-based perception systems for autonomous driving are typically developed and evaluated using fixed sensor rigs, while real-world vehicle fleets ex

Understanding How MLLMs Describe Artworks Using Token Activation Maps

ResearchDGX agent

arXiv:2606.27947v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) describe artworks with remarkable fluency, yet the visual reasoning behind their outputs remains opaque. When a

Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning

Model ReleasesDGX agent

arXiv:2606.27828v1 Announce Type: new Abstract: Recent interest in multimodal large language models (MLLMs) raises a central question: can they reason over dynamic visual evidence rather than merely r

VLM-Aware Meta-Optic Front-End Design for Frozen Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.27646v1 Announce Type: new Abstract: Conventional machine-vision pipelines typically rely on high-quality optics that produce clean, human-interpretable images, and optical design has there

VLM-Guided Visual Place Recognition for Planet-Scale Geo-Localization

AgentsDGX agent

arXiv:2507.17455v2 Announce Type: replace Abstract: Geo-localization from a single image at planet scale (essentially an advanced or extreme version of the kidnapped robot problem) is a fundamental an

Web2Grasp: Learning Functional Grasps from Web Images of Hand-Object Interactions

TutorialsDGX agent

arXiv:2505.05517v3 Announce Type: replace Abstract: Functional grasping is essential for enabling dexterous multi-finger robot hands to manipulate objects effectively. Prior work largely focuses on po

ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion Retrieval

Model ReleasesDGX agent

arXiv:2606.27708v1 Announce Type: new Abstract: Adapting a foundation vision-language encoder to a specialized retrieval task creates a fundamental tradeoff: gains on the target distribution come at t

26 Jun 2026

6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models

Model ReleasesDGX agent

arXiv:2512.04238v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly integrated into clinical workflows. However, existing benchmarks primarily assess performance on comm

A UAV-Based Multispectral and RGB Dataset for Multi-Stage Paddy Crop Monitoring in Indian Agricultural Fields

ResearchDGX agent

arXiv:2601.01084v2 Announce Type: replace Abstract: We present a large-scale unmanned aerial vehicle (UAV)-based RGB and multispectral image dataset collected over paddy fields in the Vijayawada regio

Adversarial Robustness of AI-Generated Image Detectors in the Real World

ApplicationsDGX agent

arXiv:2410.01574v4 Announce Type: replace Abstract: The rapid advancement of Generative Artificial Intelligence (GenAI) capabilities is accompanied by a concerning rise in its misuse. In particular th

Appearance-Preserving Refinement of Generated 3D Assets for Monochromatic Fabrication

ApplicationsDGX agent

arXiv:2606.26850v1 Announce Type: cross Abstract: Recent advances in 3D mesh generation have enabled the creation of visually realistic assets. However, much of their visual fidelity is encoded in tex

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

Model ReleasesDGX agent

arXiv:2606.27376v1 Announce Type: new Abstract: Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision,

Attention-Based Prototype Calibration for Multi-Rater Few-Shot Medical Image Segmentation

ResearchDGX agent

arXiv:2606.16325v2 Announce Type: replace Abstract: Few-shot medical image segmentation methods typically assume a single ground-truth annotation, overlooking systematic variability across expert rate

Beyond Aesthetics: Quantifying Information Loss in Turbid Scenes

ApplicationsDGX agent

arXiv:2606.26295v1 Announce Type: new Abstract: Visibility in underwater environments degrades rapidly under turbid conditions, yet the effects on computer-vision models remain unclear. This issue is

Beyond Single-Source Cognitive Taskonomy:Multi-Source Task Relations through fMRI Transfer Learning

ResearchDGX agent

arXiv:2606.26279v1 Announce Type: new Abstract: Cognitive tasks are organized by shared and specialized neural processes. Masked fMRI reconstruction provides a common self-supervised objective for qua

Budget-Aware Keyboardless Interaction

ResearchDGX agent

arXiv:2606.26508v1 Announce Type: cross Abstract: Interacting with computers typically relies on traditional input devices such as keyboards, mice, and monitors, which can be cumbersome for users seek

Calibrated Harmonic Overlaid Implicit Neural Representations for Multi-Dimensional Data

SafetyDGX agent

arXiv:2606.26763v1 Announce Type: new Abstract: Implicit neural representation (INR) has emerged as a powerful prior for multi-dimensional data (e.g., multispectral images and videos). However, most I

Capacity-Controlled Multi-View Stylization of 3D Gaussian Splatting

ResearchDGX agent

arXiv:2606.26754v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) provides an efficient and explicit representation for novel view synthesis, enforcing stylistic coherence across view

Circular Quasiconformal Deturbulence: Geometry-Based Restoration from Multiple Turbulent Frames

ApplicationsDGX agent

arXiv:2504.13432v3 Announce Type: replace Abstract: Imaging through inhomogeneous media often results in severe distortions, posing significant challenges to downstream image-processing tasks. The lac

Coarse-to-Fine: A Hybrid Self-Supervised Method for Non-rigid 3D Shape Matching

ResearchDGX agent

arXiv:2606.26557v1 Announce Type: new Abstract: Non-rigid 3D shape matching is a fundamental task in computer vision and graphics. In this paper, we propose a hybrid self-supervised method based on a

CogAD: Cognitive-Hierarchy Guided End-to-End Autonomous Driving

AgentsDGX agent

arXiv:2505.21581v4 Announce Type: replace-cross Abstract: While end-to-end autonomous driving has advanced significantly, prevailing methods remain fundamentally misaligned with human cognitive princi

Computer Vision for MOBA Analytics: A Dataset and Baseline for Visibility Analysis in Dota 2

ResearchDGX agent

arXiv:2606.26970v1 Announce Type: new Abstract: Introduction: Most Multiplayer Online Battle Arena (MOBA) analytics studies rely on structured data, which does not directly capture what each team coul

Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints

ResearchDGX agent

arXiv:2603.11755v2 Announce Type: replace Abstract: Controllable video generation for complex hand-object interactions is a critical step toward building visual world models. However, existing methods

CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs

Model ReleasesDGX agent

arXiv:2606.27264v1 Announce Type: new Abstract: Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form text jud

DanceDuo: Bridging Human Movement and AI Choreography

ResearchDGX agent

arXiv:2606.26507v1 Announce Type: cross Abstract: In recent years, advancements in deep learning and generative models have revolutionized music-driven dance generation. This paper introduces a novel

← Previous
1…6667686970…209
Next →