AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

CREward: A Type-Specific Creativity Reward Model

DGX agent

arXiv:2511.19995v2 Announce Type: replace Abstract: Creativity is a complex phenomenon. When it comes to representing and assessing creativity, treating it as a single undifferentiated quantity would

model-releasesarxiv-cs-cv
3 Jun 2026
Research

CropCraft: A Procedural World Generator for Robotic Simulation of Agricultural Tasks

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2511.02417v2 Announce Type: replace Abstract: The adoption of agroecological practices in modern agriculture requires robotic systems capable of operating in highly diverse and complex field env

researcharxiv-cs-cv
3 Jun 2026
Local Ai

Cross-Modality Feature Fusion Based on Structured State Space Duality for Multimodal Image Registration Network

DGX agent

arXiv:2606.03341v1 Announce Type: new Abstract: In multi-modal image registration, the primary challenge lies in shared structural information extraction. Compared to Transformers, Structured State Sp

local-aiarxiv-cs-cv
3 Jun 2026
Model Releases

Cryo-Bench: Benchmarking Foundation Models for Cryosphere Applications

DGX agent

arXiv:2603.01576v3 Announce Type: replace Abstract: Geo-Foundation Models (GFMs) have been evaluated across diverse Earth observation task including multiple domains and have demonstrated strong poten

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

Demo2Tutorial: From Human Experience to Multimodal Software Tutorials

DGX agent

arXiv:2606.03951v1 Announce Type: new Abstract: Human experience in digital environments offers a vast, underexplored resource of authentic, untrimmed interactions that contain rich procedural knowled

model-releasesarxiv-cs-cv
3 Jun 2026
Research

Depth from Dual Differential Defocus and Stereo Consensus

DGX agent

arXiv:2606.02906v1 Announce Type: cross Abstract: We introduce D^3S Consensus, a physics-based, closed-form algorithm that unifies depth-from-defocus (DfD) and stereo to achieve highly accurate depth

researcharxiv-cs-cv
3 Jun 2026
Model Releases

Diagnosis of Human Object Interaction Detectors for Real World Educational Applications

DGX agent

arXiv:2606.02789v1 Announce Type: new Abstract: Human-object interaction (HOI) recognition is critical for automatically analyzing student behavior in complex educational environments. Although state-

model-releasesarxiv-cs-cv
3 Jun 2026
Research

Diffusing in the Right Space: A Systematic Study of Latent Diffusability

DGX agent

arXiv:2606.03578v1 Announce Type: new Abstract: Latent diffusion models leverage visual tokenizers to compress images into latent spaces for efficient generative modeling. However, better reconstructi

researcharxiv-cs-cv
3 Jun 2026
Model Releases

Disentangling Visual and Factual Correctness in LVLMs' Visualization Literacy

DGX agent

arXiv:2606.03142v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) show strong visualization interpretation, yet it is unclear whether their responses reflect genuine reasoning over

model-releasesarxiv-cs-cv
3 Jun 2026
Agents

DyaPlex: Full-Duplex Speech-Motion Model for Dyadic Interaction

DGX agent

arXiv:2606.03874v1 Announce Type: new Abstract: We present DyaPlex, a streaming, full-duplex speech-and-motion model designed for dyadic interaction. To capture the continuous and reciprocal nature of

agentsarxiv-cs-cv
3 Jun 2026
Research

Electromagnetic Navigation for Femoral Osteotomy Using High-Accuracy X-ray-to-CT Registration

DGX agent

arXiv:2606.03893v1 Announce Type: new Abstract: Accurate execution of preoperative plans in corrective femoral osteotomies remains challenging. Current techniques are limited by variable accuracy, inv

researcharxiv-cs-cv
3 Jun 2026
Model Releases

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching

DGX agent

arXiv:2606.03577v1 Announce Type: new Abstract: Wide-baseline matching (WBM) requires integrating geometric understanding, viewpoint changes, fine-grained perception, and occlusion reasoning, making i

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

Enginuity: A Dataset and Benchmark for Vision-Language Understanding of Engineering Diagrams

DGX agent

arXiv:2606.03410v1 Announce Type: new Abstract: Engineering diagrams pose a distinct challenge for vision-language models: unlike natural images or general documents, they encode information through d

model-releasesarxiv-cs-cv
3 Jun 2026
Research

Estimating Central, Peripheral, and Temporal Visual Contributions to Human Decision Making in Atari Games

DGX agent

arXiv:2604.04439v2 Announce Type: replace-cross Abstract: We study how different visual information sources contribute to human decision making in dynamic visual environments. Using Atari-HEAD, a larg

researcharxiv-cs-cv
3 Jun 2026
Safety

EvoMemNav: Efficient Self-Evolving Fine-Grained Memory for Zero-Shot Embodied Navigation

DGX agent

arXiv:2606.03509v1 Announce Type: new Abstract: Building memory is essential for long-horizon planning in zero-shot embodied navigation. Detector-centric scene graphs often compress observations into

safetyarxiv-cs-cv
3 Jun 2026
Research

Exploring Easy Boosts for Lidar Semantic Scene Completion

DGX agent

arXiv:2606.03992v1 Announce Type: new Abstract: This paper investigates 'free lunch' strategies to boost the performance of lidar semantic scene completion (SSC) without requiring complex architectura

researcharxiv-cs-cv
3 Jun 2026
Agents

Face versus Body Tracking for Human-Robot Interaction: An Egocentric Dataset

DGX agent

arXiv:2606.03694v1 Announce Type: cross Abstract: To enable meaningful human-robot interaction (HRI), a robot must continuously assess engagement by consistently tracking users over time. State-of-the

agentsarxiv-cs-cv
3 Jun 2026
Safety

FAF-CD: Frequency-Aware Fusion for Change Detection under Imperfect Multimodal Remote Sensing

DGX agent

arXiv:2606.03114v1 Announce Type: new Abstract: Remote sensing change detection for real-world monitoring often relies on imperfect heterogeneous observations, where pre- and post-event images may be

safetyarxiv-cs-cv
3 Jun 2026
Tutorials

FCUS-rPPG: A Fast-Converging Unsupervised Framework for Remote Photoplethysmography via Gradient Oscillation Suppression

DGX agent

arXiv:2606.03050v1 Announce Type: new Abstract: Remote photoplethysmography (rPPG) enables non-contact extraction of blood volume pulse (BVP) signals using consumer-grade cameras. Recent unsupervised

tutorialsarxiv-cs-cv
3 Jun 2026
Safety

Follow-Your-Preference++: Rethinking Preference Alignment for Image Inpainting

DGX agent

arXiv:2606.03216v1 Announce Type: new Abstract: We study preference alignment for image inpainting. Rather than proposing yet another method, we revisit the problem from first principles and reassess

safetyarxiv-cs-cv
3 Jun 2026
Safety

FreeStreamGS: Online Feed-forward 3D Gaussian Splatting from Unposed Streaming Inputs

DGX agent

arXiv:2606.03254v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting (3DGS) allows efficient and high-fidelity novel view synthesis (NVS) from an offline recorded image sequence. However

safetyarxiv-cs-cv
3 Jun 2026
Local Ai

From 3D Perception to Safety Reasoning: A Graph-Based Framework for Real-Time Underground Mine Monitoring

DGX agent

arXiv:2606.03460v1 Announce Type: new Abstract: Underground coal mining requires personnel and heavy equipment to operate within shared, confined, and poorly illuminated spaces where hazards such as e

local-aiarxiv-cs-cv
3 Jun 2026
Model Releases

From Local Training to Large-Scale Mapping: A Comparative Assessment of Machine Learning and Deep Learning for Transferable Satellite-Derived Bathymetry

DGX agent

arXiv:2606.02764v1 Announce Type: new Abstract: Satellite-derived bathymetry (SDB) from multispectral imagery is cost-effective but scales poorly across regions, especially in optically complex coasta

model-releasesarxiv-cs-cv
3 Jun 2026
Research

From None to All: Self-Supervised 3D Reconstruction via Novel View Synthesis

DGX agent

arXiv:2603.27455v2 Announce Type: replace Abstract: In this paper, we introduce NAS3R, a self-supervised feed-forward framework that jointly learns explicit 3D geometry and camera parameters with no g

researcharxiv-cs-cv
3 Jun 2026
Research

GARDEN: Gravity-Aligned Reconstruction of Disentangled ENvironments from RGB images

DGX agent

arXiv:2606.03921v1 Announce Type: new Abstract: Converting multi-view RGB observations into simulation-ready 3D environments remains challenging because current reconstruction pipelines produce monoli

researcharxiv-cs-cv
3 Jun 2026
Model Releases

GeoDrive-Bench: Benchmarking Region-Specific Multimodal Reasoning in Autonomous Driving

DGX agent

arXiv:2606.02774v1 Announce Type: new Abstract: Vision-language models (VLMs) for autonomous driving have shown promising performance, but their ability to handle region-specific traffic rules remains

model-releasesarxiv-cs-cv
3 Jun 2026
Research

Graph Mamba Survival Analysis Based on Topology-Aware ordering

DGX agent

arXiv:2606.02602v1 Announce Type: cross Abstract: In computational pathology, Whole Slide Images (WSIs) survival analysis is crucial for patient prognosis assessment, but it faces multiple technical c

researcharxiv-cs-cv
3 Jun 2026
Local Ai

Graph Regularized Non-negative Reduced Biquaternion Matrix Factorization for Color Image Recognition

DGX agent

arXiv:2606.03654v1 Announce Type: new Abstract: Non-negative reduced biquaternion matrix factorization (NRBMF) uses the product of reduced biquaternion (RB) matrices to incorporate the non-negativity

local-aiarxiv-cs-cv
3 Jun 2026
Tutorials

GS-ROR^2: Bidirectional-guided 3DGS and SDF for Reflective Object Relighting and Reconstruction

DGX agent

arXiv:2406.18544v4 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has shown a powerful capability for novel view synthesis due to its detailed expressive ability and highly efficient re

tutorialsarxiv-cs-cv
3 Jun 2026
Applications

Hierarchical Federated Learning with Dynamic Clustering and Adaptive Regularization for Robust Infrastructure Inspection

DGX agent

arXiv:2606.03084v1 Announce Type: new Abstract: The deployment of data-driven computer vision models for structural health monitoring (SHM) is heavily constrained by the data silo dilemma due to strin

applicationsarxiv-cs-cv
3 Jun 2026
Research

How Much of a Model Do We Need? Redundancy and Slimmability in Remote Sensing Foundation Models

DGX agent

arXiv:2601.22841v2 Announce Type: replace Abstract: Large-scale foundation models (FMs) in remote sensing (RS) (denoted as RS FMs) are developed following paradigms established in computer vision (CV)

researcharxiv-cs-cv
3 Jun 2026
Applications

Hybrid Autoregressive-Diffusion Model for Real-Time Sign Language Production

DGX agent

arXiv:2507.09105v4 Announce Type: replace Abstract: Earlier Sign Language Production (SLP) models typically relied on autoregressive decoding, which naturally preserves temporal causality but suffers

applicationsarxiv-cs-cv
3 Jun 2026
Research

IdEst: Assessing Self-Supervised Learning Representations via Intrinsic Dimension

DGX agent

arXiv:2606.03338v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has emerged as a powerful paradigm for learning meaningful representations from unlabeled data. However, the standard p

researcharxiv-cs-cv
3 Jun 2026
Tutorials

IDO: Incongruity-aware Distribution Optimization for Multimodal Fake News Detection

DGX agent

arXiv:2606.03418v1 Announce Type: new Abstract: Multimodal fake news detection aims to identify the authenticity of news. Existing multimodal fake news detection methods mainly focus on cross-modal co

tutorialsarxiv-cs-cv
3 Jun 2026
Safety

Inference-Time Scaling for Joint Audio-Video Generation

DGX agent

arXiv:2606.03183v1 Announce Type: cross Abstract: Joint audio-video generation aims to synthesize realistic audio-video pairs that are both semantically aligned with text prompts and precisely synchro

safetyarxiv-cs-cv
3 Jun 2026
Research

Inverting the Generation Process of Denoising Diffusion Implicit Models: Empirical Evaluation and a Novel Method

DGX agent

arXiv:2606.03111v1 Announce Type: new Abstract: This paper studies the problem of inverting the DDIM image generation process to recover latent variables, particularly the initial noise map, from a ge

researcharxiv-cs-cv
3 Jun 2026
Model Releases

Investigating Adversarial Robustness of Multi-modal Large Language Models

DGX agent

arXiv:2606.03713v1 Announce Type: new Abstract: Multi-modal Large Language Models (MLLMs) achieve strong performance on vision-language tasks, but incorporating visual inputs through a vision encoder

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation

DGX agent

arXiv:2606.03168v1 Announce Type: new Abstract: While instruction-based video editing has seen significant progress, joint audio-visual editing remains constrained by the absence of dedicated datasets

model-releasesarxiv-cs-cv
3 Jun 2026
Safety

KC-3DGS: Kurtosis-Constrained Gaussian Splatting for High-Fidelity View Synthesis

DGX agent

arXiv:2606.03120v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) enables real-time novel view synthesis by representing scenes as collections of anisotropic Gaussians optimized via differe

safetyarxiv-cs-cv
3 Jun 2026
Model Releases

Knowledge-Preserved Model Tuning in Null-Space for Robust Spatio-Temporal Video Grounding

DGX agent

arXiv:2606.03539v1 Announce Type: new Abstract: Spatio-Temporal Video Grounding aims to localize object tubes based on textual queries. While recent methods have achieved remarkable success, they main

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

LAMP: Data-Efficient Linear Affine Weight-Space Models for Parameter-Controlled 3D Shape Generation and Extrapolation

DGX agent

arXiv:2510.22491v3 Announce Type: replace-cross Abstract: Generating high-fidelity 3D geometries under explicit parameter constraints is central to engineering design, yet current methods often requir

model-releasesarxiv-cs-cv
3 Jun 2026
Agents

Learning to See via Epiretinal Implant Stimulation in silico with Model-Based Deep Reinforcement Learning

DGX agent

arXiv:2606.03118v1 Announce Type: cross Abstract: Objective: Diseases such as age-related macular degeneration and retinitis pigmentosa cause the degradation of the photoreceptor layer. One approach t

agentsarxiv-cs-cv
3 Jun 2026
Safety

LoCAtion: Long-time Collaborative Attention Framework for High Dynamic Range Video Reconstruction

DGX agent

arXiv:2603.14377v2 Announce Type: replace Abstract: Prevailing High Dynamic Range (HDR) video reconstruction methods are fundamentally trapped in a fragile alignment-and-fusion paradigm. While explici

safetyarxiv-cs-cv
3 Jun 2026
Model Releases

Low-Frequency Shortcuts in Texture-Driven Visual Learning

DGX agent

arXiv:2606.03493v1 Announce Type: new Abstract: Neural networks suffer from shortcut learning, where learned features generalize well to the training set but not to in-distribution (ID) or out-of-dist

model-releasesarxiv-cs-cv
3 Jun 2026
Research

Low-Resolution Editing is All You Need for High-Resolution Editing

DGX agent

arXiv:2511.19945v3 Announce Type: replace Abstract: High-resolution content creation is rapidly emerging as a central challenge in both the vision and graphics communities. Images serve as the most fu

researcharxiv-cs-cv
3 Jun 2026
Research

MAdam: Metric-Aware Multi-Objective Adam

DGX agent

arXiv:2606.03904v1 Announce Type: cross Abstract: Multi-objective optimization (MOO) underlies many machine learning problems, yet MOO solvers across the loss-balancing, gradient-balancing, and Pareto

researcharxiv-cs-cv
3 Jun 2026
Research

Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation

DGX agent

arXiv:2606.03402v1 Announce Type: new Abstract: Audio-driven human motion video generation aims to synthesize realistic and temporally coherent human animations from a single static image, with applic

researcharxiv-cs-cv
3 Jun 2026
Agents

MariData: One-Step Unpaired Image Translation for Maritime Environments

DGX agent

arXiv:2606.03246v1 Announce Type: new Abstract: The development on robust perception systems for Maritime Autonomous Surface Ships (MASS) is heavily constrained by the scarcity of diverse training dat

agentsarxiv-cs-cv
3 Jun 2026
← Previous
1…122123124125126…263
Next →