AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Applications

MotionPRO: Exploring the Role of Pressure in Human MoCap and Beyond

DGX agent

arXiv:2504.05046v2 Announce Type: replace Abstract: Existing human Motion Capture (MoCap) methods mostly focus on the visual similarity while neglecting the physical plausibility. As a result, downstr

applicationsarxiv-cs-cv
27 May 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale

DGX agent

arXiv:2605.27235v1 Announce Type: new Abstract: Layered image generation and editing is a fundamental capability that enables layer-wise reuse, editing, and composition of generated visual content, an

model-releasesarxiv-cs-cv
27 May 2026
Research

MSCGC-KAN: Multi-scale Causal Graph Convolution and Kolmogorov-Arnold Feature Mapping for EEG Emotion Recognition

DGX agent

arXiv:2605.26624v1 Announce Type: new Abstract: Electroencephalogram (EEG)-based emotion recognition is an important affective computing task, and recent EEG foundation models provide useful generic r

researcharxiv-cs-cv
27 May 2026
Research

Multi-Modal Building Inspection via Perceiver IO Fusion of Satellite and Street-Level Imagery

DGX agent

arXiv:2605.26381v1 Announce Type: new Abstract: We present a multi-modal classification framework that fuses satellite and street-level imagery through a Perceiver IO architecture operating on spatial

researcharxiv-cs-cv
27 May 2026
Safety

MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation

DGX agent

arXiv:2602.09878v2 Announce Type: replace Abstract: World-model-based imagine-then-act becomes a promising paradigm for robotic manipulation, yet existing approaches typically support either purely im

safetyarxiv-cs-cv
27 May 2026
Research

Natural Human Motion Recovery by Aligning High-Order Temporal Dynamics from Monocular Videos

DGX agent

arXiv:2605.26879v1 Announce Type: new Abstract: Human motion recovered from monocular videos often appears overly smooth or dynamically inconsistent, even when joint positions are numerically accurate

researcharxiv-cs-cv
27 May 2026
Applications

NeR-SC: Adapting Neural Video Representation to Screen Content

DGX agent

arXiv:2605.27024v1 Announce Type: new Abstract: Implicit neural representations have emerged as a promising paradigm for video compression, with recent methods achieving competitive performance on nat

applicationsarxiv-cs-cv
27 May 2026
Applications

No Data? No Problem: Robust Vision-Tabular Learning with Missing Values

DGX agent

arXiv:2512.19602v2 Announce Type: replace Abstract: Large-scale medical biobanks provide imaging data complemented by extensive tabular information, such as clinical measurements or demographics. Howe

applicationsarxiv-cs-cv
27 May 2026
Research

Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos

DGX agent

arXiv:2605.26232v1 Announce Type: new Abstract: Pre-trained video large language models excel at visual reasoning. However, they struggle when videos arrive with auxiliary streams, such as audio, dept

researcharxiv-cs-cv
27 May 2026
Model Releases

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding

DGX agent

arXiv:2605.26584v1 Announce Type: new Abstract: Omnimodal large language models enable unified audio video understanding, but long joint token sequences make inference costly, and existing benchmarks

model-releasesarxiv-cs-cv
27 May 2026
Research

Object Pose and Shape Estimation for Grasping: Does it Work?

DGX agent

arXiv:2605.26944v1 Announce Type: cross Abstract: The problem of object pose and shape estimation has seen key advancements lately. Encoder-decoder (e.g., SAM3D, LRM, CRISP) and diffusion-based models

researcharxiv-cs-cv
27 May 2026
Model Releases

ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection

DGX agent

arXiv:2508.01253v2 Announce Type: replace Abstract: Existing studies typically investigate domain shift and category shift as independent problems, however, in real-world scenarios, the two types of s

model-releasesarxiv-cs-cv
27 May 2026
Local Ai

OmniGF: A Dual-Branch Vision-Language Framework for Unified Gaze Following

DGX agent

arXiv:2605.26399v1 Announce Type: new Abstract: Understanding human gaze behavior is essential for complex scene comprehension and human-computer interaction. Traditional gaze following models are typ

local-aiarxiv-cs-cv
27 May 2026
Model Releases

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation

DGX agent

arXiv:2605.26641v1 Announce Type: new Abstract: Unified multimodal embedding spaces have become the standard interface for cross-modal retrieval and multimodal RAG, and recent audio-video-text (AVT) e

model-releasesarxiv-cs-cv
27 May 2026
Research

On the Robustness of Machine Unlearning for Vision-Language Models

DGX agent

arXiv:2605.26992v1 Announce Type: new Abstract: Vision-language models (VLMs) may memorize undesirable information from training data, motivating growing interest in machine unlearning. In this work,

researcharxiv-cs-cv
27 May 2026
Research

Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning

DGX agent

arXiv:2605.26761v1 Announce Type: new Abstract: Multimodal instruction tuning is the de facto recipe for adapting vision language models (VLMs), yet instruction data are highly redundant, making data

researcharxiv-cs-cv
27 May 2026
Model Releases

OSMa-Bench++: Toward Open-Ended Benchmarking of Semantic Mapping for Manipulation with Prompt-Generated Synthetic Scenes

DGX agent

arXiv:2605.26831v1 Announce Type: new Abstract: Semantic mapping methods are increasingly used as intermediate scene representations for downstream robotic reasoning and manipulation, yet their evalua

model-releasesarxiv-cs-cv
27 May 2026
Research

PARE: Pruning and Adaptive Routing for Efficient Video Generation

DGX agent

arXiv:2605.27336v1 Announce Type: new Abstract: Video Diffusion Transformers (DiTs) generate high-quality videos but demand substantial compute due to wide blocks, deep architectures, and iterative sa

researcharxiv-cs-cv
27 May 2026
Tutorials

PILOT: A Data-Free Continual Learning Approach for Real-Time Semantic Segmentation via Boundary Guidance

DGX agent

arXiv:2605.27128v1 Announce Type: new Abstract: Real-time semantic segmentation models offer an excellent balance between accuracy and inference speed. However, deploying these models in dynamic real

tutorialsarxiv-cs-cv
27 May 2026
Research

PlayClass: Automated Play Behaviour Classification in Poultry

DGX agent

arXiv:2605.27304v1 Announce Type: new Abstract: Automated monitoring of animal welfare has largely targeted negative indicators, leaving positive welfare behaviours such as play underexplored. To addr

researcharxiv-cs-cv
27 May 2026
Model Releases

PRBench: A Standardized Probabilistic Robustness Benchmark

DGX agent

arXiv:2511.01724v3 Announce Type: replace Abstract: Deep learning models are notoriously vulnerable to imperceptible perturbations. Most existing research centers on adversarial robustness (AR), which

model-releasesarxiv-cs-cv
27 May 2026
Local Ai

Prototyping an End-to-End Multi-Modal Tiny-CNN for Cardiovascular Sensor Patches

DGX agent

arXiv:2510.18668v2 Announce Type: replace-cross Abstract: The vast majority of cardiovascular diseases may be preventable if early signs and risk factors are detected. Cardiovascular monitoring with b

local-aiarxiv-cs-cv
27 May 2026
Research

Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation

DGX agent

arXiv:2507.16116v2 Announce Type: replace Abstract: The rapid advancement of video diffusion models has been hindered by fundamental limitations in temporal modeling, particularly the rigid synchroniz

researcharxiv-cs-cv
27 May 2026
Safety

PyCAT4: A Hierarchical Vision Transformer-based Framework for 3D Human Pose Estimation

DGX agent

arXiv:2508.02806v3 Announce Type: replace Abstract: Recently, a significant improvement in the accuracy of 3D human pose estimation has been achieved by combining convolutional neural networks (CNNs)

safetyarxiv-cs-cv
27 May 2026
Research

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning

DGX agent

arXiv:2605.27318v1 Announce Type: new Abstract: Video spatial reasoning requires accumulating viewpoint-dependent evidence over time while retaining information useful to the question being asked. Exi

researcharxiv-cs-cv
27 May 2026
Research

R^3: 3D Reconstruction via Relative Regression

DGX agent

arXiv:2605.26519v1 Announce Type: new Abstract: Recent feed-forward geometry foundation models have demonstrated impressive generalization by recovering depth and poses in a single forward pass. Howev

researcharxiv-cs-cv
27 May 2026
Research

RadarSim: Simulating Single-Chip Radar via Multimodal Neural Fields

DGX agent

arXiv:2605.26328v1 Announce Type: new Abstract: Radars are an ideal complement to cameras: both are inexpensive, solid-state sensors, with cameras offering fine angular resolution, while radars provid

researcharxiv-cs-cv
27 May 2026
Research

Re-M3Dr: Rebalanced MultiModal Mean Deviation Regression

DGX agent

arXiv:2605.26513v1 Announce Type: new Abstract: Mean Deviation (MD) is a critical metric for assessing visual field loss in ophthalmology. While previous work has focused solely on predicting MD from

researcharxiv-cs-cv
27 May 2026
Model Releases

Receipt Replay OOD: A Small Benchmark for Screen Replay Detection Under Domain Shift

DGX agent

arXiv:2605.26855v1 Announce Type: new Abstract: Public datasets such as DLC-2021, SynID, and KID34K have significantly contributed to research on presentation attack detection for identity documents,

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

Resolving Ambiguity in Composed Image Retrieval via Calibrated Interaction

DGX agent

arXiv:2605.24634v2 Announce Type: replace Abstract: Composed image retrieval (CIR) searches a corpus with a reference image and a text describing how to modify it. Despite rapid progress from triplet-

model-releasesarxiv-cs-cv
27 May 2026
Research

Revealing the core dimensions underlying representations in brains, behavior and AI

DGX agent

arXiv:2605.26921v1 Announce Type: new Abstract: The study of representations is widespread across fields, including neuroscience, psychology, and artificial intelligence. While representations are oft

researcharxiv-cs-cv
27 May 2026
Local Ai

REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization

DGX agent

arXiv:2605.26861v1 Announce Type: new Abstract: Image geo-localization aims to determine where a photograph was taken, a task that often requires more than recognizing visible landmarks. Human experts

local-aiarxiv-cs-cv
27 May 2026
Model Releases

RoadGIE: Towards A Global-Scale Aerial Benchmark for Generalizable Interactive Road Extraction

DGX agent

arXiv:2605.26862v1 Announce Type: new Abstract: Accurate road segmentation from aerial imagery is fundamental to many geospatial applications. However, existing datasets often suffer from limited scen

model-releasesarxiv-cs-cv
27 May 2026
Research

RoMo: A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation

DGX agent

arXiv:2605.26241v1 Announce Type: new Abstract: Success in generative modeling across language, image, and video demonstrates that large, well-curated datasets are the key driver for building capable

researcharxiv-cs-cv
27 May 2026
Model Releases

Scheduled Style Injection: Expanding the Style-Content Pareto Frontier in Training-Free Diffusion-based Style Transfer

DGX agent

arXiv:2605.26538v1 Announce Type: new Abstract: Style transfer with pre-trained diffusion models has advanced rapidly, but a core question remains underexplored: where in the model should style inject

model-releasesarxiv-cs-cv
27 May 2026
Safety

SCKAN: Structural Consensus-based KAN Prototype Learning for Semi-Supervised Pancreas Segmentation

DGX agent

arXiv:2605.27032v1 Announce Type: new Abstract: Accurate pancreas segmentation is critical for early cancer diagnosis, where annotation scarcity necessitates Semi-Supervised Learning (SSL). However, d

safetyarxiv-cs-cv
27 May 2026
Research

Self-Intersection-Aware 3D Human Motion Generation Using an Efficient Human Sphere Proxy

DGX agent

arXiv:2605.26744v1 Announce Type: new Abstract: Human motion generation has made tremendous progress in recent years, with state-of-the-art approaches surpassing ground truth data in leading evaluatio

researcharxiv-cs-cv
27 May 2026
Tutorials

Semi-Supervised Gaze Estimation via Disentangled Subspace Contrastive Learning

DGX agent

arXiv:2605.27080v1 Announce Type: new Abstract: Appearance-based gaze estimation always suffers from poor generalization due to limited annotated samples and insufficient dataset diversity. Leading ap

tutorialsarxiv-cs-cv
27 May 2026
Model Releases

Sentinel: Embodied Cooperative Spatial Reasoning and Planning

DGX agent

arXiv:2605.26239v1 Announce Type: new Abstract: In this work, we study Cooperative Spatial Intelligence, the ability of decentralized embodied agents to coordinate effectively under dynamic environmen

model-releasesarxiv-cs-cv
27 May 2026
Tutorials

SIMPC: Learning Self-Induced Mirror-Point Consistency for Unsupervised Point Cloud Denoising

DGX agent

arXiv:2605.26894v1 Announce Type: new Abstract: In point clouds, noise directly perturbs point coordinates that encode both spatial location and geometry, making one-to-one correspondence construction

tutorialsarxiv-cs-cv
27 May 2026
Safety

SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing

DGX agent

arXiv:2512.14140v2 Announce Type: replace Abstract: Sketch editing requires jointly handling high-level semantic changes and precise local redrawing, a combination that is particularly challenging for

safetyarxiv-cs-cv
27 May 2026
Research

Sleep-stage efficient classification using a lightweight self-supervised model

DGX agent

arXiv:2605.26295v1 Announce Type: new Abstract: Accurate classification of sleep stages is crucial for diagnosing sleep disorders and automating this process can significantly enhance clinical assessm

researcharxiv-cs-cv
27 May 2026
Research

Small Object Detection in Industrial Recycling: A New Dataset and YOLO Performance Evaluation

DGX agent

arXiv:2605.26884v1 Announce Type: new Abstract: In this paper, we address the problem of detecting small, dense, and overlapping objects, a major challenge in computer vision. Our focus is on reviewin

researcharxiv-cs-cv
27 May 2026
Research

SoftCap: Soft-Budget Control for Diffusion Transformer Acceleration

DGX agent

arXiv:2605.27075v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) achieve strong visual quality, but their iterative denoising process requires many costly Transformer evaluations. Trainin

researcharxiv-cs-cv
27 May 2026
Local Ai

Source-Free Domain Adaptation for Geospatial Point Cloud Semantic Segmentation

DGX agent

arXiv:2601.08375v2 Announce Type: replace Abstract: Semantic segmentation of 3D geospatial point clouds is fundamental to remote sensing applications, yet domain shifts caused by regional and acquisit

local-aiarxiv-cs-cv
27 May 2026
Model Releases

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

DGX agent

arXiv:2510.09606v2 Announce Type: replace Abstract: With the current surge in spatial reasoning explorations, researchers have made significant progress in understanding indoor scenes, but still strug

model-releasesarxiv-cs-cv
27 May 2026
Research

Sparse-LiDAR Prompting of Monocular Geometry Foundations: An Empirical Study Toward Long-Range Driving Depth

DGX agent

arXiv:2605.26456v1 Announce Type: new Abstract: Sparse-LiDAR-prompted depth foundation models (PromptDA, Prior Depth Anything, DMD3C) have shown strong results on indoor scenes or within KITTI's stand

researcharxiv-cs-cv
27 May 2026
Model Releases

SpatialBench: Is Your Spatial Foundation Model an All-Round Player?

DGX agent

arXiv:2605.27367v1 Announce Type: new Abstract: While spatial foundation models have demonstrated impressive performance on standard datasets, a critical question remains: are they truly all-round pla

model-releasesarxiv-cs-cv
27 May 2026
← Previous
1…142143144145146…263
Next →