AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlog
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Agents

DynFlowDrive: Flow-Based Dynamic World Modeling for Autonomous Driving

DGX agent

arXiv:2603.19675v2 Announce Type: replace Abstract: Recently, world models have been incorporated into the autonomous driving systems to improve the planning reliability. Existing approaches typically

agentsarxiv-cs-cv
5 May 2026
Local Ai
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

DynoSLAM: Dynamic SLAM with Generative Graph Neural Networks for Real-World Social Navigation

DGX agent

arXiv:2605.02759v1 Announce Type: cross Abstract: Traditional Simultaneous Localization and Mapping (SLAM) algorithms rely heavily on the static environment assumption, which severely limits their app

local-aiarxiv-cs-cv
5 May 2026
Agents

EAPFusion: Intrinsic Evolving Auxiliary Prior Guidance for Infrared and Visible Image Fusion

DGX agent

arXiv:2605.01916v1 Announce Type: new Abstract: Infrared-visible image fusion aims to create an information-rich fused image by integrating the complementary thermal saliency from infrared sensing and

agentsarxiv-cs-cv
5 May 2026
Applications

ECG-biometrics-bench: A Unified Framework for Reproducible Benchmarking of ECG Biometrics

DGX agent

arXiv:2605.01548v1 Announce Type: cross Abstract: Electrocardiogram (ECG) biometrics have emerged as a promising modality for continuous, liveness-aware authentication in wearable systems. However, ma

applicationsarxiv-cs-cv
5 May 2026
Research

Edge-Efficient Image Restoration: Transformer Distillation into State-Space Models

DGX agent

arXiv:2605.02794v1 Announce Type: new Abstract: We propose a modular framework for hybrid image restoration that integrates transformer and state-space model (SSM) blocks with a focus on improving run

researcharxiv-cs-cv
5 May 2026
Model Releases

EdgeLPR: On the Deep Neural Network trade-off between Precision and Performance in LiDAR Place Recognition

DGX agent

arXiv:2605.02275v1 Announce Type: new Abstract: Place recognition is essential for long-term autonomous navigation, enabling loop closure and consistent mapping. Although deep learning has improved pe

model-releasesarxiv-cs-cv
5 May 2026
Research

EduGage: Methods and Dataset for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning

DGX agent

arXiv:2605.01238v1 Announce Type: cross Abstract: Engagement, which links to attentional, emotional, and cognitive dimensions, plays an important role in learning. In online and video-based learning e

researcharxiv-cs-cv
5 May 2026
Research

Embody4D: A Generalist 4D World Model for Embodied AI

DGX agent

arXiv:2605.01799v1 Announce Type: new Abstract: World models have made significant progress in modeling dynamic environments; however, most embodied world models are still restricted to 2D representat

researcharxiv-cs-cv
5 May 2026
Model Releases

EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness

DGX agent

arXiv:2605.01024v1 Announce Type: new Abstract: Multimodal Emotion Recognition (MER) is critical for interpreting real-world interactions. While Multimodal Large Language Models (MLLM) have shown prom

model-releasesarxiv-cs-cv
5 May 2026
Research

Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning

DGX agent

arXiv:2605.02378v1 Announce Type: new Abstract: In-context learning (ICL) allows large models to adapt to tasks using a few examples, yet its extension to vision-language models (VLMs) remains fragile

researcharxiv-cs-cv
5 May 2026
Model Releases

Evolving Token Communication with Parametric Memory Network

DGX agent

arXiv:2605.01869v1 Announce Type: cross Abstract: Token communication has emerged as a promising framework for efficient wireless transmission by representing source data as compact semantic tokens. H

model-releasesarxiv-cs-cv
5 May 2026
Safety

Exploring Data-Free LoRA Transferability for Video Diffusion Models

DGX agent

arXiv:2605.01929v1 Announce Type: new Abstract: Video diffusion models leveraging step distillation or causal distillation have achieved remarkable performance. However, adapting existing LoRAs to the

safetyarxiv-cs-cv
5 May 2026
Safety

Exploring Entropy-based Active Learning for Fair Brain Segmentation

DGX agent

arXiv:2605.01706v1 Announce Type: new Abstract: Active learning (AL) has emerged as a crucial strategy for reducing the prohibitive costs associated with medical image segmentation. However, standard

safetyarxiv-cs-cv
5 May 2026
Safety

Exploring Prompt Alignment with Clinical Factors in Zero-Shot Segmentation VLMs for NSCLC Tumor Segmentation

DGX agent

arXiv:2605.01266v1 Announce Type: new Abstract: Zero-shot vision-language models (VLMs) offer a promptable alternative to task-specific training for gross tumor volume (GTV) delineation in non-small-c

safetyarxiv-cs-cv
5 May 2026
Safety

ExpoCM: Exposure-Aware One-Step Generative Single-Image HDR Reconstruction

DGX agent

arXiv:2605.02464v1 Announce Type: new Abstract: Single-image HDR reconstruction aims to recover high dynamic range radiance from a single low dynamic range (LDR) input, but remains highly ill-posed du

safetyarxiv-cs-cv
5 May 2026
Research

FEAT: Fashion Editing and Try-On from Any Design

DGX agent

arXiv:2605.02393v1 Announce Type: new Abstract: Fashion design aims to express a designer's creative intent and to depict how garments interact with the human body. Recent methods condition on multimo

researcharxiv-cs-cv
5 May 2026
Safety

Fine-Grained Class-Conditional Distribution Balancing for Debiased Learning

DGX agent

arXiv:2505.06831v2 Announce Type: replace Abstract: Achieving group-robust generalization in the presence of spurious correlations remains a significant challenge, particularly when bias annotations a

safetyarxiv-cs-cv
5 May 2026
Model Releases

Fine-Tuning Impairs the Balancedness of Foundation Models in Long-tailed Personalized Federated Learning

DGX agent

arXiv:2605.02247v1 Announce Type: new Abstract: Personalized federated learning (PFL) with foundation models has emerged as a promising paradigm enabling clients to adapt to heterogeneous data distrib

model-releasesarxiv-cs-cv
5 May 2026
Safety

FLoRA: Fusion-Latent for Optical Reconstruction and Flood Area Segmentation via Cross-Modal Multi-Task Distillation Network

DGX agent

arXiv:2605.02137v1 Announce Type: new Abstract: Accurate flood water mapping is critical for disaster management, yet current methods struggle to fully exploit the potential of spaceborne imagery. Opt

safetyarxiv-cs-cv
5 May 2026
Agents

Flux4D: Flow-based Unsupervised 4D Reconstruction

DGX agent

arXiv:2512.03210v2 Announce Type: replace Abstract: Reconstructing large-scale dynamic scenes from visual observations is a fundamental challenge in computer vision, with critical implications for rob

agentsarxiv-cs-cv
5 May 2026
Model Releases

FoR-Net: Learning to Focus on Hard Regions for Efficient Semantic Segmentation

DGX agent

arXiv:2605.02764v1 Announce Type: new Abstract: We present FoR-Net, a lightweight architecture for semantic segmentation that focuses on identifying and enhancing hard regions. Instead of relying on h

model-releasesarxiv-cs-cv
5 May 2026
Local Ai

FractalMamba++: Scaling Vision Mamba Across Resolutions via Hilbert Fractal Geometr

DGX agent

arXiv:2505.14062v3 Announce Type: replace Abstract: Vision Mamba offers linear complexity for long visual sequences, yet its performance depends critically on how a two-dimensional patch grid is seria

local-aiarxiv-cs-cv
5 May 2026
Safety

From Concept to Capability: Evaluating 3D Gaussian Splatting for Synthetic Scene Editing in Autonomous Driving

DGX agent

arXiv:2605.01995v1 Announce Type: new Abstract: The perception of an Autonomous Driving System (ADS) critically depends on relevant, comprehensive, and diverse datasets to ensure its safety while oper

safetyarxiv-cs-cv
5 May 2026
Research

From Spherical to Gaussian: A Comparative Analysis of Point Cloud Cropping Strategies in Large-Scale 3D Environments

DGX agent

arXiv:2605.02098v1 Announce Type: new Abstract: Large-scale 3D point clouds can consist of billions of points. Even after downsampling, these point clouds are too large for modern 3D neural networks.

researcharxiv-cs-cv
5 May 2026
Model Releases

From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs

DGX agent

arXiv:2605.02130v1 Announce Type: new Abstract: Human-level agentic intelligence extends beyond low-level geometric perception, evolving from recognizing where things are to understanding what they ar

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

GameScope: A Multi-Attribute, Multi-Codec Benchmark Dataset for Gaming Video Quality Assessment

DGX agent

arXiv:2605.01272v1 Announce Type: new Abstract: The development of video game streaming has grown rapidly, with major platforms such as YouTube and Twitch using different codecs. To support quality as

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs

DGX agent

arXiv:2601.22709v3 Announce Type: replace Abstract: Vision-Language Models (VLMs) achieve strong multimodal performance but are costly to deploy, and post-training quantization often causes significan

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

GAZE: Grounded Agentic Zero-shot Evaluation with Viewer-Level Tools and Literature Retrieval on Rare Brain MRI

DGX agent

arXiv:2605.00876v1 Announce Type: cross Abstract: Vision-language models (VLMs) read an image and produce text in a single forward pass, whereas radiologists typically inspect an image several times a

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

GD-FPS: Growth-Driven Feedforward Parameter Selection for Efficient Fine-Tuning

DGX agent

arXiv:2510.27359v2 Announce Type: replace Abstract: Parameter-Efficient Fine-Tuning (PEFT) has emerged as a key strategy for adapting large-scale pre-trained models to downstream tasks, but existing a

model-releasesarxiv-cs-cv
5 May 2026
Research

GEASS: Training-Free Caption Steering for Hallucination Mitigation in Vision-Language Models

DGX agent

arXiv:2605.01733v1 Announce Type: new Abstract: Vision-Language Models (VLMs) excel at grounded reasoning but remain prone to object hallucination. Recent work treats self-generated captions as a unif

researcharxiv-cs-cv
5 May 2026
Model Releases

Gen-Searcher: Reinforcing Agentic Search for Image Generation

DGX agent

arXiv:2603.28767v2 Announce Type: replace Abstract: Recent image generation models have shown strong capabilities in generating high-fidelity and photorealistic images. However, they are fundamentally

model-releasesarxiv-cs-cv
5 May 2026
Applications

Generalized Category Discovery under Domain Shifts: From Vision to Vision-Language Models

DGX agent

arXiv:2605.00906v1 Announce Type: new Abstract: Generalized Category Discovery (GCD) aims to categorize unlabelled instances from both known and unknown classes by transferring knowledge from labelled

applicationsarxiv-cs-cv
5 May 2026
Tutorials

Generative Modeling with Orbit-Space Particle Flow Matching

DGX agent

arXiv:2605.02222v1 Announce Type: cross Abstract: We present Orbit-Space Geometric Probability Paths (OGPP), a particle-native flow-matching framework for generative modeling of particle systems. OGPP

tutorialsarxiv-cs-cv
5 May 2026
Research

GEODE: Angle-Adaptive OOD Detection with Universal Scorer Compatibility

DGX agent

arXiv:2605.01063v1 Announce Type: cross Abstract: Outlier Exposure (OE) is among the strongest training-based OOD detectors on standard benchmarks but exhibits scorer-dependent tradeoffs (e.g., strong

researcharxiv-cs-cv
5 May 2026
Tutorials

Geometry-Aware Scene Configurations for Novel View Synthesis

DGX agent

arXiv:2510.09880v2 Announce Type: replace Abstract: We propose scene-adaptive strategies to efficiently allocate representation capacity for generating immersive experiences of indoor environments fro

tutorialsarxiv-cs-cv
5 May 2026
Tutorials

GeoSAE: Geometric Prior-Guided Layer-Wise Sparse Autoencoder Annotation of Brain MRI Foundation Models

DGX agent

arXiv:2605.01829v1 Announce Type: new Abstract: Brain MRI foundation models learn rich representations of anatomy, but interpreting what clinical information they encode remains an open problem. Stand

tutorialsarxiv-cs-cv
5 May 2026
Research

Global-Local Feature Decoding with Adapter-Guided SAMv2 for Salient Object Detection

DGX agent

arXiv:2605.02616v1 Announce Type: new Abstract: Salient Object Detection (SOD) remains an essential yet underexplored task in the era of large-scale vision models. Although foundation models like SAM

researcharxiv-cs-cv
5 May 2026
Research

Graph-Augmented Topological Internalization with Dual-Stream Classifiers for Medical Report Generation

DGX agent

arXiv:2605.02376v1 Announce Type: new Abstract: Automated medical report generation, MRG, holds substantial value for alleviating radiologist workload and enhancing diagnostic efficiency. However, mai

researcharxiv-cs-cv
5 May 2026
Model Releases

Grounding Synthetic Data Generation With Vision and Language Models

DGX agent

arXiv:2603.09625v2 Announce Type: replace Abstract: Deep learning models benefit from increasing data diversity and volume, motivating synthetic data augmentation to improve existing datasets. However

model-releasesarxiv-cs-cv
5 May 2026
Research

GSDeformer: Direct, Real-time and Extensible Cage-based Deformation for 3D Gaussian Splatting

DGX agent

arXiv:2405.15491v4 Announce Type: replace Abstract: We present GSDeformer, a method that enables cage-based deformation on 3D Gaussian Splatting (3DGS). Our approach bridges cage-based deformation and

researcharxiv-cs-cv
5 May 2026
Safety

Hazard-Aware Traffic Scene Graph Generation

DGX agent

arXiv:2603.03584v2 Announce Type: replace Abstract: Maintaining situational awareness in complex driving scenarios is challenging. It requires continuously prioritizing attention among extensive scene

safetyarxiv-cs-cv
5 May 2026
Local Ai

Heterogeneous Model Fusion for Privacy-Aware Multi-Camera Surveillance via Synthetic Domain Adaptation

DGX agent

arXiv:2605.02169v1 Announce Type: new Abstract: We propose HeroCrystal, a novel privacy-preserving framework for multi-camera domain-adaptive object detection, addressing challenges such as data priva

local-aiarxiv-cs-cv
5 May 2026
Research

HiFi-Mamba: Dual-Stream W-Laplacian Enhanced Mamba for High-Fidelity MRI Reconstruction

DGX agent

arXiv:2508.09179v3 Announce Type: replace-cross Abstract: Reconstructing high-fidelity MR images from undersampled k-space data remains a challenging problem in MRI. While Mamba variants for vision ta

researcharxiv-cs-cv
5 May 2026
Local Ai

High-Fidelity Mobile Avatars with Pruned Local Blendshapes

DGX agent

arXiv:2605.01854v1 Announce Type: new Abstract: We propose a method to reconstruct high-fidelity human avatars from multi-view video that can run on mobile devices. Many works can model high-quality G

local-aiarxiv-cs-cv
5 May 2026
Research

High-Quality Spatial Reconstruction and Orthoimage Generation Using Efficient 2D Gaussian Splatting

DGX agent

arXiv:2503.19703v3 Announce Type: replace Abstract: Highly accurate geometric precision and dense image features characterize True Digital Orthophoto Maps (TDOMs), which are in great demand for applic

researcharxiv-cs-cv
5 May 2026
Safety

How Can One Choose the Best CAM-Based Explainability Method for a CNN Model?

DGX agent

arXiv:2605.02007v1 Announce Type: cross Abstract: In recent years, several advances have been observed in Deep Learning with surprising results. Models in this area have been increasingly used in nume

safetyarxiv-cs-cv
5 May 2026
Applications

Human Activity Recognition Method for Moderate Violence Detection

DGX agent

arXiv:2605.02659v1 Announce Type: new Abstract: Physical violence in public spaces is a significant public health concern, with minor incidents such as pushing often serving as precursors to more seve

applicationsarxiv-cs-cv
5 May 2026
Safety

HumanSplatHMR: Closing the Loop Between Human Mesh Recovery and Gaussian Splatting Avatar

DGX agent

arXiv:2605.02784v1 Announce Type: new Abstract: Accurately recovering human pose and appearance from video is an essential component of scene reconstruction, with applications to motion capture, motio

safetyarxiv-cs-cv
5 May 2026
← Previous
1…197198199200201…263
Next →