AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
5 May 2026

Flux4D: Flow-based Unsupervised 4D Reconstruction

AgentsDGX agent

arXiv:2512.03210v2 Announce Type: replace Abstract: Reconstructing large-scale dynamic scenes from visual observations is a fundamental challenge in computer vision, with critical implications for rob

FoR-Net: Learning to Focus on Hard Regions for Efficient Semantic Segmentation

Model ReleasesDGX agent

arXiv:2605.02764v1 Announce Type: new Abstract: We present FoR-Net, a lightweight architecture for semantic segmentation that focuses on identifying and enhancing hard regions. Instead of relying on h

FractalMamba++: Scaling Vision Mamba Across Resolutions via Hilbert Fractal Geometr

Local AiDGX agent

arXiv:2505.14062v3 Announce Type: replace Abstract: Vision Mamba offers linear complexity for long visual sequences, yet its performance depends critically on how a two-dimensional patch grid is seria


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

From Concept to Capability: Evaluating 3D Gaussian Splatting for Synthetic Scene Editing in Autonomous Driving

SafetyDGX agent

arXiv:2605.01995v1 Announce Type: new Abstract: The perception of an Autonomous Driving System (ADS) critically depends on relevant, comprehensive, and diverse datasets to ensure its safety while oper

From Spherical to Gaussian: A Comparative Analysis of Point Cloud Cropping Strategies in Large-Scale 3D Environments

ResearchDGX agent

arXiv:2605.02098v1 Announce Type: new Abstract: Large-scale 3D point clouds can consist of billions of points. Even after downsampling, these point clouds are too large for modern 3D neural networks.

From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs

Model ReleasesDGX agent

arXiv:2605.02130v1 Announce Type: new Abstract: Human-level agentic intelligence extends beyond low-level geometric perception, evolving from recognizing where things are to understanding what they ar

GameScope: A Multi-Attribute, Multi-Codec Benchmark Dataset for Gaming Video Quality Assessment

Model ReleasesDGX agent

arXiv:2605.01272v1 Announce Type: new Abstract: The development of video game streaming has grown rapidly, with major platforms such as YouTube and Twitch using different codecs. To support quality as

Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs

Model ReleasesDGX agent

arXiv:2601.22709v3 Announce Type: replace Abstract: Vision-Language Models (VLMs) achieve strong multimodal performance but are costly to deploy, and post-training quantization often causes significan

GAZE: Grounded Agentic Zero-shot Evaluation with Viewer-Level Tools and Literature Retrieval on Rare Brain MRI

Model ReleasesDGX agent

arXiv:2605.00876v1 Announce Type: cross Abstract: Vision-language models (VLMs) read an image and produce text in a single forward pass, whereas radiologists typically inspect an image several times a

GD-FPS: Growth-Driven Feedforward Parameter Selection for Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2510.27359v2 Announce Type: replace Abstract: Parameter-Efficient Fine-Tuning (PEFT) has emerged as a key strategy for adapting large-scale pre-trained models to downstream tasks, but existing a

GEASS: Training-Free Caption Steering for Hallucination Mitigation in Vision-Language Models

ResearchDGX agent

arXiv:2605.01733v1 Announce Type: new Abstract: Vision-Language Models (VLMs) excel at grounded reasoning but remain prone to object hallucination. Recent work treats self-generated captions as a unif

Gen-Searcher: Reinforcing Agentic Search for Image Generation

Model ReleasesDGX agent

arXiv:2603.28767v2 Announce Type: replace Abstract: Recent image generation models have shown strong capabilities in generating high-fidelity and photorealistic images. However, they are fundamentally

Generalized Category Discovery under Domain Shifts: From Vision to Vision-Language Models

ApplicationsDGX agent

arXiv:2605.00906v1 Announce Type: new Abstract: Generalized Category Discovery (GCD) aims to categorize unlabelled instances from both known and unknown classes by transferring knowledge from labelled

Generative Modeling with Orbit-Space Particle Flow Matching

TutorialsDGX agent

arXiv:2605.02222v1 Announce Type: cross Abstract: We present Orbit-Space Geometric Probability Paths (OGPP), a particle-native flow-matching framework for generative modeling of particle systems. OGPP

GEODE: Angle-Adaptive OOD Detection with Universal Scorer Compatibility

ResearchDGX agent

arXiv:2605.01063v1 Announce Type: cross Abstract: Outlier Exposure (OE) is among the strongest training-based OOD detectors on standard benchmarks but exhibits scorer-dependent tradeoffs (e.g., strong

Geometry-Aware Scene Configurations for Novel View Synthesis

TutorialsDGX agent

arXiv:2510.09880v2 Announce Type: replace Abstract: We propose scene-adaptive strategies to efficiently allocate representation capacity for generating immersive experiences of indoor environments fro

GeoSAE: Geometric Prior-Guided Layer-Wise Sparse Autoencoder Annotation of Brain MRI Foundation Models

TutorialsDGX agent

arXiv:2605.01829v1 Announce Type: new Abstract: Brain MRI foundation models learn rich representations of anatomy, but interpreting what clinical information they encode remains an open problem. Stand

Global-Local Feature Decoding with Adapter-Guided SAMv2 for Salient Object Detection

ResearchDGX agent

arXiv:2605.02616v1 Announce Type: new Abstract: Salient Object Detection (SOD) remains an essential yet underexplored task in the era of large-scale vision models. Although foundation models like SAM

Graph-Augmented Topological Internalization with Dual-Stream Classifiers for Medical Report Generation

ResearchDGX agent

arXiv:2605.02376v1 Announce Type: new Abstract: Automated medical report generation, MRG, holds substantial value for alleviating radiologist workload and enhancing diagnostic efficiency. However, mai

Grounding Synthetic Data Generation With Vision and Language Models

Model ReleasesDGX agent

arXiv:2603.09625v2 Announce Type: replace Abstract: Deep learning models benefit from increasing data diversity and volume, motivating synthetic data augmentation to improve existing datasets. However

GSDeformer: Direct, Real-time and Extensible Cage-based Deformation for 3D Gaussian Splatting

ResearchDGX agent

arXiv:2405.15491v4 Announce Type: replace Abstract: We present GSDeformer, a method that enables cage-based deformation on 3D Gaussian Splatting (3DGS). Our approach bridges cage-based deformation and

Hazard-Aware Traffic Scene Graph Generation

SafetyDGX agent

arXiv:2603.03584v2 Announce Type: replace Abstract: Maintaining situational awareness in complex driving scenarios is challenging. It requires continuously prioritizing attention among extensive scene

Heterogeneous Model Fusion for Privacy-Aware Multi-Camera Surveillance via Synthetic Domain Adaptation

Local AiDGX agent

arXiv:2605.02169v1 Announce Type: new Abstract: We propose HeroCrystal, a novel privacy-preserving framework for multi-camera domain-adaptive object detection, addressing challenges such as data priva

HiFi-Mamba: Dual-Stream W-Laplacian Enhanced Mamba for High-Fidelity MRI Reconstruction

ResearchDGX agent

arXiv:2508.09179v3 Announce Type: replace-cross Abstract: Reconstructing high-fidelity MR images from undersampled k-space data remains a challenging problem in MRI. While Mamba variants for vision ta

High-Fidelity Mobile Avatars with Pruned Local Blendshapes

Local AiDGX agent

arXiv:2605.01854v1 Announce Type: new Abstract: We propose a method to reconstruct high-fidelity human avatars from multi-view video that can run on mobile devices. Many works can model high-quality G

High-Quality Spatial Reconstruction and Orthoimage Generation Using Efficient 2D Gaussian Splatting

ResearchDGX agent

arXiv:2503.19703v3 Announce Type: replace Abstract: Highly accurate geometric precision and dense image features characterize True Digital Orthophoto Maps (TDOMs), which are in great demand for applic

How Can One Choose the Best CAM-Based Explainability Method for a CNN Model?

SafetyDGX agent

arXiv:2605.02007v1 Announce Type: cross Abstract: In recent years, several advances have been observed in Deep Learning with surprising results. Models in this area have been increasingly used in nume

Human Activity Recognition Method for Moderate Violence Detection

ApplicationsDGX agent

arXiv:2605.02659v1 Announce Type: new Abstract: Physical violence in public spaces is a significant public health concern, with minor incidents such as pushing often serving as precursors to more seve

HumanSplatHMR: Closing the Loop Between Human Mesh Recovery and Gaussian Splatting Avatar

SafetyDGX agent

arXiv:2605.02784v1 Announce Type: new Abstract: Accurately recovering human pose and appearance from video is an essential component of scene reconstruction, with applications to motion capture, motio

Hybrid Visual Telemetry for Bandwidth-Constrained Robotic Vision: A Pilot Study with HEVC Base Video and JPEG ROI Stills

Local AiDGX agent

arXiv:2605.01826v1 Announce Type: new Abstract: Bandwidth-constrained robotic and surveillance systems often rely on a single compressed video stream to support both continuous scene awareness and dow

Hyp2Former: Hierarchy-Aware Hyperbolic Embeddings for Open-Set Panoptic Segmentation

SafetyDGX agent

arXiv:2605.02580v1 Announce Type: new Abstract: Recognizing unknown objects is crucial for safety-critical applications such as autonomous driving and robotics. Open-Set Panoptic Segmentation (OPS) ai

IConFace: Identity-Structure Asymmetric Conditioning for Unified Reference-Aware Face Restoration

Local AiDGX agent

arXiv:2605.02814v1 Announce Type: new Abstract: Blind face restoration is highly ill-posed under severe degradation, where identity-critical details may be missing from the degraded input. Same-identi

IMPACT-HOI: Supervisory Control for Onset-Anchored Partial HOI Event Construction

ResearchDGX agent

arXiv:2605.01666v1 Announce Type: new Abstract: We present IMPACT-HOI, a mixed-initiative framework for annotating egocentric procedural video by constructing structured event graphs for Human-Object

IMPACT-Scribe: Interactive Temporal Action Segmentation with Boundary Scribbles and Query Planning

ResearchDGX agent

arXiv:2605.01668v1 Announce Type: new Abstract: Dense temporal annotation of procedural activity videos is vital for action understanding and embodied intelligence but remains labor-intensive due to r

Implicit Neural Representation-Based Continuous Single Image Super-Resolution: An Empirical Benchmark

Model ReleasesDGX agent

arXiv:2601.17723v2 Announce Type: replace Abstract: Implicit neural representation (INR) has become the standard approach for arbitrary-scale image super-resolution (ASSR). To date, no empirical study

Improving Imbalanced Multi-Label Chest X-Ray Diagnosis via CBAM-Enhanced CNN Backbones

ResearchDGX agent

arXiv:2605.02328v1 Announce Type: new Abstract: Chest radiography is a widely used imaging modality for thoracic disease diagnosis, yet its conventional interpretation remains time-consuming and heavi

Improving Model Safety by Targeted Error Correction

SafetyDGX agent

arXiv:2605.02544v1 Announce Type: cross Abstract: The widespread adoption of machine learning in critical applications demands techniques to mitigate high-consequence errors. Our method utilizes a dua

InfiltrNet: Dual-Branch CNN-Transformer Architecture for Brain Tumor Infiltration Risk Prediction

ResearchDGX agent

arXiv:2605.02230v1 Announce Type: new Abstract: Gliomas are aggressive brain tumors that infiltrate surrounding tissue beyond the visible tumor margins observed on Magnetic Resonance Imaging (MRI). Pr

InfiniteDiffusion: Bridging Learned Fidelity and Procedural Utility for Open-World Terrain Generation

HardwareDGX agent

arXiv:2512.08309v4 Announce Type: replace Abstract: For decades, procedural worlds have been built on procedural noise functions such as Perlin noise, which are fast and infinite, yet fundamentally li

InstructMoLE: Instruction-Guided Mixture of Low-rank Experts for Multi-Conditional Image Generation

Model ReleasesDGX agent

arXiv:2512.21788v3 Announce Type: replace Abstract: Parameter-Efficient Fine-Tuning of Diffusion Transformers (DiTs) for diverse, multi-conditional tasks often suffers from task interference when usin

Interactive Multi-Turn Retrieval for Health Videos

Model ReleasesDGX agent

arXiv:2605.01409v1 Announce Type: cross Abstract: The growing availability of health-related instructional videos creates new opportunities for clinical training, patient rehabilitation, and health ed

Interlaced R2D2 DNN Series for Scalable Non-Cartesian MRI with Sensitivity Self-calibration

ResearchDGX agent

arXiv:2503.09559v3 Announce Type: replace-cross Abstract: We introduce interlaced R2D2 (iR2D2), a DNN series paradigm for scalable image reconstruction from accelerated non-Cartesian k-space acquisiti

InterPhys: Physics-aware Human Motion Synthesis in a Dynamic Scene

Model ReleasesDGX agent

arXiv:2605.01036v1 Announce Type: new Abstract: This paper tackles the problem of physics-aware human motion synthesis in a dynamic scene. Unlike existing works which mainly tend to generate physicall

Intervention-Based Self-Supervised Learning: A Causal Probe Paradigm for Remote Photoplethysmography

TutorialsDGX agent

arXiv:2605.00882v1 Announce Type: new Abstract: Remote Photoplethysmography (rPPG) enables convenient non-contact physiological measurement. Existing Self-Supervised Learning (SSL) methods commonly fa

Investigating Anthropometric Fidelity in SAM 3D Body

SafetyDGX agent

arXiv:2601.06035v2 Announce Type: replace-cross Abstract: The release of SAM 3D Body is a recent development in human mesh recovery, demonstrating improved performance in producing clean, topologicall

Joint Architecture-Token-Bitwidth Multi-Axis Optimization of Vision Transformers for Semiconductor IC Packaging

Model ReleasesDGX agent

arXiv:2605.01742v1 Announce Type: new Abstract: Vision Transformers (ViTs) have achieved strong performance in visual recognition, yet their deployment in resource-constrained industrial environments

Know Yourself Better: Diverse Object-Related Features Improve Open Set Recognition

ResearchDGX agent

arXiv:2404.10370v3 Announce Type: replace Abstract: Open set recognition (OSR) is a critical aspect of machine learning, addressing the challenge of detecting novel classes during inference. Within th

LabBuilder: Protocol-Grounded 3D Layout Generation for Interactable and Safe Laboratory

Model ReleasesDGX agent

arXiv:2605.02288v1 Announce Type: new Abstract: Automated laboratories hold the promise of accelerating scientific discovery, yet their deployment is bottlenecked by the difficulty of designing safe a

Laplacian Frequency Interaction Network for Rural Thematic Road Extraction

ApplicationsDGX agent

arXiv:2605.02866v1 Announce Type: new Abstract: Rural thematic road network construction aims to extract topological road structures from movement trajectory images of agricultural machinery. However,

Latent Space Probing for Adult Content Detection in Video Generative Models

ResearchDGX agent

arXiv:2605.00874v1 Announce Type: new Abstract: The rapid proliferation of AI-powered video generation systems has introduced significant challenges in content moderation, particularly with respect to

LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images

Model ReleasesDGX agent

arXiv:2605.00899v1 Announce Type: new Abstract: We present LatentDiff, a scalable framework for semantic dataset comparison that operates directly in the latent space of pretrained vision encoders. By

Learning Equivariant Neural-Augmented Object Dynamics From Few Interactions

TutorialsDGX agent

arXiv:2605.02699v1 Announce Type: cross Abstract: Learning data-efficient object dynamics models for robotic manipulation remains challenging, especially for deformable objects. A popular approach is

Learning to Place Objects with Programs and Iterative Self Training

ResearchDGX agent

arXiv:2503.04496v2 Announce Type: replace-cross Abstract: In this work we study indoor scene object placement. Given a 3D indoor scene and an object, the task is to predict placement locations within

Less Precise Can Be More Reliable: A Systematic Evaluation of Quantization's Impact on VLMs Beyond Accuracy

SafetyDGX agent

arXiv:2509.21173v5 Announce Type: replace Abstract: Vision-Language Models (VLMs) such as CLIP have revolutionized zero-shot classification and safety-critical tasks, including Out-of-Distribution (OO

Leveraging Imperfect Medical Data: A Manifold-Consistent Spatio-Temporal Network for Sensor-based Human Activity Recognition

Model ReleasesDGX agent

arXiv:2605.00913v1 Announce Type: new Abstract: Sensor-based Human Activity Recognition (HAR) has attracted increasing attention in medical and healthcare monitoring, particularly with the growth of I

LGDWT-GS: Local and Global Discrete Wavelet-Regularized 3D Gaussian Splatting for Sparse-View Scene Reconstruction

ResearchDGX agent

arXiv:2601.17185v2 Announce Type: replace Abstract: We propose a new method for few-shot 3D reconstruction that integrates global and local frequency regularization to stabilize geometry and preserve

LIE: LiDAR-only HD Map Construction with Intensity Enhancement via Online Knowledge Distillation

AgentsDGX agent

arXiv:2605.01478v1 Announce Type: new Abstract: Online High-Definition (HD) map construction is a key component of autonomous driving. Recent methods rely on multi-view camera images for cost-effectiv

Limited-Angle Tomography Reconstruction via Projector Guided 3D Diffusion

TutorialsDGX agent

arXiv:2510.06516v2 Announce Type: replace Abstract: Limited-angle electron tomography aims to reconstruct 3D shapes from 2D projections of Transmission Electron Microscopy (TEM) within a restricted ra

Linear-Time Global Visual Modeling without Explicit Attention

Model ReleasesDGX agent

arXiv:2605.01711v1 Announce Type: new Abstract: Existing research largely attributes the global sequence modeling capability of Transformers to the explicit computation of attention weights, a process

Linearizing Vision Transformer with Test-Time Training

Local AiDGX agent

arXiv:2605.02772v1 Announce Type: new Abstract: While linear-complexity attention mechanisms offer a promising alternative to Softmax attention for overcoming the quadratic bottleneck, training such m

← Previous
1…156157158159160…209
Next →