AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Safety

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

DGX agent

arXiv:2604.07607v1 Announce Type: cross Abstract: Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human da

safetyarxiv-cs-cv
10 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Tutorials

EPIR: An Efficient Patch Tokenization, Integration and Representation Framework for Micro-expression Recognition

DGX agent

arXiv:2604.08106v1 Announce Type: new Abstract: Micro-expression recognition can obtain the real emotion of the individual at the current moment. Although deep learning-based methods, especially Trans

tutorialsarxiv-cs-cv
10 Apr 2026
Model Releases

ESOM: Efficiently Understanding Streaming Video Anomalies with Open-world Dynamic Definitions

DGX agent

arXiv:2604.07772v1 Announce Type: new Abstract: Open-world video anomaly detection (OWVAD) aims to detect and explain abnormal events under different anomaly definitions, which is important for applic

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

ETCH-X: Robustify Expressive Body Fitting to Clothed Humans with Composable Datasets

DGX agent

arXiv:2604.08548v1 Announce Type: new Abstract: Human body fitting, which aligns parametric body models such as SMPL to raw 3D point clouds of clothed humans, serves as a crucial first step for downst

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

Evaluating Low-Light Image Enhancement Across Multiple Intensity Levels

DGX agent

arXiv:2511.15496v2 Announce Type: replace Abstract: Imaging in low-light environments is challenging due to reduced scene radiance, which leads to elevated sensor noise and reduced color saturation. M

model-releasesarxiv-cs-cv
10 Apr 2026
Safety

Event-Level Detection of Surgical Instrument Handovers in Videos with Interpretable Vision Models

DGX agent

arXiv:2604.07577v1 Announce Type: new Abstract: Reliable monitoring of surgical instrument exchanges is essential for maintaining procedural efficiency and patient safety in the operating room. Automa

safetyarxiv-cs-cv
10 Apr 2026
Model Releases

Face-D(^2)CL: Multi-Domain Synergistic Representation with Dual Continual Learning for Facial DeepFake Detection

DGX agent

arXiv:2604.08159v1 Announce Type: new Abstract: The rapid advancement of facial forgery techniques poses severe threats to public trust and information security, making facial DeepFake detection a cri

model-releasesarxiv-cs-cv
10 Apr 2026
Tutorials

Face2Scene: Using Facial Degradation as an Oracle for Diffusion-Based Scene Restoration

DGX agent

arXiv:2603.16570v2 Announce Type: replace Abstract: Recent advances in image restoration have enabled high-fidelity recovery of faces from degraded inputs using reference-based face restoration models

tutorialsarxiv-cs-cv
10 Apr 2026
Model Releases

Fail2Drive: Benchmarking Closed-Loop Driving Generalization

DGX agent

arXiv:2604.08535v1 Announce Type: cross Abstract: Generalization under distribution shift remains a central bottleneck for closed-loop autonomous driving. Although simulators like CARLA enable safe an

model-releasesarxiv-cs-cv
10 Apr 2026
Safety

Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization

DGX agent

arXiv:2604.08476v1 Announce Type: new Abstract: Multimodal reasoning models (MRMs) trained with reinforcement learning with verifiable rewards (RLVR) show improved accuracy on visual reasoning benchma

safetyarxiv-cs-cv
10 Apr 2026
Tutorials

Few-Shot Incremental 3D Object Detection in Dynamic Indoor Environments

DGX agent

arXiv:2604.07997v1 Announce Type: new Abstract: Incremental 3D object perception is a critical step toward embodied intelligence in dynamic indoor environments. However, existing incremental 3D detect

tutorialsarxiv-cs-cv
10 Apr 2026
Model Releases

FireSenseNet: A Dual-Branch CNN with Cross-Attentive Feature Interaction for Next-Day Wildfire Spread Prediction

DGX agent

arXiv:2604.07675v1 Announce Type: new Abstract: Accurate prediction of next-day wildfire spread is critical for disaster response and resource allocation. Existing deep learning approaches typically c

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On

DGX agent

arXiv:2604.08526v1 Announce Type: new Abstract: Given a person and a garment image, virtual try-on (VTO) aims to synthesize a realistic image of the person wearing the garment, while preserving their

model-releasesarxiv-cs-cv
10 Apr 2026
Research

Flemme: A Flexible and Modular Learning Platform for Medical Images

DGX agent

arXiv:2408.09369v3 Announce Type: replace-cross Abstract: As the rapid development of computer vision and the emergence of powerful network backbones and architectures, the application of deep learnin

researcharxiv-cs-cv
10 Apr 2026
Model Releases

FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding

DGX agent

arXiv:2604.07879v1 Announce Type: new Abstract: Diffusion-based image generation models have advanced rapidly but pose a safety risk due to their potential to generate Not-Safe-For-Work (NSFW) content

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

FORGE:Fine-grained Multimodal Evaluation for Manufacturing Scenarios

DGX agent

arXiv:2604.07413v1 Announce Type: new Abstract: The manufacturing sector is increasingly adopting Multimodal Large Language Models (MLLMs) to transition from simple perception to autonomous execution,

model-releasesarxiv-cs-cv
10 Apr 2026
Research

From Classical Machine Learning to Tabular Foundation Models: An Empirical Investigation of Robustness and Scalability Under Class Imbalance in Emergency and Critical Care

DGX agent

arXiv:2512.21602v2 Announce Type: replace-cross Abstract: Millions of patients pass through emergency departments and intensive care units each year, where clinicians must make high-stakes decisions u

researcharxiv-cs-cv
10 Apr 2026
Research

Fundus-R1: Training a Fundus-Reading MLLM with Knowledge-Aware Reasoning on Public Data

DGX agent

arXiv:2604.08322v1 Announce Type: new Abstract: Fundus imaging such as CFP, OCT and UWF is crucial for the early detection of retinal anomalies and diseases. Fundus image understanding, due to its kno

researcharxiv-cs-cv
10 Apr 2026
Model Releases

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents

DGX agent

arXiv:2604.07429v1 Announce Type: new Abstract: Towards an embodied generalist for real-world interaction, Multimodal Large Language Model (MLLM) agents still suffer from challenging latency, sparse f

model-releasesarxiv-cs-cv
10 Apr 2026
Applications

GaussiAnimate: Reconstruct and Rig Animatable Categories with Level of Dynamics

DGX agent

arXiv:2604.08547v1 Announce Type: new Abstract: Free-form bones, that conform closely to the surface, can effectively capture non-rigid deformations, but lack a kinematic structure necessary for intui

applicationsarxiv-cs-cv
10 Apr 2026
Applications

Gaze to Insight: A Scalable AI Approach for Detecting Gaze Behaviours in Face-to-Face Collaborative Learning

DGX agent

arXiv:2604.03317v2 Announce Type: replace Abstract: Previous studies have illustrated the potential of analysing gaze behaviours in collaborative learning to provide educationally meaningful informati

applicationsarxiv-cs-cv
10 Apr 2026
Research

GEAR: GEometry-motion Alternating Refinement for Articulated Object Modeling with Gaussian Splatting

DGX agent

arXiv:2604.07728v1 Announce Type: new Abstract: High-fidelity interactive digital assets are essential for embodied intelligence and robotic interaction, yet articulated objects remain challenging to

researcharxiv-cs-cv
10 Apr 2026
Tutorials

Generalization Under Scrutiny: Cross-Domain Detection Progresses, Pitfalls, and Persistent Challenges

DGX agent

arXiv:2604.08230v1 Announce Type: new Abstract: Object detection models trained on a source domain often exhibit significant performance degradation when deployed in unseen target domains, due to vari

tutorialsarxiv-cs-cv
10 Apr 2026
Research

Generative 3D Gaussian Splatting for Arbitrary-ResolutionAtmospheric Downscaling and Forecasting

DGX agent

arXiv:2604.07928v1 Announce Type: new Abstract: While AI-based numerical weather prediction (NWP) enables rapid forecasting, generating high-resolution outputs remains computationally demanding due to

researcharxiv-cs-cv
10 Apr 2026
Applications

GenLCA: 3D Diffusion for Full-Body Avatars from In-the-Wild Videos

DGX agent

arXiv:2604.07273v2 Announce Type: replace Abstract: We present GenLCA, a diffusion-based generative model for generating and editing photorealistic full-body avatars from text and image inputs. The ge

applicationsarxiv-cs-cv
10 Apr 2026
Research

GroundingAnomaly: Spatially-Grounded Diffusion for Few-Shot Anomaly Synthesis

DGX agent

arXiv:2604.08301v1 Announce Type: new Abstract: The performance of visual anomaly inspection in industrial quality control is often constrained by the scarcity of real anomalous samples. Consequently,

researcharxiv-cs-cv
10 Apr 2026
Safety

Guiding a Diffusion Model by Swapping Its Tokens

DGX agent

arXiv:2604.08048v1 Announce Type: new Abstract: Classifier-Free Guidance (CFG) is a widely used inference-time technique to boost the image quality of diffusion models. Yet, its reliance on text condi

safetyarxiv-cs-cv
10 Apr 2026
Hardware

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models

DGX agent

arXiv:2604.07812v1 Announce Type: new Abstract: In multimodal large language models (MLLMs), the surge of visual tokens significantly increases the inference time and computational overhead, making th

hardwarearxiv-cs-cv
10 Apr 2026
Research

Hierarchical Feature Learning for Medical Point Clouds via State Space Model

DGX agent

arXiv:2504.13015v3 Announce Type: replace Abstract: Deep learning-based point cloud modeling has been widely investigated as an indispensable component of general shape analysis. Recently, transformer

researcharxiv-cs-cv
10 Apr 2026
Model Releases

HistDiT: A Structure-Aware Latent Conditional Diffusion Model for High-Fidelity Virtual Staining in Histopathology

DGX agent

arXiv:2604.08305v1 Announce Type: cross Abstract: Immunohistochemistry (IHC) is essential for assessing specific immune biomarkers like Human Epidermal growth-factor Receptor 2 (HER2) in breast cancer

model-releasesarxiv-cs-cv
10 Apr 2026
Applications

Horticultural Temporal Fruit Monitoring via 3D Instance Segmentation and Re-Identification using Colored Point Clouds

DGX agent

arXiv:2411.07799v3 Announce Type: replace Abstract: Accurate and consistent fruit monitoring over time is a key step toward automated agricultural production systems. However, this task is inherently

applicationsarxiv-cs-cv
10 Apr 2026
Research

HOTFLoc++: End-to-End Hierarchical LiDAR Place Recognition, Re-Ranking, and 6-DoF Metric Localisation in Forests

DGX agent

arXiv:2511.09170v2 Announce Type: replace Abstract: This article presents HOTFLoc++, an end-to-end hierarchical framework for LiDAR place recognition, re-ranking, and 6-DoF metric localisation in fore

researcharxiv-cs-cv
10 Apr 2026
Research

HST-HGN: Heterogeneous Spatial-Temporal Hypergraph Networks with Bidirectional State Space Models for Global Fatigue Assessment

DGX agent

arXiv:2604.08435v1 Announce Type: new Abstract: It remains challenging to assess driver fatigue from untrimmed videos under constrained computational budgets, due to the difficulty of modeling long-ra

researcharxiv-cs-cv
10 Apr 2026
Model Releases

HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents

DGX agent

arXiv:2604.07430v1 Announce Type: new Abstract: We introduce HY-Embodied-0.5, a family of foundation models specifically designed for real-world embodied agents. To bridge the gap between general Visi

model-releasesarxiv-cs-cv
10 Apr 2026
Research

Image-Guided Geometric Stylization of 3D Meshes

DGX agent

arXiv:2604.07795v1 Announce Type: new Abstract: Recent generative models can create visually plausible 3D representations of objects. However, the generation process often allows for implicit control

researcharxiv-cs-cv
10 Apr 2026
Research

Improving Image Coding for Machines through Optimizing Encoder via Auxiliary Loss

DGX agent

arXiv:2402.08267v3 Announce Type: replace Abstract: Image coding for machines (ICM) aims to compress images for machine analysis using recognition models rather than human vision. Hence, in ICM, it is

researcharxiv-cs-cv
10 Apr 2026
Research

ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks

DGX agent

arXiv:2604.07958v1 Announce Type: new Abstract: Current video editing models often rely on expensive paired video data, which limits their practical scalability. In essence, most video editing tasks c

researcharxiv-cs-cv
10 Apr 2026
Safety

Inside-Out: Measuring Generalization in Vision Transformers Through Inner Workings

DGX agent

arXiv:2604.08192v1 Announce Type: cross Abstract: Reliable generalization metrics are fundamental to the evaluation of machine learning models. Especially in high-stakes applications where labeled tar

safetyarxiv-cs-cv
10 Apr 2026
Model Releases

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding

DGX agent

arXiv:2604.08337v1 Announce Type: new Abstract: Current vision-language pre-training (VLP) paradigms excel at global scene understanding but struggle with instance-level reasoning due to global-only s

model-releasesarxiv-cs-cv
10 Apr 2026
Research

Interpretable Tau-PET Synthesis from Multimodal T1-Weighted and FLAIR MRI Using Partial Information Decomposition Guided Disentangled Quantized Half-UNet

DGX agent

arXiv:2602.22545v2 Announce Type: replace Abstract: Tau positron emission tomography (tau-PET) is an important in vivo biomarker of Alzheimer's disease, but its cost, limited availability, and acquisi

researcharxiv-cs-cv
10 Apr 2026
Model Releases

Kuramoto Oscillatory Phase Encoding: Neuro-inspired Synchronization for Improved Learning Efficiency

DGX agent

arXiv:2604.07904v1 Announce Type: cross Abstract: Spatiotemporal neural dynamics and oscillatory synchronization are widely implicated in biological information processing and have been hypothesized t

model-releasesarxiv-cs-cv
10 Apr 2026
Research

LAMP: Lift Image-Editing as General 3D Priors for Open-world Manipulation

DGX agent

arXiv:2604.08475v1 Announce Type: new Abstract: Human-like generalization in open-world remains a fundamental challenge for robotic manipulation. Existing learning-based methods, including reinforceme

researcharxiv-cs-cv
10 Apr 2026
Research

Lang2Act: Fine-Grained Visual Reasoning through Self-Emergent Linguistic Toolchains

DGX agent

arXiv:2602.13235v2 Announce Type: replace-cross Abstract: Visual Retrieval-Augmented Generation (VRAG) enhances Vision-Language Models (VLMs) by incorporating external visual documents to address a gi

researcharxiv-cs-cv
10 Apr 2026
Safety

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents

DGX agent

arXiv:2512.17445v2 Announce Type: replace Abstract: LangDriveCTRL is a natural-language-controllable framework for editing real-world driving videos to synthesize diverse traffic scenarios. It represe

safetyarxiv-cs-cv
10 Apr 2026
Research

Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models

DGX agent

arXiv:2604.07802v1 Announce Type: new Abstract: Large-scale vision-language models (VLMs) exhibit remarkable zero-shot capabilities, yet the internal mechanisms driving their anomaly detection (AD) pe

researcharxiv-cs-cv
10 Apr 2026
Agents

Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering

DGX agent

arXiv:2604.07146v2 Announce Type: replace Abstract: Knowledge-based visual question answering (KB-VQA) requires vision-language models to understand images and use external knowledge, especially for r

agentsarxiv-cs-cv
10 Apr 2026
Agents

Lighting-grounded Video Generation with Renderer-based Agent Reasoning

DGX agent

arXiv:2604.07966v1 Announce Type: new Abstract: Diffusion models have achieved remarkable progress in video generation, but their controllability remains a major limitation. Key scene factors such as

agentsarxiv-cs-cv
10 Apr 2026
Safety

LINE: LLM-based Iterative Neuron Explanations for Vision Models

DGX agent

arXiv:2604.08039v1 Announce Type: new Abstract: Interpreting the concepts encoded by individual neurons in deep neural networks is a crucial step towards understanding their complex decision-making pr

safetyarxiv-cs-cv
10 Apr 2026
← Previous
1…254255256257258259
Next →