AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
10 Apr 2026

EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World

SafetyDGX agent

arXiv:2604.07607v1 Announce Type: cross Abstract: Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human da

EPIR: An Efficient Patch Tokenization, Integration and Representation Framework for Micro-expression Recognition

TutorialsDGX agent

arXiv:2604.08106v1 Announce Type: new Abstract: Micro-expression recognition can obtain the real emotion of the individual at the current moment. Although deep learning-based methods, especially Trans

ESOM: Efficiently Understanding Streaming Video Anomalies with Open-world Dynamic Definitions

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.07772v1 Announce Type: new Abstract: Open-world video anomaly detection (OWVAD) aims to detect and explain abnormal events under different anomaly definitions, which is important for applic

ETCH-X: Robustify Expressive Body Fitting to Clothed Humans with Composable Datasets

Model ReleasesDGX agent

arXiv:2604.08548v1 Announce Type: new Abstract: Human body fitting, which aligns parametric body models such as SMPL to raw 3D point clouds of clothed humans, serves as a crucial first step for downst

Evaluating Low-Light Image Enhancement Across Multiple Intensity Levels

Model ReleasesDGX agent

arXiv:2511.15496v2 Announce Type: replace Abstract: Imaging in low-light environments is challenging due to reduced scene radiance, which leads to elevated sensor noise and reduced color saturation. M

Event-Level Detection of Surgical Instrument Handovers in Videos with Interpretable Vision Models

SafetyDGX agent

arXiv:2604.07577v1 Announce Type: new Abstract: Reliable monitoring of surgical instrument exchanges is essential for maintaining procedural efficiency and patient safety in the operating room. Automa

Face-D(^2)CL: Multi-Domain Synergistic Representation with Dual Continual Learning for Facial DeepFake Detection

Model ReleasesDGX agent

arXiv:2604.08159v1 Announce Type: new Abstract: The rapid advancement of facial forgery techniques poses severe threats to public trust and information security, making facial DeepFake detection a cri

Face2Scene: Using Facial Degradation as an Oracle for Diffusion-Based Scene Restoration

TutorialsDGX agent

arXiv:2603.16570v2 Announce Type: replace Abstract: Recent advances in image restoration have enabled high-fidelity recovery of faces from degraded inputs using reference-based face restoration models

Fail2Drive: Benchmarking Closed-Loop Driving Generalization

Model ReleasesDGX agent

arXiv:2604.08535v1 Announce Type: cross Abstract: Generalization under distribution shift remains a central bottleneck for closed-loop autonomous driving. Although simulators like CARLA enable safe an

Faithful GRPO: Improving Visual Spatial Reasoning in Multimodal Language Models via Constrained Policy Optimization

SafetyDGX agent

arXiv:2604.08476v1 Announce Type: new Abstract: Multimodal reasoning models (MRMs) trained with reinforcement learning with verifiable rewards (RLVR) show improved accuracy on visual reasoning benchma

Few-Shot Incremental 3D Object Detection in Dynamic Indoor Environments

TutorialsDGX agent

arXiv:2604.07997v1 Announce Type: new Abstract: Incremental 3D object perception is a critical step toward embodied intelligence in dynamic indoor environments. However, existing incremental 3D detect

FireSenseNet: A Dual-Branch CNN with Cross-Attentive Feature Interaction for Next-Day Wildfire Spread Prediction

Model ReleasesDGX agent

arXiv:2604.07675v1 Announce Type: new Abstract: Accurate prediction of next-day wildfire spread is critical for disaster response and resource allocation. Existing deep learning approaches typically c

FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On

Model ReleasesDGX agent

arXiv:2604.08526v1 Announce Type: new Abstract: Given a person and a garment image, virtual try-on (VTO) aims to synthesize a realistic image of the person wearing the garment, while preserving their

Flemme: A Flexible and Modular Learning Platform for Medical Images

ResearchDGX agent

arXiv:2408.09369v3 Announce Type: replace-cross Abstract: As the rapid development of computer vision and the emergence of powerful network backbones and architectures, the application of deep learnin

FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding

Model ReleasesDGX agent

arXiv:2604.07879v1 Announce Type: new Abstract: Diffusion-based image generation models have advanced rapidly but pose a safety risk due to their potential to generate Not-Safe-For-Work (NSFW) content

FORGE:Fine-grained Multimodal Evaluation for Manufacturing Scenarios

Model ReleasesDGX agent

arXiv:2604.07413v1 Announce Type: new Abstract: The manufacturing sector is increasingly adopting Multimodal Large Language Models (MLLMs) to transition from simple perception to autonomous execution,

From Classical Machine Learning to Tabular Foundation Models: An Empirical Investigation of Robustness and Scalability Under Class Imbalance in Emergency and Critical Care

ResearchDGX agent

arXiv:2512.21602v2 Announce Type: replace-cross Abstract: Millions of patients pass through emergency departments and intensive care units each year, where clinicians must make high-stakes decisions u

Fundus-R1: Training a Fundus-Reading MLLM with Knowledge-Aware Reasoning on Public Data

ResearchDGX agent

arXiv:2604.08322v1 Announce Type: new Abstract: Fundus imaging such as CFP, OCT and UWF is crucial for the early detection of retinal anomalies and diseases. Fundus image understanding, due to its kno

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents

Model ReleasesDGX agent

arXiv:2604.07429v1 Announce Type: new Abstract: Towards an embodied generalist for real-world interaction, Multimodal Large Language Model (MLLM) agents still suffer from challenging latency, sparse f

GaussiAnimate: Reconstruct and Rig Animatable Categories with Level of Dynamics

ApplicationsDGX agent

arXiv:2604.08547v1 Announce Type: new Abstract: Free-form bones, that conform closely to the surface, can effectively capture non-rigid deformations, but lack a kinematic structure necessary for intui

Gaze to Insight: A Scalable AI Approach for Detecting Gaze Behaviours in Face-to-Face Collaborative Learning

ApplicationsDGX agent

arXiv:2604.03317v2 Announce Type: replace Abstract: Previous studies have illustrated the potential of analysing gaze behaviours in collaborative learning to provide educationally meaningful informati

GEAR: GEometry-motion Alternating Refinement for Articulated Object Modeling with Gaussian Splatting

ResearchDGX agent

arXiv:2604.07728v1 Announce Type: new Abstract: High-fidelity interactive digital assets are essential for embodied intelligence and robotic interaction, yet articulated objects remain challenging to

Generalization Under Scrutiny: Cross-Domain Detection Progresses, Pitfalls, and Persistent Challenges

TutorialsDGX agent

arXiv:2604.08230v1 Announce Type: new Abstract: Object detection models trained on a source domain often exhibit significant performance degradation when deployed in unseen target domains, due to vari

Generative 3D Gaussian Splatting for Arbitrary-ResolutionAtmospheric Downscaling and Forecasting

ResearchDGX agent

arXiv:2604.07928v1 Announce Type: new Abstract: While AI-based numerical weather prediction (NWP) enables rapid forecasting, generating high-resolution outputs remains computationally demanding due to

GenLCA: 3D Diffusion for Full-Body Avatars from In-the-Wild Videos

ApplicationsDGX agent

arXiv:2604.07273v2 Announce Type: replace Abstract: We present GenLCA, a diffusion-based generative model for generating and editing photorealistic full-body avatars from text and image inputs. The ge

GroundingAnomaly: Spatially-Grounded Diffusion for Few-Shot Anomaly Synthesis

ResearchDGX agent

arXiv:2604.08301v1 Announce Type: new Abstract: The performance of visual anomaly inspection in industrial quality control is often constrained by the scarcity of real anomalous samples. Consequently,

Guiding a Diffusion Model by Swapping Its Tokens

SafetyDGX agent

arXiv:2604.08048v1 Announce Type: new Abstract: Classifier-Free Guidance (CFG) is a widely used inference-time technique to boost the image quality of diffusion models. Yet, its reliance on text condi

HAWK: Head Importance-Aware Visual Token Pruning in Multimodal Models

HardwareDGX agent

arXiv:2604.07812v1 Announce Type: new Abstract: In multimodal large language models (MLLMs), the surge of visual tokens significantly increases the inference time and computational overhead, making th

Hierarchical Feature Learning for Medical Point Clouds via State Space Model

ResearchDGX agent

arXiv:2504.13015v3 Announce Type: replace Abstract: Deep learning-based point cloud modeling has been widely investigated as an indispensable component of general shape analysis. Recently, transformer

HistDiT: A Structure-Aware Latent Conditional Diffusion Model for High-Fidelity Virtual Staining in Histopathology

Model ReleasesDGX agent

arXiv:2604.08305v1 Announce Type: cross Abstract: Immunohistochemistry (IHC) is essential for assessing specific immune biomarkers like Human Epidermal growth-factor Receptor 2 (HER2) in breast cancer

Horticultural Temporal Fruit Monitoring via 3D Instance Segmentation and Re-Identification using Colored Point Clouds

ApplicationsDGX agent

arXiv:2411.07799v3 Announce Type: replace Abstract: Accurate and consistent fruit monitoring over time is a key step toward automated agricultural production systems. However, this task is inherently

HOTFLoc++: End-to-End Hierarchical LiDAR Place Recognition, Re-Ranking, and 6-DoF Metric Localisation in Forests

ResearchDGX agent

arXiv:2511.09170v2 Announce Type: replace Abstract: This article presents HOTFLoc++, an end-to-end hierarchical framework for LiDAR place recognition, re-ranking, and 6-DoF metric localisation in fore

HST-HGN: Heterogeneous Spatial-Temporal Hypergraph Networks with Bidirectional State Space Models for Global Fatigue Assessment

ResearchDGX agent

arXiv:2604.08435v1 Announce Type: new Abstract: It remains challenging to assess driver fatigue from untrimmed videos under constrained computational budgets, due to the difficulty of modeling long-ra

HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents

Model ReleasesDGX agent

arXiv:2604.07430v1 Announce Type: new Abstract: We introduce HY-Embodied-0.5, a family of foundation models specifically designed for real-world embodied agents. To bridge the gap between general Visi

Image-Guided Geometric Stylization of 3D Meshes

ResearchDGX agent

arXiv:2604.07795v1 Announce Type: new Abstract: Recent generative models can create visually plausible 3D representations of objects. However, the generation process often allows for implicit control

Improving Image Coding for Machines through Optimizing Encoder via Auxiliary Loss

ResearchDGX agent

arXiv:2402.08267v3 Announce Type: replace Abstract: Image coding for machines (ICM) aims to compress images for machine analysis using recognition models rather than human vision. Hence, in ICM, it is

ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks

ResearchDGX agent

arXiv:2604.07958v1 Announce Type: new Abstract: Current video editing models often rely on expensive paired video data, which limits their practical scalability. In essence, most video editing tasks c

Inside-Out: Measuring Generalization in Vision Transformers Through Inner Workings

SafetyDGX agent

arXiv:2604.08192v1 Announce Type: cross Abstract: Reliable generalization metrics are fundamental to the evaluation of machine learning models. Especially in high-stakes applications where labeled tar

InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding

Model ReleasesDGX agent

arXiv:2604.08337v1 Announce Type: new Abstract: Current vision-language pre-training (VLP) paradigms excel at global scene understanding but struggle with instance-level reasoning due to global-only s

Interpretable Tau-PET Synthesis from Multimodal T1-Weighted and FLAIR MRI Using Partial Information Decomposition Guided Disentangled Quantized Half-UNet

ResearchDGX agent

arXiv:2602.22545v2 Announce Type: replace Abstract: Tau positron emission tomography (tau-PET) is an important in vivo biomarker of Alzheimer's disease, but its cost, limited availability, and acquisi

Kuramoto Oscillatory Phase Encoding: Neuro-inspired Synchronization for Improved Learning Efficiency

Model ReleasesDGX agent

arXiv:2604.07904v1 Announce Type: cross Abstract: Spatiotemporal neural dynamics and oscillatory synchronization are widely implicated in biological information processing and have been hypothesized t

LAMP: Lift Image-Editing as General 3D Priors for Open-world Manipulation

ResearchDGX agent

arXiv:2604.08475v1 Announce Type: new Abstract: Human-like generalization in open-world remains a fundamental challenge for robotic manipulation. Existing learning-based methods, including reinforceme

Lang2Act: Fine-Grained Visual Reasoning through Self-Emergent Linguistic Toolchains

ResearchDGX agent

arXiv:2602.13235v2 Announce Type: replace-cross Abstract: Visual Retrieval-Augmented Generation (VRAG) enhances Vision-Language Models (VLMs) by incorporating external visual documents to address a gi

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents

SafetyDGX agent

arXiv:2512.17445v2 Announce Type: replace Abstract: LangDriveCTRL is a natural-language-controllable framework for editing real-world driving videos to synthesize diverse traffic scenarios. It represe

Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models

ResearchDGX agent

arXiv:2604.07802v1 Announce Type: new Abstract: Large-scale vision-language models (VLMs) exhibit remarkable zero-shot capabilities, yet the internal mechanisms driving their anomaly detection (AD) pe

Learning to Search: A Decision-Based Agent for Knowledge-Based Visual Question Answering

AgentsDGX agent

arXiv:2604.07146v2 Announce Type: replace Abstract: Knowledge-based visual question answering (KB-VQA) requires vision-language models to understand images and use external knowledge, especially for r

Lighting-grounded Video Generation with Renderer-based Agent Reasoning

AgentsDGX agent

arXiv:2604.07966v1 Announce Type: new Abstract: Diffusion models have achieved remarkable progress in video generation, but their controllability remains a major limitation. Key scene factors such as

LINE: LLM-based Iterative Neuron Explanations for Vision Models

SafetyDGX agent

arXiv:2604.08039v1 Announce Type: new Abstract: Interpreting the concepts encoded by individual neurons in deep neural networks is a crucial step towards understanding their complex decision-making pr

Location Is All You Need: Continuous Spatiotemporal Neural Representations of Earth Observation Data

ResearchDGX agent

arXiv:2604.07092v2 Announce Type: replace Abstract: In this work, we present LIANet (Location Is All You Need Network), a coordinate-based neural representation that models multi-temporal spaceborne E

Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification

ResearchDGX agent

arXiv:2604.08333v1 Announce Type: new Abstract: The rise of multimodal large language models (MLLMs) has sparked an unprecedented wave of applications in the field of medical imaging analysis. However

LPM 1.0: Video-based Character Performance Model

Model ReleasesDGX agent

arXiv:2604.07823v1 Announce Type: new Abstract: Performance, the externalization of intent, emotion, and personality through visual, vocal, and temporal behavior, is what makes a character alive. Lear

LumiCtrl : Learning Illuminant Prompts for Lighting Control in Personalized Text-to-Image Models

ResearchDGX agent

arXiv:2512.17489v2 Announce Type: replace Abstract: Text-to-image (T2I) models have demonstrated remarkable progress in creative image generation, yet they still lack precise control over scene illumi

Making MLLMs Blind: Adversarial Smuggling Attacks in MLLM Content Moderation

Model ReleasesDGX agent

arXiv:2604.06950v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) are increasingly being deployed as automated content moderators. Within this landscape, we uncover a critic

Mathematical Analysis of Image Matching Techniques

ResearchDGX agent

arXiv:2604.07574v1 Announce Type: new Abstract: Image matching is a fundamental problem in Computer Vision with direct applications in robotics, remote sensing, and geospatial data analysis. We presen

MCLR: Improving Conditional Modeling via Inter-Class Likelihood-Ratio Maximization and Unifying Classifier-Free Guidance with Alignment Objectives

SafetyDGX agent

arXiv:2603.22364v2 Announce Type: replace-cross Abstract: Diffusion models have achieved state-of-the-art performance in generative modeling, but their success often relies heavily on classifier-free

MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2604.08203v1 Announce Type: new Abstract: Medical Vision-Language Models (VLMs) hold immense promise for complex clinical tasks, but their reasoning capabilities are often constrained by text-on

MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping

ResearchDGX agent

arXiv:2604.08364v1 Announce Type: new Abstract: In this paper, we introduce MegaStyle, a novel and scalable data curation pipeline that constructs an intra-style consistent, inter-style diverse and hi

Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models

Local AiDGX agent

arXiv:2601.04068v3 Announce Type: replace Abstract: Aligning text-to-video diffusion models with human preferences is crucial for generating high-quality videos. Existing Direct Preference Otimization

Mitigating Domain Drift in Multi Species Segmentation with DINOv2: A Cross-Domain Evaluation in Herbicide Research Trials

ApplicationsDGX agent

arXiv:2508.07514v3 Announce Type: replace Abstract: Reliable plant species and damage segmentation for herbicide field research trials requires models that can withstand substantial real-world variati

Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction

ResearchDGX agent

arXiv:2604.07914v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have achieved remarkable success across cross-modal tasks but remain hindered by hallucinations, producing textual

← Previous
1…203204205206207
Next →