AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
22 Apr 2026

LoViF 2026 Challenge on Real-World All-in-One Image Restoration: Methods and Results

Model ReleasesDGX agent

arXiv:2604.19445v1 Announce Type: new Abstract: This paper presents a review for the LoViF Challenge on Real-World All-in-One Image Restoration. The challenge aimed to advance research on real-world a

MAGICIAN: Efficient Long-Term Planning with Imagined Gaussians for Active Mapping

AgentsDGX agent

arXiv:2603.22650v2 Announce Type: replace Abstract: Active mapping aims to determine how an agent should move to efficiently reconstruct unknown environments. Most existing approaches rely on greedy n

Match-Any-Events: Zero-Shot Motion-Robust Feature Matching Across Wide Baselines for Event Cameras

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.18744v1 Announce Type: new Abstract: Event cameras have recently shown promising capabilities in instantaneous motion estimation due to their robustness to low light and fast motions. Howev

MedFlowSeg: Flow Matching for Medical Image Segmentation with Frequency-Aware Attention

ResearchDGX agent

arXiv:2604.19675v1 Announce Type: new Abstract: Flow matching has recently emerged as a principled framework for learning continuous-time transport maps, enabling efficient deterministic generation wi

Memory Over Maps: 3D Object Localization Without Reconstruction

Local AiDGX agent

arXiv:2603.20530v2 Announce Type: replace-cross Abstract: Target localization is a prerequisite for embodied tasks such as navigation and manipulation. Conventional approaches rely on constructing exp

Mind2Drive: Predicting Driver Intentions from EEG in Real-world On-Road Driving

SafetyDGX agent

arXiv:2604.19368v1 Announce Type: new Abstract: Predicting driver intention from neurophysiological signals offers a promising pathway for enhancing proactive safety in advanced driver assistance syst

MiTA Attention: Efficient Fast-Weight Scaling via a Mixture of Top-k Activations

ResearchDGX agent

arXiv:2602.01219v4 Announce Type: replace-cross Abstract: The attention operator in Transformers can be viewed as a two-layer fast-weight MLP, whose weights are dynamically instantiated from input tok

Mitigating Long-Tail Bias via Prompt-Controlled Diffusion Augmentation

SafetyDGX agent

arXiv:2602.04749v2 Announce Type: replace Abstract: Long-tailed class imbalance remains a fundamental obstacle in semantic segmentation of high-resolution remote-sensing imagery, where dominant classe

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation

SafetyDGX agent

arXiv:2604.19679v1 Announce Type: new Abstract: Recent advances in Diffusion Transformers (DiTs) have enabled high-quality joint audio-video generation, producing videos with synchronized audio within

MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors

ResearchDGX agent

arXiv:2512.15577v2 Announce Type: replace Abstract: In this paper, we focus on online zero-shot monocular 3D instance segmentation, a novel practical setting where existing approaches fail to perform

MOSA: Motion-Guided Semantic Alignment for Dynamic Scene Graph Generation

SafetyDGX agent

arXiv:2604.19631v1 Announce Type: new Abstract: Dynamic Scene Graph Generation (DSGG) aims to structurally model objects and their dynamic interactions in video sequences for high-level semantic under

MSDS: Deep Structural Similarity with Multiscale Representation

Model ReleasesDGX agent

arXiv:2604.19159v1 Announce Type: new Abstract: Deep-feature-based perceptual similarity models have demonstrated strong alignment with human visual perception in Image Quality Assessment (IQA). Howev

Multi-Domain Learning with Global Expert Mapping

Model ReleasesDGX agent

arXiv:2604.18842v1 Announce Type: new Abstract: Human perception generalizes well across different domains, but most vision models struggle beyond their training data. This gap motivates multi-dataset

Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes

ApplicationsDGX agent

arXiv:2604.19318v1 Announce Type: new Abstract: Multi-view crowd tracking estimates each person's tracking trajectories on the ground of the scene. Recent research works mainly rely on CNNs-based mult

On the Generalizability of Foundation Models for Crop Type Mapping

SafetyDGX agent

arXiv:2409.09451v5 Announce Type: replace Abstract: Foundation models pre-trained using self-supervised learning have shown powerful transfer learning capabilities on various downstream tasks, includi

PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving

AgentsDGX agent

arXiv:2604.19379v1 Announce Type: new Abstract: This paper presents the first study on Unsupervised Domain Adaptation (UDA) for multimodal 3D panoptic segmentation (mm-3DPS), aiming to improve general

Paparazzo: Active Mapping of Moving 3D Objects

Model ReleasesDGX agent

arXiv:2604.19556v1 Announce Type: new Abstract: Current 3D mapping pipelines generally assume static environments, which limits their ability to accurately capture and reconstruct moving objects. To a

PC2Model: ISPRS benchmark on 3D point cloud to model registration

Model ReleasesDGX agent

arXiv:2604.19596v1 Announce Type: new Abstract: Point cloud registration involves aligning one point cloud with another or with a three-dimensional (3D) model, enabling the integration of multimodal d

Personalized Embodied Navigation for Portable Object Finding

AgentsDGX agent

arXiv:2403.09905v5 Announce Type: replace-cross Abstract: Embodied navigation methods commonly operate in static environments with stationary objects. In this work, we present approaches for tackling

PhotoFramer: Multi-modal Image Composition Instruction

TutorialsDGX agent

arXiv:2512.00993v2 Announce Type: replace Abstract: Composition matters during the photo-taking process, yet many casual users struggle to frame well-composed images. To provide composition guidance,

Pixels or Positions? Benchmarking Modalities in Group Activity Recognition

Model ReleasesDGX agent

arXiv:2511.12606v3 Announce Type: replace Abstract: Group Activity Recognition (GAR) is well studied on the video modality for surveillance and indoor team sports (e.g., volleyball, basketball). Yet,

PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment

HardwareDGX agent

arXiv:2604.19129v1 Announce Type: new Abstract: Existing facial reenactment methods struggle with a trade-off between expressiveness and fine-grained controllability. Holistic facial reenactment model

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.00161v2 Announce Type: replace Abstract: Optical Character Recognition (OCR) is increasingly regarded as a foundational capability for modern vision-language models (VLMs), enabling them no

RAFT-MSF++: Temporal Geometry-Motion Feature Fusion for Self-Supervised Monocular Scene Flow

Model ReleasesDGX agent

arXiv:2604.19349v1 Announce Type: new Abstract: Monocular scene flow estimation aims to recover dense 3D motion from image sequences, yet most existing methods are limited to two-frame inputs, restric

Realistic Handwritten Multi-Digit Writer (MDW) Number Recognition Challenges

Model ReleasesDGX agent

arXiv:2512.00676v2 Announce Type: replace Abstract: Isolated digit classification has served as a motivating problem for decades of machine learning research. In real settings, numbers often occur as

Recurrent Video Masked Autoencoders

Model ReleasesDGX agent

arXiv:2512.13684v2 Announce Type: replace Abstract: We present Recurrent Video Masked-Autoencoders (RVM): a novel approach to video representation learning that leverages recurrent computation to mode

ReImagine: Rethinking Controllable High-Quality Human Video Generation via Image-First Synthesis

ResearchDGX agent

arXiv:2604.19720v1 Announce Type: new Abstract: Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint under limited multi-

RESFL: An Uncertainty-Aware Framework for Responsible Federated Learning by Balancing Privacy, Fairness and Utility

SafetyDGX agent

arXiv:2503.16251v2 Announce Type: replace-cross Abstract: Federated Learning (FL) has gained prominence in machine learning applications across critical domains by enabling collaborative model trainin

Rethinking Dataset Distillation: Hard Truths about Soft Labels

ResearchDGX agent

arXiv:2604.18811v1 Announce Type: cross Abstract: Despite the perceived success of large-scale dataset distillation (DD) methods, recent evidence finds that simple random image baselines perform on-pa

RF-HiT: Rectified Flow Hierarchical Transformer for General Medical Image Segmentation

ResearchDGX agent

arXiv:2604.19570v1 Announce Type: new Abstract: Accurate medical image segmentation requires both long-range contextual reasoning and precise boundary delineation, a task where existing transformer- a

Robust Continual Unlearning against Knowledge Erosion and Forgetting Reversal

ResearchDGX agent

arXiv:2604.19108v1 Announce Type: cross Abstract: As a means to balance the growth of the AI industry with the need for privacy protection, machine unlearning plays a crucial role in realizing the ``r

SAGE: Training-Free Semantic Evidence Composition for Edge-Cloud Inference under Hard Uplink Budgets

ResearchDGX agent

arXiv:2604.19623v1 Announce Type: cross Abstract: Edge-cloud hybrid inference offloads difficult inputs to a powerful remote model, but the uplink channel imposes hard per-request constraints on the n

Seeing Candidates at Scale: Multimodal LLMs for Visual Political Communication on Instagram

ApplicationsDGX agent

arXiv:2604.19489v1 Announce Type: new Abstract: This paper presents a computational case study that evaluates the capabilities of specialized machine learning models and emerging multimodal large lang

Silicon Aware Neural Networks

HardwareDGX agent

arXiv:2604.19334v1 Announce Type: new Abstract: Recent work in the machine learning literature has demonstrated that deep learning can train neural networks made of discrete logic gate functions to pe

SketchFaceGS: Real-Time Sketch-Driven Face Editing and Generation with Gaussian Splatting

ResearchDGX agent

arXiv:2604.19202v1 Announce Type: cross Abstract: 3D Gaussian representations have emerged as a powerful paradigm for digital head modeling, achieving photorealistic quality with real-time rendering.

SMART-Ship: A Comprehensive Synchronized Multi-modal Aligned Remote Sensing Targets Dataset and Benchmark for Berthed Ships Analysis

Model ReleasesDGX agent

arXiv:2508.02384v2 Announce Type: replace Abstract: Given the limitations of satellite orbits and imaging conditions, multi-modal remote sensing (RS) data is crucial in enabling long-term earth observ

SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing

ResearchDGX agent

arXiv:2604.19587v1 Announce Type: new Abstract: Traditional photographic image editing typically requires users to possess sufficient aesthetic understanding to provide appropriate instructions for ad

SmokeGS-R: Physics-Guided Pseudo-Clean 3DGS for Real-World Multi-View Smoke Restoration

Model ReleasesDGX agent

arXiv:2604.05301v2 Announce Type: replace Abstract: Real-world smoke simultaneously attenuates scene radiance, adds airlight, and destabilizes multi-view appearance consistency, making robust 3D recon

SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model

SafetyDGX agent

arXiv:2604.19710v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models offer a promising autonomous driving paradigm for leveraging world knowledge and reasoning capabilities, especially

StomaD2: An All-in-One System for Intelligent Stomatal Phenotype Analysis via Diffusion-Based Restoration Detection Network

ResearchDGX agent

arXiv:2604.18632v1 Announce Type: new Abstract: Stomata play a crucial role in regulating plant physiological processes and reflecting environmental responses. However, accurate and high-throughput st

Structure-Semantic Decoupled Modulation of Global Geospatial Embeddings for High-Resolution Remote Sensing Mapping

ResearchDGX agent

arXiv:2604.19591v1 Announce Type: new Abstract: Fine-grained high-resolution remote sensing mapping typically relies on localized visual features, which restricts cross-domain generalizability and oft

Task Switching Without Forgetting via Proximal Decoupling

TutorialsDGX agent

arXiv:2604.18857v1 Announce Type: cross Abstract: In continual learning, the primary challenge is to learn new information without forgetting old knowledge. A common solution addresses this trade-off

TESO: Online Tracking of Essential Matrix by Stochastic Optimization

AgentsDGX agent

arXiv:2604.19420v1 Announce Type: new Abstract: Maintaining long-term accuracy of stereo camera calibration parameters is important for autonomous systems' perception. This work proposes Online Tracki

The Essence of Balance for Self-Improving Agents in Vision-and-Language Navigation

SafetyDGX agent

arXiv:2604.19064v1 Announce Type: new Abstract: In vision-and-language navigation (VLN), self-improvement from policy-induced experience, using only standard VLN action supervision, critically depends

Thinking Before Matching: A Reinforcement Reasoning Paradigm Towards General Person Re-Identification

TutorialsDGX agent

arXiv:2604.19218v1 Announce Type: new Abstract: Learning identity-discriminative representations with multi-scene generality has become a critical objective in person re-identification (ReID). However

Toward Clinically Acceptable Chest X-ray Report Generation: A Qualitative Retrospective Pilot Study of CXRMate-2

SafetyDGX agent

arXiv:2604.18967v1 Announce Type: new Abstract: Chest X-ray (CXR) radiology report generation (RRG) models have shown rapid progress, yet their clinical utility remains uncertain due to limited evalua

Towards Reliable Human Evaluations in Gesture Generation: Insights from a Community-Driven State-of-the-Art Benchmark

Model ReleasesDGX agent

arXiv:2511.01233v3 Announce Type: replace Abstract: We review human evaluation practices in automatic, speech-driven 3D gesture generation and find a lack of standardisation and frequent use of flawed

TransSplat: Unbalanced Semantic Transport for Language-Driven 3DGS Editing

TutorialsDGX agent

arXiv:2604.19571v1 Announce Type: new Abstract: Language-driven 3D Gaussian Splatting (3DGS) editing provides a more convenient approach for modifying complex scenes in VR/AR. Standard pipelines typic

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation

ResearchDGX agent

arXiv:2604.19473v1 Announce Type: new Abstract: Generating high-quality videos from complex temporal descriptions that contain multiple sequential actions is a key unsolved problem. Existing methods a

Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items

Model ReleasesDGX agent

arXiv:2604.19748v1 Announce Type: new Abstract: Recent advances in image generation and editing have opened new opportunities for virtual try-on. However, existing methods still struggle to meet compl

Unposed-to-3D: Learning Simulation-Ready Vehicles from Real-World Images

AgentsDGX agent

arXiv:2604.19257v1 Announce Type: new Abstract: Creating realistic and simulation-ready 3D assets is crucial for autonomous driving research and virtual environment construction. However, existing 3D

Unveiling Fine-Grained Visual Traces: Evaluating Multimodal Interleaved Reasoning Chains in Multimodal STEM Tasks

Model ReleasesDGX agent

arXiv:2604.19697v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have shown promising reasoning abilities, yet evaluating their performance in specialized domains remains chall

URoPE: Universal Relative Position Embedding across Geometric Spaces

Model ReleasesDGX agent

arXiv:2604.18747v1 Announce Type: new Abstract: Relative position embedding has become a standard mechanism for encoding positional information in Transformers. However, existing formulations are typi

VDPP: Video Depth Post-Processing for Speed and Scalability

Model ReleasesDGX agent

arXiv:2604.06665v2 Announce Type: replace Abstract: Video depth estimation is essential for providing 3D scene structure in applications ranging from autonomous driving to mixed reality. Current end-t

VecHeart: Holistic Four-Chamber Cardiac Anatomy Modeling via Hybrid VecSets

Model ReleasesDGX agent

arXiv:2604.19403v1 Announce Type: new Abstract: Accurate cardiac anatomy modeling requires the model to be able to handle intricate interrelations among structures. In this paper, we propose VecHeart,

Vision-Based Human Awareness Estimation for Enhanced Safety and Efficiency of AMRs in Industrial Warehouses

SafetyDGX agent

arXiv:2604.18627v1 Announce Type: new Abstract: Ensuring human safety is of paramount importance in warehouse environments that feature mixed traffic of human workers and autonomous mobile robots (AMR

Visual Adversarial Attack on Vision-Language Models for Autonomous Driving

SafetyDGX agent

arXiv:2411.18275v2 Announce Type: replace Abstract: Vision-language models (VLMs) have significantly advanced autonomous driving (AD) by enhancing reasoning capabilities. However, these models remain

Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding

ResearchDGX agent

arXiv:2604.19609v1 Announce Type: new Abstract: Transformers have become a common foundation across deep learning, yet 3D scene understanding still relies on specialized backbones with strong domain p

Weakly supervised framework for wildlife detection and counting in challenging Arctic environments: a case study on caribou (Rangifer tarandus)

SafetyDGX agent

arXiv:2601.18891v3 Announce Type: replace Abstract: Caribou across the Arctic has declined in recent decades, motivating scalable and accurate monitoring approaches to guide evidence-based conservatio

When Can We Trust Deep Neural Networks? Towards Reliable Industrial Deployment with an Interpretability Guide

SafetyDGX agent

arXiv:2604.19206v1 Announce Type: new Abstract: The deployment of AI systems in safety-critical domains, such as industrial defect inspection, autonomous driving, and medical diagnosis, is severely ha

← Previous
1…178179180181182…209
Next →