AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
13 Apr 2026

Zero-Shot Generative De-identification: Inversion-Free Flow for Privacy-Preserving Skin Image Analysis

ResearchDGX agent

arXiv:2602.00821v2 Announce Type: replace Abstract: The secure analysis of dermatological images in clinical environments is fundamentally restricted by the critical trade-off between patient privacy

10 Apr 2026

3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experience

Model ReleasesDGX agent

arXiv:2604.08042v1 Announce Type: new Abstract: Sketching in 3D space enables expressive reasoning about shape, structure, and spatial relationships, yet generating 3D sketches through natural languag

A Geometric Algorithm for Blood Vessel Reconstruction from Skeletal Representation


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
ResearchDGX agent

arXiv:2402.12797v4 Announce Type: replace Abstract: We introduce a novel approach for the reconstruction of tubular shapes from skeletal representations. Our method processes all skeletal points as a

A Physical Agentic Loop for Language-Guided Grasping with Execution-State Monitoring

SafetyDGX agent

arXiv:2604.07395v1 Announce Type: cross Abstract: Robotic manipulation systems that follow language instructions often execute grasp primitives in a largely single-shot manner: a model proposes an act

A Spatial-Spectral-Frequency Interactive Network for Multimodal Remote Sensing Classification

Model ReleasesDGX agent

arXiv:2510.04628v2 Announce Type: replace Abstract: Deep learning-based methods have achieved significant success in remote sensing Earth observation data analysis. Numerous feature fusion techniques

ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning

ResearchDGX agent

arXiv:2604.08050v1 Announce Type: new Abstract: In this study, we focus on video captioning by fully open multimodal large language models (MLLMs). The comprehension of visual sequences is challenging

Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models

AgentsDGX agent

arXiv:2604.08545v1 Announce Type: new Abstract: The advent of agentic multimodal models has empowered systems to actively interact with external environments. However, current agents suffer from a pro

Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection

Model ReleasesDGX agent

arXiv:2511.20162v2 Announce Type: replace Abstract: Large multi-modal models (LMMs) show increasing performance in realistic visual tasks for images and, more recently, for videos. For example, given

Adapting Foundation Models for Annotation-Efficient Adnexal Mass Segmentation in Cine Images

ResearchDGX agent

arXiv:2604.08045v1 Announce Type: new Abstract: Adnexal mass evaluation via ultrasound is a challenging clinical task, often hindered by subjective interpretation and significant inter-observer variab

Adaptive Depth-converted-Scale Convolution for Self-supervised Monocular Depth Estimation

Model ReleasesDGX agent

arXiv:2604.07665v1 Announce Type: new Abstract: Self-supervised monocular depth estimation (MDE) has received increasing interests in the last few years. The objects in the scene, including the object

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding

ResearchDGX agent

arXiv:2604.08077v1 Announce Type: new Abstract: Processing long-form videos with Video Large Language Models (Video-LLMs) is computationally prohibitive. Current efficiency methods often compromise fi

Adversarial Evasion Attacks on Computer Vision using SHAP Values

ResearchDGX agent

arXiv:2601.10587v2 Announce Type: replace Abstract: The paper introduces a white-box attack on computer vision models using SHAP values. It demonstrates how adversarial evasion attacks can compromise

Adversarial Flow Models

TutorialsDGX agent

arXiv:2511.22475v2 Announce Type: replace-cross Abstract: We present adversarial flow models, a class of generative models that belongs to both the adversarial and flow families. Our method supports n

AgriChain Visually Grounded Expert Verified Reasoning for Interpretable Agricultural Vision Language Models

Model ReleasesDGX agent

arXiv:2604.07814v1 Announce Type: new Abstract: Accurate and interpretable plant disease diagnosis remains a major challenge for vision-language models (VLMs) in real-world agriculture. We introduce A

AI-Driven Marine Robotics: Emerging Trends in Underwater Perception and Ecosystem Monitoring

ResearchDGX agent

arXiv:2509.01878v2 Announce Type: replace-cross Abstract: Marine ecosystems face increasing pressure due to climate change, driving the need for scalable, AI-powered monitoring solutions to inform eff

AnchorSplat: Feed-Forward 3D Gaussian Splatting with 3D Geometric Priors

Model ReleasesDGX agent

arXiv:2604.07053v2 Announce Type: replace Abstract: Recent feed-forward Gaussian reconstruction models adopt a pixel-aligned formulation that maps each 2D pixel to a 3D Gaussian, entangling Gaussian r

AnomalyAgent: Agentic Industrial Anomaly Synthesis via Tool-Augmented Reinforcement Learning

AgentsDGX agent

arXiv:2604.07900v1 Announce Type: new Abstract: Industrial anomaly generation is a crucial method for alleviating the data scarcity problem in anomaly detection tasks. Most existing anomaly synthesis

AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors

Model ReleasesDGX agent

arXiv:2601.20524v2 Announce Type: replace Abstract: Zero-shot anomaly detection aims to detect and localise abnormal regions in the image without access to any in-domain training images. While recent

AnyImageNav: Any-View Geometry for Precise Last-Meter Image-Goal Navigation

AgentsDGX agent

arXiv:2604.05351v3 Announce Type: replace-cross Abstract: Image Goal Navigation (ImageNav) is evaluated by a coarse success criterion, the agent must stop within 1m of the target, which is sufficient

ASBench: Image Anomalies Synthesis Benchmark for Anomaly Detection

Model ReleasesDGX agent

arXiv:2510.07927v2 Announce Type: replace Abstract: Anomaly detection plays a pivotal role in manufacturing quality control, yet its application is constrained by limited abnormal samples and high man

AtlasOCR: Building the First Open-Source Darija OCR Model with Vision Language Models

Model ReleasesDGX agent

arXiv:2604.08070v1 Announce Type: new Abstract: Darija, the Moroccan Arabic dialect, is rich in visual content yet lacks specialized Optical Character Recognition (OCR) tools. This paper introduces At

BADiff: Bandwidth Adaptive Diffusion Model

TutorialsDGX agent

arXiv:2510.21366v3 Announce Type: replace Abstract: In this work, we propose a novel framework to enable diffusion models to adapt their generation quality based on real-time network bandwidth constra

Bag of Bags: Adaptive Visual Vocabularies for Genizah Join Image Retrieval

ResearchDGX agent

arXiv:2604.08138v1 Announce Type: new Abstract: A join is a set of manuscript fragments identified as originally emanating from the same manuscript. We study manuscript join retrieval: Given a query i

Balanced Diffusion-Guided Fusion for Multimodal Remote Sensing Classification

TutorialsDGX agent

arXiv:2509.23310v3 Announce Type: replace Abstract: Deep learning-based techniques for the analysis of multimodal remote sensing data have become popular due to their ability to effectively integrate

Beyond Mamba: Enhancing State-space Models with Deformable Dilated Convolutions for Multi-scale Traffic Object Detection

Model ReleasesDGX agent

arXiv:2604.08038v1 Announce Type: new Abstract: In a real-world traffic scenario, varying-scale objects are usually distributed in a cluttered background, which poses great challenges to accurate dete

Beyond Pedestrians: Caption-Guided CLIP Framework for High-Difficulty Video-based Person Re-Identification

ResearchDGX agent

arXiv:2604.07740v1 Announce Type: new Abstract: In recent years, video-based person Re-Identification (ReID) has gained attention for its ability to leverage spatiotemporal cues to match individuals a

Beyond Surface Artifacts: Capturing Shared Latent Forgery Knowledge Across Modalities

Model ReleasesDGX agent

arXiv:2604.07763v1 Announce Type: new Abstract: As generative artificial intelligence evolves, deepfake attacks have escalated from single-modality manipulations to complex, multimodal threats. Existi

Bias Redistribution in Visual Machine Unlearning: Does Forgetting One Group Harm Another?

SafetyDGX agent

arXiv:2604.08111v1 Announce Type: cross Abstract: Machine unlearning enables models to selectively forget training data, driven by privacy regulations such as GDPR and CCPA. However, its fairness impl

BLaDA: Bridging Language to Functional Dexterous Actions within 3DGS Fields

ResearchDGX agent

arXiv:2604.08410v1 Announce Type: new Abstract: In unstructured environments, functional dexterous grasping calls for the tight integration of semantic understanding, precise 3D functional localizatio

Bootstrapping Sign Language Annotations with Sign Language Models

Model ReleasesDGX agent

arXiv:2604.07606v1 Announce Type: new Abstract: AI-driven sign language interpretation is limited by a lack of high-quality annotated data. New datasets including ASL STEM Wiki and FLEURS-ASL contain

Brain3D: EEG-to-3D Decoding of Visual Representations via Multimodal Reasoning

SafetyDGX agent

arXiv:2604.08068v1 Announce Type: new Abstract: Decoding visual information from electroencephalography (EEG) has recently achieved promising results, primarily focusing on reconstructing two-dimensio

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding

Local AiDGX agent

arXiv:2604.08014v1 Announce Type: new Abstract: Spatio-Temporal Video Grounding requires jointly localizing target objects across both temporal and spatial dimensions based on natural language queries

BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook

Model ReleasesDGX agent

arXiv:2506.12040v2 Announce Type: replace-cross Abstract: Binary quantization represents the most extreme form of compression, reducing weights to +/-1 for maximal memory and computational efficiency.

CAMotion: A High-Quality Benchmark for Camouflaged Moving Object Detection in the Wild

Model ReleasesDGX agent

arXiv:2604.08287v1 Announce Type: new Abstract: Discovering camouflaged objects is a challenging task in computer vision due to the high similarity between camouflaged objects and their surroundings.

ChangeBridge: Spatiotemporal Image Generation with Multimodal Controls for Remote Sensing

ResearchDGX agent

arXiv:2507.04678v3 Announce Type: replace Abstract: Spatiotemporal image generation is a highly meaningful task, which can generate future scenes conditioned on given observations. However, existing c

CodecSight: Leveraging Video Codec Signals for Efficient Streaming VLM Inference

HardwareDGX agent

arXiv:2604.06036v3 Announce Type: replace-cross Abstract: Video streaming analytics is a crucial workload for vision-language model serving, but the high cost of multimodal inference limits scalabilit

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding

Model ReleasesDGX agent

arXiv:2603.18472v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) perform strongly on natural images, yet their ability to understand discrete visual symbols remains u

CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs

ApplicationsDGX agent

arXiv:2510.12184v2 Announce Type: replace Abstract: Recently, efficient Multimodal Large Language Models (MLLMs) have gained significant attention as a solution to their high computational complexity,

Component-Adaptive and Lesion-Level Supervision for Improved Small Structure Segmentation in Brain MRI

ResearchDGX agent

arXiv:2604.08015v1 Announce Type: new Abstract: We propose a unified objective function, termed CATMIL, that augments the base segmentation loss with two auxiliary supervision terms operating at diffe

Coordinate-Based Dual-Constrained Autoregressive Motion Generation

TutorialsDGX agent

arXiv:2604.08088v1 Announce Type: new Abstract: Text-to-motion generation has attracted increasing attention in the research community recently, with potential applications in animation, virtual reali

Cost-Efficient Multi-Scale Fovea for Semantic-Based Visual Search Attention

ResearchDGX agent

arXiv:2604.03836v2 Announce Type: replace-cross Abstract: Semantics are one of the primary sources of top-down preattentive information. Modern deep object detectors excel at extracting such valuable

CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoning

Model ReleasesDGX agent

arXiv:2604.08457v1 Announce Type: new Abstract: Cooperative autonomous driving requires traffic scene understanding from both vehicle and infrastructure perspectives. While vision-language models (VLM

Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video

TutorialsDGX agent

arXiv:2604.07786v1 Announce Type: new Abstract: Talking face generation has gained significant attention as a core application of generative models. To enhance the expressiveness and realism of synthe

CryoSplat: Gaussian Splatting for Cryo-EM Homogeneous Reconstruction

Model ReleasesDGX agent

arXiv:2508.04929v4 Announce Type: replace-cross Abstract: As a critical modality for structural biology, cryogenic electron microscopy (cryo-EM) facilitates the determination of macromolecular structu

DailyArt: Discovering Articulation from Single Static Images via Latent Dynamics

ResearchDGX agent

arXiv:2604.07758v1 Announce Type: new Abstract: Articulated objects are essential for embodied AI and world models, yet inferring their kinematics from a single closed-state image remains challenging

DBMF: A Dual-Branch Multimodal Framework for Out-of-Distribution Detection

ApplicationsDGX agent

arXiv:2604.08261v1 Announce Type: new Abstract: The complex and dynamic real-world clinical environment demands reliable deep learning (DL) systems. Out-of-distribution (OOD) detection plays a critica

Deep Learning-Powered Visual SLAM Aimed at Assisting Visually Impaired Navigation

SafetyDGX agent

arXiv:2510.20549v2 Announce Type: replace Abstract: Despite advancements in SLAM technologies, robust operation under challenging conditions such as low-texture, motion-blur, or challenging lighting r

DiffVC: A Non-autoregressive Framework Based on Diffusion Model for Video Captioning

ResearchDGX agent

arXiv:2604.08084v1 Announce Type: new Abstract: Current video captioning methods usually use an encoder-decoder structure to generate text autoregressively. However, autoregressive methods have inhere

DinoRADE: Full Spectral Radar-Camera Fusion with Vision Foundation Model Features for Multi-class Object Detection in Adverse Weather

AgentsDGX agent

arXiv:2604.08074v1 Announce Type: new Abstract: Reliable and weather-robust perception systems are essential for safe autonomous driving and typically employ multi-modal sensor configurations to achie

Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic Segmentation

Model ReleasesDGX agent

arXiv:2604.07723v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation (OVSS) aims to segment arbitrary category regions in images using open-vocabulary prompts, necessitating that exis

Distilling Specialized Orders for Visual Generation

ResearchDGX agent

arXiv:2504.17069v2 Announce Type: replace Abstract: Autoregressive (AR) image generators are becoming increasingly popular due to their ability to produce high-quality images and their scalability. Ty

DMin: Scalable Training Data Influence Estimation for Diffusion Models

ResearchDGX agent

arXiv:2412.08637v4 Announce Type: replace Abstract: Identifying the training data samples that most influence a generated image is a critical task in understanding diffusion models (DMs), yet existing

DP-DeGauss: Dynamic Probabilistic Gaussian Decomposition for Egocentric 4D Scene Reconstruction

ResearchDGX agent

arXiv:2604.07986v1 Announce Type: new Abstract: Egocentric video is crucial for next-generation 4D scene reconstruction, with applications in AR/VR and embodied AI. However, reconstructing dynamic fir

DSCA: Dynamic Subspace Concept Alignment for Lifelong VLM Editing

Model ReleasesDGX agent

arXiv:2604.07965v1 Announce Type: new Abstract: Model editing aims to update knowledge to add new concepts and change relevant information without retraining. Lifelong editing is a challenging task, p

Dual-level Modality Debiasing Learning for Unsupervised Visible-Infrared Person Re-Identification

Model ReleasesDGX agent

arXiv:2512.03745v2 Announce Type: replace Abstract: Two-stage learning pipeline has achieved promising results in unsupervised visible-infrared person re-identification (USL-VI-ReID). It first perform

E-3DPSM: A State Machine for Event-Based Egocentric 3D Human Pose Estimation

ResearchDGX agent

arXiv:2604.08543v1 Announce Type: new Abstract: Event cameras offer multiple advantages in monocular egocentric 3D human pose estimation from head-mounted devices, such as millisecond temporal resolut

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation

SafetyDGX agent

arXiv:2602.13669v4 Announce Type: replace Abstract: Recent multi-modal video generation models have achieved high visual quality, but their prohibitive latency and limited temporal stability hinder re

Ecological Legacies of Pre-Columbian Settlements Evident in Palm Clusters of Neotropical Mountain Forests

ResearchDGX agent

arXiv:2507.06949v3 Announce Type: replace Abstract: Ancient populations inhabited and transformed neotropical forests, yet the spatial extent of their ecological influence remains underexplored at hig

EditCaption: Human-Aligned Instruction Synthesis for Image Editing via Supervised Fine-Tuning and Direct Preference Optimization

Model ReleasesDGX agent

arXiv:2604.08213v1 Announce Type: new Abstract: High-quality training triplets (source-target image pairs with precise editing instructions) are a critical bottleneck for scaling instruction-guided im

EEG2Vision: A Multimodal EEG-Based Framework for 2D Visual Reconstruction in Cognitive Neuroscience

ResearchDGX agent

arXiv:2604.08063v1 Announce Type: new Abstract: Reconstructing visual stimuli from non-invasive electroencephalography (EEG) remains challenging due to its low spatial resolution and high noise, parti

← Previous
1…202203204205206207
Next →