AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
16 Apr 2026

Artificial intelligence application in lymphoma diagnosis with Vision Transformer using weakly supervised training

ResearchDGX agent

arXiv:2604.13795v1 Announce Type: new Abstract: Vision transformers (ViT) have been shown to allow for more flexible feature detection and can outperform convolutional neural network (CNN) when pre-tr

ASTER: Latent Pseudo-Anomaly Generation for Unsupervised Time-Series Anomaly Detection

Model ReleasesDGX agent

arXiv:2604.13924v1 Announce Type: cross Abstract: Time-series anomaly detection (TSAD) is critical in domains such as industrial monitoring, healthcare, and cybersecurity, but it remains challenging d

ASTRA: Enhancing Multi-Subject Generation with Retrieval-Augmented Pose Guidance and Disentangled Position Embedding

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2604.13938v1 Announce Type: new Abstract: Subject-driven image generation has shown great success in creating personalized content, but its capabilities are largely confined to single subjects i

AudioX: A Unified Framework for Anything-to-Audio Generation

Model ReleasesDGX agent

arXiv:2503.10522v4 Announce Type: replace-cross Abstract: Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a

Automatic Charge State Tuning of 300 mm FDSOI Quantum Dots Using Neural Network Segmentation of Charge Stability Diagram

HardwareDGX agent

arXiv:2604.13662v1 Announce Type: cross Abstract: Tuning of gate-defined semiconductor quantum dots (QDs) is a major bottleneck for scaling spin qubit technologies. We present a deep learning (DL) dri

Beyond Voxel 3D Editing: Learning from 3D Masks and Self-Constructed Data

Local AiDGX agent

arXiv:2604.13688v1 Announce Type: new Abstract: 3D editing refers to the ability to apply local or global modifications to 3D assets. Effective 3D editing requires maintaining semantic consistency by

Bias at the End of the Score

SafetyDGX agent

arXiv:2604.13305v1 Announce Type: new Abstract: Reward models (RMs) are inherently non-neutral value functions designed and trained to encode specific objectives, such as human preferences or text-ima

Blind Bitstream-corrupted Video Recovery via Metadata-guided Diffusion Model

ApplicationsDGX agent

arXiv:2604.13906v1 Announce Type: new Abstract: Bitstream-corrupted video recovery aims to restore realistic content degraded during video storage or transmission. Existing methods typically assume th

C^2T: Captioning-Structure and LLM-Aligned Common-Sense Reward Learning for Traffic--Vehicle Coordination

SafetyDGX agent

arXiv:2604.13098v1 Announce Type: cross Abstract: State-of-the-art (SOTA) urban traffic control increasingly employs Multi-Agent Reinforcement Learning (MARL) to coordinate Traffic Light Controllers (

Can Cross-Layer Transcoders Replace Vision Transformer Activations? An Interpretable Perspective on Vision

ResearchDGX agent

arXiv:2604.13304v1 Announce Type: new Abstract: Understanding the internal activations of Vision Transformers (ViTs) is critical for building interpretable and trustworthy models. While Sparse Autoenc

CausalDisenSeg: A Causality-Guided Disentanglement Framework with Counterfactual Reasoning for Robust Brain Tumor Segmentation Under Missing Modalities

SafetyDGX agent

arXiv:2604.13409v1 Announce Type: new Abstract: In clinical practice, the robustness of deep learning models for multimodal brain tumor segmentation is severely compromised by incomplete MRI data. Thi

CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling

SafetyDGX agent

arXiv:2604.13561v1 Announce Type: new Abstract: Vision-language models trained with contrastive learning on paired medical images and reports show strong zero-shot diagnostic capabilities, yet the eff

ClipGStream: Clip-Stream Gaussian Splatting for Any Length and Any Motion Multi-View Dynamic Scene Reconstruction

ResearchDGX agent

arXiv:2604.13746v1 Announce Type: new Abstract: Dynamic 3D scene reconstruction is essential for immersive media such as VR, MR, and XR, yet remains challenging for long multi-view sequences with larg

Context Sensitivity Improves Human-Machine Visual Alignment

SafetyDGX agent

arXiv:2604.13883v1 Announce Type: new Abstract: Modern machine learning models typically represent inputs as fixed points in a high-dimensional embedding space. While this approach has been proven pow

Creo: From One-Shot Image Generation to Progressive, Co-Creative Ideation

ResearchDGX agent

arXiv:2604.13956v1 Announce Type: cross Abstract: Text-to-image (T2I) systems enable rapid generation of high-fidelity imagery but are misaligned with how visual ideas develop. T2I systems generate ou

Cyclic 2.5D Perceptual Loss for Cross-Modal 3D Medical Image Synthesis: T1w MRI to Tau PET

ResearchDGX agent

arXiv:2406.12632v3 Announce Type: replace-cross Abstract: Positron emission tomography (PET) provides molecular biomarkers for Alzheimer's disease and related dementias (ADRD) and is increasingly used

Decoding the Delta: Unifying Remote Sensing Change Detection and Understanding with Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2604.14044v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) excel in general vision-language tasks, their application to remote sensing change understanding is hinde

Deep Spatially-Regularized and Superpixel-Based Diffusion Learning for Unsupervised Hyperspectral Image Clustering

ResearchDGX agent

arXiv:2604.13307v1 Announce Type: new Abstract: An unsupervised framework for hyperspectral image (HSI) clustering is proposed that incorporates masked deep representation learning with diffusion-base

Dehaze-then-Splat: Generative Dehazing with Physics-Informed 3D Gaussian Splatting for Smoke-Free Novel View Synthesis

ResearchDGX agent

arXiv:2604.13589v1 Announce Type: new Abstract: We present Dehaze-then-Splat, a two-stage pipeline for multi-view smoke removal and novel view synthesis developed for Track~2 of the NTIRE 2026 3D Rest

Delineate Anything Flow: Fast, Country-Level Field Boundary Detection from Any Source

ResearchDGX agent

arXiv:2511.13417v2 Announce Type: replace Abstract: Accurate delineation of agricultural field boundaries from satellite imagery is essential for land management and crop monitoring, yet existing meth

Democratising Pathology Co-Pilots: An Open Pipeline and Dataset for Whole-Slide Vision-Language Modelling

ResearchDGX agent

arXiv:2512.17326v2 Announce Type: replace Abstract: Vision-language models (VLMs) have the potential to become co-pilots for pathologists. However, most VLMs either focus on small regions of interest

Depth-Aware Image and Video Orientation Estimation

SafetyDGX agent

arXiv:2604.13995v1 Announce Type: new Abstract: This paper introduces a novel approach for image and video orientation estimation by leveraging depth distribution in natural images. The proposed metho

Depth-Resolved Coral Reef Thermal Fields from Satellite SST and Sparse In-Situ Loggers Using Physics-Informed Neural Networks

ResearchDGX agent

arXiv:2604.13131v1 Announce Type: cross Abstract: Satellite sea surface temperature (SST) products underpin global coral bleaching monitoring, yet they measure only the ocean skin. Corals inhabit dept

Design and Behavior of Sparse Mixture-of-Experts Layers in CNN-based Semantic Segmentation

ResearchDGX agent

arXiv:2604.13761v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) layers have been shown to substantially increase model capacity without a proportional increase in computational cost an

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis

Model ReleasesDGX agent

arXiv:2604.13416v1 Announce Type: new Abstract: Advances in radiance fields have enabled photorealistic novel view synthesis. In several domains, large-scale real-world datasets have been developed to

DiffMagicFace: Identity Consistent Facial Editing of Real Videos

ResearchDGX agent

arXiv:2604.13841v1 Announce Type: new Abstract: Text-conditioned image editing has greatly benefitted from the advancements in Image Diffusion Models. However, extending these techniques to facial vid

DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer

ResearchDGX agent

arXiv:2604.13509v1 Announce Type: new Abstract: Recent advances in video generation models has significantly accelerated video generation and related downstream tasks. Among these, video stylization h

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models

ResearchDGX agent

arXiv:2604.14129v1 Announce Type: new Abstract: While Audio-Visual Language Models (AVLMs) have achieved remarkable progress over recent years, their reliability is bottlenecked by cross-modal halluci

DRG-Font: Dynamic Reference-Guided Few-shot Font Generation via Contrastive Style-Content Disentanglement

ResearchDGX agent

arXiv:2604.13797v1 Announce Type: new Abstract: Few-shot Font Generation aims to generate stylistically consistent glyphs from a few reference glyphs. However, capturing complex font styles from a few

DroneScan-YOLO: Redundancy-Aware Lightweight Detection for Tiny Objects in UAV Imagery

ResearchDGX agent

arXiv:2604.13278v1 Announce Type: new Abstract: Aerial object detection in UAV imagery presents unique challenges due to the high prevalence of tiny objects, adverse environmental conditions, and stri

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models

SafetyDGX agent

arXiv:2602.20981v3 Announce Type: replace Abstract: Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and

Efficient Multi-View 3D Object Detection by Dynamic Token Selection and Fine-Tuning

Model ReleasesDGX agent

arXiv:2604.13586v1 Announce Type: new Abstract: Existing multi-view three-dimensional (3D) object detection approaches widely adopt large-scale pre-trained vision transformer (ViT)-based foundation mo

Enhanced Text-to-Image Generation by Fine-grained Multimodal Reasoning

SafetyDGX agent

arXiv:2604.13491v1 Announce Type: new Abstract: With the rapid progress of Multimodal Large Language Models (MLLMs), unified MLLMs that jointly perform image understanding and generation have advanced

Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling

ResearchDGX agent

arXiv:2604.13508v1 Announce Type: new Abstract: Sparse Upcycling provides an efficient way to initialize a Mixture-of-Experts (MoE) model from pretrained dense weights instead of training from scratch

ESCAPE: Episodic Spatial Memory and Adaptive Execution Policy for Long-Horizon Mobile Manipulation

Model ReleasesDGX agent

arXiv:2604.13633v1 Announce Type: new Abstract: Coordinating navigation and manipulation with robust performance is essential for embodied AI in complex indoor environments. However, as tasks extend o

Event-Adaptive State Transition and Gated Fusion for RGB-Event Object Tracking

ResearchDGX agent

arXiv:2604.13426v1 Announce Type: new Abstract: Existing Vision Mamba-based RGB-Event(RGBE) tracking methods suffer from using static state transition matrices, which fail to adapt to variations in ev

Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization

SafetyDGX agent

arXiv:2604.13533v1 Announce Type: cross Abstract: Achieving general-purpose robotics requires empowering robots to adapt and evolve based on their environment and feedback. Traditional methods face li

Explainable Fall Detection for Elderly Care via Temporally Stable SHAP in Skeleton-Based Human Activity Recognition

Local AiDGX agent

arXiv:2604.13279v1 Announce Type: new Abstract: Fall detection in elderly care requires not only accurate classification but also reliable explanations that clinicians can trust. However, existing pos

Failure Identification in Imitation Learning Via Statistical and Semantic Filtering

SafetyDGX agent

arXiv:2604.13788v1 Announce Type: cross Abstract: Imitation learning (IL) policies in robotics deliver strong performance in controlled settings but remain brittle in real-world deployments: rare even

Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective

ApplicationsDGX agent

arXiv:2604.14025v1 Announce Type: new Abstract: Reconstructing 3D representations from 2D inputs is a fundamental task in computer vision and graphics, serving as a cornerstone for understanding and i

FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks

Model ReleasesDGX agent

arXiv:2505.19662v3 Announce Type: replace-cross Abstract: This paper introduces FieldWorkArena, a benchmark for agentic AI targeting real-world field work. With the recent increase in demand for agent

Free Geometry: Refining 3D Reconstruction from Longer Versions of Itself

Model ReleasesDGX agent

arXiv:2604.14048v1 Announce Type: new Abstract: Feed-forward 3D reconstruction models are efficient but rigid: once trained, they perform inference in a zero-shot manner and cannot adapt to the test s

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding

ResearchDGX agent

arXiv:2604.13540v1 Announce Type: new Abstract: Unified Multimodal Models (UMMs) aim to integrate visual understanding and generation within a single structure. However, these models exhibit a notable

From Instruction to Event: Sound-Triggered Mobile Manipulation

SafetyDGX agent

arXiv:2601.21667v2 Announce Type: replace-cross Abstract: Current mobile manipulation research predominantly follows an instruction-driven paradigm, where agents rely on predefined textual commands to

From Pixels to Nucleotides: End-to-End Token-Based Video Compression for DNA Storage

ResearchDGX agent

arXiv:2604.13667v1 Announce Type: new Abstract: DNA-based storage has emerged as a promising approach to the global data crisis, offering molecular-scale density and millennial-scale stability at low

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models

Model ReleasesDGX agent

arXiv:2603.19790v3 Announce Type: replace Abstract: Modern vision-language models (VLMs) can act as generative OCR engines, yet open-ended decoding can expose rare but consequential failures. We ident

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation

ResearchDGX agent

arXiv:2604.13793v1 Announce Type: new Abstract: Exo-to-Ego video generation aims to synthesize a first-person video from a synchronized third-person view and corresponding camera poses. While paired s

Frozen Forecasting: A Unified Evaluation

ResearchDGX agent

arXiv:2507.13942v2 Announce Type: replace Abstract: Forecasting future events is a fundamental capability for general-purpose systems that plan or act across different levels of abstraction. Yet, eval

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation

Model ReleasesDGX agent

arXiv:2604.13803v1 Announce Type: new Abstract: Vision-language models are increasingly deployed in high-stakes settings, yet their susceptibility to sycophantic manipulation remains poorly understood

GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization

Model ReleasesDGX agent

arXiv:2512.02697v3 Announce Type: replace Abstract: Cross-view geo-localization infers a location by retrieving geo-tagged reference images that visually correspond to a query image. However, the trad

GeoLink: A 3D-Aware Framework Towards Better Generalization in Cross-View Geo-Localization

Local AiDGX agent

arXiv:2604.13183v1 Announce Type: new Abstract: Generalizable cross-view geo-localization aims to match the same location across views in unseen regions and conditions without GPS supervision. Its cor

Geometric Context Transformer for Streaming 3D Reconstruction

ResearchDGX agent

arXiv:2604.14141v1 Announce Type: new Abstract: Streaming 3D reconstruction aims to recover 3D information, such as camera poses and point clouds, from a video stream, which necessitates geometric acc

Getting the Numbers Rightnicode{x2014}Modelling Multi-Class Object Counting in Dense and Varied Scenes

ResearchDGX agent

arXiv:2510.02213v2 Announce Type: replace Abstract: Density map estimation enables accurate object counting in heavily occluded, and densely packed scenes where detection-based counting fails. In mult

Granularity-Aware Transfer for Tree Instance Segmentation in Synthetic and Real Forests

ResearchDGX agent

arXiv:2604.13722v1 Announce Type: new Abstract: We address the challenge of synthetic-to-real transfer in forestry perception where real data have only coarse Tree labels while synthetic data provide

Graph Propagated Projection Unlearning: A Unified Framework for Vision and Audio Discriminative Models

ResearchDGX agent

arXiv:2604.13127v1 Announce Type: new Abstract: The need to selectively and efficiently erase learned information from deep neural networks is becoming increasingly important for privacy, regulatory c

HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy

SafetyDGX agent

arXiv:2510.00695v3 Announce Type: replace-cross Abstract: Inherently, robotic manipulation tasks are history-dependent: leveraging past context could be beneficial. However, most existing Vision-Langu

Heavy-Tailed Class-Conditional Priors for Long-Tailed Generative Modeling

SafetyDGX agent

arXiv:2509.02154v2 Announce Type: replace-cross Abstract: Variational Autoencoders (VAEs) with global priors trained under an imbalanced empirical class distribution can lead to underrepresentation of

Heuristic Style Transfer for Real-Time, Efficient Weather Attribute Detection

ResearchDGX agent

arXiv:2604.13947v1 Announce Type: new Abstract: We present lightweight and efficient architectures to detect weather conditions from RGB images, predicting the weather type (sunny, rain, snow, fog) an

HiProto: Hierarchical Prototype Learning for Interpretable Object Detection Under Low-quality Conditions

ResearchDGX agent

arXiv:2604.13981v1 Announce Type: new Abstract: Interpretability is essential for deploying object detection systems in critical applications, especially under low-quality imaging conditions that degr

HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System

ApplicationsDGX agent

arXiv:2604.14125v1 Announce Type: new Abstract: While end-to-end Vision-Language-Action (VLA) models offer a promising paradigm for robotic manipulation, fine-tuning them on narrow control data often

← Previous
1…189190191192193…207
Next →