AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
12 May 2026

EMFormer: Efficient Multi-Scale Transformer for Accumulative Context Weather Forecasting

ResearchDGX agent

arXiv:2602.01194v2 Announce Type: replace Abstract: Long-term weather forecasting is critical for socioeconomic planning and disaster preparedness. While recent approaches employ finetuning to extend

EnergyLens: Interpretable Closed-Form Energy Models for Multimodal LLM Inference Serving

Model ReleasesDGX agent

arXiv:2605.10556v1 Announce Type: new Abstract: As large language models span dense, mixture-of-experts, and state-space architectures and are deployed on heterogeneous accelerators under increasingly

Enhancing Consistency Models for Multi-Agent Trajectory Prediction

AgentsDGX agent

arXiv:2605.08572v1 Announce Type: new Abstract: Diffusion models for multi-agent trajectory prediction are limited by iterative denoising, which causes inference latency that hinders their use in time


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Enhancing Few-Shot Out-of-Distribution Detection via the Refinement of Foreground and Background

ResearchDGX agent

arXiv:2601.15065v2 Announce Type: replace Abstract: CLIP-based foreground-background (FG-BG) decomposition methods have demonstrated remarkable effectiveness in improving few-shot out-of-distribution

ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning

ApplicationsDGX agent

arXiv:2605.09982v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) enable large language models (LLMs) to process high-resolution images, significantly improving real

Establishing Robust Retinal Eye Tracking: A Weakly Supervised Algorithmic Framework

ApplicationsDGX agent

arXiv:2605.09181v1 Announce Type: new Abstract: Retinal image-based eye tracking is widely used in ophthalmic imaging and vision science, and is a promising path to deliver higher gaze accuracy than t

Evidence-based Decision Modeling for Synthetic Face Detection with Uncertainty-driven Active Learning

ResearchDGX agent

arXiv:2605.09935v1 Announce Type: new Abstract: With the rapid development of deep generative models, forged facial images are massively exploited for illegal activities. Although existing synthetic f

Explanation-Aware Learning for Enhanced Interpretability in Biomedical Imaging

SafetyDGX agent

arXiv:2605.10054v1 Announce Type: new Abstract: Deep neural networks for medical image diagnosis often achieve high predictive accuracy while relying on spurious or clinically irrelevant visual cues,

Exploring 6D Object Pose Estimation with Deformation

ResearchDGX agent

arXiv:2604.06720v2 Announce Type: replace Abstract: We present DeSOPE, a large-scale dataset for 6DoF deformed objects. Most 6D object pose methods assume rigid or articulated objects, an assumption t

Exploring and Exploiting Stability in Latent Flow Matching

ResearchDGX agent

arXiv:2605.08398v1 Announce Type: cross Abstract: In this work, we show that Latent Flow-Matching (LFM) models are robust to different types of perturbations, including data reduction and model capaci

ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models

ResearchDGX agent

arXiv:2605.10045v1 Announce Type: new Abstract: Visual Autoregressive (VAR) models have emerged as a strong alternative to diffusion for image synthesis, yet their fixed training resolution prevents d

Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition

Model ReleasesDGX agent

arXiv:2605.10127v1 Announce Type: new Abstract: Recent research work on fashion outfit generation focuses on promoting visual consistency of garments by leveraging key information from reference image

Fetal Brain Imaging: A Composite Neural Network Approach for Keyframe Detection in Ultrasound Videos

ResearchDGX agent

arXiv:2605.09750v1 Announce Type: new Abstract: This article presents a novel approach to keyframe detection in ultrasound videos, with a particular focus on fetal brain imaging. The proposed model is

Few-Click-Driven Interactive 3D Segmentation with Semantic Embedding

SafetyDGX agent

arXiv:2605.08925v1 Announce Type: new Abstract: Interactive segmentation allows efficient label generation by leveraging user-provided clicks to progressively refine predictions, which is critical whe

Filtering Memorization from Parameter-Space in Diffusion Models

Model ReleasesDGX agent

arXiv:2605.10439v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become a widely used mechanism for customizing diffusion models, enabling users to inject new visual concepts or styles t

FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation

SafetyDGX agent

arXiv:2605.09430v1 Announce Type: new Abstract: Large-scale autoregressive models have demonstrated remarkable capabilities in image generation. However, their sequential raster-scan decoding relies o

FlashClear: Ultra-Fast Image Content Removal via Efficient Step Distillation and Feature Caching

Model ReleasesDGX agent

arXiv:2605.09003v1 Announce Type: new Abstract: Recently, diffusion-based object removal models have achieved impressive results in eliminating objects and their associated visual effects. However, th

FlowADMM: Plug-and-play ADMM with Flow-based Renoise-Denoise Priors

ResearchDGX agent

arXiv:2605.08640v1 Announce Type: new Abstract: Plug-and-play (PnP) methods for solving inverse problems have recently achieved strong performance by leveraging denoising priors based on powerful gene

Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models

HardwareDGX agent

arXiv:2605.09681v1 Announce Type: new Abstract: Autoregressive (AR) video diffusion models adopt a streaming generation framework, enabling long-horizon video generation with real-time responsiveness,

FPGA-Based Hardware Architecture for Contrast Maximization in Event-Based Vision

Model ReleasesDGX agent

arXiv:2605.09581v1 Announce Type: new Abstract: This paper presents a hardware architecture that implements the Contrast Maximization (CM) algorithm in Field-Programmable Gate Array (FPGA) resources f

FrameTwin: Curve-Anchored Gaussian Alignment from Sparse Views for Adaptive Wireframe 3D Printing

SafetyDGX agent

arXiv:2605.09362v1 Announce Type: cross Abstract: We present FrameTwin, a curve-anchored Gaussian alignment framework that uses sparse-view images to close the control loop for adaptive wireframe 3D p

Frequency Adapter with SAM for Generalized Medical Image Segmentation

SafetyDGX agent

arXiv:2605.09925v1 Announce Type: new Abstract: Medical image segmentation is a critical task in computer-aided diagnosis and treatment planning. However, deep learning models often struggle to genera

FrequencyCT: Frequency domain pseudo-label generation for self-supervised low-dose CT denoising

ApplicationsDGX agent

arXiv:2605.10583v1 Announce Type: new Abstract: Despite extensive research on computed tomography (CT) denoising, few studies exploit projection-domain data characteristics to mitigate noise correlati

From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation

Model ReleasesDGX agent

arXiv:2605.08712v1 Announce Type: new Abstract: Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimension

From Pixels to Concepts: Do Segmentation Models Understand What They Segment?

Model ReleasesDGX agent

arXiv:2605.09591v1 Announce Type: new Abstract: Segmentation is a fundamental vision task underlying numerous downstream applications. Recent promptable segmentation models, such as Segment Anything M

From pre-training to downstream performance: Does domain-specific pre-training make sense?

ResearchDGX agent

arXiv:2605.08819v1 Announce Type: new Abstract: Deep learning techniques have revolutionised medical imaging, improving diagnostic accuracy and enabling both more accurate and earlier disease detectio

FugSeg: Fast Uncertainty-aware Ground Segmentation for 3D Point Cloud

ResearchDGX agent

arXiv:2605.08952v1 Announce Type: new Abstract: In LiDAR-based environment perception systems, ground segmentation is a key preprocessing step supporting various applications such as mapping and navig

GemDepth: Geometry-Embedded Features for 3D-Consistent Video Depth

SafetyDGX agent

arXiv:2605.10525v1 Announce Type: new Abstract: Video depth estimation extends monocular prediction into the temporal domain to ensure coherence. However, existing methods often suffer from spatial bl

Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?

ResearchDGX agent

arXiv:2512.19115v2 Announce Type: replace Abstract: Despite the remarkable success of multimodal large language models (MLLMs) in generative tasks, we observe that they exhibit a counterintuitive defi

GenMed: A Pairwise Generative Reformulation of Medical Diagnostic Tasks

Model ReleasesDGX agent

arXiv:2605.10645v1 Announce Type: new Abstract: Data-driven medical AI is traditionally formulated as a discriminative mapping from input X to output Y via a learned function f, which does not general

Geometric Flood Depth Estimation: Fusing Transformer-Based Segmentation with Digital Elevation Models

ResearchDGX agent

arXiv:2605.08521v1 Announce Type: new Abstract: Post-disaster situational awareness relies heavily on understanding both the extent and the volume of floodwaters. While 2D semantic segmentation provid

Geometry-aware Prototype Learning for Cross-domain Few-shot Medical Image Segmentation

ResearchDGX agent

arXiv:2605.10885v1 Announce Type: new Abstract: Cross-domain few-shot medical image segmentation (CD-FSMIS) requires a model to generalise simultaneously to novel anatomical categories and unseen imag

GLEAM: A Multimodal Imaging Dataset and HAMM for Glaucoma Classification

ResearchDGX agent

arXiv:2603.12800v2 Announce Type: replace-cross Abstract: We propose glaucoma lesion evaluation and analysis with multimodal imaging (GLEAM), the first publicly available tri-modal glaucoma dataset co

GSMap: 2D Gaussians for Online HD Mapping

AgentsDGX agent

arXiv:2605.09619v1 Announce Type: new Abstract: Accurate High-Definition (HD) map construction is critical for autonomous driving, yet existing methods face a fundamental trade-off: vectorization-base

H-POPE: Hierarchical Polling-based Probing Evaluation of Hallucinations in Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2411.04077v2 Announce Type: replace Abstract: By leveraging both texts and images, large vision language models (LVLMs) have shown significant progress in various multi-modal tasks. Nevertheless

HairGPT: Strand-as-Language Autoregressive Modeling for Realistic 3D Hairstyle Synthesis

TutorialsDGX agent

arXiv:2605.08824v1 Announce Type: cross Abstract: Hair is a rich medium of visual and cultural expression, yet its digital modeling remains challenging due to the duality of fluidity and structure. Ma

Halo Separation-guided Underwater Multi-scale Image Restoration

AgentsDGX agent

arXiv:2605.10374v1 Announce Type: new Abstract: Underwater images captured by Autonomous Underwater Vehicles (AUVs) are inevitably affected by artificial light sources, which often produce halos in th

Harmonized Feature Conditioning and Frequency-Prompt Personalization for Multi-Rater Medical Segmentation

ResearchDGX agent

arXiv:2605.08210v1 Announce Type: new Abstract: Multi-rater medical image segmentation captures the inherent ambiguity of clinical interpretation, where diagnostic boundaries vary across experts and i

Heteroscedastic Diffusion for Multi-Agent Trajectory Modeling

AgentsDGX agent

arXiv:2605.10717v1 Announce Type: cross Abstract: Multi-agent trajectory modeling traditionally focuses on forecasting, often neglecting more general tasks like trajectory completion, which is essenti

HiDrive: A Closed-Loop Benchmark for High-Level Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.09972v1 Announce Type: cross Abstract: End-to-end autonomous driving has witnessed rapid progress, yet existing benchmarks are increasingly saturated, with state-of-the-art models achieving

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models

ResearchDGX agent

arXiv:2512.21815v2 Announce Type: replace Abstract: Vision-language models (VLMs) achieve remarkable performance but remain vulnerable to adversarial attacks. Entropy, as a measure of model uncertaint

HLGFA: High-Low Resolution Guided Feature Alignment for Unsupervised Anomaly Detection

SafetyDGX agent

arXiv:2602.09524v3 Announce Type: replace Abstract: Unsupervised industrial anomaly detection (UAD) is essential for modern manufacturing inspection, where defect samples are scarce and reliable detec

HPGN: Hybrid Priors-Guided Network for Compressed Low-Light Image Enhancement

TutorialsDGX agent

arXiv:2504.02373v3 Announce Type: replace-cross Abstract: In practical applications, low-light images are often compressed for efficient storage and transmission. Most existing methods disregard compr

Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation

Model ReleasesDGX agent

arXiv:2501.12202v4 Announce Type: replace Abstract: We present Hunyuan3D 2.0, an advanced large-scale 3D synthesis system for generating high-resolution textured 3D assets. This system includes two fo

HyNeuralMap: Hyperbolic Mapping of Visual Semantics to Neural Hierarchies

SafetyDGX agent

arXiv:2605.09392v1 Announce Type: new Abstract: Understanding the intricate mappings between visual stimuli and neural responses is a fundamental challenge in cognitive neuroscience. While current app

Hypergraph-Enhanced Training-Free and Language-Free Few-Shot Anomaly Detection

ResearchDGX agent

arXiv:2605.10628v1 Announce Type: new Abstract: Few-shot anomaly detection (FSAD) has made significant strides, yet existing methods still face critical challenges: (i) dependence on task- or dataset-

Hystar: Hypernetwork-driven Style-adaptive Retrieval via Dynamic SVD Modulation

Model ReleasesDGX agent

arXiv:2605.10009v1 Announce Type: new Abstract: Query-based image retrieval (QBIR) requires retrieving relevant images given diverse and often stylistically heterogeneous queries, such as sketches, ar

Illusion-Aware Visual Preprocessing and Anti-Illusion Prompting for Classic Illusion Understanding in Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.08841v1 Announce Type: new Abstract: Vision-Language Models (VLMs) exhibit systematic bias toward visual illusions, recalling memorized facts rather than perceiving actual visual difference

Improved Mean Flows: On the Challenges of Fastforward Generative Models

ResearchDGX agent

arXiv:2512.02012v2 Announce Type: replace Abstract: MeanFlow (MF) has recently been established as a framework for one-step generative modeling. However, its ``fastforward'' nature introduces key chal

Improving Generative Adversarial Networks with Self-Distillation

TutorialsDGX agent

arXiv:2605.08577v1 Announce Type: new Abstract: In modern GANs, maintaining an Exponential Moving Average (EMA) of the generator's weights is a standard practice, as such an averaged model consistentl

Improving Human Image Animation via Semantic Representation Alignment

SafetyDGX agent

arXiv:2605.10523v1 Announce Type: new Abstract: The field of image-to-video generation has made remarkable progress. However, challenges such as human limb twisting and facial distortion persist, espe

Improving Temporal Action Segmentation via Constraint-Aware Decoding

ResearchDGX agent

arXiv:2605.10149v1 Announce Type: new Abstract: Temporal action segmentation (TAS) divides untrimmed videos into labeled action segments. While fully supervised methods have advanced the field, challe

Increasing the Efficiency of DETR for Maritime High-Resolution Images

ResearchDGX agent

arXiv:2605.10269v1 Announce Type: new Abstract: Maritime object detection is critical for the safe navigation of unmanned surface vessels (USVs), requiring accurate recognition of obstacles from small

INFANiTE: Implicit Neural representation for high-resolution Fetal brain spatio-temporal Atlas learNing from clinical Thick-slicE MRI

ResearchDGX agent

arXiv:2605.09977v1 Announce Type: new Abstract: Spatio-temporal fetal brain atlases are important for characterizing normative neurodevelopment and identifying congenital anomalies. However, existing

Initiation of Interaction Detection Framework using a Nonverbal Cue for Human-Robot Interaction

Local AiDGX agent

arXiv:2605.10087v1 Announce Type: new Abstract: This paper describes an initiation of interaction(IoI) detection framework without keywords for human-robot interaction(HRI) based on audio and vision s

IPAD-CLIP: Teaching CLIP to Detect Image Local Perceptual Artifacts

Model ReleasesDGX agent

arXiv:2605.08664v1 Announce Type: new Abstract: Current image quality assessment methods are heavily biased towards global distortions (e.g., noise, blur), neglecting local perceptual artifacts such a

Is Class Signal Clustered or Routed in Task-Induced Implicit Neural Representation Weight Spaces?

SafetyDGX agent

arXiv:2605.08281v1 Announce Type: new Abstract: Implicit neural representations (INRs) encode images as neural-network weights, making image classification a problem of weight-space classifiability. A

Is Your Driving World Model an All-Around Player?

Model ReleasesDGX agent

arXiv:2605.10858v1 Announce Type: new Abstract: Today's driving world models can generate remarkably realistic dash-cam videos, yet no single model excels universally. Some generate photorealistic tex

JODA: Composable Joint Dynamics for Articulated Objects

Model ReleasesDGX agent

arXiv:2605.09954v1 Announce Type: cross Abstract: Articulated objects used in simulation and embodied AI are typically specified by geometry and kinematic structure, but lack the fine-grained dynamica

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion

ApplicationsDGX agent

arXiv:2601.22143v2 Announce Type: replace-cross Abstract: Audio-Visual Foundation Models, which are pretrained to jointly generate sound and visual content, have recently shown an unprecedented abilit

← Previous
1…143144145146147…209
Next →