AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
10 Jul 2026

HoloTetSphere: Unified TetSphere Mesh Reconstruction for Physical Simulations

ResearchDGX agent

arXiv:2607.08398v1 Announce Type: cross Abstract: Standard pipelines for physics-ready 3D reconstruction rely on a decoupled two-stage paradigm: extracting surface geometry followed by an error-prone

HSA: Hierarchical Slot Attention for Multi-granularity Scene-Decomposition

SafetyDGX agent

arXiv:2607.08249v1 Announce Type: new Abstract: Slot attention is a powerful framework for object-centric learning, decomposing visual scenes into latent slots through iterative competitive attention.

HumanForge: A Human-Centric Deepfake Video Benchmark with Multi-Agent Forgery Rationales

Model ReleasesDGX agent

arXiv:2607.08705v1 Announce Type: new Abstract: Rapid advancements in video diffusion models and temporal editing tools have enabled the generation of highly realistic human-centric videos, posing unp


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning

Model ReleasesDGX agent

arXiv:2603.12478v2 Announce Type: replace Abstract: Multimodal instruction tuning is often compute-inefficient because training budgets are spread across large mixed image-video pools whose utility is

LESV: Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding

SafetyDGX agent

arXiv:2604.01388v2 Announce Type: replace Abstract: Recent advancements in open-vocabulary 3D scene understanding heavily rely on 3D Gaussian Splatting (3DGS) to register vision-language features into

LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting

Model ReleasesDGX agent

arXiv:2607.08016v1 Announce Type: new Abstract: Video relighting requires balancing long-form temporal consistency with a physically grounded understanding of light transport, which depends on accurat

LlamaSeg: Image Segmentation via Autoregressive Mask Generation

Model ReleasesDGX agent

arXiv:2505.19422v2 Announce Type: replace Abstract: We present extbf{LlamaSeg}, a visual autoregressive framework that unifies multiple image segmentation tasks via natural language instructions. By r

LOGOS: Language-guided Oriented Object Detection in Aerial Scenes

TutorialsDGX agent

arXiv:2607.08004v1 Announce Type: new Abstract: Object detection in geospatial scenes, such as satellite and aerial imagery, poses significant challenges due to the varying orientations and densities

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models

SafetyDGX agent

arXiv:2607.08770v1 Announce Type: new Abstract: Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models

LTM: Large-scale Terrain Model for Wildfire-prone Landscapes

SafetyDGX agent

arXiv:2607.08711v1 Announce Type: new Abstract: Accurate 3D terrain maps are essential for emergency response when assessing wildfire hazards. However, wildfire-prone regions often span vast areas whe

LUMI: Tokenizer-Agnostic LLM-Based Lossless Image Compression

Model ReleasesDGX agent

arXiv:2607.08221v1 Announce Type: new Abstract: Large language model (LLM)-based lossless image compression methods typically represent pixel data through the native text interface of a pretrained mod

Metrics or Mirage? An Audit of Evaluation Inconsistencies in Colonoscopy Polyp Segmentation Benchmarks

ResearchDGX agent

arXiv:2607.08203v1 Announce Type: new Abstract: Progress in colonoscopy polyp segmentation is routinely reported through leaderboard comparisons on a small set of public benchmarks. We argue that this

Mixture of Enhanced-View Experts for Multi-Query Vehicle ReID and A Large-Scale Benchmark

Model ReleasesDGX agent

arXiv:2607.08085v1 Announce Type: new Abstract: Multi-query vehicle ReID aims to leverage complementary information from diverse views for robust feature learning. However, current methods suffer from

Multi-Resolution Feature Stem for Diabetic Retinopathy lesion segmentation

Model ReleasesDGX agent

arXiv:2607.08679v1 Announce Type: new Abstract: Diabetic Retinopathy (DR) is a leading cause of preventable blindness worldwide, requiring automated lesion segmentation using deep learning models for

Multimodal 3D LUT Generation via StatLUT with Statistical Features for Photorealistic Style Transfer

ResearchDGX agent

arXiv:2607.08227v1 Announce Type: new Abstract: Photorealistic Style Transfer (PST) aims to transfer the color and tonal style of a reference to a content image while strictly preserving its structura

Native Video-Action Pretraining for Generalizable Robot Control

SafetyDGX agent

arXiv:2607.08639v1 Announce Type: cross Abstract: The advent of video-action models offers a promising path for robot control. Nevertheless, we argue that repurposing video generative models designed

OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing

Local AiDGX agent

arXiv:2606.15920v2 Announce Type: replace Abstract: Reinforcement learning for multimodal large language models (MLLMs) is often hindered by severe reward sparsity in complex reasoning tasks. This cha

On the Design of Mixture-of-Experts for Dynamic Gaussian Splatting

ApplicationsDGX agent

arXiv:2607.08250v1 Announce Type: new Abstract: Dynamic scene reconstruction remains challenging due to the heterogeneous and spatially varying nature of real-world motion. Although recent 3D Gaussian

OPSD-V: On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators

SafetyDGX agent

arXiv:2607.08766v1 Announce Type: new Abstract: We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR v

PhyMAGIC: Physical Motion-Aware Generative Inference with Confidence-guided LLM

SafetyDGX agent

arXiv:2505.16456v3 Announce Type: replace Abstract: Recent advances in 3D content generation have amplified demand for dynamic models that are both visually realistic and physically consistent. Howeve

Post-Training in End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2607.08072v1 Announce Type: new Abstract: End-to-end models that map multimodal inputs directly to future trajectories/maneuvers have emerged as an increasingly prominent research paradigm in au

Predicting Viticulture Potential through an Ensemble of U-Net and a Geospatial Foundation Model

ResearchDGX agent

arXiv:2607.08449v1 Announce Type: new Abstract: Determining agricultural potential is fundamental to sustainable land management and agricultural planning. Remote sensing data is increasingly valuable

PRGCN: A Graph Memory Network for Cross-Sequence Pattern Reuse in 3D Human Pose Estimation

ResearchDGX agent

arXiv:2510.19475v2 Announce Type: replace Abstract: Monocular 3D human pose estimation remains a fundamentally ill-posed inverse problem due to the inherent depth ambiguity in 2D-to-3D lifting. While

Progression as Latent Drift: Generative Forecasting of Slow-Evolving Pathologies

Local AiDGX agent

arXiv:2607.08270v1 Announce Type: new Abstract: Forecasting the future anatomy of slow-evolving neurodegenerative diseases could enable earlier, more targeted intervention and improve clinical trial d

Real-World Blind Super-Resolution via Feature Matching with Implicit High-Resolution Priors

Model ReleasesDGX agent

arXiv:2202.13142v3 Announce Type: replace Abstract: A key challenge of real-world image super-resolution (SR) is to recover the missing details in low-resolution (LR) images with complex unknown degra

SAGA: Stable Acceleration Guidance for Autoregressive Video Generation

ResearchDGX agent

arXiv:2607.08020v1 Announce Type: new Abstract: Autoregressive video diffusion enables efficient streaming and long-horizon video generation, but repeatedly reusing generated latents as causal context

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation

ResearchDGX agent

arXiv:2607.08688v1 Announce Type: new Abstract: Modern Video Object Segmentation (VOS) involves tracking and segmenting user-specified targets. While recent approaches have achieved remarkable perform

SASGeo: Stability-Aware Semantic Map Localization for GNSS-Denied UAVs -- A Framework and Synthetic Proof of Concept

SafetyDGX agent

arXiv:2607.07737v1 Announce Type: cross Abstract: GNSS-denied unmanned aerial vehicles require occasional absolute position fixes to bound the drift of visual-inertial odometry. Cross-view image retri

Search-based Testing of Vision Language Models for In-Car Scene Understanding

SafetyDGX agent

arXiv:2607.02300v2 Announce Type: replace Abstract: In the automotive domain, in-car scene understanding (ISU) enables the detection of safety-critical events, such as driver distraction, and supports

SkelGen4D: Weakly-Supervised Skeleton-Based 4D Generation for Text-Driven Mesh Animation

ApplicationsDGX agent

arXiv:2607.08246v1 Announce Type: new Abstract: We study 4D generation to synthesize temporally coherent sequences of 3D geometry for animation and content creation. In contrast to existing SDS-based

SPHINX: First Explain, Then Explore

SafetyDGX agent

arXiv:2606.17482v2 Announce Type: replace Abstract: Generating adversarial driving scenarios is critical for evaluating and improving autonomous vehicle decision-making systems in simulation. Recent a

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

ApplicationsDGX agent

arXiv:2507.05240v2 Announce Type: replace-cross Abstract: Vision-and-Language Navigation (VLN) in real-world settings requires agents to process continuous visual streams and generate actions with low

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning

TutorialsDGX agent

arXiv:2607.08572v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) often follow a fixed Think-then-Answer paradigm, which is inefficient in heterogeneous multitask settings becau

Systematic Evaluation of Learning Rate Scheduling Strategies Across Heterogeneous Architectures

ResearchDGX agent

arXiv:2607.08511v1 Announce Type: cross Abstract: Choosing a learning rate scheduling strategy is critical to neural network training, but manual selection is costly and rarely exhaustive. While class

Texture Representations in Deep Vision Models: Comparing CNNs, Vision Transformers, and Human Perception

SafetyDGX agent

arXiv:2607.08321v1 Announce Type: new Abstract: In computational vision science, Convolutional Neural Networks (CNNs) have emerged as a popular model of biological vision because of the alignment they

TrackStudio: An Integrated Toolkit for Markerless Tracking

TutorialsDGX agent

arXiv:2511.07624v3 Announce Type: replace Abstract: Markerless motion tracking has advanced rapidly in the past 10 years and currently offers powerful opportunities for behavioural, clinical, and biom

Transformed ell_1 Gradient Regularization for Image Denoising

ResearchDGX agent

arXiv:2511.15060v2 Announce Type: replace-cross Abstract: Total variation (TV) regularization is a classical edge-preserving technique widely used across image recovery and reconstruction problems; ho

TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading

Model ReleasesDGX agent

arXiv:2607.08236v1 Announce Type: new Abstract: Event-based lip reading has recently emerged as a promising direction for visual speech recognition, benefiting from the high temporal resolution and mo

UAV-OVVIS: Unmanned Aerial Vehicles Also Need Open-Vocabulary Video Instance Segmentation

Model ReleasesDGX agent

arXiv:2607.08075v1 Announce Type: new Abstract: Unmanned Aerial Vehicle (UAV) videos are widely used in traffic monitoring, urban management, and emergency rescue. However, existing UAV video percepti

Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio

Model ReleasesDGX agent

arXiv:2607.08127v1 Announce Type: new Abstract: Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these p

Unified Face Attack Detection via Fine-Grained Semantic Guidance

SafetyDGX agent

arXiv:2607.08156v1 Announce Type: new Abstract: The growing applications of facial recognition systems are accompanied by increasingly diverse security threats. Existing datasets lack detailed textual

UniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery

Model ReleasesDGX agent

arXiv:2607.08267v1 Announce Type: new Abstract: Unmanned aerial vehicles (UAVs) increasingly rely on visual grounding capabilities to localize task-relevant targets from diverse instructions in comple

Unpaired Joint Distribution Modeling via Multi-Scale Image Representations

ApplicationsDGX agent

arXiv:2607.08198v1 Announce Type: new Abstract: This paper studies the problem of learning a joint distribution from marginal observations, which is inherently ill-posed due to the ambiguity of feasib

Vision-Language Memory for Spatial Reasoning

ResearchDGX agent

arXiv:2511.20644v2 Announce Type: replace Abstract: Spatial reasoning is a critical capability for intelligent robots, yet current vision-language models (VLMs) still fall short of human-level perform

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness

Model ReleasesDGX agent

arXiv:2607.08112v1 Announce Type: new Abstract: We introduce VSRo-200, the first large-scale dataset for visual speech recognition (lip reading) in Romanian, comprising 200 hours of real-world podcast

WaspMOT: A Benchmark for Long-Term Multi-Object Tracking of Trichogramma Wasps

Model ReleasesDGX agent

arXiv:2607.08729v1 Announce Type: new Abstract: Multi-object tracking (MOT) has achieved strong performance on benchmarks dominated by short video sequences. However, such datasets do not adequately e

Wat3R: Underwater 3D Geometry Learning without Annotations

ResearchDGX agent

arXiv:2607.08772v1 Announce Type: new Abstract: Estimating 3D geometry in underwater environments presents unique challenges due to light attenuation, scattering, and the absence of large-scale, high-

Whareformer: Learning to Track What is Where in Long Egocentric Videos

ResearchDGX agent

arXiv:2607.08537v1 Announce Type: new Abstract: The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wea

Who Gets Missed in the Tail? Thresholded Subgroup Underdiagnosis in Long-Tailed Chest X-ray Classification

SafetyDGX agent

arXiv:2607.07717v1 Announce Type: cross Abstract: In chest X-ray (CXR) classification, acceptable ranking performance can still leave rare-positive patients below threshold, especially within subgroup

XOV-Action: Towards Generalizable Open-Vocabulary Action Recognition

Model ReleasesDGX agent

arXiv:2403.01560v3 Announce Type: replace Abstract: Inspired by the impressive success of image-text foundation models, recent works have proposed to adapt these foundation models to video data, leadi

ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device

ResearchDGX agent

arXiv:2607.08771v1 Announce Type: new Abstract: Monocular depth estimation has seen remarkable progress through foundation models achieving robust zero-shot generalization, yet their computational dem

Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning

SafetyDGX agent

arXiv:2601.02918v3 Announce Type: replace Abstract: Image Quality Assessment (IQA) is a long-standing problem in computer vision. Previous methods typically focus on predicting numerical scores withou

9 Jul 2026

A Good Initialization is All You Need for Faithful Visual Attribution

ResearchDGX agent

arXiv:2607.06726v1 Announce Type: new Abstract: Faithful visual attribution identifies which image regions support a model prediction. Search-based perturbation methods lead the insertion--deletion fa

A Theory of Contrastive Learning with Natural Images

TutorialsDGX agent

arXiv:2607.07470v1 Announce Type: new Abstract: Why does contrastive learning with simple images and augmentations yield useful representations for downstream tasks? We address this question by analyt

AA-ViT: Anatomically Aware Vision Transformer with Structural and Frequency Guidance for Contrast Enhanced Brain MRI Synthesis

Local AiDGX agent

arXiv:2607.07553v1 Announce Type: new Abstract: Accurate tumour localization and diagnosis is a critical component of clinical care for brain cancers. Magnetic Resonance Imaging (MRI) is the most comm

Activation Quantization of Vision Encoders Needs Prefixing Registers

Local AiDGX agent

arXiv:2510.04547v5 Announce Type: replace-cross Abstract: Large pretrained vision encoders are central to multimodal intelligence, powering applications from on-device vision processing to vision-lang

AI for Cultural Heritage Textiles: Fine-Tuned Latent Diffusion for Novel Ulos Motif Synthesis

SafetyDGX agent

arXiv:2607.06590v1 Announce Type: new Abstract: Preserving and revitalising traditional textiles such as Ulos, a cultural heritage of the Batak ethnic group in North Sumatra, Indonesia, requires balan

An Edge-aware Prompt-enhanced SAM for Ultrasound Image Segmentation

TutorialsDGX agent

arXiv:2607.07240v1 Announce Type: new Abstract: Ultrasound image segmentation is essential for delineating anatomical structures and lesions, providing the foundation for accurate diagnosis. While the

ASFR-Net: Adversarial Alignment and Spatio-Frequency Refinement Network for Heterogeneous Remote Sensing Image Change Detection

Model ReleasesDGX agent

arXiv:2607.07161v1 Announce Type: new Abstract: The core challenge of heterogeneous change detection in remote sensing imagery lies in effectively decoupling genuine land-cover changes from significan

`Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation

SafetyDGX agent

arXiv:2607.07230v1 Announce Type: new Abstract: Video object segmentation (VOS) is a fundamental task in video understanding, requiring accurate delineation and consistent tracking of objects across f

← Previous
1…4344454647…209
Next →