AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
23 Jul 2026

ETPDesigner: Multi-Agent Orchestration for Interactive Multimodal Electronic Theater Program

Model ReleasesDGX agent

arXiv:2607.19947v1 Announce Type: new Abstract: Electronic Theater Programs (ETPs) serve as critical promotional media in the performing arts, comprising a multi-page collection of heterogeneous visua

Evolving Cache Schedules for Fast Diffusion Policy Inference

SafetyDGX agent

arXiv:2607.20293v1 Announce Type: new Abstract: Diffusion policies achieve strong visuomotor control by iteratively denoising action chunks, but repeated denoising makes real-time deployment computati

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.19341v1 Announce Type: new Abstract: Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven

Extending a Large View Synthesis Model for Multi-view Panoptic Segmentation

TutorialsDGX agent

arXiv:2607.19765v1 Announce Type: new Abstract: Large view synthesis models synthesize novel views through cross-view attention without explicit 3D representations, and recent studies have shown that

Factor-Informed Uncertainty Distillation for Gaze Estimation

ResearchDGX agent

arXiv:2607.20072v1 Announce Type: new Abstract: Deep gaze estimation works well in controlled capture but degrades in unconstrained settings, where systems must reject unreliable predictions. Single-p

FE-MCFormer: a novel time-frequency interpretable architecture for machinery fault diagnosis under strong noise environments

ApplicationsDGX agent

arXiv:2505.06285v3 Announce Type: replace-cross Abstract: Interpretable fault diagnosis (FD) plays a critical role in industrial manufacturing, as it improves human-machine understanding and operation

FedCC: A Low-Resource Federated Adaptation of Foundation Models for Robust Corpus Callosum localization in Fetal Ultrasound Images

Model ReleasesDGX agent

arXiv:2607.18283v1 Announce Type: cross Abstract: Accurate localization of the corpus callosum (CC) in fetal ultrasound (US) images is crucial for the early identification of neurodevelopmental abnorm

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling

Model ReleasesDGX agent

arXiv:2607.19038v1 Announce Type: new Abstract: Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-

FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility

ResearchDGX agent

arXiv:2607.19100v1 Announce Type: new Abstract: Reconstructing animatable 3D human avatars from monocular video is a fundamental problem in computer vision with broad applications in AR/VR and digital

Fluid-SDF: Ultra-Lightweight and Editable Implicit Shape Representation via Differentiable Primitives

Model ReleasesDGX agent

arXiv:2607.18646v1 Announce Type: new Abstract: Implicit Neural Representations (INRs) have become the standard for continuous 2D shape modeling, but they suffer from black-box uneditability, vulnerab

Forecasting the Number of Harvest-ready Fruits of Sweet Peppers Using Multimodal Time-Series Data

ApplicationsDGX agent

arXiv:2607.19975v1 Announce Type: new Abstract: Accurate yield forecasting at the individual-plant level is critical for precision agriculture and supply-chain planning, yet public datasets capturing

Frequency-Hierarchical Active k-Space Sampling for Diagnostic MRI

Model ReleasesDGX agent

arXiv:2607.19779v1 Announce Type: new Abstract: Active sampling for accelerated MRI must distribute a tight sampling budget across spatial frequencies that carry very different kinds of information. L

From Distances to Trajectories: Real-Time Signed Distance Function Mapping and Distance-Accelerated Motion Planning for UAVs

AgentsDGX agent

arXiv:2607.19306v1 Announce Type: cross Abstract: Autonomous flight in cluttered environments requires a robot to build a geometric map of its surroundings and plan safe, dynamically feasible trajecto

From Pixel to Prognosis: Convolutional and GLCM Feature Fusion for Automated Four-Class Cataract Severity Classification

HardwareDGX agent

arXiv:2607.18349v1 Announce Type: new Abstract: Objective: To develop a low-cost automated cataract severity classification system operating on standard consumer-grade colour photographs of the eye, w

GATE-3D: Geometry-Aware Test-time Adaptive Reranking for Open-Set 3D Shape Retrieval

Model ReleasesDGX agent

arXiv:2607.19111v1 Announce Type: new Abstract: Large pretrained vision models have substantially improved appearance-based 3D shape retrieval, but they still confuse shapes that look similar while di

GaussianSeed: Hierarchical Gaussian Seeding for High-Resolution 3D Occupancy Prediction

AgentsDGX agent

arXiv:2607.20071v1 Announce Type: new Abstract: Vision-centric 3D occupancy prediction provides dense scene representations essential for autonomous driving and robotic navigation, yet existing method

Gaze-DETR: Top-Down Guidance Through Priority Maps for Infrared Weak-Small UAV Detection with DETR

ResearchDGX agent

arXiv:2607.19040v1 Announce Type: new Abstract: Infrared small target detection (ISTD) remains challenging because tiny, low-contrast targets are easily overwhelmed by clutter, noise, or occlusion. Co

Generalized Least Squares Kernelized Tensor Factorization

Local AiDGX agent

arXiv:2412.07041v4 Announce Type: replace-cross Abstract: Recovering incomplete multidimensional tensor-structured data is a fundamental task in many real-world applications. Smoothness-constrained lo

Generative World Renderer at the Speed of Play

ResearchDGX agent

arXiv:2607.18703v1 Announce Type: new Abstract: Generative world renderer AlayaRenderer receives structured world states exported from physics engines and synthesizes RGB frames. Unlike models that ge

GeoTrace: Geometry-Aware Trajectory Token Compression for Video Large Language Models

Local AiDGX agent

arXiv:2607.09080v2 Announce Type: replace Abstract: Although Video Large Language Models (Video LLMs) have shown strong performance in video understanding, their efficiency is still limited by the lar

GH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision Tasks

Model ReleasesDGX agent

arXiv:2512.24592v3 Announce Type: replace Abstract: Systematic failures of vision models on semantically coherent subsets, known as error slices, reveal limitations in robustness and evaluation. Exist

Glass Surface Detection: Leveraging Reflection Dynamics in Flash/No-flash Imagery

Local AiDGX agent

arXiv:2511.16887v5 Announce Type: replace Abstract: Glass surfaces are ubiquitous in daily life, typically appearing colorless, transparent, and lacking distinctive features. These characteristics mak

GLID: Gated Local Intrinsic Dimension Repairs the Blind Spots of Face-Forgery Detectors

Model ReleasesDGX agent

arXiv:2607.18770v1 Announce Type: cross Abstract: Fine-tuned foundation-model detectors dominate face-forgery benchmarks, yet they stay blind to generator families absent from training. We present GLI

Global Building Area Estimation Products: How Accurate Are They?

SafetyDGX agent

arXiv:2607.19766v1 Announce Type: new Abstract: Geo-spatial rasters of building footprint area are useful for a variety of tasks, such as monitoring urbanization, improving energy efficiency, and trac

Great X: A Unified Multi-Modal Simulator Bridging the Sim2Real Gap for 6G

SafetyDGX agent

arXiv:2507.08716v4 Announce Type: replace Abstract: Large-scale, precisely synchronized multi-modal datasets are critical for data-driven sixth-generation (6G) wireless research, yet real-world collec

Group-of-Latents: Perceptual Video Compression at Extreme Bitrates via Masked Latent Generative Modeling

ResearchDGX agent

arXiv:2607.19437v1 Announce Type: cross Abstract: Most existing video compression algorithms follow a paradigm of transformation and quantization, optimizing the trade-off between distortion and bitra

Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies

SafetyDGX agent

arXiv:2607.18325v1 Announce Type: new Abstract: Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and support decision-making during emergencies. Visi

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

SafetyDGX agent

arXiv:2607.18217v2 Announce Type: replace Abstract: Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two

How Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing Inspection

ResearchDGX agent

arXiv:2607.20263v1 Announce Type: new Abstract: Housing-level urban physical examination is essential for identifying residential building problems and supporting targeted urban renewal. Existing auto

IBoxCLA: Towards Robust Box-supervised Segmentation of Polyp via Improved Box-dice and Contrastive Latent-anchors

Model ReleasesDGX agent

arXiv:2310.07248v5 Announce Type: replace Abstract: Box-supervised polyp segmentation attracts increasing attention for its cost-effective potential. Existing solutions often rely on learning-free met

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

ApplicationsDGX agent

arXiv:2607.19228v1 Announce Type: new Abstract: Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear

Image Editing Models are Numerical Solvers

ResearchDGX agent

arXiv:2607.18787v1 Announce Type: new Abstract: We investigate whether a pretrained generative image-editing model can provide a common interface for numerical simulation. Physical inputs and solution

IMMoE: Incomplete Multi-View Anomaly Detection via Mixture of View Experts Fusion

Local AiDGX agent

arXiv:2607.19032v1 Announce Type: new Abstract: Existing Multi-view Anomaly Detection (MAD) methods assume that all views are completely available and model each view separately. However, in real indu

Importance-Aware OBS Pruning for Diffusion Models

Model ReleasesDGX agent

arXiv:2607.20048v1 Announce Type: new Abstract: We propose importance-aware pruning for diffusion models, a training-free framework that prioritizes preserving parameters critical to semantically sali

In-Context Learning for Wound Classification with Small Multimodal Language Models

Model ReleasesDGX agent

arXiv:2607.18819v1 Announce Type: new Abstract: Wound image classification is often treated as a task-specific supervised learning problem, requiring substantial amounts of manually labelled data and

InstantSfM: Towards GPU-Native SfM for the Deep Learning Era

HardwareDGX agent

arXiv:2510.13310v3 Announce Type: replace Abstract: Structure-from-Motion (SfM) is a fundamental technique for recovering camera poses and scene structure from multi-view imagery, serving as a critica

InstructMixup: Instruction-Guided Salient Patch Editing for Robust Data Augmentation

ResearchDGX agent

arXiv:2607.19324v1 Announce Type: new Abstract: In image and video technologies, data augmentation is widely used to improve the generalization of deep visual models, and mixup-based strategies that i

Interactive Medical-SAM2 GUI: A Napari-based semi-automatic annotation tool for medical images

Model ReleasesDGX agent

arXiv:2602.22649v2 Announce Type: replace Abstract: Interactive Medical-SAM2 GUI is an open-source desktop application for semi-automatic annotation of 2D and 3D medical images. Built on the Napari mu

Internet-of-Things Architectures for Secure Cyber-Physical Spaces: the VISOR Experience Report

ResearchDGX agent

arXiv:2204.01531v2 Announce Type: replace-cross Abstract: Internet of things (IoT) technologies are becoming a more and more widespread part of civilian life in common urban spaces, which are rapidly

It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability

Model ReleasesDGX agent

arXiv:2607.16292v2 Announce Type: replace Abstract: Brain-encoding foundation models predict fMRI responses to video, audio, and text well enough to win the Algonauts 2025 challenge. We ask whether th

Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing

ResearchDGX agent

arXiv:2510.08532v2 Announce Type: replace Abstract: Instruction-based image editing offers a powerful and intuitive way to manipulate images through natural language. Yet, relying solely on text instr

Latent Riemannian Flow Matching for Geometry-Grounded 3D Foundation Models

ResearchDGX agent

arXiv:2607.19120v1 Announce Type: new Abstract: Geometric foundation models, such as the Visual Geometry Grounded Transformer (VGGT), provide strong 3D priors from unposed images. However, such models

LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition

Local AiDGX agent

arXiv:2607.19889v1 Announce Type: new Abstract: Understanding instrument-tissue interactions is essential for context-aware surgical AI and autonomous robotic surgery. Pretrained vision-language model

LC-SLab -- An object-based deep learning framework for large-scale land cover classification from satellite imagery and sparse in-situ labels

ResearchDGX agent

arXiv:2509.15868v2 Announce Type: replace Abstract: Large-scale land cover maps generated using deep learning play a critical role across a wide range of Earth science applications. Open in-situ datas

Lean-SAM2: Target-Anchored Memory and Encoder Acceleration for SAM2

SafetyDGX agent

arXiv:2607.19811v1 Announce Type: new Abstract: The Segment Anything Model 2 (SAM2) has advanced temporal promptable segmentation, yet its deployment remains hindered by heavy memory cross-attention o

Learning Encoding-Decoding Direction Pairs to Unveil Concepts of Influence in Deep Vision Networks

SafetyDGX agent

arXiv:2509.23926v4 Announce Type: replace Abstract: Empirical evidence shows that deep vision networks often represent concepts as directions in latent space with concept information written along dir

Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

Model ReleasesDGX agent

arXiv:2607.18924v1 Announce Type: new Abstract: Recent advances in image-to-video generation have improved visual realism, making physically grounded and controllable dynamics an important step toward

Learning Semantic-Robust Change Detection via Semantic-Invariant Self-Distillation

TutorialsDGX agent

arXiv:2607.19000v1 Announce Type: new Abstract: Change detection aims to identify semantic changes between remote sensing images. However, features from models are easily disturbed by non-semantic var

Local Label-Informed Feature Transfer for Generating Ground-Truth Medical Images: A Comparison of GAN- and Diffusion-Based Approaches

Model ReleasesDGX agent

arXiv:2607.18882v1 Announce Type: new Abstract: Validating Explainable Artificial Intelligence (XAI) methods in medical imaging requires ground-truth data with known locations of informative features.

Look Before You Edit: Attention-Guided Camera Placement and Multi-View Alignment for 3D Gaussian Splatting Editing

Local AiDGX agent

arXiv:2607.19777v1 Announce Type: new Abstract: Text-driven 3D scene editing with 3D Gaussian Splatting (3DGS) typically applies a 2D diffusion editor to views rendered from fixed training cameras, li

Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs

TutorialsDGX agent

arXiv:2607.20357v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance across vision-language tasks. However, their high inference cost,

LoRFT: Benchmarking Long-Range Vehicle Trajectory Reconstruction from Fixed Highway Cameras

Model ReleasesDGX agent

arXiv:2607.19911v1 Announce Type: new Abstract: Long-range vehicle trajectories provide important spatio-temporal evidence for traffic safety analysis, autonomous driving evaluation, and data-driven t

MagicPrompt: Ultra-Lightweight Prompt Tuning for Video Generation

Model ReleasesDGX agent

arXiv:2607.14595v2 Announce Type: replace Abstract: Large-scale video diffusion models deliver strong generation performance, but full fine-tuning for downstream tasks incurs prohibitive computational

Masked Visual Actions for Unified World Modeling

SafetyDGX agent

arXiv:2607.19343v1 Announce Type: new Abstract: Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

SafetyDGX agent

arXiv:2607.15273v2 Announce Type: replace Abstract: MeanFlow generators achieve fast few-step sampling by predicting average velocities over time intervals, making them attractive for efficient genera

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

Model ReleasesDGX agent

arXiv:2607.19235v1 Announce Type: cross Abstract: Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challen

MissingBench-Verified: Probing Vision-Language Models' Inability to Detect Missing Object Parts

Model ReleasesDGX agent

arXiv:2607.18673v2 Announce Type: new Abstract: Vision Language Models (VLMs) are well known for hallucinating non-existent objects in images. Objects with missing parts present a unique challenge for

Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval

SafetyDGX agent

arXiv:2607.19027v1 Announce Type: new Abstract: Zero-shot video moment retrieval aims to overcome the limitations of traditional approaches that require large-scale datasets annotated with text and it

MoAKE: Toward Unified All-in-One Action Quality Assessment via Mixture of Action Knowledge Experts

ApplicationsDGX agent

arXiv:2607.19826v1 Announce Type: new Abstract: Action Quality Assessment (AQA) aims to objectively evaluate performance quality from action videos. Most existing methods follow a ``one-by-one'' parad

MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement

Local AiDGX agent

arXiv:2607.17967v2 Announce Type: replace Abstract: Monocular geometry estimation has recently achieved impressive performance across diverse scenes. However, state-of-the-art models still face notabl

← Previous
1…3637383940…209
Next →