AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
11 Aug 2026

JUMP-lite: Compact, reproducible benchmarking of cell representations

Model ReleasesDGX agent

arXiv:2608.07632v1 Announce Type: cross Abstract: Image-based profiling captures rich phenotypic signatures for drug discovery and functional genomics. Large public datasets like JUMP Cell Painting no

LAD-COD: Language-Aligned Dense Perception for Camouflaged Object Detection

TutorialsDGX agent

arXiv:2608.07941v1 Announce Type: new Abstract: Camouflaged object detection (COD) aims to segment objects that exhibit high visual similarity to their surroundings, which reduces foreground-backgroun

LASA: Language-and-Source-Anchored Alignment for Domain Generalized Semantic Segmentation

SafetyDGX agent

arXiv:2608.08805v1 Announce Type: new Abstract: Domain Generalization Semantic Segmentation (DGSS) focuses on generalizing knowledge from labeled source domains to unseen target domains where data is


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents

AgentsDGX agent

arXiv:2608.07585v1 Announce Type: new Abstract: Long-video understanding requires models to efficiently acquire and reuse sparse visual evidence from long and redundant video streams. Recent video too

Learning Deep Modality-Shared Self-Expressiveness for Image Clustering with Textual Information

SafetyDGX agent

arXiv:2608.08418v1 Announce Type: new Abstract: Leveraging textual information for image clustering has emerged as a promising direction, largely owing to the powerful representations learned by Visio

Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning

Model ReleasesDGX agent

arXiv:2608.09926v1 Announce Type: new Abstract: The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely fit the pixels without modeling how

Learning human joint torques from pixels

Model ReleasesDGX agent

arXiv:2608.09083v1 Announce Type: new Abstract: Estimating human joint torques from visual observations is a key step toward bringing biomechanical analysis from controlled laboratories to real-world

Learning Physical Interaction: A Survey of Tactile- and Force-aware Robot Learning

ResearchDGX agent

arXiv:2608.07558v1 Announce Type: cross Abstract: Physically grounded robot intelligence requires robots to perceive, reason about, and regulate their interactions with the physical world. This capabi

Learning Structural Illumination for Unsupervised Low-light Enhancement

Model ReleasesDGX agent

arXiv:2608.08153v1 Announce Type: new Abstract: Existing unsupervised low-light image enhancement (LLIE) methods often estimate illumination directly from the entire low-light input, without separatin

Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs

TutorialsDGX agent

arXiv:2604.05695v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress in 2D visual tasks but still struggle to understand physical space in rea

LHSDet: High-Resolution AI-Generated Image Detection via Visual Question Answering

ResearchDGX agent

arXiv:2608.07863v1 Announce Type: new Abstract: Driven by advances in diffusion models and autoregressive models, the fidelity and resolution of AI-generated images now rival those of real images. How

LIBAD: A Multimodal Anomaly Detection Benchmark for Li-Ion Battery Electrode Manufacturing

Model ReleasesDGX agent

arXiv:2608.07958v1 Announce Type: new Abstract: Multimodal industrial anomaly detection has largely focused on discrete products using strongly correlated RGB and 3D observations, leaving continuous p

LightAIR: Lightweight Action Inversion and Riemannian Rectification for Text-based Person Anomaly Search

ResearchDGX agent

arXiv:2608.09152v1 Announce Type: new Abstract: Traditional Text-based Person Search (TPS) is typically limited to matching static appearance attributes, severely neglecting dynamic action information

LineGraph2Road: Structural Graph Reasoning on Line Graphs for Road Network Extraction

ApplicationsDGX agent

arXiv:2602.23290v2 Announce Type: replace Abstract: Extracting routable road networks from satellite imagery requires accurate topology recovery beyond pixel-level segmentation. Recent methods decompo

Linguistically-Aligned and Visually-Grounded Preference Optimization for Clinically-Augmented Medical Report Generation

SafetyDGX agent

arXiv:2608.08494v1 Announce Type: new Abstract: Despite significant advances in Medical Report Generation (MRG), the reliability remains constrained by the prevalence of factual errors. While Direct P

LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding

Model ReleasesDGX agent

arXiv:2608.07596v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that

LogiShot: Logically Coherent Cross-Shot Video Generation

Model ReleasesDGX agent

arXiv:2608.08820v1 Announce Type: new Abstract: Generating cross-shot videos that are logically connected is essential for content creation. Currently, most cross-shot video-generation workflows, such

LookAgain: Closed-Loop GUI Grounding with Visually Grounded Reflection

ResearchDGX agent

arXiv:2608.09723v1 Announce Type: new Abstract: Recent graphical user interface (GUI) grounders have significantly advanced single-shot accuracy on standard benchmarks, yet their performance degrades

LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection

ResearchDGX agent

arXiv:2608.09633v1 Announce Type: new Abstract: Face presentation attack detection (PAD) aims to reliably detect a wide range of presentation attacks. While PAD methods achieve strong performance with

Lying mirror using structured surfaces

ResearchDGX agent

arXiv:2410.15521v2 Announce Type: replace-cross Abstract: We introduce an all-optical system, termed the 'lying mirror', to hide input information by transforming it into misleading, ordinary-looking

MAGIC-SSCIL: Manifold Anchoring and Geometric Incremental Calibration for Semi-Supervised Class Incremental Learning

SafetyDGX agent

arXiv:2608.07586v1 Announce Type: new Abstract: Semi-supervised Class Incremental Learning (SSCIL) is a severe challenge for neural networks, and it is hardest in the exemplar-free setting where no pa

Marrying Optimal Transport and ODEs for Unified Continuous-Time 4D Reconstruction and Tracking

Model ReleasesDGX agent

arXiv:2608.09613v1 Announce Type: new Abstract: Existing unified 4D reconstruction and point tracking approaches typically rely on heuristic interpolations or just predict at integer timestamps, lacki

Mask-aware inference with State-Space Models

ApplicationsDGX agent

arXiv:2603.04568v2 Announce Type: replace Abstract: Many real-world computer vision tasks, such as depth completion, must handle inputs with arbitrarily shaped regions of missing or invalid data. For

MeanSR: Restoration Trajectory Learning for One-Step Perceptual Super-Resolution

ApplicationsDGX agent

arXiv:2608.09405v1 Announce Type: new Abstract: Diffusion-based super-resolution (SR) achieves strong perceptual quality but requires costly iterative denoising. Existing one-step distillation methods

Mechanistic Interpretability-Guided Selective Fine-Tuning of Vision-Language Models for Centimeter-Level Flood Depth Estimation

Model ReleasesDGX agent

arXiv:2608.07562v1 Announce Type: new Abstract: Urban flooding poses an escalating threat to transportation infrastructure, yet no operational system provides real-time, street-level flood-depth estim

MemeMind: Reference-Guided Trace Construction for Offline Context Optimization

Model ReleasesDGX agent

arXiv:2608.09316v1 Announce Type: new Abstract: Offline context optimization improves an agent by revising its instructions and examples while keeping the model frozen. This approach learns from rollo

Model the Edit, Not the Image: Visual Autoregressive Editing from a Source-Centric Perspective

SafetyDGX agent

arXiv:2608.09057v1 Announce Type: new Abstract: Next-scale visual autoregressive models (VARs) have emerged as a powerful generative paradigm, producing high-quality images through efficient coarse-to

Motion-Aware Animatable Gaussian Avatars Deblurring

ApplicationsDGX agent

arXiv:2411.16758v4 Announce Type: replace Abstract: The creation of 3D human avatars from multi-view videos is a significant yet challenging task in computer vision. However, existing techniques rely

MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling

SafetyDGX agent

arXiv:2608.08553v1 Announce Type: new Abstract: Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from

MPISuperRes-PnP: A Super-Resolution Zero-Shot Plug-and-Play Reconstruction Algorithm for Magnetic Particle Imaging

Model ReleasesDGX agent

arXiv:2608.09672v1 Announce Type: new Abstract: Magnetic Particle Imaging (MPI) is an emerging medical imaging modality. MPI is based on the non-linear response of magnetic nanoparticles to an applied

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval

Model ReleasesDGX agent

arXiv:2608.07993v1 Announce Type: new Abstract: Human motion-text retrieval provides a rigorous means of assessing cross-modal alignment. Prevailing benchmarks are dominated by homogeneous indoor moti

MRI super-resolution in ten sampling steps using a diffusion bridge model

ResearchDGX agent

arXiv:2608.08819v1 Announce Type: new Abstract: Objective. MRI provides excellent soft-tissue contrast, but long acquisition times can cause patient discomfort and lead to motion artifacts, forcing a

MSP-Net: Manifold-Guided Spectral Prompt Network for Hyperspectral Object Tracking

Model ReleasesDGX agent

arXiv:2608.09575v1 Announce Type: new Abstract: Hyperspectral object tracking leverages abundant spectral information to provide unique advantages for target discrimination in complex scenes. However,

Multi-Relational Knowledge Graph Enhanced Embedding for Trajectory-User Linking

TutorialsDGX agent

arXiv:2608.08646v1 Announce Type: cross Abstract: Trajectory-User Linking (TUL) aims to identify the owner of an anonymous trajectory from a set of candidate users, providing a basis for user mobility

Multi-Submap Implicit Neural SLAM with Local-to-Global Loop Closure for Large-Scale Scene Reconstruction

Local AiDGX agent

arXiv:2608.09146v1 Announce Type: new Abstract: Neural Radiance Fields (NeRF)-based SLAM has demonstrated impressive results in small-scale scene reconstruction, yet scaling these methods to extensive

Multimodal Skin Lesion Classification with Swin Transformer and Clinical Metadata Fusion

ResearchDGX agent

arXiv:2608.07574v1 Announce Type: new Abstract: Skin lesion classification plays an important role in supporting the early diagnosis of skin cancer. However, automated analysis remains challenging due

MultiShadow: Multi-Object Shadow Generation for Image Compositing via Diffusion Model

SafetyDGX agent

arXiv:2603.02743v4 Announce Type: replace Abstract: Realistic shadow generation is crucial for achieving seamless image compositing, yet existing methods primarily focus on single-object insertion and

MVMD: A Multi-View Approach for Enhanced Mirror Detection

ResearchDGX agent

arXiv:2608.07559v1 Announce Type: new Abstract: In 3D reconstruction, mirrors introduce significant challenges by creating distorted and fragmented spaces, resulting in inaccurate and unreliable 3D mo

NBA_Streaming: A Large-Scale Benchmark for Fine-Grained Basketball Commentary Generation in Continuous Streams

Model ReleasesDGX agent

arXiv:2608.09200v1 Announce Type: new Abstract: Live basketball commentary generation requires determining when an event is sufficiently observable and describing it before subsequent events unfold. H

NeuralDMD: Interpretable Neural Representation of Dynamics from Sparse and Noisy Measurements

ResearchDGX agent

arXiv:2507.03094v2 Announce Type: replace Abstract: Many challenges in scientific imaging involve solving ill-posed inverse problems, where the goal is to recover spatio-temporal fields from indirect,

NeuroGuard: Neural Gradient Update Aware of Representation Damage

Model ReleasesDGX agent

arXiv:2608.08068v1 Announce Type: new Abstract: Long-tailed class-incremental learning (LT-CIL) must learn new classes from imbalanced streams while retaining old classes. Existing methods mainly chan

NewtonGS: Physics-Structured Object-Level Neural Newtonian Dynamics for Gaussian Scene Animation

ResearchDGX agent

arXiv:2608.07598v1 Announce Type: new Abstract: Animating objects in a static 3D Gaussian scene requires an explicit object-level dynamic state and a controllable model of object motion. Existing dyna

Not All Frames Deserve Full Computation: Accelerating Autoregressive Video Generation via Selective Computation and Predictive Extrapolation

ResearchDGX agent

arXiv:2604.02979v2 Announce Type: replace Abstract: Autoregressive (AR) video diffusion models enable long-form video generation but remain expensive due to repeated multi-step denoising. Existing tra

NTIRE 2026 Low-light Enhancement: Twilight Cowboy Challenge

ApplicationsDGX agent

arXiv:2608.09782v1 Announce Type: new Abstract: This paper presents a review of the NTIRE 2026 Low-light Enhancement: Twilight Cowboy Challenge. The objective of the competition was to merge a set of

OccAnyScene: Towards Unified Indoor-Outdoor 3D Occupancy Predictio

ResearchDGX agent

arXiv:2608.08696v1 Announce Type: new Abstract: 3D occupancy prediction is fundamental to scene understanding, yet existing 3D semantic occupancy methods are typically specialized to fixed scene types

OGG-FR: Orthogonal Gradient Gaming and Frequency Rectification for Unmanned Aerial Vehicle Infrared Image Super-Resolution

Model ReleasesDGX agent

arXiv:2608.09150v1 Announce Type: new Abstract: Unmanned aerial vehicle (UAV) infrared image super-resolution aims to recover weak thermal structures for deployment on resource-constrained platforms;

One Model to Magnify Them All: Efficient Scale-Invariant Histopathology via Conditional Normalization and Continuous Magnification Training

ResearchDGX agent

arXiv:2608.09403v1 Announce Type: new Abstract: Whole slide images (WSIs) in digital histopathology are acquired at discrete magnification levels encoding complementary diagnostic information from glo

One-Time Training for All Grains: Open-Set Grain Recognition and Quantitative Analysis

Local AiDGX agent

arXiv:2608.09345v1 Announce Type: new Abstract: Advances in crop breeding have introduced an increasing number of grain varieties, creating a growing demand for efficient variety recognition and quant

Overcoming Data Scarcity and Confidentiality in Hardware Assurance via Synthetic Generation

ResearchDGX agent

arXiv:2608.09914v1 Announce Type: cross Abstract: Hardware assurance relies on scanning electron microscopy (SEM) to verify nanoscale structures, but assembling the large, high-quality datasets requir

PARAGraph: Pathology-Anatomy-Aware Hierarchical Graph for Diabetic Retinopathy Grading

ResearchDGX agent

arXiv:2608.08368v1 Announce Type: new Abstract: Diabetic retinopathy (DR) remains a leading cause of vision loss among working-age adults worldwide, making reliable severity grading clinically importa

Parcel2Progression: An Anatomy-aware Longitudinal Framework for Alzheimer's Disease Diagnosis

ResearchDGX agent

arXiv:2608.08753v1 Announce Type: new Abstract: Alzheimer's disease (AD) progression is a longitudinal process with subtle pathological cues in the early stages. Yet, computational constraints have li

PatchHead: Learning Spatial Patch Evidence for Generalizable AI-Generated Image Detection

ResearchDGX agent

arXiv:2608.09223v1 Announce Type: new Abstract: AI-generated image detectors generalize poorly when their training and test images originate from different generators or datasets. Despite the rich spa

PE-Mamba: Bidirectional Selective Layer Aggregation for AI-Generated Image Detection

ResearchDGX agent

arXiv:2608.07999v1 Announce Type: new Abstract: AI-generated image (AIGI) detection has become increasingly challenging due to the rapid advancement of generative models and the diminishing gap betwee

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots

SafetyDGX agent

arXiv:2608.09931v1 Announce Type: new Abstract: Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Dist

PhysX-CoT: Structured Physical Reasoning from a Single Image to Simulation-Ready 3D Assets

ResearchDGX agent

arXiv:2608.08053v1 Announce Type: cross Abstract: Simulation-ready 3D assets are central to robotics and embodied AI. Generating them from a single image is usually framed as a vision-language model t

PosBridge: Multi-View Positional Embedding Transplant for Identity-Aware Image Editing

Local AiDGX agent

arXiv:2508.17302v2 Announce Type: replace Abstract: Localized subject-driven image editing aims to seamlessly integrate user-specified objects into target scenes. As generative models continue to scal

Predict to Skip: Linear Multistep Feature Forecasting for Efficient Diffusion Transformers

ResearchDGX agent

arXiv:2602.18093v2 Announce Type: replace Abstract: Diffusion Transformers (DiT) have emerged as a widely adopted backbone for high-fidelity image and video generation, yet their iterative denoising p

Predictive Failure Detection in Network Hardware Using Thermal Imaging and Deep Learning with Sensor Fusion

ResearchDGX agent

arXiv:2608.07582v1 Announce Type: new Abstract: Unplanned network hardware malfunctions can interrupt services and result in expensive downtime in data centers. A deep learning-based predictive mainte

Preserve More Details: Mitigating Content Drift in Real-World Image Super-Resolution

ApplicationsDGX agent

arXiv:2608.09373v1 Announce Type: new Abstract: Real-world image super-resolution (Real-ISR) aims to reconstruct high-quality (HQ) images from low-quality (LQ) inputs subject to diverse real-world deg

PressureMesh: 3D Human Mesh Estimation from Multi-Device Pressure Images

ResearchDGX agent

arXiv:2608.09550v1 Announce Type: new Abstract: Human pose monitoring is crucial in fields such as rehabilitation assessment and human-computer interaction. Due to its privacy-preserving nature, press

← Previous
1…34567…207
Next →