AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
24 Jul 2026

FSB-Net: Frequency-Spatial Boundary Network for Brain Stroke Lesion Segmentation in Non-Contrast CT

ResearchDGX agent

arXiv:2607.20955v1 Announce Type: new Abstract: Accurate segmentation of brain stroke lesions in non-contrast computed tomography (NCCT) scans is critical for rapid clinical decision-making, yet remai

Future Rendering neq Future Surface: A Benchmark and Dataset for Dynamic Surface Reconstruction Beyond the Observed Window

Model ReleasesDGX agent

arXiv:2607.21471v1 Announce Type: new Abstract: Dynamic-scene reconstruction is almost always evaluated inside the observed time window, yet deployment settings such as AR overlays, robot interaction,

Geo3R: Mitigating Spatial Reasoning Hallucination in Multimodal Large Language Models

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.21085v1 Announce Type: new Abstract: Despite remarkable progress in visual understanding, Multimodal Large Language Models (MLLMs) remain prone to hallucinations when reasoning about spatia

GeoThreat: Transferable Targeted Adversarial Attacks on Large Vision-Language Models for Remote Sensing Image Interpretation

ResearchDGX agent

arXiv:2607.21036v1 Announce Type: new Abstract: Adversarial attacks against large vision-language models (LVLMs) serve as an effective means of assessing their robustness in cross-modal semantic under

GLAM-SLAM: Real-time Gaussian Large-scale Mapping via Flow Densification and Spatial Decomposition

SafetyDGX agent

arXiv:2607.21416v1 Announce Type: cross Abstract: Existing Gaussian-splatting-based monocular Simultaneous Localization and Mapping (SLAM) systems are either tailored to short sequences, are not real-

GrainGS: Gradient-Decoupled Gaussian Splatting for Efficient Dynamic Novel View Synthesis

Model ReleasesDGX agent

arXiv:2607.21448v1 Announce Type: new Abstract: Dynamic scene reconstruction with 3D Gaussian Splatting requires a balance between fine-grained motion modeling, structural stability, and compact repre

GroupVideo: Multi-Identity Customized Text-to-Video Generation

SafetyDGX agent

arXiv:2607.21027v1 Announce Type: new Abstract: Current identity customized video generation methodologies are predominantly limited to single-identity scenarios, as the lack of explicit identity sepa

HalluScope: Fine-grained Hallucination Diagnosis for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2607.21105v1 Announce Type: new Abstract: Although Multimodal Large Language Models have achieved strong performance across a wide range of vision-language tasks, they still suffer from hallucin

HGeo-TopoMap: Boosting Topological Mapping with Hierarchical Geometric Priors

AgentsDGX agent

arXiv:2607.21281v1 Announce Type: new Abstract: Topological maps are key outputs of autonomous driving perception systems, delivering essential road information for path planning. They identify instan

HyperImageNet: A Large-Scale High-Spatial Resolution Hyperspectral Imagery Classification Benchmark

Model ReleasesDGX agent

arXiv:2607.21050v1 Announce Type: new Abstract: We present HyperImageNet, a large-scale benchmark for fine-grained hyperspectral land-cover understanding. The dataset contains 26,084 airborne hyperspe

Incremental Optimal Assignment for Real-Time Crowd Tracking

ResearchDGX agent

arXiv:2607.21368v1 Announce Type: new Abstract: Multi-object tracking in dense crowds requires solving a bipartite assignment problem between detections and trajectories at every video frame. The clas

Inference-Time Scaling of Diffusion Models via Progressive Seed Pruning

SafetyDGX agent

arXiv:2607.21591v1 Announce Type: new Abstract: Diffusion and flow-matching models dominate conditional image generation, yet inference-time scaling for these models is far less developed than for aut

KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

HardwareDGX agent

arXiv:2607.21446v1 Announce Type: cross Abstract: Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear l

Latent Variable-Mediated Cross-Learning for Few-Shot Acoustic Impedance Imaging

ResearchDGX agent

arXiv:2607.20989v1 Announce Type: new Abstract: Acoustic impedance imaging is a fundamental yet severely ill-posed problem in subsurface analysis: the seismic wavelet is unknown, observations are band

Learning-based Seam Correspondence Reconstruction in Sewing Patterns

ResearchDGX agent

arXiv:2607.21213v1 Announce Type: new Abstract: Digital sewing patterns typically consist of disjoint 2D panels without explicit stitch annotations, making downstream 3D modeling reliant on labor-inte

Learning to Navigate Efficiently with Only 0.58M Trainable Parameters

Model ReleasesDGX agent

arXiv:2607.11029v2 Announce Type: replace-cross Abstract: Recent progress in visual navigation has largely been driven by scale: end-to-end policies with hundreds of millions of parameters trained on

Lessons and Open Questions from a Unified Study of Camera-Trap Species Recognition Over Time

Model ReleasesDGX agent

arXiv:2603.20509v2 Announce Type: replace Abstract: Camera traps are vital for large-scale biodiversity monitoring, yet accurate automated analysis remains challenging due to diverse deployment enviro

Loss Landscape Topology Reveals Why Simple Baselines are Competitive at 3D Point Cloud Segmentation Under Class Imbalance

ResearchDGX agent

arXiv:2607.21089v1 Announce Type: new Abstract: Semantic segmentation of 3D point clouds faces severe class imbalance, yet the effectiveness of specialized imbalance-aware methods from 2D computer vis

MAGE-Vein: Multi-Instance Age and Gender Estimation from Finger Vein Images

ResearchDGX agent

arXiv:2607.20897v1 Announce Type: new Abstract: Age estimation from finger vein images has been widely considered impractical due to severe demographic biases in public datasets and physiological conf

MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-Transfer

Model ReleasesDGX agent

arXiv:2607.20924v1 Announce Type: new Abstract: Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Despite advances in full-face editing by diffusion

Masked Topology Modeling for Self-Supervised Learning on Parametric CAD

ResearchDGX agent

arXiv:2607.20642v1 Announce Type: new Abstract: Computer aided design (CAD) is ubiquitous: virtually any modern object was designed using editable CAD tools. However, with the shortage of available CA

Ms. Forcing: Efficient Streaming Video Generation with Multi-Scale Patchification and Attention

HardwareDGX agent

arXiv:2607.20940v1 Announce Type: new Abstract: Streaming video diffusion models have made substantial progress toward interactive and dynamic world simulation, but the nested autoregressive and denoi

Multivariate Planar Curves: A Statistical Framework for Shape Analysis in Images

SafetyDGX agent

arXiv:2508.11780v3 Announce Type: replace-cross Abstract: Recent developments in computer vision have made segmented images widely available across many domains, such as medicine, where segmented radi

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement

Model ReleasesDGX agent

arXiv:2607.21061v1 Announce Type: new Abstract: Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward

Ocular Verification for Virtual Reality

ResearchDGX agent

arXiv:2607.20790v1 Announce Type: new Abstract: Virtual reality (VR) headsets (e.g., Meta Quest, Apple Vision Pro) provide a seamless user experience due to their fast, frictionless interaction with t

ODeform: Learning Continuous 4D Motion for Shape Deformation with Neural ODEs

Model ReleasesDGX agent

arXiv:2607.20670v1 Announce Type: new Abstract: Modeling continuous object deformation is important for many computer vision and robotics tasks, such as manipulation and simulation. Existing approache

Out of Sight, Still in Mind: Token Compression for Omni-LLMs

ResearchDGX agent

arXiv:2607.21179v1 Announce Type: new Abstract: The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at inference time. Omni-LLMs reason jointly ove

PersonaGesture: Single-Reference Co-Speech Gesture Personalization for Unseen Speakers

ResearchDGX agent

arXiv:2605.06064v2 Announce Type: replace Abstract: We propose PersonaGesture, a diffusion-based pipeline for single-reference co-speech gesture personalization of unseen speakers. Given target speech

PhysCoRe: Physics-Corrected Residual World Models for Material-Aware Deformable Dynamics

ResearchDGX agent

arXiv:2607.20653v1 Announce Type: cross Abstract: Predicting how deformable objects evolve under robotic manipulation is a longstanding challenge. Existing approaches typically rely on per-object opti

Physics-Informed Deep Learning Model for Cross-Modality Super-Resolution in Fluorescence Microscopy

ResearchDGX agent

arXiv:2607.21190v1 Announce Type: new Abstract: Cross-modality image translation offers a route to super-resolution fluorescence microscopy from low-resolution images while reducing phototoxicity and

ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning

Model ReleasesDGX agent

arXiv:2607.21022v1 Announce Type: new Abstract: Improving video captioning quality typically demands retraining large vision-language models, an expensive and often impractical requirement. Existing t

QATMA: Quantization-Aware Training with Multimodal Alignment for Open-Vocabulary Object Detection

SafetyDGX agent

arXiv:2603.05964v3 Announce Type: replace Abstract: Quantizing open-vocabulary object detection (OVOD) models reduces their memory and computational costs, but extremely low-bit quantization severely

Quality-Aware Multimodal Fusion Reveals Implicit Identity in Valence-Arousal Features

TutorialsDGX agent

arXiv:2607.21347v1 Announce Type: new Abstract: Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with expression variation, occlusion, and poor lig

RECO: Region-Aware Compensation for Extrinsic Perturbations in Roadside 3D Detection

AgentsDGX agent

arXiv:2607.20947v1 Announce Type: new Abstract: In intelligent transportation systems, roadside 3D object detection provides wide-area perception crucial for traffic understanding, cooperative early w

Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation

ResearchDGX agent

arXiv:2607.21485v1 Announce Type: new Abstract: We study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural representations (INRs). Our analysis reveal

Rethinking Open-World Video Anomaly Detection: Diagnosing Definition Blindness

Local AiDGX agent

arXiv:2607.20780v1 Announce Type: new Abstract: Open-world video anomaly detection (OWVAD) is expected to detect events that match a user-specified definition of abnormality. This requirement is stron

Safety-oriented sidewalk and road segmentation for smartphone-based assistive navigation

SafetyDGX agent

arXiv:2607.21137v1 Announce Type: new Abstract: Independent sidewalk mobility is essential for blind and visually impaired pedestrians (BVIPs), yet smartphone-based assistive navigation requires perce

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

HardwareDGX agent

arXiv:2607.21553v1 Announce Type: new Abstract: We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate h

Scale Up Strategically: Learning Compositional Generalization via Bias-Aware Evaluation and Data Collection for Robotic Manipulation

SafetyDGX agent

arXiv:2607.21582v1 Announce Type: cross Abstract: Compositional generalization is essential for robot to follow diverse instructions. However, pretrained policies are known to take shortcuts, deferrin

Scene Parameter Saliency via Differentiable Light Transport

Model ReleasesDGX agent

arXiv:2607.21562v1 Announce Type: new Abstract: Gradient-based saliency methods reveal which input features most influence a neural network's output, and are a standard tool for model interpretability

Self-Supervised Learning of Structured Dynamics from Videos

SafetyDGX agent

arXiv:2607.21576v1 Announce Type: new Abstract: Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Model ReleasesDGX agent

arXiv:2607.21072v1 Announce Type: new Abstract: Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks a

Sidewalk Moments: Are Richer Representations Always More Human-Aligned? Evidence from City-Walk Videos

SafetyDGX agent

arXiv:2607.20903v1 Announce Type: new Abstract: We examine whether richer visual representations yield more human-aligned measures of urban engagement, using 61 first-person city-walk videos from YouT

SlerpFlow: Spherical Trajectory Correction for Rectified Flow Inversion

SafetyDGX agent

arXiv:2607.21326v1 Announce Type: new Abstract: Rectified-flow-based diffusion transformers, particularly FLUX, have demonstrated outstanding performance in high-quality image generation. However, ach

SoccerSynth Field: enhancing field detection with synthetic data from virtual soccer simulator

ApplicationsDGX agent

arXiv:2503.13969v2 Announce Type: replace Abstract: Field detection in team sports is an essential task in sports video analysis. However, collecting large-scale and diverse real-world datasets for tr

SPDCN: Strip-based Deformable Convolutional Network for Steel Surface Defect Segmentation

ResearchDGX agent

arXiv:2607.21456v1 Announce Type: new Abstract: Steel surface defect segmentation is critical for industrial quality inspection, yet existing methods struggle with elongated, anisotropic defects such

Spectral-Spatial Synergistic Guided Network for Hyperspectral Salient Object Detection

Model ReleasesDGX agent

arXiv:2607.21032v1 Announce Type: new Abstract: Hyperspectral salient object detection aims to identify visually salient regions from hyperspectral images. Existing methods often fail because they fun

Stokes-Informed Diffusion for Robust Linear Polarization Estimation

SafetyDGX agent

arXiv:2607.21239v1 Announce Type: new Abstract: Polarization cues benefit applications such as material detection and de-reflection, yet acquiring them typically requires dedicated hardware. This moti

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

AgentsDGX agent

arXiv:2607.21594v1 Announce Type: new Abstract: Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evo

SubSplat: High-Resolution Pixel-aligned 3DGS via Sub-pixel Gaussian Reparameterization

ResearchDGX agent

arXiv:2607.20813v1 Announce Type: new Abstract: Pixel-aligned Gaussian splatting enables efficient and generalizable novel-view synthesis. However, high-resolution rendering faces a critical trade-off

SuperFlow: Training Flow Matching Models with RL on the Fly

SafetyDGX agent

arXiv:2512.17951v3 Announce Type: replace Abstract: Recent progress in flow-based generative models and reinforcement learning (RL) has improved text-image alignment and visual quality. However, curre

T-STAR: A Large-Scale Benchmark for Spatio-Temporal Panoptic Scene Graph Generation in Satellite Video

Model ReleasesDGX agent

arXiv:2607.21228v1 Announce Type: new Abstract: Structured understanding of satellite video is essential for advancing dynamic geospatial scene analysis from low-level perception to high-level cogniti

Texture++: Elevating 3D Asset Texture Resolution with a Region-Aware Diffusion Model

ResearchDGX agent

arXiv:2607.21504v1 Announce Type: new Abstract: Numerous 3D assets are discarded due to low texture resolution, while current super-resolution models ignore texture maps and focus on natural images. A

The RealDefocus Benchmark for Defocus Deblurring

Model ReleasesDGX agent

arXiv:2607.21078v1 Announce Type: new Abstract: Single-Image Defocus Deblurring (SIDD) aims to recover an all-in-focus image from a single defocused observation, but rigorous and reproducible evaluati

The Second LoViF 2026 Challenge on Real-World All-in-One Image Restoration: Methods and Results

Model ReleasesDGX agent

arXiv:2607.21118v1 Announce Type: new Abstract: This paper presents a review of the second LoViF Challenge on Real-World All-in-One Image Restoration. The challenge aims to advance unified image resto

Towards Privacy-Preserving Federated Prompt Tuning under Data Heterogeneity: A Subspace-Decomposed Expert Approach

Local AiDGX agent

arXiv:2607.21417v1 Announce Type: new Abstract: Federated prompt tuning (FPT) enables collaborative adaptation of vision--language models (VLMs) using lightweight prompts. Existing methods often addre

Towards Robust Iris Recognition Through Occlusion Identification and Conditional Diffusion-Based Reconstruction

ResearchDGX agent

arXiv:2607.21545v1 Announce Type: new Abstract: Iris recognition is a reliable biometric approach that identifies individuals using the distinctive and stable texture of the iris. However, recognition

Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers

HardwareDGX agent

arXiv:2512.16615v2 Announce Type: replace Abstract: Diffusion Transformers (DiTs) set the state of the art in visual generation, yet their quadratic self-attention cost fundamentally limits scaling to

TransBiolab: A Real-World Multi-View Dataset of Cluttered Transparent Biomedical Objects

Model ReleasesDGX agent

arXiv:2607.21071v1 Announce Type: new Abstract: Autonomous biomedical laboratories increasingly rely on visual perception to recognize, localize, and manipulate transparent plasticware, yet high-quali

UnDA: Unpaired Domain Alignment for Cross-Modal Knowledge Transfer in Medical Imaging

SafetyDGX agent

arXiv:2607.21546v1 Announce Type: new Abstract: Multimodal based approaches often outperform single modality approaches in downstream tasks as the different modalities provide complementary informatio

← Previous
1…3435363738…209
Next →