AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Hardware

KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

DGX agent

arXiv:2607.21446v1 Announce Type: cross Abstract: Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear l

hardwarearxiv-cs-cv
24 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

Latent Variable-Mediated Cross-Learning for Few-Shot Acoustic Impedance Imaging

DGX agent

arXiv:2607.20989v1 Announce Type: new Abstract: Acoustic impedance imaging is a fundamental yet severely ill-posed problem in subsurface analysis: the seismic wavelet is unknown, observations are band

researcharxiv-cs-cv
24 Jul 2026
Research

Learning-based Seam Correspondence Reconstruction in Sewing Patterns

DGX agent

arXiv:2607.21213v1 Announce Type: new Abstract: Digital sewing patterns typically consist of disjoint 2D panels without explicit stitch annotations, making downstream 3D modeling reliant on labor-inte

researcharxiv-cs-cv
24 Jul 2026
Model Releases

Learning to Navigate Efficiently with Only 0.58M Trainable Parameters

DGX agent

arXiv:2607.11029v2 Announce Type: replace-cross Abstract: Recent progress in visual navigation has largely been driven by scale: end-to-end policies with hundreds of millions of parameters trained on

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

Lessons and Open Questions from a Unified Study of Camera-Trap Species Recognition Over Time

DGX agent

arXiv:2603.20509v2 Announce Type: replace Abstract: Camera traps are vital for large-scale biodiversity monitoring, yet accurate automated analysis remains challenging due to diverse deployment enviro

model-releasesarxiv-cs-cv
24 Jul 2026
Research

Loss Landscape Topology Reveals Why Simple Baselines are Competitive at 3D Point Cloud Segmentation Under Class Imbalance

DGX agent

arXiv:2607.21089v1 Announce Type: new Abstract: Semantic segmentation of 3D point clouds faces severe class imbalance, yet the effectiveness of specialized imbalance-aware methods from 2D computer vis

researcharxiv-cs-cv
24 Jul 2026
Research

MAGE-Vein: Multi-Instance Age and Gender Estimation from Finger Vein Images

DGX agent

arXiv:2607.20897v1 Announce Type: new Abstract: Age estimation from finger vein images has been widely considered impractical due to severe demographic biases in public datasets and physiological conf

researcharxiv-cs-cv
24 Jul 2026
Model Releases

MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-Transfer

DGX agent

arXiv:2607.20924v1 Announce Type: new Abstract: Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Despite advances in full-face editing by diffusion

model-releasesarxiv-cs-cv
24 Jul 2026
Research

Masked Topology Modeling for Self-Supervised Learning on Parametric CAD

DGX agent

arXiv:2607.20642v1 Announce Type: new Abstract: Computer aided design (CAD) is ubiquitous: virtually any modern object was designed using editable CAD tools. However, with the shortage of available CA

researcharxiv-cs-cv
24 Jul 2026
Hardware

Ms. Forcing: Efficient Streaming Video Generation with Multi-Scale Patchification and Attention

DGX agent

arXiv:2607.20940v1 Announce Type: new Abstract: Streaming video diffusion models have made substantial progress toward interactive and dynamic world simulation, but the nested autoregressive and denoi

hardwarearxiv-cs-cv
24 Jul 2026
Safety

Multivariate Planar Curves: A Statistical Framework for Shape Analysis in Images

DGX agent

arXiv:2508.11780v3 Announce Type: replace-cross Abstract: Recent developments in computer vision have made segmented images widely available across many domains, such as medicine, where segmented radi

safetyarxiv-cs-cv
24 Jul 2026
Model Releases

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement

DGX agent

arXiv:2607.21061v1 Announce Type: new Abstract: Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward

model-releasesarxiv-cs-cv
24 Jul 2026
Research

Ocular Verification for Virtual Reality

DGX agent

arXiv:2607.20790v1 Announce Type: new Abstract: Virtual reality (VR) headsets (e.g., Meta Quest, Apple Vision Pro) provide a seamless user experience due to their fast, frictionless interaction with t

researcharxiv-cs-cv
24 Jul 2026
Model Releases

ODeform: Learning Continuous 4D Motion for Shape Deformation with Neural ODEs

DGX agent

arXiv:2607.20670v1 Announce Type: new Abstract: Modeling continuous object deformation is important for many computer vision and robotics tasks, such as manipulation and simulation. Existing approache

model-releasesarxiv-cs-cv
24 Jul 2026
Research

Out of Sight, Still in Mind: Token Compression for Omni-LLMs

DGX agent

arXiv:2607.21179v1 Announce Type: new Abstract: The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at inference time. Omni-LLMs reason jointly ove

researcharxiv-cs-cv
24 Jul 2026
Research

PersonaGesture: Single-Reference Co-Speech Gesture Personalization for Unseen Speakers

DGX agent

arXiv:2605.06064v2 Announce Type: replace Abstract: We propose PersonaGesture, a diffusion-based pipeline for single-reference co-speech gesture personalization of unseen speakers. Given target speech

researcharxiv-cs-cv
24 Jul 2026
Research

PhysCoRe: Physics-Corrected Residual World Models for Material-Aware Deformable Dynamics

DGX agent

arXiv:2607.20653v1 Announce Type: cross Abstract: Predicting how deformable objects evolve under robotic manipulation is a longstanding challenge. Existing approaches typically rely on per-object opti

researcharxiv-cs-cv
24 Jul 2026
Research

Physics-Informed Deep Learning Model for Cross-Modality Super-Resolution in Fluorescence Microscopy

DGX agent

arXiv:2607.21190v1 Announce Type: new Abstract: Cross-modality image translation offers a route to super-resolution fluorescence microscopy from low-resolution images while reducing phototoxicity and

researcharxiv-cs-cv
24 Jul 2026
Model Releases

ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning

DGX agent

arXiv:2607.21022v1 Announce Type: new Abstract: Improving video captioning quality typically demands retraining large vision-language models, an expensive and often impractical requirement. Existing t

model-releasesarxiv-cs-cv
24 Jul 2026
Safety

QATMA: Quantization-Aware Training with Multimodal Alignment for Open-Vocabulary Object Detection

DGX agent

arXiv:2603.05964v3 Announce Type: replace Abstract: Quantizing open-vocabulary object detection (OVOD) models reduces their memory and computational costs, but extremely low-bit quantization severely

safetyarxiv-cs-cv
24 Jul 2026
Tutorials

Quality-Aware Multimodal Fusion Reveals Implicit Identity in Valence-Arousal Features

DGX agent

arXiv:2607.21347v1 Announce Type: new Abstract: Conventional face recognition relies on static appearance cues and degrades in unconstrained settings with expression variation, occlusion, and poor lig

tutorialsarxiv-cs-cv
24 Jul 2026
Agents

RECO: Region-Aware Compensation for Extrinsic Perturbations in Roadside 3D Detection

DGX agent

arXiv:2607.20947v1 Announce Type: new Abstract: In intelligent transportation systems, roadside 3D object detection provides wide-area perception crucial for traffic understanding, cooperative early w

agentsarxiv-cs-cv
24 Jul 2026
Research

Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation

DGX agent

arXiv:2607.21485v1 Announce Type: new Abstract: We study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural representations (INRs). Our analysis reveal

researcharxiv-cs-cv
24 Jul 2026
Local Ai

Rethinking Open-World Video Anomaly Detection: Diagnosing Definition Blindness

DGX agent

arXiv:2607.20780v1 Announce Type: new Abstract: Open-world video anomaly detection (OWVAD) is expected to detect events that match a user-specified definition of abnormality. This requirement is stron

local-aiarxiv-cs-cv
24 Jul 2026
Safety

Safety-oriented sidewalk and road segmentation for smartphone-based assistive navigation

DGX agent

arXiv:2607.21137v1 Announce Type: new Abstract: Independent sidewalk mobility is essential for blind and visually impaired pedestrians (BVIPs), yet smartphone-based assistive navigation requires perce

safetyarxiv-cs-cv
24 Jul 2026
Hardware

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

DGX agent

arXiv:2607.21553v1 Announce Type: new Abstract: We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate h

hardwarearxiv-cs-cv
24 Jul 2026
Safety

Scale Up Strategically: Learning Compositional Generalization via Bias-Aware Evaluation and Data Collection for Robotic Manipulation

DGX agent

arXiv:2607.21582v1 Announce Type: cross Abstract: Compositional generalization is essential for robot to follow diverse instructions. However, pretrained policies are known to take shortcuts, deferrin

safetyarxiv-cs-cv
24 Jul 2026
Model Releases

Scene Parameter Saliency via Differentiable Light Transport

DGX agent

arXiv:2607.21562v1 Announce Type: new Abstract: Gradient-based saliency methods reveal which input features most influence a neural network's output, and are a standard tool for model interpretability

model-releasesarxiv-cs-cv
24 Jul 2026
Safety

Self-Supervised Learning of Structured Dynamics from Videos

DGX agent

arXiv:2607.21576v1 Announce Type: new Abstract: Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion

safetyarxiv-cs-cv
24 Jul 2026
Model Releases

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

DGX agent

arXiv:2607.21072v1 Announce Type: new Abstract: Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks a

model-releasesarxiv-cs-cv
24 Jul 2026
Safety

Sidewalk Moments: Are Richer Representations Always More Human-Aligned? Evidence from City-Walk Videos

DGX agent

arXiv:2607.20903v1 Announce Type: new Abstract: We examine whether richer visual representations yield more human-aligned measures of urban engagement, using 61 first-person city-walk videos from YouT

safetyarxiv-cs-cv
24 Jul 2026
Safety

SlerpFlow: Spherical Trajectory Correction for Rectified Flow Inversion

DGX agent

arXiv:2607.21326v1 Announce Type: new Abstract: Rectified-flow-based diffusion transformers, particularly FLUX, have demonstrated outstanding performance in high-quality image generation. However, ach

safetyarxiv-cs-cv
24 Jul 2026
Applications

SoccerSynth Field: enhancing field detection with synthetic data from virtual soccer simulator

DGX agent

arXiv:2503.13969v2 Announce Type: replace Abstract: Field detection in team sports is an essential task in sports video analysis. However, collecting large-scale and diverse real-world datasets for tr

applicationsarxiv-cs-cv
24 Jul 2026
Research

SPDCN: Strip-based Deformable Convolutional Network for Steel Surface Defect Segmentation

DGX agent

arXiv:2607.21456v1 Announce Type: new Abstract: Steel surface defect segmentation is critical for industrial quality inspection, yet existing methods struggle with elongated, anisotropic defects such

researcharxiv-cs-cv
24 Jul 2026
Model Releases

Spectral-Spatial Synergistic Guided Network for Hyperspectral Salient Object Detection

DGX agent

arXiv:2607.21032v1 Announce Type: new Abstract: Hyperspectral salient object detection aims to identify visually salient regions from hyperspectral images. Existing methods often fail because they fun

model-releasesarxiv-cs-cv
24 Jul 2026
Safety

Stokes-Informed Diffusion for Robust Linear Polarization Estimation

DGX agent

arXiv:2607.21239v1 Announce Type: new Abstract: Polarization cues benefit applications such as material detection and de-reflection, yet acquiring them typically requires dedicated hardware. This moti

safetyarxiv-cs-cv
24 Jul 2026
Agents

Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

DGX agent

arXiv:2607.21594v1 Announce Type: new Abstract: Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evo

agentsarxiv-cs-cv
24 Jul 2026
Research

SubSplat: High-Resolution Pixel-aligned 3DGS via Sub-pixel Gaussian Reparameterization

DGX agent

arXiv:2607.20813v1 Announce Type: new Abstract: Pixel-aligned Gaussian splatting enables efficient and generalizable novel-view synthesis. However, high-resolution rendering faces a critical trade-off

researcharxiv-cs-cv
24 Jul 2026
Safety

SuperFlow: Training Flow Matching Models with RL on the Fly

DGX agent

arXiv:2512.17951v3 Announce Type: replace Abstract: Recent progress in flow-based generative models and reinforcement learning (RL) has improved text-image alignment and visual quality. However, curre

safetyarxiv-cs-cv
24 Jul 2026
Model Releases

T-STAR: A Large-Scale Benchmark for Spatio-Temporal Panoptic Scene Graph Generation in Satellite Video

DGX agent

arXiv:2607.21228v1 Announce Type: new Abstract: Structured understanding of satellite video is essential for advancing dynamic geospatial scene analysis from low-level perception to high-level cogniti

model-releasesarxiv-cs-cv
24 Jul 2026
Research

Texture++: Elevating 3D Asset Texture Resolution with a Region-Aware Diffusion Model

DGX agent

arXiv:2607.21504v1 Announce Type: new Abstract: Numerous 3D assets are discarded due to low texture resolution, while current super-resolution models ignore texture maps and focus on natural images. A

researcharxiv-cs-cv
24 Jul 2026
Model Releases

The RealDefocus Benchmark for Defocus Deblurring

DGX agent

arXiv:2607.21078v1 Announce Type: new Abstract: Single-Image Defocus Deblurring (SIDD) aims to recover an all-in-focus image from a single defocused observation, but rigorous and reproducible evaluati

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

The Second LoViF 2026 Challenge on Real-World All-in-One Image Restoration: Methods and Results

DGX agent

arXiv:2607.21118v1 Announce Type: new Abstract: This paper presents a review of the second LoViF Challenge on Real-World All-in-One Image Restoration. The challenge aims to advance unified image resto

model-releasesarxiv-cs-cv
24 Jul 2026
Local Ai

Towards Privacy-Preserving Federated Prompt Tuning under Data Heterogeneity: A Subspace-Decomposed Expert Approach

DGX agent

arXiv:2607.21417v1 Announce Type: new Abstract: Federated prompt tuning (FPT) enables collaborative adaptation of vision--language models (VLMs) using lightweight prompts. Existing methods often addre

local-aiarxiv-cs-cv
24 Jul 2026
Research

Towards Robust Iris Recognition Through Occlusion Identification and Conditional Diffusion-Based Reconstruction

DGX agent

arXiv:2607.21545v1 Announce Type: new Abstract: Iris recognition is a reliable biometric approach that identifies individuals using the distinctive and stable texture of the iris. However, recognition

researcharxiv-cs-cv
24 Jul 2026
Hardware

Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers

DGX agent

arXiv:2512.16615v2 Announce Type: replace Abstract: Diffusion Transformers (DiTs) set the state of the art in visual generation, yet their quadratic self-attention cost fundamentally limits scaling to

hardwarearxiv-cs-cv
24 Jul 2026
Model Releases

TransBiolab: A Real-World Multi-View Dataset of Cluttered Transparent Biomedical Objects

DGX agent

arXiv:2607.21071v1 Announce Type: new Abstract: Autonomous biomedical laboratories increasingly rely on visual perception to recognize, localize, and manipulate transparent plasticware, yet high-quali

model-releasesarxiv-cs-cv
24 Jul 2026
Safety

UnDA: Unpaired Domain Alignment for Cross-Modal Knowledge Transfer in Medical Imaging

DGX agent

arXiv:2607.21546v1 Announce Type: new Abstract: Multimodal based approaches often outperform single modality approaches in downstream tasks as the different modalities provide complementary informatio

safetyarxiv-cs-cv
24 Jul 2026
← Previous
1…4344454647…261
Next →