AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Model Releases

Inference-time Trajectory Optimization for Structure-Preserving Manga Image Editing

DGX agent

arXiv:2603.27790v2 Announce Type: replace Abstract: We present a lightweight, training-free trajectory correction method that adapts a pretrained image editing model to each input manga image using on

model-releasesarxiv-cs-cv
3 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Is It Time for the Renaissance of Salient Object Detection in the Era of MLLMs?

DGX agent

arXiv:2607.29222v1 Announce Type: new Abstract: The zero-shot capabilities of multimodal large language models (MLLMs) are pushing salient object detection (SOD) beyond task-specific supervision. To d

model-releasesarxiv-cs-cv
3 Aug 2026
Safety

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

DGX agent

arXiv:2512.13677v2 Announce Type: replace Abstract: In this paper, we present JoVA, a streamlined framework that unifies joint video-audio generation and editing. While existing methods often rely on

safetyarxiv-cs-cv
3 Aug 2026
Research

Learning from Adversity: Semantic-Aware Mask Refinement through Adversarial Perturbation

DGX agent

arXiv:2607.29059v1 Announce Type: new Abstract: Despite significant advances in image segmentation, even state-of-the-art models produce masks with imperfect boundaries, semantic inconsistencies, and

researcharxiv-cs-cv
3 Aug 2026
Applications

Learning Manifolds in High-D Point Embedding for Anisotropic Surface Approximation from Unstructured Point Clouds

DGX agent

arXiv:2607.28855v1 Announce Type: cross Abstract: Dense 3D sensors in various real-world fields produce point clouds that are geometrically redundant for real-time processing. In this paper, we propos

applicationsarxiv-cs-cv
3 Aug 2026
Research

LegoQ: Density-Matrix Representation Learning with Spectral-Spatial State Transitions for Hyperspectral Classification

DGX agent

arXiv:2607.28970v1 Announce Type: new Abstract: Hyperspectral image classification is complicated by mixed pixels, spectral ambiguity, class imbalance, and limited annotations. Most current classifier

researcharxiv-cs-cv
3 Aug 2026
Research

Leveraging Transfer Learning with Class-Specific Decoders for Laparoscopic Segmentation

DGX agent

arXiv:2607.29509v1 Announce Type: new Abstract: Effective multi-organ segmentation in surgical data requires learning the intricate anatomical features and alleviating the challenge of class imbalance

researcharxiv-cs-cv
3 Aug 2026
Applications

Lightweight Neural Networks for Affordance Segmentation: Enhancement of the Decoder Module

DGX agent

arXiv:2607.29473v1 Announce Type: new Abstract: The deployment of deep neural networks for visual affordance segmentation on wearable robots poses may prove critical, due to some conflicting aspects o

applicationsarxiv-cs-cv
3 Aug 2026
Model Releases

Locally Consistent Transductive Information Maximization for Few-Shot Remote Sensing Scene Classification

DGX agent

arXiv:2607.29192v1 Announce Type: new Abstract: Remote sensing scene classification is increasingly relying on foundation models pre-trained on large-scale Earth-observation data. Moreover, transducti

model-releasesarxiv-cs-cv
3 Aug 2026
Research

Meshy T2: Fast Native Mesh Generation with Flow Matching

DGX agent

arXiv:2607.28675v1 Announce Type: cross Abstract: Polygonal meshes are the standard surface representation of modern 3D pipelines, and generating high-quality meshes with artist-style topology is esse

researcharxiv-cs-cv
3 Aug 2026
Research

MHRGait: Gait Recognition from Momentum Human Rig Pose

DGX agent

arXiv:2607.29083v1 Announce Type: new Abstract: Gait recognition is shaped by its input representation. Silhouettes encode projected body shape, skeletons encode sparse joint coordinates, and 3D meshe

researcharxiv-cs-cv
3 Aug 2026
Safety

Mirror Learning

DGX agent

arXiv:2607.28737v1 Announce Type: cross Abstract: We investigate imitation learning through the lens of third-person observation and propose a framework for mirror learning: acquiring actionable polic

safetyarxiv-cs-cv
3 Aug 2026
Local Ai

Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift

DGX agent

arXiv:2607.28696v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) can retain high observed marginal coverage after clinical shift while substantially under-covering an individual

local-aiarxiv-cs-cv
3 Aug 2026
Research

Moment kernels: a simple and scalable approach for equivariance to rotations and reflections in deep convolutional networks

DGX agent

arXiv:2505.21736v2 Announce Type: replace Abstract: Translation equivariance is a central reason convolutional neural networks have been successful in computer vision. Other symmetries, such as rotati

researcharxiv-cs-cv
3 Aug 2026
Model Releases

MoPET: Parameter-Efficient Mixture-of-Experts for Unified Medical Image Classification

DGX agent

arXiv:2607.29462v1 Announce Type: cross Abstract: Adapting deep learning models to profound clinical heterogeneity typically relies on parameter-efficient fine-tuning (PEFT) to avoid the severe overfi

model-releasesarxiv-cs-cv
3 Aug 2026
Model Releases

MoRoute: Dynamic Routing for In-Context Multimodal Video Generation

DGX agent

arXiv:2607.29545v1 Announce Type: new Abstract: Multimodal video generation aims to generate and edit videos conditioned on arbitrary combinations of text, images, and videos within a single model, al

model-releasesarxiv-cs-cv
3 Aug 2026
Safety

Multi-Modal Object Re-Identification with Dual Semantic Guidance and Global-Local Mutual Modulation

DGX agent

arXiv:2607.29207v1 Announce Type: new Abstract: Multi-modal object Re-Identification (ReID) aims to retrieve target instances by leveraging complementary information across modalities. However, existi

safetyarxiv-cs-cv
3 Aug 2026
Safety

Multi-Source Multi-View Graph Domain Adaptation with Hyperbolic Residual Encoding for Cross-Site MDD Identification from rs-fMRI

DGX agent

arXiv:2607.29531v1 Announce Type: new Abstract: Cross-site identification of major depressive disorder (MDD) from resting-state functional magnetic resonance imaging (rs-fMRI) is hindered by inter-sit

safetyarxiv-cs-cv
3 Aug 2026
Research

OASIS: Occlusion-aware Single-image Hand Avatar Reconstruction via 3D Gaussian Splatting

DGX agent

arXiv:2607.29633v1 Announce Type: new Abstract: Single-image 3D hand avatar reconstruction is fundamentally ill-posed and particularly challenging due to limited visual evidence under severe self-occl

researcharxiv-cs-cv
3 Aug 2026
Safety

On the Efficacy of Self-Supervised Point Cloud Encoders for Efficient 3D Large Language Models

DGX agent

arXiv:2607.29136v1 Announce Type: new Abstract: 3D point cloud-language models (3D-LLMs) enable 3D understanding by pairing point cloud encoders with large language models, but existing methods rely o

safetyarxiv-cs-cv
3 Aug 2026
Research

Optical Flow Sensor: A Direction-Selective Bionic Retina Design

DGX agent

arXiv:2607.28686v1 Announce Type: cross Abstract: Optical flow characterizes motion in the visual field and is fundamental to motion perception and tracking in biological and artificial vision systems

researcharxiv-cs-cv
3 Aug 2026
Model Releases

OSAGEN: Object-Aware Mask Priors and Multistage Decoupled Diffusion for Industrial Anomaly Generation

DGX agent

arXiv:2607.29533v1 Announce Type: new Abstract: Industrial anomaly detection and localization are limited by scarce real anomalies and pixel-level annotations, a bottleneck that synthetic image-mask p

model-releasesarxiv-cs-cv
3 Aug 2026
Model Releases

OSEF: One-Step Evidence Fusion for Cross-Video Scene Procedure Planning

DGX agent

arXiv:2607.29401v1 Announce Type: new Abstract: Video Scene Procedure Planning (VSPP) supplies the target start-goal observations in advance, leaving open how a planner should act when the evidence mu

model-releasesarxiv-cs-cv
3 Aug 2026
Model Releases

Parameter-Efficient Fine-Tuning for Spiking Point Cloud Models

DGX agent

arXiv:2607.29048v1 Announce Type: new Abstract: Spiking Neural Networks (SNNs) offer energy-efficient solutions for point cloud analysis on resource-constrained devices through event-driven computatio

model-releasesarxiv-cs-cv
3 Aug 2026
Safety

Physics-Aligned Self-Supervised Learning for Scientific Imaging

DGX agent

arXiv:2607.28868v1 Announce Type: new Abstract: Data augmentations define the invariances learned by self-supervised learning (SSL). Standard augmentation pipelines were designed for natural images, y

safetyarxiv-cs-cv
3 Aug 2026
Research

PhyUnfold-Net: Advancing Remote Sensing Change Detection with Physics-Guided Deep Unfolding

DGX agent

arXiv:2603.19566v3 Announce Type: replace Abstract: Bi-temporal change detection is highly sensitive to acquisition discrepancies, including illumination, season, and atmosphere, which often cause fal

researcharxiv-cs-cv
3 Aug 2026
Research

Progressive Checkerboards for Autoregressive Multiscale Image Generation

DGX agent

arXiv:2602.03811v3 Announce Type: replace Abstract: A key challenge in autoregressive image generation is to efficiently sample independent locations in parallel, while still modeling mutual dependenc

researcharxiv-cs-cv
3 Aug 2026
Local Ai

Progressive Decision-Making for Localizing Open-Ended AI-Generated Image Forgeries

DGX agent

arXiv:2607.29156v1 Announce Type: new Abstract: AI-generated image forgeries are becoming increasingly realistic and difficult to characterize with fixed manipulation patterns. As generative models co

local-aiarxiv-cs-cv
3 Aug 2026
Model Releases

RayViT: Ray-Conditioned Visual Representations for Viewpoint-Robust Imitation Learning

DGX agent

arXiv:2607.29622v1 Announce Type: cross Abstract: Visual imitation learning enables robots to acquire visuomotor skills directly from images, yet RGB observations lack explicit geometric cues, making

model-releasesarxiv-cs-cv
3 Aug 2026
Model Releases

ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding

DGX agent

arXiv:2607.28751v1 Announce Type: new Abstract: Universal multimodal embedding (UME) maps heterogeneous multimodal inputs into a shared embedding space. Existing UME models either form embeddings thro

model-releasesarxiv-cs-cv
3 Aug 2026
Research

ReMoE: Report-Guided Mixture-of-Experts for Multimodal OCT/OCTA Anomaly Detection

DGX agent

arXiv:2607.29039v1 Announce Type: new Abstract: Multimodal medical anomaly detection identifies samples deviating from normal patterns, where scarce abnormal cases make normality modeling from normal

researcharxiv-cs-cv
3 Aug 2026
Safety

Rethinking Detection Calibration: A Coordinate and Direction Perspective

DGX agent

arXiv:2607.29040v1 Announce Type: new Abstract: Deep learning based object detectors require trustworthiness beyond competitive detection performance, but deep neural networks are prone to overconfide

safetyarxiv-cs-cv
3 Aug 2026
Research

Role-Break in Attention Heads: Understanding and Detecting Hallucinations in VLMs

DGX agent

arXiv:2607.29412v1 Announce Type: new Abstract: Despite remarkable progress in vision-language generation, Vision-Language Models (VLMs) remain prone to hallucinations, producing content that is incon

researcharxiv-cs-cv
3 Aug 2026
Local Ai

SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs

DGX agent

arXiv:2607.28969v1 Announce Type: new Abstract: Although Large Language Models (LLMs) have demonstrated promising safety performance, extending them to Multimodal Large Language Models (MLLMs) exposes

local-aiarxiv-cs-cv
3 Aug 2026
Model Releases

SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models via Depth-Routed LoRA and Depth Shifting

DGX agent

arXiv:2607.29033v1 Announce Type: new Abstract: Existing methods for adapting 2D foundation models such as SAM to 3D volumes either process slices independently---ignoring inter-slice context---or req

model-releasesarxiv-cs-cv
3 Aug 2026
Safety

SatEdit: Mask-Conditioned Image Editing via VLM-Guided Segment Annotation

DGX agent

arXiv:2607.29367v1 Announce Type: new Abstract: Satellite image editing requires spatially precise object-level control, but supervised editing datasets for overhead imagery are costly to build becaus

safetyarxiv-cs-cv
3 Aug 2026
Model Releases

Scaling Properties of Text Conditioning in Visual Generation

DGX agent

arXiv:2607.29679v1 Announce Type: new Abstract: We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does

model-releasesarxiv-cs-cv
3 Aug 2026
Model Releases

SciFigPlag-Bench: A Benchmark for Provenance-Aware Scientific Figure Plagiarism Detection

DGX agent

arXiv:2607.29124v1 Announce Type: new Abstract: Scientific figures often encode the visual evidence behind scientific findings, yet figure plagiarism remains underexplored as a benchmarked multimodal

model-releasesarxiv-cs-cv
3 Aug 2026
Research

Simulative Anomaly Detection using 2D Tomography

DGX agent

arXiv:2607.28701v1 Announce Type: cross Abstract: We present a novel technique for predicting the imaging quality of anomalies such as cancer cells located inside organic tissues. This technique is us

researcharxiv-cs-cv
3 Aug 2026
Model Releases

So-Fake: Benchmarking and Explaining Social Media Image Forgery Detection

DGX agent

arXiv:2505.18660v5 Announce Type: replace Abstract: Recent advances in AI-powered generative models have enabled the creation of increasingly realistic synthetic images, posing significant risks to in

model-releasesarxiv-cs-cv
3 Aug 2026
Applications

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

DGX agent

arXiv:2607.28993v1 Announce Type: cross Abstract: World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance

applicationsarxiv-cs-cv
3 Aug 2026
Model Releases

StraightDP: Geometry-Aware Differential Privacy for Rectified-Flow Transformers

DGX agent

arXiv:2607.29100v1 Announce Type: cross Abstract: Differentially private (DP) training of text-conditioned generative models suffers a utility cliff at strong privacy. We revisit this problem through

model-releasesarxiv-cs-cv
3 Aug 2026
Model Releases

SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift

DGX agent

arXiv:2607.28996v1 Announce Type: new Abstract: RGB imagery offers a practical, low-cost option for Unmanned Aerial/Ground Vehicle (UAV/UGV) survey support in surface-landmine detection, but object de

model-releasesarxiv-cs-cv
3 Aug 2026
Model Releases

Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution

DGX agent

arXiv:2605.25333v2 Announce Type: replace Abstract: Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption.

model-releasesarxiv-cs-cv
3 Aug 2026
Safety

The K-Space Signature: Frequency-Domain Representation Learning for Medical Deepfake Detection

DGX agent

arXiv:2607.29541v1 Announce Type: new Abstract: In medical imaging, generative models are increasingly deployed to synthesize realistic data and augment limited datasets. Unfortunately, while benefici

safetyarxiv-cs-cv
3 Aug 2026
Tutorials

TOOD: Task-Aware Out-of-Distribution Score Calibration for Continual Learners

DGX agent

arXiv:2607.29592v1 Announce Type: new Abstract: The primary challenge of continual learning (CL) systems is to learn new tasks while remaining performant on previously learned tasks. A similarly impor

tutorialsarxiv-cs-cv
3 Aug 2026
Research

Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark

DGX agent

arXiv:2607.29684v1 Announce Type: new Abstract: Robust low-light imaging remains challenging for the community. Recent studies have explored fusing Near-Infrared (NIR) with noisy RGB to achieve improv

researcharxiv-cs-cv
3 Aug 2026
Applications

Towards Automated Initial Probe Placement in Transthoracic Teleultrasound Using Human Mesh and Skeleton Recovery

DGX agent

arXiv:2603.11257v2 Announce Type: replace Abstract: Cardiac and lung ultrasound are technically demanding because operators must identify patient-specific intercostal acoustic windows and then navigat

applicationsarxiv-cs-cv
3 Aug 2026
← Previous
1…2728293031…261
Next →