AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
3 Aug 2026

Multi-Source Multi-View Graph Domain Adaptation with Hyperbolic Residual Encoding for Cross-Site MDD Identification from rs-fMRI

SafetyDGX agent

arXiv:2607.29531v1 Announce Type: new Abstract: Cross-site identification of major depressive disorder (MDD) from resting-state functional magnetic resonance imaging (rs-fMRI) is hindered by inter-sit

OASIS: Occlusion-aware Single-image Hand Avatar Reconstruction via 3D Gaussian Splatting

ResearchDGX agent

arXiv:2607.29633v1 Announce Type: new Abstract: Single-image 3D hand avatar reconstruction is fundamentally ill-posed and particularly challenging due to limited visual evidence under severe self-occl

On the Efficacy of Self-Supervised Point Cloud Encoders for Efficient 3D Large Language Models

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.29136v1 Announce Type: new Abstract: 3D point cloud-language models (3D-LLMs) enable 3D understanding by pairing point cloud encoders with large language models, but existing methods rely o

Optical Flow Sensor: A Direction-Selective Bionic Retina Design

ResearchDGX agent

arXiv:2607.28686v1 Announce Type: cross Abstract: Optical flow characterizes motion in the visual field and is fundamental to motion perception and tracking in biological and artificial vision systems

OSAGEN: Object-Aware Mask Priors and Multistage Decoupled Diffusion for Industrial Anomaly Generation

Model ReleasesDGX agent

arXiv:2607.29533v1 Announce Type: new Abstract: Industrial anomaly detection and localization are limited by scarce real anomalies and pixel-level annotations, a bottleneck that synthetic image-mask p

OSEF: One-Step Evidence Fusion for Cross-Video Scene Procedure Planning

Model ReleasesDGX agent

arXiv:2607.29401v1 Announce Type: new Abstract: Video Scene Procedure Planning (VSPP) supplies the target start-goal observations in advance, leaving open how a planner should act when the evidence mu

Parameter-Efficient Fine-Tuning for Spiking Point Cloud Models

Model ReleasesDGX agent

arXiv:2607.29048v1 Announce Type: new Abstract: Spiking Neural Networks (SNNs) offer energy-efficient solutions for point cloud analysis on resource-constrained devices through event-driven computatio

Physics-Aligned Self-Supervised Learning for Scientific Imaging

SafetyDGX agent

arXiv:2607.28868v1 Announce Type: new Abstract: Data augmentations define the invariances learned by self-supervised learning (SSL). Standard augmentation pipelines were designed for natural images, y

PhyUnfold-Net: Advancing Remote Sensing Change Detection with Physics-Guided Deep Unfolding

ResearchDGX agent

arXiv:2603.19566v3 Announce Type: replace Abstract: Bi-temporal change detection is highly sensitive to acquisition discrepancies, including illumination, season, and atmosphere, which often cause fal

Progressive Checkerboards for Autoregressive Multiscale Image Generation

ResearchDGX agent

arXiv:2602.03811v3 Announce Type: replace Abstract: A key challenge in autoregressive image generation is to efficiently sample independent locations in parallel, while still modeling mutual dependenc

Progressive Decision-Making for Localizing Open-Ended AI-Generated Image Forgeries

Local AiDGX agent

arXiv:2607.29156v1 Announce Type: new Abstract: AI-generated image forgeries are becoming increasingly realistic and difficult to characterize with fixed manipulation patterns. As generative models co

RayViT: Ray-Conditioned Visual Representations for Viewpoint-Robust Imitation Learning

Model ReleasesDGX agent

arXiv:2607.29622v1 Announce Type: cross Abstract: Visual imitation learning enables robots to acquire visuomotor skills directly from images, yet RGB observations lack explicit geometric cues, making

ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding

Model ReleasesDGX agent

arXiv:2607.28751v1 Announce Type: new Abstract: Universal multimodal embedding (UME) maps heterogeneous multimodal inputs into a shared embedding space. Existing UME models either form embeddings thro

ReMoE: Report-Guided Mixture-of-Experts for Multimodal OCT/OCTA Anomaly Detection

ResearchDGX agent

arXiv:2607.29039v1 Announce Type: new Abstract: Multimodal medical anomaly detection identifies samples deviating from normal patterns, where scarce abnormal cases make normality modeling from normal

Rethinking Detection Calibration: A Coordinate and Direction Perspective

SafetyDGX agent

arXiv:2607.29040v1 Announce Type: new Abstract: Deep learning based object detectors require trustworthiness beyond competitive detection performance, but deep neural networks are prone to overconfide

Role-Break in Attention Heads: Understanding and Detecting Hallucinations in VLMs

ResearchDGX agent

arXiv:2607.29412v1 Announce Type: new Abstract: Despite remarkable progress in vision-language generation, Vision-Language Models (VLMs) remain prone to hallucinations, producing content that is incon

SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs

Local AiDGX agent

arXiv:2607.28969v1 Announce Type: new Abstract: Although Large Language Models (LLMs) have demonstrated promising safety performance, extending them to Multimodal Large Language Models (MLLMs) exposes

SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models via Depth-Routed LoRA and Depth Shifting

Model ReleasesDGX agent

arXiv:2607.29033v1 Announce Type: new Abstract: Existing methods for adapting 2D foundation models such as SAM to 3D volumes either process slices independently---ignoring inter-slice context---or req

SatEdit: Mask-Conditioned Image Editing via VLM-Guided Segment Annotation

SafetyDGX agent

arXiv:2607.29367v1 Announce Type: new Abstract: Satellite image editing requires spatially precise object-level control, but supervised editing datasets for overhead imagery are costly to build becaus

Scaling Properties of Text Conditioning in Visual Generation

Model ReleasesDGX agent

arXiv:2607.29679v1 Announce Type: new Abstract: We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does

SciFigPlag-Bench: A Benchmark for Provenance-Aware Scientific Figure Plagiarism Detection

Model ReleasesDGX agent

arXiv:2607.29124v1 Announce Type: new Abstract: Scientific figures often encode the visual evidence behind scientific findings, yet figure plagiarism remains underexplored as a benchmarked multimodal

Simulative Anomaly Detection using 2D Tomography

ResearchDGX agent

arXiv:2607.28701v1 Announce Type: cross Abstract: We present a novel technique for predicting the imaging quality of anomalies such as cancer cells located inside organic tissues. This technique is us

So-Fake: Benchmarking and Explaining Social Media Image Forgery Detection

Model ReleasesDGX agent

arXiv:2505.18660v5 Announce Type: replace Abstract: Recent advances in AI-powered generative models have enabled the creation of increasingly realistic synthetic images, posing significant risks to in

ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

ApplicationsDGX agent

arXiv:2607.28993v1 Announce Type: cross Abstract: World Action Models (WAMs) have emerged as a promising paradigm by jointly modeling robot actions and future visual dynamics. However, their reliance

StraightDP: Geometry-Aware Differential Privacy for Rectified-Flow Transformers

Model ReleasesDGX agent

arXiv:2607.29100v1 Announce Type: cross Abstract: Differentially private (DP) training of text-conditioned generative models suffers a utility cliff at strong privacy. We revisit this problem through

SULAND v2: A Refined RGB Dataset and Deep Learning Object Detection Benchmark for UAV/UGV-Based SUrface LANDmine Detection Under Domain Shift

Model ReleasesDGX agent

arXiv:2607.28996v1 Announce Type: new Abstract: RGB imagery offers a practical, low-cost option for Unmanned Aerial/Ground Vehicle (UAV/UGV) survey support in surface-landmine detection, but object de

Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution

Model ReleasesDGX agent

arXiv:2605.25333v2 Announce Type: replace Abstract: Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption.

The K-Space Signature: Frequency-Domain Representation Learning for Medical Deepfake Detection

SafetyDGX agent

arXiv:2607.29541v1 Announce Type: new Abstract: In medical imaging, generative models are increasingly deployed to synthesize realistic data and augment limited datasets. Unfortunately, while benefici

TOOD: Task-Aware Out-of-Distribution Score Calibration for Continual Learners

TutorialsDGX agent

arXiv:2607.29592v1 Announce Type: new Abstract: The primary challenge of continual learning (CL) systems is to learn new tasks while remaining performant on previously learned tasks. A similarly impor

Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark

ResearchDGX agent

arXiv:2607.29684v1 Announce Type: new Abstract: Robust low-light imaging remains challenging for the community. Recent studies have explored fusing Near-Infrared (NIR) with noisy RGB to achieve improv

Towards Automated Initial Probe Placement in Transthoracic Teleultrasound Using Human Mesh and Skeleton Recovery

ApplicationsDGX agent

arXiv:2603.11257v2 Announce Type: replace Abstract: Cardiac and lung ultrasound are technically demanding because operators must identify patient-specific intercostal acoustic windows and then navigat

TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking

SafetyDGX agent

arXiv:2604.01207v2 Announce Type: replace Abstract: Existing 3D Gaussian Splatting (3DGS) editing methods primarily focus on appearance modification and often struggle to support flexible geometry edi

Training-Free Entity-Level Few-Shot Segmentation of Remote Sensing Images with Advection Refinement

ResearchDGX agent

arXiv:2607.29278v1 Announce Type: new Abstract: Existing cross-domain few-shot segmentation approaches suffer from high training costs due to source-domain episodic training and pixel-wise dense predi

UltraSAM3: A Concept-Driven Foundation Model for Universal Ultrasound Image Segmentation

AgentsDGX agent

arXiv:2607.29200v1 Announce Type: new Abstract: Ultrasound imaging has become increasingly widespread in clinical practice due to its portability, low cost and real-time capability, making ultrasound

Uncertainty-Aware Deepfake Detection via Multi-View Structural Learning

ResearchDGX agent

arXiv:2607.28769v1 Announce Type: new Abstract: Security-critical biometric and forensic applications require accurate predictions and reliable confidence estimates, particularly under distribution sh

VFAD: Variational Semantic Prompting Meets Frequency-Adaptive Representation Learning for Zero-Shot Anomaly Detection

SafetyDGX agent

arXiv:2607.29370v1 Announce Type: new Abstract: Zero-shot anomaly detection (ZSAD) aims to detect and localize anomalies in unseen categories without access to target-specific training data. Although

Visual Distribution Anchoring for Efficient Prompt Tuning

ResearchDGX agent

arXiv:2607.28967v1 Announce Type: new Abstract: Prompt tuning adapts vision--language models with few trainable parameters, but existing approaches trade off efficiency and adaptation: static textual

WaMo: Wavelet-Enhanced Multi-Frequency Trajectory Analysis for Fine-Grained Text-Motion Retrieval

SafetyDGX agent

arXiv:2508.03343v2 Announce Type: replace Abstract: Text-Motion Retrieval (TMR) aims to retrieve 3D motion sequences semantically relevant to text descriptions. However, matching 3D motions with text

Weight-Space Mixture-of-Experts for Implicit Neural Representation Classification

ResearchDGX agent

arXiv:2607.29463v1 Announce Type: new Abstract: Implicit Neural Representations (INRs) encode signals as the weights of a coordinate-based neural network and have recently been proposed as an alternat

31 Jul 2026

4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans

ResearchDGX agent

arXiv:2607.27634v1 Announce Type: new Abstract: Generating high-quality 360-degree dynamic human assets from text prompts is challenging. Existing methods usually synthesize monocular or multi-view vi

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

Model ReleasesDGX agent

arXiv:2607.28625v1 Announce Type: new Abstract: Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, o

AdaAnchor4D: Anchor-Conditioned Spatiotemporal Feature Aggregation for Monocular UAV 4D Reconstruction

ResearchDGX agent

arXiv:2607.28320v1 Announce Type: new Abstract: Monocular UAV videos provide valuable observations for dynamic reconstruction of complex urban scenes. However, such scenes exhibit pronounced spatiotem

ARD-REFSM: Enhancing Reflection Symmetry Detection with Asymmetric Denoising and Rotation Equivariance

Model ReleasesDGX agent

arXiv:2607.27927v1 Announce Type: new Abstract: Reflection symmetry detection remains challenging due to interference from asymmetric regions and arbitrary orientations of symmetric patterns. Asymmetr

Articulated Object Reconstruction from Rest-State Observation

ResearchDGX agent

arXiv:2607.27749v1 Announce Type: new Abstract: Building interactive digital twins requires recovering both 3D geometry and the kinematic structures that govern how objects articulate. Yet existing me

AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans

TutorialsDGX agent

arXiv:2607.28487v1 Announce Type: new Abstract: Fine-grained segmentation of auricular structures in CT is challenging because the ear occupies a small image region, cartilage boundaries are highly ir

Backbone-Agnostic Stochastic Perturbation Learning for End-to-End Real-World Image Dehazing

TutorialsDGX agent

arXiv:2607.11623v3 Announce Type: replace Abstract: Real-world paired image dehazing remains challenging because haze degradation is spatially non-uniform, illumination-dependent, and physically ambig

BCNet: Bronchus Classification via Structure Guided Representation Learning

Model ReleasesDGX agent

arXiv:2205.06947v3 Announce Type: replace-cross Abstract: CT-based bronchial tree analysis is essential for diagnosing lung and airway diseases, yet automatic bronchus classification remains challengi

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

AgentsDGX agent

arXiv:2607.28595v1 Announce Type: new Abstract: The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather tha

Benchmarking Foundation and Large Language Models for Few-Shot Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2607.27856v1 Announce Type: new Abstract: Few-shot medical image segmentation (FS-MIS) aims to segment novel regions of interest (ROIs) from a few annotated support examples. Despite rapid progr

Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures

ResearchDGX agent

arXiv:2607.28007v1 Announce Type: new Abstract: Pathology foundation models (FMs) are models trained on vast amounts of typically unlabeled data and have been shown to yield regularized latent spaces

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding

Model ReleasesDGX agent

arXiv:2607.28516v1 Announce Type: new Abstract: Long-video understanding commonly compresses videos into a small set of frames or visual tokens for answer generation. Existing compact pipelines focus

Beyond Visual Ambiguity: Guiding Robust Monocular Depth Estimation in Challenging Scenarios via Detailed Long Captions

ResearchDGX agent

arXiv:2607.28285v1 Announce Type: new Abstract: Monocular depth estimation (MDE) faces challenges with non-Lambertian surfaces and adverse weather conditions due to the visual ambiguities inherent in

BladeYOLO: Wind Turbine Blade Defect Detection with Limited Annotations and Weak-Saliency Awareness

ApplicationsDGX agent

arXiv:2607.28065v1 Announce Type: new Abstract: Wind turbine blade defect detection remains highly challenging in real-world inspection scenarios due to limited on-site data and the subtle visual char

BlindPSNR: A No-Reference Fidelity Predictor for Low-Light Image Enhancement

Model ReleasesDGX agent

arXiv:2607.27628v1 Announce Type: new Abstract: Low-light image enhancement (LLIE) methods involve tunable parameters that are typically fixed, often leading to performance degradation when applied ac

Bunraku: Turning a Single Illustration into an Editable Live2D Character

Model ReleasesDGX agent

arXiv:2607.27348v1 Announce Type: new Abstract: Live2D is the dominant 2D character-animation format for anime characters and virtual avatars, representing each character as a stack of RGBA layers dri

Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs

ResearchDGX agent

arXiv:2607.27700v1 Announce Type: new Abstract: Large Vision-Language Models (VLMs) suffer from prohibitive inference overhead due to long sequences of visual tokens. However, existing visual token re

Can Vision-Language Models Reason about AI Edits in Images?

Local AiDGX agent

arXiv:2607.28464v1 Announce Type: new Abstract: Detection and localization of AI-tampered images are critical for trustworthy AI, yet modern generative models have made such manipulations increasingly

Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models

Local AiDGX agent

arXiv:2607.28341v1 Announce Type: new Abstract: While visual token pruning is essential for efficient Multimodal Large Language Models (MLLMs), existing training-free methods suffer from a critical li

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

Model ReleasesDGX agent

arXiv:2607.28611v1 Announce Type: new Abstract: Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibi

Collaborative Feature Aggregation for Face Super-Resolution and Robust Re-Identification

ResearchDGX agent

arXiv:2607.28130v1 Announce Type: new Abstract: We propose a novel collaborative approach for face super-resolution (SR) and robust person re-identification from sequential or multi-view facial images

← Previous
1…2021222324…207
Next →