AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
4 Aug 2026

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs

SafetyDGX agent

arXiv:2608.00076v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from images and text. W

WiFuse: An Attention Mechanism for Human Activity Recognition using Fused CSI Amplitude and Delay-Doppler Channel Features

ResearchDGX agent

arXiv:2608.00642v1 Announce Type: new Abstract: Recently, Wi-Fi sensing has played a significant role in Human Activity Recognition (HAR), as it enables the detection of various activities using only

WorldDynCache: Risk-Controlled Latent Dynamics Approximation for Diffusion World Model

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.01845v1 Announce Type: cross Abstract: Diffusion world models generate high-quality futures, but re- peated transformer evaluations make inference prohibitively slow. Existing caches reuse

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

Model ReleasesDGX agent

arXiv:2608.02603v1 Announce Type: new Abstract: Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the appa

WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting

ResearchDGX agent

arXiv:2510.10726v2 Announce Type: replace Abstract: We present WorldMirror, a unified feed-forward model for comprehensive 3D geometric prediction tasks. Unlike existing methods constrained to image-o

XSPA: Crafting Imperceptible X-Shaped Sparse Adversarial Perturbations for Transferable Attacks on VLMs

Model ReleasesDGX agent

arXiv:2603.28568v2 Announce Type: replace Abstract: Vision-language models (VLMs) share visual-textual representations across zero-shot classification, image captioning, and visual question answering

Zero-Cost Virtual RNA: Approximating Immunotherapy Signatures via Cross-Modal WSI Retrieval

ResearchDGX agent

arXiv:2608.00544v1 Announce Type: new Abstract: Identifying the ``Inflamed'' immunophenotype in Gastric Adenocarcinoma predicts immunotherapy response but requires an expensive 10-gene RNA signature.

3 Aug 2026

A Biometric Sensor Network to Enable Real-Time Measurement of Individual Student Engagement in STEM Lecture Environments

Local AiDGX agent

arXiv:2607.28944v1 Announce Type: cross Abstract: Student engagement (SE) is a critical predictor of academic performance and retention in STEM education, yet existing measurement approaches are often

A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples

Model ReleasesDGX agent

arXiv:2607.29122v1 Announce Type: new Abstract: Pixel-space diffusion models aim to learn an end-to-end generator directly over raw pixels. This is challenging because a single model must capture both

Adaptive Emotional Video Captioning via Affective Heterogeneous Graph Reasoning and Multi-task Joint Learning

SafetyDGX agent

arXiv:2607.29045v1 Announce Type: new Abstract: Emotional video captioning (EVC) aims to describe a video with both factual correctness and affective expressiveness. It requires a model to perceive su

AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models

ResearchDGX agent

arXiv:2505.20255v3 Announce Type: replace Abstract: Recent advances in video diffusion models have substantially enhanced character animation techniques. However, existing methods primarily depend on

Automated classification method of COVID-19 cases from chest CT volumes using 2D and 3D hybrid CNN for anisotropic volumes

ResearchDGX agent

arXiv:2607.28950v1 Announce Type: new Abstract: This paper proposes an automated classification method of chest CT volumes based on likelihood of COVID-19 cases. Novel coronavirus disease 2019 (COVID-

BWM: A Low-Cost High-Fidelity World Simulator for Robot Learning

Model ReleasesDGX agent

arXiv:2607.29302v1 Announce Type: cross Abstract: Reliable robot learning requires a world simulator that can predict action consequences before execution on physical hardware, including risky and fai

CAER: Conflict-Aware Evidence Routing with Dual Prefix Experts for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2607.28991v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in multimodal understanding and generation. However, when textual inp

CALM-AH: An ABAW11-Calibrated Multimodal Ensemble with Reliability-Gated Multi-Expert Consensus for Video-Level Ambivalence and Hesitancy Recognition

ResearchDGX agent

arXiv:2607.29310v1 Announce Type: new Abstract: Ambivalence and hesitancy (A/H) are subtle behavioural states that may be expressed through language, voice, facial activity, and other non-verbal cues.

Can Synthetic Data Overcome the Generalization Limits of AI-Based Flower and Pod Detection Across Cowpea Breeding Genotypes and Environments?

ResearchDGX agent

arXiv:2607.28796v1 Announce Type: new Abstract: High-throughput phenotyping requires AI-enabled computer vision models that generalize across genotypes, locations, and growing seasons, yet such models

CBCT-IQ: A Publicly Available Annotated Cone-Beam CT Dataset for Image Quality Assessment and Benchmarking

Model ReleasesDGX agent

arXiv:2607.29253v1 Announce Type: cross Abstract: Medical image quality plays a critical role in diagnostic accuracy, especially in X-ray-based imaging modalities such as cone-beam computed tomography

Classification of COVID-19 cases from chest CT volumes using hybrid model of 3D CNN and 3D MLP-Mixer

Local AiDGX agent

arXiv:2607.28978v1 Announce Type: new Abstract: This paper proposes an automated classification method of COVID-19 chest CT volumes using improved 3D MLP-Mixer. Novel coronavirus disease 2019 (COVID-1

CoDe-SSM: Context-Detail Decoupled State Space Model for Efficient UHD Image Restoration

Local AiDGX agent

arXiv:2607.29595v1 Announce Type: new Abstract: Ultra-high-definition (UHD) image restoration must balance the aggregation of spatially recurring degradation cues with the preservation of localized im

CodeShrink: Adaptive Visual Compression for Efficient Multimodal Code Understanding

AgentsDGX agent

arXiv:2607.29637v1 Announce Type: new Abstract: Rendering source code as images offers a promising way to reduce the input costs of Multimodal Large Language Models (MLLMs). Adjusting image resolution

Contrastive Learning for Image Complexity Representation

TutorialsDGX agent

arXiv:2408.03230v2 Announce Type: replace Abstract: Quantifying and evaluating image complexity can be instrumental in enhancing the performance of various computer vision tasks. Supervised learning c

CorrelationFlow: A Training-Free Geometric Approach for LiDAR Scene Flow Estimation

ResearchDGX agent

arXiv:2607.29237v1 Announce Type: new Abstract: LiDAR scene flow estimation has settled into a monoculture: nearly all recent methods share the same feed-forward architecture and the same family of se

Deformable Medical Image Registration with KAN-based Implicit Neural Representations

SafetyDGX agent

arXiv:2509.22874v2 Announce Type: replace Abstract: Deformable image registration (DIR) is central to medical image analysis, supporting spatial alignment for longitudinal studies and multi-modal fusi

Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation

SafetyDGX agent

arXiv:2603.13415v2 Announce Type: replace Abstract: Valence-arousal (VA) estimation is crucial for capturing the nuanced nature of human emotions in naturalistic environments. While pre-trained vision

Do Medical Foundation Models Generalize on the African Brain?

SafetyDGX agent

arXiv:2607.28771v1 Announce Type: new Abstract: Medical foundation models (FMs) are increasingly used for brain MRI analysis. However, their evaluation remains dominated by high-resource datasets, lea

Domain-Adaptive Deep Joint Source-Channel Coding for Image Classification

SafetyDGX agent

arXiv:2607.28907v1 Announce Type: cross Abstract: Deep joint source--channel coding (Deep JSCC) enables visual semantic transmission by mapping inputs directly to channel symbols and task outputs, but

Domain-Division based Progressive Learning for Source-Free Domain Adaptation

SafetyDGX agent

arXiv:2607.29202v1 Announce Type: new Abstract: With growing privacy and portability concerns, source-free domain adaptation requires only a source pre-trained model and an unlabeled target domain, al

DuET: Dual Expert Trajectories for Diffusion Image Editing

ResearchDGX agent

arXiv:2606.13303v2 Announce Type: replace Abstract: Recent diffusion editors perform diverse instruction-based edits while conditioning on the source image at every denoising step. Yet persistent sour

DuetHOI: Language-Guided Bimanual Hand--Object Motion Generation with Articulation Planning and Contact Refinement

SafetyDGX agent

arXiv:2603.08390v3 Announce Type: replace-cross Abstract: Bimanual articulated-object interaction generation requires a model to capture the evolution of object articulation, coordination between the

DynoDINO: Harnessing Dynamic Latent Information from DINO Features for Multi-Phase Medical Image Segmentation

SafetyDGX agent

arXiv:2607.29568v1 Announce Type: new Abstract: Multi-phase Contrast-Enhanced Computed Tomography (CECT) plays a central role in the diagnosis and characterization of focal lesions by capturing tempor

EMAG: Self-Rectifying Diffusion Sampling with Exponential Moving Average Guidance

ResearchDGX agent

arXiv:2512.17303v2 Announce Type: replace Abstract: In diffusion and flow-matching generative models, guidance techniques are widely used to improve sample quality and consistency. Classifier-free gui

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

Model ReleasesDGX agent

arXiv:2607.29025v1 Announce Type: new Abstract: While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency

Explaining AI-Image Detection: What the Heatmap Actually Shows

ResearchDGX agent

arXiv:2607.29581v1 Announce Type: new Abstract: A marketplace review photograph is a document: platforms approve refunds on it, and generative models drove the cost of forging one to zero. We study th

FibVLA: An Efficient Temporal Vision-Language-Action Model with Fibonacci Sampling

ApplicationsDGX agent

arXiv:2607.29596v1 Announce Type: cross Abstract: Vision-language-action models (VLAs), which leverage the cognition of multimodal information to infer physical-world actions, provide a generalized so

FillGS: Filling Observation Gaps in 4D Gaussian Splatting via Viewpoint-Time Selection and Generative Refinement

ResearchDGX agent

arXiv:2607.29284v1 Announce Type: new Abstract: 4D Gaussian Splatting (4DGS) can render dynamic scenes photorealistically. However, with limited viewpoint coverage, some spatiotemporal regions remain

First Investigation of Deep Learning for Intraoperative Gauze Segmentation in Minimally Invasive Abdominal Surgery

SafetyDGX agent

arXiv:2607.29132v1 Announce Type: new Abstract: Surgical gauze is an essential part of surgical procedures, primarily used for controlling bleeding and absorbing bodily fluids. The post-surgical reten

FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control

Model ReleasesDGX agent

arXiv:2607.29627v1 Announce Type: new Abstract: Generative video compositing, which involves inserting external assets seamlessly into existing video sequences, is essential for content creation and v

FocusGS: Spatial Delta Layers for Local Repair and Deterministic Editing of Trained 3D Gaussian Assets

ResearchDGX agent

arXiv:2607.28834v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) is evolving from one-time reconstruction into deliverable, inspectable, and maintainable visual assets. Existing workflows

Forwardrobe: Garment-Aware Gaussian Avatars from a Single Image

ResearchDGX agent

arXiv:2607.29106v1 Announce Type: new Abstract: Reconstructing animatable 3D human avatars from a single image remains particularly challenging for loose garments, whose geometry and motion cannot be

GO-PRE: Goal-Oriented Next-Best-View Selection via Predictive Rendering Entropy for Active 3D Reconstruction

Model ReleasesDGX agent

arXiv:2607.29037v1 Announce Type: new Abstract: Active 3D reconstruction relies on active view selection to maximize reconstruction fidelity under limited capture budgets. However, most existing metho

Group-wise Supervision with Focal-Dice Loss for Long-Tailed Indoor Semantic Occupancy Prediction

TutorialsDGX agent

arXiv:2607.28935v1 Announce Type: new Abstract: Recently, 3D semantic occupancy prediction has garnered increasing attention for understanding the indoor scene. However, unlike structured outdoor envi

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering

SafetyDGX agent

arXiv:2607.29638v1 Announce Type: new Abstract: Multi-page document visual question answering requires locating sparse evidence at both the page and region levels. Existing approaches typically emphas

I3DM: Implicit 3D-aware Memory Retrieval and Injection for Consistent Video Scene Generation

ResearchDGX agent

arXiv:2603.23413v2 Announce Type: replace Abstract: Despite remarkable progress in video generation, maintaining long-term scene consistency upon revisiting previously explored areas remains challengi

Inference-time Trajectory Optimization for Structure-Preserving Manga Image Editing

Model ReleasesDGX agent

arXiv:2603.27790v2 Announce Type: replace Abstract: We present a lightweight, training-free trajectory correction method that adapts a pretrained image editing model to each input manga image using on

Is It Time for the Renaissance of Salient Object Detection in the Era of MLLMs?

Model ReleasesDGX agent

arXiv:2607.29222v1 Announce Type: new Abstract: The zero-shot capabilities of multimodal large language models (MLLMs) are pushing salient object detection (SOD) beyond task-specific supervision. To d

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

SafetyDGX agent

arXiv:2512.13677v2 Announce Type: replace Abstract: In this paper, we present JoVA, a streamlined framework that unifies joint video-audio generation and editing. While existing methods often rely on

Learning from Adversity: Semantic-Aware Mask Refinement through Adversarial Perturbation

ResearchDGX agent

arXiv:2607.29059v1 Announce Type: new Abstract: Despite significant advances in image segmentation, even state-of-the-art models produce masks with imperfect boundaries, semantic inconsistencies, and

Learning Manifolds in High-D Point Embedding for Anisotropic Surface Approximation from Unstructured Point Clouds

ApplicationsDGX agent

arXiv:2607.28855v1 Announce Type: cross Abstract: Dense 3D sensors in various real-world fields produce point clouds that are geometrically redundant for real-time processing. In this paper, we propos

LegoQ: Density-Matrix Representation Learning with Spectral-Spatial State Transitions for Hyperspectral Classification

ResearchDGX agent

arXiv:2607.28970v1 Announce Type: new Abstract: Hyperspectral image classification is complicated by mixed pixels, spectral ambiguity, class imbalance, and limited annotations. Most current classifier

Leveraging Transfer Learning with Class-Specific Decoders for Laparoscopic Segmentation

ResearchDGX agent

arXiv:2607.29509v1 Announce Type: new Abstract: Effective multi-organ segmentation in surgical data requires learning the intricate anatomical features and alleviating the challenge of class imbalance

Lightweight Neural Networks for Affordance Segmentation: Enhancement of the Decoder Module

ApplicationsDGX agent

arXiv:2607.29473v1 Announce Type: new Abstract: The deployment of deep neural networks for visual affordance segmentation on wearable robots poses may prove critical, due to some conflicting aspects o

Locally Consistent Transductive Information Maximization for Few-Shot Remote Sensing Scene Classification

Model ReleasesDGX agent

arXiv:2607.29192v1 Announce Type: new Abstract: Remote sensing scene classification is increasingly relying on foundation models pre-trained on large-scale Earth-observation data. Moreover, transducti

Meshy T2: Fast Native Mesh Generation with Flow Matching

ResearchDGX agent

arXiv:2607.28675v1 Announce Type: cross Abstract: Polygonal meshes are the standard surface representation of modern 3D pipelines, and generating high-quality meshes with artist-style topology is esse

MHRGait: Gait Recognition from Momentum Human Rig Pose

ResearchDGX agent

arXiv:2607.29083v1 Announce Type: new Abstract: Gait recognition is shaped by its input representation. Silhouettes encode projected body shape, skeletons encode sparse joint coordinates, and 3D meshe

Mirror Learning

SafetyDGX agent

arXiv:2607.28737v1 Announce Type: cross Abstract: We investigate imitation learning through the lens of third-person observation and propose a framework for mirror learning: acquiring actionable polic

Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift

Local AiDGX agent

arXiv:2607.28696v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) can retain high observed marginal coverage after clinical shift while substantially under-covering an individual

Moment kernels: a simple and scalable approach for equivariance to rotations and reflections in deep convolutional networks

ResearchDGX agent

arXiv:2505.21736v2 Announce Type: replace Abstract: Translation equivariance is a central reason convolutional neural networks have been successful in computer vision. Other symmetries, such as rotati

MoPET: Parameter-Efficient Mixture-of-Experts for Unified Medical Image Classification

Model ReleasesDGX agent

arXiv:2607.29462v1 Announce Type: cross Abstract: Adapting deep learning models to profound clinical heterogeneity typically relies on parameter-efficient fine-tuning (PEFT) to avoid the severe overfi

MoRoute: Dynamic Routing for In-Context Multimodal Video Generation

Model ReleasesDGX agent

arXiv:2607.29545v1 Announce Type: new Abstract: Multimodal video generation aims to generate and edit videos conditioned on arbitrary combinations of text, images, and videos within a single model, al

Multi-Modal Object Re-Identification with Dual Semantic Guidance and Global-Local Mutual Modulation

SafetyDGX agent

arXiv:2607.29207v1 Announce Type: new Abstract: Multi-modal object Re-Identification (ReID) aims to retrieve target instances by leveraging complementary information across modalities. However, existi

← Previous
1…1920212223…207
Next →