AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Safety

WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA

DGX agent

arXiv:2608.01035v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deployment is sev

safetyarxiv-cs-cv
4 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

What Carries the Signal in Pathology Foundation-Model Atlases? A Patient-Level Controlled Benchmark in Breast Cancer

DGX agent

arXiv:2608.00105v1 Announce Type: new Abstract: Pathology foundation models are reported to encode molecular programmes in tissue morphology, but the evidence is usually a cohort-wide ranked gene list

model-releasesarxiv-cs-cv
4 Aug 2026
Research

When Extreme Darkness Meets Motion Blur: MeanFlow for Unified RAW Restoration

DGX agent

arXiv:2608.01720v1 Announce Type: new Abstract: Extremely low-light RAW enhancement aims to recover severely attenuated sensor signals, yet existing methods often focus on illumination and noise while

researcharxiv-cs-cv
4 Aug 2026
Model Releases

When Measurement Conventions Masquerade as Calibration Gains in Cardiac Digital Twins

DGX agent

arXiv:2608.01602v1 Announce Type: new Abstract: Cardiac digital twins convert clinical images into physiological measurements through observation operators, yet calibration studies often assume a fixe

model-releasesarxiv-cs-cv
4 Aug 2026
Local Ai

Where Does Generative Difficulty Reside? An Empirical Study of Target Representations

DGX agent

arXiv:2608.00626v1 Announce Type: new Abstract: The target representation defines the distribution an image generator must learn, yet it is often treated as an interchangeable interface. This assumpti

local-aiarxiv-cs-cv
4 Aug 2026
Safety

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs

DGX agent

arXiv:2608.00076v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from images and text. W

safetyarxiv-cs-cv
4 Aug 2026
Research

WiFuse: An Attention Mechanism for Human Activity Recognition using Fused CSI Amplitude and Delay-Doppler Channel Features

DGX agent

arXiv:2608.00642v1 Announce Type: new Abstract: Recently, Wi-Fi sensing has played a significant role in Human Activity Recognition (HAR), as it enables the detection of various activities using only

researcharxiv-cs-cv
4 Aug 2026
Research

WorldDynCache: Risk-Controlled Latent Dynamics Approximation for Diffusion World Model

DGX agent

arXiv:2608.01845v1 Announce Type: cross Abstract: Diffusion world models generate high-quality futures, but re- peated transformer evaluations make inference prohibitively slow. Existing caches reuse

researcharxiv-cs-cv
4 Aug 2026
Model Releases

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

DGX agent

arXiv:2608.02603v1 Announce Type: new Abstract: Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the appa

model-releasesarxiv-cs-cv
4 Aug 2026
Research

WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting

DGX agent

arXiv:2510.10726v2 Announce Type: replace Abstract: We present WorldMirror, a unified feed-forward model for comprehensive 3D geometric prediction tasks. Unlike existing methods constrained to image-o

researcharxiv-cs-cv
4 Aug 2026
Model Releases

XSPA: Crafting Imperceptible X-Shaped Sparse Adversarial Perturbations for Transferable Attacks on VLMs

DGX agent

arXiv:2603.28568v2 Announce Type: replace Abstract: Vision-language models (VLMs) share visual-textual representations across zero-shot classification, image captioning, and visual question answering

model-releasesarxiv-cs-cv
4 Aug 2026
Research

Zero-Cost Virtual RNA: Approximating Immunotherapy Signatures via Cross-Modal WSI Retrieval

DGX agent

arXiv:2608.00544v1 Announce Type: new Abstract: Identifying the ``Inflamed'' immunophenotype in Gastric Adenocarcinoma predicts immunotherapy response but requires an expensive 10-gene RNA signature.

researcharxiv-cs-cv
4 Aug 2026
Local Ai

A Biometric Sensor Network to Enable Real-Time Measurement of Individual Student Engagement in STEM Lecture Environments

DGX agent

arXiv:2607.28944v1 Announce Type: cross Abstract: Student engagement (SE) is a critical predictor of academic performance and retention in STEM education, yet existing measurement approaches are often

local-aiarxiv-cs-cv
3 Aug 2026
Model Releases

A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples

DGX agent

arXiv:2607.29122v1 Announce Type: new Abstract: Pixel-space diffusion models aim to learn an end-to-end generator directly over raw pixels. This is challenging because a single model must capture both

model-releasesarxiv-cs-cv
3 Aug 2026
Safety

Adaptive Emotional Video Captioning via Affective Heterogeneous Graph Reasoning and Multi-task Joint Learning

DGX agent

arXiv:2607.29045v1 Announce Type: new Abstract: Emotional video captioning (EVC) aims to describe a video with both factual correctness and affective expressiveness. It requires a model to perceive su

safetyarxiv-cs-cv
3 Aug 2026
Research

AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models

DGX agent

arXiv:2505.20255v3 Announce Type: replace Abstract: Recent advances in video diffusion models have substantially enhanced character animation techniques. However, existing methods primarily depend on

researcharxiv-cs-cv
3 Aug 2026
Research

Automated classification method of COVID-19 cases from chest CT volumes using 2D and 3D hybrid CNN for anisotropic volumes

DGX agent

arXiv:2607.28950v1 Announce Type: new Abstract: This paper proposes an automated classification method of chest CT volumes based on likelihood of COVID-19 cases. Novel coronavirus disease 2019 (COVID-

researcharxiv-cs-cv
3 Aug 2026
Model Releases

BWM: A Low-Cost High-Fidelity World Simulator for Robot Learning

DGX agent

arXiv:2607.29302v1 Announce Type: cross Abstract: Reliable robot learning requires a world simulator that can predict action consequences before execution on physical hardware, including risky and fai

model-releasesarxiv-cs-cv
3 Aug 2026
Model Releases

CAER: Conflict-Aware Evidence Routing with Dual Prefix Experts for Multimodal Large Language Models

DGX agent

arXiv:2607.28991v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in multimodal understanding and generation. However, when textual inp

model-releasesarxiv-cs-cv
3 Aug 2026
Research

CALM-AH: An ABAW11-Calibrated Multimodal Ensemble with Reliability-Gated Multi-Expert Consensus for Video-Level Ambivalence and Hesitancy Recognition

DGX agent

arXiv:2607.29310v1 Announce Type: new Abstract: Ambivalence and hesitancy (A/H) are subtle behavioural states that may be expressed through language, voice, facial activity, and other non-verbal cues.

researcharxiv-cs-cv
3 Aug 2026
Research

Can Synthetic Data Overcome the Generalization Limits of AI-Based Flower and Pod Detection Across Cowpea Breeding Genotypes and Environments?

DGX agent

arXiv:2607.28796v1 Announce Type: new Abstract: High-throughput phenotyping requires AI-enabled computer vision models that generalize across genotypes, locations, and growing seasons, yet such models

researcharxiv-cs-cv
3 Aug 2026
Model Releases

CBCT-IQ: A Publicly Available Annotated Cone-Beam CT Dataset for Image Quality Assessment and Benchmarking

DGX agent

arXiv:2607.29253v1 Announce Type: cross Abstract: Medical image quality plays a critical role in diagnostic accuracy, especially in X-ray-based imaging modalities such as cone-beam computed tomography

model-releasesarxiv-cs-cv
3 Aug 2026
Local Ai

Classification of COVID-19 cases from chest CT volumes using hybrid model of 3D CNN and 3D MLP-Mixer

DGX agent

arXiv:2607.28978v1 Announce Type: new Abstract: This paper proposes an automated classification method of COVID-19 chest CT volumes using improved 3D MLP-Mixer. Novel coronavirus disease 2019 (COVID-1

local-aiarxiv-cs-cv
3 Aug 2026
Local Ai

CoDe-SSM: Context-Detail Decoupled State Space Model for Efficient UHD Image Restoration

DGX agent

arXiv:2607.29595v1 Announce Type: new Abstract: Ultra-high-definition (UHD) image restoration must balance the aggregation of spatially recurring degradation cues with the preservation of localized im

local-aiarxiv-cs-cv
3 Aug 2026
Agents

CodeShrink: Adaptive Visual Compression for Efficient Multimodal Code Understanding

DGX agent

arXiv:2607.29637v1 Announce Type: new Abstract: Rendering source code as images offers a promising way to reduce the input costs of Multimodal Large Language Models (MLLMs). Adjusting image resolution

agentsarxiv-cs-cv
3 Aug 2026
Tutorials

Contrastive Learning for Image Complexity Representation

DGX agent

arXiv:2408.03230v2 Announce Type: replace Abstract: Quantifying and evaluating image complexity can be instrumental in enhancing the performance of various computer vision tasks. Supervised learning c

tutorialsarxiv-cs-cv
3 Aug 2026
Research

CorrelationFlow: A Training-Free Geometric Approach for LiDAR Scene Flow Estimation

DGX agent

arXiv:2607.29237v1 Announce Type: new Abstract: LiDAR scene flow estimation has settled into a monoculture: nearly all recent methods share the same feed-forward architecture and the same family of se

researcharxiv-cs-cv
3 Aug 2026
Safety

Deformable Medical Image Registration with KAN-based Implicit Neural Representations

DGX agent

arXiv:2509.22874v2 Announce Type: replace Abstract: Deformable image registration (DIR) is central to medical image analysis, supporting spatial alignment for longitudinal studies and multi-modal fusi

safetyarxiv-cs-cv
3 Aug 2026
Safety

Distance-aware Soft Prompt Guidance for Multimodal Valence-Arousal Estimation

DGX agent

arXiv:2603.13415v2 Announce Type: replace Abstract: Valence-arousal (VA) estimation is crucial for capturing the nuanced nature of human emotions in naturalistic environments. While pre-trained vision

safetyarxiv-cs-cv
3 Aug 2026
Safety

Do Medical Foundation Models Generalize on the African Brain?

DGX agent

arXiv:2607.28771v1 Announce Type: new Abstract: Medical foundation models (FMs) are increasingly used for brain MRI analysis. However, their evaluation remains dominated by high-resource datasets, lea

safetyarxiv-cs-cv
3 Aug 2026
Safety

Domain-Adaptive Deep Joint Source-Channel Coding for Image Classification

DGX agent

arXiv:2607.28907v1 Announce Type: cross Abstract: Deep joint source--channel coding (Deep JSCC) enables visual semantic transmission by mapping inputs directly to channel symbols and task outputs, but

safetyarxiv-cs-cv
3 Aug 2026
Safety

Domain-Division based Progressive Learning for Source-Free Domain Adaptation

DGX agent

arXiv:2607.29202v1 Announce Type: new Abstract: With growing privacy and portability concerns, source-free domain adaptation requires only a source pre-trained model and an unlabeled target domain, al

safetyarxiv-cs-cv
3 Aug 2026
Research

DuET: Dual Expert Trajectories for Diffusion Image Editing

DGX agent

arXiv:2606.13303v2 Announce Type: replace Abstract: Recent diffusion editors perform diverse instruction-based edits while conditioning on the source image at every denoising step. Yet persistent sour

researcharxiv-cs-cv
3 Aug 2026
Safety

DuetHOI: Language-Guided Bimanual Hand--Object Motion Generation with Articulation Planning and Contact Refinement

DGX agent

arXiv:2603.08390v3 Announce Type: replace-cross Abstract: Bimanual articulated-object interaction generation requires a model to capture the evolution of object articulation, coordination between the

safetyarxiv-cs-cv
3 Aug 2026
Safety

DynoDINO: Harnessing Dynamic Latent Information from DINO Features for Multi-Phase Medical Image Segmentation

DGX agent

arXiv:2607.29568v1 Announce Type: new Abstract: Multi-phase Contrast-Enhanced Computed Tomography (CECT) plays a central role in the diagnosis and characterization of focal lesions by capturing tempor

safetyarxiv-cs-cv
3 Aug 2026
Research

EMAG: Self-Rectifying Diffusion Sampling with Exponential Moving Average Guidance

DGX agent

arXiv:2512.17303v2 Announce Type: replace Abstract: In diffusion and flow-matching generative models, guidance techniques are widely used to improve sample quality and consistency. Classifier-free gui

researcharxiv-cs-cv
3 Aug 2026
Model Releases

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

DGX agent

arXiv:2607.29025v1 Announce Type: new Abstract: While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency

model-releasesarxiv-cs-cv
3 Aug 2026
Research

Explaining AI-Image Detection: What the Heatmap Actually Shows

DGX agent

arXiv:2607.29581v1 Announce Type: new Abstract: A marketplace review photograph is a document: platforms approve refunds on it, and generative models drove the cost of forging one to zero. We study th

researcharxiv-cs-cv
3 Aug 2026
Applications

FibVLA: An Efficient Temporal Vision-Language-Action Model with Fibonacci Sampling

DGX agent

arXiv:2607.29596v1 Announce Type: cross Abstract: Vision-language-action models (VLAs), which leverage the cognition of multimodal information to infer physical-world actions, provide a generalized so

applicationsarxiv-cs-cv
3 Aug 2026
Research

FillGS: Filling Observation Gaps in 4D Gaussian Splatting via Viewpoint-Time Selection and Generative Refinement

DGX agent

arXiv:2607.29284v1 Announce Type: new Abstract: 4D Gaussian Splatting (4DGS) can render dynamic scenes photorealistically. However, with limited viewpoint coverage, some spatiotemporal regions remain

researcharxiv-cs-cv
3 Aug 2026
Safety

First Investigation of Deep Learning for Intraoperative Gauze Segmentation in Minimally Invasive Abdominal Surgery

DGX agent

arXiv:2607.29132v1 Announce Type: new Abstract: Surgical gauze is an essential part of surgical procedures, primarily used for controlling bleeding and absorbing bodily fluids. The post-surgical reten

safetyarxiv-cs-cv
3 Aug 2026
Model Releases

FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control

DGX agent

arXiv:2607.29627v1 Announce Type: new Abstract: Generative video compositing, which involves inserting external assets seamlessly into existing video sequences, is essential for content creation and v

model-releasesarxiv-cs-cv
3 Aug 2026
Research

FocusGS: Spatial Delta Layers for Local Repair and Deterministic Editing of Trained 3D Gaussian Assets

DGX agent

arXiv:2607.28834v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) is evolving from one-time reconstruction into deliverable, inspectable, and maintainable visual assets. Existing workflows

researcharxiv-cs-cv
3 Aug 2026
Research

Forwardrobe: Garment-Aware Gaussian Avatars from a Single Image

DGX agent

arXiv:2607.29106v1 Announce Type: new Abstract: Reconstructing animatable 3D human avatars from a single image remains particularly challenging for loose garments, whose geometry and motion cannot be

researcharxiv-cs-cv
3 Aug 2026
Model Releases

GO-PRE: Goal-Oriented Next-Best-View Selection via Predictive Rendering Entropy for Active 3D Reconstruction

DGX agent

arXiv:2607.29037v1 Announce Type: new Abstract: Active 3D reconstruction relies on active view selection to maximize reconstruction fidelity under limited capture budgets. However, most existing metho

model-releasesarxiv-cs-cv
3 Aug 2026
Tutorials

Group-wise Supervision with Focal-Dice Loss for Long-Tailed Indoor Semantic Occupancy Prediction

DGX agent

arXiv:2607.28935v1 Announce Type: new Abstract: Recently, 3D semantic occupancy prediction has garnered increasing attention for understanding the indoor scene. However, unlike structured outdoor envi

tutorialsarxiv-cs-cv
3 Aug 2026
Safety

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering

DGX agent

arXiv:2607.29638v1 Announce Type: new Abstract: Multi-page document visual question answering requires locating sparse evidence at both the page and region levels. Existing approaches typically emphas

safetyarxiv-cs-cv
3 Aug 2026
Research

I3DM: Implicit 3D-aware Memory Retrieval and Injection for Consistent Video Scene Generation

DGX agent

arXiv:2603.23413v2 Announce Type: replace Abstract: Despite remarkable progress in video generation, maintaining long-term scene consistency upon revisiting previously explored areas remains challengi

researcharxiv-cs-cv
3 Aug 2026
← Previous
1…2627282930…261
Next →