AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
31 Jul 2026

Enhancing Scene Transition Awareness in Video Generation via Post-Training

TutorialsDGX agent

arXiv:2507.18046v2 Announce Type: replace Abstract: Recent advances in AI-generated video have shown strong performance on text-to-video tasks, particularly for short clips depicting a single scene. H

Explaining Image Similarity with Automatically Extracted Concept Activation Vectors

Local AiDGX agent

arXiv:2607.28386v1 Announce Type: new Abstract: Image similarity underlies many computer vision applications, yet it is often unclear why two images receive a high or low similarity score. Existing ex

Face and Voice Cross-modal Association with Learning Convex Feature Embedding

ResearchDGX agent

arXiv:2607.28129v1 Announce Type: new Abstract: Face-and-voice association learning is one of the most challenging tasks in deep learning. In this paper, we propose a simple but powerful cross-modal f


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification

AgentsDGX agent

arXiv:2607.28225v1 Announce Type: new Abstract: Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, h

FDDWAN: A Frequency-Decoupled Diffusion Network for Watermarking Attack

ResearchDGX agent

arXiv:2607.27800v1 Announce Type: new Abstract: Existing invisible watermark removal methods often struggle to accurately capture the watermark-bearing features, leading to an unfavorable trade-off be

FeatFix: Reuse What You Verify through Local Exact-Feature Correction for Faster Cached Diffusion Inference

Local AiDGX agent

arXiv:2607.27842v1 Announce Type: new Abstract: Diffusion models are widely used to generate high-quality images and videos, but their iterative denoising process remains computationally intensive. A

Filling the Pareto-Optimal Front for Affordance Segmentation on Embedded Devices Using RGB-D Cameras

ApplicationsDGX agent

arXiv:2607.28293v1 Announce Type: new Abstract: While depth sensors have the potential to complement RGB data for affordance segmentation in wearable robots, their usage seems to remain underexplored.

Finding Change in Satellite Archives from Text: How to Combine Before-and-After Images Efficiently

TutorialsDGX agent

arXiv:2607.28571v1 Announce Type: new Abstract: Operational Earth observation increasingly calls for answering queries such as ``find the image pairs where a new building appeared.'' This means search

FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval

SafetyDGX agent

arXiv:2607.27959v1 Announce Type: new Abstract: Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significa

FlexiGrad: Adaptive Gradient Modulation for Hierarchical Fine-Grained Classification

Model ReleasesDGX agent

arXiv:2607.17563v2 Announce Type: replace Abstract: Many fine-grained recognition tasks contain hierarchical labels such as order, family and species. Although this supervision should be beneficial, j

FootprintNet: State-Transition-Guided Dynamic Footprint Learning for Multi-temporal Remote Sensing Change Detection

TutorialsDGX agent

arXiv:2607.27969v1 Announce Type: new Abstract: Despite substantial progress in remote sensing multi-temporal change detection (MTCD), most existing MTCD methods still represent the dynamic process at

GVR-Coder: A Visual-Feedback Framework for Structured SVG Generation in Complex Document and Meeting Scenarios

AgentsDGX agent

arXiv:2607.28073v1 Announce Type: cross Abstract: In demanding professional environments and meeting review scenarios, lengthy text often imposes a high cognitive load. To facilitate efficient informa

Hallucinations Leave a Grounding Signature:Verifier-Guided Decoding for Selective Object Correction

ResearchDGX agent

arXiv:2607.27823v1 Announce Type: new Abstract: Large vision-language models (LVLMs) often hallucinate objects that are absent from an image. Despite recent progress, existing mitigation methods still

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer

ResearchDGX agent

arXiv:2607.28394v1 Announce Type: new Abstract: Hand-object interaction (HOI) modeling remains challenging because it requires joint reasoning about hand articulation, object geometry, contact, semant

Human Mesh Modeling for Anny Body

ResearchDGX agent

arXiv:2511.03589v3 Announce Type: replace Abstract: Parametric body models provide the structural basis for many human-centric tasks, yet existing models often rely on costly 3D scans and learned shap

ID-Guard: A Universal Framework for Combating Facial Manipulation via Breaking Identification

ResearchDGX agent

arXiv:2409.13349v3 Announce Type: replace Abstract: The misuse of deep learning-based facial manipulation poses a serious threat to civil rights. To prevent such fraud at its source, proactive defense

IGME: Efficient Chained Method Ensemble for Transferable Semantic Segmentation Attacks

ResearchDGX agent

arXiv:2607.27465v1 Announce Type: new Abstract: Semantic segmentation models are vulnerable to transferable adversarial perturbations, yet evaluating transfer attacks on dense prediction models can be

Improved Classification of Nitrogen Stress Severity in Plants Under Combined Stress Conditions Using Spatio-Temporal Deep Learning Framework

ResearchDGX agent

arXiv:2509.06625v3 Announce Type: replace Abstract: Plants in their natural habitats endure an array of interacting stresses, both biotic and abiotic, that rarely occur in isolation. Nutrient stress-p

Inference-Time Agentic Decision Rules Beat Longer Evolving Search for Multi-Image Medical Reasoning

SafetyDGX agent

arXiv:2607.27564v1 Announce Type: new Abstract: Multi-image medical VQA is not merely a prompt-length problem; it is a fundamental challenge of agentic decision-making. Medical vision-language agents

Isolating to Harness: Cross-Division Distillation for Fully Unsupervised Anomaly Detection

TutorialsDGX agent

arXiv:2508.18007v2 Announce Type: replace Abstract: Fully Unsupervised Anomaly Detection (FUAD) addresses the practical scenario where training data is contaminated with unlabeled anomalies. This sett

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Model ReleasesDGX agent

arXiv:2607.27670v1 Announce Type: new Abstract: Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that creat

Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image Classification

Local AiDGX agent

arXiv:2607.28428v1 Announce Type: cross Abstract: We introduce Kohn--Sham Spectral Embedding (KSSE), a physics-inspired energy-based model replacing dense CNN classifiers with a sparse-graph spectral

Landmark shape spaces with induced metrics

ResearchDGX agent

arXiv:2607.28064v1 Announce Type: new Abstract: We present a unification of Kendall's landmark shape spaces, where rigid motions are factored out and scale fixed on landmark configurations equipped wi

Large scale cross-regional remote sensing flood monitoring framework for operative mapping and impact analysis

ResearchDGX agent

arXiv:2607.28401v1 Announce Type: new Abstract: Effective flood monitoring is critical for minimizing the impacts of flood disasters on populations and infrastructure. Yet reliable remote sensing acro

LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference

ResearchDGX agent

arXiv:2607.27952v1 Announce Type: new Abstract: Multimodal foundation models are reshaping edge-cloud visual intelligence from task-specific feature pipelines into token-based interfaces, where edge d

Learning Color Grading, No Photo Sharing: Federated Aesthetic Preference Learning for Personalized Image Enhancement

Model ReleasesDGX agent

arXiv:2607.27659v1 Announce Type: new Abstract: Personalized image enhancement should reflect individual aesthetic taste, yet learning such preferences commonly depends on private photos and ratings t

Learning to Understand Body Language from Flight through Robust 3D Avatar Placing

TutorialsDGX agent

arXiv:2607.27865v1 Announce Type: new Abstract: Perceiving human motion and intent at long range is a prerequisite for socially intelligent aerial robots, yet the data to learn it barely exists. We in

LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA

Model ReleasesDGX agent

arXiv:2607.27806v1 Announce Type: new Abstract: In clinical practice, patients often undergo multiple imaging examinations over successive visits, yielding longitudinal data. Modeling such temporal in

MarkushGlyph and OCSRGlyph: Improved Chemical Structure Recognition

ResearchDGX agent

arXiv:2607.28532v1 Announce Type: new Abstract: Chemical structures appear in patents and the scientific literature as images. For programmatic usage, such as indexing in databases or constructing mac

MedXplore: Towards Reliable and Unbiased Generalized Category Discovery in Medical Imaging

SafetyDGX agent

arXiv:2607.27620v1 Announce Type: new Abstract: Deep learning has shown strong potential in medical image analysis, but most existing methods rely on large-scale annotations and a closed-world assumpt

MeshFM: 2D Features Are All You Need for 3D Shape Understanding

ResearchDGX agent

arXiv:2607.27592v1 Announce Type: new Abstract: We present MeshFM, an efficient feedforward framework for extracting rich features from 3D inputs. Our method distills 2D features from visual foundatio

MetaRank: Task-Aware Metric Selection for Model Transferability Estimation

TutorialsDGX agent

arXiv:2511.21007v2 Announce Type: replace Abstract: Selecting an appropriate pre-trained source model is a critical, yet computationally expensive, task in transfer learning. Model Transferability Est

MIND: Multimodal Intent-Driven Network via Diffusion Transformers for Medical Image Fusion

SafetyDGX agent

arXiv:2607.28565v1 Announce Type: new Abstract: Medical image fusion aims to integrate complementary information from diverse imaging modalities to support clinical diagnosis. Existing methods typical

MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision Transformers

ResearchDGX agent

arXiv:2607.28589v1 Announce Type: new Abstract: Post-training quantization (PTQ) has emerged as an effective solution for deploying Vision Transformers (ViTs) on resource-constrained devices. However,

MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos

Model ReleasesDGX agent

arXiv:2607.27895v1 Announce Type: cross Abstract: Mental health understanding in long-form videos requires nuanced reasoning over observable behavior, interpersonal context, and latent psychological s

MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2607.27637v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or sh

mmRadarTwin: A Measurement-Calibrated Signal-Level Digital Twin Platform for Indoor mmWave Radar

ResearchDGX agent

arXiv:2607.28108v1 Announce Type: new Abstract: Indoor mmWave radar perception is difficult to reproduce because measured range-angle responses depend on scene geometry, material response, multipath,

MonoVoc: Decoupling Geometry and Semantics for Lightweight Monocular Open-Vocabulary 3D Gaussians

ResearchDGX agent

arXiv:2607.28300v1 Announce Type: new Abstract: Open vocabulary 3D scene understanding is essential for next-generation interactive systems, empowering users to intuitively query and navigate reconstr

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

Model ReleasesDGX agent

arXiv:2511.12449v3 Announce Type: replace Abstract: Recent Multimodal Large Language Models (MLLMs) have significantly advanced e-commerce product understanding. However, they still face three challen

Morphological Detection and Classification of Microplastics and Nanoplastics Emerged from Consumer Products by Deep Learning

ResearchDGX agent

arXiv:2409.13688v2 Announce Type: replace Abstract: Plastic pollution presents an escalating global issue, impacting health and environmental systems, with micro- and nanoplastics found across mediums

MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing

Model ReleasesDGX agent

arXiv:2607.27616v1 Announce Type: new Abstract: Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into share

MRD: Using Physically Based Differentiable Rendering to Probe Vision Models for 3D Scene Understanding

ResearchDGX agent

arXiv:2512.12307v5 Announce Type: replace Abstract: While deep learning methods have achieved impressive success in many vision benchmarks, it remains difficult to understand and explain the represent

MSCM-net: A hyperspectral image classiffcation method based on multi-scale convolution and Mamba

Model ReleasesDGX agent

arXiv:2607.28277v1 Announce Type: new Abstract: Hyperspectral imaging is widely used in remote sensing and engineering. Therefore, research on its classification methods is crucial. While CNN and Tran

MUL-T: Decoding Spatial Cellular Architecture in Multiplexed Tissue Images

ResearchDGX agent

arXiv:2607.28030v1 Announce Type: cross Abstract: Understanding tissue organisation in multiplexed imaging requires modelling both cellular phenotypes and their spatial context. Existing approaches ty

Negative controls reveal volume-driven confounding in radiomics and imaging foundation model features

ResearchDGX agent

arXiv:2607.28423v1 Announce Type: new Abstract: Radiomics and imaging foundation models promise non-invasive biomarkers of tumour biology, yet predictive signatures may reflect tumour volume or acquis

Now You Have My Healthy Attention: A U-DiT for Brain-MRI Inpainting

Local AiDGX agent

arXiv:2607.27974v1 Announce Type: new Abstract: The ASNR-MICCAI BraTS Local Synthesis (Inpainting) task asks for the anatomically plausible completion of healthy brain tissue within a masked region of

Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA

Model ReleasesDGX agent

arXiv:2607.27566v1 Announce Type: new Abstract: Multi-frame medical VQA appears to reward increasingly complex adaptation: controller-style inference, localization-aware reranking, static hard-negativ

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

HardwareDGX agent

arXiv:2607.28312v1 Announce Type: new Abstract: Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches prima

ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow

SafetyDGX agent

arXiv:2607.27924v1 Announce Type: cross Abstract: In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning paradigms for world modeling are lar

One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting

Model ReleasesDGX agent

arXiv:2607.27902v1 Announce Type: new Abstract: Scene text spotting requires high-precision alignment between textual recognition and spatial localization. While visual-token grounding has emerged as

OPLD: On-Policy Latent Distillation for Multimodal Reasoning

SafetyDGX agent

arXiv:2607.28154v1 Announce Type: new Abstract: Interleaved multimodal Chain-of-Thought (CoT) improves visual reasoning by incorporating auxiliary visual evidence into intermediate reasoning. However,

OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation

Model ReleasesDGX agent

arXiv:2607.27278v1 Announce Type: new Abstract: Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchm

PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology

Model ReleasesDGX agent

arXiv:2607.27378v1 Announce Type: new Abstract: Accurate evaluation of multimodal large language models (MLLMs) in dental panoramic radiography (orthopantomogram, OPG) is limited by the lack of fine-g

PhiZero: A World Model Built Around Physical Language

TutorialsDGX agent

arXiv:2607.28624v1 Announce Type: new Abstract: We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing phys

Physical prior guided cooperative learning framework for joint turbulence degradation estimation and infrared video restoration

ResearchDGX agent

arXiv:2408.04227v2 Announce Type: replace-cross Abstract: Infrared imaging and turbulence strength measurements are in widespread demand in many fields. This paper introduces a Physical Prior Guided C

PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation

SafetyDGX agent

arXiv:2506.21076v4 Announce Type: replace Abstract: Pose stylization, which aims to synthesize stylized content aligning with target poses, serves as a fundamental task across 2D, 3D, and video domain

Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models

ResearchDGX agent

arXiv:2607.27304v1 Announce Type: cross Abstract: Medical vision-language models (VLMs) generate chain-of-thought (CoT) reasoning before answering clinical questions, but whether this reasoning causal

PrintAnything: Learning an Intermediate Representation for 3D printing G-code Generation

ResearchDGX agent

arXiv:2607.27729v1 Announce Type: new Abstract: Point clouds are one of the most fundamental and widely used 3D representations, serving as the most basic geometric representation of 3D shapes. Nevert

Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation

Model ReleasesDGX agent

arXiv:2607.27764v1 Announce Type: new Abstract: Publishing private face recognition~(FR) training datasets is privacy-sensitive because faces expose identity information. Private FR training dataset p

ProgFormer: Hierarchical Voxel Diffusion Transformer for Longitudinal Brain MRI Prediction

ResearchDGX agent

arXiv:2607.27537v1 Announce Type: new Abstract: Predicting future structural MRI of a brain is challenging because longitudinal changes are often subtle and confined to specific anatomical regions, wh

← Previous
1…2324252627…209
Next →