AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
16 Apr 2026

Training-Free Semantic Multi-Object Tracking with Vision-Language Models

ResearchDGX agent

arXiv:2604.14074v1 Announce Type: new Abstract: Semantic Multi-Object Tracking (SMOT) extends multi-object tracking with semantic outputs such as video summaries, instance-level captions, and interact

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing

ResearchDGX agent

arXiv:2604.13565v1 Announce Type: new Abstract: Ultra-high-resolution (UHR) remote sensing imagery couples kilometer-scale context with query-critical evidence that may occupy only a few pixels. Such

UNBOX: Unveiling Black-box visual models with Natural-language

SafetyDGX agent

arXiv:2603.08639v2 Announce Type: replace Abstract: Ensuring trustworthiness in open-world visual recognition requires models that are interpretable, fair, and robust to distribution shifts. Yet moder


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

UniBlendNet: Unified Global, Multi-Scale, and Region-Adaptive Modeling for Ambient Lighting Normalization

Model ReleasesDGX agent

arXiv:2604.13383v1 Announce Type: new Abstract: Ambient Lighting Normalization (ALN) aims to restore images degraded by complex, spatially varying illumination conditions. Existing methods, such as IF

UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes

Model ReleasesDGX agent

arXiv:2511.23332v2 Announce Type: replace Abstract: Instruction-driven segmentation in remote sensing generates masks from guidance, offering great potential for accessible and generalizable applicati

VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentation

Model ReleasesDGX agent

arXiv:2604.13596v1 Announce Type: new Abstract: Instance-level object segmentation across disparate egocentric and exocentric views is a fundamental challenge in visual understanding, critical for app

VibeFlow: Versatile Video Chroma-Lux Editing through Self-Supervised Learning

ResearchDGX agent

arXiv:2604.13425v1 Announce Type: new Abstract: Video chroma-lux editing, which aims to modify illumination and color while preserving structural and temporal fidelity, remains a significant challenge

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body

Model ReleasesDGX agent

arXiv:2512.14234v2 Announce Type: replace Abstract: Human communication is inherently multimodal and social: words, prosody, and body language jointly carry intent. Yet most prior systems model human

Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images

SafetyDGX agent

arXiv:2603.08486v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) face safety misalignment, where visual inputs enable harmful outputs. To address this, existing methods req

Visual Sparse Steering (VS2): Unsupervised Adaptation for Image Classification using Sparsity-Guided Steering Vectors

ResearchDGX agent

arXiv:2506.01247v2 Announce Type: replace Abstract: Steering vision foundation models at test time, without updating foundation-model weights or using labeled target data, is a desirable yet challengi

VRAG-DFD: Verifiable Retrieval-Augmentation for MLLM-based Deepfake Detection

SafetyDGX agent

arXiv:2604.13660v1 Announce Type: new Abstract: In Deepfake Detection (DFD) tasks, researchers proposed two types of MLLM-based methods: complementary combination with small DFD detectors, or static f

What Are We Really Measuring? Rethinking Dataset Bias in Web-Scale Natural Image Collections via Unsupervised Semantic Clustering

SafetyDGX agent

arXiv:2604.13610v1 Announce Type: new Abstract: In computer vision, a prevailing method for quantifying dataset bias is to train a model to distinguish between datasets. High classification accuracy i

Why MLLMs Struggle to Determine Object Orientations

SafetyDGX agent

arXiv:2604.13321v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) struggle with tasks that require reasoning about 2D object orientation in images, as documented in prior work.

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks

SafetyDGX agent

arXiv:2604.13403v1 Announce Type: new Abstract: In-context learning (ICL) enables models to adapt to new tasks via inference-time demonstrations. Despite its success in large language models, the exte

X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations

TutorialsDGX agent

arXiv:2511.04671v2 Announce Type: replace-cross Abstract: Human videos are a scalable source of training data for robot learning. However, humans and robots significantly differ in embodiment, making

ZoomSpec: A Physics-Guided Coarse-to-Fine Framework for Wideband Spectrum Sensing

ApplicationsDGX agent

arXiv:2604.13568v1 Announce Type: new Abstract: Wideband spectrum sensing for low-altitude monitoring is critical yet challenging due to heterogeneous protocols,large bandwidths, and non-stationary SN

15 Apr 2026

A Dataset and Evaluation for Complex 4D Markerless Human Motion Capture

SafetyDGX agent

arXiv:2604.12765v1 Announce Type: new Abstract: Marker-based motion capture (MoCap) systems have long been the gold standard for accurate 4D human modeling, yet their reliance on specialized hardware

A Hybrid Architecture for Benign-Malignant Classification of Mammography ROIs

Local AiDGX agent

arXiv:2604.12437v1 Announce Type: new Abstract: Accurate characterization of suspicious breast lesions in mammography is important for early diagnosis and treatment planning. While Convolutional Neura

A Multi-Agent Feedback System for Detecting and Describing News Events in Satellite Imagery

AgentsDGX agent

arXiv:2604.12772v1 Announce Type: new Abstract: Changes in satellite imagery often occur over multiple time steps. Despite the emergence of bi-temporal change captioning datasets, there is a lack of m

A Sanity Check on Composed Image Retrieval

Model ReleasesDGX agent

arXiv:2604.12904v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) aims to retrieve a target image based on a query composed of a reference image, and a relative caption that specifies the

A Workflow to Efficiently Generate Dense Tissue Ground Truth Masks for Digital Breast Tomosynthesis

ResearchDGX agent

arXiv:2604.11927v1 Announce Type: new Abstract: Digital breast tomosynthesis (DBT) is now the standard of care for breast cancer screening in the USA. Accurate segmentation of fibroglandular tissue in

AbdomenGen: Sequential Volume-Conditioned Diffusion Framework for Abdominal Anatomy Generation

ResearchDGX agent

arXiv:2604.12969v1 Announce Type: new Abstract: Computational phantoms are widely used in medical imaging research, yet current systems to generate controlled, clinically meaningful anatomical variati

Adaptive Data Dropout: Towards Self-Regulated Learning in Deep Neural Networks

Model ReleasesDGX agent

arXiv:2604.12945v1 Announce Type: cross Abstract: Deep neural networks are typically trained by uniformly sampling large datasets across epochs, despite evidence that not all samples contribute equall

AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition

Model ReleasesDGX agent

arXiv:2604.12735v1 Announce Type: new Abstract: LLM-based multimodal emotion recognition relies on static parametric memory and often hallucinates when interpreting nuanced affective states. In this p

Agentic Discovery with Active Hypothesis Exploration for Visual Recognition

AgentsDGX agent

arXiv:2604.12999v1 Announce Type: new Abstract: We introduce HypoExplore, an agentic framework that formulates neural architecture discovery for visual recognition as a hypothesis-driven scientific in

AGMA: Adaptive Gaussian Mixture Anchors for Prior-Guided Multimodal Human Trajectory Forecasting

ResearchDGX agent

arXiv:2602.04204v2 Announce Type: replace Abstract: Human trajectory forecasting requires capturing the multimodal nature of pedestrian behavior. However, existing approaches suffer from prior misalig

All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding

ApplicationsDGX agent

arXiv:2604.12335v1 Announce Type: new Abstract: Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object coun

ArtifactWorld: Scaling 3D Gaussian Splatting Artifact Restoration via Video Generation Models

ApplicationsDGX agent

arXiv:2604.12251v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) delivers high-fidelity real-time rendering but suffers from geometric and photometric degradations under sparse-view constr

ASTRA: Let Arbitrary Subjects Transform in Video Editing

Model ReleasesDGX agent

arXiv:2510.01186v2 Announce Type: replace Abstract: While existing video editing methods excel with single subjects, they struggle in dense, multi-subject scenes, frequently suffering from attention d

BarbieGait: An Identity-Consistent Synthetic Human Dataset with Versatile Cloth-Changing for Gait Recognition

ApplicationsDGX agent

arXiv:2604.12221v1 Announce Type: new Abstract: Gait recognition, as a reliable biometric technology, has seen rapid development in recent years while it faces significant challenges caused by diverse

Beyond Perception Errors: Semantic Fixation in Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.12119v1 Announce Type: new Abstract: Large vision-language models (VLMs) often rely on familiar semantic priors, but existing evaluations do not cleanly separate perception failures from ru

Boosting Robust AIGI Detection with LoRA-based Pairwise Training

ResearchDGX agent

arXiv:2604.12307v1 Announce Type: new Abstract: The proliferation of highly realistic AI-Generated Image (AIGI) has necessitated the development of practical detection methods. While current AIGI dete

Boosting Visual Instruction Tuning with Self-Supervised Guidance

ResearchDGX agent

arXiv:2604.12966v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform well on many vision-language tasks but often struggle with vision-centric problems that require fine-gr

Brain-DiT: A Universal Multi-state fMRI Foundation Model with Metadata-Conditioned Pretraining

SafetyDGX agent

arXiv:2604.12683v1 Announce Type: new Abstract: Current fMRI foundation models primarily rely on a limited range of brain states and mismatched pretraining tasks, restricting their ability to learn ge

Bridging the Micro--Macro Gap: Frequency-Aware Semantic Alignment for Image Manipulation Localization

SafetyDGX agent

arXiv:2604.12341v1 Announce Type: new Abstract: As generative image editing advances, image manipulation localization (IML) must handle both traditional manipulations with conspicuous forensic artifac

Causal Fingerprints of AI Generative Models

ResearchDGX agent

arXiv:2509.15406v2 Announce Type: replace Abstract: AI generative models leave implicit traces in their generated images, which are commonly referred to as model fingerprints and are exploited for sou

CBAM-Enhanced DenseNet121 for Multi-Class Chest X-Ray Classification with Grad-CAM Explainability

ResearchDGX agent

arXiv:2604.12305v1 Announce Type: cross Abstract: Pneumonia remains a leading cause of childhood mortality worldwide, with a heavy burden in low-resource settings such as Bangladesh where radiologist

Cell Instance Segmentation via Multi-Task Image-to-Image Schrodinger Bridge

ResearchDGX agent

arXiv:2604.12318v1 Announce Type: new Abstract: Existing cell instance segmentation pipelines typically combine deterministic predictions with post-processing, which imposes limited explicit constrain

Challenging Vision-Language Models with Physically Deployable Multimodal Semantic Lighting Attacks

SafetyDGX agent

arXiv:2604.12833v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown remarkable performance, yet their security remains insufficiently understood. Existing adversarial studies focu

CoD-Lite: Real-Time Diffusion-Based Generative Image Compression

Model ReleasesDGX agent

arXiv:2604.12525v1 Announce Type: new Abstract: Recent advanced diffusion methods typically derive strong generative priors by scaling diffusion transformers. However, scaling fails to generalize when

CoLA: A Choice Leakage Attack Framework to Expose Privacy Risks in Subset Training

ResearchDGX agent

arXiv:2604.12342v1 Announce Type: cross Abstract: Training models on a carefully chosen portion of data rather than the full dataset is now a standard preprocess for modern ML. From vision coreset sel

Combating Pattern and Content Bias: Adversarial Feature Learning for Generalized AI-Generated Image Detection

SafetyDGX agent

arXiv:2604.12353v1 Announce Type: new Abstract: In recent years, the rapid development of generative artificial intelligence technology has significantly lowered the barrier to creating high-quality f

Conflated Inverse Modeling to Generate Diverse and Temperature-Change Inducing Urban Vegetation Patterns

ResearchDGX agent

arXiv:2604.13028v1 Announce Type: new Abstract: Urban areas are increasingly vulnerable to thermal extremes driven by rapid urbanization and climate change. Traditionally, thermal extremes have been m

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing

SafetyDGX agent

arXiv:2604.12292v1 Announce Type: cross Abstract: Movie dubbing aims to synthesize speech that preserves the vocal identity of a reference audio while synchronizing with the lip movements in a target

CREG: Compass Relational Evidence Graph for Characterizing Directional Structure in VLM Spatial-Reasoning Attribution

Local AiDGX agent

arXiv:2603.20475v3 Announce Type: replace Abstract: Standard attribution heatmaps show where a vision-language model (VLM) focuses, but they do not reveal whether the recovered evidence is organized b

Cross-Attentive Multiview Fusion of Vision-Language Embeddings

ResearchDGX agent

arXiv:2604.12551v1 Announce Type: new Abstract: Vision-language models have been key to the development of open-vocabulary 2D semantic segmentation. Lifting these models from 2D images to 3D scenes, h

Cross-Modal Knowledge Distillation for PET-Free Amyloid-Beta Detection from MRI

SafetyDGX agent

arXiv:2604.12574v1 Announce Type: new Abstract: Detecting amyloid-eta (Aeta) positivity is crucial for early diagnosis of Alzheimer's disease but typically requires PET imaging, which is costly, invas

DAV-GSWT: Diffusion-Active-View Sampling for Data-Efficient Gaussian Splatting Wang Tiles

ResearchDGX agent

arXiv:2602.15355v3 Announce Type: replace Abstract: The emergence of 3D Gaussian Splatting has fundamentally redefined the capabilities of photorealistic neural rendering by enabling high-throughput s

DC-TTA: Divide-and-Conquer Framework for Test-Time Adaptation of Interactive Segmentation

ResearchDGX agent

arXiv:2506.23104v2 Announce Type: replace Abstract: Interactive segmentation (IS) allows users to iteratively refine object boundaries with minimal cues, such as positive and negative clicks. While th

Deep Learning using Rectified Linear Units (ReLU)

ResearchDGX agent

arXiv:1803.08375v3 Announce Type: replace-cross Abstract: The Rectified Linear Unit (ReLU) is a foundational activation function in artficial neural networks. Recent literature frequently misattribute

DeferredSeg: A Multi-Expert Deferral Framework for Trustworthy Medical Image Segmentation

ResearchDGX agent

arXiv:2604.12411v1 Announce Type: new Abstract: Segmentation models based on deep neural networks demonstrate strong generalization for medical image segmentation. However, they often exhibit overconf

Detecting Precise Hand Touch Moments in Egocentric Video

ResearchDGX agent

arXiv:2604.12343v1 Announce Type: new Abstract: We address the challenging task of detecting the precise moment when hands make contact with objects in egocentric videos. This frame-level detection is

DiffusionPrint: Learning Generative Fingerprints for Diffusion-Based Inpainting Localization

Local AiDGX agent

arXiv:2604.12443v1 Announce Type: new Abstract: Modern diffusion-based inpainting models pose significant challenges for image forgery localization (IFL), as their full regeneration pipelines reconstr

DINO-Explorer: Active Underwater Discovery via Ego-Motion Compensated Semantic Predictive Coding

AgentsDGX agent

arXiv:2604.12933v1 Announce Type: cross Abstract: Marine ecosystem degradation necessitates continuous, scientifically selective underwater monitoring. However, most autonomous underwater vehicles (AU

Direct Discrepancy Replay: Distribution-Discrepancy Condensation and Manifold-Consistent Replay for Continual Face Forgery Detection

TutorialsDGX agent

arXiv:2604.12941v1 Announce Type: new Abstract: Continual face forgery detection (CFFD) requires detectors to learn emerging forgery paradigms without forgetting previously seen manipulations. Existin

Does Visual Token Pruning Improve Calibration? An Empirical Study on Confidence in MLLMs

ResearchDGX agent

arXiv:2604.12035v1 Announce Type: new Abstract: Visual token pruning is a widely used strategy for efficient inference in multimodal large language models (MLLMs), but existing work mainly evaluates i

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs

Model ReleasesDGX agent

arXiv:2604.12896v1 Announce Type: new Abstract: Multimodal language models (MLLMs) are increasingly paired with vision tools (e.g., depth, flow, correspondence) to enhance visual reasoning. However, d

DPC-VQA: Decoupling Quality Perception and Residual Calibration for Video Quality Assessment

Model ReleasesDGX agent

arXiv:2604.12813v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) have shown promising performance on video quality assessment (VQA) tasks. However, adapting them to new

DreamStereo: Towards Real-Time Stereo Inpainting for HD Videos

HardwareDGX agent

arXiv:2604.12270v1 Announce Type: new Abstract: Stereo video inpainting, which aims to fill the occluded regions of warped videos with visually coherent content while maintaining temporal consistency,

Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off

Model ReleasesDGX agent

arXiv:2603.22607v2 Announce Type: replace Abstract: Recent advances in Virtual Try-On (VTON) and Virtual Try-Off (VTOFF) have greatly improved photo-realistic fashion synthesis and garment reconstruct

← Previous
1…191192193194195…207
Next →