AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Model Releases

Towards Generalizable Robotic Manipulation in Dynamic Environments

DGX agent

arXiv:2603.15620v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models excel in static manipulation but struggle in dynamic environments with moving targets. This performance gap prim

model-releasesarxiv-cs-cv
16 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Agents

Towards Multi-Object-Tracking with Radar on a Fast Moving Vehicle: On the Potential of Processing Radar in the Frequency Domain

DGX agent

arXiv:2604.14013v1 Announce Type: cross Abstract: We promote in this paper the processing of radar data in the frequency domain to achieve higher robustness against noise and structural errors, especi

agentsarxiv-cs-cv
16 Apr 2026
Research

Towards Patient-Specific Deformable Registration in Laparoscopic Surgery

DGX agent

arXiv:2604.13186v1 Announce Type: new Abstract: Unsafe surgical care is a critical health concern, often linked to limitations in surgeon experience, skills, and situational awareness. Integrating pat

researcharxiv-cs-cv
16 Apr 2026
Model Releases

Towards Successful Implementation of Automated Raveling Detection: Effects of Training Data Size, Illumination Difference, and Spatial Shift

DGX agent

arXiv:2604.13322v1 Announce Type: new Abstract: Raveling, the loss of aggregates, is a major form of asphalt pavement surface distress, especially on highways. While research has shown that machine le

model-releasesarxiv-cs-cv
16 Apr 2026
Research

Towards Unconstrained Human-Object Interaction

DGX agent

arXiv:2604.14069v1 Announce Type: new Abstract: Human-Object Interaction (HOI) detection is a longstanding computer vision problem concerned with predicting the interaction between humans and objects.

researcharxiv-cs-cv
16 Apr 2026
Research

Training-Free Semantic Multi-Object Tracking with Vision-Language Models

DGX agent

arXiv:2604.14074v1 Announce Type: new Abstract: Semantic Multi-Object Tracking (SMOT) extends multi-object tracking with semantic outputs such as video summaries, instance-level captions, and interact

researcharxiv-cs-cv
16 Apr 2026
Research

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing

DGX agent

arXiv:2604.13565v1 Announce Type: new Abstract: Ultra-high-resolution (UHR) remote sensing imagery couples kilometer-scale context with query-critical evidence that may occupy only a few pixels. Such

researcharxiv-cs-cv
16 Apr 2026
Safety

UNBOX: Unveiling Black-box visual models with Natural-language

DGX agent

arXiv:2603.08639v2 Announce Type: replace Abstract: Ensuring trustworthiness in open-world visual recognition requires models that are interpretable, fair, and robust to distribution shifts. Yet moder

safetyarxiv-cs-cv
16 Apr 2026
Model Releases

UniBlendNet: Unified Global, Multi-Scale, and Region-Adaptive Modeling for Ambient Lighting Normalization

DGX agent

arXiv:2604.13383v1 Announce Type: new Abstract: Ambient Lighting Normalization (ALN) aims to restore images degraded by complex, spatially varying illumination conditions. Existing methods, such as IF

model-releasesarxiv-cs-cv
16 Apr 2026
Model Releases

UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes

DGX agent

arXiv:2511.23332v2 Announce Type: replace Abstract: Instruction-driven segmentation in remote sensing generates masks from guidance, offering great potential for accessible and generalizable applicati

model-releasesarxiv-cs-cv
16 Apr 2026
Model Releases

VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentation

DGX agent

arXiv:2604.13596v1 Announce Type: new Abstract: Instance-level object segmentation across disparate egocentric and exocentric views is a fundamental challenge in visual understanding, critical for app

model-releasesarxiv-cs-cv
16 Apr 2026
Research

VibeFlow: Versatile Video Chroma-Lux Editing through Self-Supervised Learning

DGX agent

arXiv:2604.13425v1 Announce Type: new Abstract: Video chroma-lux editing, which aims to modify illumination and color while preserving structural and temporal fidelity, remains a significant challenge

researcharxiv-cs-cv
16 Apr 2026
Model Releases

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body

DGX agent

arXiv:2512.14234v2 Announce Type: replace Abstract: Human communication is inherently multimodal and social: words, prosody, and body language jointly carry intent. Yet most prior systems model human

model-releasesarxiv-cs-cv
16 Apr 2026
Safety

Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images

DGX agent

arXiv:2603.08486v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) face safety misalignment, where visual inputs enable harmful outputs. To address this, existing methods req

safetyarxiv-cs-cv
16 Apr 2026
Research

Visual Sparse Steering (VS2): Unsupervised Adaptation for Image Classification using Sparsity-Guided Steering Vectors

DGX agent

arXiv:2506.01247v2 Announce Type: replace Abstract: Steering vision foundation models at test time, without updating foundation-model weights or using labeled target data, is a desirable yet challengi

researcharxiv-cs-cv
16 Apr 2026
Safety

VRAG-DFD: Verifiable Retrieval-Augmentation for MLLM-based Deepfake Detection

DGX agent

arXiv:2604.13660v1 Announce Type: new Abstract: In Deepfake Detection (DFD) tasks, researchers proposed two types of MLLM-based methods: complementary combination with small DFD detectors, or static f

safetyarxiv-cs-cv
16 Apr 2026
Safety

What Are We Really Measuring? Rethinking Dataset Bias in Web-Scale Natural Image Collections via Unsupervised Semantic Clustering

DGX agent

arXiv:2604.13610v1 Announce Type: new Abstract: In computer vision, a prevailing method for quantifying dataset bias is to train a model to distinguish between datasets. High classification accuracy i

safetyarxiv-cs-cv
16 Apr 2026
Safety

Why MLLMs Struggle to Determine Object Orientations

DGX agent

arXiv:2604.13321v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) struggle with tasks that require reasoning about 2D object orientation in images, as documented in prior work.

safetyarxiv-cs-cv
16 Apr 2026
Safety

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks

DGX agent

arXiv:2604.13403v1 Announce Type: new Abstract: In-context learning (ICL) enables models to adapt to new tasks via inference-time demonstrations. Despite its success in large language models, the exte

safetyarxiv-cs-cv
16 Apr 2026
Tutorials

X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations

DGX agent

arXiv:2511.04671v2 Announce Type: replace-cross Abstract: Human videos are a scalable source of training data for robot learning. However, humans and robots significantly differ in embodiment, making

tutorialsarxiv-cs-cv
16 Apr 2026
Applications

ZoomSpec: A Physics-Guided Coarse-to-Fine Framework for Wideband Spectrum Sensing

DGX agent

arXiv:2604.13568v1 Announce Type: new Abstract: Wideband spectrum sensing for low-altitude monitoring is critical yet challenging due to heterogeneous protocols,large bandwidths, and non-stationary SN

applicationsarxiv-cs-cv
16 Apr 2026
Safety

A Dataset and Evaluation for Complex 4D Markerless Human Motion Capture

DGX agent

arXiv:2604.12765v1 Announce Type: new Abstract: Marker-based motion capture (MoCap) systems have long been the gold standard for accurate 4D human modeling, yet their reliance on specialized hardware

safetyarxiv-cs-cv
15 Apr 2026
Local Ai

A Hybrid Architecture for Benign-Malignant Classification of Mammography ROIs

DGX agent

arXiv:2604.12437v1 Announce Type: new Abstract: Accurate characterization of suspicious breast lesions in mammography is important for early diagnosis and treatment planning. While Convolutional Neura

local-aiarxiv-cs-cv
15 Apr 2026
Agents

A Multi-Agent Feedback System for Detecting and Describing News Events in Satellite Imagery

DGX agent

arXiv:2604.12772v1 Announce Type: new Abstract: Changes in satellite imagery often occur over multiple time steps. Despite the emergence of bi-temporal change captioning datasets, there is a lack of m

agentsarxiv-cs-cv
15 Apr 2026
Model Releases

A Sanity Check on Composed Image Retrieval

DGX agent

arXiv:2604.12904v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) aims to retrieve a target image based on a query composed of a reference image, and a relative caption that specifies the

model-releasesarxiv-cs-cv
15 Apr 2026
Research

A Workflow to Efficiently Generate Dense Tissue Ground Truth Masks for Digital Breast Tomosynthesis

DGX agent

arXiv:2604.11927v1 Announce Type: new Abstract: Digital breast tomosynthesis (DBT) is now the standard of care for breast cancer screening in the USA. Accurate segmentation of fibroglandular tissue in

researcharxiv-cs-cv
15 Apr 2026
Research

AbdomenGen: Sequential Volume-Conditioned Diffusion Framework for Abdominal Anatomy Generation

DGX agent

arXiv:2604.12969v1 Announce Type: new Abstract: Computational phantoms are widely used in medical imaging research, yet current systems to generate controlled, clinically meaningful anatomical variati

researcharxiv-cs-cv
15 Apr 2026
Model Releases

Adaptive Data Dropout: Towards Self-Regulated Learning in Deep Neural Networks

DGX agent

arXiv:2604.12945v1 Announce Type: cross Abstract: Deep neural networks are typically trained by uniformly sampling large datasets across epochs, despite evidence that not all samples contribute equall

model-releasesarxiv-cs-cv
15 Apr 2026
Model Releases

AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition

DGX agent

arXiv:2604.12735v1 Announce Type: new Abstract: LLM-based multimodal emotion recognition relies on static parametric memory and often hallucinates when interpreting nuanced affective states. In this p

model-releasesarxiv-cs-cv
15 Apr 2026
Agents

Agentic Discovery with Active Hypothesis Exploration for Visual Recognition

DGX agent

arXiv:2604.12999v1 Announce Type: new Abstract: We introduce HypoExplore, an agentic framework that formulates neural architecture discovery for visual recognition as a hypothesis-driven scientific in

agentsarxiv-cs-cv
15 Apr 2026
Research

AGMA: Adaptive Gaussian Mixture Anchors for Prior-Guided Multimodal Human Trajectory Forecasting

DGX agent

arXiv:2602.04204v2 Announce Type: replace Abstract: Human trajectory forecasting requires capturing the multimodal nature of pedestrian behavior. However, existing approaches suffer from prior misalig

researcharxiv-cs-cv
15 Apr 2026
Applications

All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding

DGX agent

arXiv:2604.12335v1 Announce Type: new Abstract: Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object coun

applicationsarxiv-cs-cv
15 Apr 2026
Applications

ArtifactWorld: Scaling 3D Gaussian Splatting Artifact Restoration via Video Generation Models

DGX agent

arXiv:2604.12251v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) delivers high-fidelity real-time rendering but suffers from geometric and photometric degradations under sparse-view constr

applicationsarxiv-cs-cv
15 Apr 2026
Model Releases

ASTRA: Let Arbitrary Subjects Transform in Video Editing

DGX agent

arXiv:2510.01186v2 Announce Type: replace Abstract: While existing video editing methods excel with single subjects, they struggle in dense, multi-subject scenes, frequently suffering from attention d

model-releasesarxiv-cs-cv
15 Apr 2026
Applications

BarbieGait: An Identity-Consistent Synthetic Human Dataset with Versatile Cloth-Changing for Gait Recognition

DGX agent

arXiv:2604.12221v1 Announce Type: new Abstract: Gait recognition, as a reliable biometric technology, has seen rapid development in recent years while it faces significant challenges caused by diverse

applicationsarxiv-cs-cv
15 Apr 2026
Model Releases

Beyond Perception Errors: Semantic Fixation in Large Vision-Language Models

DGX agent

arXiv:2604.12119v1 Announce Type: new Abstract: Large vision-language models (VLMs) often rely on familiar semantic priors, but existing evaluations do not cleanly separate perception failures from ru

model-releasesarxiv-cs-cv
15 Apr 2026
Research

Boosting Robust AIGI Detection with LoRA-based Pairwise Training

DGX agent

arXiv:2604.12307v1 Announce Type: new Abstract: The proliferation of highly realistic AI-Generated Image (AIGI) has necessitated the development of practical detection methods. While current AIGI dete

researcharxiv-cs-cv
15 Apr 2026
Research

Boosting Visual Instruction Tuning with Self-Supervised Guidance

DGX agent

arXiv:2604.12966v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform well on many vision-language tasks but often struggle with vision-centric problems that require fine-gr

researcharxiv-cs-cv
15 Apr 2026
Safety

Brain-DiT: A Universal Multi-state fMRI Foundation Model with Metadata-Conditioned Pretraining

DGX agent

arXiv:2604.12683v1 Announce Type: new Abstract: Current fMRI foundation models primarily rely on a limited range of brain states and mismatched pretraining tasks, restricting their ability to learn ge

safetyarxiv-cs-cv
15 Apr 2026
Safety

Bridging the Micro--Macro Gap: Frequency-Aware Semantic Alignment for Image Manipulation Localization

DGX agent

arXiv:2604.12341v1 Announce Type: new Abstract: As generative image editing advances, image manipulation localization (IML) must handle both traditional manipulations with conspicuous forensic artifac

safetyarxiv-cs-cv
15 Apr 2026
Research

Causal Fingerprints of AI Generative Models

DGX agent

arXiv:2509.15406v2 Announce Type: replace Abstract: AI generative models leave implicit traces in their generated images, which are commonly referred to as model fingerprints and are exploited for sou

researcharxiv-cs-cv
15 Apr 2026
Research

CBAM-Enhanced DenseNet121 for Multi-Class Chest X-Ray Classification with Grad-CAM Explainability

DGX agent

arXiv:2604.12305v1 Announce Type: cross Abstract: Pneumonia remains a leading cause of childhood mortality worldwide, with a heavy burden in low-resource settings such as Bangladesh where radiologist

researcharxiv-cs-cv
15 Apr 2026
Research

Cell Instance Segmentation via Multi-Task Image-to-Image Schrodinger Bridge

DGX agent

arXiv:2604.12318v1 Announce Type: new Abstract: Existing cell instance segmentation pipelines typically combine deterministic predictions with post-processing, which imposes limited explicit constrain

researcharxiv-cs-cv
15 Apr 2026
Safety

Challenging Vision-Language Models with Physically Deployable Multimodal Semantic Lighting Attacks

DGX agent

arXiv:2604.12833v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown remarkable performance, yet their security remains insufficiently understood. Existing adversarial studies focu

safetyarxiv-cs-cv
15 Apr 2026
Model Releases

CoD-Lite: Real-Time Diffusion-Based Generative Image Compression

DGX agent

arXiv:2604.12525v1 Announce Type: new Abstract: Recent advanced diffusion methods typically derive strong generative priors by scaling diffusion transformers. However, scaling fails to generalize when

model-releasesarxiv-cs-cv
15 Apr 2026
Research

CoLA: A Choice Leakage Attack Framework to Expose Privacy Risks in Subset Training

DGX agent

arXiv:2604.12342v1 Announce Type: cross Abstract: Training models on a carefully chosen portion of data rather than the full dataset is now a standard preprocess for modern ML. From vision coreset sel

researcharxiv-cs-cv
15 Apr 2026
Safety

Combating Pattern and Content Bias: Adversarial Feature Learning for Generalized AI-Generated Image Detection

DGX agent

arXiv:2604.12353v1 Announce Type: new Abstract: In recent years, the rapid development of generative artificial intelligence technology has significantly lowered the barrier to creating high-quality f

safetyarxiv-cs-cv
15 Apr 2026
Research

Conflated Inverse Modeling to Generate Diverse and Temperature-Change Inducing Urban Vegetation Patterns

DGX agent

arXiv:2604.13028v1 Announce Type: new Abstract: Urban areas are increasingly vulnerable to thermal extremes driven by rapid urbanization and climate change. Traditionally, thermal extremes have been m

researcharxiv-cs-cv
15 Apr 2026
← Previous
1…241242243244245…261
Next →