AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Safety

EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations

DGX agent

arXiv:2604.18320v1 Announce Type: new Abstract: Self-evolution of multimodal large language models (MLLMs) remains a critical challenge: pseudo-label-based methods suffer from progressive quality degr

safetyarxiv-cs-cv
21 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling

DGX agent

arXiv:2604.17087v1 Announce Type: new Abstract: Recent Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language understanding tasks, yet their inference efficie

researcharxiv-cs-cv
21 Apr 2026
Tutorials

Expert-Annotated Embryo Image Dataset with Natural Language Descriptions for Evidence-Based Patient Communication in IVF

DGX agent

arXiv:2604.16528v1 Announce Type: new Abstract: Embryo selection is one of multiple crucial steps in in-vitro fertilization, commonly based on morphological assessment by clinical embryologists. Altho

tutorialsarxiv-cs-cv
21 Apr 2026
Applications

Explaining Uncertainty in Multiple Sclerosis Cortical Lesion Segmentation Beyond Prediction Errors

DGX agent

arXiv:2504.04814v3 Announce Type: replace-cross Abstract: Trustworthy artificial intelligence (AI) is essential in healthcare, particularly for high-stakes tasks like medical image segmentation. Expla

applicationsarxiv-cs-cv
21 Apr 2026
Model Releases

Exploring Boundary-Aware Spatial-Frequency Fusion for Camouflaged Object Detection

DGX agent

arXiv:2604.17879v1 Announce Type: new Abstract: Camouflaged Object Detection is challenging due to the high degree of similarity between camouflaged objects and their surrounding backgrounds. Current

model-releasesarxiv-cs-cv
21 Apr 2026
Research

Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation

DGX agent

arXiv:2502.13637v2 Announce Type: replace Abstract: Human affordance learning investigates contextually relevant novel pose prediction such that the estimated pose represents a valid human action with

researcharxiv-cs-cv
21 Apr 2026
Research

Expressive yet Efficient Feature Expansion with Adaptive Cross-Hadamard Products

DGX agent

arXiv:2505.22226v2 Announce Type: replace Abstract: Recent theoretical advances reveal that the Hadamard product induces nonlinear representations and implicit high-dimensional mappings for the field

researcharxiv-cs-cv
21 Apr 2026
Research

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation

DGX agent

arXiv:2604.18168v1 Announce Type: new Abstract: Few-step generation has been a long-standing goal, with recent one-step generation methods exemplified by MeanFlow achieving remarkable results. Existin

researcharxiv-cs-cv
21 Apr 2026
Safety

FairNVT: Improving Fairness via Noise Injection in Vision Transformers

DGX agent

arXiv:2604.16780v1 Announce Type: new Abstract: This paper presents FairNVT, a lightweight debiasing framework for pretrained transformer-based encoders that improves both representation and predictio

safetyarxiv-cs-cv
21 Apr 2026
Research

Fast Online 3D Multi-Camera Multi-Object Tracking and Pose Estimation

DGX agent

arXiv:2604.16522v1 Announce Type: new Abstract: This paper proposes a fast and online method for jointly performing 3D multi-object tracking and pose estimation using multiple monocular cameras. Our a

researcharxiv-cs-cv
21 Apr 2026
Model Releases

FireScope: Wildfire Risk Prediction with a Chain-of-Thought Oracle

DGX agent

arXiv:2511.17171v4 Announce Type: replace Abstract: Predicting wildfire risk is a reasoning-intensive spatial problem that requires the integration of visual, climatic, and geographic factors to infer

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

FlashFPS: Efficient Farthest Point Sampling for Large-Scale Point Clouds via Pruning and Caching

DGX agent

arXiv:2604.17720v1 Announce Type: cross Abstract: Point-based Neural Networks (PNNs) have become a key approach for point cloud processing. However, a core operation in these models, Farthest Point Sa

model-releasesarxiv-cs-cv
21 Apr 2026
Tutorials

FlowC2S: Flowing from Current to Succeeding Frames for Fast and Memory-Efficient Video Continuation

DGX agent

arXiv:2604.17625v1 Announce Type: new Abstract: This paper introduces a novel methodology for generating fast and memory-efficient video continuations. Our method, dubbed FlowC2S, fine-tunes a pre-tra

tutorialsarxiv-cs-cv
21 Apr 2026
Research

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts

DGX agent

arXiv:2603.19857v2 Announce Type: replace-cross Abstract: Recent Video-to-Audio (V2A) methods have achieved remarkable progress, enabling the synthesis of realistic, high-quality audio. However, they

researcharxiv-cs-cv
21 Apr 2026
Research

Fractal Characterization of Low-Correlation Signals in AI-Generated Image Detection

DGX agent

arXiv:2604.17268v1 Announce Type: new Abstract: AI-generated imagery has reached near-photorealistic fidelity, yet this technology poses significant threats to information security and societal trust.

researcharxiv-cs-cv
21 Apr 2026
Research

FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation

DGX agent

arXiv:2603.09721v2 Announce Type: replace Abstract: High-fidelity video generation remains challenging for diffusion models due to the difficulty of modeling complex spatio-temporal dynamics efficient

researcharxiv-cs-cv
21 Apr 2026
Research

FrameVGGT: Geometry-Aligned Frame-Level Memory for Bounded Streaming VGGT

DGX agent

arXiv:2603.07690v2 Announce Type: replace Abstract: Streaming Visual Geometry Transformers such as StreamVGGT enable strong online 3D perception, but their KV-cache grows unbounded over long streams,

researcharxiv-cs-cv
21 Apr 2026
Tutorials

Frequency-Decomposed INR for NIR-Assisted Low-Light RGB Image Denoising

DGX agent

arXiv:2604.16800v1 Announce Type: new Abstract: Addressing the issues of severe noise and high frequency structural degradation in visible images under low-light conditions, this paper proposes a Near

tutorialsarxiv-cs-cv
21 Apr 2026
Research

Frequency-guided Multi-level Reasoning for Scene Graph Generation in Video

DGX agent

arXiv:2604.17298v1 Announce Type: new Abstract: Video Scene Graph Generation aims to obtain structured semantic representations of objects and their relationships in videos for high-level understandin

researcharxiv-cs-cv
21 Apr 2026
Local Ai

Fringe Projection Based Vision Pipeline for Autonomous Hard Drive Disassembly

DGX agent

arXiv:2604.17231v1 Announce Type: new Abstract: Unrecovered e-waste represents a significant economic loss. Hard disk drives (HDDs) comprise a valuable e-waste stream necessitating robotic disassembly

local-aiarxiv-cs-cv
21 Apr 2026
Tutorials

From Adaptation to Generalization: Adaptive Visual Prompting for Medical Image Segmentation

DGX agent

arXiv:2604.17455v1 Announce Type: new Abstract: Visual prompting has emerged as a powerful method for adapting pre-trained models to new domains without updating model parameters. However, existing pr

tutorialsarxiv-cs-cv
21 Apr 2026
Agents

From Clinical Intent to Clinical Model: An Autonomous Coding-Agent Framework for Clinician-driven AI Development

DGX agent

arXiv:2604.17110v1 Announce Type: new Abstract: Clinical AI development has traditionally followed a collaborative paradigm that depends on close interaction between clinicians and specialized AI team

agentsarxiv-cs-cv
21 Apr 2026
Model Releases

From Handwriting to Structured Data: Benchmarking AI Digitisation of Handwritten Forms

DGX agent

arXiv:2604.16504v1 Announce Type: new Abstract: Manual digitisation of structured handwritten documents is slow and costly. We benchmark 17 leading frontier multi-modal large language models and open-

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference Acceleration

DGX agent

arXiv:2604.16462v1 Announce Type: new Abstract: High-resolution Multimodal Large Language Models (MLLMs) face prohibitive computational costs during inference due to the explosion of visual tokens. Ex

model-releasesarxiv-cs-cv
21 Apr 2026
Applications

Frozen Vision Transformers for Dense Prediction on Small Datasets: A Case Study in Arrow Localization

DGX agent

arXiv:2604.16758v1 Announce Type: new Abstract: We present a system for automated detection, localization, and scoring of arrow punctures on 40,cm indoor archery target faces, trained on only 48 annot

applicationsarxiv-cs-cv
21 Apr 2026
Research

GeGS-PCR: Effective and Robust 3D Point Cloud Registration with Two-Stage Color-Enhanced Geometric-3DGS Fusion

DGX agent

arXiv:2604.17721v1 Announce Type: new Abstract: We address the challenge of point cloud registration using color information, where traditional methods relying solely on geometric features often strug

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Generalizable Face Forgery Detection via Separable Prompt Learning

DGX agent

arXiv:2604.17307v1 Announce Type: new Abstract: Detecting face forgeries using CLIP has recently emerged as a promising and increasingly popular research direction. Owing to its rich visual knowledge

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

Generative Semantic Communication via Alternating Dual-Domain Posterior Sampling

DGX agent

arXiv:2604.16796v1 Announce Type: new Abstract: Generative semantic communication (SemCom) harnesses pretrained generative priors to improve the perceptual quality of wireless image transmission. Exis

safetyarxiv-cs-cv
21 Apr 2026
Local Ai

Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering

DGX agent

arXiv:2604.16487v1 Announce Type: new Abstract: CLIP retrieval is typically framed as a pointwise similarity problem in a shared embedding space. While CLIP achieves strong global cross-modal alignmen

local-aiarxiv-cs-cv
21 Apr 2026
Research

Geometry-Guided 3D Visual Token Pruning for Video-Language Models

DGX agent

arXiv:2604.18260v1 Announce Type: new Abstract: Multimodal large language models have demonstrated remarkable capabilities in 2D vision, motivating their extension to 3D scene understanding. Recent st

researcharxiv-cs-cv
21 Apr 2026
Model Releases

GR4CIL: Gap-compensated Routing for CLIP-based Class Incremental Learning

DGX agent

arXiv:2604.17822v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) aims to continuously acquire new categories while preserving previously learned knowledge. Recently, Contrastive Langua

model-releasesarxiv-cs-cv
21 Apr 2026
Tutorials

Graph neural network for colliding particles with an application to sea ice floe modeling

DGX agent

arXiv:2602.16213v2 Announce Type: replace-cross Abstract: This paper introduces a novel approach to sea ice modeling using Graph Neural Networks (GNNs), utilizing the natural graph structure of sea ic

tutorialsarxiv-cs-cv
21 Apr 2026
Safety

GS-STVSR: Ultra-Efficient Continuous Spatio-Temporal Video Super-Resolution via 2D Gaussian Splatting

DGX agent

arXiv:2604.18047v1 Announce Type: new Abstract: Continuous Spatio-Temporal Video Super-Resolution (C-STVSR) aims to simultaneously enhance the spatial resolution and frame rate of videos by arbitrary

safetyarxiv-cs-cv
21 Apr 2026
Research

HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image Retrieval

DGX agent

arXiv:2604.18037v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) is a flexible image retrieval paradigm that enables users to accurately locate the target image through a multimodal quer

researcharxiv-cs-cv
21 Apr 2026
Local Ai

Hierarchical Vision Transformer Enhanced by Graph Convolutional Network for Image Classification

DGX agent

arXiv:2604.16823v1 Announce Type: new Abstract: Vision Transformer (ViT) has brought new breakthroughs to the field of image classification by introducing the self-attention mechanism and Graph Convol

local-aiarxiv-cs-cv
21 Apr 2026
Safety

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

DGX agent

arXiv:2507.05920v2 Announce Type: replace Abstract: State-of-the-art large multi-modal models (LMMs) face challenges when processing high-resolution images, as these inputs are converted into enormous

safetyarxiv-cs-cv
21 Apr 2026
Research

HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models

DGX agent

arXiv:2508.00553v3 Announce Type: replace Abstract: Vision-Language Models (VLMs) encode images and videos into abundant tokens, which contain substantial redundancy and computation cost. While visual

researcharxiv-cs-cv
21 Apr 2026
Model Releases

HQA-VLAttack: Towards High Quality Adversarial Attack on Vision-Language Pre-Trained Models

DGX agent

arXiv:2604.16499v1 Announce Type: new Abstract: Black-box adversarial attack on vision-language pre-trained models is a practical and challenging task, as text and image perturbations need to be consi

model-releasesarxiv-cs-cv
21 Apr 2026
Tutorials

HSG: Hyperbolic Scene Graph

DGX agent

arXiv:2604.17454v1 Announce Type: new Abstract: Scene graph representations enable structured visual understanding by modeling objects and their relationships, and have been widely used for multiview

tutorialsarxiv-cs-cv
21 Apr 2026
Agents

Human Cognition in Machines: A Unified Perspective of World Models

DGX agent

arXiv:2604.16592v1 Announce Type: cross Abstract: This comprehensive report distinguishes prior works by the cognitive functions they innovate. Many works claim an almost 'human-like' cognitive capabi

agentsarxiv-cs-cv
21 Apr 2026
Safety

Hybrid Multi-Dimensional MRI Prostate Cancer Detection via Hadamard Network-Based Bias Correction and Residual Networks

DGX agent

arXiv:2604.17107v1 Announce Type: new Abstract: Magnetic Resonance Imaging (MRI) is vital for prostate cancer (PCa) diagnosis. While advanced techniques such as Hybrid Multi-dimensional MRI (HM-MRI) h

safetyarxiv-cs-cv
21 Apr 2026
Applications

Hybrid Quantum Neural Networks for Enhanced Breast Cancer Thermographic Classification: A Novel Quantum-Classical Integration Approach

DGX agent

arXiv:2604.16953v1 Announce Type: cross Abstract: Breast cancer diagnosis through thermographic image analysis remains a critical challenge in medical AI, with classical deep learning approaches facin

applicationsarxiv-cs-cv
21 Apr 2026
Model Releases

Hybrid-Vector Retrieval for Visually Rich Documents: Combining Single-Vector Efficiency and Multi-Vector Accuracy

DGX agent

arXiv:2510.22215v2 Announce Type: replace-cross Abstract: Retrieval over visually rich documents is essential for tasks such as legal discovery, scientific search, and enterprise knowledge management.

model-releasesarxiv-cs-cv
21 Apr 2026
Research

HyKey: Hyperspectral Keypoint Detection and Matching in Minimally Invasive Surgery

DGX agent

arXiv:2604.17446v1 Announce Type: new Abstract: Purpose: 3D reconstruction in minimally invasive surgery (MIS) enables enhanced surgical guidance through improved visualisation, tool tracking, and aug

researcharxiv-cs-cv
21 Apr 2026
Safety

Hyperbolic Enhanced Representation Learning for Incomplete Multi-view Clustering

DGX agent

arXiv:2604.16959v1 Announce Type: cross Abstract: Incomplete Multi-View Clustering (IMVC) faces the challenge of learning discriminative representations from fragmentary observations while maintaining

safetyarxiv-cs-cv
21 Apr 2026
Research

Hyperspectral Unmixing Hierarchies

DGX agent

arXiv:2604.16969v1 Announce Type: new Abstract: Unmixing reveals the spatial distribution and spectral details of different constituents, called endmembers, in a hyperspectral image. Because unmixing

researcharxiv-cs-cv
21 Apr 2026
Model Releases

ICAT: Incident-Case-Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models

DGX agent

arXiv:2604.16405v1 Announce Type: cross Abstract: Video-generative world models are increasingly used as neural simulators for embodied planning and policy learning, yet their ability to predict physi

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

Identifying Ethical Biases in Action Recognition Models

DGX agent

arXiv:2604.17971v1 Announce Type: new Abstract: Human Action Recognition (HAR) models are increasingly deployed in high-stakes environments, yet their fairness across different human appearances has n

safetyarxiv-cs-cv
21 Apr 2026
← Previous
1…227228229230231…261
Next →