AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Agents

PoseFM: Relative Camera Pose Estimation Through Flow Matching

DGX agent

arXiv:2604.22350v1 Announce Type: new Abstract: Monocular visual odometry (VO) is a fundamental computer vision problem with applications in autonomous navigation, augmented reality and more. While de

agentsarxiv-cs-cv
27 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Railway Artificial Intelligence Learning Benchmark (RAIL-BENCH): A Benchmark Suite for Perception in the Railway Domain

DGX agent

arXiv:2604.22507v1 Announce Type: new Abstract: Automated train operation on existing railway infrastructure requires robust camera-based perception, yet the railway domain lacks public benchmark suit

model-releasesarxiv-cs-cv
27 Apr 2026
Model Releases

Recent Advances in Multi-Agent Human Trajectory Prediction: A Comprehensive Review

DGX agent

arXiv:2506.14831v3 Announce Type: replace Abstract: With the emergence of powerful data-driven methods in human trajectory prediction (HTP), gaining a finer understanding of multi-agent interactions l

model-releasesarxiv-cs-cv
27 Apr 2026
Model Releases

Region Matters: Efficient and Reliable Region-Aware Visual Place Recognition

DGX agent

arXiv:2604.22390v1 Announce Type: new Abstract: Visual Place Recognition (VPR) determines a query image's geographic location by matching it against geotagged databases. However, existing methods stru

model-releasesarxiv-cs-cv
27 Apr 2026
Research

ReLIC-SGG: Relation Lattice Completion for Open-Vocabulary Scene Graph Generation

DGX agent

arXiv:2604.22546v1 Announce Type: new Abstract: Open-vocabulary scene graph generation (SGG) aims to describe visual scenes with flexible relation phrases beyond a fixed predicate set. Existing method

researcharxiv-cs-cv
27 Apr 2026
Research

Rethinking Token Pruning for Historical Screenshots in GUI Visual Agents: Semantic, Spatial, and Temporal Perspectives

DGX agent

arXiv:2603.26041v3 Announce Type: replace Abstract: In recent years, GUI visual agents built upon Multimodal Large Language Models (MLLMs) have demonstrated strong potential in navigation tasks. Howev

researcharxiv-cs-cv
27 Apr 2026
Research

Revisiting Geometric Obfuscation with Dual Convergent Lines for Privacy-Preserving Image Queries in Visual Localization

DGX agent

arXiv:2604.22310v1 Announce Type: new Abstract: Privacy-Preserving Image Queries (PPIQ) are an emerging mechanism for cloud-based visual localization, enabling pose estimation from obfuscated features

researcharxiv-cs-cv
27 Apr 2026
Applications

Robust Camera-to-Mocap Calibration and Verification for Large-Scale Multi-Camera Data Capture

DGX agent

arXiv:2604.22118v1 Announce Type: new Abstract: Optical motion capture (mocap) systems are widely used for ground-truth capture in AR/VR, SLAM and robotics datasets. These datasets require extrinsic c

applicationsarxiv-cs-cv
27 Apr 2026
Research

SAMIDARE: Advanced Tracking-by-Segmentation for Dense Scenarios

DGX agent

arXiv:2604.22162v1 Announce Type: new Abstract: Automated sports analysis demands robust multi-object tracking (MOT), yet segmentation-based methods often struggle with mask errors and ID switches in

researcharxiv-cs-cv
27 Apr 2026
Model Releases

Score-based Membership Inference on Diffusion Models

DGX agent

arXiv:2509.25003v2 Announce Type: replace-cross Abstract: Membership inference attacks (MIAs) against Diffusion Models (DMs) raise pressing privacy concerns by revealing whether a sample was part of t

model-releasesarxiv-cs-cv
27 Apr 2026
Applications

Segment Any-Quality Images with Generative Latent Space Enhancement

DGX agent

arXiv:2503.12507v3 Announce Type: replace Abstract: Despite their success, Segment Anything Models (SAMs) experience significant performance drops on severely degraded, low-quality images, limiting th

applicationsarxiv-cs-cv
27 Apr 2026
Research

Segmentation of Gray Matters and White Matters from Brain MRI data

DGX agent

arXiv:2603.29171v3 Announce Type: replace Abstract: Accurate segmentation of brain tissues such as gray matter and white matter from magnetic resonance imaging is essential for studying brain anatomy,

researcharxiv-cs-cv
27 Apr 2026
Model Releases

Selective Depthwise Separable Convolution for Lightweight Joint Source-Channel Coding in Wireless Image Transmission

DGX agent

arXiv:2604.22338v1 Announce Type: cross Abstract: Depthwise separable convolutional (DSConv) layers have been successfully applied to deep learning (DL)-based joint source-channel coding (JSCC) scheme

model-releasesarxiv-cs-cv
27 Apr 2026
Model Releases

Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging

DGX agent

arXiv:2510.05971v3 Announce Type: replace Abstract: The generalization of the Transformer architecture via MetaFormer has reshaped our understanding of its success in computer vision. By replacing sel

model-releasesarxiv-cs-cv
27 Apr 2026
Hardware

SIE3D: Single-Image Expressive 3D Avatar Generation via Semantic Embedding and Perceptual Expression Loss

DGX agent

arXiv:2509.24004v2 Announce Type: replace Abstract: Generating high-fidelity 3D head avatars from a single image is challenging, as current methods lack fine-grained, intuitive control over expression

hardwarearxiv-cs-cv
27 Apr 2026
Hardware

Soft Anisotropic Diagrams for Differentiable Image Representation

DGX agent

arXiv:2604.21984v1 Announce Type: new Abstract: We introduce Soft Anisotropic Diagrams (SAD), an explicit and differentiable image representation parameterized by a set of adaptive sites in the image

hardwarearxiv-cs-cv
27 Apr 2026
Model Releases

SpaMEM: Benchmarking Dynamic Spatial Reasoning via Perception-Memory Integration in Embodied Environments

DGX agent

arXiv:2604.22409v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have advanced static visual--spatial reasoning, yet they often fail to preserve long-horizon spatial coherence

model-releasesarxiv-cs-cv
27 Apr 2026
Research

SS3D: End2End Self-Supervised 3D from Web Videos

DGX agent

arXiv:2604.22686v1 Announce Type: new Abstract: We present SS3D, a web-scale SfM-based self-supervision pretraining pipeline for feed-forward 3D estimation from monocular video. Our model jointly pred

researcharxiv-cs-cv
27 Apr 2026
Tutorials

Structure-Guided Diffusion Model for EEG-Based Visual Cognition Reconstruction

DGX agent

arXiv:2604.22649v1 Announce Type: cross Abstract: Objective: Decoding visual information from electroencephalography (EEG) is an important problem in neuroscience and brain-computer interface (BCI) re

tutorialsarxiv-cs-cv
27 Apr 2026
Model Releases

Sum-of-Checks: Structured Reasoning for Surgical Safety with Large Vision-Language Models

DGX agent

arXiv:2604.22156v1 Announce Type: cross Abstract: Purpose: Accurate assessment of the Critical View of Safety (CVS) during laparoscopic cholecystectomy is essential to prevent bile duct injury, a comp

model-releasesarxiv-cs-cv
27 Apr 2026
Safety

Thermal background reduction for mid-infrared imaging by low-rank background and sparse point-source modelling

DGX agent

arXiv:2604.22351v1 Announce Type: cross Abstract: Mid-infrared astronomy from the ground faces critical challenges in accurately detecting and quantifying sources due to the dominant spatially and tim

safetyarxiv-cs-cv
27 Apr 2026
Model Releases

Towards Temporal Compositional Reasoning in Long-Form Sports Videos

DGX agent

arXiv:2604.22226v1 Announce Type: new Abstract: Sports videos are a challenging domain for multimodal understanding because they involve complex and dynamic human activities. Despite rapid progress in

model-releasesarxiv-cs-cv
27 Apr 2026
Safety

Transferable Physical-World Adversarial Patches Against Pedestrian Detection Models

DGX agent

arXiv:2604.22552v1 Announce Type: new Abstract: Physical adversarial patch attacks critically threaten pedestrian detection, causing surveillance and autonomous driving systems to miss pedestrians and

safetyarxiv-cs-cv
27 Apr 2026
Research

Uni-Encoder Meets Multi-Encoders: Representation Before Fusion for Brain Tumor Segmentation with Missing Modalities

DGX agent

arXiv:2604.22177v1 Announce Type: new Abstract: Multimodal MRI offers complementary information for brain tumor segmentation, but clinical scans often lack one or more modalities, which degrades segme

researcharxiv-cs-cv
27 Apr 2026
Safety

Unlocking Optical Prior: Spectrum-Guided Knowledge Transfer for SAR Generalized Category Discovery

DGX agent

arXiv:2604.22174v1 Announce Type: new Abstract: Generalized Category Discovery (GCD) holds significant promise for the label-scarce Synthetic Aperture Radar (SAR) domain, yet its efficacy is severely

safetyarxiv-cs-cv
27 Apr 2026
Tutorials

Useful nonrobust features are ubiquitous in biomedical images

DGX agent

arXiv:2604.22579v1 Announce Type: cross Abstract: We study whether deep networks for medical imaging learn useful nonrobust features - predictive input patterns that are not human interpretable and hi

tutorialsarxiv-cs-cv
27 Apr 2026
Research

V-MAGE: A Game Evaluation Framework for Assessing Vision-Centric Capabilities in Multimodal Large Language Models

DGX agent

arXiv:2504.06148v3 Announce Type: replace Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in visual-text processing. However, existi

researcharxiv-cs-cv
27 Apr 2026
Applications

Video Analysis and Generation via a Semantic Progress Function

DGX agent

arXiv:2604.22554v1 Announce Type: new Abstract: Transformations produced by image and video generation models often evolve in a highly non-linear manner: long stretches where the content barely change

applicationsarxiv-cs-cv
27 Apr 2026
Applications

ViFiCon: Vision and Wireless Association Via Self-Supervised Contrastive Learning

DGX agent

arXiv:2210.05513v2 Announce Type: replace Abstract: We introduce ViFiCon, a self-supervised contrastive scheme which learns a cross-modal association between vision and wireless modalities. Specifical

applicationsarxiv-cs-cv
27 Apr 2026
Research

When LoRA Betrays: Backdooring Text-to-Image Models by Masquerading as Benign Adapters

DGX agent

arXiv:2602.21977v4 Announce Type: replace Abstract: Low-Rank Adaptation (LoRA) has emerged as a leading technique for efficiently fine-tuning text-to-image diffusion models, and its widespread adoptio

researcharxiv-cs-cv
27 Apr 2026
Research

2L-LSH: A Locality-Sensitive Hash Function-Based Method For Rapid Point Cloud Indexing

DGX agent

arXiv:2604.21442v1 Announce Type: new Abstract: The development of 3D scanning technology has enabled the acquisition of massive point cloud models with diverse structures and large scales, thereby pr

researcharxiv-cs-cv
24 Apr 2026
Applications

A Probabilistic Framework for Improving Dense Object Detection in Underwater Image Data via Annealing-Based Data Augmentation

DGX agent

arXiv:2604.21198v1 Announce Type: new Abstract: Object detection models typically perform well on images captured in controlled environments with stable lighting, water clarity, and viewpoint, but the

applicationsarxiv-cs-cv
24 Apr 2026
Safety

Adaptive Moments are Surprisingly Effective for Plug-and-Play Diffusion Sampling

DGX agent

arXiv:2603.16797v2 Announce Type: replace-cross Abstract: Guided diffusion sampling relies on approximating often intractable likelihood scores, which introduces significant noise into the sampling dy

safetyarxiv-cs-cv
24 Apr 2026
Local Ai

an interpretable vision transformer framework for automated brain tumor classification

DGX agent

arXiv:2604.21311v1 Announce Type: new Abstract: Brain tumors represent one of the most critical neurological conditions, where early and accurate diagnosis is directly correlated with patient survival

local-aiarxiv-cs-cv
24 Apr 2026
Research

Anatomy-Aware Text-Visual Fusion with Dual-Perspective Prompts for Fine-Grained Lumbar Spine Segmentation

DGX agent

arXiv:2504.03476v2 Announce Type: replace Abstract: Accurate lumbar spine segmentation is crucial for diagnosing spinal disorders. Existing methods typically use coarse-grained segmentation strategies

researcharxiv-cs-cv
24 Apr 2026
Model Releases

APCoTTA: Continual Test-Time Adaptation for Semantic Segmentation of Airborne LiDAR Point Clouds

DGX agent

arXiv:2505.09971v3 Announce Type: replace Abstract: Airborne laser scanning (ALS) point cloud semantic segmentation is a fundamental task for large-scale 3D scene understanding. Fixed models deployed

model-releasesarxiv-cs-cv
24 Apr 2026
Model Releases

ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response

DGX agent

arXiv:2604.21199v1 Announce Type: cross Abstract: Time series question-answering (TSQA), in which we ask natural language questions to infer and reason about properties of time series, is a promising

model-releasesarxiv-cs-cv
24 Apr 2026
Safety

ATATA: One Algorithm to Align Them All

DGX agent

arXiv:2601.11194v2 Announce Type: replace Abstract: We suggest a new multi-modal algorithm for joint inference of paired structurally aligned samples with Rectified Flow models. While some existing me

safetyarxiv-cs-cv
24 Apr 2026
Safety

AttDiff-GAN: A Hybrid Diffusion-GAN Framework for Facial Attribute Editing

DGX agent

arXiv:2604.21289v1 Announce Type: new Abstract: Facial attribute editing aims to modify target attributes while preserving attribute-irrelevant content and overall image fidelity. Existing GAN-based m

safetyarxiv-cs-cv
24 Apr 2026
Local Ai

AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe

DGX agent

arXiv:2604.20936v1 Announce Type: cross Abstract: We present AttentionBender, a tool that manipulates cross-attention in Video Diffusion Transformers to help artists probe the internal mechanics of bl

local-aiarxiv-cs-cv
24 Apr 2026
Safety

Automated Annotation of Shearographic Measurements Enabling Weakly Supervised Defect Detection

DGX agent

arXiv:2512.06171v2 Announce Type: replace Abstract: Shearography is an interferometric technique sensitive to surface displacement gradients, providing high sensitivity for detecting subsurface defect

safetyarxiv-cs-cv
24 Apr 2026
Research

Back to Source: Open-Set Continual Test-Time Adaptation via Domain Compensation

DGX agent

arXiv:2604.21772v1 Announce Type: new Abstract: Test-Time Adaptation (TTA) aims to mitigate distributional shifts between training and test domains during inference time. However, existing TTA methods

researcharxiv-cs-cv
24 Apr 2026
Research

Beyond the Frame: Generating 360 Panoramic Videos from Perspective Videos

DGX agent

arXiv:2504.07940v3 Announce Type: replace Abstract: 360{eg} videos have emerged as a promising medium to represent our dynamic visual world. Compared to the 'tunnel vision' of standard cameras, their

researcharxiv-cs-cv
24 Apr 2026
Safety

BiTDiff: Fine-Grained 3D Conducting Motion Generation via BiMamba-Transformer Diffusion

DGX agent

arXiv:2604.04395v2 Announce Type: replace Abstract: 3D conducting motion generation aims to synthesize fine-grained conductor motions from music, with broad potential in music education, virtual perfo

safetyarxiv-cs-cv
24 Apr 2026
Applications

Bridging Supervision Gaps: A Unified Framework for Remote Sensing Change Detection

DGX agent

arXiv:2601.17747v2 Announce Type: replace Abstract: Change detection (CD) aims to identify surface changes from multi-temporal remote sensing imagery. In real-world scenarios, Pixel-level change label

applicationsarxiv-cs-cv
24 Apr 2026
Safety

CHRep: Cross-modal Histology Representation and Post-hoc Calibration for Spatial Gene Expression Prediction

DGX agent

arXiv:2604.21573v1 Announce Type: new Abstract: Spatial transcriptomics (ST) enables spatially resolved gene profiling but remains expensive and low-throughput, limiting large-cohort studies and routi

safetyarxiv-cs-cv
24 Apr 2026
Research

Clinically-Informed Modeling for Pediatric Brain Tumor Classification from Whole-Slide Histopathology Images

DGX agent

arXiv:2604.21060v1 Announce Type: new Abstract: Accurate diagnosis of pediatric brain tumors, starting with histopathology, presents unique challenges for deep learning, including severe data scarcity

researcharxiv-cs-cv
24 Apr 2026
Research

Component-Based Out-of-Distribution Detection

DGX agent

arXiv:2604.21546v1 Announce Type: new Abstract: Out-of-Distribution (OOD) detection requires sensitivity to subtle shifts without overreacting to natural In-Distribution (ID) diversity. However, from

researcharxiv-cs-cv
24 Apr 2026
← Previous
1…216217218219220…261
Next →