AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
27 Apr 2026

Unlocking Optical Prior: Spectrum-Guided Knowledge Transfer for SAR Generalized Category Discovery

SafetyDGX agent

arXiv:2604.22174v1 Announce Type: new Abstract: Generalized Category Discovery (GCD) holds significant promise for the label-scarce Synthetic Aperture Radar (SAR) domain, yet its efficacy is severely

Useful nonrobust features are ubiquitous in biomedical images

TutorialsDGX agent

arXiv:2604.22579v1 Announce Type: cross Abstract: We study whether deep networks for medical imaging learn useful nonrobust features - predictive input patterns that are not human interpretable and hi

V-MAGE: A Game Evaluation Framework for Assessing Vision-Centric Capabilities in Multimodal Large Language Models

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2504.06148v3 Announce Type: replace Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in visual-text processing. However, existi

Video Analysis and Generation via a Semantic Progress Function

ApplicationsDGX agent

arXiv:2604.22554v1 Announce Type: new Abstract: Transformations produced by image and video generation models often evolve in a highly non-linear manner: long stretches where the content barely change

ViFiCon: Vision and Wireless Association Via Self-Supervised Contrastive Learning

ApplicationsDGX agent

arXiv:2210.05513v2 Announce Type: replace Abstract: We introduce ViFiCon, a self-supervised contrastive scheme which learns a cross-modal association between vision and wireless modalities. Specifical

When LoRA Betrays: Backdooring Text-to-Image Models by Masquerading as Benign Adapters

ResearchDGX agent

arXiv:2602.21977v4 Announce Type: replace Abstract: Low-Rank Adaptation (LoRA) has emerged as a leading technique for efficiently fine-tuning text-to-image diffusion models, and its widespread adoptio

24 Apr 2026

2L-LSH: A Locality-Sensitive Hash Function-Based Method For Rapid Point Cloud Indexing

ResearchDGX agent

arXiv:2604.21442v1 Announce Type: new Abstract: The development of 3D scanning technology has enabled the acquisition of massive point cloud models with diverse structures and large scales, thereby pr

A Probabilistic Framework for Improving Dense Object Detection in Underwater Image Data via Annealing-Based Data Augmentation

ApplicationsDGX agent

arXiv:2604.21198v1 Announce Type: new Abstract: Object detection models typically perform well on images captured in controlled environments with stable lighting, water clarity, and viewpoint, but the

Adaptive Moments are Surprisingly Effective for Plug-and-Play Diffusion Sampling

SafetyDGX agent

arXiv:2603.16797v2 Announce Type: replace-cross Abstract: Guided diffusion sampling relies on approximating often intractable likelihood scores, which introduces significant noise into the sampling dy

an interpretable vision transformer framework for automated brain tumor classification

Local AiDGX agent

arXiv:2604.21311v1 Announce Type: new Abstract: Brain tumors represent one of the most critical neurological conditions, where early and accurate diagnosis is directly correlated with patient survival

Anatomy-Aware Text-Visual Fusion with Dual-Perspective Prompts for Fine-Grained Lumbar Spine Segmentation

ResearchDGX agent

arXiv:2504.03476v2 Announce Type: replace Abstract: Accurate lumbar spine segmentation is crucial for diagnosing spinal disorders. Existing methods typically use coarse-grained segmentation strategies

APCoTTA: Continual Test-Time Adaptation for Semantic Segmentation of Airborne LiDAR Point Clouds

Model ReleasesDGX agent

arXiv:2505.09971v3 Announce Type: replace Abstract: Airborne laser scanning (ALS) point cloud semantic segmentation is a fundamental task for large-scale 3D scene understanding. Fixed models deployed

ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response

Model ReleasesDGX agent

arXiv:2604.21199v1 Announce Type: cross Abstract: Time series question-answering (TSQA), in which we ask natural language questions to infer and reason about properties of time series, is a promising

ATATA: One Algorithm to Align Them All

SafetyDGX agent

arXiv:2601.11194v2 Announce Type: replace Abstract: We suggest a new multi-modal algorithm for joint inference of paired structurally aligned samples with Rectified Flow models. While some existing me

AttDiff-GAN: A Hybrid Diffusion-GAN Framework for Facial Attribute Editing

SafetyDGX agent

arXiv:2604.21289v1 Announce Type: new Abstract: Facial attribute editing aims to modify target attributes while preserving attribute-irrelevant content and overall image fidelity. Existing GAN-based m

AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe

Local AiDGX agent

arXiv:2604.20936v1 Announce Type: cross Abstract: We present AttentionBender, a tool that manipulates cross-attention in Video Diffusion Transformers to help artists probe the internal mechanics of bl

Automated Annotation of Shearographic Measurements Enabling Weakly Supervised Defect Detection

SafetyDGX agent

arXiv:2512.06171v2 Announce Type: replace Abstract: Shearography is an interferometric technique sensitive to surface displacement gradients, providing high sensitivity for detecting subsurface defect

Back to Source: Open-Set Continual Test-Time Adaptation via Domain Compensation

ResearchDGX agent

arXiv:2604.21772v1 Announce Type: new Abstract: Test-Time Adaptation (TTA) aims to mitigate distributional shifts between training and test domains during inference time. However, existing TTA methods

Beyond the Frame: Generating 360 Panoramic Videos from Perspective Videos

ResearchDGX agent

arXiv:2504.07940v3 Announce Type: replace Abstract: 360{eg} videos have emerged as a promising medium to represent our dynamic visual world. Compared to the 'tunnel vision' of standard cameras, their

BiTDiff: Fine-Grained 3D Conducting Motion Generation via BiMamba-Transformer Diffusion

SafetyDGX agent

arXiv:2604.04395v2 Announce Type: replace Abstract: 3D conducting motion generation aims to synthesize fine-grained conductor motions from music, with broad potential in music education, virtual perfo

Bridging Supervision Gaps: A Unified Framework for Remote Sensing Change Detection

ApplicationsDGX agent

arXiv:2601.17747v2 Announce Type: replace Abstract: Change detection (CD) aims to identify surface changes from multi-temporal remote sensing imagery. In real-world scenarios, Pixel-level change label

CHRep: Cross-modal Histology Representation and Post-hoc Calibration for Spatial Gene Expression Prediction

SafetyDGX agent

arXiv:2604.21573v1 Announce Type: new Abstract: Spatial transcriptomics (ST) enables spatially resolved gene profiling but remains expensive and low-throughput, limiting large-cohort studies and routi

Clinically-Informed Modeling for Pediatric Brain Tumor Classification from Whole-Slide Histopathology Images

ResearchDGX agent

arXiv:2604.21060v1 Announce Type: new Abstract: Accurate diagnosis of pediatric brain tumors, starting with histopathology, presents unique challenges for deep learning, including severe data scarcity

Component-Based Out-of-Distribution Detection

ResearchDGX agent

arXiv:2604.21546v1 Announce Type: new Abstract: Out-of-Distribution (OOD) detection requires sensitivity to subtle shifts without overreacting to natural In-Distribution (ID) diversity. However, from

Context Unrolling in Omni Models

ResearchDGX agent

arXiv:2604.21921v1 Announce Type: new Abstract: We present Omni, a unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representati

Cross-Distribution Diffusion Priors-Driven Iterative Reconstruction for Sparse-View CT

ResearchDGX agent

arXiv:2509.13576v2 Announce Type: replace-cross Abstract: Sparse-View CT (SVCT) reconstruction enhances temporal resolution and reduces radiation dose, yet its clinical use is hindered by artifacts du

DAVIS: OOD Detection via Dominant Activations and Variance for Increased Separation

Model ReleasesDGX agent

arXiv:2601.22703v2 Announce Type: replace Abstract: Detecting out-of-distribution (OOD) inputs is a critical safeguard for deploying machine learning models in the real world. However, most post-hoc d

DCMorph: Face Morphing via Dual-Stream Cross-Attention Diffusion

ResearchDGX agent

arXiv:2604.21627v1 Announce Type: new Abstract: Advancing face morphing attack techniques is crucial to anticipate evolving threats and develop robust defensive mechanisms for identity verification sy

Deep kernel video approximation for unsupervised action segmentation

ResearchDGX agent

arXiv:2604.21572v1 Announce Type: new Abstract: This work focuses on per-video unsupervised action segmentation, which is of interest to applications where storing large datasets is either not possibl

Demystifying Action Space Design for Robotic Manipulation Policies

SafetyDGX agent

arXiv:2602.23408v2 Announce Type: replace-cross Abstract: The specification of the action space plays a pivotal role in imitation-based robotic manipulation policy learning, fundamentally shaping the

DepthMaster: Taming Diffusion Models for Monocular Depth Estimation

SafetyDGX agent

arXiv:2501.02576v2 Announce Type: replace Abstract: Monocular depth estimation within the diffusion-denoising paradigm demonstrates impressive generalization ability but suffers from low inference spe

DiffNR: Diffusion-Enhanced Neural Representation Optimization for Sparse-View 3D Tomographic Reconstruction

ResearchDGX agent

arXiv:2604.21518v1 Announce Type: cross Abstract: Neural representations (NRs), such as neural fields and 3D Gaussians, effectively model volumetric data in computed tomography (CT) but suffer from se

Directional Confusions Reveal Divergent Inductive Biases Through Rate-Distortion Geometry in Human and Machine Vision

SafetyDGX agent

arXiv:2604.21909v1 Announce Type: new Abstract: Humans and modern vision models can reach similar classification accuracy while making systematically different kinds of mistakes - differing not in how

Discriminative-Generative Synergy for Occlusion Robust 3D Human Mesh Recovery

ApplicationsDGX agent

arXiv:2604.21712v1 Announce Type: new Abstract: 3D human mesh recovery from monocular RGB images aims to estimate anatomically plausible 3D human models for downstream applications, but remains challe

Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision

Model ReleasesDGX agent

arXiv:2604.21461v1 Announce Type: new Abstract: Egocentric AI agents, such as smart glasses, rely on pointing gestures to resolve referential ambiguities in natural language commands. However, despite

DualSplat: Robust 3D Gaussian Splatting via Pseudo-Mask Bootstrapping from Reconstruction Failures

TutorialsDGX agent

arXiv:2604.21631v1 Announce Type: new Abstract: While 3D Gaussian Splatting (3DGS) achieves real-time photorealistic rendering, its performance degrades significantly when training images contain tran

EdgeFormer: local patch-based edge detection transformer on point clouds

Local AiDGX agent

arXiv:2604.21387v1 Announce Type: new Abstract: Edge points on 3D point clouds can clearly convey 3D geometry and surface characteristics, therefore, edge detection is widely used in many vision appli

Efficient Multi-Source Knowledge Transfer by Model Merging

Model ReleasesDGX agent

arXiv:2508.19353v2 Announce Type: replace-cross Abstract: While transfer learning is an effective strategy, it often overlooks the opportunity to leverage knowledge from numerous available models onli

Encoder-Free Human Motion Understanding via Structured Motion Descriptions

SafetyDGX agent

arXiv:2604.21668v1 Announce Type: new Abstract: The world knowledge and reasoning capabilities of text-based large language models (LLMs) are advancing rapidly, yet current approaches to human motion

Flow Matching for Conditional MRI-CT and CBCT-CT Image Synthesis

Model ReleasesDGX agent

arXiv:2510.04823v2 Announce Type: replace Abstract: Generating synthetic CT (sCT) from MRI or CBCT plays a crucial role in enabling MRI-only and CBCT-based adaptive radiotherapy, improving treatment p

Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models

ResearchDGX agent

arXiv:2604.21079v1 Announce Type: new Abstract: Vision-language models benefit from high-resolution images, but the increase in visual-token count incurs high compute overhead. Humans resolve this ten

From Codebooks to VLMs: Evaluating Automated Visual Discourse Analysis for Climate Change on Social Media

Model ReleasesDGX agent

arXiv:2604.21786v1 Announce Type: new Abstract: Social media platforms have become primary arenas for climate communication, generating millions of images and posts that - if systematically analysed -

Frozen LLMs as Map-Aware Spatio-Temporal Reasoners for Vehicle Trajectory Prediction

Local AiDGX agent

arXiv:2604.21479v1 Announce Type: new Abstract: Large language models (LLMs) have recently demonstrated strong reasoning capabilities and attracted increasing research attention in the field of autono

FryNet: Dual-Stream Adversarial Fusion for Non-Destructive Frying Oil Oxidation Assessment

SafetyDGX agent

arXiv:2604.21321v1 Announce Type: new Abstract: Monitoring frying oil degradation is critical for food safety, yet current practice relies on destructive wet-chemistry assays that provide no spatial i

GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure

Model ReleasesDGX agent

arXiv:2512.22274v2 Announce Type: replace Abstract: We introduce GeCo, a geometry-grounded metric for jointly detecting geometric deformation and occlusion-inconsistency artifacts in static scenes. By

Geometry-aided Vision-based Localization of Future Mars Helicopters in Challenging Illumination Conditions

Local AiDGX agent

arXiv:2502.09795v3 Announce Type: replace Abstract: Planetary exploration using aerial assets has the potential for unprecedented scientific discoveries on Mars. While NASA's Mars helicopter Ingenuity

Gmd: Gaussian mixture descriptor for pair matching of 3D fragments

Local AiDGX agent

arXiv:2604.21519v1 Announce Type: new Abstract: In the automatic reassembly of fragments acquired using laser scanners to reconstruct objects, a crucial step is the matching of fractured surfaces. In

GraphLeap: Decoupling Graph Construction and Convolution for Vision GNN Acceleration on FPGA

HardwareDGX agent

arXiv:2604.21290v1 Announce Type: new Abstract: Vision Graph Neural Networks (ViGs) represent an image as a graph of patch tokens, enabling adaptive, feature-driven neighborhoods. Unlike CNNs with fix

Grounding Video Reasoning in Physical Signals

Model ReleasesDGX agent

arXiv:2604.21873v1 Announce Type: new Abstract: Physical video understanding requires more than naming an event correctly. A model can answer a question about pouring, sliding, or collision from textu

HyperFM: An Efficient Hyperspectral Foundation Model with Spectral Grouping

Model ReleasesDGX agent

arXiv:2604.21127v1 Announce Type: new Abstract: The NASA PACE mission provides unprecedented hyperspectral observations of ocean color, aerosols, and clouds, offering new insights into how these compo

ID-Eraser: Proactive Defense Against Face Swapping via Identity Perturbation

ResearchDGX agent

arXiv:2604.21465v1 Announce Type: new Abstract: Deepfake technologies have rapidly advanced with modern generative AI, and face swapping in particular poses serious threats to privacy and digital secu

ImageHD: Energy-Efficient On-Device Continual Learning of Visual Representations via Hyperdimensional Computing

Local AiDGX agent

arXiv:2604.21280v1 Announce Type: new Abstract: On-device continual learning (CL) is critical for edge AI systems operating on non-stationary data streams, but most existing methods rely on backpropag

Information Bottleneck-Guided Heterogeneous Graph Learning for Interpretable Neurodevelopmental Disorder Diagnosis

TutorialsDGX agent

arXiv:2502.20769v3 Announce Type: replace Abstract: Developing interpretable models for neurodevelopmental disorders (NDDs) diagnosis presents significant challenges in effectively encoding, decoding,

Instance-level Visual Active Tracking with Occlusion-Aware Planning

ApplicationsDGX agent

arXiv:2604.21453v1 Announce Type: new Abstract: Visual Active Tracking (VAT) aims to control cameras to follow a target in 3D space, which is critical for applications like drone navigation and securi

Interpretable facial dynamics as behavioral and perceptual traces of deepfakes

Model ReleasesDGX agent

arXiv:2604.21760v1 Announce Type: new Abstract: Deepfake detection research has largely converged on deep learning approaches that, despite strong benchmark performance, offer limited insight into wha

KD-CVG: A Knowledge-Driven Approach for Creative Video Generation

SafetyDGX agent

arXiv:2604.21362v1 Announce Type: new Abstract: Creative Generation (CG) leverages generative models to automatically produce advertising content that highlights product features, and it has been a si

Latent Denoising Improves Visual Alignment in Large Multimodal Models

Model ReleasesDGX agent

arXiv:2604.21343v1 Announce Type: new Abstract: Large Multimodal Models (LMMs) such as LLaVA are typically trained with an autoregressive language modeling objective, providing only indirect supervisi

LatRef-Diff: Latent and Reference-Guided Diffusion for Facial Attribute Editing and Style Manipulation

ResearchDGX agent

arXiv:2604.21279v1 Announce Type: new Abstract: Facial attribute editing and style manipulation are crucial for applications like virtual avatars and photo editing. However, achieving precise control

Linear Image Generation by Synthesizing Exposure Brackets

ResearchDGX agent

arXiv:2604.21008v1 Announce Type: new Abstract: The life of a photo begins with photons striking the sensor, whose signals are passed through a sophisticated image signal processing (ISP) pipeline to

LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval

Model ReleasesDGX agent

arXiv:2505.15269v2 Announce Type: replace Abstract: Recent developments in Video Large Language Models (Video LLMs) have enabled models to process hour-long videos and exhibit exceptional performance.

← Previous
1…173174175176177…209
Next →