AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
28 May 2026

Adaptive Temporal Gating of Longitudinal Magnetic Resonance Imaging for Alzheimer's Prediction

ResearchDGX agent

arXiv:2605.28397v1 Announce Type: new Abstract: Predicting conversion from Mild Cognitive Impairment (MCI) to Alzheimer's Disease (AD) is critical for early intervention. Current deep learning paradig

Alterbute: Editing Intrinsic Attributes of Objects in Images

ResearchDGX agent

arXiv:2601.10714v2 Announce Type: replace Abstract: We introduce Alterbute, a diffusion-based method for editing an object's intrinsic attributes in an image. We allow changing color, texture, materia

An analytic theory of convolutional neural network inverse problems solvers

ResearchDGX agent

arXiv:2601.10334v2 Announce Type: replace Abstract: Supervised convolutional neural networks (CNNs) are widely used to solve imaging inverse problems, achieving state-of-the-art performance in numerou


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

An Empirical Study on Variance-based MC Dropout Uncertainty-Error Correlation in 2D Brain Tumor Segmentation

ResearchDGX agent

arXiv:2510.15541v2 Announce Type: replace-cross Abstract: Accurate brain tumor segmentation from MRI is vital for diagnosis and treatment planning. Although Monte Carlo (MC) Dropout is widely used to

AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source Applications

Model ReleasesDGX agent

arXiv:2605.27761v1 Announce Type: new Abstract: The rapid development of GUI foundation models and mobile GUI agents has spurred numerous evaluation benchmarks, yet most rely on simulated environments

Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?

Model ReleasesDGX agent

arXiv:2508.11011v2 Announce Type: replace Abstract: Construction safety inspections typically involve a human inspector identifying safety concerns on-site. With the rise of powerful Vision Language M

AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning

SafetyDGX agent

arXiv:2605.28809v1 Announce Type: new Abstract: Class-Incremental Learning (CIL) is important in building real-world learning systems. In CLIP-based CIL, the model performs classification by comparing

Artemis: Structured Visual Reasoning for Perception Policy Learning

SafetyDGX agent

arXiv:2512.01988v2 Announce Type: replace Abstract: Recent reinforcement-learning frameworks for visual perception policy usually incorporate intermediate reasoning chains expressed in natural languag

Asynchronous Remote Sensing Time-Series Fusion for Cloud Removal and Anytime Reconstruction

Model ReleasesDGX agent

arXiv:2605.27726v1 Announce Type: new Abstract: Frequent cloud cover severely limits the usability of Sentinel-2 (S2) optical time series for Earth surface monitoring. Sentinel-1 (S1) SAR provides all

Automated Estimation of Impact Time, Impact Location, and Shuttlecock Speed in Badminton Smashes Using Event Cameras

SafetyDGX agent

arXiv:2605.28011v1 Announce Type: new Abstract: Quantifying impact phenomena in badminton smashes is important for evaluating both athletic performance and equipment; however, conventional measurement

Automatic Pruning Discovery for Large Language Models

TutorialsDGX agent

arXiv:2511.15390v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved remarkable performance on a wide range of tasks, hindering real-world deployment due to their massive siz

Benchmarking Ultrasound Foundation Models for Fetal Plane Classification

Model ReleasesDGX agent

arXiv:2605.27796v1 Announce Type: cross Abstract: Ultrasound is widely used in obstetric care due to its safety, accessibility, and real-time imaging. However, interpretation remains operator-dependen

Beyond Surrogate Gradients: Fully Differentiable Token Pruning for Vision-Language Models

ResearchDGX agent

arXiv:2605.28051v1 Announce Type: new Abstract: Visual token pruning reduces the computational cost of Vision-Language Models (VLMs) by removing redundant visual tokens. Existing methods typically rel

Bias Leaves a Gradient Trail: Label-Free Bias Identification via Gradient Probes on Concept Decompositions

Model ReleasesDGX agent

arXiv:2605.28780v1 Announce Type: new Abstract: Vision classifiers can exploit spurious correlations, achieving high in-distribution accuracy yet failing under distribution shift. Existing approaches

Bound-Constrained Sparse Representation for Electrical Impedance Tomography

ResearchDGX agent

arXiv:2605.28392v1 Announce Type: new Abstract: This study proposes a bound-constrained sparse representation (BC-SR) framework for electrical impedance tomography (EIT), aimed at improving conductivi

Bounded-Compute Multimodal Regression for Product-Rating Prediction

Model ReleasesDGX agent

arXiv:2605.27737v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly attractive for multimodal quality assessment, but their default reliance on autoregressive text generatio

Bridging the Generalization Gap in Adverse Weather Segmentation: A Training Recipe Perspective

ResearchDGX agent

arXiv:2605.27962v1 Announce Type: new Abstract: This paper describes our approach for the 8th UG2+ Workshop (CVPR 2026) Track~2, which targets semantic segmentation of outdoor scenes degraded by five

Bridging the Sampling Distribution Shift in Radio Map Estimation: A Trajectory-Aware Paradigm

ResearchDGX agent

arXiv:2605.28234v1 Announce Type: new Abstract: Learning-based radio map estimation (RME) plays a critical role in UAV-assisted wireless sensing, enabling tasks such as coverage prediction and network

Category-Level 3D Correspondence in Camera Space via Morphable Object Priors

Model ReleasesDGX agent

arXiv:2605.28257v1 Announce Type: new Abstract: Understanding 3D objects from images is fundamental to robotics and AR/VR applications. While recent work has made progress in category-level pose estim

Chirpy3D: Part-Aware Multi-View Diffusion for Creative Fine-Grained Object Generation

Model ReleasesDGX agent

arXiv:2501.04144v3 Announce Type: replace Abstract: Understanding and generating the fine-grained structure of objects -- such as birds with species-specific beaks, wings, and tails -- is a long-stand

CLEAR-NeRF: Collinearity and Local-region Enhanced Accurate 3D Reconstruction in Unbounded Scenes

TutorialsDGX agent

arXiv:2605.28125v1 Announce Type: new Abstract: Many real-world 3D reconstruction applications demand photorealism and metric accuracy across unbounded, complex scenes with challenging lighting and im

ClothTransformer: Unified Latent-Space Transformers for Scalable Cloth Simulation

ResearchDGX agent

arXiv:2605.27852v1 Announce Type: cross Abstract: Unified and scalable Transformers have recently achieved remarkable success in modeling diverse phenomena traditionally associated with computer graph

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning

Model ReleasesDGX agent

arXiv:2605.28056v1 Announce Type: new Abstract: Portrait animation methods have achieved substantial visual quality and lip synchronization, but fine-grained manipulation of the eye region still faces

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification

Model ReleasesDGX agent

arXiv:2508.21046v3 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) models built on pre-trained Vision-Language Models (VLMs) require extensive post-training, resulting in high com

Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization

SafetyDGX agent

arXiv:2605.28615v1 Announce Type: new Abstract: Despite the rapid progress of text-to-image (T2I) models, generating images that accurately reflect complex compositional prompts (covering attribute bi

Con-DSO: Learning Short-Horizon Consistency Priors for RGB-D Direct Sparse Odometry

SafetyDGX agent

arXiv:2605.27952v1 Announce Type: new Abstract: Visual odometry (VO) is a fundamental component in robotics and augmented reality. RGB-D direct VO benefits from metric depth measurements, but it can d

CPPO: Contrastive Perception Policy Optimization for VLM Agents

SafetyDGX agent

arXiv:2601.00501v2 Announce Type: replace Abstract: We introduce CPPO, a Contrastive Perception Policy Optimization method for finetuning vision--language models (VLMs). Reliable perception is a core

Cross-Modal Action Recognition in Egocentric Video Using Mamba: Integrating RGB and Hand Skeleton Streams via CLS Token Fusion Strategies

ResearchDGX agent

arXiv:2605.24302v2 Announce Type: replace Abstract: Egocentric action recognition is a challenging task due to erratic camera motion, frequent hand occlusion, and the difficulty of maintaining consist

CuriosAI Submission to the CASTLE Challenge at EgoVis 2026

ResearchDGX agent

arXiv:2605.27800v1 Announce Type: new Abstract: CASTLE 2026 asks 185 multiple-choice questions over 600+ hours of synchronized multi-view egocentric video. We explore two approaches on top of a shared

D^2Turb: Depth-Aware Simulation and Decoupled Learning for Single-Frame Atmospheric Turbulence Mitigation

TutorialsDGX agent

arXiv:2605.27460v1 Announce Type: new Abstract: Single-frame atmospheric turbulence mitigation is inherently ill-posed due to spatially varying blur coupled with non-rigid geometric distortion. Existi

DebFilter: Eradicating Biases Stashed in Value

SafetyDGX agent

arXiv:2605.28167v1 Announce Type: new Abstract: Text-to-image diffusion models, which are theoretically equivalent to score-based generative models, generate images through a multi-step denoising proc

Decoupled Training with Local Reinforcement Fine-Tuning in Federated Learning

Model ReleasesDGX agent

arXiv:2605.27900v1 Announce Type: new Abstract: Federated Learning (FL) with pre-trained Vision-Language Models (VLMs) has emerged as a promising paradigm for various downstream tasks. By leveraging i

Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized Distillation

Model ReleasesDGX agent

arXiv:2605.28587v1 Announce Type: new Abstract: Understanding dynamic 3D environments is essential for safe autonomous driving, particularly when reasoning about human-centric, nonrigid agents. Howeve

DiscoForcing: A Unified Framework for Real-Time Audio-Driven Character Control with Diffusion Forcing

SafetyDGX agent

arXiv:2605.28491v1 Announce Type: new Abstract: We study real-time audio-responsive character control as a deployment-faithful problem: strictly causal, bounded-latency streaming that must generate co

DODO: Discrete OCR Diffusion Models

ResearchDGX agent

arXiv:2602.16872v2 Announce Type: replace Abstract: Optical Character Recognition (OCR) is a fundamental task for digitizing information, serving as a critical bridge between visual data and textual u

DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.28544v1 Announce Type: new Abstract: Pretrained foundation models have become an important basis for end-to-end autonomous driving. In contrast to vision-language models pretrained primaril

Dual-branch Distilled Transformer for Efficient Asymmetric UAV Tracking

Local AiDGX agent

arXiv:2605.28018v1 Announce Type: new Abstract: Given the real-time demands of UAV tracking, many methods simplify the backbone to reduce computation, but this often weakens feature representation and

EchoAvatar: Real-time Generative Avatar Animation from Audio Streams

ResearchDGX agent

arXiv:2605.28272v1 Announce Type: new Abstract: Real-time synthesis of high-fidelity 3D character motion from audio is a pivotal component for next-generation interactive avatars and virtual assistant

EgoRelight: Egocentric Human Capture and Illumination Recovery for Relightable and Photoreal Avatar Rendering

ResearchDGX agent

arXiv:2605.28401v1 Announce Type: new Abstract: Mixed Reality (MR) headsets promise a future of immersive telepresence where virtual humans blend indistinguishably into real or virtual surroundings. A

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning

Model ReleasesDGX agent

arXiv:2510.27266v2 Announce Type: replace Abstract: Autonomous graphical user interface (GUI) agents rely on accurate GUI grounding, which maps language instructions to on-screen coordinates, to execu

Enhancing Ultra-low-field MRI with Segmentation-guided Adversarial Learning

ResearchDGX agent

arXiv:2605.28016v1 Announce Type: new Abstract: Ultra-low-field (ULF) MRI offers portable and low-cost imaging but suffers from poor image quality. To address this, we present our submission to the 20

EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection

SafetyDGX agent

arXiv:2605.28630v1 Announce Type: new Abstract: Zero-Shot Anomaly Detection (ZSAD) aims to detect anomalies in unseen domains without target-domain adaptation. Recent CLIP-based methods have shown pro

Evaluating the Feasibility of Inferring Dietary Behavior Change Receptivity from Egocentric Images of Eating Environment

ResearchDGX agent

arXiv:2605.27950v1 Announce Type: new Abstract: Accurately assessing dietary behavior change receptivity is essential for designing effective just-in-time adaptive interventions (JITAIs) that promote

Event-based Motion & Appearance Fusion for 6D Object Pose Tracking

ResearchDGX agent

arXiv:2603.08264v2 Announce Type: replace Abstract: Object pose tracking is a fundamental and essential task for robotics to perform tasks in the home and industrial settings. The most commonly used s

EventShiftFlow: Towards Hardware-efficient FPGA-based Flow Estimation

Model ReleasesDGX agent

arXiv:2605.28312v1 Announce Type: cross Abstract: Event-based vision sensors offer asynchronous, high-temporal-resolution measurements that are attractive for low-latency robotic perception, but many

Every9D-21M: Large-Scale Real-World 9D Canonicalization of Everyday Objects

Model ReleasesDGX agent

arXiv:2605.28270v1 Announce Type: new Abstract: Estimating the 9D pose of everyday objects from a single real-world image remains challenging. This is largely due to the lack of large-scale supervisio

Explaining Digital Pathology Models via Clustering Activations

ResearchDGX agent

arXiv:2511.14558v2 Announce Type: replace Abstract: We present a clustering-based explainability technique for digital pathology models based on convolutional neural networks. Unlike commonly used met

Explicit Critic Guidance for Aligning Diffusion Models

ResearchDGX agent

arXiv:2605.27736v1 Announce Type: cross Abstract: Online reinforcement learning is becoming increasingly important for aligning diffusion models with non-differentiable objectives. However, existing m

FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models

ResearchDGX agent

arXiv:2512.16483v2 Announce Type: replace Abstract: Visual Autoregressive (VAR) modeling departs from the next-token prediction paradigm of traditional Autoregressive (AR) models through next-scale pr

Fine-Tuning Vision-Language Models for Understanding Current Damage and Scoring Priority with Quality Guard Agent

AgentsDGX agent

arXiv:2605.27452v1 Announce Type: new Abstract: Bridge inspection in Japan requires mandatory visual assessments every five years, yet qualitative damage ratings (levels a-e) assigned by different eng

ForestHG-Trace: Traceable Long-Horizon Ecological Reasoning over Large-Scale Forest Scenes

Model ReleasesDGX agent

arXiv:2605.27590v1 Announce Type: new Abstract: Remote sensing question answering (RS-QA) often requires more than direct semantic prediction, especially in large-scale forest scenes where ecological

From Affect to Complex Behavior: Advancing Multimodal Human-Centered AI at the 10th ABAW Workshop & Competition

SafetyDGX agent

arXiv:2605.27451v1 Announce Type: new Abstract: The 10th Affective & Behavior Analysis in-the-Wild (ABAW) Workshop and Competition, held at CVPR 2026, continues to advance research on modelling, analy

From Kellgren-Lawrence to Calcium Pyrophosphate Crystal Deposition: A Soft-Labelling Framework for Knee Osteoarthritis Assessmen

ResearchDGX agent

arXiv:2605.28176v1 Announce Type: new Abstract: Background and objective. Conventional Deep Learning (DL) approaches for Knee Osteoarthritis (KOA) grading rely on one-hot labels, which fail to capture

From Pixels to Words -- Towards Native One-Vision Models at Scale

SafetyDGX agent

arXiv:2605.28820v1 Announce Type: new Abstract: Current vision-language models (VLMs) typically stitch together separate image encoders and language decoders via multi-stage alignment, a modular frame

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players

Model ReleasesDGX agent

arXiv:2605.28816v1 Announce Type: new Abstract: World models for interactive video generation have largely focused on single-agent settings, where future observations are generated from a single contr

GEM: Generative Supervision Helps Embodied Intelligence

ApplicationsDGX agent

arXiv:2605.28548v1 Announce Type: new Abstract: Embodied Vision-Language Models (VLMs) have demonstrated impressive performance and generalization in robotics, particularly within Vision-Language-Acti

HarmoVid: Relightful Video Portrait Harmonization

Local AiDGX agent

arXiv:2605.28811v1 Announce Type: new Abstract: We present a method for harmonizing the lighting of a foreground video to match a target background scene, adjusting shadows, color tone, and illuminati

Hierarchical Relation-augmented Representation Generalization for Few-shot Action Recognition

TutorialsDGX agent

arXiv:2504.10079v4 Announce Type: replace Abstract: Few-shot action recognition (FSAR) aims to recognize novel action categories with few exemplars. Existing methods typically learn frame-level repres

HiRQA: Hierarchical Ranking and Quality Alignment for Opinion-Unaware Image Quality Assessment

SafetyDGX agent

arXiv:2508.15130v2 Announce Type: replace Abstract: Despite significant progress in no-reference image quality assessment (NR-IQA), dataset biases and reliance on subjective labels continue to hinder

IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction

Model ReleasesDGX agent

arXiv:2510.06928v2 Announce Type: replace Abstract: Autoregressive models have emerged as a powerful paradigm for visual content creation, but often overlook the intrinsic structural properties of vis

← Previous
1…110111112113114…211
Next →