AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
3 Jun 2026

Video-Mirai: Autoregressive Video Diffusion Models Need Foresight

TutorialsDGX agent

arXiv:2606.03971v1 Announce Type: new Abstract: Causal video generators must predict from the past, but they need not learn only from it. In streaming autoregressive video diffusion, each emitted segm

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2512.22539v2 Announce Type: replace-cross Abstract: While Vision-Language-Action models (VLAs) are rapidly advancing towards generalist robot policies, it remains difficult to quantitatively und

VLESA: Vision-Language Embodied Safety Agent for Human Activity Monitoring

Model ReleasesDGX agent

arXiv:2606.03954v1 Announce Type: new Abstract: As AI systems increasingly assist humans in physical tasks, ensuring safety becomes paramount -- physical actions carry immediate and irreversible conse


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Weak Diffusion Priors Can Still Achieve Strong Inverse-Problem Performance

ResearchDGX agent

arXiv:2601.22443v2 Announce Type: replace-cross Abstract: Can a diffusion model trained on bedrooms recover human faces? Diffusion models are widely used as priors for inverse problems, but standard a

Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation?

Model ReleasesDGX agent

arXiv:2606.03837v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) and probing enable adaptation of foundation models using only a small number of trainable parameters, making it a

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation

ResearchDGX agent

arXiv:2606.03100v1 Announce Type: new Abstract: Recently, zero-shot 3D scene understanding via 2D Vision-Language Models (VLMs) has gained increasing research interest due to their promising spatial r

2 Jun 2026

3D Segment Anything Model with Visual Mamba for Diagnosing Placenta Accreta Spectrum

Model ReleasesDGX agent

arXiv:2606.00489v1 Announce Type: new Abstract: Placenta Accreta Spectrum (PAS) is a rare but highly dangerous obstetric disease. Early and accurate PAS diagnosis is critical for maternal health. Trad

3rd Place at CVPR 2026 CASTLE Challenge: Agentic Multi-View Long-Context Video Understanding via Hierarchical Knowledge Graph Retrieval

Model ReleasesDGX agent

arXiv:2606.01933v1 Announce Type: new Abstract: This paper presents our winning methodology for the CASTLE 2026 Challenge at the CVPR 2026 EgoVis Workshop, where our team secured third place globally.

4D Radar Meets LiDAR and Camera: Cooperative Perception under Adverse Weather

AgentsDGX agent

arXiv:2606.00416v1 Announce Type: new Abstract: Cooperative perception is important for autonomous driving but remains fragile when cameras and LiDAR degrade in adverse weather. We address this challe

A Closer Look at In-Distribution vs. Out-of-Distribution Accuracy for Open-Set Test-time Adaptation

Model ReleasesDGX agent

arXiv:2606.01973v1 Announce Type: cross Abstract: Open-set test-time adaptation (TTA) updates models on new data in the presence of input shifts and unknown output classes. While recent methods have m

A combination of noise and bilateral filters achieve supralinear and scalable adversarial robustness in CNNs

ApplicationsDGX agent

arXiv:2606.02267v1 Announce Type: cross Abstract: The vulnerability of deep neural networks to adversarial examples poses a significant challenge for real-world deployment. Existing techniques to enha

A Hypertoroidal Covering for Perfect Color Equivariance

ResearchDGX agent

arXiv:2603.04256v3 Announce Type: replace Abstract: When the color distribution of input images changes at inference, the performance of conventional neural network architectures drops considerably. A

A Modelling and Evaluation Framework for EuroCrops-Driven Sentinel-2 Crop Segmentation

SafetyDGX agent

arXiv:2606.00676v1 Announce Type: new Abstract: This work presents a configurable pipeline for generating semantic-segmentation-ready agricultural datasets from Sentinel-2 imagery and EuroCrops parcel

A Multiscale Network with Supervised Contrastive Learning for Real-Time Facial Emotion Recognition

ResearchDGX agent

arXiv:2606.01069v1 Announce Type: new Abstract: Real-time emotion recognition from facial expressions is a challenging task, particularly in video-based scenarios where multiple emotional states may o

A Systematic Benchmark of Intraoperative Ultrasound-to-MR Synthesis for Brain Tumour Surgery

Model ReleasesDGX agent

arXiv:2606.00630v1 Announce Type: new Abstract: Intraoperative ultrasound (ioUS) is a versatile, cost-effective modality in brain tumour surgery, but its interpretation is difficult: acquisition plane

Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.02459v1 Announce Type: new Abstract: Enabling Vision-Language Models (VLMs) to perform spatial reasoning remains challenging. Existing approaches treat VLMs as passive observers, which is d

ActMVS: Active Scene Reconstruction with Monocular Multi-View Stereo

ResearchDGX agent

arXiv:2606.01367v1 Announce Type: cross Abstract: Active scene reconstruction enables robots/UAVs to autonomously plan trajectories and reconstruct environments without costly manual data acquisition.

Adaptive Dense Evidence Refinement for Video Relational Reasoning for VRR-QA Challenge

ResearchDGX agent

arXiv:2606.01104v1 Announce Type: new Abstract: VRR-QA evaluates whether video-language systems can infer spatial, temporal, viewpoint, depth, and visibility relations that are not always resolved by

Advances in Neural 3D Mesh Texturing: A Survey

TutorialsDGX agent

arXiv:2606.00137v1 Announce Type: new Abstract: Texturing 3D meshes plays a vital role in determining the visual realism of digital objects and scenes. Although recent generative 3D approaches based o

Adversarial Attacks on Robot Localization Systems via Deep Feature Perturbation

SafetyDGX agent

arXiv:2606.01892v1 Announce Type: new Abstract: Robot localization systems are critical for autonomous navigation and safety. Adversarial perturbations can mislead these systems, resulting in mislocal

AFUN: Towards an Affordance Foundation Model for Functionality Understanding

Local AiDGX agent

arXiv:2606.02551v1 Announce Type: cross Abstract: Affordance understanding bridges visual perception and physical action, serving as an explainable interface for robot manipulation in open and unstruc

Agent Skills Should Go Beyond Text: The Case for Visual Skills

AgentsDGX agent

arXiv:2606.01414v1 Announce Type: new Abstract: Reusable skills are a key mechanism for extending agent capabilities, allowing agents to accumulate experience and solve increasingly complex tasks. Yet

AlbedoEdit: Unified Instance-Level Video Editing with Albedo Guidance

ApplicationsDGX agent

arXiv:2606.01362v1 Announce Type: cross Abstract: Video generative models have achieved remarkable progress in synthesizing photorealistic video sequences. However, enabling broader and more creative

{alpha}Depth: Learning Single-Pass Soft Boundary Decomposition for Stereo Conversion

Local AiDGX agent

arXiv:2606.00386v1 Announce Type: new Abstract: Accurately modeling soft boundaries, e.g., hair and defocus blur, is a fundamental challenge in stereo conversion due to the ambiguous blending of foreg

An Attribute-Based Measure of Video Complexity

ResearchDGX agent

arXiv:2606.00640v1 Announce Type: new Abstract: A new framework for the estimation of the complexity posed by video-question pairs to video-LLMs, Video Attribute-Based Complexity (VideoABC), is propos

An Effective Solution for the CVPR 2026 8th UG2+ Challenge Track 3: Dynamic Object Segmentation in Turbulence

ResearchDGX agent

arXiv:2606.00522v1 Announce Type: new Abstract: In this work, we present our solution for the 8th UG2+ Challenge (CVPR 2026) Track 3: Dynamic Object Segmentation in Turbulence (DOST). Our method is bu

An explainable hierarchical self attention-based approach for tremor detection in the time domain

TutorialsDGX agent

arXiv:2606.00461v1 Announce Type: new Abstract: Tremor is a common movement disorder associated with conditions like Parkinson's disease and Essential tremor, traditionally diagnosed through expert cl

Analysis of Ethnic Disparities in Autism Spectrum Disorder among Toddlers

ResearchDGX agent

arXiv:2606.01217v1 Announce Type: new Abstract: Autism Spectrum Disorder (ASD) is a neurodevelopmental disorder characterized by challenges in communication and behavior. This study examines the relat

APE: Agentic Prompt Enhancer for Image Generation and Editing

Model ReleasesDGX agent

arXiv:2606.00204v1 Announce Type: new Abstract: Natural language has become a powerful interface for image generation and editing, yet text-guided visual systems remain highly sensitive to prompt form

ASAP: Advancing Medical Volumetric Representation Learning with Anatomy-aware Semantically-adaptive Pre-training

Model ReleasesDGX agent

arXiv:2606.00602v1 Announce Type: new Abstract: Learning transferable and interpretable representations from medical volumetric scans remains challenging due to complex anatomical structures and weak,

Ask4VG: Risk-Aware Question Selection for Reducing Prior-Driven Answers in Medical VQA

TutorialsDGX agent

arXiv:2606.01044v1 Announce Type: new Abstract: Medical visual question answering requires models to ground their responses in image evidence, because visually unsupported answers can mislead downstre

Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning

ResearchDGX agent

arXiv:2606.01558v1 Announce Type: new Abstract: The effectiveness of Chain-of-Thought (CoT) prompting in Multimodal Large Language Models (MLLMs) remains uncertain: across several visual reasoning ben

Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation

ApplicationsDGX agent

arXiv:2606.01900v1 Announce Type: new Abstract: Generative video models have achieved remarkable visual fidelity and temporal coherence, yet intentional camera control remains elusive. Existing framew

AutoFFS: Adversarial Deformations for Facial Feminization Surgery Planning

SafetyDGX agent

arXiv:2603.02288v2 Announce Type: replace Abstract: Facial feminization surgery (FFS) is a key component of gender affirmation for transgender and gender diverse patients, aiming to reshape craniofaci

AutoIQ: An Ensemble Framework for Automatic Assessment of Geometric Distortion in Prostate Diffusion-Weighted Imaging

Local AiDGX agent

arXiv:2606.00393v1 Announce Type: cross Abstract: Geometric distortion in prostate diffusion-weighted imaging (DWI) can impair lesion localization and reduce the reliability of MRI-based clinical asse

Automated Erythrocyte Detection and Tracking for Retinal Blood Flow Quantification in Erythrocyte-Mediated Angiography

ResearchDGX agent

arXiv:2606.01006v1 Announce Type: new Abstract: Capillary-level retinal blood flow (RBF) has strong potential as a biomarker for various ocular diseases. However, modalities for measuring capillary-le

Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanations

ResearchDGX agent

arXiv:2511.20295v2 Announce Type: replace Abstract: Counterfactual explanations (CFEs) are minimal and semantically meaningful modifications of the input of a model that alter the model predictions. T

Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation

SafetyDGX agent

arXiv:2605.25195v2 Announce Type: replace Abstract: Current open-source diffusion models struggle to generate stable and synchronized audio-visual content, particularly in scenarios demanding complex

Bayesian meta-learning for modeling Alzheimer's disease progression

ApplicationsDGX agent

arXiv:2606.02228v1 Announce Type: cross Abstract: Predicting whether an individual with Alzheimer's disease will experience mild or severe disease progression is essential for personalized treatment.

Belief Consistency Between Foundation-Model Evidence and Geometric Perception in Persistent Robotic Maps

AgentsDGX agent

arXiv:2606.00318v1 Announce Type: cross Abstract: Persistent maps used by autonomous robots increasingly fuse a geometric perception stack whose assertions are well-characterized with a foundation-mod

Beyond Low-Rank: Low-Rank Sparse Prompting via Spiking Neural Network and Prompt Factorization

ResearchDGX agent

arXiv:2606.01945v1 Announce Type: new Abstract: Visual Prompting (VP) has emerged as an efficient paradigm for adapting large-scale pre-trained vision models to downstream tasks by incorporating learn

Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification

ResearchDGX agent

arXiv:2510.24078v2 Announce Type: replace Abstract: Text-to-image (T2I) models are increasingly used for synthetic dataset generation, but generating effective synthetic training data for classificati

Beyond Rigid: Benchmarking Non-Rigid Video Editing

Model ReleasesDGX agent

arXiv:2601.18340v2 Announce Type: replace Abstract: As video generation models are increasingly expected to manipulate physical dynamics, there is a growing need to move evaluation beyond appearance f

Beyond Static Gaussians: An Empirical Investigation of Architectural Paradigms for Dynamic 3D Scene Reconstruction

Model ReleasesDGX agent

arXiv:2606.00452v1 Announce Type: new Abstract: Dynamic scene reconstruction via 3D Gaussian Splatting (3DGS) has emerged as a compelling approach for representing evolving environments, yet understan

Beyond the Simplex: Balanced Prototype Geometry for Scorer-Agnostic Open-Set Recognition

Model ReleasesDGX agent

arXiv:2606.01883v1 Announce Type: cross Abstract: Open-set recognition (OSR) requires a classifier to reject inputs from unseen classes which is essential in safety-critical settings such as medical i

Boundary-Protection W8A8 HiFloat8 Quantization for Large-Scale Text-to-Video Diffusion Transformers

Model ReleasesDGX agent

arXiv:2606.00957v1 Announce Type: new Abstract: We present a post-training quantization (PTQ) approach for Wan2.1-T2V-14B, a 14-billion-parameter text-to-video diffusion transformer, targeting the W8A

Braille to Text Translation for Bengali Language: A Geometric Approach

ResearchDGX agent

arXiv:2012.01494v2 Announce Type: replace Abstract: Braille is the only system to visually impaired people for reading and writing. However, general people cannot read Braille. So, teachers and relati

Bridging Topology and Deep Representation Learning: A TDA-ViT Fusion Model for Four-Class Brain Tumor Classification

ResearchDGX agent

arXiv:2606.00927v1 Announce Type: new Abstract: Accurate brain tumor classification from magnetic resonance imaging (MRI) is a key requirement for early diagnosis and clinical decision-making. Vision

C-LEAD: Contrastive Learning for Enhanced Adversarial Defense

TutorialsDGX agent

arXiv:2510.27249v2 Announce Type: replace Abstract: Deep neural networks (DNNs) have achieved remarkable success in computer vision tasks such as image classification, segmentation, and object detecti

CanonCGT: Reference-Based Color Grading via Canonical Pivot Representation

SafetyDGX agent

arXiv:2606.01638v1 Announce Type: new Abstract: Reference-based color grading aims to reproduce the tonal mood and lighting of a reference while preserving color harmony and scene structure. Existing

CASTLE2026 Team WDL Technical Report

Model ReleasesDGX agent

arXiv:2606.00712v1 Announce Type: new Abstract: The CASTLE Challenge @ EgoVis 2026 evaluates long-form egocentric video question answering over 600+ hours of multi-perspective recordings. Each four-ch

Chameleon: Style-Content Disentangled Framework for Cross-Domain Object Compositing

Model ReleasesDGX agent

arXiv:2606.01079v1 Announce Type: new Abstract: Image compositing aims to seamlessly insert a foreground object into a background image, and recent advances in diffusion models have significantly enha

ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats

Model ReleasesDGX agent

arXiv:2606.01348v1 Announce Type: new Abstract: Charts are a primary medium for conveying quantitative and relational information, yet systematically evaluating chart parsing models remains difficult.

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation

ResearchDGX agent

arXiv:2602.06442v2 Announce Type: replace Abstract: Unified multimodal models (UMMs) have achieved remarkable progress yet remain constrained by a single-turn interaction paradigm, effectively functio

Chroma Clues: Leveraging Color Statistics to Detect Synthetic Images

TutorialsDGX agent

arXiv:2606.02224v1 Announce Type: new Abstract: The evolution and dissemination of AI-synthesized images is occurring at an unprecedented rate. Image generators are making rapid progress in their goal

ChWDTA: Channel-wise Wavelet-Domain Transformer Attention and Entropy Modeling for Learned Image Compression

ResearchDGX agent

arXiv:2606.00111v1 Announce Type: cross Abstract: State-of-the-art learned image compression (LIC) schemes are increasingly based on hybrid CNN-transformer architectures. To further improve rate-disto

CLIP-like Model as a Foundational Density Ratio Estimator

TutorialsDGX agent

arXiv:2506.22881v3 Announce Type: replace Abstract: Density ratio estimation is a core concept in statistical machine learning because it provides a unified mechanism for tasks such as importance weig

CloSE: A Geometric Shape-Agnostic Cloth State Representation

ResearchDGX agent

arXiv:2504.05033v3 Announce Type: replace-cross Abstract: Cloth manipulation is a difficult problem mainly because of the non-rigid nature of cloth, which makes a good representation of deformation es

Closing the Alignment-Maturity Gap in Federated Prototype Learning

SafetyDGX agent

arXiv:2606.02172v1 Announce Type: cross Abstract: Learning discriminative visual representations from distributed, heterogeneous data is a fundamental challenge in Federated Learning (FL). Prototype-b

Cohort-Scale Neural Atlases of Ultrasound Video

HardwareDGX agent

arXiv:2606.00890v1 Announce Type: new Abstract: Ultrasound is the most widely used real-time imaging modality in clinical practice, yet per-frame video annotation remains a major bottleneck: expert la

← Previous
1…99100101102103…211
Next →