AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
28 May 2026

VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking

SafetyDGX agent

arXiv:2605.28083v1 Announce Type: new Abstract: While Vision-Language-Action (VLA) models have emerged as powerful generalist policies, their severe vulnerability to adversarial patches significantly

WeatherCity: Urban Scene Reconstruction with Controllable Multi-Weather Transformation

Model ReleasesDGX agent

arXiv:2602.22096v2 Announce Type: replace Abstract: Editable high-fidelity 4D scenes are crucial for autonomous driving, as they can be applied to end-to-end training and closed-loop simulation. Howev

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.27589v1 Announce Type: new Abstract: Video generation models are increasingly used as world simulators for tasks like driving and robotic manipulation. What matters in these settings is not

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models

TutorialsDGX agent

arXiv:2605.28132v1 Announce Type: new Abstract: Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this,

XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the Edge

Model ReleasesDGX agent

arXiv:2506.22726v4 Announce Type: replace Abstract: Deep learning for human sensing on edge systems presents significant potential for smart applications. However, its training and development are hin

27 May 2026

3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language Navigation

AgentsDGX agent

arXiv:2605.26500v1 Announce Type: new Abstract: Vision-language navigation (VLN) requires an agent to traverse complex 3D environments based on natural language instructions, necessitating a thorough

A Dynamic Programming Framework for Discovering Count and Values of Multilevel Image Thresholding

ResearchDGX agent

arXiv:2605.27287v1 Announce Type: new Abstract: Multilevel Image thresholding is an important preprocessing algorithm in computer vision applications nowadays. Since most common thresholding methods t

A multifractal-based masked auto-encoder: an application to medical images

ResearchDGX agent

arXiv:2605.26287v1 Announce Type: new Abstract: Masked autoencoders (MAE) have shown great promise in medical image classification. However, the random masking strategy employed by traditional MAEs ma

A Unified Framework for Diffusion Model Unlearning with f-Divergence

ResearchDGX agent

arXiv:2509.21167v2 Announce Type: replace-cross Abstract: Most existing methods for concept unlearning in text-to-image diffusion models minimize a mean squared error (MSE) loss between the denoiser o

AD-H: Language-guided Autonomous Driving with Hierarchical Agents

AgentsDGX agent

arXiv:2406.03474v2 Announce Type: replace Abstract: Language-guided autonomous driving requires bridging a large abstraction gap between high-level natural-language instructions and low-level vehicle

Adaptation-Free Heterogeneous Collaborative Perception with Unseen Agent Configurations

AgentsDGX agent

arXiv:2605.26642v1 Announce Type: new Abstract: Collaborative perception improves 3D object detection by enabling agents to share complementary observations, but most existing methods assume fixed or

Advancing Metallic Surface Defect Detection via Anomaly-Guided Pretraining on a Large Industrial Dataset

ResearchDGX agent

arXiv:2509.18919v2 Announce Type: replace Abstract: The pretraining-finetuning paradigm is a crucial strategy in metallic surface defect detection for mitigating the challenges posed by data scarcity.

AI-T2I: Aggregating-and-Isolating Cross-Attention to Diffusion Models for Text-to-Image Synthesis

SafetyDGX agent

arXiv:2605.25763v2 Announce Type: replace Abstract: Text-to-image synthesis has made significant progress, benefiting from the strong generative capabilities of diffusion models. However, these models

Align & Invert: Solving Inverse Problems with Diffusion and Flow-based Models via Representation Alignment

SafetyDGX agent

arXiv:2511.16870v3 Announce Type: replace Abstract: Enforcing alignment between the internal representations of diffusion or flow-based generative models and those of pretrained self-supervised encode

An uncertainty-aware Bayesian framework for machine learning classification models: A case study in land cover classification

Model ReleasesDGX agent

arXiv:2503.21510v3 Announce Type: replace-cross Abstract: Ensuring that predictions of machine learning (ML) classification models are accompanied by uncertainty estimates is one of the main pillars o

AnySurf: Any Surface Generation with Directed Edge

ResearchDGX agent

arXiv:2605.26149v1 Announce Type: cross Abstract: Open surface components prevail in real industrial 3D content and support rendering, physical simulation and geometric editing. Garments serve as a ty

Attenuation-Resilient Alternating Optimization for Laparoscopic Liver Landmark Detection

ResearchDGX agent

arXiv:2605.26630v1 Announce Type: new Abstract: Liver surface landmark detection is a fundamental prerequisite for anatomical guidance in laparoscopic liver surgery. However, it remains unreliable in

Axial-Centric Cross-Plane Attention for 3D Medical Image Classification

Model ReleasesDGX agent

arXiv:2602.21636v2 Announce Type: replace Abstract: Abridged: Clinicians commonly interpret 3D medical images by examining multiple anatomical planes rather than relying on volumetric views. In clinic

BEAT: Rhythm-Elastic Alignment for Agentic Music-guided Movie Trailer Generation

Model ReleasesDGX agent

arXiv:2605.27067v1 Announce Type: new Abstract: Automatic movie trailer generation must select shots from a full-length film and synchronize them with background music. Existing methods either relegat

Benchmarking Convolutional, Transformer, Hybrid, and Vision Language Models for Multi Disease Retinal Screening

Model ReleasesDGX agent

arXiv:2605.26283v1 Announce Type: new Abstract: Modern deep learning offers powerful tools for automated retinal screening, but it remains unclear how different visual model families compare in realis

Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models

SafetyDGX agent

arXiv:2605.26491v1 Announce Type: cross Abstract: Preference optimization has emerged as an efficient alternative to online reinforcement learning from human feedback (RLHF) for aligning text-to-image

Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models

AgentsDGX agent

arXiv:2605.27243v1 Announce Type: new Abstract: Large vision-language models increasingly rely on long-context modeling to reason over documents, hour-level videos, and long-horizon agent trajectories

Cesarean Scar Defect Segmentation in Transvaginal Ultrasound Images: a Dataset and Benchmark

Model ReleasesDGX agent

arXiv:2605.26774v1 Announce Type: new Abstract: Cesarean Scar Defect (CSD) is one of the most prevalent complications following cesarean delivery. Transvaginal ultrasonography is widely used for prima

Chaos-SSL: An Attention-Based Self-Supervised Learning Framework with Chaotic Transformation for Medical Image Classification

TutorialsDGX agent

arXiv:2605.27146v1 Announce Type: new Abstract: Self-Supervised Learning (SSL) has emerged as a powerful paradigm to mitigate the reliance on large, annotated datasets, a common bottleneck in medical

ChartAct: A Benchmark for Dynamic Chart Understanding

Model ReleasesDGX agent

arXiv:2605.26994v1 Announce Type: new Abstract: Charts are widely used to present complex data for analysis and decision making. Existing chart understanding benchmarks mainly focus on static charts,

CIRCLED: A Multi-turn CIR Dataset with Consistent Dialogues across Domains

Model ReleasesDGX agent

arXiv:2605.26734v1 Announce Type: new Abstract: Existing Multi-Turn Composed Image Retrieval (MTCIR) datasets lack dialogue-history consistency and are restricted to the fashion domain. To address the

Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis

Model ReleasesDGX agent

arXiv:2605.26483v1 Announce Type: new Abstract: Medical video diagnosis involves inferring clinical decisions from dynamic tissue responses throughout examination processes. Existing methods rely on a

CNNs, Transformers, Hybrid, and Vision Language Models for Skin Cancer Detection

Model ReleasesDGX agent

arXiv:2605.26294v1 Announce Type: new Abstract: Skin cancer is a common and fast rising malignancy worldwide. Early detection is critical for improving outcomes. Deep learning models trained on dermos

CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning

Model ReleasesDGX agent

arXiv:2605.26967v1 Announce Type: new Abstract: Existing video captioning methods struggle to balance visual fidelity and redundancy: holistic captions are compact but lose fine-grained evidence, wher

ControlLight: Towards Controllable, Consistent, and Generalizable Low-Light Enhancement

ApplicationsDGX agent

arXiv:2605.25569v2 Announce Type: replace Abstract: Existing deep learning-based low-light enhancement methods are typically trained on limited datasets with single enhancement targets, which restrict

COVD: Continual Open-Vocabulary Object Detection with Novel Concept Injection

Model ReleasesDGX agent

arXiv:2605.27116v1 Announce Type: new Abstract: Open-vocabulary object detection (OVD) has made significant progress, enabling detectors to generalize from seen to unseen categories. However, real-wor

CRoFT: Robust Fine-Tuning with Concurrent Optimization for OOD Generalization and Open-Set OOD Detection

TutorialsDGX agent

arXiv:2405.16417v2 Announce Type: replace Abstract: Recent vision-language pre-trained models (VL-PTMs) have shown remarkable success in open-vocabulary tasks. However, downstream use cases often invo

Datasets for Lane Detection in Autonomous Driving: A Comprehensive Review

AgentsDGX agent

arXiv:2504.08540v2 Announce Type: replace Abstract: Accurate lane detection is essential for automated driving, enabling safe and reliable vehicle navigation across a variety of road scenarios. Numero

DeepInterestGR: Mining Deep Multi-Interest Using Multi-Modal LLMs for Generative Recommendation

ResearchDGX agent

arXiv:2602.18907v2 Announce Type: replace-cross Abstract: We introduce DeepInterestGR, a novel framework that integrates deep interest mining into the generative recommendation pipeline. This addresse

DelowlightSplat: Feed-Forward Gaussian Splatting for Lowlight 3D Scene Reconstruction

Model ReleasesDGX agent

arXiv:2605.26629v1 Announce Type: new Abstract: Novel-view synthesis and 3D reconstruction from sparse posed images are central to robotics and AR/VR. Yet, feed-forward 3D Gaussian reconstruction fail

Design First, Code Later: Aesthetically Pleasing Template-Free Slides Generation

AgentsDGX agent

arXiv:2605.26451v1 Announce Type: cross Abstract: Producing presentation slides automatically entails coordinating narrative structure with page-level graphic design under strict spatial constraints.

Detail Consistent Stage-Wise Distillation for Efficient 3D MRI Segmentation

ResearchDGX agent

arXiv:2605.26382v1 Announce Type: new Abstract: Deploying high-performing 3D medical image segmenters (e.g., nnU-Net) is often limited by memory footprint and inference latency. Compression is therefo

Dimensional Distribution Emotion State: Leveraging Valence and Arousal as a Common Embedding Space for Visual Emotion Analysis

SafetyDGX agent

arXiv:2605.26262v1 Announce Type: new Abstract: Museums are important sites for the dissemination of culture and art. They are institutions rooted in history and tradition; their exhibitions are often

DinoComplete: 3D Shape Completion with Distilled Semantic Priors and State Space Models

ApplicationsDGX agent

arXiv:2605.26949v1 Announce Type: new Abstract: 3D shape completion from partial scans remains challenging for unseen categories and noisy real-world observations, where geometry alone is often insuff

DirectFisheye-GS: Enabling Native Fisheye Input in Gaussian Splatting with Cross-View Joint Optimization

ResearchDGX agent

arXiv:2604.00648v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has enabled efficient 3D scene reconstruction from everyday images with real-time, high-fidelity rendering, greatly adv

Do Modern Post-Hoc Watermarking Methods Beat Broken-Arrows?

ApplicationsDGX agent

arXiv:2605.27135v1 Announce Type: cross Abstract: With the rapid proliferation of generative models, such as diffusion models, digital watermarking has emerged as a crucial solution for identifying AI

Dual-Thresholded Heatmap-Guided Proposal Clustering and Negative Certainty Supervision with Enhanced Base Network for Weakly Supervised Object Detection

ResearchDGX agent

arXiv:2509.08289v3 Announce Type: replace Abstract: Weakly supervised object detection (WSOD) has attracted significant attention in recent years, as it does not require box-level annotations. State-o

DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture Generation

SafetyDGX agent

arXiv:2605.26236v1 Announce Type: new Abstract: Co-speech gesture generation requires both semantic expressivity and biomechanically plausible rhythmic motion. Existing holistic gesture models mix lex

DV-SFT: Direct Vision Supervision for Fine-Grained Visual Understanding

SafetyDGX agent

arXiv:2605.26656v1 Announce Type: new Abstract: Multimodal large language models are typically trained end-to-end to predict ground-truth answers, yet supervision signals are applied exclusively to te

Efficient All-Pairs Correlation Volume Sampling for Optical Flow Estimation

Local AiDGX agent

arXiv:2505.16942v2 Announce Type: replace Abstract: Recent optical flow estimation methods often employ local cost sampling from a dense all-pairs correlation volume. This results in quadratic computa

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy

Model ReleasesDGX agent

arXiv:2605.24456v2 Announce Type: replace Abstract: Humans constantly reason about 3D proximity, the relations between their body and surrounding objects, to guide perception and action in daily life.

Feedforward 3D Editing Learns from Semantic-Part Transformation

Local AiDGX agent

arXiv:2605.27351v1 Announce Type: new Abstract: 3D editing is a fundamental capability for scalable 3D content creation. While image editing has rapidly evolved toward large-scale feedforward generati

Frequency-Guided Fusion For RGB-Thermal Semantic Segmentation

ResearchDGX agent

arXiv:2605.26273v1 Announce Type: new Abstract: Semantic segmentation in complex environments such as urban driving scenes remains challenging under adverse lighting conditions, where RGB images alone

From Contrast to Consistency: Rethinking Event-based Continuous-Time Optical Flow Estimation

ResearchDGX agent

arXiv:2605.25570v2 Announce Type: replace Abstract: Estimating continuous optical flow is a fundamental yet challenging problem in dynamic visual perception. Event-based cameras, with microsecond late

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling

SafetyDGX agent

arXiv:2605.26601v1 Announce Type: new Abstract: Vision-language models have progressed rapidly, but Tibetan remains a severely underserved low-resource language due to the lack of reproducible trainin

G3T Up! Gravity Aligned Coordinate Frames Simplify Pointmap Processing

ApplicationsDGX agent

arXiv:2605.27372v1 Announce Type: new Abstract: Modern feed-forward 3D reconstruction methods like VGGT predict pixel-aligned pointmaps in camera-centric coordinate frames. However, this choice of coo

Garment Particles: A 2D--3D Symmetric Garment Representation for Generation and Editing

ResearchDGX agent

arXiv:2605.26391v1 Announce Type: cross Abstract: Practical garment design spans two modes: intuitive creation from high-level intent, such as a reference image or text description, and complex low-le

Gaussian-Voxel Duet: A Dual-Scaffolding Hybrid Representation for Fast and Accurate Monocular Surface Reconstruction

ApplicationsDGX agent

arXiv:2605.26616v1 Announce Type: new Abstract: While 3D Gaussian Splatting has achieved remarkable success in photorealistic novel view synthesis, its pursuit of fast and high-fidelity 3D reconstruct

Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini

Model ReleasesDGX agent

arXiv:2605.27295v1 Announce Type: new Abstract: We introduce Gemini Embedding 2, a native multimodal embedding model that allows embedding video, audio, image, and text modalities in a unified represe

Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction

Model ReleasesDGX agent

arXiv:2605.26230v1 Announce Type: new Abstract: Multi-view 3D reconstruction has achieved remarkable progress with the advent of feed-forward 3D reconstruction models. However, these models are typica

GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision

TutorialsDGX agent

arXiv:2603.09551v2 Announce Type: replace Abstract: While Vision-Language Models (VLMs) have significantly advanced remote sensing interpretation, enabling them to perform complex, step-by-step reason

Global Structure-from-Motion Meets Feedforward Reconstruction

ResearchDGX agent

arXiv:2605.26103v2 Announce Type: replace Abstract: Structure-from-Motion -- the process of simultaneously estimating camera poses and 3D scene structure from a collection of images -- remains a centr

GS-CLIP: Zero-shot 3D Anomaly Detection by Geometry-Aware Prompt and Synergistic View Representation Learning

ResearchDGX agent

arXiv:2602.19206v3 Announce Type: replace Abstract: Zero-shot 3D Anomaly Detection is an emerging task that aims to detect anomalies in a target dataset without any target training data, which is part

Guiding Token-Sparse Diffusion Models

Model ReleasesDGX agent

arXiv:2601.01608v2 Announce Type: replace Abstract: Diffusion models deliver high quality in image synthesis but remain expensive during training and inference. Recent works have leveraged the inheren

How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning

TutorialsDGX agent

arXiv:2605.27310v1 Announce Type: new Abstract: Cross-view spatial reasoning remains a weak spot for vision-language models (VLMs): they often reason in language and lose the fine-grained geometry nee

← Previous
1…112113114115116…211
Next →