AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
3 Jun 2026

AmbientEye: A Dataset for Pupil Segmentation under Natural Ambient Infrared Illumination

Model ReleasesDGX agent

arXiv:2606.03774v1 Announce Type: new Abstract: Eye tracking is essential for smart glasses, as it provides insight into user attention for ambient intelligence applications. However, most existing ey

An Attention-Based Denoising Model for Diffusion Weighted Imaging

ResearchDGX agent

arXiv:2606.03903v1 Announce Type: new Abstract: Diffusion-weighted imaging (DWI) is used for whole-body cancer screening, but it typically requires a long acquisition time. When the scan time is reduc

An Improved Method for Personalizing Diffusion Models

ResearchDGX agent

arXiv:2407.05312v2 Announce Type: replace Abstract: Diffusion models have demonstrated impressive image generation capabilities. Personalized approaches, such as textual inversion and Dreambooth, enha


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Any2Poster: Any-Source Poster Generation Across Modalities and Domains

Model ReleasesDGX agent

arXiv:2606.02915v1 Announce Type: new Abstract: Visual posters are a compact medium for communicating dense information, yet progress on automatic poster generation remains difficult to measure becaus

Ask When It Pays: Cost-Aware Open-Ended Interaction for Instance Goal Navigation

Model ReleasesDGX agent

arXiv:2606.03175v1 Announce Type: new Abstract: Instance Goal Navigation (IGN) requires an embodied agent to find a specific object instance among distractors from an underspecified natural-language d

ATLAS: A Large-Scale Evaluation Benchmark for Adversarial LiDAR Perception

Model ReleasesDGX agent

arXiv:2606.02924v1 Announce Type: new Abstract: Autonomous driving perception is typically evaluated on clean benchmark data, yet real-world deployment requires robustness to rare, structured, and pot

Attend to Anything: Foundation Model for Unified Human Attention Modeling

ApplicationsDGX agent

arXiv:2606.03540v1 Announce Type: new Abstract: Existing human attention (saliency) modeling methods persist as highly fragmented across modalities, scenes, and task formulations. Consequently, even w

Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models

Local AiDGX agent

arXiv:2604.06052v2 Announce Type: replace Abstract: Text-to-image diffusion models exhibit remarkable generative capabilities, yet their internal operations remain opaque, particularly when handling p

Automated Report-Derived Oncology VQA Benchmark for Evaluating Vision-Language Models on 3D Medical Imaging

Model ReleasesDGX agent

arXiv:2606.02809v1 Announce Type: new Abstract: Evaluating vision-language models (VLMs) on medical images requires benchmarks that are clinically grounded, scalable, and controlled for evaluation con

AvatarMix: Identity-Preserving Cross-Avatar Composition for Outfit Personalization

ResearchDGX agent

arXiv:2606.03506v1 Announce Type: new Abstract: Existing 3D avatar outfit transfer methods face distinct challenges: approaches that lift 2D edits to 3D often suffer from outfit or identity quality de

BA-T: An Iterative Transformer for Two-View Bundle Adjustment

ResearchDGX agent

arXiv:2606.03287v1 Announce Type: new Abstract: Feed-forward models for 3D reconstruction have achieved strong performance using deep cross-view attention to exchange information across images. Howeve

BEAST3D: Animal behavioral analysis and neural encoding from multi-view video via Gaussian splatting

ResearchDGX agent

arXiv:2606.02937v1 Announce Type: cross Abstract: Multi-view video recordings are increasingly used to capture the 3D movements of animals in experimental settings, yet extracting rich 3D representati

Benchmarking Visual State Tracking in Multimodal Video Understanding

Model ReleasesDGX agent

arXiv:2606.03920v1 Announce Type: new Abstract: Understanding a video requires more than recognizing isolated moments, as humans continuously track entities, states, and events over time. This capacit

Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching

SafetyDGX agent

arXiv:2602.12221v2 Announce Type: replace Abstract: We propose UniDFlow, a unified discrete flow-matching framework for multimodal understanding, generation, and editing. It decouples understanding an

Beyond Compression: Quantifying Spectral Accessibility in Vision Representations

ResearchDGX agent

arXiv:2606.03795v1 Announce Type: new Abstract: Vision-language models map visual features into a shared embedding space through learned projection layers, yet it remains unclear how these transformat

Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models

ResearchDGX agent

arXiv:2606.03730v1 Announce Type: new Abstract: Vision-language models (VLMs) such as CLIP show strong zero-shot generalization but remain highly vulnerable to adversarial attacks. Adversarial trainin

Beyond Single Solution: Multi-Hypothesis Collaborative Deep Unfolding Network for Image Compressive Sensing

ResearchDGX agent

arXiv:2606.03666v1 Announce Type: new Abstract: Recent deep unfolding networks (DUNs) have advanced Compressive Sensing (CS) by effectively integrating iterative optimization with deep learning archit

Bootstrap Your Generator: Unpaired Visual Editing with Flow Matching

ResearchDGX agent

arXiv:2606.03911v1 Announce Type: new Abstract: Modern generative models possess a deep understanding of visual content, yet training them for image editing typically requires massive datasets of pair

BYORn: Bootstrap Your Own Responses to Defend Large Vision-Language Models Against Backdoor Attacks

ResearchDGX agent

arXiv:2606.02947v1 Announce Type: cross Abstract: Supervised fine-tuning is the predominant approach for adapting autoregressive vision-language models to downstream tasks. Recent work has shown that

CAD-to-CT Registration of Cylindrical Objects via Ellipse-Based Axis Estimation

ResearchDGX agent

arXiv:2606.02935v1 Announce Type: new Abstract: Accurate registration of CAD models to CT scans is essential for establishing ground truth geometry in volumetric imaging. Obtaining reliable object mas

Characterizing Detectability in 3DGS Poisoning: A Stage-wise Benchmark

Model ReleasesDGX agent

arXiv:2606.03499v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has rapidly emerged as a leading representation for real-time novel view synthesis, but recent work shows it is vulnerable

COD10K-C: Benchmarking Robustness of Camouflaged Object Detection Under Natural Image Corruptions

Model ReleasesDGX agent

arXiv:2606.02603v1 Announce Type: new Abstract: Camouflaged object detection has improved substantially, but most standard benchmarks evaluate models only on clean images. This is not realistic becaus

Consistent Yet Wrong: Evidence Insensitivity in Spatial Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.02742v1 Announce Type: new Abstract: Spatial reasoning is fundamental to robotics, autonomy, and embodied AI, yet modern vision-language models (VLMs) remain unreliable on metric distance q

CoralBay: A Self-Supervised CT Foundation Model

Model ReleasesDGX agent

arXiv:2606.03888v1 Announce Type: new Abstract: Self-supervised learning has enabled large-scale pre-training on 2D natural images, producing general-purpose visual representations that transfer effec

CREward: A Type-Specific Creativity Reward Model

Model ReleasesDGX agent

arXiv:2511.19995v2 Announce Type: replace Abstract: Creativity is a complex phenomenon. When it comes to representing and assessing creativity, treating it as a single undifferentiated quantity would

CropCraft: A Procedural World Generator for Robotic Simulation of Agricultural Tasks

ResearchDGX agent

arXiv:2511.02417v2 Announce Type: replace Abstract: The adoption of agroecological practices in modern agriculture requires robotic systems capable of operating in highly diverse and complex field env

Cross-Modality Feature Fusion Based on Structured State Space Duality for Multimodal Image Registration Network

Local AiDGX agent

arXiv:2606.03341v1 Announce Type: new Abstract: In multi-modal image registration, the primary challenge lies in shared structural information extraction. Compared to Transformers, Structured State Sp

Cryo-Bench: Benchmarking Foundation Models for Cryosphere Applications

Model ReleasesDGX agent

arXiv:2603.01576v3 Announce Type: replace Abstract: Geo-Foundation Models (GFMs) have been evaluated across diverse Earth observation task including multiple domains and have demonstrated strong poten

Demo2Tutorial: From Human Experience to Multimodal Software Tutorials

Model ReleasesDGX agent

arXiv:2606.03951v1 Announce Type: new Abstract: Human experience in digital environments offers a vast, underexplored resource of authentic, untrimmed interactions that contain rich procedural knowled

Depth from Dual Differential Defocus and Stereo Consensus

ResearchDGX agent

arXiv:2606.02906v1 Announce Type: cross Abstract: We introduce D^3S Consensus, a physics-based, closed-form algorithm that unifies depth-from-defocus (DfD) and stereo to achieve highly accurate depth

Diagnosis of Human Object Interaction Detectors for Real World Educational Applications

Model ReleasesDGX agent

arXiv:2606.02789v1 Announce Type: new Abstract: Human-object interaction (HOI) recognition is critical for automatically analyzing student behavior in complex educational environments. Although state-

Diffusing in the Right Space: A Systematic Study of Latent Diffusability

ResearchDGX agent

arXiv:2606.03578v1 Announce Type: new Abstract: Latent diffusion models leverage visual tokenizers to compress images into latent spaces for efficient generative modeling. However, better reconstructi

Disentangling Visual and Factual Correctness in LVLMs' Visualization Literacy

Model ReleasesDGX agent

arXiv:2606.03142v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) show strong visualization interpretation, yet it is unclear whether their responses reflect genuine reasoning over

DyaPlex: Full-Duplex Speech-Motion Model for Dyadic Interaction

AgentsDGX agent

arXiv:2606.03874v1 Announce Type: new Abstract: We present DyaPlex, a streaming, full-duplex speech-and-motion model designed for dyadic interaction. To capture the continuous and reciprocal nature of

Electromagnetic Navigation for Femoral Osteotomy Using High-Accuracy X-ray-to-CT Registration

ResearchDGX agent

arXiv:2606.03893v1 Announce Type: new Abstract: Accurate execution of preoperative plans in corrective femoral osteotomies remains challenging. Current techniques are limited by variable accuracy, inv

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching

Model ReleasesDGX agent

arXiv:2606.03577v1 Announce Type: new Abstract: Wide-baseline matching (WBM) requires integrating geometric understanding, viewpoint changes, fine-grained perception, and occlusion reasoning, making i

Enginuity: A Dataset and Benchmark for Vision-Language Understanding of Engineering Diagrams

Model ReleasesDGX agent

arXiv:2606.03410v1 Announce Type: new Abstract: Engineering diagrams pose a distinct challenge for vision-language models: unlike natural images or general documents, they encode information through d

Estimating Central, Peripheral, and Temporal Visual Contributions to Human Decision Making in Atari Games

ResearchDGX agent

arXiv:2604.04439v2 Announce Type: replace-cross Abstract: We study how different visual information sources contribute to human decision making in dynamic visual environments. Using Atari-HEAD, a larg

EvoMemNav: Efficient Self-Evolving Fine-Grained Memory for Zero-Shot Embodied Navigation

SafetyDGX agent

arXiv:2606.03509v1 Announce Type: new Abstract: Building memory is essential for long-horizon planning in zero-shot embodied navigation. Detector-centric scene graphs often compress observations into

Exploring Easy Boosts for Lidar Semantic Scene Completion

ResearchDGX agent

arXiv:2606.03992v1 Announce Type: new Abstract: This paper investigates 'free lunch' strategies to boost the performance of lidar semantic scene completion (SSC) without requiring complex architectura

Face versus Body Tracking for Human-Robot Interaction: An Egocentric Dataset

AgentsDGX agent

arXiv:2606.03694v1 Announce Type: cross Abstract: To enable meaningful human-robot interaction (HRI), a robot must continuously assess engagement by consistently tracking users over time. State-of-the

FAF-CD: Frequency-Aware Fusion for Change Detection under Imperfect Multimodal Remote Sensing

SafetyDGX agent

arXiv:2606.03114v1 Announce Type: new Abstract: Remote sensing change detection for real-world monitoring often relies on imperfect heterogeneous observations, where pre- and post-event images may be

FCUS-rPPG: A Fast-Converging Unsupervised Framework for Remote Photoplethysmography via Gradient Oscillation Suppression

TutorialsDGX agent

arXiv:2606.03050v1 Announce Type: new Abstract: Remote photoplethysmography (rPPG) enables non-contact extraction of blood volume pulse (BVP) signals using consumer-grade cameras. Recent unsupervised

Follow-Your-Preference++: Rethinking Preference Alignment for Image Inpainting

SafetyDGX agent

arXiv:2606.03216v1 Announce Type: new Abstract: We study preference alignment for image inpainting. Rather than proposing yet another method, we revisit the problem from first principles and reassess

FreeStreamGS: Online Feed-forward 3D Gaussian Splatting from Unposed Streaming Inputs

SafetyDGX agent

arXiv:2606.03254v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting (3DGS) allows efficient and high-fidelity novel view synthesis (NVS) from an offline recorded image sequence. However

From 3D Perception to Safety Reasoning: A Graph-Based Framework for Real-Time Underground Mine Monitoring

Local AiDGX agent

arXiv:2606.03460v1 Announce Type: new Abstract: Underground coal mining requires personnel and heavy equipment to operate within shared, confined, and poorly illuminated spaces where hazards such as e

From Local Training to Large-Scale Mapping: A Comparative Assessment of Machine Learning and Deep Learning for Transferable Satellite-Derived Bathymetry

Model ReleasesDGX agent

arXiv:2606.02764v1 Announce Type: new Abstract: Satellite-derived bathymetry (SDB) from multispectral imagery is cost-effective but scales poorly across regions, especially in optically complex coasta

From None to All: Self-Supervised 3D Reconstruction via Novel View Synthesis

ResearchDGX agent

arXiv:2603.27455v2 Announce Type: replace Abstract: In this paper, we introduce NAS3R, a self-supervised feed-forward framework that jointly learns explicit 3D geometry and camera parameters with no g

GARDEN: Gravity-Aligned Reconstruction of Disentangled ENvironments from RGB images

ResearchDGX agent

arXiv:2606.03921v1 Announce Type: new Abstract: Converting multi-view RGB observations into simulation-ready 3D environments remains challenging because current reconstruction pipelines produce monoli

GeoDrive-Bench: Benchmarking Region-Specific Multimodal Reasoning in Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.02774v1 Announce Type: new Abstract: Vision-language models (VLMs) for autonomous driving have shown promising performance, but their ability to handle region-specific traffic rules remains

Graph Mamba Survival Analysis Based on Topology-Aware ordering

ResearchDGX agent

arXiv:2606.02602v1 Announce Type: cross Abstract: In computational pathology, Whole Slide Images (WSIs) survival analysis is crucial for patient prognosis assessment, but it faces multiple technical c

Graph Regularized Non-negative Reduced Biquaternion Matrix Factorization for Color Image Recognition

Local AiDGX agent

arXiv:2606.03654v1 Announce Type: new Abstract: Non-negative reduced biquaternion matrix factorization (NRBMF) uses the product of reduced biquaternion (RB) matrices to incorporate the non-negativity

GS-ROR^2: Bidirectional-guided 3DGS and SDF for Reflective Object Relighting and Reconstruction

TutorialsDGX agent

arXiv:2406.18544v4 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has shown a powerful capability for novel view synthesis due to its detailed expressive ability and highly efficient re

Hierarchical Federated Learning with Dynamic Clustering and Adaptive Regularization for Robust Infrastructure Inspection

ApplicationsDGX agent

arXiv:2606.03084v1 Announce Type: new Abstract: The deployment of data-driven computer vision models for structural health monitoring (SHM) is heavily constrained by the data silo dilemma due to strin

How Much of a Model Do We Need? Redundancy and Slimmability in Remote Sensing Foundation Models

ResearchDGX agent

arXiv:2601.22841v2 Announce Type: replace Abstract: Large-scale foundation models (FMs) in remote sensing (RS) (denoted as RS FMs) are developed following paradigms established in computer vision (CV)

Hybrid Autoregressive-Diffusion Model for Real-Time Sign Language Production

ApplicationsDGX agent

arXiv:2507.09105v4 Announce Type: replace Abstract: Earlier Sign Language Production (SLP) models typically relied on autoregressive decoding, which naturally preserves temporal causality but suffers

IdEst: Assessing Self-Supervised Learning Representations via Intrinsic Dimension

ResearchDGX agent

arXiv:2606.03338v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has emerged as a powerful paradigm for learning meaningful representations from unlabeled data. However, the standard p

IDO: Incongruity-aware Distribution Optimization for Multimodal Fake News Detection

TutorialsDGX agent

arXiv:2606.03418v1 Announce Type: new Abstract: Multimodal fake news detection aims to identify the authenticity of news. Existing multimodal fake news detection methods mainly focus on cross-modal co

Inference-Time Scaling for Joint Audio-Video Generation

SafetyDGX agent

arXiv:2606.03183v1 Announce Type: cross Abstract: Joint audio-video generation aims to synthesize realistic audio-video pairs that are both semantically aligned with text prompts and precisely synchro

Inverting the Generation Process of Denoising Diffusion Implicit Models: Empirical Evaluation and a Novel Method

ResearchDGX agent

arXiv:2606.03111v1 Announce Type: new Abstract: This paper studies the problem of inverting the DDIM image generation process to recover latent variables, particularly the initial noise map, from a ge

← Previous
1…979899100101…211
Next →