AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
3 Jun 2026

Investigating Adversarial Robustness of Multi-modal Large Language Models

Model ReleasesDGX agent

arXiv:2606.03713v1 Announce Type: new Abstract: Multi-modal Large Language Models (MLLMs) achieve strong performance on vision-language tasks, but incorporating visual inputs through a vision encoder

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation

Model ReleasesDGX agent

arXiv:2606.03168v1 Announce Type: new Abstract: While instruction-based video editing has seen significant progress, joint audio-visual editing remains constrained by the absence of dedicated datasets

KC-3DGS: Kurtosis-Constrained Gaussian Splatting for High-Fidelity View Synthesis

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.03120v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) enables real-time novel view synthesis by representing scenes as collections of anisotropic Gaussians optimized via differe

Knowledge-Preserved Model Tuning in Null-Space for Robust Spatio-Temporal Video Grounding

Model ReleasesDGX agent

arXiv:2606.03539v1 Announce Type: new Abstract: Spatio-Temporal Video Grounding aims to localize object tubes based on textual queries. While recent methods have achieved remarkable success, they main

LAMP: Data-Efficient Linear Affine Weight-Space Models for Parameter-Controlled 3D Shape Generation and Extrapolation

Model ReleasesDGX agent

arXiv:2510.22491v3 Announce Type: replace-cross Abstract: Generating high-fidelity 3D geometries under explicit parameter constraints is central to engineering design, yet current methods often requir

Learning to See via Epiretinal Implant Stimulation in silico with Model-Based Deep Reinforcement Learning

AgentsDGX agent

arXiv:2606.03118v1 Announce Type: cross Abstract: Objective: Diseases such as age-related macular degeneration and retinitis pigmentosa cause the degradation of the photoreceptor layer. One approach t

LoCAtion: Long-time Collaborative Attention Framework for High Dynamic Range Video Reconstruction

SafetyDGX agent

arXiv:2603.14377v2 Announce Type: replace Abstract: Prevailing High Dynamic Range (HDR) video reconstruction methods are fundamentally trapped in a fragile alignment-and-fusion paradigm. While explici

Low-Frequency Shortcuts in Texture-Driven Visual Learning

Model ReleasesDGX agent

arXiv:2606.03493v1 Announce Type: new Abstract: Neural networks suffer from shortcut learning, where learned features generalize well to the training set but not to in-distribution (ID) or out-of-dist

Low-Resolution Editing is All You Need for High-Resolution Editing

ResearchDGX agent

arXiv:2511.19945v3 Announce Type: replace Abstract: High-resolution content creation is rapidly emerging as a central challenge in both the vision and graphics communities. Images serve as the most fu

MAdam: Metric-Aware Multi-Objective Adam

ResearchDGX agent

arXiv:2606.03904v1 Announce Type: cross Abstract: Multi-objective optimization (MOO) underlies many machine learning problems, yet MOO solvers across the loss-balancing, gradient-balancing, and Pareto

Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation

ResearchDGX agent

arXiv:2606.03402v1 Announce Type: new Abstract: Audio-driven human motion video generation aims to synthesize realistic and temporally coherent human animations from a single static image, with applic

MariData: One-Step Unpaired Image Translation for Maritime Environments

AgentsDGX agent

arXiv:2606.03246v1 Announce Type: new Abstract: The development on robust perception systems for Maritime Autonomous Surface Ships (MASS) is heavily constrained by the scarcity of diverse training dat

MARIO: Motion-Augmented Real-Time Multi-Sensor Inertial Odometry

Model ReleasesDGX agent

arXiv:2606.02996v1 Announce Type: cross Abstract: Inertial odometry (IO) using only Inertial Measurement Units (IMUs) provides a lightweight solution for human motion tracking in augmented reality (AR

MemoGen: Can Past Experience Improve Future Text-to-Image Generation?

Model ReleasesDGX agent

arXiv:2606.03243v1 Announce Type: new Abstract: Modern text-to-image models have achieved strong visual synthesis, yet remain unreliable when prompts require implicit visual constraints, relational re

Mixed-Modality Dual Face-Hair Retrieval

Model ReleasesDGX agent

arXiv:2606.03470v1 Announce Type: new Abstract: We introduce Dual Face-Hair Retrieval (DFHR), a new mixed-modality dual-reference task in image retrieval where a query consists of a face image specify

MLP Splatting: Object-Centric Neural Fields

ResearchDGX agent

arXiv:2606.03877v1 Announce Type: new Abstract: 3D representations are fundamental to scene rendering, understanding, and interaction. Recent approaches, such as 3D Gaussian Splatting and Neural Radia

Neural Fields as World Models

Local AiDGX agent

arXiv:2602.18690v2 Announce Type: replace-cross Abstract: Humans rehearse possible futures offline, as in mental practice and perhaps dreaming, suggesting that world models may support task learning a

NewtPhys: Do Foundation Models Understand Newtonian Physics?

ApplicationsDGX agent

arXiv:2606.03986v1 Announce Type: new Abstract: Previous work has evaluated physics reasoning in foundation models using synthetic or semi-synthetic scenes and visual question-answering tasks. However

Optimizing Neuro-Fuzzy and Colonial Competition Algorithms for Skin Cancer Diagnosis in Dermatoscopic Images

ResearchDGX agent

arXiv:2505.08886v2 Announce Type: replace Abstract: The rising incidence of skin cancer, coupled with limited public awareness and a shortfall in clinical expertise, underscores an urgent need for adv

OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance

ApplicationsDGX agent

arXiv:2603.18639v3 Announce Type: replace Abstract: Recent progress in video generation has led to substantial improvements in visual fidelity, yet ensuring physically consistent motion remains a fund

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs

Model ReleasesDGX agent

arXiv:2606.03890v1 Announce Type: new Abstract: Multimodal agents in robotics, AR, and autonomous driving must reason about places and layouts from continuous egocentric streams, often using evidence

PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training

ResearchDGX agent

arXiv:2606.03264v1 Announce Type: new Abstract: We introduce PaddleOCR-VL-1.6, an upgraded compact document parsing model built upon PaddleOCR-VL-1.5. Although PaddleOCR-VL-1.5 establishes a strong 0.

PatchScene: Patch-based Voxel Diffusion for Large-Scale Scene Completion

Model ReleasesDGX agent

arXiv:2606.03915v1 Announce Type: new Abstract: We propose PatchScene, a novel diffusion-based framework for large-scale LiDAR scene completion. Unlike existing methods that rely on global latent repr

Pathway-Structured Privileged Distillation for Deployable Computational Pathology

SafetyDGX agent

arXiv:2606.02877v1 Announce Type: new Abstract: Integrating transcriptomics and histopathology can improve cancer risk modelling, yet practical use is constrained by the limited availability of RNA pr

PersistGS: Differentiable Physics for Object Permanence in 4D Gaussian Splatting

ResearchDGX agent

arXiv:2606.03479v1 Announce Type: new Abstract: Dynamic 3D Gaussian Splatting (3DGS) methods reconstruct time-varying scenes from synchronized multi-camera video using photometric supervision. When a

PHAF-Personalized Hand Avatars in a Flash

SafetyDGX agent

arXiv:2606.03420v1 Announce Type: new Abstract: We present PHAF-Personalized Hand Avatars in a Flash, a personalized photo-realistic hand avatar which provides high quality multi-view renders from jus

Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance

SafetyDGX agent

arXiv:2511.10055v2 Announce Type: replace Abstract: The performance of image generation has been significantly improved in recent years. However, the study of image screening is rare, and its performa

Pixel Cube: Diffusion-based Portrait Video Relighting Through Realistic Lighting Reproduction

ResearchDGX agent

arXiv:2606.02919v1 Announce Type: new Abstract: We present a diffusion-based method for relighting dynamic portrait videos with photorealism and temporal consistency. Our method is fueled by a hybrid

PixVOD: Pixel-Distributed Direct Visual Odometry and Depth Estimation

ResearchDGX agent

arXiv:2606.03989v1 Announce Type: new Abstract: Images composed of 2D pixel arrays are the standard input to computer vision algorithms, yet many underlying computations can be distributed across pixe

Principled Reflection Separation via Nonlinear Superposition and Feature Interaction

ApplicationsDGX agent

arXiv:2606.02831v1 Announce Type: new Abstract: Single-image reflection separation is fundamentally challenged by the entanglement of transmission and reflection layers under complex image formation p

PRISM: Rethinking Atmospheric Scattering Reconstruction as a Unified Understanding and Restoration Model for Real-world Dehazing

TutorialsDGX agent

arXiv:2604.07048v2 Announce Type: replace Abstract: Real-world image dehazing (RID) aims to remove haze-induced degradation from real scenes. This task remains challenging due to non-uniform haze dist

PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction

Model ReleasesDGX agent

arXiv:2512.10888v3 Announce Type: replace Abstract: Table extraction (TE) is a key challenge in document understanding. Traditional approaches detect tables first, then recognize their structure. Rece

Reconstructing Objects along Hand Interaction Timelines in Egocentric Video

Model ReleasesDGX agent

arXiv:2512.07394v2 Announce Type: replace Abstract: We introduce the task of Reconstructing Objects along Hand Interaction Timelines (ROHIT). We first define the Hand Interaction Timeline (HIT) from a

ResCLIP: Residual Attention for Training-free Dense Vision-language Inference

Local AiDGX agent

arXiv:2411.15851v2 Announce Type: replace Abstract: While vision-language models like CLIP have shown remarkable success in open-vocabulary tasks, their application is currently confined to image-leve

SaluNet: Enabling Total Plasticity in Normalization-Free Deep Networks

TutorialsDGX agent

arXiv:2606.02927v1 Announce Type: new Abstract: Normalization layers such as BatchNorm and LayerNorm have long been considered essential for stable training in deep networks. This work demonstrates th

SAMatcher: Co-Visibility Modeling with Segment Anything for Robust Feature Matching

Local AiDGX agent

arXiv:2606.03406v1 Announce Type: new Abstract: Reliable correspondence estimation is a fundamental problem in image processing, underpinning applications such as Structure from Motion, visual localiz

SEAOTTER: Sensor Embedded Autoencoding with One-Time Transcode for Efficient Reconstruction

Local AiDGX agent

arXiv:2606.03940v1 Announce Type: cross Abstract: In robotics systems, vast amounts of visual data are easily captured at high resolution using low-cost, low-power hardware. Yet, limited bandwidth and

Seg2Track++: Probabilistic Track Validation and Data Association for Multi-Object Tracking and Segmentation

AgentsDGX agent

arXiv:2606.03875v1 Announce Type: new Abstract: Autonomous systems require robust Multi-Object Tracking and Segmentation (MOTS) to operate reliably in dynamic environments, ensuring consistent object

SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image

SafetyDGX agent

arXiv:2606.03994v1 Announce Type: new Abstract: Reconstructing interactive, simulation-ready 3D scenes from a single image is a critical bottleneck for robotic manipulation. While recent single-image

SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive Continuation

Local AiDGX agent

arXiv:2603.18599v2 Announce Type: replace Abstract: Speculative Jacobi Decoding (SJD) offers a draft-model-free approach to accelerate autoregressive text-to-image synthesis. However, the high-entropy

SkelHCC: A Hyperbolic CLIP-Driven Cache Adaptation Framework for Skeleton-based One-Shot Action Recognition

SafetyDGX agent

arXiv:2606.03610v1 Announce Type: new Abstract: Skeleton-based action recognition aims to understand human behaviors from body joint sequences and is especially challenging in the one-shot setting, wh

SLU-2K: A Question-Based Benchmark for Semantic Evaluation of Sign Language Translation

Model ReleasesDGX agent

arXiv:2606.03788v1 Announce Type: new Abstract: Sign Language Translation (SLT) is typically evaluated with surface-form metrics such as BLEU and ROUGE, which reward lexical overlap but do not directl

SparseStreet: Sparse Gaussian Splatting for Real-Time Street Scene Simulation

ResearchDGX agent

arXiv:2606.03909v1 Announce Type: new Abstract: While 3D Gaussian Splatting has shown promising results in street scene reconstruction, existing methods require massive numbers of Gaussian primitives

SRENet: Spectral Re-Entry Network for Point Cloud Action Recognition

AgentsDGX agent

arXiv:2606.03160v1 Announce Type: new Abstract: Recognizing human actions from point cloud sequences is critical for 3D perception driven applications such as autonomous driving and human-computer int

Structure-Guided Mixed Masked Pretraining and Spatial Continuity Regularization for Printed Circuit Board Defect Detection

ResearchDGX agent

arXiv:2606.03508v1 Announce Type: new Abstract: Printed circuit board (PCB) defect detection is an essential part of automated optical inspection (AOI); yet it remains challenging in practice because

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

Model ReleasesDGX agent

arXiv:2512.21094v2 Announce Type: replace Abstract: Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet it

TASE: Truncation-Aware Semantic Embeddings for 3D Scene Understanding and Editing

AgentsDGX agent

arXiv:2606.03314v1 Announce Type: new Abstract: High-fidelity semantic 3D scene representations are crucial for numerous applications, including robotics, autonomous driving, and simulation. Beyond th

Template Collapse and Information-Theoretic Limits in Camera rPPG Pulse Morphology Restoration

ResearchDGX agent

arXiv:2606.03802v1 Announce Type: new Abstract: Objective: Consumer face camera remote photoplethysmography (rPPG) enables passive cardiovascular monitoring, but whether single-cycle waveform morpholo

TeX-1500: A Paired Real-World LWIR Hyperspectral Dataset and Benchmark for Temperature-Emissivity-Texture Decomposition

Model ReleasesDGX agent

arXiv:2606.03806v1 Announce Type: new Abstract: Temperature-emissivity-texture (TeX) decomposition seeks to recover object heat state, material spectral response, and visible-like geometric texture fr

Text-to-Image Models Need Less from Text Encoders Than You Think

TutorialsDGX agent

arXiv:2606.03715v1 Announce Type: new Abstract: Text-to-image models rely on text prompts as their primary interface to human intent. Prompts are encoded by a text encoder into embeddings that conditi

TGV-KV: Text-Grounded KV Eviction for Vision-Language Models

SafetyDGX agent

arXiv:2606.03075v1 Announce Type: new Abstract: Vision-Language Models (VLMs) inherit the auto-regressive generation paradigm and cache the keys and values (KV) of all previous tokens to accelerate in

The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset

Local AiDGX agent

arXiv:2606.02956v1 Announce Type: new Abstract: Existing autonomous driving datasets have enabled major progress, but fall short in sensor fidelity, map completeness, or geographic diversity. We prese

Tiny Collaborative Inference for Occlusion-Robust Object Detection

Local AiDGX agent

arXiv:2606.02894v1 Announce Type: new Abstract: Small edge devices such as IoT surveillance nodes and search-and-rescue (SAR) platforms are increasingly expected to run computer vision locally. On ult

Towards Blind Lens Aberration Correction via Large LensLib Pre-training and Discrete Degradation Priors

ApplicationsDGX agent

arXiv:2511.17126v4 Announce Type: replace-cross Abstract: Emerging deep-learning-based lens library pre-training (LensLib-PT) pipeline offers a new avenue for blind lens aberration correction by train

Towards Characterizing Scientific Image Utility and Upgradability

Model ReleasesDGX agent

arXiv:2606.03401v1 Announce Type: new Abstract: Scientific images function as critical evidence in research communication, yet their integrity faces unprecedented threats from AI-generated content tha

TrAction: Action Recognition with Sparse Trajectories

ResearchDGX agent

arXiv:2606.03490v1 Announce Type: new Abstract: Modern action recognition models operate on memory- and compute-intensive dense RGB video volumes and frequently exploit appearance and background short

Training-Free Multi-Concept LoRA Composition with Prompt-Aware Weighting

ApplicationsDGX agent

arXiv:2606.03792v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) successfully enables personalization in text-to-image generation by adapting pre-trained diffusion models to specific visual

Transformer-Guided Content-Adaptive Graph Learning for Hyperspectral Unmixing

TutorialsDGX agent

arXiv:2509.03376v2 Announce Type: replace Abstract: Hyperspectral unmixing (HU) targets to decompose each mixed pixel in remote sensing images into a set of endmembers and their corresponding abundanc

Unified Video-Action Joint Denoising for Dexterous Action and Data Generation

SafetyDGX agent

arXiv:2606.03868v1 Announce Type: new Abstract: Recent world action models leverage video foundation models by aligning broad visual-dynamics priors with executable robot actions. We revisit this alig

UnsOcc: 3D Semantic Occupancy Prediction in Unstructured Scene via Rendering Fusion

SafetyDGX agent

arXiv:2606.03581v1 Announce Type: new Abstract: Unstructured scenes present unique challenges for autonomous driving, as irregular obstacles and sparse scene layouts undermine the effectiveness of tra

← Previous
1…9899100101102…211
Next →