AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
15 May 2026

Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds

ApplicationsDGX agent

arXiv:2602.00807v2 Announce Type: replace Abstract: Existing Vision-Language-Action (VLA) models typically take 2D images as visual input, which limits their spatial understanding in complex scenes. H

AnyBand-Diff: A Unified Remote Sensing Image Generation and Band Repair Framework with Spectral Priors

ResearchDGX agent

arXiv:2605.14341v1 Announce Type: new Abstract: Existing diffusion models have made significant progress in generating realistic images. However, their direct adaptation to remote sensing imagery ofte

ArcGate: Adaptive Arctangent Gated Activation

ResearchDGX agent

arXiv:2605.14518v1 Announce Type: new Abstract: Activation functions are central to deep networks, influencing non-linearity, feature learning, convergence, and robustness. This paper proposes the Ada


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Architecture-Aware Explanation Auditing for Industrial Visual Inspection

ResearchDGX agent

arXiv:2605.14255v1 Announce Type: cross Abstract: Industrial visual inspection systems increasingly rely on deep classifiers whose heatmap explanations may appear visually plausible while failing to i

Are Candidate Models Really Needed for Active Learning?

ResearchDGX agent

arXiv:2605.14689v1 Announce Type: new Abstract: Deep learning has profoundly impacted domains such as computer vision and natural language processing by uncovering complex patterns in vast datasets. H

Articraft: An Agentic System for Scalable Articulated 3D Asset Generation

AgentsDGX agent

arXiv:2605.15187v1 Announce Type: new Abstract: A bottleneck in learning to understand articulated 3D objects is the lack of large and diverse datasets. In this paper, we propose to leverage large lan

AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting

ResearchDGX agent

arXiv:2506.01015v2 Announce Type: replace Abstract: Segment Anything Model 2 (SAM2) exhibits strong generalisation for promptable segmentation in video clips; however, its integration with the audio m

Automatic Landmark-Based Segmentation of Human Subcortical Structures in MRI

Local AiDGX agent

arXiv:2605.14221v1 Announce Type: new Abstract: Precise segmentation of brain structures in magnetic resonance imaging (MRI) is essential for reliable neuroimaging analysis, yet voxel-wise deep models

AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2603.14851v3 Announce Type: replace Abstract: Integrating vision-language models (VLMs) into end-to-end (E2E) autonomous driving (AD) systems has shown promise in improving scene understanding.

Before the Body Moves: Learning Anticipatory Joint Intent for Language-Conditioned Humanoid Control

SafetyDGX agent

arXiv:2605.14417v1 Announce Type: cross Abstract: Natural language is an intuitive interface for humanoid robots, yet streaming whole-body control requires control representations that are executable

Beyond Instance-Level Self-Supervision in 3D Multi-Modal Medical Imaging

SafetyDGX agent

arXiv:2605.14654v1 Announce Type: new Abstract: Self-supervised pre-training methods in medical imaging typically treat each individual as an isolated instance, learning representations through augmen

BioHuman: Learning Biomechanical Human Representations from Video

Model ReleasesDGX agent

arXiv:2605.14772v1 Announce Type: new Abstract: Understanding human motion beyond surface kinematics is crucial for motion analysis, rehabilitation, and injury risk assessment. However, progress in th

Breaking Dual Bottlenecks: Evolving Unified Multimodal Models into Self-Adaptive Interleaved Visual Reasoners

Model ReleasesDGX agent

arXiv:2605.14709v1 Announce Type: new Abstract: Recent unified models integrate multimodal understanding and generation within a single framework. However, an 'understanding-generation gap' persists,

Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction

ResearchDGX agent

arXiv:2605.14569v1 Announce Type: new Abstract: Reconstructing dynamic visual experiences as videos from functional magnetic resonance imaging (fMRI) is pivotal for advancing the understanding of neur

CalibAnyView: Beyond Single-View Camera Calibration in the Wild

ApplicationsDGX agent

arXiv:2605.14615v1 Announce Type: new Abstract: Camera calibration is a fundamental prerequisite for reliable geometric perception, yet classical approaches rely on controlled acquisition setups that

Can Visual Mamba Improve AI-Generated Image Detection? An In-Depth Investigation

Model ReleasesDGX agent

arXiv:2605.14799v1 Announce Type: new Abstract: In recent years, computer vision has witnessed remarkable progress, fueled by the development of innovative architectures such as Convolutional Neural N

Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation

ResearchDGX agent

arXiv:2605.15141v1 Announce Type: new Abstract: Real-time interactive video generation requires low-latency, streaming, and controllable rollout. Existing autoregressive (AR) diffusion distillation me

CC-Pan: Channel-wise Compression based Diffusion for Efficient Pan-Sharpening

ResearchDGX agent

arXiv:2602.04473v2 Announce Type: replace Abstract: Recently, diffusion models have brought novel insights to pan-sharpening and notably boosted fusion precision. However, most existing models perform

Characterizing the visual representation of objects from the child's view

TutorialsDGX agent

arXiv:2605.14990v1 Announce Type: new Abstract: Children acquire object category representations from their everyday experiences in the first few years of life. What do the inputs to this learning pro

CHASM: Cross-frequency Harmonized Axis-Separable Mixing for Spectral Token Operators

SafetyDGX agent

arXiv:2605.14727v1 Announce Type: new Abstract: Spectral token mixers based on Fourier transforms provide an efficient way to model global interactions in visual feature maps. Existing designs often e

ClickRemoval: An Interactive Open-Source Tool for Object Removal in Diffusion Models

ResearchDGX agent

arXiv:2605.14461v1 Announce Type: new Abstract: Existing object removal tools often rely on manual masks or text prompts, making precise removal difficult for non-expert users in complex scenes and of

Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers

ResearchDGX agent

arXiv:2511.14751v2 Announce Type: replace Abstract: We propose Confidence-Guided Token Merging (Co-Me), an acceleration mechanism for visual geometric transformers without retraining or finetuning the

COAL: Counterfactual and Observation-Enhanced Alignment Learning for Discriminative Referring Multi-Object Tracking

SafetyDGX agent

arXiv:2605.14795v1 Announce Type: new Abstract: Referring Multi-Object Tracking (RMOT) faces a fundamental structural contradiction between the high-discriminability demand and the sparse semantic sup

CoCoEdit: Content-Consistent Image Editing via Region Regularized Reinforcement Learning

Model ReleasesDGX agent

arXiv:2602.14068v2 Announce Type: replace Abstract: Image editing has achieved impressive results with the development of large-scale generative models. However, existing models mainly focus on the ed

Compositional Video Generation via Inference-Time Guidance

ResearchDGX agent

arXiv:2605.14988v1 Announce Type: new Abstract: Text-to-video diffusion models generate realistic videos, but often fail on prompts requiring fine-grained compositional understanding, such as relation

Computational Imaging Priors for Wireless Capsule Endoscopy: Monte Carlo-Guided Hemoglobin Mapping for Rare-Anomaly Detection

ResearchDGX agent

arXiv:2605.15062v1 Announce Type: new Abstract: Background. RGB-trained capsule-endoscopy classifiers underperform on small-vessel vascular findings by conflating hemoglobin contrast with bile and ill

Contrastive Multi-Modal Hypergraph Reasoning for 3D Crowd Mesh Recovery

ApplicationsDGX agent

arXiv:2605.13854v1 Announce Type: new Abstract: Multi-person 3D reconstruction is pivotal for real-world interaction analysis, yet remains challenging due to severe occlusions and depth ambiguity. Cur

CoralLite: {mu}CT Reconstruction of Coral Colonies from Individual Corallites

ResearchDGX agent

arXiv:2605.15093v1 Announce Type: new Abstract: The life history of an individual coral is archived within the accreting skeleton of the colony. While reef-forming coral colonies (e.g. massive Porites

CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding

ResearchDGX agent

arXiv:2605.14310v1 Announce Type: new Abstract: Streaming video understanding with large vision-language models (VLMs) requires a compact memory that can support future reasoning over an ever-growing

CoReDiT: Spatial Coherence-Guided Token Pruning and Reconstruction for Efficient Diffusion Transformers

Local AiDGX agent

arXiv:2605.14191v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) deliver remarkable image and video generation quality but incur high computational cost, limiting scalability and on-devic

Covariance-aware sampling for Diffusion Models

ResearchDGX agent

arXiv:2605.13910v1 Announce Type: cross Abstract: We present a covariance-aware sampler that improves the quality of pixel-space Diffusion Model (DM) sampling in the few-step regime. We hypothesize th

CreFlow: Corrective Reflow for Sparse-Reward Embodied Video Diffusion RL

ResearchDGX agent

arXiv:2605.14274v1 Announce Type: new Abstract: Video generation models trained on heterogeneous data with likelihood-surrogate objectives can produce visually plausible rollouts that violate physical

Cross-Domain Few-Shot Segmentation via Ordinary Differential Equations over Time Intervals

ApplicationsDGX agent

arXiv:2509.01299v2 Announce Type: replace Abstract: Cross-domain few-shot segmentation (CD-FSS) aims to segment unseen categories with very limited samples while alleviating the negative effects of do

CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves

Model ReleasesDGX agent

arXiv:2605.14068v1 Announce Type: new Abstract: We introduce CurveBench, a benchmark for hierarchical topological reasoning from visual input. CurveBench consists of extbf{756 images} of pairwise non-

D2-CDIG: Controlled Diffusion Remote Sensing Image Generation with Dual Priors of DEM and Cloud-Fog

ResearchDGX agent

arXiv:2605.14326v1 Announce Type: new Abstract: Remote sensing image generation provides a reliable data foundation for remote sensing large models and downstream tasks. However, existing controllable

DAPL: Integration of Positive and Negative Descriptions in Text-Based Person Search

SafetyDGX agent

arXiv:2405.07459v3 Announce Type: replace Abstract: Text-based person search (TBPS) aims to retrieve specific images of individuals from large datasets using textual descriptions. Existing TBPS method

Deep Image Segmentation via Discriminant Feature Learning

Model ReleasesDGX agent

arXiv:2605.14609v1 Announce Type: new Abstract: Accurate image segmentation remains challenging, particularly in generating sharp, confident boundaries. While modern architectures have advanced the fi

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation

SafetyDGX agent

arXiv:2605.14382v1 Announce Type: new Abstract: Interactive real-time autoregressive video generation is essential for applications such as content creation and world modeling, where visual content mu

Denoising-GS: Gaussian Splatting with Spatial-aware Denoising

Model ReleasesDGX agent

arXiv:2605.14880v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting (3DGS) have achieved remarkable success in high-fidelity Novel View Synthesis (NVS), yet the optimization proce

DermAgent: A Self-Reflective Agentic System for Dermatological Image Analysis with Multi-Tool Reasoning and Traceable Decision-Making

AgentsDGX agent

arXiv:2605.14403v1 Announce Type: new Abstract: Dermatological diagnosis requires integrating fine-grained visual perception with expert clinical knowledge. Although Multimodal Large Language Models (

Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers

ResearchDGX agent

arXiv:2605.14270v1 Announce Type: new Abstract: Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-to-image generation, yet they frequently suffer from concept omiss

DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models

SafetyDGX agent

arXiv:2605.15055v1 Announce Type: cross Abstract: Reinforcement learning has emerged as a powerful tool for improving diffusion-based text-to-image models, but existing methods are largely limited to

DIVER: Reinforced Diffusion Breaks Imitation Bottlenecks in End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2507.04049v4 Announce Type: replace Abstract: Most end-to-end autonomous driving methods rely on imitation learning from single expert demonstrations, often leading to conservative and homogeneo

Do-Undo Bench: Reversibility for Action Understanding in Image Generation

Model ReleasesDGX agent

arXiv:2512.13609v2 Announce Type: replace Abstract: We introduce the Do-Undo task and benchmark to address a critical gap in vision-language models: understanding and generating plausible scene transf

Does Synthetic Layered Design Data Benefit Layered Design Decomposition?

ApplicationsDGX agent

arXiv:2605.15167v1 Announce Type: new Abstract: Recent advances in image generation have made it easy to produce high-quality images. However, these outputs are inherently flattened, entangling foregr

DriveCtrl: Conditioned Sim-to-Real Driving Video Generation

AgentsDGX agent

arXiv:2605.15116v1 Announce Type: new Abstract: Large-scale labelled driving video data is essential for training autonomous driving systems. Although simulation offers scalable and fully annotated da

Dual-Latent Collaborative Decoding for Fidelity-Perception Balanced Image Compression

ResearchDGX agent

arXiv:2605.14391v1 Announce Type: new Abstract: Learned image compression (LIC) increasingly requires reconstructions that balance distortion fidelity and perceptual realism across a wide range of bit

DUET: Dual-Paradigm Adaptive Expert Triage with Single-cell Inductive Prior for Spatial Transcriptomics Prediction

ResearchDGX agent

arXiv:2605.14104v1 Announce Type: new Abstract: Inferring spatially resolved gene expression from histology images offers a cost-effective complement to spatial transcriptomics (ST). However, existing

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding

ResearchDGX agent

arXiv:2605.14742v1 Announce Type: new Abstract: Understanding human--environment interactions from egocentric vision is essential for assistive robotics and embodied intelligent agents, yet existing m

Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis

Model ReleasesDGX agent

arXiv:2605.14842v1 Announce Type: new Abstract: Humans naturally communicate through abstract concepts like 'mood'. However, current image editing benchmarks focus primarily on explicit, literal comma

Efficient Dense Matching for Enhanced Gaussian Splatting Using AV1 Motion Vectors

ResearchDGX agent

arXiv:2605.14629v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a prominent framework for real-time, photorealistic scene reconstruction, offering significant speed-ups o

Enhancing Few-Shot Classification of Benchmark and Disaster Imagery with ABHFA-Net

Model ReleasesDGX agent

arXiv:2510.18326v3 Announce Type: replace Abstract: The rising incidence of natural and human-induced disasters necessitates robust visual recognition systems capable of operating under limited labele

EponaV2: Driving World Model with Comprehensive Future Reasoning

SafetyDGX agent

arXiv:2605.14696v1 Announce Type: new Abstract: Data scaling plays a pivotal role in the pursuit of general intelligence. However, the prevailing perception-planning paradigm in autonomous driving rel

Every Subtlety Counts: Fine-grained Person Independence Micro-Action Recognition via Distributionally Robust Optimization

SafetyDGX agent

arXiv:2509.21261v3 Announce Type: replace Abstract: Micro-action Recognition is vital for psychological assessment and human-computer interaction. However, existing methods often fail in real-world sc

Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model

SafetyDGX agent

arXiv:2605.14950v1 Announce Type: new Abstract: Vision-Language-Action models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action gener

Evolving Layer-Specific Scalar Functions for Hardware-Aware Transformer Adaptation

SafetyDGX agent

arXiv:2605.14047v1 Announce Type: new Abstract: Vision Transformers (ViTs) achieve state-of-the-art performance on challenging vision tasks, but their deployment on edge devices is severely hindered b

Exploring Vision-Language Models for Online Signature Verification: A Zero-Shot Capability Study

Model ReleasesDGX agent

arXiv:2605.14845v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) have demonstrated strong capabilities in general visual reasoning, yet their applicability to rigor

FALO: Fast and Accurate LiDAR 3D Object Detection on Resource-Constrained Devices

HardwareDGX agent

arXiv:2506.04499v2 Announce Type: replace Abstract: Existing LiDAR 3D object detection methods predominantely rely on sparse convolutions and/or transformers, which can be challenging to run on resour

FedStain: Modeling Higher-Order Stain Statistics for Federated Domain Generalization in Computational Pathology

Model ReleasesDGX agent

arXiv:2605.14590v1 Announce Type: new Abstract: Robust whole-slide image (WSI) analysis under strict data-governance remains challenging due to substantial cross-institutional stain heterogeneity. Dom

FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching

Model ReleasesDGX agent

arXiv:2604.06757v2 Announce Type: replace Abstract: Multimodal generation has long been dominated by text-driven pipelines where language dictates vision but cannot reason or create within it. We chal

← Previous
1…134135136137138…211
Next →