AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

Breaking Dual Bottlenecks: Evolving Unified Multimodal Models into Self-Adaptive Interleaved Visual Reasoners

DGX agent

arXiv:2605.14709v1 Announce Type: new Abstract: Recent unified models integrate multimodal understanding and generation within a single framework. However, an 'understanding-generation gap' persists,

model-releasesarxiv-cs-cv
15 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction

DGX agent

arXiv:2605.14569v1 Announce Type: new Abstract: Reconstructing dynamic visual experiences as videos from functional magnetic resonance imaging (fMRI) is pivotal for advancing the understanding of neur

researcharxiv-cs-cv
15 May 2026
Applications

CalibAnyView: Beyond Single-View Camera Calibration in the Wild

DGX agent

arXiv:2605.14615v1 Announce Type: new Abstract: Camera calibration is a fundamental prerequisite for reliable geometric perception, yet classical approaches rely on controlled acquisition setups that

applicationsarxiv-cs-cv
15 May 2026
Model Releases

Can Visual Mamba Improve AI-Generated Image Detection? An In-Depth Investigation

DGX agent

arXiv:2605.14799v1 Announce Type: new Abstract: In recent years, computer vision has witnessed remarkable progress, fueled by the development of innovative architectures such as Convolutional Neural N

model-releasesarxiv-cs-cv
15 May 2026
Research

Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation

DGX agent

arXiv:2605.15141v1 Announce Type: new Abstract: Real-time interactive video generation requires low-latency, streaming, and controllable rollout. Existing autoregressive (AR) diffusion distillation me

researcharxiv-cs-cv
15 May 2026
Research

CC-Pan: Channel-wise Compression based Diffusion for Efficient Pan-Sharpening

DGX agent

arXiv:2602.04473v2 Announce Type: replace Abstract: Recently, diffusion models have brought novel insights to pan-sharpening and notably boosted fusion precision. However, most existing models perform

researcharxiv-cs-cv
15 May 2026
Tutorials

Characterizing the visual representation of objects from the child's view

DGX agent

arXiv:2605.14990v1 Announce Type: new Abstract: Children acquire object category representations from their everyday experiences in the first few years of life. What do the inputs to this learning pro

tutorialsarxiv-cs-cv
15 May 2026
Safety

CHASM: Cross-frequency Harmonized Axis-Separable Mixing for Spectral Token Operators

DGX agent

arXiv:2605.14727v1 Announce Type: new Abstract: Spectral token mixers based on Fourier transforms provide an efficient way to model global interactions in visual feature maps. Existing designs often e

safetyarxiv-cs-cv
15 May 2026
Research

ClickRemoval: An Interactive Open-Source Tool for Object Removal in Diffusion Models

DGX agent

arXiv:2605.14461v1 Announce Type: new Abstract: Existing object removal tools often rely on manual masks or text prompts, making precise removal difficult for non-expert users in complex scenes and of

researcharxiv-cs-cv
15 May 2026
Research

Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers

DGX agent

arXiv:2511.14751v2 Announce Type: replace Abstract: We propose Confidence-Guided Token Merging (Co-Me), an acceleration mechanism for visual geometric transformers without retraining or finetuning the

researcharxiv-cs-cv
15 May 2026
Safety

COAL: Counterfactual and Observation-Enhanced Alignment Learning for Discriminative Referring Multi-Object Tracking

DGX agent

arXiv:2605.14795v1 Announce Type: new Abstract: Referring Multi-Object Tracking (RMOT) faces a fundamental structural contradiction between the high-discriminability demand and the sparse semantic sup

safetyarxiv-cs-cv
15 May 2026
Model Releases

CoCoEdit: Content-Consistent Image Editing via Region Regularized Reinforcement Learning

DGX agent

arXiv:2602.14068v2 Announce Type: replace Abstract: Image editing has achieved impressive results with the development of large-scale generative models. However, existing models mainly focus on the ed

model-releasesarxiv-cs-cv
15 May 2026
Research

Compositional Video Generation via Inference-Time Guidance

DGX agent

arXiv:2605.14988v1 Announce Type: new Abstract: Text-to-video diffusion models generate realistic videos, but often fail on prompts requiring fine-grained compositional understanding, such as relation

researcharxiv-cs-cv
15 May 2026
Research

Computational Imaging Priors for Wireless Capsule Endoscopy: Monte Carlo-Guided Hemoglobin Mapping for Rare-Anomaly Detection

DGX agent

arXiv:2605.15062v1 Announce Type: new Abstract: Background. RGB-trained capsule-endoscopy classifiers underperform on small-vessel vascular findings by conflating hemoglobin contrast with bile and ill

researcharxiv-cs-cv
15 May 2026
Applications

Contrastive Multi-Modal Hypergraph Reasoning for 3D Crowd Mesh Recovery

DGX agent

arXiv:2605.13854v1 Announce Type: new Abstract: Multi-person 3D reconstruction is pivotal for real-world interaction analysis, yet remains challenging due to severe occlusions and depth ambiguity. Cur

applicationsarxiv-cs-cv
15 May 2026
Research

CoralLite: {mu}CT Reconstruction of Coral Colonies from Individual Corallites

DGX agent

arXiv:2605.15093v1 Announce Type: new Abstract: The life history of an individual coral is archived within the accreting skeleton of the colony. While reef-forming coral colonies (e.g. massive Porites

researcharxiv-cs-cv
15 May 2026
Research

CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding

DGX agent

arXiv:2605.14310v1 Announce Type: new Abstract: Streaming video understanding with large vision-language models (VLMs) requires a compact memory that can support future reasoning over an ever-growing

researcharxiv-cs-cv
15 May 2026
Local Ai

CoReDiT: Spatial Coherence-Guided Token Pruning and Reconstruction for Efficient Diffusion Transformers

DGX agent

arXiv:2605.14191v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) deliver remarkable image and video generation quality but incur high computational cost, limiting scalability and on-devic

local-aiarxiv-cs-cv
15 May 2026
Research

Covariance-aware sampling for Diffusion Models

DGX agent

arXiv:2605.13910v1 Announce Type: cross Abstract: We present a covariance-aware sampler that improves the quality of pixel-space Diffusion Model (DM) sampling in the few-step regime. We hypothesize th

researcharxiv-cs-cv
15 May 2026
Research

CreFlow: Corrective Reflow for Sparse-Reward Embodied Video Diffusion RL

DGX agent

arXiv:2605.14274v1 Announce Type: new Abstract: Video generation models trained on heterogeneous data with likelihood-surrogate objectives can produce visually plausible rollouts that violate physical

researcharxiv-cs-cv
15 May 2026
Applications

Cross-Domain Few-Shot Segmentation via Ordinary Differential Equations over Time Intervals

DGX agent

arXiv:2509.01299v2 Announce Type: replace Abstract: Cross-domain few-shot segmentation (CD-FSS) aims to segment unseen categories with very limited samples while alleviating the negative effects of do

applicationsarxiv-cs-cv
15 May 2026
Model Releases

CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves

DGX agent

arXiv:2605.14068v1 Announce Type: new Abstract: We introduce CurveBench, a benchmark for hierarchical topological reasoning from visual input. CurveBench consists of extbf{756 images} of pairwise non-

model-releasesarxiv-cs-cv
15 May 2026
Research

D2-CDIG: Controlled Diffusion Remote Sensing Image Generation with Dual Priors of DEM and Cloud-Fog

DGX agent

arXiv:2605.14326v1 Announce Type: new Abstract: Remote sensing image generation provides a reliable data foundation for remote sensing large models and downstream tasks. However, existing controllable

researcharxiv-cs-cv
15 May 2026
Safety

DAPL: Integration of Positive and Negative Descriptions in Text-Based Person Search

DGX agent

arXiv:2405.07459v3 Announce Type: replace Abstract: Text-based person search (TBPS) aims to retrieve specific images of individuals from large datasets using textual descriptions. Existing TBPS method

safetyarxiv-cs-cv
15 May 2026
Model Releases

Deep Image Segmentation via Discriminant Feature Learning

DGX agent

arXiv:2605.14609v1 Announce Type: new Abstract: Accurate image segmentation remains challenging, particularly in generating sharp, confident boundaries. While modern architectures have advanced the fi

model-releasesarxiv-cs-cv
15 May 2026
Safety

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation

DGX agent

arXiv:2605.14382v1 Announce Type: new Abstract: Interactive real-time autoregressive video generation is essential for applications such as content creation and world modeling, where visual content mu

safetyarxiv-cs-cv
15 May 2026
Model Releases

Denoising-GS: Gaussian Splatting with Spatial-aware Denoising

DGX agent

arXiv:2605.14880v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting (3DGS) have achieved remarkable success in high-fidelity Novel View Synthesis (NVS), yet the optimization proce

model-releasesarxiv-cs-cv
15 May 2026
Agents

DermAgent: A Self-Reflective Agentic System for Dermatological Image Analysis with Multi-Tool Reasoning and Traceable Decision-Making

DGX agent

arXiv:2605.14403v1 Announce Type: new Abstract: Dermatological diagnosis requires integrating fine-grained visual perception with expert clinical knowledge. Although Multimodal Large Language Models (

agentsarxiv-cs-cv
15 May 2026
Research

Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers

DGX agent

arXiv:2605.14270v1 Announce Type: new Abstract: Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-to-image generation, yet they frequently suffer from concept omiss

researcharxiv-cs-cv
15 May 2026
Safety

DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models

DGX agent

arXiv:2605.15055v1 Announce Type: cross Abstract: Reinforcement learning has emerged as a powerful tool for improving diffusion-based text-to-image models, but existing methods are largely limited to

safetyarxiv-cs-cv
15 May 2026
Safety

DIVER: Reinforced Diffusion Breaks Imitation Bottlenecks in End-to-End Autonomous Driving

DGX agent

arXiv:2507.04049v4 Announce Type: replace Abstract: Most end-to-end autonomous driving methods rely on imitation learning from single expert demonstrations, often leading to conservative and homogeneo

safetyarxiv-cs-cv
15 May 2026
Model Releases

Do-Undo Bench: Reversibility for Action Understanding in Image Generation

DGX agent

arXiv:2512.13609v2 Announce Type: replace Abstract: We introduce the Do-Undo task and benchmark to address a critical gap in vision-language models: understanding and generating plausible scene transf

model-releasesarxiv-cs-cv
15 May 2026
Applications

Does Synthetic Layered Design Data Benefit Layered Design Decomposition?

DGX agent

arXiv:2605.15167v1 Announce Type: new Abstract: Recent advances in image generation have made it easy to produce high-quality images. However, these outputs are inherently flattened, entangling foregr

applicationsarxiv-cs-cv
15 May 2026
Agents

DriveCtrl: Conditioned Sim-to-Real Driving Video Generation

DGX agent

arXiv:2605.15116v1 Announce Type: new Abstract: Large-scale labelled driving video data is essential for training autonomous driving systems. Although simulation offers scalable and fully annotated da

agentsarxiv-cs-cv
15 May 2026
Research

Dual-Latent Collaborative Decoding for Fidelity-Perception Balanced Image Compression

DGX agent

arXiv:2605.14391v1 Announce Type: new Abstract: Learned image compression (LIC) increasingly requires reconstructions that balance distortion fidelity and perceptual realism across a wide range of bit

researcharxiv-cs-cv
15 May 2026
Research

DUET: Dual-Paradigm Adaptive Expert Triage with Single-cell Inductive Prior for Spatial Transcriptomics Prediction

DGX agent

arXiv:2605.14104v1 Announce Type: new Abstract: Inferring spatially resolved gene expression from histology images offers a cost-effective complement to spatial transcriptomics (ST). However, existing

researcharxiv-cs-cv
15 May 2026
Research

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding

DGX agent

arXiv:2605.14742v1 Announce Type: new Abstract: Understanding human--environment interactions from egocentric vision is essential for assistive robotics and embodied intelligent agents, yet existing m

researcharxiv-cs-cv
15 May 2026
Model Releases

Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis

DGX agent

arXiv:2605.14842v1 Announce Type: new Abstract: Humans naturally communicate through abstract concepts like 'mood'. However, current image editing benchmarks focus primarily on explicit, literal comma

model-releasesarxiv-cs-cv
15 May 2026
Research

Efficient Dense Matching for Enhanced Gaussian Splatting Using AV1 Motion Vectors

DGX agent

arXiv:2605.14629v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a prominent framework for real-time, photorealistic scene reconstruction, offering significant speed-ups o

researcharxiv-cs-cv
15 May 2026
Model Releases

Enhancing Few-Shot Classification of Benchmark and Disaster Imagery with ABHFA-Net

DGX agent

arXiv:2510.18326v3 Announce Type: replace Abstract: The rising incidence of natural and human-induced disasters necessitates robust visual recognition systems capable of operating under limited labele

model-releasesarxiv-cs-cv
15 May 2026
Safety

EponaV2: Driving World Model with Comprehensive Future Reasoning

DGX agent

arXiv:2605.14696v1 Announce Type: new Abstract: Data scaling plays a pivotal role in the pursuit of general intelligence. However, the prevailing perception-planning paradigm in autonomous driving rel

safetyarxiv-cs-cv
15 May 2026
Safety

Every Subtlety Counts: Fine-grained Person Independence Micro-Action Recognition via Distributionally Robust Optimization

DGX agent

arXiv:2509.21261v3 Announce Type: replace Abstract: Micro-action Recognition is vital for psychological assessment and human-computer interaction. However, existing methods often fail in real-world sc

safetyarxiv-cs-cv
15 May 2026
Safety

Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model

DGX agent

arXiv:2605.14950v1 Announce Type: new Abstract: Vision-Language-Action models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action gener

safetyarxiv-cs-cv
15 May 2026
Safety

Evolving Layer-Specific Scalar Functions for Hardware-Aware Transformer Adaptation

DGX agent

arXiv:2605.14047v1 Announce Type: new Abstract: Vision Transformers (ViTs) achieve state-of-the-art performance on challenging vision tasks, but their deployment on edge devices is severely hindered b

safetyarxiv-cs-cv
15 May 2026
Model Releases

Exploring Vision-Language Models for Online Signature Verification: A Zero-Shot Capability Study

DGX agent

arXiv:2605.14845v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) have demonstrated strong capabilities in general visual reasoning, yet their applicability to rigor

model-releasesarxiv-cs-cv
15 May 2026
Hardware

FALO: Fast and Accurate LiDAR 3D Object Detection on Resource-Constrained Devices

DGX agent

arXiv:2506.04499v2 Announce Type: replace Abstract: Existing LiDAR 3D object detection methods predominantely rely on sparse convolutions and/or transformers, which can be challenging to run on resour

hardwarearxiv-cs-cv
15 May 2026
Model Releases

FedStain: Modeling Higher-Order Stain Statistics for Federated Domain Generalization in Computational Pathology

DGX agent

arXiv:2605.14590v1 Announce Type: new Abstract: Robust whole-slide image (WSI) analysis under strict data-governance remains challenging due to substantial cross-institutional stain heterogeneity. Dom

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching

DGX agent

arXiv:2604.06757v2 Announce Type: replace Abstract: Multimodal generation has long been dominated by text-driven pipelines where language dictates vision but cannot reason or create within it. We chal

model-releasesarxiv-cs-cv
15 May 2026
← Previous
1…168169170171172…263
Next →