AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Research

Delineate Anything Flow: Fast, Country-Level Field Boundary Detection from Any Source

DGX agent

arXiv:2511.13417v2 Announce Type: replace Abstract: Accurate delineation of agricultural field boundaries from satellite imagery is essential for land management and crop monitoring, yet existing meth

researcharxiv-cs-cv
16 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

Democratising Pathology Co-Pilots: An Open Pipeline and Dataset for Whole-Slide Vision-Language Modelling

DGX agent

arXiv:2512.17326v2 Announce Type: replace Abstract: Vision-language models (VLMs) have the potential to become co-pilots for pathologists. However, most VLMs either focus on small regions of interest

researcharxiv-cs-cv
16 Apr 2026
Safety

Depth-Aware Image and Video Orientation Estimation

DGX agent

arXiv:2604.13995v1 Announce Type: new Abstract: This paper introduces a novel approach for image and video orientation estimation by leveraging depth distribution in natural images. The proposed metho

safetyarxiv-cs-cv
16 Apr 2026
Research

Depth-Resolved Coral Reef Thermal Fields from Satellite SST and Sparse In-Situ Loggers Using Physics-Informed Neural Networks

DGX agent

arXiv:2604.13131v1 Announce Type: cross Abstract: Satellite sea surface temperature (SST) products underpin global coral bleaching monitoring, yet they measure only the ocean skin. Corals inhabit dept

researcharxiv-cs-cv
16 Apr 2026
Research

Design and Behavior of Sparse Mixture-of-Experts Layers in CNN-based Semantic Segmentation

DGX agent

arXiv:2604.13761v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) layers have been shown to substantially increase model capacity without a proportional increase in computational cost an

researcharxiv-cs-cv
16 Apr 2026
Model Releases

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis

DGX agent

arXiv:2604.13416v1 Announce Type: new Abstract: Advances in radiance fields have enabled photorealistic novel view synthesis. In several domains, large-scale real-world datasets have been developed to

model-releasesarxiv-cs-cv
16 Apr 2026
Research

DiffMagicFace: Identity Consistent Facial Editing of Real Videos

DGX agent

arXiv:2604.13841v1 Announce Type: new Abstract: Text-conditioned image editing has greatly benefitted from the advancements in Image Diffusion Models. However, extending these techniques to facial vid

researcharxiv-cs-cv
16 Apr 2026
Research

DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer

DGX agent

arXiv:2604.13509v1 Announce Type: new Abstract: Recent advances in video generation models has significantly accelerated video generation and related downstream tasks. Among these, video stylization h

researcharxiv-cs-cv
16 Apr 2026
Research

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models

DGX agent

arXiv:2604.14129v1 Announce Type: new Abstract: While Audio-Visual Language Models (AVLMs) have achieved remarkable progress over recent years, their reliability is bottlenecked by cross-modal halluci

researcharxiv-cs-cv
16 Apr 2026
Research

DRG-Font: Dynamic Reference-Guided Few-shot Font Generation via Contrastive Style-Content Disentanglement

DGX agent

arXiv:2604.13797v1 Announce Type: new Abstract: Few-shot Font Generation aims to generate stylistically consistent glyphs from a few reference glyphs. However, capturing complex font styles from a few

researcharxiv-cs-cv
16 Apr 2026
Research

DroneScan-YOLO: Redundancy-Aware Lightweight Detection for Tiny Objects in UAV Imagery

DGX agent

arXiv:2604.13278v1 Announce Type: new Abstract: Aerial object detection in UAV imagery presents unique challenges due to the high prevalence of tiny objects, adverse environmental conditions, and stri

researcharxiv-cs-cv
16 Apr 2026
Safety

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models

DGX agent

arXiv:2602.20981v3 Announce Type: replace Abstract: Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and

safetyarxiv-cs-cv
16 Apr 2026
Model Releases

Efficient Multi-View 3D Object Detection by Dynamic Token Selection and Fine-Tuning

DGX agent

arXiv:2604.13586v1 Announce Type: new Abstract: Existing multi-view three-dimensional (3D) object detection approaches widely adopt large-scale pre-trained vision transformer (ViT)-based foundation mo

model-releasesarxiv-cs-cv
16 Apr 2026
Safety

Enhanced Text-to-Image Generation by Fine-grained Multimodal Reasoning

DGX agent

arXiv:2604.13491v1 Announce Type: new Abstract: With the rapid progress of Multimodal Large Language Models (MLLMs), unified MLLMs that jointly perform image understanding and generation have advanced

safetyarxiv-cs-cv
16 Apr 2026
Research

Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling

DGX agent

arXiv:2604.13508v1 Announce Type: new Abstract: Sparse Upcycling provides an efficient way to initialize a Mixture-of-Experts (MoE) model from pretrained dense weights instead of training from scratch

researcharxiv-cs-cv
16 Apr 2026
Model Releases

ESCAPE: Episodic Spatial Memory and Adaptive Execution Policy for Long-Horizon Mobile Manipulation

DGX agent

arXiv:2604.13633v1 Announce Type: new Abstract: Coordinating navigation and manipulation with robust performance is essential for embodied AI in complex indoor environments. However, as tasks extend o

model-releasesarxiv-cs-cv
16 Apr 2026
Research

Event-Adaptive State Transition and Gated Fusion for RGB-Event Object Tracking

DGX agent

arXiv:2604.13426v1 Announce Type: new Abstract: Existing Vision Mamba-based RGB-Event(RGBE) tracking methods suffer from using static state transition matrices, which fail to adapt to variations in ev

researcharxiv-cs-cv
16 Apr 2026
Safety

Evolvable Embodied Agent for Robotic Manipulation via Long Short-Term Reflection and Optimization

DGX agent

arXiv:2604.13533v1 Announce Type: cross Abstract: Achieving general-purpose robotics requires empowering robots to adapt and evolve based on their environment and feedback. Traditional methods face li

safetyarxiv-cs-cv
16 Apr 2026
Local Ai

Explainable Fall Detection for Elderly Care via Temporally Stable SHAP in Skeleton-Based Human Activity Recognition

DGX agent

arXiv:2604.13279v1 Announce Type: new Abstract: Fall detection in elderly care requires not only accurate classification but also reliable explanations that clinicians can trust. However, existing pos

local-aiarxiv-cs-cv
16 Apr 2026
Safety

Failure Identification in Imitation Learning Via Statistical and Semantic Filtering

DGX agent

arXiv:2604.13788v1 Announce Type: cross Abstract: Imitation learning (IL) policies in robotics deliver strong performance in controlled settings but remain brittle in real-world deployments: rare even

safetyarxiv-cs-cv
16 Apr 2026
Applications

Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective

DGX agent

arXiv:2604.14025v1 Announce Type: new Abstract: Reconstructing 3D representations from 2D inputs is a fundamental task in computer vision and graphics, serving as a cornerstone for understanding and i

applicationsarxiv-cs-cv
16 Apr 2026
Model Releases

FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks

DGX agent

arXiv:2505.19662v3 Announce Type: replace-cross Abstract: This paper introduces FieldWorkArena, a benchmark for agentic AI targeting real-world field work. With the recent increase in demand for agent

model-releasesarxiv-cs-cv
16 Apr 2026
Model Releases

Free Geometry: Refining 3D Reconstruction from Longer Versions of Itself

DGX agent

arXiv:2604.14048v1 Announce Type: new Abstract: Feed-forward 3D reconstruction models are efficient but rigid: once trained, they perform inference in a zero-shot manner and cannot adapt to the test s

model-releasesarxiv-cs-cv
16 Apr 2026
Research

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding

DGX agent

arXiv:2604.13540v1 Announce Type: new Abstract: Unified Multimodal Models (UMMs) aim to integrate visual understanding and generation within a single structure. However, these models exhibit a notable

researcharxiv-cs-cv
16 Apr 2026
Safety

From Instruction to Event: Sound-Triggered Mobile Manipulation

DGX agent

arXiv:2601.21667v2 Announce Type: replace-cross Abstract: Current mobile manipulation research predominantly follows an instruction-driven paradigm, where agents rely on predefined textual commands to

safetyarxiv-cs-cv
16 Apr 2026
Research

From Pixels to Nucleotides: End-to-End Token-Based Video Compression for DNA Storage

DGX agent

arXiv:2604.13667v1 Announce Type: new Abstract: DNA-based storage has emerged as a promising approach to the global data crisis, offering molecular-scale density and millennial-scale stability at low

researcharxiv-cs-cv
16 Apr 2026
Model Releases

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models

DGX agent

arXiv:2603.19790v3 Announce Type: replace Abstract: Modern vision-language models (VLMs) can act as generative OCR engines, yet open-ended decoding can expose rare but consequential failures. We ident

model-releasesarxiv-cs-cv
16 Apr 2026
Research

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation

DGX agent

arXiv:2604.13793v1 Announce Type: new Abstract: Exo-to-Ego video generation aims to synthesize a first-person video from a synchronized third-person view and corresponding camera poses. While paired s

researcharxiv-cs-cv
16 Apr 2026
Research

Frozen Forecasting: A Unified Evaluation

DGX agent

arXiv:2507.13942v2 Announce Type: replace Abstract: Forecasting future events is a fundamental capability for general-purpose systems that plan or act across different levels of abstraction. Yet, eval

researcharxiv-cs-cv
16 Apr 2026
Model Releases

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation

DGX agent

arXiv:2604.13803v1 Announce Type: new Abstract: Vision-language models are increasingly deployed in high-stakes settings, yet their susceptibility to sycophantic manipulation remains poorly understood

model-releasesarxiv-cs-cv
16 Apr 2026
Model Releases

GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization

DGX agent

arXiv:2512.02697v3 Announce Type: replace Abstract: Cross-view geo-localization infers a location by retrieving geo-tagged reference images that visually correspond to a query image. However, the trad

model-releasesarxiv-cs-cv
16 Apr 2026
Local Ai

GeoLink: A 3D-Aware Framework Towards Better Generalization in Cross-View Geo-Localization

DGX agent

arXiv:2604.13183v1 Announce Type: new Abstract: Generalizable cross-view geo-localization aims to match the same location across views in unseen regions and conditions without GPS supervision. Its cor

local-aiarxiv-cs-cv
16 Apr 2026
Research

Geometric Context Transformer for Streaming 3D Reconstruction

DGX agent

arXiv:2604.14141v1 Announce Type: new Abstract: Streaming 3D reconstruction aims to recover 3D information, such as camera poses and point clouds, from a video stream, which necessitates geometric acc

researcharxiv-cs-cv
16 Apr 2026
Research

Getting the Numbers Rightnicode{x2014}Modelling Multi-Class Object Counting in Dense and Varied Scenes

DGX agent

arXiv:2510.02213v2 Announce Type: replace Abstract: Density map estimation enables accurate object counting in heavily occluded, and densely packed scenes where detection-based counting fails. In mult

researcharxiv-cs-cv
16 Apr 2026
Research

Granularity-Aware Transfer for Tree Instance Segmentation in Synthetic and Real Forests

DGX agent

arXiv:2604.13722v1 Announce Type: new Abstract: We address the challenge of synthetic-to-real transfer in forestry perception where real data have only coarse Tree labels while synthetic data provide

researcharxiv-cs-cv
16 Apr 2026
Research

Graph Propagated Projection Unlearning: A Unified Framework for Vision and Audio Discriminative Models

DGX agent

arXiv:2604.13127v1 Announce Type: new Abstract: The need to selectively and efficiently erase learned information from deep neural networks is becoming increasingly important for privacy, regulatory c

researcharxiv-cs-cv
16 Apr 2026
Safety

HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy

DGX agent

arXiv:2510.00695v3 Announce Type: replace-cross Abstract: Inherently, robotic manipulation tasks are history-dependent: leveraging past context could be beneficial. However, most existing Vision-Langu

safetyarxiv-cs-cv
16 Apr 2026
Safety

Heavy-Tailed Class-Conditional Priors for Long-Tailed Generative Modeling

DGX agent

arXiv:2509.02154v2 Announce Type: replace-cross Abstract: Variational Autoencoders (VAEs) with global priors trained under an imbalanced empirical class distribution can lead to underrepresentation of

safetyarxiv-cs-cv
16 Apr 2026
Research

Heuristic Style Transfer for Real-Time, Efficient Weather Attribute Detection

DGX agent

arXiv:2604.13947v1 Announce Type: new Abstract: We present lightweight and efficient architectures to detect weather conditions from RGB images, predicting the weather type (sunny, rain, snow, fog) an

researcharxiv-cs-cv
16 Apr 2026
Research

HiProto: Hierarchical Prototype Learning for Interpretable Object Detection Under Low-quality Conditions

DGX agent

arXiv:2604.13981v1 Announce Type: new Abstract: Interpretability is essential for deploying object detection systems in critical applications, especially under low-quality imaging conditions that degr

researcharxiv-cs-cv
16 Apr 2026
Applications

HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System

DGX agent

arXiv:2604.14125v1 Announce Type: new Abstract: While end-to-end Vision-Language-Action (VLA) models offer a promising paradigm for robotic manipulation, fine-tuning them on narrow control data often

applicationsarxiv-cs-cv
16 Apr 2026
Research

Hybrid Approach for Enhancing Lesion Segmentation in Fundus Images

DGX agent

arXiv:2509.25549v2 Announce Type: replace Abstract: Choroidal nevi are common benign pigmented lesions in the eye, with a small risk of transforming into melanoma. Early detection is critical to impro

researcharxiv-cs-cv
16 Apr 2026
Hardware

Hydra: Unifying Document Retrieval and Generation in a Single Vision-Language Model

DGX agent

arXiv:2603.28554v2 Announce Type: replace Abstract: Visual document understanding typically requires separate retrieval and generation models, doubling memory and system complexity. We present Hydra,

hardwarearxiv-cs-cv
16 Apr 2026
Safety

Learning Class Difficulty in Imbalanced Histopathology Segmentation via Dynamic Focal Attention

DGX agent

arXiv:2604.13479v1 Announce Type: cross Abstract: Semantic segmentation of histopathology images under class imbalance is typically addressed through frequency-based loss reweighting, which implicitly

safetyarxiv-cs-cv
16 Apr 2026
Research

Learning Sewing Patterns via Latent Flow Matching of Implicit Fields

DGX agent

arXiv:2601.17740v2 Announce Type: replace Abstract: Sewing patterns define the structural foundation of garments and are essential for applications such as fashion design, fabrication, and physical si

researcharxiv-cs-cv
16 Apr 2026
Applications

Lite Any Stereo: Efficient Zero-Shot Stereo Matching

DGX agent

arXiv:2511.16555v3 Announce Type: replace Abstract: Recent advances in stereo matching have focused on accuracy, often at the cost of significantly increased model size. Traditionally, the community h

applicationsarxiv-cs-cv
16 Apr 2026
Model Releases

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding

DGX agent

arXiv:2602.20913v2 Announce Type: replace Abstract: This paper addresses the critical and underexplored challenge of long video understanding with low computational budgets. We propose LongVideo-R1, a

model-releasesarxiv-cs-cv
16 Apr 2026
Hardware

MaMe & MaRe: Matrix-Based Token Merging and Restoration for Efficient Visual Perception and Synthesis

DGX agent

arXiv:2604.13432v1 Announce Type: new Abstract: Token compression is crucial for mitigating the quadratic complexity of self-attention mechanisms in Vision Transformers (ViTs), which often involve num

hardwarearxiv-cs-cv
16 Apr 2026
← Previous
1…239240241242243…261
Next →