AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Model Releases

RobustSpring: Benchmarking Robustness to Image Corruptions for Optical Flow, Scene Flow and Stereo

DGX agent

arXiv:2505.09368v2 Announce Type: replace Abstract: Standard benchmarks for optical flow, scene flow, and stereo vision algorithms generally focus on model accuracy rather than robustness to image cor

model-releasesarxiv-cs-cv
14 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

rPPG-VQA: A Video Quality Assessment Framework for Unsupervised rPPG Training

DGX agent

arXiv:2604.11156v1 Announce Type: new Abstract: Unsupervised remote photoplethysmography (rPPG) promises to leverage unlabeled video data, but its potential is hindered by a critical challenge: traini

researcharxiv-cs-cv
14 Apr 2026
Research

S4M: 4-points to Segment Anything

DGX agent

arXiv:2503.05534v3 Announce Type: replace Abstract: Purpose: The Segment Anything Model (SAM) promises to ease the annotation bottleneck in medical segmentation, but overlapping anatomy and blurred bo

researcharxiv-cs-cv
14 Apr 2026
Hardware

SatReg: Regression-based Neural Architecture Search for Lightweight Satellite Image Segmentation

DGX agent

arXiv:2604.10306v1 Announce Type: new Abstract: As Earth-observation workloads move toward onboard and edge processing, remote-sensing segmentation models must operate under tight latency and energy c

hardwarearxiv-cs-cv
14 Apr 2026
Research

Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes

DGX agent

arXiv:2512.12982v2 Announce Type: replace Abstract: The pursuit of a universal AI-generated image (AIGI) detector often relies on aggregating data from numerous generators to improve generalization. H

researcharxiv-cs-cv
14 Apr 2026
Applications

Scene Change Detection with Vision-Language Representation Learning

DGX agent

arXiv:2604.11402v1 Announce Type: new Abstract: Scene change detection (SCD) is crucial for urban monitoring and navigation but remains challenging in real-world environments due to lighting variation

applicationsarxiv-cs-cv
14 Apr 2026
Research

SciPostLayoutTree: A Dataset for Structural Analysis of Scientific Posters

DGX agent

arXiv:2511.18329v3 Announce Type: replace Abstract: Scientific posters play a vital role in academic communication by presenting ideas through visual summaries. Analyzing reading order and parent-chil

researcharxiv-cs-cv
14 Apr 2026
Model Releases

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding

DGX agent

arXiv:2604.11244v1 Announce Type: new Abstract: Advances in Multimodal Large Language Models (MLLMs) are transforming video captioning from a descriptive endpoint into a semantic interface for both vi

model-releasesarxiv-cs-cv
14 Apr 2026
Local Ai

Search-MIND: Training-Free Multi-Modal Medical Image Registration

DGX agent

arXiv:2604.09743v1 Announce Type: cross Abstract: Multi-modal image registration plays a critical role in precision medicine but faces challenges from non-linear intensity relationships and local opti

local-aiarxiv-cs-cv
14 Apr 2026
Safety

See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment

DGX agent

arXiv:2604.09749v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) frequently hallucinate objects that are absent from the visual input, often because attention during decoding i

safetyarxiv-cs-cv
14 Apr 2026
Model Releases

Seeing Through the Tool: A Controlled Benchmark for Occlusion Robustness in Foundation Segmentation Models

DGX agent

arXiv:2604.11711v1 Announce Type: new Abstract: Occlusion, where target structures are partially hidden by surgical instruments or overlapping tissues, remains a critical yet underexplored challenge f

model-releasesarxiv-cs-cv
14 Apr 2026
Local Ai

Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions

DGX agent

arXiv:2604.11579v1 Announce Type: new Abstract: We address the problem of tactile localization, where the goal is to identify image regions that share the same material properties as a tactile input.

local-aiarxiv-cs-cv
14 Apr 2026
Model Releases

Seg2Change: Adapting Open-Vocabulary Semantic Segmentation Model for Remote Sensing Change Detection

DGX agent

arXiv:2604.11231v1 Announce Type: new Abstract: Change detection is a fundamental task in remote sensing, aiming to quantify the impacts of human activities and ecological dynamics on land-cover chang

model-releasesarxiv-cs-cv
14 Apr 2026
Model Releases

Self-supervised Pretraining of Cell Segmentation Models

DGX agent

arXiv:2604.10609v1 Announce Type: new Abstract: Instance segmentation enables the analysis of spatial and temporal properties of cells in microscopy images by identifying the pixels belonging to each

model-releasesarxiv-cs-cv
14 Apr 2026
Research

Sharpness-Aware Surrogate Training for On-Sensor Spiking Neural Networks

DGX agent

arXiv:2604.09696v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) are a natural computational model for on-sensor and near-sensor vision, where event driven processors must operate unde

researcharxiv-cs-cv
14 Apr 2026
Model Releases

SignReasoner: Compositional Reasoning for Complex Traffic Sign Understanding via Functional Structure Units

DGX agent

arXiv:2604.10436v1 Announce Type: new Abstract: Accurate semantic understanding of complex traffic signs-including those with intricate layouts, multi-lingual text, and composite symbols-is critical f

model-releasesarxiv-cs-cv
14 Apr 2026
Safety

SIMPLER: H&E-Informed Representation Learning for Structured Illumination Microscopy

DGX agent

arXiv:2604.10334v1 Announce Type: new Abstract: Structured Illumination Microscopy (SIM) enables rapid, high-contrast optical sectioning of fresh tissue without staining or physical sectioning, making

safetyarxiv-cs-cv
14 Apr 2026
Research

SinkTrack: Attention Sink based Context Anchoring for Large Language Models

DGX agent

arXiv:2604.10027v1 Announce Type: new Abstract: Large language models (LLMs) suffer from hallucination and context forgetting. Prior studies suggest that attention drift is a primary cause of these pr

researcharxiv-cs-cv
14 Apr 2026
Model Releases

SMFormer: Empowering Self-supervised Stereo Matching via Foundation Models and Data Augmentation

DGX agent

arXiv:2604.10218v1 Announce Type: new Abstract: Recent self-supervised stereo matching methods have made significant progress. They typically rely on the photometric consistency assumption, which pres

model-releasesarxiv-cs-cv
14 Apr 2026
Research

Sparse Hypergraph-Enhanced Frame-Event Object Detection with Fine-Grained MoE

DGX agent

arXiv:2604.11140v1 Announce Type: new Abstract: Integrating frame-based RGB cameras with event streams offers a promising solution for robust object detection under challenging dynamic conditions. How

researcharxiv-cs-cv
14 Apr 2026
Applications

Spatio-Temporal Difference Guided Motion Deblurring with the Complementary Vision Sensor

DGX agent

arXiv:2604.10554v1 Announce Type: new Abstract: Motion blur arises when rapid scene changes occur during the exposure period, collapsing rich intra-exposure motion into a single RGB frame. Without exp

applicationsarxiv-cs-cv
14 Apr 2026
Tutorials

Specificity-aware reinforcement learning for fine-grained open-world classification

DGX agent

arXiv:2603.03197v3 Announce Type: replace Abstract: Classifying fine-grained visual concepts under open-world settings, i.e., without a predefined label set, demands models to be both accurate and spe

tutorialsarxiv-cs-cv
14 Apr 2026
Local Ai

SpotFormer: Multi-Scale Spatio-Temporal Transformer for Facial Expression Spotting

DGX agent

arXiv:2407.20799v3 Announce Type: replace Abstract: Facial expression spotting, identifying periods where facial expressions occur in a video, is a significant yet challenging task in facial expressio

local-aiarxiv-cs-cv
14 Apr 2026
Tutorials

Spotlight and Shadow: Attention-Guided Dual-Anchor Introspective Decoding for MLLM Hallucination Mitigation

DGX agent

arXiv:2604.10071v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities yet continue to suffer from hallucination, where generated

tutorialsarxiv-cs-cv
14 Apr 2026
Safety

StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation

DGX agent

arXiv:2510.05057v2 Announce Type: replace-cross Abstract: A fundamental challenge in embodied intelligence is developing expressive and compact state representations for efficient world modeling and d

safetyarxiv-cs-cv
14 Apr 2026
Research

STGV: Spatio-Temporal Hash Encoding for Gaussian-based Video Representation

DGX agent

arXiv:2604.10910v1 Announce Type: new Abstract: 2D Gaussian Splatting (2DGS) has recently become a promising paradigm for high-quality video representation. However, existing methods employ content-ag

researcharxiv-cs-cv
14 Apr 2026
Tutorials

Structured State-Space Regularization for Compact and Generation-Friendly Image Tokenization

DGX agent

arXiv:2604.11089v1 Announce Type: new Abstract: Image tokenizers are central to modern vision models as they often operate in latent spaces. An ideal latent space must be simultaneously compact and ge

tutorialsarxiv-cs-cv
14 Apr 2026
Research

STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding

DGX agent

arXiv:2604.11637v1 Announce Type: new Abstract: 4D point cloud videos capture rich spatial and temporal dynamics of scenes which possess unique values in various 4D understanding tasks. However, most

researcharxiv-cs-cv
14 Apr 2026
Research

SwinTextUNet: Integrating CLIP-Based Text Guidance into Swin Transformer U-Nets for Medical Image Segmentation

DGX agent

arXiv:2604.10000v1 Announce Type: new Abstract: Precise medical image segmentation is fundamental for enabling computer aided diagnosis and effective treatment planning. Traditional models that rely s

researcharxiv-cs-cv
14 Apr 2026
Model Releases

Switch-JustDance: Benchmarking Whole Body Motion Tracking Controllers Using a Commercial Console Game

DGX agent

arXiv:2511.17925v3 Announce Type: replace-cross Abstract: Recent advances in whole-body robot control have enabled humanoid and legged robots to perform increasingly agile and coordinated motions. How

model-releasesarxiv-cs-cv
14 Apr 2026
Model Releases

SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation

DGX agent

arXiv:2604.03723v2 Announce Type: replace Abstract: Controlling both camera motion and object dynamics is essential for coherent and expressive video generation, yet current methods typically handle o

model-releasesarxiv-cs-cv
14 Apr 2026
Research

SyncFix: Fixing 3D Reconstructions via Multi-View Synchronization

DGX agent

arXiv:2604.11797v1 Announce Type: new Abstract: We present SyncFix, a framework that enforces cross-view consistency during the diffusion-based refinement of reconstructed scenes. SyncFix formulates r

researcharxiv-cs-cv
14 Apr 2026
Model Releases

TAG-Head: Time-Aligned Graph Head for Plug-and-Play Fine-grained Action Recognition

DGX agent

arXiv:2604.11498v1 Announce Type: new Abstract: Fine-grained human action recognition (FHAR) is challenging because visually similar actions differ by subtle spatio-temporal cues. Many recent systems

model-releasesarxiv-cs-cv
14 Apr 2026
Research

TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering

DGX agent

arXiv:2509.04123v2 Announce Type: replace Abstract: Text-to-story visualization is challenging due to the need for consistent interaction among multiple characters across frames. Existing methods stru

researcharxiv-cs-cv
14 Apr 2026
Research

TAMISeg: Text-Aligned Multi-scale Medical Image Segmentation with Semantic Encoder Distillation

DGX agent

arXiv:2604.10912v1 Announce Type: new Abstract: Medical image segmentation remains challenging due to limited fine-grained annotations, complex anatomical structures, and image degradation from noise,

researcharxiv-cs-cv
14 Apr 2026
Research

TAPNext++: What's Next for Tracking Any Point (TAP)?

DGX agent

arXiv:2604.10582v1 Announce Type: new Abstract: Tracking-Any-Point (TAP) models aim to track any point through a video which is a crucial task in AR/XR and robotics applications. The recently introduc

researcharxiv-cs-cv
14 Apr 2026
Safety

TCSA-UDA: Text-Driven Cross-Semantic Alignment for Unsupervised Domain Adaptation in Medical Image Segmentation

DGX agent

arXiv:2511.05782v2 Announce Type: replace Abstract: Unsupervised domain adaptation for medical image segmentation remains a significant challenge due to substantial domain shifts across imaging modali

safetyarxiv-cs-cv
14 Apr 2026
Research

Temporal-Aware Spiking Transformer Hashing Based on 3D-DWT

DGX agent

arXiv:2501.06786v2 Announce Type: replace Abstract: With the rapid growth of dynamic vision sensor (DVS) data, constructing a low-energy, efficient data retrieval system has become an urgent task. Has

researcharxiv-cs-cv
14 Apr 2026
Research

TerraSky3D: Multi-View Reconstructions of European Landmarks in 4K

DGX agent

arXiv:2603.28287v2 Announce Type: replace Abstract: Despite the growing need for data of more and more sophisticated 3D reconstruction pipelines, we can still observe a scarcity of suitable public dat

researcharxiv-cs-cv
14 Apr 2026
Research

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images

DGX agent

arXiv:2604.11025v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) have begun to support Thinking with Images by invoking visual tools such as zooming and cropping during

researcharxiv-cs-cv
14 Apr 2026
Agents

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents

DGX agent

arXiv:2604.09781v1 Announce Type: new Abstract: Vision-Language Models (VLMs) exhibit strong visual reasoning capabilities, yet they still struggle with 3D understanding. In particular, VLMs often fai

agentsarxiv-cs-cv
14 Apr 2026
Model Releases

Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities

DGX agent

arXiv:2504.06313v5 Announce Type: replace Abstract: This paper investigates how popular text-to-image (T2I) models, DALL-E 3 and Gemini 3 Pro Preview, depict people from 206 nationalities when prompte

model-releasesarxiv-cs-cv
14 Apr 2026
Agents

The Devil is in the Details -- From OCR for Old Church Slavonic to Purely Visual Stemma Reconstruction

DGX agent

arXiv:2604.11724v1 Announce Type: new Abstract: The age of artificial intelligence has brought many new possibilities and pitfalls in many fields and tasks. The devil is in the details, and those come

agentsarxiv-cs-cv
14 Apr 2026
Local Ai

The Impact of Federated Learning on Distributed Remote Sensing Archives

DGX agent

arXiv:2604.11562v1 Announce Type: new Abstract: Remote sensing archives are inherently distributed: Earth observation missions such as Sentinel-1, Sentinel-2, and Sentinel-3 have collectively accumula

local-aiarxiv-cs-cv
14 Apr 2026
Applications

The Second Challenge on Real-World Face Restoration at NTIRE 2026: Methods and Results

DGX agent

arXiv:2604.10532v1 Announce Type: new Abstract: This paper provides a review of the NTIRE 2026 challenge on real-world face restoration, highlighting the proposed solutions and the resulting outcomes.

applicationsarxiv-cs-cv
14 Apr 2026
Safety

THOM: Generating Physically Plausible Hand-Object Meshes From Text

DGX agent

arXiv:2604.02736v3 Announce Type: replace Abstract: Generating photorealistic 3D hand-object interactions (HOIs) from text is important for applications like robotic grasping and AR/VR content creatio

safetyarxiv-cs-cv
14 Apr 2026
Model Releases

TinyGaze: Lightweight Gaze-Gesture Recognition on Commodity Mobile Devices

DGX agent

arXiv:2604.09658v1 Announce Type: cross Abstract: Gaze gestures can provide hands free input on mobile devices, but practical use requires (i) gestures users can learn and recall and (ii) recognition

model-releasesarxiv-cs-cv
14 Apr 2026
Model Releases

Topo-ADV: Generating Topology-Driven Imperceptible Adversarial Point Clouds

DGX agent

arXiv:2604.09879v1 Announce Type: new Abstract: Deep neural networks for 3D point cloud understanding have achieved remarkable success in object classification and recognition, yet recent work shows t

model-releasesarxiv-cs-cv
14 Apr 2026
← Previous
1…248249250251252…259
Next →