AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
14 Apr 2026

RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought

SafetyDGX agent

arXiv:2506.16796v4 Announce Type: replace Abstract: Real-World Image Super-Resolution is one of the most challenging task in image restoration. However, existing methods struggle with an accurate unde

ReContraster: Making Your Posters Stand Out with Regional Contrast

Model ReleasesDGX agent

arXiv:2604.10442v1 Announce Type: new Abstract: Effective poster design requires rapidly capturing attention and clearly conveying messages. Inspired by the ``contrast effects'' principle, we propose

Rein3D: Reinforced 3D Indoor Scene Generation with Panoramic Video Diffusion Models

Local AiDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.10578v1 Announce Type: new Abstract: The growing demand for Embodied AI and VR applications has highlighted the need for synthesizing high-quality 3D indoor scenes from sparse inputs. Howev

ReplicateAnyScene: Zero-Shot Video-to-3D Composition via Textual-Visual-Spatial Alignment

Model ReleasesDGX agent

arXiv:2604.10789v1 Announce Type: new Abstract: Humans exhibit an innate capacity to rapidly perceive and segment objects from video observations, and even mentally assemble them into structured 3D sc

Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction

AgentsDGX agent

arXiv:2604.11707v1 Announce Type: new Abstract: Accurate future video prediction requires both high visual fidelity and consistent scene semantics, particularly in complex dynamic environments such as

RESP: Reference-guided Sequential Prompting for Visual Glitch Detection in Video Games

ApplicationsDGX agent

arXiv:2604.11082v1 Announce Type: new Abstract: Visual glitches in video games degrade player experience and perceived quality, yet manual quality assurance cannot scale to the growing test surface of

Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent Purification

ApplicationsDGX agent

arXiv:2604.10695v1 Announce Type: new Abstract: Recent Audio-Visual Question Answering (AVQA) methods have advanced significantly. However, most AVQA methods lack effective mechanisms for handling mis

Revisiting the Scale Loss Function and Gaussian-Shape Convolution for Infrared Small Target Detection

Model ReleasesDGX agent

arXiv:2604.09991v1 Announce Type: new Abstract: Infrared small target detection still faces two persistent challenges: training instability from non-monotonic scale loss functions, and inadequate spat

RL makes MLLMs see better than SFT

Model ReleasesDGX agent

arXiv:2510.16333v2 Announce Type: replace Abstract: A dominant assumption in Multimodal Language Model (MLLM) research is that its performance is largely inherited from the LLM backbone, given its imm

RoboStereo: Dual-Tower 4D Embodied World Models for Unified Policy Optimization

SafetyDGX agent

arXiv:2603.12639v2 Announce Type: replace Abstract: Scalable Embodied AI faces fundamental constraints due to prohibitive costs and safety risks of real-world interaction. While Embodied World Models

Robust Fair Disease Diagnosis in CT Images

Model ReleasesDGX agent

arXiv:2604.09710v1 Announce Type: new Abstract: Automated diagnosis from chest CT has improved considerably with deep learning, but models trained on skewed datasets tend to perform unevenly across pa

RobustMedSAM: Degradation-Resilient Medical Image Segmentation via Robust Foundation Model Adaptation

Model ReleasesDGX agent

arXiv:2604.09814v1 Announce Type: new Abstract: Medical image segmentation models built on Segment Anything Model (SAM) achieve strong performance on clean benchmarks, yet their reliability often degr

RobustSpring: Benchmarking Robustness to Image Corruptions for Optical Flow, Scene Flow and Stereo

Model ReleasesDGX agent

arXiv:2505.09368v2 Announce Type: replace Abstract: Standard benchmarks for optical flow, scene flow, and stereo vision algorithms generally focus on model accuracy rather than robustness to image cor

rPPG-VQA: A Video Quality Assessment Framework for Unsupervised rPPG Training

ResearchDGX agent

arXiv:2604.11156v1 Announce Type: new Abstract: Unsupervised remote photoplethysmography (rPPG) promises to leverage unlabeled video data, but its potential is hindered by a critical challenge: traini

S4M: 4-points to Segment Anything

ResearchDGX agent

arXiv:2503.05534v3 Announce Type: replace Abstract: Purpose: The Segment Anything Model (SAM) promises to ease the annotation bottleneck in medical segmentation, but overlapping anatomy and blurred bo

SatReg: Regression-based Neural Architecture Search for Lightweight Satellite Image Segmentation

HardwareDGX agent

arXiv:2604.10306v1 Announce Type: new Abstract: As Earth-observation workloads move toward onboard and edge processing, remote-sensing segmentation models must operate under tight latency and energy c

Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes

ResearchDGX agent

arXiv:2512.12982v2 Announce Type: replace Abstract: The pursuit of a universal AI-generated image (AIGI) detector often relies on aggregating data from numerous generators to improve generalization. H

Scene Change Detection with Vision-Language Representation Learning

ApplicationsDGX agent

arXiv:2604.11402v1 Announce Type: new Abstract: Scene change detection (SCD) is crucial for urban monitoring and navigation but remains challenging in real-world environments due to lighting variation

SciPostLayoutTree: A Dataset for Structural Analysis of Scientific Posters

ResearchDGX agent

arXiv:2511.18329v3 Announce Type: replace Abstract: Scientific posters play a vital role in academic communication by presenting ideas through visual summaries. Analyzing reading order and parent-chil

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding

Model ReleasesDGX agent

arXiv:2604.11244v1 Announce Type: new Abstract: Advances in Multimodal Large Language Models (MLLMs) are transforming video captioning from a descriptive endpoint into a semantic interface for both vi

Search-MIND: Training-Free Multi-Modal Medical Image Registration

Local AiDGX agent

arXiv:2604.09743v1 Announce Type: cross Abstract: Multi-modal image registration plays a critical role in precision medicine but faces challenges from non-linear intensity relationships and local opti

See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment

SafetyDGX agent

arXiv:2604.09749v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) frequently hallucinate objects that are absent from the visual input, often because attention during decoding i

Seeing Through the Tool: A Controlled Benchmark for Occlusion Robustness in Foundation Segmentation Models

Model ReleasesDGX agent

arXiv:2604.11711v1 Announce Type: new Abstract: Occlusion, where target structures are partially hidden by surgical instruments or overlapping tissues, remains a critical yet underexplored challenge f

Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions

Local AiDGX agent

arXiv:2604.11579v1 Announce Type: new Abstract: We address the problem of tactile localization, where the goal is to identify image regions that share the same material properties as a tactile input.

Seg2Change: Adapting Open-Vocabulary Semantic Segmentation Model for Remote Sensing Change Detection

Model ReleasesDGX agent

arXiv:2604.11231v1 Announce Type: new Abstract: Change detection is a fundamental task in remote sensing, aiming to quantify the impacts of human activities and ecological dynamics on land-cover chang

Self-supervised Pretraining of Cell Segmentation Models

Model ReleasesDGX agent

arXiv:2604.10609v1 Announce Type: new Abstract: Instance segmentation enables the analysis of spatial and temporal properties of cells in microscopy images by identifying the pixels belonging to each

Sharpness-Aware Surrogate Training for On-Sensor Spiking Neural Networks

ResearchDGX agent

arXiv:2604.09696v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) are a natural computational model for on-sensor and near-sensor vision, where event driven processors must operate unde

SignReasoner: Compositional Reasoning for Complex Traffic Sign Understanding via Functional Structure Units

Model ReleasesDGX agent

arXiv:2604.10436v1 Announce Type: new Abstract: Accurate semantic understanding of complex traffic signs-including those with intricate layouts, multi-lingual text, and composite symbols-is critical f

SIMPLER: H&E-Informed Representation Learning for Structured Illumination Microscopy

SafetyDGX agent

arXiv:2604.10334v1 Announce Type: new Abstract: Structured Illumination Microscopy (SIM) enables rapid, high-contrast optical sectioning of fresh tissue without staining or physical sectioning, making

SinkTrack: Attention Sink based Context Anchoring for Large Language Models

ResearchDGX agent

arXiv:2604.10027v1 Announce Type: new Abstract: Large language models (LLMs) suffer from hallucination and context forgetting. Prior studies suggest that attention drift is a primary cause of these pr

SMFormer: Empowering Self-supervised Stereo Matching via Foundation Models and Data Augmentation

Model ReleasesDGX agent

arXiv:2604.10218v1 Announce Type: new Abstract: Recent self-supervised stereo matching methods have made significant progress. They typically rely on the photometric consistency assumption, which pres

Sparse Hypergraph-Enhanced Frame-Event Object Detection with Fine-Grained MoE

ResearchDGX agent

arXiv:2604.11140v1 Announce Type: new Abstract: Integrating frame-based RGB cameras with event streams offers a promising solution for robust object detection under challenging dynamic conditions. How

Spatio-Temporal Difference Guided Motion Deblurring with the Complementary Vision Sensor

ApplicationsDGX agent

arXiv:2604.10554v1 Announce Type: new Abstract: Motion blur arises when rapid scene changes occur during the exposure period, collapsing rich intra-exposure motion into a single RGB frame. Without exp

Specificity-aware reinforcement learning for fine-grained open-world classification

TutorialsDGX agent

arXiv:2603.03197v3 Announce Type: replace Abstract: Classifying fine-grained visual concepts under open-world settings, i.e., without a predefined label set, demands models to be both accurate and spe

SpotFormer: Multi-Scale Spatio-Temporal Transformer for Facial Expression Spotting

Local AiDGX agent

arXiv:2407.20799v3 Announce Type: replace Abstract: Facial expression spotting, identifying periods where facial expressions occur in a video, is a significant yet challenging task in facial expressio

Spotlight and Shadow: Attention-Guided Dual-Anchor Introspective Decoding for MLLM Hallucination Mitigation

TutorialsDGX agent

arXiv:2604.10071v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities yet continue to suffer from hallucination, where generated

StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation

SafetyDGX agent

arXiv:2510.05057v2 Announce Type: replace-cross Abstract: A fundamental challenge in embodied intelligence is developing expressive and compact state representations for efficient world modeling and d

STGV: Spatio-Temporal Hash Encoding for Gaussian-based Video Representation

ResearchDGX agent

arXiv:2604.10910v1 Announce Type: new Abstract: 2D Gaussian Splatting (2DGS) has recently become a promising paradigm for high-quality video representation. However, existing methods employ content-ag

Structured State-Space Regularization for Compact and Generation-Friendly Image Tokenization

TutorialsDGX agent

arXiv:2604.11089v1 Announce Type: new Abstract: Image tokenizers are central to modern vision models as they often operate in latent spaces. An ideal latent space must be simultaneously compact and ge

STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding

ResearchDGX agent

arXiv:2604.11637v1 Announce Type: new Abstract: 4D point cloud videos capture rich spatial and temporal dynamics of scenes which possess unique values in various 4D understanding tasks. However, most

SwinTextUNet: Integrating CLIP-Based Text Guidance into Swin Transformer U-Nets for Medical Image Segmentation

ResearchDGX agent

arXiv:2604.10000v1 Announce Type: new Abstract: Precise medical image segmentation is fundamental for enabling computer aided diagnosis and effective treatment planning. Traditional models that rely s

Switch-JustDance: Benchmarking Whole Body Motion Tracking Controllers Using a Commercial Console Game

Model ReleasesDGX agent

arXiv:2511.17925v3 Announce Type: replace-cross Abstract: Recent advances in whole-body robot control have enabled humanoid and legged robots to perform increasingly agile and coordinated motions. How

SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation

Model ReleasesDGX agent

arXiv:2604.03723v2 Announce Type: replace Abstract: Controlling both camera motion and object dynamics is essential for coherent and expressive video generation, yet current methods typically handle o

SyncFix: Fixing 3D Reconstructions via Multi-View Synchronization

ResearchDGX agent

arXiv:2604.11797v1 Announce Type: new Abstract: We present SyncFix, a framework that enforces cross-view consistency during the diffusion-based refinement of reconstructed scenes. SyncFix formulates r

TAG-Head: Time-Aligned Graph Head for Plug-and-Play Fine-grained Action Recognition

Model ReleasesDGX agent

arXiv:2604.11498v1 Announce Type: new Abstract: Fine-grained human action recognition (FHAR) is challenging because visually similar actions differ by subtle spatio-temporal cues. Many recent systems

TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering

ResearchDGX agent

arXiv:2509.04123v2 Announce Type: replace Abstract: Text-to-story visualization is challenging due to the need for consistent interaction among multiple characters across frames. Existing methods stru

TAMISeg: Text-Aligned Multi-scale Medical Image Segmentation with Semantic Encoder Distillation

ResearchDGX agent

arXiv:2604.10912v1 Announce Type: new Abstract: Medical image segmentation remains challenging due to limited fine-grained annotations, complex anatomical structures, and image degradation from noise,

TAPNext++: What's Next for Tracking Any Point (TAP)?

ResearchDGX agent

arXiv:2604.10582v1 Announce Type: new Abstract: Tracking-Any-Point (TAP) models aim to track any point through a video which is a crucial task in AR/XR and robotics applications. The recently introduc

TCSA-UDA: Text-Driven Cross-Semantic Alignment for Unsupervised Domain Adaptation in Medical Image Segmentation

SafetyDGX agent

arXiv:2511.05782v2 Announce Type: replace Abstract: Unsupervised domain adaptation for medical image segmentation remains a significant challenge due to substantial domain shifts across imaging modali

Temporal-Aware Spiking Transformer Hashing Based on 3D-DWT

ResearchDGX agent

arXiv:2501.06786v2 Announce Type: replace Abstract: With the rapid growth of dynamic vision sensor (DVS) data, constructing a low-energy, efficient data retrieval system has become an urgent task. Has

TerraSky3D: Multi-View Reconstructions of European Landmarks in 4K

ResearchDGX agent

arXiv:2603.28287v2 Announce Type: replace Abstract: Despite the growing need for data of more and more sophisticated 3D reconstruction pipelines, we can still observe a scarcity of suitable public dat

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images

ResearchDGX agent

arXiv:2604.11025v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) have begun to support Thinking with Images by invoking visual tools such as zooming and cropping during

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents

AgentsDGX agent

arXiv:2604.09781v1 Announce Type: new Abstract: Vision-Language Models (VLMs) exhibit strong visual reasoning capabilities, yet they still struggle with 3D understanding. In particular, VLMs often fai

Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities

Model ReleasesDGX agent

arXiv:2504.06313v5 Announce Type: replace Abstract: This paper investigates how popular text-to-image (T2I) models, DALL-E 3 and Gemini 3 Pro Preview, depict people from 206 nationalities when prompte

The Devil is in the Details -- From OCR for Old Church Slavonic to Purely Visual Stemma Reconstruction

AgentsDGX agent

arXiv:2604.11724v1 Announce Type: new Abstract: The age of artificial intelligence has brought many new possibilities and pitfalls in many fields and tasks. The devil is in the details, and those come

The Impact of Federated Learning on Distributed Remote Sensing Archives

Local AiDGX agent

arXiv:2604.11562v1 Announce Type: new Abstract: Remote sensing archives are inherently distributed: Earth observation missions such as Sentinel-1, Sentinel-2, and Sentinel-3 have collectively accumula

The Second Challenge on Real-World Face Restoration at NTIRE 2026: Methods and Results

ApplicationsDGX agent

arXiv:2604.10532v1 Announce Type: new Abstract: This paper provides a review of the NTIRE 2026 challenge on real-world face restoration, highlighting the proposed solutions and the resulting outcomes.

THOM: Generating Physically Plausible Hand-Object Meshes From Text

SafetyDGX agent

arXiv:2604.02736v3 Announce Type: replace Abstract: Generating photorealistic 3D hand-object interactions (HOIs) from text is important for applications like robotic grasping and AR/VR content creatio

TinyGaze: Lightweight Gaze-Gesture Recognition on Commodity Mobile Devices

Model ReleasesDGX agent

arXiv:2604.09658v1 Announce Type: cross Abstract: Gaze gestures can provide hands free input on mobile devices, but practical use requires (i) gestures users can learn and recall and (ii) recognition

Topo-ADV: Generating Topology-Driven Imperceptible Adversarial Point Clouds

Model ReleasesDGX agent

arXiv:2604.09879v1 Announce Type: new Abstract: Deep neural networks for 3D point cloud understanding have achieved remarkable success in object classification and recognition, yet recent work shows t

← Previous
1…198199200201202…207
Next →