AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
18 May 2026

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding

Model ReleasesDGX agent

arXiv:2605.15342v1 Announce Type: new Abstract: Video reasoning models are a core component of egocentric and embodied agents. However, standard benchmarks for assessing models provide only evaluation

MorphoHELM: A Comprehensive Benchmark for Evaluating Representations for Microscopy-Based Morphology Assays

Model ReleasesDGX agent

arXiv:2605.15383v1 Announce Type: new Abstract: Microscopy images contain rich information about how cells respond to perturbations, making them essential to applications like drug screening. To quant

mRadNet: A Compact Radar Object Detector with MetaFormer

ResearchDGX agent

arXiv:2509.16223v3 Announce Type: replace-cross Abstract: Frequency-modulated continuous wave radars have gained increasing popularity in the automotive industry. Their robustness against adverse weat


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Multimodal Object Detection Under Sparse Forest-Canopy Occlusion

ResearchDGX agent

arXiv:2605.15326v1 Announce Type: new Abstract: Reliable detection of humans beneath forest canopy remains a difficult remote-sensing challenge due to sparse, structured, and viewpoint-dependent occlu

Navigating the Challenges of AI-Generated Image Detection in the Wild: What Truly Matters?

Model ReleasesDGX agent

arXiv:2507.10236v2 Announce Type: replace Abstract: As generative Artificial Intelligence (AI) advances, the realism of AI generated imagery has reached a threshold capable of deceiving even vigilant

Neurosymbolic Object-Centric Learning with Distant Supervision

ResearchDGX agent

arXiv:2506.16129v2 Announce Type: replace Abstract: Neurosymbolic learning can use symbolic rules to provide supervision for latent concepts from weak labels, but it commonly assumes that the entities

Neutral-Reference Prompting for Vision-Language Models

SafetyDGX agent

arXiv:2605.15615v1 Announce Type: new Abstract: Efficient transfer learning of vision-language models (VLMs) commonly suffers from a Base-New Trade-off (BNT): improving performance on unseen (new) cla

Not All Tasks Quantize Equally: Fisher-Guided Quantization for Visual Geometry Transformer

Local AiDGX agent

arXiv:2605.15828v1 Announce Type: new Abstract: Feed-forward 3D reconstruction models, represented by Visual Geometry Grounded Transformer (VGGT), jointly predict multiple visual geometry tasks such a

On RGB-TIR Stereo Calibration under Extreme Resolution Asymmetry

Model ReleasesDGX agent

arXiv:2605.15860v1 Announce Type: new Abstract: Accurate geometric calibration of RGB-thermal infrared (TIR) stereo camera systems is essential for multimodal building envelope analysis, yet remains c

One Pass Is Not Enough: Recursive Latent Refinement for Generative Models

ResearchDGX agent

arXiv:2605.15309v1 Announce Type: new Abstract: Despite remarkable progress, image generation is far from solved. The dominant metric, FID, conflates sample fidelity with mode coverage and is close to

OpenFrontier: General Navigation with Visual-Language Grounded Frontiers

SafetyDGX agent

arXiv:2603.05377v2 Announce Type: replace-cross Abstract: Open-world navigation requires robots to make decisions in complex everyday environments while adapting to flexible task requirements. Convent

Overlap-aware segmentation for topological reconstruction of obscured objects

ResearchDGX agent

arXiv:2510.06194v2 Announce Type: replace-cross Abstract: The separation of overlapping objects presents a significant challenge in scientific imaging. While deep learning segmentation-regression algo

PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models

Model ReleasesDGX agent

arXiv:2512.01843v2 Announce Type: replace Abstract: Driven by the growing capacity and training scale, Text-to-Video (T2V) generation models have recently achieved substantial progress in video qualit

Preprocessing Algorithm Leveraging Geometric Modeling for Scale Correction in Hyperspectral Images for Improved Unmixing Performance

ApplicationsDGX agent

arXiv:2508.08431v3 Announce Type: replace-cross Abstract: Spectral variability significantly impacts the accuracy and convergence of hyperspectral unmixing algorithms. Many methods address complex spe

Probabilistic Dating of Historical Manuscripts via Evidential Deep Regression on Visual Script Features

Model ReleasesDGX agent

arXiv:2605.06475v1 Announce Type: cross Abstract: We introduce a probabilistic approach for dating historical manuscript pages from visual features alone. Instead of aggregating centuries into classes

ReactiveGWM: Steering NPC in Reactive Game World Models

SafetyDGX agent

arXiv:2605.15256v1 Announce Type: new Abstract: Current game world models simulate environments from a subjective, player-centric perspective. However, by treating the Non-Player Character (NPC) merel

ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation

SafetyDGX agent

arXiv:2605.16080v1 Announce Type: new Abstract: The rise of AI-generated images (AIGIs) poses growing challenges for digital authenticity, prompting the need for efficient, generalizable image forgery

RealRep: Generalized SDR-to-HDR Conversion via Attribute-Disentangled Representation Learning

ApplicationsDGX agent

arXiv:2505.07322v4 Announce Type: replace Abstract: High-Dynamic-Range Wide-Color-Gamut (HDR-WCG) technology is becoming increasingly widespread, driving a growing need for converting Standard Dynamic

Registers Matter for Pixel-Space Diffusion Transformers

Model ReleasesDGX agent

arXiv:2605.16147v1 Announce Type: new Abstract: Vision Transformers (ViTs) are known to exhibit high-norm patch-token outliers that degrade feature map quality, a problem effectively mitigated by exti

Res^2CLIP: Few-Shot Generalist Anomaly Detection with Residual-to-Residual Alignment

SafetyDGX agent

arXiv:2605.16171v1 Announce Type: new Abstract: Few-shot Generalist Anomaly Detection requires models to generalize to novel categories without retraining, posing significant challenges in real-world

Reversing the Flow: Generation-to-Understanding Synergy in Large Multimodal Models

SafetyDGX agent

arXiv:2605.15792v1 Announce Type: new Abstract: The long-standing goal of multimodal AI is to build unified models in which visual understanding and visual generation mutually enhance one another. Des

RoiMAM: Region-of-Interest Medical Attention Model for Efficient Vision-Language Understanding

ResearchDGX agent

arXiv:2605.15561v1 Announce Type: new Abstract: Vision-Language Models (VLMs) facilitate medical visual question answering (MedVQA) by jointly interpreting images and text. However, existing models ty

Seeing What Matters: Visual Preference Policy Optimization for Visual Generation

SafetyDGX agent

arXiv:2511.18719v4 Announce Type: replace Abstract: Reinforcement learning (RL) has become a powerful tool for post-training visual generative models, with Group Relative Policy Optimization (GRPO) in

Segmentation, Detection and Explanation: A Unified Framework for CT Appearance Reasoning

Local AiDGX agent

arXiv:2605.15997v1 Announce Type: new Abstract: Recent progress in deep learning has significantly advanced CT image analysis, particularly for segmentation tasks. However, these advances are largely

Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning

ResearchDGX agent

arXiv:2605.15523v1 Announce Type: new Abstract: Scene text editing aims to modify text in a target region of an image while preserving surrounding background style and texture. Existing methods rely s

Self-Supervised ImageNet Representations for In Vivo Confocal Microscopy: Tortuosity Grading without Segmentation Maps

ResearchDGX agent

arXiv:2603.15269v2 Announce Type: replace Abstract: The tortuosity of corneal nerve fibers are used as indication for different diseases. Current state-of-the-art methods for grading the tortuosity he

Self-Supervised Learning by Curvature Alignment

SafetyDGX agent

arXiv:2511.17426v2 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) has recently advanced through non-contrastive methods that couple an invariance term with variance, covariance,

Semi-MedRef: Semi-Supervised Medical Referring Image Segmentation with Cross-Modal Alignment

SafetyDGX agent

arXiv:2605.15720v1 Announce Type: new Abstract: Medical referring image segmentation (MRIS) requires pixel-level masks aligned with textual descriptions of anatomical locations, making annotation cost

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting

ResearchDGX agent

arXiv:2511.18127v2 Announce Type: replace Abstract: Real-time 3D hand forecasting is a critical component for fluid human-computer interaction in applications like AR and assistive robotics. However,

SkyLink: A Large Vision-Language Model Driven Re-ranking Framework for Cross-View UAV geolocalization

Model ReleasesDGX agent

arXiv:2603.08063v3 Announce Type: replace Abstract: Cross-view UAV geolocalization is fundamentally a challenging large-scale image retrieval task, aiming to determine the geographic coordinates of Un

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning

Model ReleasesDGX agent

arXiv:2512.15693v2 Announce Type: replace Abstract: The misuse of AI-driven video generation technologies has raised serious social concerns, highlighting the urgent need for reliable AI-generated vid

Social-Mamba: Socially-Aware Trajectory Forecasting with State-Space Models

Model ReleasesDGX agent

arXiv:2605.15424v1 Announce Type: new Abstract: Human trajectory forecasting is crucial for safe navigation in crowded environments, requiring models that balance accuracy with computational efficienc

SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval

Model ReleasesDGX agent

arXiv:2605.15868v1 Announce Type: new Abstract: In this work, we address the critical yet underexplored challenge of symmetric multimodal-to-multimodal (MM2MM) retrieval, where queries and contexts ar

Sound Sparks Motion: Audio and Text Tuning for Video Editing

Local AiDGX agent

arXiv:2605.15307v1 Announce Type: cross Abstract: Motion-centric video editing remains difficult for large generative video models, which often respond well to appearance changes but struggle to produ

Sparse ActionGen: Accelerating Diffusion Policy with Real-time Pruning

Model ReleasesDGX agent

arXiv:2601.12894v2 Announce Type: replace-cross Abstract: Diffusion Policy has dominated action generation due to its strong capabilities for modeling multi-modal action distributions, but its multi-s

Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models

ResearchDGX agent

arXiv:2605.15961v1 Announce Type: new Abstract: Large-scale pre-trained vision-language models like CLIP demonstrate remarkable zero-shot performance across diverse tasks. However, fine-tuning these m

STABLE: Simulation-Ready Tabletop Layout Generation via a Semantics-Physics Dual System

SafetyDGX agent

arXiv:2605.16137v1 Announce Type: new Abstract: Generating simulation-ready tabletop scenes from task instructions is an intriguing and promising research direction in the field of Embodied AI. Howeve

StippleDiffusion: Capacity-Constrained Stippling using Controlled Diffusion

Model ReleasesDGX agent

arXiv:2605.15816v1 Announce Type: cross Abstract: Stipple patterns, point sets whose local density tracks a target image, are traditionally produced by per-density iterative optimizers, which are slow

SynthRender and IRIS: Open-Source Framework and Dataset for Bidirectional Sim-Real Transfer in Industrial Object Perception

Model ReleasesDGX agent

arXiv:2602.21141v2 Announce Type: replace Abstract: Object perception is fundamental for tasks such as robotic material handling and quality inspection. However, modern supervised deep-learning models

Text-RSIR: A Text-Guided Framework for Efficient Remote Sensing Image Transmission and Reconstruction

ResearchDGX agent

arXiv:2605.15558v1 Announce Type: cross Abstract: High-resolution remote sensing imagery is critical for environmental monitoring, urban mapping, and land cover analysis, but its transmission is often

TSBOW -- Traffic Surveillance Benchmark for Occluded Vehicles Under Various Weather Conditions

Model ReleasesDGX agent

arXiv:2602.05414v2 Announce Type: replace Abstract: Global warming has intensified the frequency and severity of extreme weather events, which degrade CCTV signal and video quality while disrupting tr

TVRN: Invertible Neural Networks for Compression-Aware Temporal Video Rescaling

ApplicationsDGX agent

arXiv:2605.15579v1 Announce Type: cross Abstract: To fit diverse display and bandwidth constraints, high-frame-rate videos are temporally downscaled to low-frame-rate (LFR) and later upscaled, requiri

U-SEG: Uncertainty in SEGmentation -- A systematic multi-variable exploration

ResearchDGX agent

arXiv:2605.15421v1 Announce Type: new Abstract: In this study, we explore in depth a few under-studied topics at the intersection of uncertainty estimation and segmentation. Prior work has shown that

Unlocking Dense Metric Depth Estimation in VLMs

Model ReleasesDGX agent

arXiv:2605.15876v1 Announce Type: new Abstract: Vision-Language Models (VLMs) excel at 2D tasks such as grounding and captioning, yet remain limited in 3D understanding. A key limitation is their text

Unsupervised 3D Human Pose Estimation via Conditional Multi-view Ancestral Sampling

ResearchDGX agent

arXiv:2605.15583v1 Announce Type: new Abstract: We propose a method of estimating a 3D human pose from a single view without 3D supervision. The key to our method is to leverage the 2D diffusion prior

Video Models Can Reason with Verifiable Rewards

SafetyDGX agent

arXiv:2605.15458v1 Announce Type: new Abstract: Video diffusion models have made rapid progress in perceptual realism and temporal coherence, but they remain primarily optimized for plausible generati

VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?

Model ReleasesDGX agent

arXiv:2510.08398v4 Announce Type: replace Abstract: The recent rapid advancement of Text-to-Video (T2V) generation technologies are engaging the trained models with more world model ability, making th

ViewBridge: Curriculum Knowledge Distillation for Activity View-Invariance Under Extreme Viewpoint Changes

ResearchDGX agent

arXiv:2504.05451v2 Announce Type: replace Abstract: Traditional methods for view-invariant learning rely on controlled multi-view training data with minimal scene clutter. However, they struggle with

Visual Compositional Tuning

ResearchDGX agent

arXiv:2504.21850v3 Announce Type: replace Abstract: Visual instruction tuning (VIT) datasets have grown rapidly in scale, yet the informativeness of individual training samples has largely been overlo

WeatherOcc3D: VLM-Assisted Adverse Weather Aware 3D Semantic Occupancy Prediction

Model ReleasesDGX agent

arXiv:2605.16127v1 Announce Type: new Abstract: While multi-modal 3D semantic occupancy prediction typically enhances robustness by fusing camera and LiDAR inputs, its effectiveness is fundamentally c

When Does Sparse MoE Help in Vision? The Role of Backbone Compute Leverage in Sparse Routing

ResearchDGX agent

arXiv:2605.15484v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) networks promise favorable accuracy-compute trade-offs, yet practical vision deployments are hindered by expert collapse and li

Where to Perch in a Tree: Vision-Guidance for Tree-Grasping Drones

AgentsDGX agent

arXiv:2605.15430v1 Announce Type: cross Abstract: This study demonstrates a method to locate an ideal perch location on a tree for vision-guided autonomous tree-perching drones. Various image processi

WorldAct: Activating Monolithic 3D Worlds into Interactive-Ready Object-Centric Scenes

AgentsDGX agent

arXiv:2605.15843v1 Announce Type: new Abstract: Recent 3D world modeling systems based on generative scene synthesis, such as Marble, can create coherent and explorable 3D environments, yet their outp

WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation

AgentsDGX agent

arXiv:2605.15964v1 Announce Type: cross Abstract: Aerial vision-language navigation (VLN) requires agents to follow natural-language instructions through closed-loop perception and action in 3D enviro

15 May 2026

3D Skew-Normal Splatting

Model ReleasesDGX agent

arXiv:2605.15010v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has emerged as a leading representation for real-time novel view synthesis and been widely adopted in various downstream ap

A CUBS-Compatible Ultrasound Morphology and Uncertainty-Aware Baseline for Carotid Intima-Media Segmentation and Preliminary Risk Prediction

ResearchDGX agent

arXiv:2605.14949v1 Announce Type: new Abstract: Carotid atherosclerosis is a major contributor to ischemic stroke and transient ischemic attack. Conventional ultrasound assessment is commonly based on

ACE-LoRA: Adaptive Orthogonal Decoupling for Continual Image Editing

Model ReleasesDGX agent

arXiv:2605.14948v1 Announce Type: new Abstract: State-of-the-art diffusion models often rely on parameter-efficient fine-tuning to perform specialized image editing tasks. However, real-world applicat

Aligning Latent Geometry for Spherical Flow Matching in Image Generation

SafetyDGX agent

arXiv:2605.15193v1 Announce Type: new Abstract: Latent flow matching for image generation usually transports Gaussian noise to variational autoencoder latents along linear paths. Both endpoints, howev

Analogical Trajectory Transfer

ResearchDGX agent

arXiv:2605.14393v1 Announce Type: new Abstract: We study analogical trajectory transfer, where the goal is to translate motion trajectories in one 3D environment to a semantically analogous location i

AnchorRoute: Human Motion Synthesis with Interval-Routed Sparse Contro

Model ReleasesDGX agent

arXiv:2605.14716v1 Announce Type: cross Abstract: Sparse anchors provide a compact interface for human motion authoring: users specify a few root positions, planar trajectory samples, or body-point ta

← Previous
1…133134135136137…211
Next →