AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
12 May 2026

KEPIL: Knowledge-Enhanced Prompt-Image Learning for Prompt-Robust Disease Detection

Model ReleasesDGX agent

arXiv:2605.09132v1 Announce Type: new Abstract: Vision--language models (VLMs) show promise for clinical decision support in radiology because they enable joint reasoning over radiological images and

KeyframeFace: Language-Driven Facial Animation via Semantic Keyframes

SafetyDGX agent

arXiv:2512.11321v3 Announce Type: replace Abstract: Facial animation is a core component for creating digital characters in Computer Graphics (CG) industry. A typical production workflow relies on spa

Kinematics-Driven Gaussian Shape Deformation for Blurry Monocular Dynamic Scenes

ApplicationsDGX agent

arXiv:2605.08635v1 Announce Type: new Abstract: Reconstructing dynamic 3D scenes from blurry monocular videos is challenging as motion-induced blur entangles object motion and geometry, hindering geom


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

L2A: Learning to Accumulate Pose History for Accurate 3D Human Pose Estimation

ResearchDGX agent

arXiv:2605.08806v1 Announce Type: new Abstract: Existing 2D-3D lifting human pose estimation methods have achieved strong performance. But the utilization of historical pose representations across net

LCGNav: Local Candidate-Aware Geometric Enhancement for General Topological Planning in Vision-Language Navigation

AgentsDGX agent

arXiv:2605.09053v1 Announce Type: new Abstract: Online topological planning has become an effective paradigm for Vision-Language Navigation in Continuous Environments (VLN-CE), but existing methods st

Learning-Augmented Scalable Linear Assignment Problem Optimization via Neural Dual Warm-Starts

ApplicationsDGX agent

arXiv:2605.09382v1 Announce Type: cross Abstract: The Linear Assignment Problem (LAP) is a fundamental combinatorial optimization task with applications ranging from computer vision to logistics. Clas

Learning to Align Generative Appearance Priors for Fine-grained Image Retrieval

SafetyDGX agent

arXiv:2605.09859v1 Announce Type: new Abstract: Fine-grained image retrieval (FGIR) typically relies on supervision from seen categories to learn discriminative embeddings for retrieving unseen catego

Learning to Perceive 'Where': Spatial Pretext Tasks for Robust Self-Supervised Learning

Model ReleasesDGX agent

arXiv:2605.09963v1 Announce Type: new Abstract: Existing self-supervised learning (SSL) methods primarily learn object-invariant representations but often neglect the spatial structure and relationshi

LightAVSeg: Lightweight Audio-Visual Segmentation

Model ReleasesDGX agent

arXiv:2605.08805v1 Announce Type: new Abstract: Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-mo

LimeCross: Context-Conditioned Layered Image Editing with Structural Consistency

Model ReleasesDGX agent

arXiv:2605.10319v1 Announce Type: new Abstract: Layered image assets are widely used in real-world creative workflows, enabling non-destructive iteration and flexible re-composition. Recent advances i

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?

Model ReleasesDGX agent

arXiv:2605.08985v1 Announce Type: new Abstract: Visual encoding constitutes a major computational bottleneck in Multimodal Large Language Models (MLLMs), especially for high-resolution image inputs. T

Loom: Hybrid Retrieval-Scoring Outfit Recommendation with Semantic Material Compatibility and Occasion-Aware Embedding Priors

ResearchDGX agent

arXiv:2605.09830v1 Announce Type: cross Abstract: We present Loom, an outfit recommendation system that combines neural embedding retrieval with structured domain scoring to generate complete, coheren

Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.08787v1 Announce Type: new Abstract: Recent advances in 3D medical vision-language models have enabled joint reasoning over volumetric images and text, showing strong performance in medical

Low-Cost Neural Radiance Fields

ResearchDGX agent

arXiv:2605.09312v1 Announce Type: new Abstract: Neural Radiance Fields (NeRF) achieve high-quality novel-view synthesis, but their long training times and reliance on dense input views limit accessibi

Low-Cost Stereo Vision for Robust 3D Positioning of Thin Radiata Pine Branches in Autonomous Drone Pruning

AgentsDGX agent

arXiv:2605.08213v1 Announce Type: new Abstract: Manual pruning of radiata pine, a species of major economic importance to New Zealand forestry, is hazardous, labour-intensive, and increasingly constra

M^2E-UAV: A Benchmark and Analysis for Onboard Motion-on-Motion Event-Based Tiny UAV Detection

Model ReleasesDGX agent

arXiv:2605.10496v1 Announce Type: new Abstract: Tiny UAV detection from an onboard event camera is difficult when the observer and target move at the same time. In this motion-on-motion regime, ego-mo

Machine Unlearning on Pre-trained Models by Residual Feature Alignment Using LoRA

SafetyDGX agent

arXiv:2411.08443v2 Announce Type: replace-cross Abstract: Machine unlearning is an emerging technology that removes a subset of the training data from a trained model without significantly affecting t

MAG-VLAQ: Multi-modal Aerial-Ground Query Aggregation for Cross-View Place Recognition

SafetyDGX agent

arXiv:2605.09418v1 Announce Type: new Abstract: Multi-modal cross-view place recognition remains a fundamental challenge in computer vision and robotics due to the severe viewpoint, modality, and spat

Markerless Head Tracking for Accurate and Accessible Neuronavigation

TutorialsDGX agent

arXiv:2602.07052v2 Announce Type: replace Abstract: Neuronavigation is widely used in biomedical research and interventions to guide the precise placement of instruments around the head to support pro

Masked Generative Transformer Is What You Need for Image Editing

ResearchDGX agent

arXiv:2605.10859v1 Announce Type: new Abstract: Diffusion models dominate image editing, yet their global denoising mechanism entangles edited regions with surrounding context, causing modifications t

Measurement-Adapted Eigentask Representations for Photon-Limited Optical Readout

ResearchDGX agent

arXiv:2605.10008v1 Announce Type: cross Abstract: Optical readout in low-light imaging is fundamentally limited by measurement noise, including photon shot noise, detector noise, and quantization erro

Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.10002v1 Announce Type: new Abstract: Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet inco

MedFL-Stress: A Systematic Robustness Evaluation of Federated Brain Tumor Segmentation under Cross-Hospital MRI Appearance Shift

SafetyDGX agent

arXiv:2605.09025v1 Announce Type: new Abstract: Federated learning enables hospitals to collaboratively train segmentation models without sharing patient data. However, current evaluation protocols re

MFVLR: Multi-domain Fine-grained Vision-Language Reconstruction for Generalizable Diffusion Face Forgery Detection and Localization

Local AiDGX agent

arXiv:2605.10071v1 Announce Type: new Abstract: The swift advancement in photo-realistic face generation technology has sparked considerable concerns across society and academia, emphasizing the requi

MicroDiffuse3D: A Foundation Model for 3D Microscopy Imaging Restoration

ResearchDGX agent

arXiv:2605.08566v1 Announce Type: new Abstract: Chemical imaging enables label-free visualization of cells, tissues and living systems while providing direct biochemical information that is difficult

MicroViTv2: Beyond the FLOPS for Edge Energy-Friendly Vision Transformers

ResearchDGX agent

arXiv:2605.10148v1 Announce Type: new Abstract: The Vision Transformer (ViT) achieves remarkable accuracy across visual tasks but remains computationally expensive for edge deployment. This paper pres

ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality

ResearchDGX agent

arXiv:2605.09479v1 Announce Type: cross Abstract: We study full-reference image quality assessment from a machine-centric perspective, where images are evaluated by how well they preserve information

Model-based Dynamic 3D MRI Reconstructions using Neural Fields and Tensor Product Expansions

ResearchDGX agent

arXiv:2605.08275v1 Announce Type: cross Abstract: Conventional MRI reconstruction methods treat images and coil sensitivities as discrete objects, leading to high memory demands and limited structural

Modular Retrieval-Augmented Generalization for Human Action Recognition

ApplicationsDGX agent

arXiv:2605.08117v1 Announce Type: cross Abstract: Inertial Measurement Unit (IMU)-based Human Activity Recognition (HAR) aims to interpret and classify user behaviors from temporal motion signals. Rec

MOTOR-Bench: A Real-world Dataset and Multi-agent Framework for Zero-shot Human Mental State Understanding

Model ReleasesDGX agent

arXiv:2605.09703v1 Announce Type: new Abstract: Understanding human mental states from natural behavior is crucial for intelligent systems in the real world. However, most current research focuses on

MultiAnimate: Pose-Guided Image Animation Made Extensible

ResearchDGX agent

arXiv:2602.21581v2 Announce Type: replace Abstract: Pose-guided human image animation aims to synthesize realistic videos of a reference character driven by a sequence of poses. While diffusion-based

MultiMedVision: Multi-Modal Medical Vision Framework

ResearchDGX agent

arXiv:2605.09151v1 Announce Type: new Abstract: Multi-modal medical imaging enables comprehensive diagnostics, yet current foundation models process 2D (e.g. X-ray) and 3D (e.g. CT) data with separate

Multimodal Emotion Recognition via Causal-Diffusion Bridge (Affect-Diff)

ResearchDGX agent

arXiv:2605.08252v1 Announce Type: new Abstract: Multimodal emotion recognition on CMU-MOSEI faces an extreme imbalance as Happy accounts for 65.9% of samples while three Ekman categories collectively

MUSDA: Multi-source Multi-modality Unsupervised Domain Adaptive 3D Object Detection for Autonomous Driving

AgentsDGX agent

arXiv:2605.10026v1 Announce Type: new Abstract: With the advancement of autonomous driving, numerous annotated multi-modality datasets have become available. This presents an opportunity to develop do

Nano-U: Efficient Terrain Segmentation for Tiny Robot Navigation

AgentsDGX agent

arXiv:2605.10210v1 Announce Type: cross Abstract: Terrain segmentation is a fundamental capability for autonomous mobile robots operating in unstructured outdoor environments. However, state-of-the-ar

NEO: No-Optimization Test-Time Adaptation through Latent Re-Centering

SafetyDGX agent

arXiv:2510.05635v2 Announce Type: replace-cross Abstract: Test-Time Adaptation (TTA) methods are often computationally expensive, require a large amount of data for effective adaptation, or are brittl

Neuromorphic Monocular Depth Estimation with Uncertainty Modeling

ResearchDGX agent

arXiv:2605.10675v1 Announce Type: new Abstract: Event cameras offer distinct advantages over conventional frame-based sensors, including microsecond-level temporal resolution, high dynamic range, and

NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-Identification

ApplicationsDGX agent

arXiv:2505.20001v5 Announce Type: replace Abstract: Multi-modal object Re-IDentification (ReID) aims to obtain complete identity features across heterogeneous modalities. However, most existing method

NICE FACT: Diagnosing and Calibrating VLMs in Quantitative Reasoning for Kinematic Physics

TutorialsDGX agent

arXiv:2605.08452v1 Announce Type: new Abstract: The ability to derive precise spatial and physical insights is a cornerstone of vision-language models (VLMs), yet their poor performances in related sp

Nix and Fix: Targeting 1000x Compression of 3D Gaussian Splatting with Diffusion Models

ResearchDGX agent

arXiv:2602.04549v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) revolutionized novel view rendering. Instead of inferring from dense spatial points, as implicit representations do, 3D

Noise-Started One-Step Real-World Super-Resolution via LR-Conditioned SplitMeanFlow and GAN Refinement

ApplicationsDGX agent

arXiv:2605.09328v1 Announce Type: new Abstract: Pre-trained text-to-image (T2I) diffusion models have shown strong potential for real-world image super-resolution (Real-ISR), owing to their noise-star

Not Blind but Silenced: Rebalancing Vision and Language via Adversarial Counter-Commonsense Equilibrium

ResearchDGX agent

arXiv:2605.10676v1 Announce Type: new Abstract: During MLLM decoding, attention often abnormally concentrates on irrelevant image tokens. While existing research dismisses this as invalid noise and fo

NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces

ApplicationsDGX agent

arXiv:2510.03895v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models represent a pivotal advance in embodied intelligence, yet they confront critical barriers to real-world de

OCP-GN: A Scalable Second-order Optimizer for Stochastic Optimization

ResearchDGX agent

arXiv:2512.24552v2 Announce Type: replace Abstract: This paper proposes a novel second-order optimization algorithm based on the Optimal Control Principle (OCP), applicable to large-scale optimization

Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs

SafetyDGX agent

arXiv:2605.09433v1 Announce Type: new Abstract: Existing preference datasets for text-to-image models typically store only the final winner/loser images. This representation is insufficient for rectif

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

Model ReleasesDGX agent

arXiv:2605.09996v1 Announce Type: new Abstract: While multimodal large language models have advanced across text, image, and audio, personalization research has remained primarily vision-language, wit

On-Policy Distillation with Best-of-N Teacher Rollout Selection

SafetyDGX agent

arXiv:2605.09725v1 Announce Type: new Abstract: On-policy distillation (OPD), which supervises a student on its own sampled trajectories, has emerged as a data-efficient post-training method for impro

On the Generation and Mitigation of Harmful Geometry in Image-to-3D Models

SafetyDGX agent

arXiv:2605.09606v1 Announce Type: cross Abstract: Recent advances in image-to-3D models have significantly improved the fidelity and accessibility of 3D content creation. Such a powerful reconstructio

One-step Latent-free Image Generation with Pixel Mean Flows

ResearchDGX agent

arXiv:2601.22158v3 Announce Type: replace Abstract: Modern diffusion/flow-based models for image generation typically exhibit two core characteristics: (i) using multi-step sampling, and (ii) operatin

Only Train Once: Uncertainty-Aware One-Class Learning for Face Authenticity Detection

ResearchDGX agent

arXiv:2605.10040v1 Announce Type: new Abstract: The rapid evolution of generative paradigms has enabled the creation of highly realistic imagery, which escalating the risks of identity fraud and the d

OpenSGA: Efficient 3D Scene Graph Alignment in the Open World

Model ReleasesDGX agent

arXiv:2605.10484v1 Announce Type: new Abstract: Scene graph alignment establishes object correspondences between two 3D scene graphs constructed from partially overlapping observations. This enables e

OsteoFlow: Lyapunov-Guided Flow Distillation for Predicting Bone Remodeling after Mandibular Reconstruction

ResearchDGX agent

arXiv:2603.22421v2 Announce Type: replace Abstract: Predicting long-term bone remodeling after mandibular reconstruction would be of great clinical benefit, yet standard generative models struggle to

Overcoming Catastrophic Forgetting in Visual Continual Learning with Reinforcement Fine-Tuning

SafetyDGX agent

arXiv:2605.09640v1 Announce Type: new Abstract: Recent studies suggest that Reinforcement Fine-Tuning (RFT) is inherently more resilient to catastrophic forgetting than Supervised Fine-Tuning (SFT). H

OZ-TAL: Online Zero-Shot Temporal Action Localization

ResearchDGX agent

arXiv:2605.09976v1 Announce Type: new Abstract: Online Temporal Action Localization (On-TAL) aims to detect the occurrence time and category of actions in untrimmed streaming videos immediately upon t

P-Flow: Proxy-gradient Flows for Linear Inverse Problems

ResearchDGX agent

arXiv:2605.08328v1 Announce Type: cross Abstract: Generative models based on flow matching have emerged as a powerful paradigm for inverse problems, offering straighter trajectories and faster samplin

PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers

ResearchDGX agent

arXiv:2605.08371v1 Announce Type: new Abstract: Visual Geometry Transformer (VGGT) is a strong feed-forward model for multiple 3D tasks, but its Alternating-Attention (AA) stack scales quadratically i

PaMoSplat: Part-Aware Motion-Guided Gaussian Splatting for Dynamic Scene Reconstruction

TutorialsDGX agent

arXiv:2605.10307v1 Announce Type: new Abstract: Dynamic scene reconstruction represents a fundamental yet demanding challenge in computer vision and robotics. While recent progress in 3DGS-based metho

PermuQuant: Lowering Per-Group Quantization Error by Reordering Channels for Diffusion Models

HardwareDGX agent

arXiv:2605.09503v1 Announce Type: new Abstract: Large-scale visual generative models have achieved remarkable performance. However, their high computational and memory costs make deployment challengin

Personal Visual Context Learning in Large Multimodal Models

Model ReleasesDGX agent

arXiv:2605.10936v1 Announce Type: new Abstract: As wearable devices like smart glasses integrate Large Multimodal Models (LMMs) into the continuous first-person visual streams of individual users, the

PGID: Progressive Guided Inversion and Denoising for Robust Watermark Detection

ResearchDGX agent

arXiv:2605.09319v1 Announce Type: new Abstract: With the proliferation of AI-generated images, digital watermarking has become an essential safeguard for protecting intellectual property and mitigatin

← Previous
1…144145146147148…209
Next →