AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlog
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

MOTOR-Bench: A Real-world Dataset and Multi-agent Framework for Zero-shot Human Mental State Understanding

DGX agent

arXiv:2605.09703v1 Announce Type: new Abstract: Understanding human mental states from natural behavior is crucial for intelligent systems in the real world. However, most current research focuses on

model-releasesarxiv-cs-cv
12 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

MultiAnimate: Pose-Guided Image Animation Made Extensible

DGX agent

arXiv:2602.21581v2 Announce Type: replace Abstract: Pose-guided human image animation aims to synthesize realistic videos of a reference character driven by a sequence of poses. While diffusion-based

researcharxiv-cs-cv
12 May 2026
Research

MultiMedVision: Multi-Modal Medical Vision Framework

DGX agent

arXiv:2605.09151v1 Announce Type: new Abstract: Multi-modal medical imaging enables comprehensive diagnostics, yet current foundation models process 2D (e.g. X-ray) and 3D (e.g. CT) data with separate

researcharxiv-cs-cv
12 May 2026
Research

Multimodal Emotion Recognition via Causal-Diffusion Bridge (Affect-Diff)

DGX agent

arXiv:2605.08252v1 Announce Type: new Abstract: Multimodal emotion recognition on CMU-MOSEI faces an extreme imbalance as Happy accounts for 65.9% of samples while three Ekman categories collectively

researcharxiv-cs-cv
12 May 2026
Agents

MUSDA: Multi-source Multi-modality Unsupervised Domain Adaptive 3D Object Detection for Autonomous Driving

DGX agent

arXiv:2605.10026v1 Announce Type: new Abstract: With the advancement of autonomous driving, numerous annotated multi-modality datasets have become available. This presents an opportunity to develop do

agentsarxiv-cs-cv
12 May 2026
Agents

Nano-U: Efficient Terrain Segmentation for Tiny Robot Navigation

DGX agent

arXiv:2605.10210v1 Announce Type: cross Abstract: Terrain segmentation is a fundamental capability for autonomous mobile robots operating in unstructured outdoor environments. However, state-of-the-ar

agentsarxiv-cs-cv
12 May 2026
Safety

NEO: No-Optimization Test-Time Adaptation through Latent Re-Centering

DGX agent

arXiv:2510.05635v2 Announce Type: replace-cross Abstract: Test-Time Adaptation (TTA) methods are often computationally expensive, require a large amount of data for effective adaptation, or are brittl

safetyarxiv-cs-cv
12 May 2026
Research

Neuromorphic Monocular Depth Estimation with Uncertainty Modeling

DGX agent

arXiv:2605.10675v1 Announce Type: new Abstract: Event cameras offer distinct advantages over conventional frame-based sensors, including microsecond-level temporal resolution, high dynamic range, and

researcharxiv-cs-cv
12 May 2026
Applications

NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-Identification

DGX agent

arXiv:2505.20001v5 Announce Type: replace Abstract: Multi-modal object Re-IDentification (ReID) aims to obtain complete identity features across heterogeneous modalities. However, most existing method

applicationsarxiv-cs-cv
12 May 2026
Tutorials

NICE FACT: Diagnosing and Calibrating VLMs in Quantitative Reasoning for Kinematic Physics

DGX agent

arXiv:2605.08452v1 Announce Type: new Abstract: The ability to derive precise spatial and physical insights is a cornerstone of vision-language models (VLMs), yet their poor performances in related sp

tutorialsarxiv-cs-cv
12 May 2026
Research

Nix and Fix: Targeting 1000x Compression of 3D Gaussian Splatting with Diffusion Models

DGX agent

arXiv:2602.04549v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) revolutionized novel view rendering. Instead of inferring from dense spatial points, as implicit representations do, 3D

researcharxiv-cs-cv
12 May 2026
Applications

Noise-Started One-Step Real-World Super-Resolution via LR-Conditioned SplitMeanFlow and GAN Refinement

DGX agent

arXiv:2605.09328v1 Announce Type: new Abstract: Pre-trained text-to-image (T2I) diffusion models have shown strong potential for real-world image super-resolution (Real-ISR), owing to their noise-star

applicationsarxiv-cs-cv
12 May 2026
Research

Not Blind but Silenced: Rebalancing Vision and Language via Adversarial Counter-Commonsense Equilibrium

DGX agent

arXiv:2605.10676v1 Announce Type: new Abstract: During MLLM decoding, attention often abnormally concentrates on irrelevant image tokens. While existing research dismisses this as invalid noise and fo

researcharxiv-cs-cv
12 May 2026
Applications

NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces

DGX agent

arXiv:2510.03895v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models represent a pivotal advance in embodied intelligence, yet they confront critical barriers to real-world de

applicationsarxiv-cs-cv
12 May 2026
Research

OCP-GN: A Scalable Second-order Optimizer for Stochastic Optimization

DGX agent

arXiv:2512.24552v2 Announce Type: replace Abstract: This paper proposes a novel second-order optimization algorithm based on the Optimal Control Principle (OCP), applicable to large-scale optimization

researcharxiv-cs-cv
12 May 2026
Safety

Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs

DGX agent

arXiv:2605.09433v1 Announce Type: new Abstract: Existing preference datasets for text-to-image models typically store only the final winner/loser images. This representation is insufficient for rectif

safetyarxiv-cs-cv
12 May 2026
Model Releases

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

DGX agent

arXiv:2605.09996v1 Announce Type: new Abstract: While multimodal large language models have advanced across text, image, and audio, personalization research has remained primarily vision-language, wit

model-releasesarxiv-cs-cv
12 May 2026
Safety

On-Policy Distillation with Best-of-N Teacher Rollout Selection

DGX agent

arXiv:2605.09725v1 Announce Type: new Abstract: On-policy distillation (OPD), which supervises a student on its own sampled trajectories, has emerged as a data-efficient post-training method for impro

safetyarxiv-cs-cv
12 May 2026
Safety

On the Generation and Mitigation of Harmful Geometry in Image-to-3D Models

DGX agent

arXiv:2605.09606v1 Announce Type: cross Abstract: Recent advances in image-to-3D models have significantly improved the fidelity and accessibility of 3D content creation. Such a powerful reconstructio

safetyarxiv-cs-cv
12 May 2026
Research

One-step Latent-free Image Generation with Pixel Mean Flows

DGX agent

arXiv:2601.22158v3 Announce Type: replace Abstract: Modern diffusion/flow-based models for image generation typically exhibit two core characteristics: (i) using multi-step sampling, and (ii) operatin

researcharxiv-cs-cv
12 May 2026
Research

Only Train Once: Uncertainty-Aware One-Class Learning for Face Authenticity Detection

DGX agent

arXiv:2605.10040v1 Announce Type: new Abstract: The rapid evolution of generative paradigms has enabled the creation of highly realistic imagery, which escalating the risks of identity fraud and the d

researcharxiv-cs-cv
12 May 2026
Model Releases

OpenSGA: Efficient 3D Scene Graph Alignment in the Open World

DGX agent

arXiv:2605.10484v1 Announce Type: new Abstract: Scene graph alignment establishes object correspondences between two 3D scene graphs constructed from partially overlapping observations. This enables e

model-releasesarxiv-cs-cv
12 May 2026
Research

OsteoFlow: Lyapunov-Guided Flow Distillation for Predicting Bone Remodeling after Mandibular Reconstruction

DGX agent

arXiv:2603.22421v2 Announce Type: replace Abstract: Predicting long-term bone remodeling after mandibular reconstruction would be of great clinical benefit, yet standard generative models struggle to

researcharxiv-cs-cv
12 May 2026
Safety

Overcoming Catastrophic Forgetting in Visual Continual Learning with Reinforcement Fine-Tuning

DGX agent

arXiv:2605.09640v1 Announce Type: new Abstract: Recent studies suggest that Reinforcement Fine-Tuning (RFT) is inherently more resilient to catastrophic forgetting than Supervised Fine-Tuning (SFT). H

safetyarxiv-cs-cv
12 May 2026
Research

OZ-TAL: Online Zero-Shot Temporal Action Localization

DGX agent

arXiv:2605.09976v1 Announce Type: new Abstract: Online Temporal Action Localization (On-TAL) aims to detect the occurrence time and category of actions in untrimmed streaming videos immediately upon t

researcharxiv-cs-cv
12 May 2026
Research

P-Flow: Proxy-gradient Flows for Linear Inverse Problems

DGX agent

arXiv:2605.08328v1 Announce Type: cross Abstract: Generative models based on flow matching have emerged as a powerful paradigm for inverse problems, offering straighter trajectories and faster samplin

researcharxiv-cs-cv
12 May 2026
Research

PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers

DGX agent

arXiv:2605.08371v1 Announce Type: new Abstract: Visual Geometry Transformer (VGGT) is a strong feed-forward model for multiple 3D tasks, but its Alternating-Attention (AA) stack scales quadratically i

researcharxiv-cs-cv
12 May 2026
Tutorials

PaMoSplat: Part-Aware Motion-Guided Gaussian Splatting for Dynamic Scene Reconstruction

DGX agent

arXiv:2605.10307v1 Announce Type: new Abstract: Dynamic scene reconstruction represents a fundamental yet demanding challenge in computer vision and robotics. While recent progress in 3DGS-based metho

tutorialsarxiv-cs-cv
12 May 2026
Hardware

PermuQuant: Lowering Per-Group Quantization Error by Reordering Channels for Diffusion Models

DGX agent

arXiv:2605.09503v1 Announce Type: new Abstract: Large-scale visual generative models have achieved remarkable performance. However, their high computational and memory costs make deployment challengin

hardwarearxiv-cs-cv
12 May 2026
Model Releases

Personal Visual Context Learning in Large Multimodal Models

DGX agent

arXiv:2605.10936v1 Announce Type: new Abstract: As wearable devices like smart glasses integrate Large Multimodal Models (LMMs) into the continuous first-person visual streams of individual users, the

model-releasesarxiv-cs-cv
12 May 2026
Research

PGID: Progressive Guided Inversion and Denoising for Robust Watermark Detection

DGX agent

arXiv:2605.09319v1 Announce Type: new Abstract: With the proliferation of AI-generated images, digital watermarking has become an essential safeguard for protecting intellectual property and mitigatin

researcharxiv-cs-cv
12 May 2026
Research

PIDNet: Progressive Implicit Decouple Network for Multimodal Action Quality Assessment

DGX agent

arXiv:2605.08945v1 Announce Type: new Abstract: Action quality assessment (AQA) aims to automatically quantify the execution quality of human actions in videos and is valuable for applications such as

researcharxiv-cs-cv
12 May 2026
Model Releases

Pix2Fact: When Vision Is Not Enough -- Benchmarking Fine-Grained VQA with Web Verification on High-Resolution Real-World Scenes

DGX agent

arXiv:2602.00593v2 Announce Type: replace Abstract: Despite progress on general tasks, vision-language models (VLMs) still struggle with challenges that demand both fine-grained visual grounding and e

model-releasesarxiv-cs-cv
12 May 2026
Research

Pixal3D: Pixel-Aligned 3D Generation from Images

DGX agent

arXiv:2605.10922v1 Announce Type: new Abstract: Recent advances in 3D generative models have rapidly improved image-to-3D synthesis quality, enabling higher-resolution geometry and more realistic appe

researcharxiv-cs-cv
12 May 2026
Applications

PixelFlowCast: Latent-Free Precipitation Nowcasting via Pixel Mean Flows

DGX agent

arXiv:2605.10046v1 Announce Type: new Abstract: Precipitation nowcasting aims to forecast short-term radar echo sequences for extreme weather warning, where both prediction fidelity and inference effi

applicationsarxiv-cs-cv
12 May 2026
Model Releases

PolarVSR: A Unified Framework and Benchmark for Continuous Space-Time Polarization Video Reconstruction

DGX agent

arXiv:2605.10275v1 Announce Type: new Abstract: Polarimetric imaging captures surface polarization characteristics, such as the Degree of Linear Polarization (DoLP) and the Angle of Polarization (AoP)

model-releasesarxiv-cs-cv
12 May 2026
Local Ai

Polygon-mamba: Retinal vessel segmentation using polygon scanning mamba and space-frequency collaborative attention

DGX agent

arXiv:2605.10581v1 Announce Type: new Abstract: Retinal vessel segmentation is crucial for diagnosis and assessment of ocular diseases. Notably, segmentation of small retinal vessels has been consiste

local-aiarxiv-cs-cv
12 May 2026
Research

Position: Life-Logging Video Streams Make the Privacy-Utility Trade-off Inevitable

DGX agent

arXiv:2605.10404v1 Announce Type: new Abstract: With the growing prevalence of always-on hardware such as smart glasses, body cameras, and home security systems, life-logging visual sensing is becomin

researcharxiv-cs-cv
12 May 2026
Research

Post-hoc Selective Classification for Reliable Synthetic Image Detection

DGX agent

arXiv:2605.08574v1 Announce Type: new Abstract: As synthetic images become increasingly realistic, reliable synthetic image detection techniques are of pressing need to prevent their misuse. Despite s

researcharxiv-cs-cv
12 May 2026
Local Ai

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping

DGX agent

arXiv:2605.10937v1 Announce Type: new Abstract: Recently, post-training methods based on reinforcement learning, with a particular focus on Group Relative Policy Optimization (GRPO), have emerged as t

local-aiarxiv-cs-cv
12 May 2026
Tutorials

Predicting 3D structure by latent posterior sampling

DGX agent

arXiv:2605.10830v1 Announce Type: new Abstract: The remarkable achievements of both generative models of 2D images and neural field representations for 3D scenes present a compelling opportunity to in

tutorialsarxiv-cs-cv
12 May 2026
Model Releases

Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional Drift

DGX agent

arXiv:2505.19519v3 Announce Type: replace Abstract: Personalizing text-to-image diffusion models involves integrating novel visual concepts from a small set of reference images while retaining the mod

model-releasesarxiv-cs-cv
12 May 2026
Research

Probability-Flow Distillation: Exact Wasserstein Gradient Flow for High-Fidelity 3D Generation

DGX agent

arXiv:2605.09071v1 Announce Type: new Abstract: Score Distillation Sampling (SDS) and its variants have been widely used for text-to-3D generation by distilling 2D image diffusion priors. However, the

researcharxiv-cs-cv
12 May 2026
Research

ProDG: Prototypes for Data-Free Generative Post-Hoc Explainability

DGX agent

arXiv:2605.08858v1 Announce Type: new Abstract: Ante-hoc interpretability methods based on prototypes provide highly accurate explanations by utilizing the intuitive 'this looks like that' reasoning p

researcharxiv-cs-cv
12 May 2026
Model Releases

Product-of-Gaussian-Mixture Diffusion Models for Joint Nonlinear MRI Reconstruction

DGX agent

arXiv:2605.10629v1 Announce Type: new Abstract: Recently, diffusion models have attracted considerable attention for magnetic resonance image reconstruction due to their high sample quality. However,

model-releasesarxiv-cs-cv
12 May 2026
Research

Progressive Photorealistic Simplification

DGX agent

arXiv:2605.10409v1 Announce Type: new Abstract: Existing image simplification techniques often rely on Non-Photorealistic Rendering (NPR), transforming photographs into stylized sketches, cartoons, or

researcharxiv-cs-cv
12 May 2026
Model Releases

Prompt Estimation from Prototypes for Federated Prompt Tuning of Vision Transformers

DGX agent

arXiv:2510.25372v2 Announce Type: replace Abstract: Visual Prompt Tuning (VPT) of pre-trained Vision Transformers (ViTs) has proven highly effective as a parameter-efficient fine-tuning technique for

model-releasesarxiv-cs-cv
12 May 2026
Research

QueST: Persistent Queries as Semantic Monitors for Drift Suppression in Long-Horizon Tracking

DGX agent

arXiv:2605.09513v1 Announce Type: new Abstract: Tracking points in videos is typically formulated as frame-to-frame correspondence, where each point is matched locally to the next frame. While this wo

researcharxiv-cs-cv
12 May 2026
← Previous
1…183184185186187…263
Next →