AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Local Ai

One-Step Distillation of Discrete Diffusion Image Generators via Fixed-Point Iteration

DGX agent

arXiv:2605.21484v1 Announce Type: new Abstract: Discrete diffusion models excel at visual synthesis but rely on slow, iterative decoding. Existing single-step distillation methods attempt to bypass th

local-aiarxiv-cs-cv
21 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Local Ai

Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation

DGX agent

arXiv:2509.09946v2 Announce Type: replace Abstract: Multi-Target Multi-Camera Tracking (MTMC) is an essential computer vision task for automating large-scale surveillance. With camera calibration and

local-aiarxiv-cs-cv
21 May 2026
Research

Oracle Supervision Transfers for Hyperparameter Prediction in Model-Based Image Denoising

DGX agent

arXiv:2605.20479v1 Announce Type: new Abstract: Hyperparameter prediction is a critical practical bottleneck for model-based image denoisers, ranging from classical TV/TGV variational solvers to moder

researcharxiv-cs-cv
21 May 2026
Research

Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition

DGX agent

arXiv:2605.21417v1 Announce Type: new Abstract: Blended emotion recognition is challenging because emotions are often expressed as mixtures of subtle and overlapping multimodal cues rather than a sing

researcharxiv-cs-cv
21 May 2026
Local Ai

OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026

DGX agent

arXiv:2605.20818v1 Announce Type: new Abstract: In this report, we present our champion solutions for the Natural Language Queries and GoalStep tracks of the Ego4D Episodic Memory Challenge at CVPR 20

local-aiarxiv-cs-cv
21 May 2026
Local Ai

PaintCopilot: Modeling Painting as Autonomous Artistic Continuation

DGX agent

arXiv:2605.20941v1 Announce Type: new Abstract: We present PaintCopilot, a co-creative neural painting assistant that models painting as an open-ended autoregressive artistic behavior conditioned on e

local-aiarxiv-cs-cv
21 May 2026
Model Releases

Parameters as Experts: Adapting Vision Models with Dynamic Parameter Routing

DGX agent

arXiv:2602.06862v2 Announce Type: replace Abstract: Adapting pre-trained vision models using parameter-efficient fine-tuning (PEFT) remains challenging, as it aims to achieve performance comparable to

model-releasesarxiv-cs-cv
21 May 2026
Agents

ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning

DGX agent

arXiv:2605.20342v1 Announce Type: new Abstract: Training large multimodal models (LMMs) via reinforcement learning (RL) to natively invoke video-processing tools (e.g., cropping) has become a promisin

agentsarxiv-cs-cv
21 May 2026
Safety

Pareto-Enhanced Portrait Generation: Vision-Aligned Text Supervision for Alignment, Realism, and Aesthetics

DGX agent

arXiv:2605.20640v1 Announce Type: new Abstract: Text-to-image diffusion models often face a severe trilemma in human portrait generation: text-image alignment, photorealism, and human-perceived aesthe

safetyarxiv-cs-cv
21 May 2026
Research

PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation

DGX agent

arXiv:2602.04876v2 Announce Type: replace Abstract: We introduce PerpetualWonder, a hybrid generative simulator that enables long-horizon, action-conditioned 4D scene generation from a single image. C

researcharxiv-cs-cv
21 May 2026
Model Releases

PGC: Peak-Guided Calibration for Generalizable AI-Generated Image Detection

DGX agent

arXiv:2605.21207v1 Announce Type: new Abstract: The rapid evolution of generative AI, from GANs to modern diffusion models, has resulted in increasingly subtle discriminative clues. These fine-grained

model-releasesarxiv-cs-cv
21 May 2026
Research

Pixel Wised Lesion Prediction on COVID-19 CT Imagery: A Comparative Analysis of Automated Image Segmentation Architectures

DGX agent

arXiv:2605.20459v1 Announce Type: new Abstract: In recent years, there has been a notable increase in the level of attention that is given to algorithms based on deep learning in the context of medica

researcharxiv-cs-cv
21 May 2026
Research

Platonic Representations in the Human Brain: Unsupervised Recovery of Universal Geometry

DGX agent

arXiv:2605.20496v1 Announce Type: cross Abstract: The Strong Platonic Representation Hypothesis suggests that representational convergence in artificial neural networks can be harnessed constructively

researcharxiv-cs-cv
21 May 2026
Safety

PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction

DGX agent

arXiv:2605.21414v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation by leveraging large pretrained vision-languag

safetyarxiv-cs-cv
21 May 2026
Research

PREF: Phasorial Embedding Fields for Compact Neural Representations

DGX agent

arXiv:2205.13524v4 Announce Type: replace Abstract: We present an efficient frequency-based neural representation termed PREF: a shallow MLP augmented with a phasor volume that covers significant bord

researcharxiv-cs-cv
21 May 2026
Model Releases

Preserve, Reveal, Expand: Faithful 4D Video Editing with Region-Aware Conditioning

DGX agent

arXiv:2605.20961v1 Announce Type: new Abstract: Existing 4D-driven video diffusion models primarily target plausible generation, but faithful 4D editing requires preserving source-observed regions whi

model-releasesarxiv-cs-cv
21 May 2026
Safety

Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization

DGX agent

arXiv:2510.21583v2 Announce Type: replace Abstract: Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated st

safetyarxiv-cs-cv
21 May 2026
Agents

ProCrit: Self-Elicited Multi-Perspective Reasoning with Critic-Guided Revision for Multimodal Sarcasm Detection

DGX agent

arXiv:2605.20867v1 Announce Type: cross Abstract: Multimodal sarcasm detection requires reasoning over cross-modal incongruities between literal expression and intended meaning, yet the specific analy

agentsarxiv-cs-cv
21 May 2026
Research

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers

DGX agent

arXiv:2602.06886v3 Announce Type: replace Abstract: Multimodal Diffusion Transformers (MMDiTs) for text-to-image generation maintain separate text and image branches, with bidirectional information fl

researcharxiv-cs-cv
21 May 2026
Research

ProtoPathway: Biologically Structured Prototype-Pathway Fusion for Multimodal Cancer Survival Prediction

DGX agent

arXiv:2605.21454v1 Announce Type: new Abstract: We introduce ProtoPathway, an interpretable-by-design multimodal framework for cancer survival prediction that unifies whole slide imaging and transcrip

researcharxiv-cs-cv
21 May 2026
Research

Q-ARVD: Quantizing Autoregressive Video Diffusion Models

DGX agent

arXiv:2605.21072v1 Announce Type: new Abstract: Autoregressive video diffusion models (ARVDs) have emerged as a promising architecture for streaming video generation, paving the way for real-time inte

researcharxiv-cs-cv
21 May 2026
Model Releases

Q-DiT4SR: Exploration of Detail-Preserving Diffusion Transformer Quantization for Real-World Image Super-Resolution

DGX agent

arXiv:2602.01273v4 Announce Type: replace Abstract: Recently, Diffusion Transformers (DiTs) have emerged in Real-World Image Super-Resolution (Real-ISR) to generate high-quality textures, yet their he

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

Query-Calibrated Segmental Admission for Descriptor-Agnostic LiDAR Loop Closure in Repetitive Environments

DGX agent

arXiv:2512.09447v2 Announce Type: replace-cross Abstract: Structurally repetitive environments produce visually plausible but aliased LiDAR loop candidates that can destabilize pose-graph optimization

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

QwenSafe: Multimodal Content Rating Description Identification via Preference-Aligned VLMs

DGX agent

arXiv:2605.20584v1 Announce Type: new Abstract: Mobile app marketplaces require developers to disclose standardized content rating descriptors (CRDs) to inform users about potentially sensitive or res

model-releasesarxiv-cs-cv
21 May 2026
Local Ai

R2AoP: Reliable and Robust Angle of Progression Estimation from Intrapartum Ultrasound

DGX agent

arXiv:2605.21099v1 Announce Type: new Abstract: Accurate estimation of the Angle of Progression (AoP) from intrapartum transperineal ultrasound is critical for objective assessment of labor progressio

local-aiarxiv-cs-cv
21 May 2026
Model Releases

RadProPoser: Probabilistic Radar Tensor Human Pose Estimation That Knows Its Limits

DGX agent

arXiv:2508.03578v2 Announce Type: replace Abstract: Radar-based human pose estimation enables privacy-preserving motion tracking for ambient intelligence, yet the noisy nature of radar sensing makes u

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution

DGX agent

arXiv:2605.21195v1 Announce Type: new Abstract: Discrete autoregressive (AR) text-to-image (T2I) models pair a VQ tokenizer with an AR policy, and current post-training pipelines optimize only the pol

model-releasesarxiv-cs-cv
21 May 2026
Model Releases

RCGDet3D: Rethinking 4D Radar-Camera Fusion-based 3D Object Detection with Enhanced Radar Feature Encoding

DGX agent

arXiv:2605.21112v1 Announce Type: new Abstract: 4D automotive radar is indispensable for autonomous driving due to its low cost and robustness, yet its point cloud sparsity challenges 3D object detect

model-releasesarxiv-cs-cv
21 May 2026
Tutorials

Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens

DGX agent

arXiv:2605.21300v1 Announce Type: new Abstract: Object hallucination is a significant challenge that hinders the application of large vision-language models (LVLMs) in practice. We hypothesize that on

tutorialsarxiv-cs-cv
21 May 2026
Safety

Regulating Anatomy-Aware Rewards via Trajectory-Integral Feedback for Volumetric Computed Tomography Analysis

DGX agent

arXiv:2605.20277v1 Announce Type: new Abstract: Medical vision-language models (VLMs) have rapidly advanced as general-purpose multimodal assistants, yet their deployment in 3D Computed Tomography (CT

safetyarxiv-cs-cv
21 May 2026
Research

RelWitness: Open-Vocabulary 3D Scene Graph Generation with Visual-Geometric Relation Witnesses

DGX agent

arXiv:2605.20823v1 Announce Type: new Abstract: Open-vocabulary 3D scene graph generation seeks to describe object instances and their relations with flexible natural-language predicates. The central

researcharxiv-cs-cv
21 May 2026
Research

ReMATF: Recurrent Motion-Adaptive Multi-scale Turbulence Mitigation for Dynamic Scenes

DGX agent

arXiv:2605.21440v1 Announce Type: new Abstract: Atmospheric turbulence severely degrades video quality by introducing distortions such as geometric warping, blur, and temporal flickering, posing signi

researcharxiv-cs-cv
21 May 2026
Local Ai

RePCM: Region-Specific and Phenotype-Adaptive Bi-Ventricular Cardiac Motion Synthesis

DGX agent

arXiv:2605.21237v1 Announce Type: new Abstract: Cardiac motion over a cardiac cycle is crucial for quantifying regional function and is strongly affected by cardiovascular diseases. Since temporally d

local-aiarxiv-cs-cv
21 May 2026
Research

ResNet-50 with Class Reweighting and Anatomy-Guided Temporal Decoding for Gastrointestinal Video Analysis

DGX agent

arXiv:2603.17784v2 Announce Type: replace Abstract: We developed a multi-label gastrointestinal video analysis pipeline based on a ResNet-50 frame classifier followed by anatomy-guided temporal event

researcharxiv-cs-cv
21 May 2026
Model Releases

Resolving Long-Tail Ambiguity in Unsupervised 3D Point Cloud Segmentation with Language Priors

DGX agent

arXiv:2605.20737v1 Announce Type: new Abstract: Existing approaches for unsupervised 3D point cloud segmentation predominantly rely on a purely visual similarity-based learning-by-clustering paradigm,

model-releasesarxiv-cs-cv
21 May 2026
Safety

Rethinking Cross-Layer Information Routing in Diffusion Transformers

DGX agent

arXiv:2605.20708v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have become a de facto backbone of modern visual generation, and nearly every major axis of their design -- tokenization,

safetyarxiv-cs-cv
21 May 2026
Research

RISE: Reliable Improvement in Self-Evolving Vision-Language Models

DGX agent

arXiv:2605.20914v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong multimodal reasoning capabilities, but further improving them still relies heavily on large-scale hum

researcharxiv-cs-cv
21 May 2026
Research

RoadTones: Tone Controllable Text Generation from Road Event Videos

DGX agent

arXiv:2605.21411v1 Announce Type: new Abstract: Existing video-language models can generate factual descriptions of road events but lack control over how these events are expressed: their tone, urgenc

researcharxiv-cs-cv
21 May 2026
Research

ROAR-3D: Routing Arbitrary Views for High-Fidelity 3D Generation

DGX agent

arXiv:2605.21121v1 Announce Type: new Abstract: Single-image-to-3D generative models can now produce high-quality geometry, yet conditioning on a single view inevitably introduces ambiguity about unse

researcharxiv-cs-cv
21 May 2026
Research

RoPeSLR: 3D RoPE-driven Sparse-LowRank Attention for Efficient Diffusion Transformers

DGX agent

arXiv:2605.20659v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have revolutionized high-fidelity video generation, yet their O(L^2) attention complexity poses a formidable bottleneck fo

researcharxiv-cs-cv
21 May 2026
Safety

SAM-Sode: Towards Faithful Explanations for Tiny Bacteria Detection

DGX agent

arXiv:2605.21186v1 Announce Type: new Abstract: Interpretability in object detection provides crucial confidence support for clinical auxiliary diagnosis. However, in tiny bacteria detection, traditio

safetyarxiv-cs-cv
21 May 2026
Research

SAVER: Selective As-Needed Vision Evidence for Multimodal Information Extraction

DGX agent

arXiv:2605.20713v1 Announce Type: new Abstract: Multimodal IE in social media is difficult because a post may attach multiple images that are weakly related, redundant, or even misleading with respect

researcharxiv-cs-cv
21 May 2026
Research

SDM: A Powerful Tool for Evaluating Model Robustness

DGX agent

arXiv:2605.20308v1 Announce Type: new Abstract: Gradient-based attacks are important methods for evaluating model robustness. However, since the proposal of APGD, it has been difficult for such method

researcharxiv-cs-cv
21 May 2026
Model Releases

Seeing Through Fog: Towards Fog-Invariant Action Recognition

DGX agent

arXiv:2605.20645v1 Announce Type: new Abstract: Foggy conditions are commonly encountered in real-world applications; however, existing action recognition approaches typically assume favorable weather

model-releasesarxiv-cs-cv
21 May 2026
Safety

Self-Refining Video Sampling

DGX agent

arXiv:2601.18577v2 Announce Type: replace Abstract: Modern video generators still struggle with complex physical dynamics, often falling short of physical realism. Existing approaches address this usi

safetyarxiv-cs-cv
21 May 2026
Local Ai

Semantic Granularity Navigation in Image Editing

DGX agent

arXiv:2605.21190v1 Announce Type: new Abstract: Despite the generative capabilities of diffusion and flow models, real-image editing remains constrained by a persistent trade-off between semantic edit

local-aiarxiv-cs-cv
21 May 2026
Model Releases

ShadeBench: A Benchmark Dataset for Building Shade Simulation in Sustainable Society

DGX agent

arXiv:2605.20510v1 Announce Type: new Abstract: Urban heat exposure is becoming an increasingly critical challenge due to the intensifying urban heat island effect. Fine-grained shade patterns, especi

model-releasesarxiv-cs-cv
21 May 2026
Applications

ShowMak3r: Compositional TV Show Reconstruction

DGX agent

arXiv:2504.19584v3 Announce Type: replace Abstract: Reconstructing dynamic radiance fields from video clips is challenging, especially when entertainment videos like TV shows are given. Many challenge

applicationsarxiv-cs-cv
21 May 2026
← Previous
1…152153154155156…263
Next →