AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Local Ai

Multi-view Pyramid Transformer: Look Coarser to See Broader

DGX agent

arXiv:2512.07806v2 Announce Type: replace Abstract: We propose Multi-view Pyramid Transformer (MVP), a scalable multi-view transformer architecture that directly reconstructs large 3D scenes from tens

local-aiarxiv-cs-cv
2 Jun 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Multimodal Action Diffusion for Robust End-to-End Autonomous Driving

DGX agent

arXiv:2606.02105v1 Announce Type: new Abstract: End-to-End Autonomous Driving (E2E-AD) systems have largely converged on predicting intermediate trajectory waypoints, delegating final control to hand-

model-releasesarxiv-cs-cv
2 Jun 2026
Agents

MUSCLE-NET: Predicted-Multiscale-Aware Network for Pedestrian Trajectory Forecasting

DGX agent

arXiv:2606.00471v1 Announce Type: new Abstract: Accurate pedestrian trajectory prediction is essential for safe navigation in autonomous driving and intelligent transportation systems. Despite substan

agentsarxiv-cs-cv
2 Jun 2026
Applications

NAPPure: Adversarial Purification for Robust Image Classification under Non-Additive Perturbations

DGX agent

arXiv:2510.14025v2 Announce Type: replace Abstract: Adversarial purification has achieved great success in combating adversarial image perturbations, which are usually assumed to be additive. However,

applicationsarxiv-cs-cv
2 Jun 2026
Applications

Neural Acquisition & Representation of Subsurface Scattering

DGX agent

arXiv:2606.02292v1 Announce Type: new Abstract: We present a method to acquire and estimate the sub-surface scattering properties of light transport at a highly detailed level by learning the pixel fo

applicationsarxiv-cs-cv
2 Jun 2026
Model Releases

Non-Learning Low-Light Stereo Vision

DGX agent

arXiv:2606.00379v1 Announce Type: new Abstract: We present a non-learning stereo framework for disparity estimation from severely noisy images. Using the Field of Junctions (FoJ), it retains coarse vi

model-releasesarxiv-cs-cv
2 Jun 2026
Research

Normality-Preserving Continual Industrial Anomaly Detection via Orthogonal LoRA Banks

DGX agent

arXiv:2606.02042v1 Announce Type: new Abstract: Continual industrial anomaly detection with diffusion models suffers from historical normality prior drift and catastrophic forgetting. Existing continu

researcharxiv-cs-cv
2 Jun 2026
Tutorials

Not All Points Are Equal: Uncertainty-Aware 4D LiDAR Scene Synthesis

DGX agent

arXiv:2606.02510v1 Announce Type: new Abstract: Constructing faithful 4D worlds from LiDAR-acquired sequences is crucial for embodied AI, yet current generative frameworks apply uniform modeling capac

tutorialsarxiv-cs-cv
2 Jun 2026
Safety

ObjEmbed: Towards Universal Multimodal Object Embeddings

DGX agent

arXiv:2602.01753v3 Announce Type: replace Abstract: Aligning objects with corresponding textual descriptions is a fundamental challenge and a realistic requirement in vision-language understanding. Wh

safetyarxiv-cs-cv
2 Jun 2026
Local Ai

One Channel to Rule Them All: Rethinking Input Representation for Visual Place Recognition

DGX agent

arXiv:2606.00936v1 Announce Type: new Abstract: Visual Place Recognition (VPR) is fundamental to long-term robot localization and SLAM, yet current systems overwhelmingly rely on RGB input, implicitly

local-aiarxiv-cs-cv
2 Jun 2026
Local Ai

One-Shot Crowd Counting With Density Guidance For Scene Adaptation

DGX agent

arXiv:2602.07955v2 Announce Type: replace Abstract: Crowd scenes captured by cameras at different locations vary greatly, and existing crowd models have limited generalization for unseen surveillance

local-aiarxiv-cs-cv
2 Jun 2026
Research

OP-LoRA: The Blessing of Dimensionality

DGX agent

arXiv:2412.10362v2 Announce Type: replace-cross Abstract: Low-rank adapters (LoRA) enable finetuning of large models with only a small number of parameters. However, they often suffer from an ill-cond

researcharxiv-cs-cv
2 Jun 2026
Model Releases

OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing Imagery

DGX agent

arXiv:2603.27645v2 Announce Type: replace Abstract: Open-vocabulary change detection (OVCD) seeks to recognize arbitrary changes of interest by enabling generalization beyond a fixed set of predefined

model-releasesarxiv-cs-cv
2 Jun 2026
Research

Optimizing 3D Gaussian Splatting via Point Cloud Upsampling

DGX agent

arXiv:2606.00450v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) is a technique for creating and rendering 3D scenes, however its performance depends heavily on the quality of initial seed

researcharxiv-cs-cv
2 Jun 2026
Research

OptiWorld: Optimal Control for Video World Generation under Physical Constraints

DGX agent

arXiv:2606.00499v1 Announce Type: new Abstract: Video generation models are becoming a scalable form of world models, but they mainly generate plausible motion rather than proactively control or optim

researcharxiv-cs-cv
2 Jun 2026
Research

PaCX-MAE: Physiology-Augmented Chest X-Ray Masked Autoencoder

DGX agent

arXiv:2606.01537v1 Announce Type: new Abstract: Clinical diagnosis often requires combining imaging with physiological measurements, yet deployed models typically operate on unimodal data. We present

researcharxiv-cs-cv
2 Jun 2026
Research

PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion

DGX agent

arXiv:2606.01399v1 Announce Type: new Abstract: We present PAI-Studio, a new reference-conditioned video synthesis task that addresses a long-standing challenge in cinematic background replacement: ge

researcharxiv-cs-cv
2 Jun 2026
Agents

PairedGTA: Generating Driving Datasets for Controlled Photometric Shift Analysis

DGX agent

arXiv:2606.01192v1 Announce Type: new Abstract: Evaluating the performance of visual perception systems for autonomous driving is essential to ensure reliable operation across diverse environmental sc

agentsarxiv-cs-cv
2 Jun 2026
Research

PathAR: Structure-First Autoregressive Synthesis of Multimodal Pathology Images

DGX agent

arXiv:2606.01543v1 Announce Type: new Abstract: Data scarcity in multimodal pathology motivates unified generative models that synthesize modality-specific appearance while preserving anatomically coh

researcharxiv-cs-cv
2 Jun 2026
Safety

Pave-GRPO: Beyond Instantaneous Guidance through Principled Average Velocity Decomposition

DGX agent

arXiv:2606.01636v1 Announce Type: new Abstract: Post-training via Group Relative Policy Optimization (GRPO) has emerged as a powerful paradigm for aligning flow-based generative models with human pref

safetyarxiv-cs-cv
2 Jun 2026
Research

Paving the Way for Point Cloud Video Representation Learning Using A PDE Model

DGX agent

arXiv:2606.01604v1 Announce Type: new Abstract: Investigating spatial-temporal correlations, specifically how spatial points vary over time, is crucial for understanding point cloud videos. Traditiona

researcharxiv-cs-cv
2 Jun 2026
Research

PerBite: A Curated Diagnostic Workflow for Bite-Aware Food Volume Estimation

DGX agent

arXiv:2606.02021v1 Announce Type: new Abstract: Can a visually plausible food mesh be trusted to estimate the volume of consumed food? method investigates this question using selected paired before- a

researcharxiv-cs-cv
2 Jun 2026
Model Releases

Perception First: A Frontier Native-Video Model with Self-Consistency for Implicit Video Question Answering

DGX agent

arXiv:2606.01485v1 Announce Type: new Abstract: We describe our submission to the VRR Challenge @ CVPR 2026, built on the ImplicitQA / VRR-QA benchmark~ite{implicitqa}: multiple-choice video question

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Personalized 3D Myocardial Infarct Geometry Reconstruction from Cine MRI for Cardiac Digital Twins

DGX agent

arXiv:2606.01808v1 Announce Type: new Abstract: Accurate 3D geometric characterization of myocardial infarction (MI) is essential for building cardiac digital twins (CDTs) to precisely simulate infarc

model-releasesarxiv-cs-cv
2 Jun 2026
Safety

PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation

DGX agent

arXiv:2606.01649v1 Announce Type: new Abstract: Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The chal

safetyarxiv-cs-cv
2 Jun 2026
Research

Physical Object Understanding with a Physically Controllable World Model

DGX agent

arXiv:2606.00439v1 Announce Type: new Abstract: A central challenge in visual intelligence is learning the physical structure of scenes from raw videos: how regions form objects and the laws that gove

researcharxiv-cs-cv
2 Jun 2026
Research

Physics-Aware Linearized ADMM and Its Unrolling

DGX agent

arXiv:2606.01652v1 Announce Type: cross Abstract: Recently, partial differential equations (PDEs) have been used to directly model the measurement process in signal processing, although their evaluati

researcharxiv-cs-cv
2 Jun 2026
Model Releases

Physics from Video: Identifiability of Time-Invariant Second-Order ODEs under Minimal Trajectory Conditions

DGX agent

arXiv:2606.00115v1 Announce Type: new Abstract: Bridging the gap between visual realism and physical understanding is a core challenge for video-based world models. We study the structural identifiabi

model-releasesarxiv-cs-cv
2 Jun 2026
Agents

PillarDETR: YOLO-Backbone and RT-DETR Head for Real-Time 3D Object Detection

DGX agent

arXiv:2606.01757v1 Announce Type: new Abstract: Real-time 3D object detection is a critical component for the safe operation of autonomous driving systems and robotics. While LiDAR point clouds provid

agentsarxiv-cs-cv
2 Jun 2026
Research

PINNOCHIO: Physics-Informed Neural Network for Coupled Hyperelastic Interface-Volume Simulation in Orthognathic Surgery

DGX agent

arXiv:2606.01572v1 Announce Type: cross Abstract: Predicting patient-specific facial soft-tissue deformation is critical for iterative orthognathic surgery planning. However, current computational met

researcharxiv-cs-cv
2 Jun 2026
Research

Pinterest Canvas: Large-Scale Image Generation at Pinterest

DGX agent

arXiv:2603.06453v2 Announce Type: replace Abstract: While recent image generation models demonstrate a remarkable ability to handle a wide variety of image generation tasks, this flexibility makes the

researcharxiv-cs-cv
2 Jun 2026
Research

Places in the Wild: A Large, High-Resolution RAW Photograph Dataset for Ecologically Valid Vision Research

DGX agent

arXiv:2606.02481v1 Announce Type: new Abstract: Large image datasets have accelerated progress in cognitive neuroscience and computer vision. However, most datasets are low-resolution, internet-source

researcharxiv-cs-cv
2 Jun 2026
Model Releases

Plan-R1: Safe and Feasible Trajectory Planning as Language Modeling

DGX agent

arXiv:2505.17659v4 Announce Type: replace-cross Abstract: Safe and feasible trajectory planning is critical for real-world autonomous driving systems. However, existing learning-based planners rely he

model-releasesarxiv-cs-cv
2 Jun 2026
Agents

PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps

DGX agent

arXiv:2606.01788v1 Announce Type: new Abstract: Embodied visual navigation, where an agent perceives a complex environment and acts to reach a goal from raw sensory input, underpins a wide range of ap

agentsarxiv-cs-cv
2 Jun 2026
Research

Polaris: Scaling Up Instruction-Guided Image Generation Towards Millions of Personalized Style Needs

DGX agent

arXiv:2606.01858v1 Announce Type: new Abstract: Users increasingly expect image generation models to quickly adapt to highly diverse and personalized requirements, such as producing images with distin

researcharxiv-cs-cv
2 Jun 2026
Safety

Policy-based Foveated Imaging and Perception

DGX agent

arXiv:2606.02565v1 Announce Type: new Abstract: Ultra-high-resolution image sensors offer the potential to capture fine spatial details critical for many visual perception tasks, but acquiring and pro

safetyarxiv-cs-cv
2 Jun 2026
Safety

Pool-Select-Refine: Allocation-Aware Generative Dataset Distillation with Soft-Label-Guided Latent Refinement

DGX agent

arXiv:2606.01920v1 Announce Type: new Abstract: Diffusion-based dataset distillation has recently emerged as a promising paradigm for condensing large-scale datasets into compact synthetic sets. By le

safetyarxiv-cs-cv
2 Jun 2026
Research

Positional Encodings Anchor Spatial Structure in Vision Transformers: A Geometric Perspective on Robustness

DGX agent

arXiv:2606.00124v1 Announce Type: new Abstract: Positional embeddings (PEs) in Vision Transformers (ViTs) are known to impact performance and robustness, but their role in shaping internal spatial rep

researcharxiv-cs-cv
2 Jun 2026
Research

PRIMA: Boosting Animal Mesh Recovery with Biological Priors and Test-Time Adaptation

DGX agent

arXiv:2606.02366v1 Announce Type: new Abstract: We present PRIMA (*PRI*ors for *M*esh *A*daptation), a framework for robust 3D quadruped mesh recovery under severe species and pose imbalance. Existing

researcharxiv-cs-cv
2 Jun 2026
Model Releases

Princeton365: A Diverse Dataset with Accurate Camera Pose

DGX agent

arXiv:2506.09035v2 Announce Type: replace Abstract: We introduce Princeton365, a large-scale diverse dataset of 365 videos with accurate camera pose. Our dataset bridges the gap between accuracy and d

model-releasesarxiv-cs-cv
2 Jun 2026
Research

Private and Stable Test-Time Adaptation with Differential Privacy

DGX agent

arXiv:2606.01908v1 Announce Type: cross Abstract: Test-time adaptation (TTA) can reduce error on new and different data by updating the model on these inputs during inference. However, these updates r

researcharxiv-cs-cv
2 Jun 2026
Applications

ProtoAda: Prototype-Guided Adaptive Adapter Expansion and Geometric Consolidation for Multimodal Continual Instruction Tuning

DGX agent

arXiv:2606.02576v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) achieve strong performance through instruction tuning, but real-world deployment requires them to continually a

applicationsarxiv-cs-cv
2 Jun 2026
Model Releases

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning

DGX agent

arXiv:2507.08064v3 Announce Type: replace-cross Abstract: As multimedia content expands, the demand for unified multimodal retrieval (UMR) in real-world applications increases. Recent work leverages m

model-releasesarxiv-cs-cv
2 Jun 2026
Research

Quality-Guided Semi-Supervised Learning for Medical Image Segmentation

DGX agent

arXiv:2606.01753v1 Announce Type: new Abstract: Training accurate medical image segmentation models requires large amounts of densely annotated data, which is costly and time-consuming to obtain. Semi

researcharxiv-cs-cv
2 Jun 2026
Model Releases

Question-Aware Evidence Ledgers for Video Relational Reasoning

DGX agent

arXiv:2606.02506v1 Announce Type: new Abstract: The VRR-QA challenge evaluates visual relational reasoning in videos, where answers often depend on implicit spatial relations, event boundaries, target

model-releasesarxiv-cs-cv
2 Jun 2026
Research

R^3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking

DGX agent

arXiv:2606.01113v1 Announce Type: new Abstract: The CoVR-R challenge evaluates composed video retrieval, where a system must retrieve a target video from a large gallery given a reference video and a

researcharxiv-cs-cv
2 Jun 2026
Safety

RAIGen: Rare Attribute Identification in Text-to-Image Generative Models

DGX agent

arXiv:2602.06806v2 Announce Type: replace Abstract: Text-to-image diffusion models achieve impressive generation quality but inherit and amplify training-data biases, skewing coverage of semantic attr

safetyarxiv-cs-cv
2 Jun 2026
Research

Random Erasing vs. Model Inversion: A Promising Defense or a False Hope?

DGX agent

arXiv:2409.01062v4 Announce Type: replace-cross Abstract: Model Inversion (MI) attacks pose a significant privacy threat by reconstructing private training data from machine learning models. While exi

researcharxiv-cs-cv
2 Jun 2026
← Previous
1…128129130131132…263
Next →