AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
2 Jun 2026

MipSLAM: Alias-Free Gaussian Splatting SLAM

ResearchDGX agent

arXiv:2603.06989v3 Announce Type: replace Abstract: This paper introduces MipSLAM, a frequency-aware 3D Gaussian Splatting (3DGS) SLAM framework capable of high-fidelity anti-aliased novel view synthe

MixerSENet: A Lightweight Framework for Efficient Hyperspectral Image Classification

Model ReleasesDGX agent

arXiv:2606.01700v1 Announce Type: new Abstract: In this paper, a novel framework, MixerSENet, is introduced for hyperspectral image (HSI) classification, designed to address the challenges of computat

MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue

Model ReleasesDGX agent

arXiv:2606.00622v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) demonstrate remarkable visual understanding, yet their reliability in interactive settings is severely undermin


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

MMDG-Bench: A Benchmark for Multimodal Domain Generalization

Model ReleasesDGX agent

arXiv:2606.00891v1 Announce Type: new Abstract: Multi-modal Domain Generalization (MMDG) seeks to leverage complementary modalities to enhance model robustness on unseen domains. Despite extensive pro

MMTalker: Multiresolution 3D Talking Head Synthesis with Multimodal Feature Fusion

ResearchDGX agent

arXiv:2604.02941v2 Announce Type: replace Abstract: Speech-driven three-dimensional (3D) facial animation synthesis aims to build a mapping from one-dimensional (1D) speech signals to time-varying 3D

Modeling Robotics Dataset Construction as an Artifact-Based Build Process

ResearchDGX agent

arXiv:2606.00162v1 Announce Type: cross Abstract: Robotic systems generate large volumes of multimodal sensor data, but converting ROS bag recordings into machine learning datasets is often handled by

MORPHOS: Autoregressive 4D Generation with Temporal Structured Latents

ResearchDGX agent

arXiv:2606.02491v1 Announce Type: new Abstract: We present MORPHOS, a novel autoregressive framework that generates dynamic 3D assets from videos across diverse representations, including meshes, 3D G

Motion-aware Event Suppression for Event Cameras

Model ReleasesDGX agent

arXiv:2602.23204v3 Announce Type: replace Abstract: In this work, we introduce the first framework for Motion-aware Event Suppression, which learns to filter events triggered by IMOs and ego-motion in

MotionDreamer: Universal Skeletal Motion Generation for 3D Rigged Shapes

Model ReleasesDGX agent

arXiv:2606.01518v1 Announce Type: new Abstract: Motion generation for rigged shapes is vital for scalable 4D asset production. However, template-based methods are limited by specific topologies and fa

MPMWorlds: Material-Point-Method Simulations for Inferring and Extrapolating Physical Dynamics

ResearchDGX agent

arXiv:2606.01538v1 Announce Type: cross Abstract: To study the ability to infer physical dynamics from videos and extrapolate them forward in time, we assemble a dataset of 2D Material Point Method (M

MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching

Model ReleasesDGX agent

arXiv:2606.01985v1 Announce Type: new Abstract: Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing de

Multi-modal Video Representation Alignment for Robust Self-supervised Driver Distraction Detection

SafetyDGX agent

arXiv:2606.02352v1 Announce Type: new Abstract: Robust self-supervised learning of multi-modal video representations is critical for real-world applications such as driver distraction detection, where

Multi-view Pyramid Transformer: Look Coarser to See Broader

Local AiDGX agent

arXiv:2512.07806v2 Announce Type: replace Abstract: We propose Multi-view Pyramid Transformer (MVP), a scalable multi-view transformer architecture that directly reconstructs large 3D scenes from tens

Multimodal Action Diffusion for Robust End-to-End Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.02105v1 Announce Type: new Abstract: End-to-End Autonomous Driving (E2E-AD) systems have largely converged on predicting intermediate trajectory waypoints, delegating final control to hand-

MUSCLE-NET: Predicted-Multiscale-Aware Network for Pedestrian Trajectory Forecasting

AgentsDGX agent

arXiv:2606.00471v1 Announce Type: new Abstract: Accurate pedestrian trajectory prediction is essential for safe navigation in autonomous driving and intelligent transportation systems. Despite substan

NAPPure: Adversarial Purification for Robust Image Classification under Non-Additive Perturbations

ApplicationsDGX agent

arXiv:2510.14025v2 Announce Type: replace Abstract: Adversarial purification has achieved great success in combating adversarial image perturbations, which are usually assumed to be additive. However,

Neural Acquisition & Representation of Subsurface Scattering

ApplicationsDGX agent

arXiv:2606.02292v1 Announce Type: new Abstract: We present a method to acquire and estimate the sub-surface scattering properties of light transport at a highly detailed level by learning the pixel fo

Non-Learning Low-Light Stereo Vision

Model ReleasesDGX agent

arXiv:2606.00379v1 Announce Type: new Abstract: We present a non-learning stereo framework for disparity estimation from severely noisy images. Using the Field of Junctions (FoJ), it retains coarse vi

Normality-Preserving Continual Industrial Anomaly Detection via Orthogonal LoRA Banks

ResearchDGX agent

arXiv:2606.02042v1 Announce Type: new Abstract: Continual industrial anomaly detection with diffusion models suffers from historical normality prior drift and catastrophic forgetting. Existing continu

Not All Points Are Equal: Uncertainty-Aware 4D LiDAR Scene Synthesis

TutorialsDGX agent

arXiv:2606.02510v1 Announce Type: new Abstract: Constructing faithful 4D worlds from LiDAR-acquired sequences is crucial for embodied AI, yet current generative frameworks apply uniform modeling capac

ObjEmbed: Towards Universal Multimodal Object Embeddings

SafetyDGX agent

arXiv:2602.01753v3 Announce Type: replace Abstract: Aligning objects with corresponding textual descriptions is a fundamental challenge and a realistic requirement in vision-language understanding. Wh

One Channel to Rule Them All: Rethinking Input Representation for Visual Place Recognition

Local AiDGX agent

arXiv:2606.00936v1 Announce Type: new Abstract: Visual Place Recognition (VPR) is fundamental to long-term robot localization and SLAM, yet current systems overwhelmingly rely on RGB input, implicitly

One-Shot Crowd Counting With Density Guidance For Scene Adaptation

Local AiDGX agent

arXiv:2602.07955v2 Announce Type: replace Abstract: Crowd scenes captured by cameras at different locations vary greatly, and existing crowd models have limited generalization for unseen surveillance

OP-LoRA: The Blessing of Dimensionality

ResearchDGX agent

arXiv:2412.10362v2 Announce Type: replace-cross Abstract: Low-rank adapters (LoRA) enable finetuning of large models with only a small number of parameters. However, they often suffer from an ill-cond

OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing Imagery

Model ReleasesDGX agent

arXiv:2603.27645v2 Announce Type: replace Abstract: Open-vocabulary change detection (OVCD) seeks to recognize arbitrary changes of interest by enabling generalization beyond a fixed set of predefined

Optimizing 3D Gaussian Splatting via Point Cloud Upsampling

ResearchDGX agent

arXiv:2606.00450v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) is a technique for creating and rendering 3D scenes, however its performance depends heavily on the quality of initial seed

OptiWorld: Optimal Control for Video World Generation under Physical Constraints

ResearchDGX agent

arXiv:2606.00499v1 Announce Type: new Abstract: Video generation models are becoming a scalable form of world models, but they mainly generate plausible motion rather than proactively control or optim

PaCX-MAE: Physiology-Augmented Chest X-Ray Masked Autoencoder

ResearchDGX agent

arXiv:2606.01537v1 Announce Type: new Abstract: Clinical diagnosis often requires combining imaging with physiological measurements, yet deployed models typically operate on unimodal data. We present

PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion

ResearchDGX agent

arXiv:2606.01399v1 Announce Type: new Abstract: We present PAI-Studio, a new reference-conditioned video synthesis task that addresses a long-standing challenge in cinematic background replacement: ge

PairedGTA: Generating Driving Datasets for Controlled Photometric Shift Analysis

AgentsDGX agent

arXiv:2606.01192v1 Announce Type: new Abstract: Evaluating the performance of visual perception systems for autonomous driving is essential to ensure reliable operation across diverse environmental sc

PathAR: Structure-First Autoregressive Synthesis of Multimodal Pathology Images

ResearchDGX agent

arXiv:2606.01543v1 Announce Type: new Abstract: Data scarcity in multimodal pathology motivates unified generative models that synthesize modality-specific appearance while preserving anatomically coh

Pave-GRPO: Beyond Instantaneous Guidance through Principled Average Velocity Decomposition

SafetyDGX agent

arXiv:2606.01636v1 Announce Type: new Abstract: Post-training via Group Relative Policy Optimization (GRPO) has emerged as a powerful paradigm for aligning flow-based generative models with human pref

Paving the Way for Point Cloud Video Representation Learning Using A PDE Model

ResearchDGX agent

arXiv:2606.01604v1 Announce Type: new Abstract: Investigating spatial-temporal correlations, specifically how spatial points vary over time, is crucial for understanding point cloud videos. Traditiona

PerBite: A Curated Diagnostic Workflow for Bite-Aware Food Volume Estimation

ResearchDGX agent

arXiv:2606.02021v1 Announce Type: new Abstract: Can a visually plausible food mesh be trusted to estimate the volume of consumed food? method investigates this question using selected paired before- a

Perception First: A Frontier Native-Video Model with Self-Consistency for Implicit Video Question Answering

Model ReleasesDGX agent

arXiv:2606.01485v1 Announce Type: new Abstract: We describe our submission to the VRR Challenge @ CVPR 2026, built on the ImplicitQA / VRR-QA benchmark~ite{implicitqa}: multiple-choice video question

Personalized 3D Myocardial Infarct Geometry Reconstruction from Cine MRI for Cardiac Digital Twins

Model ReleasesDGX agent

arXiv:2606.01808v1 Announce Type: new Abstract: Accurate 3D geometric characterization of myocardial infarction (MI) is essential for building cardiac digital twins (CDTs) to precisely simulate infarc

PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation

SafetyDGX agent

arXiv:2606.01649v1 Announce Type: new Abstract: Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The chal

Physical Object Understanding with a Physically Controllable World Model

ResearchDGX agent

arXiv:2606.00439v1 Announce Type: new Abstract: A central challenge in visual intelligence is learning the physical structure of scenes from raw videos: how regions form objects and the laws that gove

Physics-Aware Linearized ADMM and Its Unrolling

ResearchDGX agent

arXiv:2606.01652v1 Announce Type: cross Abstract: Recently, partial differential equations (PDEs) have been used to directly model the measurement process in signal processing, although their evaluati

Physics from Video: Identifiability of Time-Invariant Second-Order ODEs under Minimal Trajectory Conditions

Model ReleasesDGX agent

arXiv:2606.00115v1 Announce Type: new Abstract: Bridging the gap between visual realism and physical understanding is a core challenge for video-based world models. We study the structural identifiabi

PillarDETR: YOLO-Backbone and RT-DETR Head for Real-Time 3D Object Detection

AgentsDGX agent

arXiv:2606.01757v1 Announce Type: new Abstract: Real-time 3D object detection is a critical component for the safe operation of autonomous driving systems and robotics. While LiDAR point clouds provid

PINNOCHIO: Physics-Informed Neural Network for Coupled Hyperelastic Interface-Volume Simulation in Orthognathic Surgery

ResearchDGX agent

arXiv:2606.01572v1 Announce Type: cross Abstract: Predicting patient-specific facial soft-tissue deformation is critical for iterative orthognathic surgery planning. However, current computational met

Pinterest Canvas: Large-Scale Image Generation at Pinterest

ResearchDGX agent

arXiv:2603.06453v2 Announce Type: replace Abstract: While recent image generation models demonstrate a remarkable ability to handle a wide variety of image generation tasks, this flexibility makes the

Places in the Wild: A Large, High-Resolution RAW Photograph Dataset for Ecologically Valid Vision Research

ResearchDGX agent

arXiv:2606.02481v1 Announce Type: new Abstract: Large image datasets have accelerated progress in cognitive neuroscience and computer vision. However, most datasets are low-resolution, internet-source

Plan-R1: Safe and Feasible Trajectory Planning as Language Modeling

Model ReleasesDGX agent

arXiv:2505.17659v4 Announce Type: replace-cross Abstract: Safe and feasible trajectory planning is critical for real-world autonomous driving systems. However, existing learning-based planners rely he

PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps

AgentsDGX agent

arXiv:2606.01788v1 Announce Type: new Abstract: Embodied visual navigation, where an agent perceives a complex environment and acts to reach a goal from raw sensory input, underpins a wide range of ap

Polaris: Scaling Up Instruction-Guided Image Generation Towards Millions of Personalized Style Needs

ResearchDGX agent

arXiv:2606.01858v1 Announce Type: new Abstract: Users increasingly expect image generation models to quickly adapt to highly diverse and personalized requirements, such as producing images with distin

Policy-based Foveated Imaging and Perception

SafetyDGX agent

arXiv:2606.02565v1 Announce Type: new Abstract: Ultra-high-resolution image sensors offer the potential to capture fine spatial details critical for many visual perception tasks, but acquiring and pro

Pool-Select-Refine: Allocation-Aware Generative Dataset Distillation with Soft-Label-Guided Latent Refinement

SafetyDGX agent

arXiv:2606.01920v1 Announce Type: new Abstract: Diffusion-based dataset distillation has recently emerged as a promising paradigm for condensing large-scale datasets into compact synthetic sets. By le

Positional Encodings Anchor Spatial Structure in Vision Transformers: A Geometric Perspective on Robustness

ResearchDGX agent

arXiv:2606.00124v1 Announce Type: new Abstract: Positional embeddings (PEs) in Vision Transformers (ViTs) are known to impact performance and robustness, but their role in shaping internal spatial rep

PRIMA: Boosting Animal Mesh Recovery with Biological Priors and Test-Time Adaptation

ResearchDGX agent

arXiv:2606.02366v1 Announce Type: new Abstract: We present PRIMA (*PRI*ors for *M*esh *A*daptation), a framework for robust 3D quadruped mesh recovery under severe species and pose imbalance. Existing

Princeton365: A Diverse Dataset with Accurate Camera Pose

Model ReleasesDGX agent

arXiv:2506.09035v2 Announce Type: replace Abstract: We introduce Princeton365, a large-scale diverse dataset of 365 videos with accurate camera pose. Our dataset bridges the gap between accuracy and d

Private and Stable Test-Time Adaptation with Differential Privacy

ResearchDGX agent

arXiv:2606.01908v1 Announce Type: cross Abstract: Test-time adaptation (TTA) can reduce error on new and different data by updating the model on these inputs during inference. However, these updates r

ProtoAda: Prototype-Guided Adaptive Adapter Expansion and Geometric Consolidation for Multimodal Continual Instruction Tuning

ApplicationsDGX agent

arXiv:2606.02576v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) achieve strong performance through instruction tuning, but real-world deployment requires them to continually a

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning

Model ReleasesDGX agent

arXiv:2507.08064v3 Announce Type: replace-cross Abstract: As multimedia content expands, the demand for unified multimodal retrieval (UMR) in real-world applications increases. Recent work leverages m

Quality-Guided Semi-Supervised Learning for Medical Image Segmentation

ResearchDGX agent

arXiv:2606.01753v1 Announce Type: new Abstract: Training accurate medical image segmentation models requires large amounts of densely annotated data, which is costly and time-consuming to obtain. Semi

Question-Aware Evidence Ledgers for Video Relational Reasoning

Model ReleasesDGX agent

arXiv:2606.02506v1 Announce Type: new Abstract: The VRR-QA challenge evaluates visual relational reasoning in videos, where answers often depend on implicit spatial relations, event boundaries, target

R^3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking

ResearchDGX agent

arXiv:2606.01113v1 Announce Type: new Abstract: The CoVR-R challenge evaluates composed video retrieval, where a system must retrieve a target video from a large gallery given a reference video and a

RAIGen: Rare Attribute Identification in Text-to-Image Generative Models

SafetyDGX agent

arXiv:2602.06806v2 Announce Type: replace Abstract: Text-to-image diffusion models achieve impressive generation quality but inherit and amplify training-data biases, skewing coverage of semantic attr

Random Erasing vs. Model Inversion: A Promising Defense or a False Hope?

ResearchDGX agent

arXiv:2409.01062v4 Announce Type: replace-cross Abstract: Model Inversion (MI) attacks pose a significant privacy threat by reconstructing private training data from machine learning models. While exi

← Previous
1…102103104105106…211
Next →