AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
1 Jul 2026

MemLearner: Learning to Query Context memory for Video World Models

ApplicationsDGX agent

arXiv:2606.31734v1 Announce Type: new Abstract: Video World Models are interactive video generation models that predict future world states based on user actions and history video frames. A critical c

Mesh BDF: Barycentric Dominance Field for 3D Native Mesh Generation

ResearchDGX agent

arXiv:2606.31777v1 Announce Type: new Abstract: Autoregressive (AR) modeling has recently achieved remarkable progress in native 3D mesh generation, largely due to its natural ability to handle variab

MetricHMSR:Metric Human Mesh and Scene Recovery from Monocular Images

Local AiDGX agent

arXiv:2506.09919v4 Announce Type: replace Abstract: We introduce MetricHMSR, a novel framework for recovering metric human meshes and 3D scenes from a single monocular image. Existing methods struggle


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

MS-Resampler: Multi-Scope Visual Resampling for Efficient Multimodal LLMs

ResearchDGX agent

arXiv:2606.31383v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) typically employ resampling-based projectors to transform dense visual features into a compact token sequence f

MSNN-LINet: Cross-Modal Learning via Continuous Linear Integration

ResearchDGX agent

arXiv:2606.31135v1 Announce Type: new Abstract: We present LINet (Linear Integration Network), a Multi-Stream Neural Network (MSNN) for RGB-D scene classification. Current multi-modal architectures tr

Multi-Channel Uncertainty-Weighted Score Matching for Conditional Diffusion in Medical UDA

ResearchDGX agent

arXiv:2509.22476v2 Announce Type: replace Abstract: Robust medical image segmentation across modalities remains challenging due to severe domain shifts and the lack of target-domain labels. While diff

Multimodal Benchmark for Safety Assessment in Industrial Inspection Scenarios

Model ReleasesDGX agent

arXiv:2601.21173v2 Announce Type: replace-cross Abstract: With the rapid development of industrial intelligence and unmanned inspection, reliable perception and safety assessment for AI systems in com

MuSViT: A Foundation Vision Model for Sheet Music Representation

ApplicationsDGX agent

arXiv:2606.31811v1 Announce Type: new Abstract: Foundation models have transformed vision and language processing by providing rich, reusable representations that transfer across diverse tasks. Sheet

MV-GEL: Language-Driven Multi-View Geometric Entity Localization on Meshes

ResearchDGX agent

arXiv:2606.31533v1 Announce Type: new Abstract: Identifying and grounding precise geometric entities, such as edges, planar regions, and curved surfaces within 3D objects, is foundational to computer-

No Adaptation Without Observation: Observability-Constrained Test-Time Prompt Tuning for LiDAR Semantic Segmentation

SafetyDGX agent

arXiv:2606.30937v1 Announce Type: new Abstract: LiDAR semantic segmentation often degrades under real-world deployment due to evolving sensing conditions, while collecting new annotations for retraini

No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs

Model ReleasesDGX agent

arXiv:2606.31933v1 Announce Type: new Abstract: We introduce VidPair-Halluc, a new benchmark for evaluating video hallucination in large video models (LVMs) under rigorous and controlled conditions. U

No Prompt, No Leaks: A Robust Generative Steganography Framework via Prompt-Free Diffusion

ResearchDGX agent

arXiv:2606.31427v1 Announce Type: new Abstract: Generative image steganography synthesizes stego images directly from secret information to achieve inherent security advantages. Latent Diffusion Model

NURBS Splatting: A Unified Differentiable Rendering Framework for Vector Graphics

Model ReleasesDGX agent

arXiv:2606.31764v1 Announce Type: cross Abstract: Differentiable rendering of planar rational splines remains largely underexplored, despite their widespread use in vector graphics and design. Existin

Off the Rails: Hijacking the Scoring Head in Generative End-to-End Driving Planners with Safety-Violating Adversarial Perturbations

SafetyDGX agent

arXiv:2606.30807v1 Announce Type: cross Abstract: Generative models have recently seen rapid adoption in End-to-End (E2E) autonomous driving (AD), with diffusion-based denoising and vocabulary-based r

On the Role of Rotation Equivariance in Monocular 2D-to-3D Human Pose Lifting

ResearchDGX agent

arXiv:2601.13913v2 Announce Type: replace Abstract: Estimating 3D from 2D is one of the central tasks in computer vision. In this work, we consider the monocular setting, i.e. single-view input, for 3

One Video, One World: Turning Monocular Video into Physical 4D Scenes

Model ReleasesDGX agent

arXiv:2606.31388v1 Announce Type: new Abstract: We introduce extbf{OVOW}, the first training-free system that reconstructs instance-level, simulation-ready 4D mesh scenes from a single monocular video

Online TT-ALS for Streaming Tensor Decomposition with Incremental Orthogonalization

ResearchDGX agent

arXiv:2606.31061v1 Announce Type: cross Abstract: Tensor Train (TT) decomposition is a powerful technique for analyzing high-dimensional data. Existing algorithms for computing TT decompositions can b

PA-VAD: Diffusion-Based Pseudo-Only Video Anomaly Detection via Domain-Aligned Memory Updates

SafetyDGX agent

arXiv:2512.06845v2 Announce Type: replace Abstract: Deploying video anomaly detection (VAD) in the real world is often constrained by the scarcity, privacy, and cost of collecting real abnormal footag

Pano3D: Unified 3D Reconstruction and Panoptic Segmentation

ResearchDGX agent

arXiv:2606.14307v2 Announce Type: replace Abstract: Recent advances in 3D feedforward reconstruction neural networks have achieved remarkable success in dense reconstruction from images without any ca

Patient-Level Elbow Abnormality Detection: Leakage-Aware Evaluation of Learned Preprocessing, Calibration, and Triage-Oriented Operating Points

ResearchDGX agent

arXiv:2606.31348v1 Announce Type: new Abstract: In this study, we examine learned preprocessing pipelines in the context of triage-oriented orthopedic abnormality detection task using elbow radiograph

Phantom: A Unified Face-Swap Deepfake Protection Framework with Latent and Spatial Constraints

TutorialsDGX agent

arXiv:2606.31703v1 Announce Type: new Abstract: Face-swapping deepfakes pose an escalating threat to personal privacy by enabling unauthorized identity manipulation. While adversarial approaches have

Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer

ResearchDGX agent

arXiv:2511.19778v2 Announce Type: replace Abstract: Rotary positional embeddings (RoPE) are widely used in diffusion transformers (DiTs) to encode spatial relationships, yet their behavior with mixed-

{Phi}eat: Physically Grounded Material Feature Representation

ResearchDGX agent

arXiv:2511.11270v2 Announce Type: replace Abstract: While foundation models have emerged as general-purpose visual backbones, their representations are primarily optimized for semantics and lack expli

PhotoQuilt: Training-Free Arbitrary-Resolution Photomosaics via Bootstrapped Tiled Denoising

ResearchDGX agent

arXiv:2606.30968v1 Announce Type: new Abstract: Photomosaics are large images whose local regions are seen as independent tiles while their overall arrangement forms a coherent scene. Generating them

PiLoT v2: Pixel-to-Orthogonal Map Alignment for Free-view UAV Geo-localization

SafetyDGX agent

arXiv:2606.31098v1 Announce Type: new Abstract: Real-time, drift-free UAV geo-localization is essential for autonomous missions in GNSS-denied environments. The pioneering system, PiLoT, achieves high

Planar-SfM: Camera Pose Estimation via Homography Graph Embeddings

Model ReleasesDGX agent

arXiv:2606.31979v1 Announce Type: new Abstract: Structure from Motion (SfM) systems traditionally struggle with planar scenes, where standard epipolar geometry-based methods become degenerate. Rather

PointSplat: Compact Gaussian Splatting via Human-Centric Prediction

TutorialsDGX agent

arXiv:2606.32036v1 Announce Type: new Abstract: Producing 3D human representations from input views on the fly is essential for immersive live streaming systems, where representation compactness is as

PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh Generation

Model ReleasesDGX agent

arXiv:2606.30673v1 Announce Type: cross Abstract: Autoregressive Transformers dominate high-quality mesh generation by producing artist-worthy topologies, yet their inherent sequential decoding induce

PoseGravity: Pose Estimation from Points and Lines with Axis Prior

ResearchDGX agent

arXiv:2405.12646v3 Announce Type: replace Abstract: This paper presents a new algorithm to estimate absolute camera pose given an axis of the camera's rotation matrix. Current algorithms solve the pro

Practical High-Fidelity Novel-View Synthesis of Mounted Lepidoptera

ResearchDGX agent

arXiv:2606.31679v1 Announce Type: cross Abstract: Mounted butterflies are among the most striking objects in natural history collections. However, their beauty is notoriously hard to digitize in 3D: t

PriorEye: Geospatial Visual Priors for End-to-End Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.31830v1 Announce Type: new Abstract: Most end-to-end autonomous driving methods rely solely on instantaneous sensor observations, limiting them to reactive behavior without the anticipatory

PrISM-IQA: Image Quality Assessment Made Practical for Smartphone Photography

Model ReleasesDGX agent

arXiv:2606.31626v1 Announce Type: new Abstract: Existing smartphone image quality assessment (IQA) methods commonly reduce perceptual quality to a single score. However, this scalar formulation is poo

PRISM: Latent Composition Consistency for Single-Image Reflection Removal

ResearchDGX agent

arXiv:2606.31513v1 Announce Type: new Abstract: Single-image reflection removal (SIRR) seeks to recover the transmission layer from a mixture corrupted by reflections -- a severely ill-posed problem.

Proxy-GS: Unified Occlusion Priors for Training and Inference in Structured 3D Gaussian Splatting

ResearchDGX agent

arXiv:2509.24421v5 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has emerged as an efficient approach for achieving photorealistic rendering. Recent MLP-based variants further improve

PSHuman: Photorealistic Single-image 3D Human Reconstruction using Cross-Scale Multiview Diffusion and Explicit Remeshing

Local AiDGX agent

arXiv:2409.10141v3 Announce Type: replace Abstract: Detailed and photorealistic 3D human modeling is essential for various applications and has seen tremendous progress. However, full-body reconstruct

RASR: Retrieval-Augmented Semantic Reasoning for Fake News Video Detection

ResearchDGX agent

arXiv:2604.06687v2 Announce Type: replace Abstract: Multimodal fake news video detection is a crucial research direction for maintaining the credibility of online information. Existing studies primari

RCL-Mamba: A Dual-domain State Space Model for Measurement-oriented Image Restoration in Rotational Sparse-View Scanning Computed Laminography

ApplicationsDGX agent

arXiv:2606.31353v1 Announce Type: new Abstract: Rotational Scanning Computed Laminography (RCL) is widely utilized for the Non-Destructive Testing(NDT) of large planar components. However, to facilita

Reasoning-aware Speculative Decoding for Efficient Vision-Language-Action Models in Autonomous Driving

AgentsDGX agent

arXiv:2606.31160v1 Announce Type: new Abstract: Modern Vision-Language-Action (VLA) planners for autonomous driving emit a chain-of-causation (CoC) reasoning step before producing a trajectory. The re

Reasoning in machine vision by learning fast and slow thinking

Local AiDGX agent

arXiv:2506.22075v2 Announce Type: replace Abstract: Reasoning is a hallmark of human intelligence, enabling adaptive decision-making in complex unfamiliar scenarios. In contrast, machine intelligence

REDI: Corpus Aware Patch Ranking for DINOv3 Token Reduction

ResearchDGX agent

arXiv:2606.31676v1 Announce Type: new Abstract: Most token reduction methods for Vision Transformers seek favorable tradeoffs between accuracy and efficiency by pruning, merging, or pooling patch toke

Reference-Free Image Quality Assessment for Virtual Try-On via Human Feedback

Model ReleasesDGX agent

arXiv:2603.13057v2 Announce Type: replace Abstract: As virtual try-on (VTON) systems become increasingly important in fashion e-commerce, there is a growing need for reliable reference-free evaluation

Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation

TutorialsDGX agent

arXiv:2506.23102v2 Announce Type: replace-cross Abstract: Current CT report generation frameworks predominantly rely on global feature representations, often failing to capture region-specific details

Registering the 4D Millimeter Wave Radar Point Clouds Via Generalized Method of Moments

ApplicationsDGX agent

arXiv:2508.02187v3 Announce Type: replace-cross Abstract: 4D millimeter wave radars (4D radars) are new emerging sensors that provide point clouds of objects with both position and radial velocity mea

RESOLVE: A Multi-Resolution and Multi-Modal Dataset for Roadside Cooperative Perception

Model ReleasesDGX agent

arXiv:2606.31895v1 Announce Type: new Abstract: LiDAR has increasingly been integrated into traffic cameras to expand coverage and mitigate occlusion in roadside cooperative perception. However, how u

Rethinking Foundation Model Collaboration: Enhancing Specialized Models through Proxy Task Reasoning

ResearchDGX agent

arXiv:2606.31157v1 Announce Type: new Abstract: Foundation models are increasingly integrated into embodied intelligence systems, but directly assigning them structured prediction tasks requires preci

Rethinking the Role of Feature Engineering and Learning Strategies in Few-Shot Hidden Emotion Recognition

Model ReleasesDGX agent

arXiv:2606.31249v1 Announce Type: new Abstract: In this paper, we present the solution developed by our team, XInsight Lab, which achieved first place in Track 3 of the 4th EI-MIGA-IJCAI Challenge wit

RGBT-GroundBench: Visual Grounding Beyond RGB in Complex Real-World Scenarios

Model ReleasesDGX agent

arXiv:2512.24561v2 Announce Type: replace Abstract: Visual grounding (VG) localizes target objects in an image from natural-language expressions. In real-world perception, RGB cues often degrade under

Rhythm-Structured Predictive Learning for Remote Photoplethysmography

SafetyDGX agent

arXiv:2606.31736v1 Announce Type: new Abstract: Remote photoplethysmography (rPPG) estimates physiological signals from facial videos by analyzing subtle pulse induced skin color variations. Despite r

Robust Autonomous UAV Landing on Maritime Platforms via Multimodal Agentic AI and Active Wave Compensation

AgentsDGX agent

arXiv:2606.31613v1 Announce Type: new Abstract: Autonomous aerial inspection of marine infrastructure is frequently compromised by stochastic sea states, introducing risks of high-kinetic impacts, pos

SAMBA: A Scatter-Guided Masked Bidirectional Mamba Foundation Model for SAR Target Recognition

ResearchDGX agent

arXiv:2606.31668v1 Announce Type: new Abstract: Synthetic aperture radar automatic target recognition (SAR ATR) is critical for Earth observation and defense, but its practical deployment is constrain

Seeing Through the Weights: Privacy Leakage in Scene Coordinate Regression

ResearchDGX agent

arXiv:2606.31164v1 Announce Type: new Abstract: Scene Coordinate Regression (SCR) methods are increasingly adopted for visual localization. In these approaches, the scene is implicitly encoded within

Self-Supervised Temporal Regularization for Landmark-Based Cardiac Segmentation with Automatic AHA Regional Mapping

ResearchDGX agent

arXiv:2606.31785v1 Announce Type: new Abstract: Graph-based cardiac segmentation with implicit anatomical correspondences provides topological guarantees and population-level analysis capabilities, bu

Semantic-Aware Multiple Access via Spatial Redundancy Exploitation for Uplink-Dominant 6G Use Cases

ResearchDGX agent

arXiv:2606.31715v1 Announce Type: cross Abstract: Emerging uplink-dominant 6G use cases, such as cooperative vehicular streaming, require efficient transmission of high-volume visual data over limited

Semantic Occupancy Prediction with Dual Range-Voxel Representation

AgentsDGX agent

arXiv:2606.31688v1 Announce Type: new Abstract: LiDAR-based 3D semantic occupancy prediction, which aims to provide accurate and comprehensive scene representation, is crucial for autonomous driving s

SENSE-VAD: Sentient and Semantic Video Anomaly Detection for Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.31875v1 Announce Type: new Abstract: Autonomous vehicles (AVs) must navigate not only motion-based hazards but also socially complex situations whose danger is constituted by inter-agent re

ShellMaker: Language-Guided Exterior Completion under Structural Constraints

ResearchDGX agent

arXiv:2606.31680v1 Announce Type: new Abstract: Despite advances in indoor scene generation, synthesizing coherent building exteriors consistent with generated interiors remains largely unexplored. Ex

SHMoAReg: Spark Deformable Image Registration via Spatial Heterogeneous Mixture of Experts and Attention Heads

ResearchDGX agent

arXiv:2509.20073v2 Announce Type: replace Abstract: Encoder-Decoder architectures are widely used in deep learning-based Deformable Image Registration (DIR), where the encoder extracts multi-scale fea

Simple Supervision Is Hard to Beat: A Bitter Lesson from Sparse Target Labels in Domain-Adaptive Object Detection

ResearchDGX agent

arXiv:2606.30795v1 Announce Type: new Abstract: Source-free domain adaptive object detection adapts a source-trained detector to an unlabeled target domain, typically through teacher-student self-trai

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search

Model ReleasesDGX agent

arXiv:2606.31504v1 Announce Type: new Abstract: We present SimpleSearch-VL, an efficient, reliable, and practical framework for multimodal agentic search. Its core idea is to improve the agent's own s

SpectralSplats: Robust Differentiable Tracking via Spectral Moment Supervision

SafetyDGX agent

arXiv:2603.24036v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) enables real-time, photorealistic novel view synthesis, making it a highly attractive representation for model-based vi

← Previous
1…5859606162…209
Next →