AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
25 Jun 2026

Neural Network Quantization by Learning Low-Loss Subspaces

ResearchDGX agent

arXiv:2606.25087v1 Announce Type: new Abstract: Neural network quantization aims to find a discrete representation of parameters that preserves the performance of a full-precision (FP) model as faithf

Noise-Aware Boundary-Enhanced Generative Learning for Ultrasound Speckle Reduction

ResearchDGX agent

arXiv:2606.25009v1 Announce Type: new Abstract: Ultrasound is a non-invasive, real-time, and cost-effective imaging technique widely used in clinical diagnosis. However, its diagnostic efficacy is oft

OracleAnalyser: Analysing Implicit Semantics of Oracle Bone Scripts through MLLMs with Post-training

Model ReleasesDGX agent

arXiv:2606.25906v1 Announce Type: new Abstract: With the advancement of artificial intelligence, research on oracle bone scripts has entered a new era. However, existing methods and benchmarks remain


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

OrthoTrack: Continuous 6-DoF UAV Trajectory Estimation Anchored in Public Orthophotos

Model ReleasesDGX agent

arXiv:2606.25245v1 Announce Type: new Abstract: Continuous 6-DoF pose estimation is essential for autonomous UAV operations. Yet, existing visual odometry and SLAM methods accumulate drift and yield o

PatchINR: Patch-Based Implicit Neural Representations for Efficient and Scalable Inference

Model ReleasesDGX agent

arXiv:2606.25534v1 Announce Type: new Abstract: Implicit Neural Representation (INR) provides an effective approach for continuous signal modeling, but classical per-pixel inference results in quadrat

PhaseWin: An Efficient Search Algorithm for Faithful Visual Attribution

Local AiDGX agent

arXiv:2606.18008v2 Announce Type: replace Abstract: Visual attribution is a fundamental tool for interpreting modern vision and vision-language models, particularly when their decisions must be inspec

PhyGile: Physics-Prefix Guided Motion Generation for Agile General Humanoid Motion Tracking

ApplicationsDGX agent

arXiv:2603.19305v2 Announce Type: replace-cross Abstract: Humanoid robots are expected to execute agile and expressive whole-body motions in real-world settings. Existing text-to-motion generation mod

Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation

Model ReleasesDGX agent

arXiv:2606.25306v1 Announce Type: new Abstract: Video generation models are increasingly capable of producing realistic videos, but they still struggle to generate videos that follow basic physical la

Point Cloud Diffusion with Global and Local Reconstruction for Instance-Level 3D Anomaly Detection

SafetyDGX agent

arXiv:2606.25740v1 Announce Type: new Abstract: 3D anomaly detection in point clouds is critical for high-precision industrial manufacturing. Reconstruction-based methods have laid a strong foundation

Pre-Warm: Input-Conditioned Weight Initialization for Convolutional Neural Networks

Model ReleasesDGX agent

arXiv:2606.25256v1 Announce Type: new Abstract: We introduce Pre-Warm, a simple yet effective zero-training-cost method for data-conditioned initialization of the first convolutional layer. Before the

PRISM: Feed-Forward Single-Image 3D Reconstruction via Geometric Warp-Residual Modeling

Model ReleasesDGX agent

arXiv:2606.25430v1 Announce Type: new Abstract: Reconstructing 3D scenes from a single image is a fundamental challenge in computer vision, with broad applications in virtual reality, robotics, and co

Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need

Model ReleasesDGX agent

arXiv:2606.25956v1 Announce Type: new Abstract: Risk stratification for pulmonary embolism (PE) is critical for clinical decision-making. Stratification guidelines are based on patient medical records

Re-mixing Embeddings for Patient Augmentation in Data Scarce Multiple Instance Learning

ResearchDGX agent

arXiv:2606.25770v1 Announce Type: cross Abstract: Data scarcity is a major bottleneck in medical Multiple Instance Learning (MIL), especially for rare diseases or expensive modalities. We introduce a

ReaDy-Go: Real-to-Sim Dynamic 3D Gaussian Splatting Simulation for Environment-Specific Visual Navigation with Moving Obstacles

ApplicationsDGX agent

arXiv:2602.11575v3 Announce Type: replace-cross Abstract: Visual navigation models often struggle in real-world dynamic environments due to limited robustness to the sim-to-real gap and the difficulty

Reflective VLA: In-Context Action Consequences Make VLAs Generalize

SafetyDGX agent

arXiv:2606.25215v1 Announce Type: new Abstract: Most vision-language-action (VLA) models are reactive: they predict the next action from the current instruction and observation, implicitly assuming th

REViT: Roto-reflection Equivariant Convolutional Vision Transformer

ResearchDGX agent

arXiv:2606.25318v1 Announce Type: new Abstract: In this paper, we propose a discrete roto-reflection group equivariant vision transformer with convolutional attention. Roto-reflection equivariant netw

RoboAtlas: Contextual Active SLAM

Model ReleasesDGX agent

arXiv:2606.26046v1 Announce Type: cross Abstract: We present RoboAtlas, a contextual Active SLAM framework that adaptively balances geometric exploration and semantic reasoning using a scalable 3D sem

S^{2}-FracMix: Label-Preserving Self-Saliency Mixup Augmentation

ResearchDGX agent

arXiv:2606.25784v1 Announce Type: new Abstract: Data augmentation is known to improve generalization of deep visual models. Recent methods favor mixup strategies that generate interpolated samples to

SAC^2-Net: Semantic Anchoring and Complementary-Consensus Fusion for Multimodal Micro-Expression Recognition

SafetyDGX agent

arXiv:2606.25542v1 Announce Type: new Abstract: Micro-expression recognition (MER) is challenging due to subtle facial movements, limited data, and the ambiguous relationship between Action Units (AUs

ScaleHP: Estimating Hand Pose in Metric Space

ApplicationsDGX agent

arXiv:2606.25619v1 Announce Type: new Abstract: Accurate metric-space hand pose estimation (HPE) is essential for immersive human-computer interaction and robotics. However, most existing methods pred

ScalingAR: Scaling Confidence for Autoregressive Image Generation

SafetyDGX agent

arXiv:2509.26376v3 Announce Type: replace Abstract: Test-time strategies have shown remarkable success in improving large language models, but their application to next-token prediction (NTP) autoregr

Semantic Allocation in Ordered Bottlenecks: Predictive Residual Inference for Visual Representation Learning

ResearchDGX agent

arXiv:2606.25232v1 Announce Type: cross Abstract: Ordered bottlenecks aim to provide utility at flexible budgets by assigning coarse information to early tokens and task-relevant detail to later ones.

SEMIR: Topology-Preserving Graph Minors for Thin-Structure Segmentation

ResearchDGX agent

arXiv:2606.24935v1 Announce Type: new Abstract: Thin-structure segmentation--power lines, cracks, lane markings at 1-3 pixel width--requires preserving connectivity that standard representations precl

Shift Variant Image Degradation and Restoration Using Singular Value Decomposition

ResearchDGX agent

arXiv:2606.25818v1 Announce Type: new Abstract: Shift-variant image degradation is frequently encountered in practical imaging systems where the point spread function (PSF) varies across the image fie

ShutterMuse: Capture-Time Photography Guidance with MLLMs

Model ReleasesDGX agent

arXiv:2606.25763v1 Announce Type: new Abstract: Real-world photography requires capture-time guidance for both camera framing and subject pose. Yet existing aesthetic cropping benchmarks mainly evalua

SparseGS: Sparse View Synthesis using 3D Gaussian Splatting

ResearchDGX agent

arXiv:2312.00206v4 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has recently enabled real-time rendering of unbounded 3D scenes for novel view synthesis. However, this technique requi

Spatio-Temporal Mixture-of-Modality-Experts Diffusion for Quantitative DCE-MRI Synthesis from Incomplete MR Sequences

Model ReleasesDGX agent

arXiv:2606.25535v1 Announce Type: new Abstract: Quantitative maps from dynamic contrast-enhanced MRI (DCE-MRI) are essential for tumor assessment but are often unavailable due to contrast-agent risks

SplatPainter: Interactive Authoring of 3D Gaussians from 2D Edits via Test-Time Training

ResearchDGX agent

arXiv:2512.05354v2 Announce Type: replace Abstract: The rise of 3D Gaussian Splatting has revolutionized photorealistic 3D asset creation, yet a critical gap remains for their interactive refinement a

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity

Model ReleasesDGX agent

arXiv:2606.25634v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown remarkable progress in single-image perception, yet their ability to reason about complex cross-view

State Space Models Meet Remote Sensing: A Survey

ResearchDGX agent

arXiv:2606.25329v1 Announce Type: new Abstract: State Space Models (SSMs), designed for long-range modeling, offer linear computational complexity and strong capabilities in capturing long-range depen

Steering Vision-Language Models with Joint Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2606.25657v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) have shown promise for analyzing language models, but applying them to vision-language models (VLMs) often yields representat

Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds

ResearchDGX agent

arXiv:2606.25234v1 Announce Type: new Abstract: What is the geometry of a visual percept? The most widely used protocols for decomposing neural network representations into interpretable parts treat c

StyleFusion360: View-Consistent Head Stylization via Adaptive Style Modulation

ResearchDGX agent

arXiv:2511.22411v2 Announce Type: replace Abstract: 3D head stylization enables expressive reimagining of human faces for creative visual experiences in digital media. Existing 3D-aware methods often

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery

Model ReleasesDGX agent

arXiv:2606.25905v1 Announce Type: new Abstract: We introduce SurgAtlas, the largest surgical video-language dataset to date, comprising 15,291 videos (2,391 hours) spanning 18 surgical specialties and

TACO: Towards Task-Consistent Open-Vocabulary Adaptation in Video Recognition

Model ReleasesDGX agent

arXiv:2606.25478v1 Announce Type: new Abstract: Adapting CLIP for open-vocabulary video recognition necessitates a delicate balance between newly acquired video knowledge and the pretrained generaliza

Taxonomy-aware deep learning for hierarchical marine species classification in underwater imagery

SafetyDGX agent

arXiv:2606.25989v1 Announce Type: new Abstract: Automated classification of marine species from underwater imagery is essential for scalable ocean biodiversity monitoring and conservation policy. Exis

Teach-to-Reason: Competition-Guided Reasoning with a Self-Improving Teacher

ResearchDGX agent

arXiv:2606.25407v1 Announce Type: new Abstract: Chest X-ray visual question answering (CXR VQA) requires models not only to predict correct answers, but also to produce reliable medical reasoning. How

Tensorion: A Tensor-Aware Generalization of the Muon Optimizer

Model ReleasesDGX agent

arXiv:2606.25975v1 Announce Type: cross Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight

TensorLDM: A Component-Wise Latent Diffusion Model for Volumetric DTI Reconstruction from Sparse DWIs

ResearchDGX agent

arXiv:2606.25545v1 Announce Type: new Abstract: Reconstructing diffusion tensors from sparse DWIs is critical for accelerating Diffusion Tensor Imaging (DTI) in clinical settings, yet current deep lea

Test-Time Adaptation in Optical Coherence Tomography Using Trajectory-Aligned Time-Independent Flow

ApplicationsDGX agent

arXiv:2606.18876v2 Announce Type: replace Abstract: Optical coherence tomography (OCT) is essential in ophthalmology, but inconsistent image quality especially in low-cost devices hinders automated an

To View Transform or Not to View Transform: NeRF-based Pre-training Perspective

AgentsDGX agent

arXiv:2603.28090v2 Announce Type: replace Abstract: Neural radiance fields (NeRFs) have emerged as a prominent pre-training paradigm for vision-centric autonomous driving, which enhances 3D geometry a

Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding

Model ReleasesDGX agent

arXiv:2606.25160v1 Announce Type: cross Abstract: The rapid rise of Vision-Language Models (VLMs) in egocentric visual understanding has made low-latency inference in human-robot collaborative (HRC) t

Towards a Dynamic and Fixed-budget Memory Bank for Efficient Streaming Video Understanding

ResearchDGX agent

arXiv:2606.25658v1 Announce Type: new Abstract: Currently, streaming video understanding is still a daunting task for existing multimodal large language models (MLLMs). Its difficulties not only lie i

Transferable Attack against Face Swapping in an Extended Space

ResearchDGX agent

arXiv:2606.25376v1 Announce Type: new Abstract: Although deep Face Swapping (FS) models may benefit the entertainment industry, they pose severe threats to privacy and security. Existing protections,

TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs

Model ReleasesDGX agent

arXiv:2606.26029v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) demonstrate strong performance on standard visual question answering benchmarks, yet their scalability under co

TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy

ResearchDGX agent

arXiv:2606.26092v1 Announce Type: new Abstract: While Video Virtual Try-on (VVT) has achieved remarkable progress in synthesizing realistic garment overlays on dynamic subjects, existing paradigms rem

TTSA3R: Training-Free Temporal-Spatial Adaptive Persistent State for Streaming 3D Reconstruction

SafetyDGX agent

arXiv:2601.22615v3 Announce Type: replace Abstract: Streaming recurrent models enable efficient 3D reconstruction by maintaining persistent state representations. However, they suffer from catastrophi

UniTeD: Unified Temporal Diffusion for Joint Perception and Planning in Autonomous Driving

AgentsDGX agent

arXiv:2606.25736v1 Announce Type: new Abstract: Diffusion models have shown strong potential for multi-modal planning in end-to-end autonomous driving. However, most existing methods confine diffusion

USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning

Model ReleasesDGX agent

arXiv:2606.25880v1 Announce Type: new Abstract: Embodied Visual Tracking (EVT) requires an agent to continuously follow a specified target while actively moving through dynamic environments. However,

V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning

Model ReleasesDGX agent

arXiv:2606.25319v1 Announce Type: new Abstract: Fine-grained visual reasoning requires multimodal large language models (MLLMs) to identify task-relevant visual evidence and ground their reasoning in

VENI: Variational Encoder for Natural Illumination

ResearchDGX agent

arXiv:2601.14079v2 Announce Type: replace Abstract: Inverse rendering is an ill-posed problem, but priors such as illumination priors can help simplify it. Existing work either disregards the spherica

VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction

SafetyDGX agent

arXiv:2509.19297v3 Announce Type: replace Abstract: Feed-forward 3D Gaussian Splatting (3DGS) has emerged as a highly effective solution for novel view synthesis. Existing methods predominantly rely o

VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks

Model ReleasesDGX agent

arXiv:2606.25592v1 Announce Type: new Abstract: Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfac

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

ResearchDGX agent

arXiv:2606.25041v1 Announce Type: new Abstract: We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency, full-duplex

What Does the Brain See? Multiview Neural Representations to Demystify the Brain-Visual Alignment

Model ReleasesDGX agent

arXiv:2606.25718v1 Announce Type: new Abstract: Zero-shot visual decoding from electroencephalography (EEG) aims to infer visual semantics from non-invasive neural recordings, but remains challenging

Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety

SafetyDGX agent

arXiv:2606.25034v1 Announce Type: new Abstract: General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversaria

24 Jun 2026

3D Masked Autoencoders are Robust Learners of Volumetric and Multimodal Cellular Representations for Microscopy

Local AiDGX agent

arXiv:2606.23964v1 Announce Type: cross Abstract: Self-supervised learning in fluorescence microscopy often relies on 2D projections, despite the inherently three-dimensional nature of cells. We prese

3DCarGen: Scalable 3D Car Generation via 3D-consistent Multi-view Synthesis

Model ReleasesDGX agent

arXiv:2606.24257v1 Announce Type: new Abstract: High-quality 3D vehicle assets are essential for autonomous driving simulation. Although multi-view diffusion-based paradigms enable controllable single

A Benchmark of State-Space Models vs. Transformers and BiLSTM-based Models for Historical Newspaper OCR

Model ReleasesDGX agent

arXiv:2604.00725v2 Announce Type: replace Abstract: End-to-end OCR for historical newspapers remains challenging, as models must handle long text sequences, degraded print quality, and complex layouts

A Dual Edge Spatial Jacobian Image Graph for Interpretable Diabetic Retinopathy Grading

ResearchDGX agent

arXiv:2606.24168v1 Announce Type: cross Abstract: Automated diabetic retinopathy (DR) grading from colour fundus photographs can achieve strong predictive performance, but clinical interpretation requ

← Previous
1…7273747576…211
Next →