AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Model Releases

Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity

DGX agent

arXiv:2606.25343v1 Announce Type: new Abstract: Vision Language Models have achieved near-human performance on single-document Visual Question Answering, yet their effectiveness degrades significantly

model-releasesarxiv-cs-cv
25 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

KidRisk: Benchmark Dataset for Children Dangerous Action Recognition

DGX agent

arXiv:2606.25298v1 Announce Type: new Abstract: Children are naturally energetic, and during their spontaneous activities, they often encounter potentially dangerous situations, especially when lackin

model-releasesarxiv-cs-cv
25 Jun 2026
Research

Latent Space Analysis for Interpretable Uncertainty in Melanoma Classification

DGX agent

arXiv:2506.18414v3 Announce Type: replace Abstract: Melanoma is a highly aggressive skin cancer, making early and accurate diagnosis critical. While deep learning excels in skin lesion classification,

researcharxiv-cs-cv
25 Jun 2026
Safety

Learning Action Priors for Cross-embodiment Robot Manipulation

DGX agent

arXiv:2606.26095v1 Announce Type: cross Abstract: Most Vision-Language-Action (VLA) models build on a Vision-Language Model (VLM) backbone by attaching an action module and optimizing the full policy

safetyarxiv-cs-cv
25 Jun 2026
Model Releases

LEVIRDet: A Million-Scale 159-Category Dataset and Foundation Model for Universal Remote Sensing Object Detection

DGX agent

arXiv:2606.25312v1 Announce Type: new Abstract: Remote sensing object detection has advanced rapidly with the development of large-scale benchmarks and modern detection architectures. However, existin

model-releasesarxiv-cs-cv
25 Jun 2026
Applications

LinStereo: Linear-Complexity Global Attention for Multi-Scale Iterative Stereo Matching

DGX agent

arXiv:2606.25437v1 Announce Type: new Abstract: Existing Vision Foundation Model (VFM)-based iterative stereo pipelines under-exploit three information pathways: multi-scale backbone features are coll

applicationsarxiv-cs-cv
25 Jun 2026
Research

MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation

DGX agent

arXiv:2606.26016v1 Announce Type: new Abstract: Normalizing Flows (NFs) are powerful generative models capable of exact density estimation and sampling. However, their strict invertibility often force

researcharxiv-cs-cv
25 Jun 2026
Research

Minimalist Preprocessing Approach for Image Synthesis Detection

DGX agent

arXiv:2606.25297v1 Announce Type: new Abstract: Generative models have significantly advanced image generation, resulting in synthesized images that are increasingly indistinguishable from authentic o

researcharxiv-cs-cv
25 Jun 2026
Tutorials

MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning

DGX agent

arXiv:2606.25225v1 Announce Type: new Abstract: Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning. Since audio and visual strea

tutorialsarxiv-cs-cv
25 Jun 2026
Applications

MRI2Rep: Autoregressive Structured Report Generation for 3D Liver MRI

DGX agent

arXiv:2606.25279v1 Announce Type: new Abstract: Manual reporting of 3D MRI studies is time-consuming, yet end-to-end structured report generation for 3D liver MRI remains underexplored due to volumetr

applicationsarxiv-cs-cv
25 Jun 2026
Research

MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation

DGX agent

arXiv:2606.26087v1 Announce Type: new Abstract: Synthesizing a novel-view video from a monocular reference video along a target camera trajectory requires both geometric consistency and motion fidelit

researcharxiv-cs-cv
25 Jun 2026
Applications

Naturalness Predicts but Does Not Cause Transferability in Image Encodings of Real-World Streams

DGX agent

arXiv:2606.25844v1 Announce Type: new Abstract: A common practice converts a one-dimensional signal into an image so that a vision backbone pretrained on natural photographs can be reused for recognit

applicationsarxiv-cs-cv
25 Jun 2026
Research

Neural Network Quantization by Learning Low-Loss Subspaces

DGX agent

arXiv:2606.25087v1 Announce Type: new Abstract: Neural network quantization aims to find a discrete representation of parameters that preserves the performance of a full-precision (FP) model as faithf

researcharxiv-cs-cv
25 Jun 2026
Research

Noise-Aware Boundary-Enhanced Generative Learning for Ultrasound Speckle Reduction

DGX agent

arXiv:2606.25009v1 Announce Type: new Abstract: Ultrasound is a non-invasive, real-time, and cost-effective imaging technique widely used in clinical diagnosis. However, its diagnostic efficacy is oft

researcharxiv-cs-cv
25 Jun 2026
Model Releases

OracleAnalyser: Analysing Implicit Semantics of Oracle Bone Scripts through MLLMs with Post-training

DGX agent

arXiv:2606.25906v1 Announce Type: new Abstract: With the advancement of artificial intelligence, research on oracle bone scripts has entered a new era. However, existing methods and benchmarks remain

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

OrthoTrack: Continuous 6-DoF UAV Trajectory Estimation Anchored in Public Orthophotos

DGX agent

arXiv:2606.25245v1 Announce Type: new Abstract: Continuous 6-DoF pose estimation is essential for autonomous UAV operations. Yet, existing visual odometry and SLAM methods accumulate drift and yield o

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

PatchINR: Patch-Based Implicit Neural Representations for Efficient and Scalable Inference

DGX agent

arXiv:2606.25534v1 Announce Type: new Abstract: Implicit Neural Representation (INR) provides an effective approach for continuous signal modeling, but classical per-pixel inference results in quadrat

model-releasesarxiv-cs-cv
25 Jun 2026
Local Ai

PhaseWin: An Efficient Search Algorithm for Faithful Visual Attribution

DGX agent

arXiv:2606.18008v2 Announce Type: replace Abstract: Visual attribution is a fundamental tool for interpreting modern vision and vision-language models, particularly when their decisions must be inspec

local-aiarxiv-cs-cv
25 Jun 2026
Applications

PhyGile: Physics-Prefix Guided Motion Generation for Agile General Humanoid Motion Tracking

DGX agent

arXiv:2603.19305v2 Announce Type: replace-cross Abstract: Humanoid robots are expected to execute agile and expressive whole-body motions in real-world settings. Existing text-to-motion generation mod

applicationsarxiv-cs-cv
25 Jun 2026
Model Releases

Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation

DGX agent

arXiv:2606.25306v1 Announce Type: new Abstract: Video generation models are increasingly capable of producing realistic videos, but they still struggle to generate videos that follow basic physical la

model-releasesarxiv-cs-cv
25 Jun 2026
Safety

Point Cloud Diffusion with Global and Local Reconstruction for Instance-Level 3D Anomaly Detection

DGX agent

arXiv:2606.25740v1 Announce Type: new Abstract: 3D anomaly detection in point clouds is critical for high-precision industrial manufacturing. Reconstruction-based methods have laid a strong foundation

safetyarxiv-cs-cv
25 Jun 2026
Model Releases

Pre-Warm: Input-Conditioned Weight Initialization for Convolutional Neural Networks

DGX agent

arXiv:2606.25256v1 Announce Type: new Abstract: We introduce Pre-Warm, a simple yet effective zero-training-cost method for data-conditioned initialization of the first convolutional layer. Before the

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

PRISM: Feed-Forward Single-Image 3D Reconstruction via Geometric Warp-Residual Modeling

DGX agent

arXiv:2606.25430v1 Announce Type: new Abstract: Reconstructing 3D scenes from a single image is a fundamental challenge in computer vision, with broad applications in virtual reality, robotics, and co

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need

DGX agent

arXiv:2606.25956v1 Announce Type: new Abstract: Risk stratification for pulmonary embolism (PE) is critical for clinical decision-making. Stratification guidelines are based on patient medical records

model-releasesarxiv-cs-cv
25 Jun 2026
Research

Re-mixing Embeddings for Patient Augmentation in Data Scarce Multiple Instance Learning

DGX agent

arXiv:2606.25770v1 Announce Type: cross Abstract: Data scarcity is a major bottleneck in medical Multiple Instance Learning (MIL), especially for rare diseases or expensive modalities. We introduce a

researcharxiv-cs-cv
25 Jun 2026
Applications

ReaDy-Go: Real-to-Sim Dynamic 3D Gaussian Splatting Simulation for Environment-Specific Visual Navigation with Moving Obstacles

DGX agent

arXiv:2602.11575v3 Announce Type: replace-cross Abstract: Visual navigation models often struggle in real-world dynamic environments due to limited robustness to the sim-to-real gap and the difficulty

applicationsarxiv-cs-cv
25 Jun 2026
Safety

Reflective VLA: In-Context Action Consequences Make VLAs Generalize

DGX agent

arXiv:2606.25215v1 Announce Type: new Abstract: Most vision-language-action (VLA) models are reactive: they predict the next action from the current instruction and observation, implicitly assuming th

safetyarxiv-cs-cv
25 Jun 2026
Research

REViT: Roto-reflection Equivariant Convolutional Vision Transformer

DGX agent

arXiv:2606.25318v1 Announce Type: new Abstract: In this paper, we propose a discrete roto-reflection group equivariant vision transformer with convolutional attention. Roto-reflection equivariant netw

researcharxiv-cs-cv
25 Jun 2026
Model Releases

RoboAtlas: Contextual Active SLAM

DGX agent

arXiv:2606.26046v1 Announce Type: cross Abstract: We present RoboAtlas, a contextual Active SLAM framework that adaptively balances geometric exploration and semantic reasoning using a scalable 3D sem

model-releasesarxiv-cs-cv
25 Jun 2026
Research

S^{2}-FracMix: Label-Preserving Self-Saliency Mixup Augmentation

DGX agent

arXiv:2606.25784v1 Announce Type: new Abstract: Data augmentation is known to improve generalization of deep visual models. Recent methods favor mixup strategies that generate interpolated samples to

researcharxiv-cs-cv
25 Jun 2026
Safety

SAC^2-Net: Semantic Anchoring and Complementary-Consensus Fusion for Multimodal Micro-Expression Recognition

DGX agent

arXiv:2606.25542v1 Announce Type: new Abstract: Micro-expression recognition (MER) is challenging due to subtle facial movements, limited data, and the ambiguous relationship between Action Units (AUs

safetyarxiv-cs-cv
25 Jun 2026
Applications

ScaleHP: Estimating Hand Pose in Metric Space

DGX agent

arXiv:2606.25619v1 Announce Type: new Abstract: Accurate metric-space hand pose estimation (HPE) is essential for immersive human-computer interaction and robotics. However, most existing methods pred

applicationsarxiv-cs-cv
25 Jun 2026
Safety

ScalingAR: Scaling Confidence for Autoregressive Image Generation

DGX agent

arXiv:2509.26376v3 Announce Type: replace Abstract: Test-time strategies have shown remarkable success in improving large language models, but their application to next-token prediction (NTP) autoregr

safetyarxiv-cs-cv
25 Jun 2026
Research

Semantic Allocation in Ordered Bottlenecks: Predictive Residual Inference for Visual Representation Learning

DGX agent

arXiv:2606.25232v1 Announce Type: cross Abstract: Ordered bottlenecks aim to provide utility at flexible budgets by assigning coarse information to early tokens and task-relevant detail to later ones.

researcharxiv-cs-cv
25 Jun 2026
Research

SEMIR: Topology-Preserving Graph Minors for Thin-Structure Segmentation

DGX agent

arXiv:2606.24935v1 Announce Type: new Abstract: Thin-structure segmentation--power lines, cracks, lane markings at 1-3 pixel width--requires preserving connectivity that standard representations precl

researcharxiv-cs-cv
25 Jun 2026
Research

Shift Variant Image Degradation and Restoration Using Singular Value Decomposition

DGX agent

arXiv:2606.25818v1 Announce Type: new Abstract: Shift-variant image degradation is frequently encountered in practical imaging systems where the point spread function (PSF) varies across the image fie

researcharxiv-cs-cv
25 Jun 2026
Model Releases

ShutterMuse: Capture-Time Photography Guidance with MLLMs

DGX agent

arXiv:2606.25763v1 Announce Type: new Abstract: Real-world photography requires capture-time guidance for both camera framing and subject pose. Yet existing aesthetic cropping benchmarks mainly evalua

model-releasesarxiv-cs-cv
25 Jun 2026
Research

SparseGS: Sparse View Synthesis using 3D Gaussian Splatting

DGX agent

arXiv:2312.00206v4 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has recently enabled real-time rendering of unbounded 3D scenes for novel view synthesis. However, this technique requi

researcharxiv-cs-cv
25 Jun 2026
Model Releases

Spatio-Temporal Mixture-of-Modality-Experts Diffusion for Quantitative DCE-MRI Synthesis from Incomplete MR Sequences

DGX agent

arXiv:2606.25535v1 Announce Type: new Abstract: Quantitative maps from dynamic contrast-enhanced MRI (DCE-MRI) are essential for tumor assessment but are often unavailable due to contrast-agent risks

model-releasesarxiv-cs-cv
25 Jun 2026
Research

SplatPainter: Interactive Authoring of 3D Gaussians from 2D Edits via Test-Time Training

DGX agent

arXiv:2512.05354v2 Announce Type: replace Abstract: The rise of 3D Gaussian Splatting has revolutionized photorealistic 3D asset creation, yet a critical gap remains for their interactive refinement a

researcharxiv-cs-cv
25 Jun 2026
Model Releases

SSMNBench: Diagnosing Image-based Cross-View Human-Object Understanding via Single-View Sufficiency and Multi-View Necessity

DGX agent

arXiv:2606.25634v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown remarkable progress in single-image perception, yet their ability to reason about complex cross-view

model-releasesarxiv-cs-cv
25 Jun 2026
Research

State Space Models Meet Remote Sensing: A Survey

DGX agent

arXiv:2606.25329v1 Announce Type: new Abstract: State Space Models (SSMs), designed for long-range modeling, offer linear computational complexity and strong capabilities in capturing long-range depen

researcharxiv-cs-cv
25 Jun 2026
Model Releases

Steering Vision-Language Models with Joint Sparse Autoencoders

DGX agent

arXiv:2606.25657v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) have shown promise for analyzing language models, but applying them to vision-language models (VLMs) often yields representat

model-releasesarxiv-cs-cv
25 Jun 2026
Research

Structuring Sparsity: Block-Sparse Featurizers Capture Visual Concept Manifolds

DGX agent

arXiv:2606.25234v1 Announce Type: new Abstract: What is the geometry of a visual percept? The most widely used protocols for decomposing neural network representations into interpretable parts treat c

researcharxiv-cs-cv
25 Jun 2026
Research

StyleFusion360: View-Consistent Head Stylization via Adaptive Style Modulation

DGX agent

arXiv:2511.22411v2 Announce Type: replace Abstract: 3D head stylization enables expressive reimagining of human faces for creative visual experiences in digital media. Existing 3D-aware methods often

researcharxiv-cs-cv
25 Jun 2026
Model Releases

SurgAtlas: A Large-Scale Surgical Video-Language Dataset with 2,391 Hours of Open and Minimally Invasive Surgery

DGX agent

arXiv:2606.25905v1 Announce Type: new Abstract: We introduce SurgAtlas, the largest surgical video-language dataset to date, comprising 15,291 videos (2,391 hours) spanning 18 surgical specialties and

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

TACO: Towards Task-Consistent Open-Vocabulary Adaptation in Video Recognition

DGX agent

arXiv:2606.25478v1 Announce Type: new Abstract: Adapting CLIP for open-vocabulary video recognition necessitates a delicate balance between newly acquired video knowledge and the pretrained generaliza

model-releasesarxiv-cs-cv
25 Jun 2026
Safety

Taxonomy-aware deep learning for hierarchical marine species classification in underwater imagery

DGX agent

arXiv:2606.25989v1 Announce Type: new Abstract: Automated classification of marine species from underwater imagery is essential for scalable ocean biodiversity monitoring and conservation policy. Exis

safetyarxiv-cs-cv
25 Jun 2026
← Previous
1…9091929394…263
Next →