AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlog
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Agents

Semantic Occupancy Prediction with Dual Range-Voxel Representation

DGX agent

arXiv:2606.31688v1 Announce Type: new Abstract: LiDAR-based 3D semantic occupancy prediction, which aims to provide accurate and comprehensive scene representation, is crucial for autonomous driving s

agentsarxiv-cs-cv
1 Jul 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

SENSE-VAD: Sentient and Semantic Video Anomaly Detection for Autonomous Driving

DGX agent

arXiv:2606.31875v1 Announce Type: new Abstract: Autonomous vehicles (AVs) must navigate not only motion-based hazards but also socially complex situations whose danger is constituted by inter-agent re

model-releasesarxiv-cs-cv
1 Jul 2026
Research

ShellMaker: Language-Guided Exterior Completion under Structural Constraints

DGX agent

arXiv:2606.31680v1 Announce Type: new Abstract: Despite advances in indoor scene generation, synthesizing coherent building exteriors consistent with generated interiors remains largely unexplored. Ex

researcharxiv-cs-cv
1 Jul 2026
Research

SHMoAReg: Spark Deformable Image Registration via Spatial Heterogeneous Mixture of Experts and Attention Heads

DGX agent

arXiv:2509.20073v2 Announce Type: replace Abstract: Encoder-Decoder architectures are widely used in deep learning-based Deformable Image Registration (DIR), where the encoder extracts multi-scale fea

researcharxiv-cs-cv
1 Jul 2026
Research

Simple Supervision Is Hard to Beat: A Bitter Lesson from Sparse Target Labels in Domain-Adaptive Object Detection

DGX agent

arXiv:2606.30795v1 Announce Type: new Abstract: Source-free domain adaptive object detection adapts a source-trained detector to an unlabeled target domain, typically through teacher-student self-trai

researcharxiv-cs-cv
1 Jul 2026
Model Releases

SimpleSearch-VL: A Simple Recipe for Multimodal Agentic Deep Search

DGX agent

arXiv:2606.31504v1 Announce Type: new Abstract: We present SimpleSearch-VL, an efficient, reliable, and practical framework for multimodal agentic search. Its core idea is to improve the agent's own s

model-releasesarxiv-cs-cv
1 Jul 2026
Safety

SpectralSplats: Robust Differentiable Tracking via Spectral Moment Supervision

DGX agent

arXiv:2603.24036v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) enables real-time, photorealistic novel view synthesis, making it a highly attractive representation for model-based vi

safetyarxiv-cs-cv
1 Jul 2026
Research

SPFSplatV2: Efficient Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views

DGX agent

arXiv:2509.17246v2 Announce Type: replace Abstract: We introduce SPFSplatV2, an efficient feed-forward framework for 3D Gaussian splatting from sparse multi-view images, requiring no ground-truth pose

researcharxiv-cs-cv
1 Jul 2026
Research

SpheRoPE: Zero-Shot Optimization-Free 360 Panorama Generation with Spherical RoPE

DGX agent

arXiv:2606.32033v1 Announce Type: new Abstract: We present a zero-shot, training-free and optimization-free framework for generating 360 panoramic images and videos by directly injecting spherical pri

researcharxiv-cs-cv
1 Jul 2026
Safety

Stealthy Multi-Task Adversarial Attacks

DGX agent

arXiv:2411.17936v2 Announce Type: replace-cross Abstract: Deep neural networks are highly vulnerable to adversarial perturbations, raising serious safety concerns in the real-world systems. While prio

safetyarxiv-cs-cv
1 Jul 2026
Research

StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation

DGX agent

arXiv:2602.23721v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models integrate visual observations and language instructions to predict robot actions, demonstrating promising

researcharxiv-cs-cv
1 Jul 2026
Tutorials

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance

DGX agent

arXiv:2506.20995v4 Announce Type: replace Abstract: We propose a step-by-step video-to-audio (V2A) generation method that provides finer control over the generation process and more realistic audio sy

tutorialsarxiv-cs-cv
1 Jul 2026
Research

Streaming Gaussian Encoding for 4D Panoptic Occupancy Tracking

DGX agent

arXiv:2606.30754v1 Announce Type: new Abstract: Camera-based 4D panoptic occupancy tracking (4D-POT) is a promising paradigm for holistic scene understanding from multi-view imagery, enabling joint re

researcharxiv-cs-cv
1 Jul 2026
Research

Structured SIR: Efficient and Expressive Importance-Weighted Inference for High-Dimensional Image Registration

DGX agent

arXiv:2603.17415v2 Announce Type: replace-cross Abstract: Image registration is an ill-posed dense vision task, where multiple solutions achieve similar loss values, motivating probabilistic inference

researcharxiv-cs-cv
1 Jul 2026
Safety

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation

DGX agent

arXiv:2606.30849v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have significantly advanced audio-driven portrait animation, but their high computational cost leads to substantial infere

safetyarxiv-cs-cv
1 Jul 2026
Research

T-QPM: Enabling Temporal Out-Of-Distribution Detection and Domain Generalization for Vision-Language Models in Open-World

DGX agent

arXiv:2603.18481v2 Announce Type: replace Abstract: Out-of-distribution (OOD) detection remains a critical challenge in open-world learning, where models must adapt to evolving data distributions. Whi

researcharxiv-cs-cv
1 Jul 2026
Model Releases

TaxoMIL: Taxonomy-Constrained Learning for Hierarchical Whole Slide Image Analysis

DGX agent

arXiv:2606.31100v1 Announce Type: new Abstract: Whole slide image (WSI) analysis is central to computational pathology, with multiple instance learning (MIL) emerging as the standard pipeline for slid

model-releasesarxiv-cs-cv
1 Jul 2026
Research

Technical Report of RoboSpatial Challenge at CVPR 2026: Selective Reasoning Activation and Reference-Frame Disambiguation for Embodied Spatial Reasoning

DGX agent

arXiv:2606.31645v1 Announce Type: new Abstract: Vision-language models achieve strong general perception but often struggle with the spatial reasoning required for embodied tasks. We present RoboSpati

researcharxiv-cs-cv
1 Jul 2026
Research

Temporal Training Strategies for Left Atrium and Left Atrial Appendage Segmentation in Dynamic Contrast 4DCT

DGX agent

arXiv:2606.31444v1 Announce Type: new Abstract: Dynamic contrast-enhanced cardiac CT enables time-resolved analysis of contrast filling and washout in the left atrium (LA) and left atrial appendage (L

researcharxiv-cs-cv
1 Jul 2026
Local Ai

TerraDiT-Omega: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial Primitive

DGX agent

arXiv:2606.31029v1 Announce Type: new Abstract: Generative models have achieved remarkable progress, yet applying them to satellite imagery remains challenging. Unlike natural imagery, satellite scene

local-aiarxiv-cs-cv
1 Jul 2026
Research

Think While You Map: Asynchronous Vision-Language Agents for Incremental 3D Scene Graphs

DGX agent

arXiv:2606.31471v1 Announce Type: new Abstract: Open-vocabulary 3D scene graph methods typically operate in two stages: first reconstruct, then enrich with vision-language models, leaving the graph un

researcharxiv-cs-cv
1 Jul 2026
Safety

TORA: Topological Representation Alignment for 3D Shape Assembly

DGX agent

arXiv:2604.04050v2 Announce Type: replace Abstract: Flow-matching methods for 3D shape assembly learn point-wise velocity fields that transport parts toward assembled configurations, yet they receive

safetyarxiv-cs-cv
1 Jul 2026
Research

TotalFM: An Organ-Separated 3D-CT Foundation Model Leveraging Large-Scale Routine Clinical Radiology Data

DGX agent

arXiv:2601.00260v2 Announce Type: replace Abstract: While foundation models in radiology are expected to be applied to various clinical tasks, computational cost constraints remain a major challenge w

researcharxiv-cs-cv
1 Jul 2026
Research

Towards a foundational model for recognising diastematic Gregorian notation

DGX agent

arXiv:2606.31454v1 Announce Type: new Abstract: Optical recognition of Gregorian notation has recently been attempted with end-to-end methods, with four datasets introduced. However, each of these dat

researcharxiv-cs-cv
1 Jul 2026
Research

Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation

DGX agent

arXiv:2606.31088v1 Announce Type: new Abstract: Conversational talking face generation has recently attracted increasing attention, aiming to synthesize interactive talking videos where characters spe

researcharxiv-cs-cv
1 Jul 2026
Research

Towards Voxel Spacing Consistency for Medical Image Segmentation

DGX agent

arXiv:2606.31839v1 Announce Type: new Abstract: Volumetric medical image segmentation is essential for both preoperative diagnosis and intraoperative guidance. While recent years have witnessed rapid

researcharxiv-cs-cv
1 Jul 2026
Research

UHD-MFF: Shattering Barriers in Multi-Focus Ultra-High-Definition Image Fusion via Learnable Lookup Tables

DGX agent

arXiv:2606.31242v1 Announce Type: new Abstract: With the advancement of imaging technology, ultra-high-definition images have become increasingly essential in modern visual applications. However, exis

researcharxiv-cs-cv
1 Jul 2026
Model Releases

UniCoder: Unified Visual-to-Code Generation via Symbolic Rewards and Reference-Guided Code Optimization

DGX agent

arXiv:2606.31732v1 Announce Type: new Abstract: Visual-to-Code generation, which transforms scientific plots, vector graphics, and webpages into executable scripts, demands a level of pixel-precise al

model-releasesarxiv-cs-cv
1 Jul 2026
Safety

Unsupervised Data-Efficient Cross-Modal Retrieval with Global-Neighborhood Alignment Hashing

DGX agent

arXiv:2606.31517v1 Announce Type: cross Abstract: Compared to supervised cross-modal hashing (CMH), unsupervised CMH reduces the reliance on manual labeling by learning binary codes from unlabeled ima

safetyarxiv-cs-cv
1 Jul 2026
Agents

Unveiling Transferability in Trajectory Prediction via Latent Scene Embeddings

DGX agent

arXiv:2606.30777v1 Announce Type: new Abstract: The growing availability of trajectory datasets has fueled major advances in data-driven motion prediction. Yet, models trained on one dataset often fai

agentsarxiv-cs-cv
1 Jul 2026
Safety

VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment

DGX agent

arXiv:2603.16271v3 Announce Type: replace Abstract: Video diffusion models lack explicit geometric supervision during training, leading to inconsistency artifacts such as object deformation, spatial d

safetyarxiv-cs-cv
1 Jul 2026
Research

VS3R: Robust Full-frame Video Stabilization via Deep 3D Reconstruction

DGX agent

arXiv:2603.05851v2 Announce Type: replace Abstract: Video stabilization aims to mitigate camera shake but faces a fundamental trade-off between geometric robustness and full-frame consistency. While 2

researcharxiv-cs-cv
1 Jul 2026
Research

WAFT-Stereo: Warping-Alone Field Transforms for Stereo Matching

DGX agent

arXiv:2603.24836v3 Announce Type: replace Abstract: We introduce WAFT-Stereo, a simple and effective warping-based method for stereo matching. WAFT-Stereo demonstrates that cost volumes, a common desi

researcharxiv-cs-cv
1 Jul 2026
Research

WarpHammer: Densifying Scene Warps with 3D Object Priors for Extreme View Synthesis

DGX agent

arXiv:2606.31258v1 Announce Type: new Abstract: Projection-conditioned novel view synthesis (NVS) warps an explicit 3D reconstruction of the input view into the target camera and conditions a generato

researcharxiv-cs-cv
1 Jul 2026
Research

WarpI2I: Image Warping for Image-to-Image Translation

DGX agent

arXiv:2606.31018v1 Announce Type: new Abstract: Image-to-image (I2I) translation has achieved strong results in tasks like human relighting and driving scene translation using latent diffusion models

researcharxiv-cs-cv
1 Jul 2026
Research

WaterGen: Decoupling Scene and Medium in Underwater Image Generation

DGX agent

arXiv:2606.31147v1 Announce Type: new Abstract: Underwater computer vision tasks, such as detection, restoration, and segmentation, are limited by the scarcity of large-scale and diverse training data

researcharxiv-cs-cv
1 Jul 2026
Applications

Wavelet-Optimized Pseudo-3D Accelerated Diffusion Model for Truncated Computed Laminography

DGX agent

arXiv:2606.31318v1 Announce Type: new Abstract: Computed Laminography (CL) is a key technology for the nondestructive testing of large plate-shaped objects. However, field-of-view (FOV) limitations in

applicationsarxiv-cs-cv
1 Jul 2026
Model Releases

What Memory Do GUI Agents Really Need? From Passive Records to Active Task-Driving States

DGX agent

arXiv:2606.31612v1 Announce Type: new Abstract: Mobile GUI agents increasingly face long-horizon tasks that require reading, updating, and reusing task-relevant data across pages and applications. Exi

model-releasesarxiv-cs-cv
1 Jul 2026
Research

When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models

DGX agent

arXiv:2604.03316v2 Announce Type: replace Abstract: Attention sinks are defined as tokens that attract disproportionate attention. While these have been studied in single modality transformers, their

researcharxiv-cs-cv
1 Jul 2026
Model Releases

WIDER-FAIR: An Annotated Version of the WIDER-FACE Dataset for Fairness Evaluation

DGX agent

arXiv:2606.31704v1 Announce Type: new Abstract: The deployment of face detection models in real-world applications raises important fairness concerns, as these systems may showcase performance dispari

model-releasesarxiv-cs-cv
1 Jul 2026
Research

WildProp: Visual Estimation of Wildlife Body Proportions at Scale

DGX agent

arXiv:2606.31125v1 Announce Type: new Abstract: Population-level morphometric measurements underpin ecological and evolutionary studies but traditionally require controlled imaging or physical specime

researcharxiv-cs-cv
1 Jul 2026
Tutorials

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration

DGX agent

arXiv:2606.31946v1 Announce Type: new Abstract: The fundamental obstacle to industrial grade video generation is the lack of controllability: existing models treat video as a pixel distribution sampli

tutorialsarxiv-cs-cv
1 Jul 2026
Research

3D Field of Junctions: A Noise-Robust, Training-Free Structural Prior for Volumetric Inverse Problems

DGX agent

arXiv:2603.02149v2 Announce Type: replace Abstract: Volume denoising is a foundational problem in computational imaging, as many 3D imaging inverse problems face high levels of measurement noise. Insp

researcharxiv-cs-cv
30 Jun 2026
Model Releases

3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

DGX agent

arXiv:2606.30514v1 Announce Type: new Abstract: Human image animation, which aims to generate a video of a reference subject following a provided action sequence, has received increasing research inte

model-releasesarxiv-cs-cv
30 Jun 2026
Research

A Classifier-Agnostic Zero-Shot Adversarial Attack Detection via CLIP

DGX agent

arXiv:2606.30342v1 Announce Type: new Abstract: Adversarial attacks pose a challenge to the reliability of deep learning models, motivating effective detection methods. Existing techniques often rely

researcharxiv-cs-cv
30 Jun 2026
Research

A Dual-domain Refinement Network with FBP-based Jacobian Learning for Sparse-view Dual-Energy CT Material Decomposition

DGX agent

arXiv:2606.30159v1 Announce Type: new Abstract: Dual-energy CT (DECT) exploits attenuation differences across different X-ray spectra to provide richer material information and has been widely used in

researcharxiv-cs-cv
30 Jun 2026
Research

A Morse-Bott Framework for Blind Inverse Problems: Local Recovery Guarantees and the Failure of the MAP

DGX agent

arXiv:2508.02923v3 Announce Type: replace Abstract: Maximum A Posteriori (MAP) estimation is a cornerstone framework for blind inverse problems, where an image and a forward operator are jointly estim

researcharxiv-cs-cv
30 Jun 2026
Model Releases

A multi-architecture study of specificity refinement and false-positive mechanism analysis in prostate MRI

DGX agent

arXiv:2606.29977v1 Announce Type: cross Abstract: Objectives: To characterize residual false positives in prostate MRI detection, and to evaluate a lightweight post-hoc refinement head for case-level

model-releasesarxiv-cs-cv
30 Jun 2026
← Previous
1…7677787980…263
Next →