AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlog
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Learning Efficient 4D Gaussian Representations from Monocular Videos with Flow Splatting

DGX agent

arXiv:2606.29976v1 Announce Type: new Abstract: Reconstructing dynamic 3D scenes from monocular videos is challenging due to scene complexity and temporal dynamics. With the advancement of 3D Gaussian

researcharxiv-cs-cv
30 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

Learning from Acquisition: Metadata-driven Multimodal Pre-training for Cardiac MRI

DGX agent

arXiv:2606.28991v1 Announce Type: new Abstract: Cardiac magnetic resonance imaging (CMR) routinely records structured acquisition metadata, yet most CMR foundation models rely primarily on image-only

researcharxiv-cs-cv
30 Jun 2026
Model Releases

Learning from Reliable Latent Prompts for Visual Recognition with Missing Modalities

DGX agent

arXiv:2606.30597v1 Announce Type: new Abstract: Large-scale multimodal models (LMMs) have achieved superior performance in visual recognition by synergizing information across diverse, massive-scale p

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Learning to Segment Liquids in Real-world Images

DGX agent

arXiv:2601.00940v2 Announce Type: replace Abstract: Liquids like water, wine and medicine are everywhere. However, limited attention has been given to the task of segmenting liquids, hindering the abi

model-releasesarxiv-cs-cv
30 Jun 2026
Local Ai

Learning Where and When: Patch-Based Spatiotemporal Localization in Weakly Supervised Video Anomaly Detection

DGX agent

arXiv:2606.29498v1 Announce Type: new Abstract: Weakly supervised video anomaly detection (WSVAD) has predominantly focused on temporal localization, identifying when anomalies occur while largely neg

local-aiarxiv-cs-cv
30 Jun 2026
Model Releases

LEIQ-Assessor: Multi-dimensional Quality Assessment of Low-light Enhanced Images via Multi-task Learning

DGX agent

arXiv:2606.29752v1 Announce Type: new Abstract: Low-light image enhancement algorithms (LIEAs) aim to improve the visibility of images captured under poor illumination. However, the enhancement proces

model-releasesarxiv-cs-cv
30 Jun 2026
Research

LETT-NeXt: A Lightweight RECIST-Guided Model for 3D CT Lesion Segmentation

DGX agent

arXiv:2606.30108v1 Announce Type: new Abstract: RECIST diameter measurements are widely used for tumor response assessment, but they provide only a limited 2D description of lesion extent. We present

researcharxiv-cs-cv
30 Jun 2026
Research

LogiCo: A Unified Framework for Logical and Structural Anomaly Detection

DGX agent

arXiv:2606.28688v1 Announce Type: new Abstract: Current anomaly detection methods primarily focus on structural anomalies, while paying insufficient attention to anomalies that violate logical constra

researcharxiv-cs-cv
30 Jun 2026
Model Releases

LoGSAM: Parameter-Efficient Cross-Modal Grounding for MRI Segmentation

DGX agent

arXiv:2603.17576v3 Announce Type: replace Abstract: Precise localization and delineation of brain tumors using magnetic resonance imaging (MRI) are essential for planning therapy and guiding surgical

model-releasesarxiv-cs-cv
30 Jun 2026
Local Ai

MatchAttention: Embedding Explicit Matching Constraints into Attention for Efficient Stereo Matching

DGX agent

arXiv:2510.14260v3 Announce Type: replace Abstract: Standard attention mechanisms are not well suited to stereo matching. Global attention scales quadratically and provides no explicit matching constr

local-aiarxiv-cs-cv
30 Jun 2026
Agents

MAVIN: Multi-Shot Audio-Visual Generation with Narrative Control

DGX agent

arXiv:2606.29473v1 Announce Type: new Abstract: While recent generative models produce high-fidelity videos, they struggle with the complex narrative control required for coherent multi-shot audio-vis

agentsarxiv-cs-cv
30 Jun 2026
Research

Measured-Subspace Consistency: A Plug-and-Play Operator for Diffusion Posterior Sampling in Accelerated MRI Reconstruction

DGX agent

arXiv:2606.28448v1 Announce Type: cross Abstract: Diffusion posterior samplers for accelerated MRI can reconstruct accurately yet still disagree on the acquired k-space across samples, placing posteri

researcharxiv-cs-cv
30 Jun 2026
Applications

Meshtryoshka: Differentiable Rendering of Real-World Scenes via Mesh Rasterization

DGX agent

arXiv:2606.28622v1 Announce Type: new Abstract: Differentiable rendering has emerged as a powerful approach for 3D reconstruction and novel view synthesis. State-of-the-art differentiable rendering me

applicationsarxiv-cs-cv
30 Jun 2026
Safety

Meta-learning as a principle for human-like visual representations

DGX agent

arXiv:2606.28399v1 Announce Type: new Abstract: The structure of human visual representations underpins our capacity for adaptive behaviour. While pretrained neural networks model human visual represe

safetyarxiv-cs-cv
30 Jun 2026
Local Ai

MF-UAVPose6D: A Model-Free Monocular 6-DoF Pose Estimation Framework for Fixed-Wing UAVs

DGX agent

arXiv:2606.29697v1 Announce Type: new Abstract: For uncrewed aerial vehicles (UAVs), estimating six-degree-of-freedom (6-DoF) poses is essential for airspace situational awareness, target tracking, an

local-aiarxiv-cs-cv
30 Jun 2026
Safety

MIRROR: Aligning Semantic Relations from Language to Image via Gromov--Wasserstein

DGX agent

arXiv:2606.29462v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) inherit rich relational priors from their language backbones, yet often fail when asked to apply these relation

safetyarxiv-cs-cv
30 Jun 2026
Research

MirrorPPR: Exemplar-Based Portrait Photo Retouching

DGX agent

arXiv:2606.29308v1 Announce Type: new Abstract: While text-guided image editing has made remarkable progress, it remains limited in structural portrait retouching. Textual descriptions struggle to con

researcharxiv-cs-cv
30 Jun 2026
Research

Miti360: A Comprehensive Dataset for Improved Reforestation Monitoring

DGX agent

arXiv:2606.29447v1 Announce Type: new Abstract: Over the past decade, interest in applying machine learning (ML) to automate forest monitoring has grown significantly. However, existing training datas

researcharxiv-cs-cv
30 Jun 2026
Safety

MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning

DGX agent

arXiv:2510.03142v2 Announce Type: replace-cross Abstract: Visual navigation policy is widely regarded as a promising direction, as it mimics humans by using egocentric visual observations for navigati

safetyarxiv-cs-cv
30 Jun 2026
Research

Momentum Guidance: Plug-and-Play Guidance for Flow Models

DGX agent

arXiv:2602.20360v2 Announce Type: replace-cross Abstract: Flow-based generative methods offer a simple and effective framework for high-fidelity generation, yet pretrained flow models are rarely used

researcharxiv-cs-cv
30 Jun 2026
Agents

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images

DGX agent

arXiv:2511.19119v2 Announce Type: replace Abstract: Spatial reasoning (SR), the ability to infer 3D spatial information from 2D inputs, is essential for real-world applications such as embodied AI and

agentsarxiv-cs-cv
30 Jun 2026
Model Releases

Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting

DGX agent

arXiv:2606.30017v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting have demonstrated unprecedented success in novel view synthesis. However, the substantial inference and storage

model-releasesarxiv-cs-cv
30 Jun 2026
Safety

MR-IQA: A Unified Margin View of Regression and Ranking for Blind Image Quality Assessment

DGX agent

arXiv:2606.29760v1 Announce Type: new Abstract: Blind image quality assessment (BIQA) is commonly built on two basic learning paradigms: regression and ranking. Regression calibrates absolute scores,

safetyarxiv-cs-cv
30 Jun 2026
Model Releases

MSA-UNet3+: Multi-Scale Attention UNet3+ with New Supervised Prototypical Contrastive Loss for Coronary DSA Image Segmentation

DGX agent

arXiv:2504.05184v4 Announce Type: replace-cross Abstract: Accurate segmentation of coronary Digital Subtraction Angiography (DSA) images is essential for diagnosing and treating coronary artery diseas

model-releasesarxiv-cs-cv
30 Jun 2026
Safety

muFlow: Leveraging Average Images for Improving Generalisation of Deepfake Faces Detectors

DGX agent

arXiv:2606.30528v1 Announce Type: new Abstract: Current generative models, including GANs and diffusion models, have reached an outstanding level of photorealism, posing significant risks to privacy a

safetyarxiv-cs-cv
30 Jun 2026
Model Releases

Multi-scale Object-Aware Gaze Estimation via Geometric Reasoning

DGX agent

arXiv:2606.29334v1 Announce Type: new Abstract: Gaze target estimation aims to predict the semantic object an observer fixates upon within an image, a task deeply rooted in the object-oriented nature

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Multimodal Graph RAG for Long-range Visually Rich Document Understanding

DGX agent

arXiv:2606.28780v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are widely applied to visual document understanding. However, comprehending long documents remains an issue b

model-releasesarxiv-cs-cv
30 Jun 2026
Safety

Multimodal Large Language Model driven Radiology Report Generation with Clinical Knowledge Enhancement

DGX agent

arXiv:2403.06728v2 Announce Type: replace Abstract: Radiology report generation (RRG) has attracted significant attention due to its potential to reduce the workload of radiologists. The performance o

safetyarxiv-cs-cv
30 Jun 2026
Tutorials

Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers

DGX agent

arXiv:2606.29013v1 Announce Type: new Abstract: Leveraging capabilities of large language models (LLMs) in text-to-image (T2I) synthesis is an important research direction. In this work we investigate

tutorialsarxiv-cs-cv
30 Jun 2026
Model Releases

MUSE: Unlocking Timestep as Native Task Steering for One-Step Dense Prediction

DGX agent

arXiv:2606.30370v1 Announce Type: new Abstract: Monocular dense prediction has recently seen remarkable success by repurposing pre-trained diffusion models. This opens a promising yet challenging aven

model-releasesarxiv-cs-cv
30 Jun 2026
Agents

NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation

DGX agent

arXiv:2606.29395v1 Announce Type: new Abstract: Recently, Large Language Models (LLMs) have emerged as promising layout agents for 3D scene generation. Existing layout agents still suffer from implaus

agentsarxiv-cs-cv
30 Jun 2026
Model Releases

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis

DGX agent

arXiv:2606.29814v1 Announce Type: new Abstract: We propose Nemotron-Labs-Diffusion-Image, a state-of-the-art masked discrete diffusion model (MDM) for high-resolution text-to-image synthesis. Compared

model-releasesarxiv-cs-cv
30 Jun 2026
Local Ai

Neural Gate: Mitigating Privacy Risks in LVLMs via Neuron-Level Gradient Gating

DGX agent

arXiv:2603.12598v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) have shown remarkable potential across a wide array of vision-language tasks, leading to their adoption in crit

local-aiarxiv-cs-cv
30 Jun 2026
Agents

Neural Stereo Video Compression with Hybrid Disparity Compensation

DGX agent

arXiv:2504.20383v3 Announce Type: replace Abstract: Disparity compensation represents the primary strategy in stereo video compression (SVC) for exploiting cross-view redundancy. These mechanisms can

agentsarxiv-cs-cv
30 Jun 2026
Model Releases

Nonlinear mixture model motivated subspace clustering

DGX agent

arXiv:2606.29261v1 Announce Type: cross Abstract: We derive the linear union-of-subspaces (UoS) model for subspace clustering (SC) from the nonlinear mixture model (NMM) used in blind source separatio

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Obliviate: Erasing Concepts from Autoregressive Image Generation Models

DGX agent

arXiv:2606.28643v1 Announce Type: new Abstract: The widespread adoption of generative AI models has intensified concerns about misuse, including the creation of unsafe or disturbing imagery. To mitiga

model-releasesarxiv-cs-cv
30 Jun 2026
Applications

Occlusion-Robust Multi-Object Decoupling for Physics-Based Interaction

DGX agent

arXiv:2606.29303v1 Announce Type: new Abstract: We propose a mask-free method for lossless multi-object 3D reconstruction from sparse and occluded real-world views, enabling physically plausible inter

applicationsarxiv-cs-cv
30 Jun 2026
Model Releases

OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning

DGX agent

arXiv:2606.30378v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated promising spatial reasoning capabilities, while these abilities remain underexplored in the e

model-releasesarxiv-cs-cv
30 Jun 2026
Research

OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data

DGX agent

arXiv:2606.30019v1 Announce Type: new Abstract: Music-driven dance video generation aims to synthesize expressive human motion that is temporally aligned with music while maintaining high visual fidel

researcharxiv-cs-cv
30 Jun 2026
Research

On Test-Time Scaling for Vision-Language Models

DGX agent

arXiv:2606.28864v1 Announce Type: new Abstract: Test-time scaling is a paradigm where large models use additional compute at inference to achieve better performance, without changing model weights. Wh

researcharxiv-cs-cv
30 Jun 2026
Model Releases

On the Vulnerability of Parameter-Level Defenses to Model Merging

DGX agent

arXiv:2606.30360v1 Announce Type: cross Abstract: The training-free integration of expert models via model merging has exposed significant security risks, enabling free-riders to combine specialized m

model-releasesarxiv-cs-cv
30 Jun 2026
Local Ai

One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding

DGX agent

arXiv:2606.30084v1 Announce Type: new Abstract: MLLM-based GUI grounding methods commonly formulate target localization as autoregressive coordinate generation, enabling models to leverage the strong

local-aiarxiv-cs-cv
30 Jun 2026
Model Releases

OP3DSG: Open-Vocabulary Part-Aware 3D Scene Graph Generation for Real-World Environments

DGX agent

arXiv:2606.29786v1 Announce Type: new Abstract: 3D scene graphs (3DSGs) provide a compact and structured abstraction of 3D environments. Although advances in foundation models have enabled open-vocabu

model-releasesarxiv-cs-cv
30 Jun 2026
Tutorials

Open-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D Detectors

DGX agent

arXiv:2606.30638v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has emerged at the forefront of 3D scene reconstruction. Extending 3DGS with language-driven, open-vocabulary understanding

tutorialsarxiv-cs-cv
30 Jun 2026
Research

Optimizing Image Preparation and Compression for Face Recognition within 1024 Bytes

DGX agent

arXiv:2606.30321v1 Announce Type: new Abstract: ICAO-compliant machine readable travel documents enable automated biometric face verification. The biometric reference is stored on an RFID chip include

researcharxiv-cs-cv
30 Jun 2026
Research

Orca: The World is in Your Mind

DGX agent

arXiv:2606.30534v1 Announce Type: new Abstract: We introduce Orca, an initial instantiation of a general world foundation model. Orca learns a unified world latent space from multimodal world signals

researcharxiv-cs-cv
30 Jun 2026
Safety

OWMDrive: Causality-Aware End-to-End Autonomous Driving via 4D Occupancy World Model

DGX agent

arXiv:2606.30421v1 Announce Type: new Abstract: Autonomous driving systems are steadily moving toward end-to-end paradigms to mitigate the limited adaptability of rule-based pipelines in complex traff

safetyarxiv-cs-cv
30 Jun 2026
Research

PCP-GAN: Property-Constrained Pore-scale image reconstruction via conditional Generative Adversarial Networks

DGX agent

arXiv:2510.19465v2 Announce Type: replace Abstract: Obtaining truly representative pore-scale images that match bulk formation properties remains a fundamental challenge in subsurface characterization

researcharxiv-cs-cv
30 Jun 2026
← Previous
1…8081828384…263
Next →