AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
30 Jun 2026

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models

ResearchDGX agent

arXiv:2606.30168v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) often fail in fine-grained visual reasoning, as question-relevant visual cues are diluted by dense and redundan

LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion

Model ReleasesDGX agent

arXiv:2603.14526v2 Announce Type: replace Abstract: The recent success of inference-time scaling in large language models has inspired similar explorations in video diffusion. In particular, motivated

LaVPR: Benchmarking Language and Vision for Place Recognition

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2602.03253v2 Announce Type: replace Abstract: Visual Place Recognition (VPR) often fails under extreme environmental changes and perceptual aliasing. Beyond these limitations, standard systems c

LeAD-M3D: Leveraging Asymmetric Distillation for Real-Time Monocular 3D Detection

ResearchDGX agent

arXiv:2512.05663v3 Announce Type: replace Abstract: Real-time monocular 3D object detection remains challenging due to severe depth ambiguity, viewpoint shifts, and the high computational cost of 3D r

Learning Cross-view Correspondences for Geo-localization on Planetary Surfaces

Model ReleasesDGX agent

arXiv:2606.29821v1 Announce Type: new Abstract: Maintaining global position awareness is a fundamental challenge for planetary surface exploration, since satellite-based positioning systems are unavai

Learning Efficient 4D Gaussian Representations from Monocular Videos with Flow Splatting

ResearchDGX agent

arXiv:2606.29976v1 Announce Type: new Abstract: Reconstructing dynamic 3D scenes from monocular videos is challenging due to scene complexity and temporal dynamics. With the advancement of 3D Gaussian

Learning from Acquisition: Metadata-driven Multimodal Pre-training for Cardiac MRI

ResearchDGX agent

arXiv:2606.28991v1 Announce Type: new Abstract: Cardiac magnetic resonance imaging (CMR) routinely records structured acquisition metadata, yet most CMR foundation models rely primarily on image-only

Learning from Reliable Latent Prompts for Visual Recognition with Missing Modalities

Model ReleasesDGX agent

arXiv:2606.30597v1 Announce Type: new Abstract: Large-scale multimodal models (LMMs) have achieved superior performance in visual recognition by synergizing information across diverse, massive-scale p

Learning to Segment Liquids in Real-world Images

Model ReleasesDGX agent

arXiv:2601.00940v2 Announce Type: replace Abstract: Liquids like water, wine and medicine are everywhere. However, limited attention has been given to the task of segmenting liquids, hindering the abi

Learning Where and When: Patch-Based Spatiotemporal Localization in Weakly Supervised Video Anomaly Detection

Local AiDGX agent

arXiv:2606.29498v1 Announce Type: new Abstract: Weakly supervised video anomaly detection (WSVAD) has predominantly focused on temporal localization, identifying when anomalies occur while largely neg

LEIQ-Assessor: Multi-dimensional Quality Assessment of Low-light Enhanced Images via Multi-task Learning

Model ReleasesDGX agent

arXiv:2606.29752v1 Announce Type: new Abstract: Low-light image enhancement algorithms (LIEAs) aim to improve the visibility of images captured under poor illumination. However, the enhancement proces

LETT-NeXt: A Lightweight RECIST-Guided Model for 3D CT Lesion Segmentation

ResearchDGX agent

arXiv:2606.30108v1 Announce Type: new Abstract: RECIST diameter measurements are widely used for tumor response assessment, but they provide only a limited 2D description of lesion extent. We present

LogiCo: A Unified Framework for Logical and Structural Anomaly Detection

ResearchDGX agent

arXiv:2606.28688v1 Announce Type: new Abstract: Current anomaly detection methods primarily focus on structural anomalies, while paying insufficient attention to anomalies that violate logical constra

LoGSAM: Parameter-Efficient Cross-Modal Grounding for MRI Segmentation

Model ReleasesDGX agent

arXiv:2603.17576v3 Announce Type: replace Abstract: Precise localization and delineation of brain tumors using magnetic resonance imaging (MRI) are essential for planning therapy and guiding surgical

MatchAttention: Embedding Explicit Matching Constraints into Attention for Efficient Stereo Matching

Local AiDGX agent

arXiv:2510.14260v3 Announce Type: replace Abstract: Standard attention mechanisms are not well suited to stereo matching. Global attention scales quadratically and provides no explicit matching constr

MAVIN: Multi-Shot Audio-Visual Generation with Narrative Control

AgentsDGX agent

arXiv:2606.29473v1 Announce Type: new Abstract: While recent generative models produce high-fidelity videos, they struggle with the complex narrative control required for coherent multi-shot audio-vis

Measured-Subspace Consistency: A Plug-and-Play Operator for Diffusion Posterior Sampling in Accelerated MRI Reconstruction

ResearchDGX agent

arXiv:2606.28448v1 Announce Type: cross Abstract: Diffusion posterior samplers for accelerated MRI can reconstruct accurately yet still disagree on the acquired k-space across samples, placing posteri

Meshtryoshka: Differentiable Rendering of Real-World Scenes via Mesh Rasterization

ApplicationsDGX agent

arXiv:2606.28622v1 Announce Type: new Abstract: Differentiable rendering has emerged as a powerful approach for 3D reconstruction and novel view synthesis. State-of-the-art differentiable rendering me

Meta-learning as a principle for human-like visual representations

SafetyDGX agent

arXiv:2606.28399v1 Announce Type: new Abstract: The structure of human visual representations underpins our capacity for adaptive behaviour. While pretrained neural networks model human visual represe

MF-UAVPose6D: A Model-Free Monocular 6-DoF Pose Estimation Framework for Fixed-Wing UAVs

Local AiDGX agent

arXiv:2606.29697v1 Announce Type: new Abstract: For uncrewed aerial vehicles (UAVs), estimating six-degree-of-freedom (6-DoF) poses is essential for airspace situational awareness, target tracking, an

MIRROR: Aligning Semantic Relations from Language to Image via Gromov--Wasserstein

SafetyDGX agent

arXiv:2606.29462v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) inherit rich relational priors from their language backbones, yet often fail when asked to apply these relation

MirrorPPR: Exemplar-Based Portrait Photo Retouching

ResearchDGX agent

arXiv:2606.29308v1 Announce Type: new Abstract: While text-guided image editing has made remarkable progress, it remains limited in structural portrait retouching. Textual descriptions struggle to con

Miti360: A Comprehensive Dataset for Improved Reforestation Monitoring

ResearchDGX agent

arXiv:2606.29447v1 Announce Type: new Abstract: Over the past decade, interest in applying machine learning (ML) to automate forest monitoring has grown significantly. However, existing training datas

MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning

SafetyDGX agent

arXiv:2510.03142v2 Announce Type: replace-cross Abstract: Visual navigation policy is widely regarded as a promising direction, as it mimics humans by using egocentric visual observations for navigati

Momentum Guidance: Plug-and-Play Guidance for Flow Models

ResearchDGX agent

arXiv:2602.20360v2 Announce Type: replace-cross Abstract: Flow-based generative methods offer a simple and effective framework for high-fidelity generation, yet pretrained flow models are rarely used

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images

AgentsDGX agent

arXiv:2511.19119v2 Announce Type: replace Abstract: Spatial reasoning (SR), the ability to infer 3D spatial information from 2D inputs, is essential for real-world applications such as embodied AI and

Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splatting

Model ReleasesDGX agent

arXiv:2606.30017v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting have demonstrated unprecedented success in novel view synthesis. However, the substantial inference and storage

MR-IQA: A Unified Margin View of Regression and Ranking for Blind Image Quality Assessment

SafetyDGX agent

arXiv:2606.29760v1 Announce Type: new Abstract: Blind image quality assessment (BIQA) is commonly built on two basic learning paradigms: regression and ranking. Regression calibrates absolute scores,

MSA-UNet3+: Multi-Scale Attention UNet3+ with New Supervised Prototypical Contrastive Loss for Coronary DSA Image Segmentation

Model ReleasesDGX agent

arXiv:2504.05184v4 Announce Type: replace-cross Abstract: Accurate segmentation of coronary Digital Subtraction Angiography (DSA) images is essential for diagnosing and treating coronary artery diseas

muFlow: Leveraging Average Images for Improving Generalisation of Deepfake Faces Detectors

SafetyDGX agent

arXiv:2606.30528v1 Announce Type: new Abstract: Current generative models, including GANs and diffusion models, have reached an outstanding level of photorealism, posing significant risks to privacy a

Multi-scale Object-Aware Gaze Estimation via Geometric Reasoning

Model ReleasesDGX agent

arXiv:2606.29334v1 Announce Type: new Abstract: Gaze target estimation aims to predict the semantic object an observer fixates upon within an image, a task deeply rooted in the object-oriented nature

Multimodal Graph RAG for Long-range Visually Rich Document Understanding

Model ReleasesDGX agent

arXiv:2606.28780v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are widely applied to visual document understanding. However, comprehending long documents remains an issue b

Multimodal Large Language Model driven Radiology Report Generation with Clinical Knowledge Enhancement

SafetyDGX agent

arXiv:2403.06728v2 Announce Type: replace Abstract: Radiology report generation (RRG) has attracted significant attention due to its potential to reduce the workload of radiologists. The performance o

Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers

TutorialsDGX agent

arXiv:2606.29013v1 Announce Type: new Abstract: Leveraging capabilities of large language models (LLMs) in text-to-image (T2I) synthesis is an important research direction. In this work we investigate

MUSE: Unlocking Timestep as Native Task Steering for One-Step Dense Prediction

Model ReleasesDGX agent

arXiv:2606.30370v1 Announce Type: new Abstract: Monocular dense prediction has recently seen remarkable success by repurposing pre-trained diffusion models. This opens a promising yet challenging aven

NaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene Generation

AgentsDGX agent

arXiv:2606.29395v1 Announce Type: new Abstract: Recently, Large Language Models (LLMs) have emerged as promising layout agents for 3D scene generation. Existing layout agents still suffer from implaus

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis

Model ReleasesDGX agent

arXiv:2606.29814v1 Announce Type: new Abstract: We propose Nemotron-Labs-Diffusion-Image, a state-of-the-art masked discrete diffusion model (MDM) for high-resolution text-to-image synthesis. Compared

Neural Gate: Mitigating Privacy Risks in LVLMs via Neuron-Level Gradient Gating

Local AiDGX agent

arXiv:2603.12598v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) have shown remarkable potential across a wide array of vision-language tasks, leading to their adoption in crit

Neural Stereo Video Compression with Hybrid Disparity Compensation

AgentsDGX agent

arXiv:2504.20383v3 Announce Type: replace Abstract: Disparity compensation represents the primary strategy in stereo video compression (SVC) for exploiting cross-view redundancy. These mechanisms can

Nonlinear mixture model motivated subspace clustering

Model ReleasesDGX agent

arXiv:2606.29261v1 Announce Type: cross Abstract: We derive the linear union-of-subspaces (UoS) model for subspace clustering (SC) from the nonlinear mixture model (NMM) used in blind source separatio

Obliviate: Erasing Concepts from Autoregressive Image Generation Models

Model ReleasesDGX agent

arXiv:2606.28643v1 Announce Type: new Abstract: The widespread adoption of generative AI models has intensified concerns about misuse, including the creation of unsafe or disturbing imagery. To mitiga

Occlusion-Robust Multi-Object Decoupling for Physics-Based Interaction

ApplicationsDGX agent

arXiv:2606.29303v1 Announce Type: new Abstract: We propose a mask-free method for lossless multi-object 3D reconstruction from sparse and occluded real-world views, enabling physically plausible inter

OmniCoT: A Benchmark for Global and Multi-Step Panoramic Reasoning

Model ReleasesDGX agent

arXiv:2606.30378v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated promising spatial reasoning capabilities, while these abilities remain underexplored in the e

OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data

ResearchDGX agent

arXiv:2606.30019v1 Announce Type: new Abstract: Music-driven dance video generation aims to synthesize expressive human motion that is temporally aligned with music while maintaining high visual fidel

On Test-Time Scaling for Vision-Language Models

ResearchDGX agent

arXiv:2606.28864v1 Announce Type: new Abstract: Test-time scaling is a paradigm where large models use additional compute at inference to achieve better performance, without changing model weights. Wh

On the Vulnerability of Parameter-Level Defenses to Model Merging

Model ReleasesDGX agent

arXiv:2606.30360v1 Announce Type: cross Abstract: The training-free integration of expert models via model merging has exposed significant security risks, enabling free-riders to combine specialized m

One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding

Local AiDGX agent

arXiv:2606.30084v1 Announce Type: new Abstract: MLLM-based GUI grounding methods commonly formulate target localization as autoregressive coordinate generation, enabling models to leverage the strong

OP3DSG: Open-Vocabulary Part-Aware 3D Scene Graph Generation for Real-World Environments

Model ReleasesDGX agent

arXiv:2606.29786v1 Announce Type: new Abstract: 3D scene graphs (3DSGs) provide a compact and structured abstraction of 3D environments. Although advances in foundation models have enabled open-vocabu

Open-Vocabulary and Referring Segmentation for 3D Gaussians Using 2D Detectors

TutorialsDGX agent

arXiv:2606.30638v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has emerged at the forefront of 3D scene reconstruction. Extending 3DGS with language-driven, open-vocabulary understanding

Optimizing Image Preparation and Compression for Face Recognition within 1024 Bytes

ResearchDGX agent

arXiv:2606.30321v1 Announce Type: new Abstract: ICAO-compliant machine readable travel documents enable automated biometric face verification. The biometric reference is stored on an RFID chip include

Orca: The World is in Your Mind

ResearchDGX agent

arXiv:2606.30534v1 Announce Type: new Abstract: We introduce Orca, an initial instantiation of a general world foundation model. Orca learns a unified world latent space from multimodal world signals

OWMDrive: Causality-Aware End-to-End Autonomous Driving via 4D Occupancy World Model

SafetyDGX agent

arXiv:2606.30421v1 Announce Type: new Abstract: Autonomous driving systems are steadily moving toward end-to-end paradigms to mitigate the limited adaptability of rule-based pipelines in complex traff

PCP-GAN: Property-Constrained Pore-scale image reconstruction via conditional Generative Adversarial Networks

ResearchDGX agent

arXiv:2510.19465v2 Announce Type: replace Abstract: Obtaining truly representative pore-scale images that match bulk formation properties remains a fundamental challenge in subsurface characterization

Personalizing MLLMs via Reinforced Multimodal Reference Game

ResearchDGX agent

arXiv:2606.28845v1 Announce Type: new Abstract: Personalizing Multimodal Large Language Models (MLLMs) aims to recognize users' unique concepts from visual data and provide personalized responses. Alt

PGE-SAM: Prompt-Guided Feature Enhancement for Interactive Segmentation under Degradation

Model ReleasesDGX agent

arXiv:2606.30477v1 Announce Type: new Abstract: Segment Anything Model (SAM) has revolutionized promptable image segmentation with strong zero-shot generalization. However, its performance degrades su

Physics-Grounded Disentangled Flow Modeling for Brain Disease Progression Trajectory

TutorialsDGX agent

arXiv:2606.28630v1 Announce Type: new Abstract: Forecasting longitudinal brain lesion evolution is critical for disease monitoring and treatment planning. Existing approaches typically learn a direct

PLOT: Pseudo-Labeling via Object Tracking for Monocular 3D Object Detection

AgentsDGX agent

arXiv:2507.02393v2 Announce Type: replace Abstract: Monocular 3D object detection is crucial for scalable perception across fields like autonomous driving, robotics, and surveillance. However, progres

Pointer-CAD v2: Plan-Then-Construct CAD Generation with Dimension-Aware Parametric Precision

Model ReleasesDGX agent

arXiv:2606.29301v1 Announce Type: new Abstract: Computer-aided design (CAD) plays a fundamental role in modern manufacturing by providing the high precision required for industrial production. Recent

PolarAPP: Beyond Polarization Demosaicking for Polarimetric Applications

SafetyDGX agent

arXiv:2603.23071v2 Announce Type: replace Abstract: Polarimetric imaging enables advanced vision applications such as normal estimation and de-reflection by capturing unique surface-material interacti

PoseShield: Neural Collision Fields for Human Self-Collision Resolution

Model ReleasesDGX agent

arXiv:2606.29686v1 Announce Type: new Abstract: Self-collision remains a persistent challenge in SMPL-based human pose estimation and motion generation. Under extreme articulations or stochastic motio

← Previous
1…6263646566…209
Next →