AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
8 Jul 2026

From RGB Generation to Dense Field Readout: Pixel-Space Dense Prediction with Text-to-Image Models

ResearchDGX agent

arXiv:2607.06553v1 Announce Type: new Abstract: Large-scale text-to-image models are attractive backbones for dense prediction because RGB generation pretraining learns rich semantic, structural, and

FUSE: A Flow-based Mapping Between Shapes

ResearchDGX agent

arXiv:2511.13431v2 Announce Type: replace Abstract: We introduce a novel neural representation for maps between 3D shapes based on flow-matching models, which is computationally efficient and supports

GaussFusion: Towards Multimodal 3D Gaussian Pretraining

SafetyDGX agent

arXiv:2607.05906v1 Announce Type: new Abstract: 3D Gaussian Splatting provides an explicit representation that jointly models geometry and appearance, serving as a scalable foundation for 3D represent


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

GEM-Occ: From Visual Geometry Evidence to Embodied Semantic Occupancy Memory

Model ReleasesDGX agent

arXiv:2607.05543v1 Announce Type: cross Abstract: Semantic occupancy provides a structured spatial memory for embodied indoor agents by jointly representing occupied regions, observed free space, unkn

Generalized Synthetic Image Detection with Enhanced RGB-Noise Representation Learning

Model ReleasesDGX agent

arXiv:2607.06354v1 Announce Type: new Abstract: The rapid advancement of large-scale generative models has accelerated the spread of highly deceptive AI-generated images, making generalized synthetic

GraspIT: A Dataset Bridging the Sim-to-Real gap and back for Validated Grasping SE(3) Pose Generation

SafetyDGX agent

arXiv:2607.05869v1 Announce Type: cross Abstract: Robust robotic grasping of novel objects requires datasets that simultaneously provide photorealistic RGB-D observations, physically validated grasp q

Ground3D-LMM: Fine-Grained 3D Point Grounding and Spatial Reasoning with LMM

ApplicationsDGX agent

arXiv:2607.05493v1 Announce Type: new Abstract: Natural-language queries about 3D environments become actionable when responses are verifiable and metric. Verifiability requires explicit grounding to

High-Resolution Artwork Outpainting with Global Blueprint Guidance and Layout Control

Local AiDGX agent

arXiv:2607.06162v1 Announce Type: new Abstract: Image outpainting extends an image beyond its original borders, requiring seamless style integration and globally coherent scene completion. Building on

HoloCount: A Holistic Visual Counting Benchmark for MLLMs

Model ReleasesDGX agent

arXiv:2607.06420v1 Announce Type: new Abstract: Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning. Wh

Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator

ApplicationsDGX agent

arXiv:2607.05765v1 Announce Type: new Abstract: Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world.

Imbalance-Robust and Sampling-Efficient Continuous Conditional GANs via Adaptive Vicinal Learning and Auxiliary Regularization

Local AiDGX agent

arXiv:2508.01725v5 Announce Type: replace-cross Abstract: Recent advances in continuous conditional generative modeling, including Continuous conditional Generative Adversarial Network (CcGAN) and Con

KOAL: Knowledge-Driven Prostate Cancer Grading with Ordinal-Aware Learning

SafetyDGX agent

arXiv:2607.06019v1 Announce Type: new Abstract: Non-invasive prediction of Gleason Grade Group (GGG) in prostate cancer using multiparametric MRI (mpMRI) is clinically vital for reducing unnecessary b

Label Hierarchy Transition: Delving into Class Hierarchies to Enhance Deep Classifiers

Model ReleasesDGX agent

arXiv:2112.02353v3 Announce Type: replace Abstract: Hierarchical classification aims to sort the object into a hierarchical structure of categories. For example, a bird can be categorized according to

LaViDa-R1: Advancing Reasoning for Unified Multimodal Diffusion Language Models

ResearchDGX agent

arXiv:2602.14147v2 Announce Type: replace Abstract: Diffusion language models (dLLMs) recently emerged as a promising alternative to auto-regressive LLMs. The latest works further extended it to multi

Learning to Throw Objects Safely in Multi-Obstacle Environments

SafetyDGX agent

arXiv:2607.06388v1 Announce Type: cross Abstract: Robotic throwing enables fast and efficient object placement beyond the robot's immediate workspace, but reliable throwing in cluttered environments r

Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation

ApplicationsDGX agent

arXiv:2607.06564v1 Announce Type: cross Abstract: Recently, Vision-Language-Action (VLA) models have demonstrated strong generalization across diverse tasks. However, effective robotic manipulation in

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory

HardwareDGX agent

arXiv:2607.05511v1 Announce Type: new Abstract: Agentic video understanding equips models with long-term memory to autonomously process and respond to continuous, long-horizon multimodal streams. Howe

LingDT-VL-OCR: Structure-Aware Document-Level Parsing with Fine-Grained Visual Reference

Model ReleasesDGX agent

arXiv:2603.11044v2 Announce Type: replace Abstract: In this paper, we propose LingDT-VL-OCR, a document parsing system tailored to financial-domain documents, transforming ultra-long financial PDFs in

LLM-Driven Neural Network Generation with Same-Family Architecture Guidance: Disentangling Transfer and Adaptation

Model ReleasesDGX agent

arXiv:2607.05704v1 Announce Type: cross Abstract: Large language models (LLMs) can generate neural-network modifications, but unrestricted generation is often invalid or harmful. This paper studies a

MAC-XA: Multi-view Anatomy-Correspondence Fusion for Coronary Stenosis Reporting from X-ray Angiography

SafetyDGX agent

arXiv:2607.06268v1 Announce Type: new Abstract: Multi-view reasoning in coronary X-ray angiography is inherently a cross-projection geometric problem, yet automated report generation in this setting r

mathbf{lambda}-VAE: Variance Equalization for Posterior Collapse

ResearchDGX agent

arXiv:2607.05531v1 Announce Type: cross Abstract: Variational Autoencoders (VAEs) frequently suffer from posterior collapse, a failure mode in which the approximate posterior converges to the prior, r

Mitigating Domain Shift in Conditioned Floor Plan Generation: Synthetic Pre-training for Data-Efficient Adaptation

ApplicationsDGX agent

arXiv:2607.06483v1 Announce Type: new Abstract: Robustness to domain shift is a key requirement for floor plan generative models to be applicable beyond the single dataset they were trained on, as flo

MobileWan: Closing the Quality Gap for Mobile Video Diffusion

Model ReleasesDGX agent

arXiv:2607.06173v1 Announce Type: new Abstract: Recent advances in video diffusion have been driven by scaling transformer-based architectures to billions of parameters, substantially improving visual

MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation

Model ReleasesDGX agent

arXiv:2607.06552v1 Announce Type: new Abstract: Infrared remote-sensing imagery captures intensity structure, object-background contrast, and illumination-invariant cues often invisible in RGB imagery

MorphGS: Morphology-Adaptive Articulated 3D Motion Transfer from Videos

ApplicationsDGX agent

arXiv:2601.02716v3 Announce Type: replace Abstract: Transferring articulated motion from monocular videos to rigged 3D characters is challenging due to pose ambiguity in 2D observations and morphologi

MoWorld: A Flash World Model

AgentsDGX agent

arXiv:2607.06216v1 Announce Type: new Abstract: The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate infe

MSA-DCNN: A Data-Efficient Multi-Scale Deformable CNN for Medical Image Classification

ResearchDGX agent

arXiv:2607.06083v1 Announce Type: new Abstract: Existing deep learning methods perform well in medical image classification but struggle with multi-scale morphology and limited annotations due to fixe

Multi-Teacher Contrastive Distillation for Edge-Efficient Pathology Foundation Models

Model ReleasesDGX agent

arXiv:2607.05533v1 Announce Type: new Abstract: Computational pathology foundation models (PFMs) have advanced whole-slide image analysis. However, their size and inference cost hinder local deploymen

NAMD: Virtual Follow-up Computed Tomography Synthesis via Nodule-Aligned Multimodal Diffusion Models for Early Lung Cancer Diagnosis

ResearchDGX agent

arXiv:2603.15932v2 Announce Type: replace Abstract: Lung cancer remains the leading cause of cancer-related mortality worldwide, with survival outcomes critically dependent on early and accurate detec

NumGrad-Pull: Numerical Gradient Guided Tri-plane Representation for Surface Reconstruction from Point Clouds

Local AiDGX agent

arXiv:2411.17392v3 Announce Type: replace Abstract: Reconstructing continuous surfaces from unoriented and unordered 3D points is a fundamental challenge in computer vision and graphics. Recent advanc

O3N: Omnidirectional Open-Vocabulary Occupancy Prediction

SafetyDGX agent

arXiv:2603.12144v2 Announce Type: replace Abstract: Understanding and reconstructing the 3D world through omnidirectional perception is becoming increasingly important for autonomous agents and embodi

OBBSeg: Irregular Lesion Segmentation under Oriented Bounding Box Annotations

SafetyDGX agent

arXiv:2607.06007v1 Announce Type: new Abstract: Pixel-level annotation remains a major bottleneck in medical image segmentation, making weak supervision an attractive yet under-constrained alternative

On the Redundancy of Timestep Embeddings in Diffusion Models

ResearchDGX agent

arXiv:2606.20416v2 Announce Type: replace-cross Abstract: Diffusion models rely heavily on explicit timestep embeddings to modulate the denoising process across various noise scales. In this work, we

Optimized Adaptive Loop Filter in Versatile Video Coding

Model ReleasesDGX agent

arXiv:2607.05737v1 Announce Type: new Abstract: In the Versatile Video Coding~(VVC) standard, adaptive loop filter~(ALF), including Geometry transformation-based Adaptive Loop Filter~(GALF) and Cross

OrchardBench: A Physically-Grounded, GPU-Parallel Apple-Orchard Simulation Benchmark for Agricultural Robotics

Model ReleasesDGX agent

arXiv:2607.06337v1 Announce Type: cross Abstract: Robotic tree-fruit harvesting is a flagship problem for agricultural automation, but progress is bottlenecked by the cost and irreproducibility of fie

Partial Symmetry Detection for 3D Geometry using Contrastive Learning with Geodesic Point Cloud Patches

Model ReleasesDGX agent

arXiv:2312.08230v2 Announce Type: replace Abstract: Detecting partial extrinsic symmetry in 3D geometry is a fundamental yet persistent challenge in computer vision and graphics, critical for tasks ra

Patch Knowledge Transfer for Efficient AI-Generated Image Quality Assessment

Local AiDGX agent

arXiv:2607.05605v1 Announce Type: new Abstract: With the rapid advancement of image generation technologies, perceptual quality assessment of AI-generated images has emerged as a crucial research dire

PhyMRI-SR: Toward Physics-Aware MRI Image Super-Resolution

ApplicationsDGX agent

arXiv:2607.06238v1 Announce Type: new Abstract: Magnetic resonance imaging (MRI) super-resolution is vital for improving diagnostic accessibility, yet most methods treat it as a deterministic mapping

PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation

Model ReleasesDGX agent

arXiv:2607.06440v1 Announce Type: new Abstract: Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personaliz

Point as Skeleton: Accumulated Point Cloud Enhanced Autoregressive Generation for Closed-Loop Autonomous Driving Simulation

AgentsDGX agent

arXiv:2607.06516v1 Announce Type: new Abstract: Evaluating end-to-end autonomous driving (E2E-AD) remains challenging, as existing driving simulation methods often trade off closed-loop interactivity

Pro-Pose: Unpaired Full-Body Portrait Synthesis via Canonical UV Maps

TutorialsDGX agent

arXiv:2512.17143v3 Announce Type: replace Abstract: Photographs of people taken by professional photographers typically present the person in beautiful lighting, with an interesting pose, and flatteri

Progressive Reasoning with Primitive Correction for Compositional Zero-Shot Learning

ResearchDGX agent

arXiv:2607.05911v1 Announce Type: new Abstract: Compositional Zero-Shot Learning (CZSL) aims to combine known attributes and objects as primitives for recognizing previously unseen attribute-object pa

ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation

Local AiDGX agent

arXiv:2607.06555v1 Announce Type: new Abstract: Tracking the six-degree-of-freedom (6-DoF) pose of objects and surfaces from monocular video is a long-standing problem in computer vision. To tackle th

RayRoPE: Projective Ray Positional Encoding for Multi-view Attention

ResearchDGX agent

arXiv:2601.15275v3 Announce Type: replace Abstract: We study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes pa

Realistic Compound-Lens Defocus Blur Synthesis

SafetyDGX agent

arXiv:2607.05837v1 Announce Type: new Abstract: Defocus blur degrades fine image structures and limits visual perception, which can adversely affect downstream vision tasks. Although recent deep learn

Recovering Cloud Microstructures with Cascaded Diffusion Inversion

SafetyDGX agent

arXiv:2607.05637v1 Announce Type: new Abstract: High-resolution satellite imagery is critical for observing fine-scale cloud structures that inform weather modification strategies like cloud seeding f

Reliable Mislabel Detection for Video Capsule Endoscopy Data

ResearchDGX agent

arXiv:2602.06938v2 Announce Type: replace Abstract: The classification performance of deep neural networks relies strongly on access to large, accurately annotated datasets. In medical imaging, howeve

Revisiting Scene Graph Generation from the Perspective of Detector-Conditioned Reachability

ResearchDGX agent

arXiv:2607.06176v1 Announce Type: new Abstract: Scene graph generation (SGG) approaches can be broadly classified into detector-based and query-based methods according to their underlying reasoning me

REVIVE: A Multi-Modal Framework for Vandalism Detection and Recovery in Autonomous Vehicles

AgentsDGX agent

arXiv:2607.05649v1 Announce Type: new Abstract: Autonomous vehicles (AVs) face increasing threats from vandalism-induced occlusion attacks (VOAs) that compromise camera-based perception. While detecti

RFHNet: Relational and Frequency-Aware Hashing Network for Large-Scale Fine-Grained Food Image Retrieval

Model ReleasesDGX agent

arXiv:2607.06148v1 Announce Type: new Abstract: Fine-grained food image retrieval is a key task in computational gastronomy, with applications in food traceability, dietary monitoring, and smart cater

RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes

ResearchDGX agent

arXiv:2601.05249v3 Announce Type: replace Abstract: Nighttime color constancy still remains a challenging problem in computational photography due to low-light noise and complex illumination condition

Robust Face Super-Resolution and Recognition Through Multi-Feature Aggregation in Diffusion Models

TutorialsDGX agent

arXiv:2607.05702v1 Announce Type: new Abstract: Images acquired in surveillance environments often suffer from conditions such as low resolution, variations in pose, irregular illumination, and occlus

RS-Agent: Automating Remote Sensing Tasks through Intelligent Agent

AgentsDGX agent

arXiv:2406.07089v4 Announce Type: replace Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have shown promise for remote sensing tasks such as visual question answering and scene

SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition

Model ReleasesDGX agent

arXiv:2509.25723v4 Announce Type: replace Abstract: Visual Place Recognition (VPR) requires robust retrieval of geotagged images despite large appearance, viewpoint, and environmental variation. Prior

SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs

Local AiDGX agent

arXiv:2607.05727v1 Announce Type: new Abstract: Pre-trained Vision-Language Models (VLMs) like CLIP have proven highly effective as foundation models for various downstream applications. However, prom

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models

ResearchDGX agent

arXiv:2607.05716v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong perception and reasoning capabilities. However, most existing models focus on isolated

Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots

Model ReleasesDGX agent

arXiv:2509.24966v2 Announce Type: replace Abstract: Understanding how people interact with their surroundings and each other is essential for enabling robots to act in socially compliant and context-a

SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation

ResearchDGX agent

arXiv:2607.05994v1 Announce Type: new Abstract: Human-Object Interaction (HOI) video generation aims to synthesize realistic videos of humans manipulating diverse objects, serving as a promising avenu

SpecTrack: Spectral Prompt Guided Adaptive Experts for Multispectral Object Tracking

ResearchDGX agent

arXiv:2607.05988v1 Announce Type: new Abstract: Multispectral image(MSI) and hyperspectral image(HSI) object tracking object tracking exploits recorded band-wise observations to improve target--backgr

SSA-3DGS: Unsupervised Removal of Screen-Space Artifacts for 3D Gaussian Splatting

ApplicationsDGX agent

arXiv:2607.05598v1 Announce Type: cross Abstract: Novel View Synthesis (NVS) methods, such as 3D Gaussian Splatting (3DGS), rely heavily on the assumption of clean, multi-view consistent, posed input

← Previous
1…4647484950…209
Next →