AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Model Releases

Label Hierarchy Transition: Delving into Class Hierarchies to Enhance Deep Classifiers

DGX agent

arXiv:2112.02353v3 Announce Type: replace Abstract: Hierarchical classification aims to sort the object into a hierarchical structure of categories. For example, a bird can be categorized according to

model-releasesarxiv-cs-cv
8 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

LaViDa-R1: Advancing Reasoning for Unified Multimodal Diffusion Language Models

DGX agent

arXiv:2602.14147v2 Announce Type: replace Abstract: Diffusion language models (dLLMs) recently emerged as a promising alternative to auto-regressive LLMs. The latest works further extended it to multi

researcharxiv-cs-cv
8 Jul 2026
Safety

Learning to Throw Objects Safely in Multi-Obstacle Environments

DGX agent

arXiv:2607.06388v1 Announce Type: cross Abstract: Robotic throwing enables fast and efficient object placement beyond the robot's immediate workspace, but reliable throwing in cluttered environments r

safetyarxiv-cs-cv
8 Jul 2026
Applications

Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation

DGX agent

arXiv:2607.06564v1 Announce Type: cross Abstract: Recently, Vision-Language-Action (VLA) models have demonstrated strong generalization across diverse tasks. However, effective robotic manipulation in

applicationsarxiv-cs-cv
8 Jul 2026
Hardware

Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory

DGX agent

arXiv:2607.05511v1 Announce Type: new Abstract: Agentic video understanding equips models with long-term memory to autonomously process and respond to continuous, long-horizon multimodal streams. Howe

hardwarearxiv-cs-cv
8 Jul 2026
Model Releases

LingDT-VL-OCR: Structure-Aware Document-Level Parsing with Fine-Grained Visual Reference

DGX agent

arXiv:2603.11044v2 Announce Type: replace Abstract: In this paper, we propose LingDT-VL-OCR, a document parsing system tailored to financial-domain documents, transforming ultra-long financial PDFs in

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

LLM-Driven Neural Network Generation with Same-Family Architecture Guidance: Disentangling Transfer and Adaptation

DGX agent

arXiv:2607.05704v1 Announce Type: cross Abstract: Large language models (LLMs) can generate neural-network modifications, but unrestricted generation is often invalid or harmful. This paper studies a

model-releasesarxiv-cs-cv
8 Jul 2026
Safety

MAC-XA: Multi-view Anatomy-Correspondence Fusion for Coronary Stenosis Reporting from X-ray Angiography

DGX agent

arXiv:2607.06268v1 Announce Type: new Abstract: Multi-view reasoning in coronary X-ray angiography is inherently a cross-projection geometric problem, yet automated report generation in this setting r

safetyarxiv-cs-cv
8 Jul 2026
Research

mathbf{lambda}-VAE: Variance Equalization for Posterior Collapse

DGX agent

arXiv:2607.05531v1 Announce Type: cross Abstract: Variational Autoencoders (VAEs) frequently suffer from posterior collapse, a failure mode in which the approximate posterior converges to the prior, r

researcharxiv-cs-cv
8 Jul 2026
Applications

Mitigating Domain Shift in Conditioned Floor Plan Generation: Synthetic Pre-training for Data-Efficient Adaptation

DGX agent

arXiv:2607.06483v1 Announce Type: new Abstract: Robustness to domain shift is a key requirement for floor plan generative models to be applicable beyond the single dataset they were trained on, as flo

applicationsarxiv-cs-cv
8 Jul 2026
Model Releases

MobileWan: Closing the Quality Gap for Mobile Video Diffusion

DGX agent

arXiv:2607.06173v1 Announce Type: new Abstract: Recent advances in video diffusion have been driven by scaling transformer-based architectures to billions of parameters, substantially improving visual

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

MonoIR-RS: Infrared Remote Sensing Vision-Language Learning with CLIP and VLM Adaptation

DGX agent

arXiv:2607.06552v1 Announce Type: new Abstract: Infrared remote-sensing imagery captures intensity structure, object-background contrast, and illumination-invariant cues often invisible in RGB imagery

model-releasesarxiv-cs-cv
8 Jul 2026
Applications

MorphGS: Morphology-Adaptive Articulated 3D Motion Transfer from Videos

DGX agent

arXiv:2601.02716v3 Announce Type: replace Abstract: Transferring articulated motion from monocular videos to rigged 3D characters is challenging due to pose ambiguity in 2D observations and morphologi

applicationsarxiv-cs-cv
8 Jul 2026
Agents

MoWorld: A Flash World Model

DGX agent

arXiv:2607.06216v1 Announce Type: new Abstract: The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate infe

agentsarxiv-cs-cv
8 Jul 2026
Research

MSA-DCNN: A Data-Efficient Multi-Scale Deformable CNN for Medical Image Classification

DGX agent

arXiv:2607.06083v1 Announce Type: new Abstract: Existing deep learning methods perform well in medical image classification but struggle with multi-scale morphology and limited annotations due to fixe

researcharxiv-cs-cv
8 Jul 2026
Model Releases

Multi-Teacher Contrastive Distillation for Edge-Efficient Pathology Foundation Models

DGX agent

arXiv:2607.05533v1 Announce Type: new Abstract: Computational pathology foundation models (PFMs) have advanced whole-slide image analysis. However, their size and inference cost hinder local deploymen

model-releasesarxiv-cs-cv
8 Jul 2026
Research

NAMD: Virtual Follow-up Computed Tomography Synthesis via Nodule-Aligned Multimodal Diffusion Models for Early Lung Cancer Diagnosis

DGX agent

arXiv:2603.15932v2 Announce Type: replace Abstract: Lung cancer remains the leading cause of cancer-related mortality worldwide, with survival outcomes critically dependent on early and accurate detec

researcharxiv-cs-cv
8 Jul 2026
Local Ai

NumGrad-Pull: Numerical Gradient Guided Tri-plane Representation for Surface Reconstruction from Point Clouds

DGX agent

arXiv:2411.17392v3 Announce Type: replace Abstract: Reconstructing continuous surfaces from unoriented and unordered 3D points is a fundamental challenge in computer vision and graphics. Recent advanc

local-aiarxiv-cs-cv
8 Jul 2026
Safety

O3N: Omnidirectional Open-Vocabulary Occupancy Prediction

DGX agent

arXiv:2603.12144v2 Announce Type: replace Abstract: Understanding and reconstructing the 3D world through omnidirectional perception is becoming increasingly important for autonomous agents and embodi

safetyarxiv-cs-cv
8 Jul 2026
Safety

OBBSeg: Irregular Lesion Segmentation under Oriented Bounding Box Annotations

DGX agent

arXiv:2607.06007v1 Announce Type: new Abstract: Pixel-level annotation remains a major bottleneck in medical image segmentation, making weak supervision an attractive yet under-constrained alternative

safetyarxiv-cs-cv
8 Jul 2026
Research

On the Redundancy of Timestep Embeddings in Diffusion Models

DGX agent

arXiv:2606.20416v2 Announce Type: replace-cross Abstract: Diffusion models rely heavily on explicit timestep embeddings to modulate the denoising process across various noise scales. In this work, we

researcharxiv-cs-cv
8 Jul 2026
Model Releases

Optimized Adaptive Loop Filter in Versatile Video Coding

DGX agent

arXiv:2607.05737v1 Announce Type: new Abstract: In the Versatile Video Coding~(VVC) standard, adaptive loop filter~(ALF), including Geometry transformation-based Adaptive Loop Filter~(GALF) and Cross

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

OrchardBench: A Physically-Grounded, GPU-Parallel Apple-Orchard Simulation Benchmark for Agricultural Robotics

DGX agent

arXiv:2607.06337v1 Announce Type: cross Abstract: Robotic tree-fruit harvesting is a flagship problem for agricultural automation, but progress is bottlenecked by the cost and irreproducibility of fie

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

Partial Symmetry Detection for 3D Geometry using Contrastive Learning with Geodesic Point Cloud Patches

DGX agent

arXiv:2312.08230v2 Announce Type: replace Abstract: Detecting partial extrinsic symmetry in 3D geometry is a fundamental yet persistent challenge in computer vision and graphics, critical for tasks ra

model-releasesarxiv-cs-cv
8 Jul 2026
Local Ai

Patch Knowledge Transfer for Efficient AI-Generated Image Quality Assessment

DGX agent

arXiv:2607.05605v1 Announce Type: new Abstract: With the rapid advancement of image generation technologies, perceptual quality assessment of AI-generated images has emerged as a crucial research dire

local-aiarxiv-cs-cv
8 Jul 2026
Applications

PhyMRI-SR: Toward Physics-Aware MRI Image Super-Resolution

DGX agent

arXiv:2607.06238v1 Announce Type: new Abstract: Magnetic resonance imaging (MRI) super-resolution is vital for improving diagnostic accessibility, yet most methods treat it as a deterministic mapping

applicationsarxiv-cs-cv
8 Jul 2026
Model Releases

PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation

DGX agent

arXiv:2607.06440v1 Announce Type: new Abstract: Recent text-to-image models such as DALLE-3 excel at following diverse prompts yet remain blind to individual aesthetic preferences. We study personaliz

model-releasesarxiv-cs-cv
8 Jul 2026
Agents

Point as Skeleton: Accumulated Point Cloud Enhanced Autoregressive Generation for Closed-Loop Autonomous Driving Simulation

DGX agent

arXiv:2607.06516v1 Announce Type: new Abstract: Evaluating end-to-end autonomous driving (E2E-AD) remains challenging, as existing driving simulation methods often trade off closed-loop interactivity

agentsarxiv-cs-cv
8 Jul 2026
Tutorials

Pro-Pose: Unpaired Full-Body Portrait Synthesis via Canonical UV Maps

DGX agent

arXiv:2512.17143v3 Announce Type: replace Abstract: Photographs of people taken by professional photographers typically present the person in beautiful lighting, with an interesting pose, and flatteri

tutorialsarxiv-cs-cv
8 Jul 2026
Research

Progressive Reasoning with Primitive Correction for Compositional Zero-Shot Learning

DGX agent

arXiv:2607.05911v1 Announce Type: new Abstract: Compositional Zero-Shot Learning (CZSL) aims to combine known attributes and objects as primitives for recognizing previously unseen attribute-object pa

researcharxiv-cs-cv
8 Jul 2026
Local Ai

ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation

DGX agent

arXiv:2607.06555v1 Announce Type: new Abstract: Tracking the six-degree-of-freedom (6-DoF) pose of objects and surfaces from monocular video is a long-standing problem in computer vision. To tackle th

local-aiarxiv-cs-cv
8 Jul 2026
Research

RayRoPE: Projective Ray Positional Encoding for Multi-view Attention

DGX agent

arXiv:2601.15275v3 Announce Type: replace Abstract: We study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes pa

researcharxiv-cs-cv
8 Jul 2026
Safety

Realistic Compound-Lens Defocus Blur Synthesis

DGX agent

arXiv:2607.05837v1 Announce Type: new Abstract: Defocus blur degrades fine image structures and limits visual perception, which can adversely affect downstream vision tasks. Although recent deep learn

safetyarxiv-cs-cv
8 Jul 2026
Safety

Recovering Cloud Microstructures with Cascaded Diffusion Inversion

DGX agent

arXiv:2607.05637v1 Announce Type: new Abstract: High-resolution satellite imagery is critical for observing fine-scale cloud structures that inform weather modification strategies like cloud seeding f

safetyarxiv-cs-cv
8 Jul 2026
Research

Reliable Mislabel Detection for Video Capsule Endoscopy Data

DGX agent

arXiv:2602.06938v2 Announce Type: replace Abstract: The classification performance of deep neural networks relies strongly on access to large, accurately annotated datasets. In medical imaging, howeve

researcharxiv-cs-cv
8 Jul 2026
Research

Revisiting Scene Graph Generation from the Perspective of Detector-Conditioned Reachability

DGX agent

arXiv:2607.06176v1 Announce Type: new Abstract: Scene graph generation (SGG) approaches can be broadly classified into detector-based and query-based methods according to their underlying reasoning me

researcharxiv-cs-cv
8 Jul 2026
Agents

REVIVE: A Multi-Modal Framework for Vandalism Detection and Recovery in Autonomous Vehicles

DGX agent

arXiv:2607.05649v1 Announce Type: new Abstract: Autonomous vehicles (AVs) face increasing threats from vandalism-induced occlusion attacks (VOAs) that compromise camera-based perception. While detecti

agentsarxiv-cs-cv
8 Jul 2026
Model Releases

RFHNet: Relational and Frequency-Aware Hashing Network for Large-Scale Fine-Grained Food Image Retrieval

DGX agent

arXiv:2607.06148v1 Announce Type: new Abstract: Fine-grained food image retrieval is a key task in computational gastronomy, with applications in food traceability, dietary monitoring, and smart cater

model-releasesarxiv-cs-cv
8 Jul 2026
Research

RL-AWB: Deep Reinforcement Learning for Auto White Balance Correction in Low-Light Night-time Scenes

DGX agent

arXiv:2601.05249v3 Announce Type: replace Abstract: Nighttime color constancy still remains a challenging problem in computational photography due to low-light noise and complex illumination condition

researcharxiv-cs-cv
8 Jul 2026
Tutorials

Robust Face Super-Resolution and Recognition Through Multi-Feature Aggregation in Diffusion Models

DGX agent

arXiv:2607.05702v1 Announce Type: new Abstract: Images acquired in surveillance environments often suffer from conditions such as low resolution, variations in pose, irregular illumination, and occlus

tutorialsarxiv-cs-cv
8 Jul 2026
Agents

RS-Agent: Automating Remote Sensing Tasks through Intelligent Agent

DGX agent

arXiv:2406.07089v4 Announce Type: replace Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have shown promise for remote sensing tasks such as visual question answering and scene

agentsarxiv-cs-cv
8 Jul 2026
Model Releases

SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition

DGX agent

arXiv:2509.25723v4 Announce Type: replace Abstract: Visual Place Recognition (VPR) requires robust retrieval of geotagged images despite large appearance, viewpoint, and environmental variation. Prior

model-releasesarxiv-cs-cv
8 Jul 2026
Local Ai

SAMPLe: SAM-based Optimizer for Prompt Learning in VLMs

DGX agent

arXiv:2607.05727v1 Announce Type: new Abstract: Pre-trained Vision-Language Models (VLMs) like CLIP have proven highly effective as foundation models for various downstream applications. However, prom

local-aiarxiv-cs-cv
8 Jul 2026
Research

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models

DGX agent

arXiv:2607.05716v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated strong perception and reasoning capabilities. However, most existing models focus on isolated

researcharxiv-cs-cv
8 Jul 2026
Model Releases

Social 3D Scene Graphs: Modeling Human Actions and Relations for Interactive Service Robots

DGX agent

arXiv:2509.24966v2 Announce Type: replace Abstract: Understanding how people interact with their surroundings and each other is essential for enabling robots to act in socially compliant and context-a

model-releasesarxiv-cs-cv
8 Jul 2026
Research

SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation

DGX agent

arXiv:2607.05994v1 Announce Type: new Abstract: Human-Object Interaction (HOI) video generation aims to synthesize realistic videos of humans manipulating diverse objects, serving as a promising avenu

researcharxiv-cs-cv
8 Jul 2026
Research

SpecTrack: Spectral Prompt Guided Adaptive Experts for Multispectral Object Tracking

DGX agent

arXiv:2607.05988v1 Announce Type: new Abstract: Multispectral image(MSI) and hyperspectral image(HSI) object tracking object tracking exploits recorded band-wise observations to improve target--backgr

researcharxiv-cs-cv
8 Jul 2026
Applications

SSA-3DGS: Unsupervised Removal of Screen-Space Artifacts for 3D Gaussian Splatting

DGX agent

arXiv:2607.05598v1 Announce Type: cross Abstract: Novel View Synthesis (NVS) methods, such as 3D Gaussian Splatting (3DGS), rely heavily on the assumption of clean, multi-view consistent, posed input

applicationsarxiv-cs-cv
8 Jul 2026
← Previous
1…5859606162…261
Next →