AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Model Releases

Diverse-Intent Multi-Turn Fashion Image Retrieval

DGX agent

arXiv:2607.20291v1 Announce Type: new Abstract: Real-world fashion search involves interactive retrieval across multiple turns. However, existing multi-turn retrieval methods are built on a restrictiv

model-releasesarxiv-cs-cv
23 Jul 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization

DGX agent

arXiv:2607.18988v1 Announce Type: new Abstract: Medical imaging is a cornerstone of diagnostics, yet automated chest X-ray report generation struggles with structural adherence, anatomical completenes

model-releasesarxiv-cs-cv
23 Jul 2026
Research

Domain Shift in Echocardiography: Interpretable Quantification and Prediction of Cross-Dataset Left Ventricular Segmentation

DGX agent

arXiv:2607.19643v1 Announce Type: new Abstract: Cross-dataset generalisation remains a major barrier to clinical deployment of echocardiographic left ventricular segmentation, yet the sources of this

researcharxiv-cs-cv
23 Jul 2026
Model Releases

DRGBT-1K: A Large-scale High-quality Benchmark for Dynamic RGBT Tracking

DGX agent

arXiv:2607.19772v1 Announce Type: new Abstract: Dynamic RGBT (DRGBT) tracking aims to continuously localize a target when the available sensing modalities and observation platforms vary over time. Com

model-releasesarxiv-cs-cv
23 Jul 2026
Safety

Dual Adversarial Fine-tuning for Enhancing Robustness of Large Vision Language Model

DGX agent

arXiv:2607.18958v1 Announce Type: new Abstract: While Large Vision-Language Models (LVLMs), represented by LLaVA and GPT-4V, have demonstrated remarkable capabilities, their visual inputs remain vulne

safetyarxiv-cs-cv
23 Jul 2026
Research

Dual-Edged Homogeneous-Modality Similarity: Towards Visible-Infrared Modality-Incomplete Person Re-Identification with Modality Adaptive Matching

DGX agent

arXiv:2607.18688v1 Announce Type: new Abstract: Visible-Infrared Person Re-Identification (VI-ReID) operates under a closed-world assumption, where queries and galleries are from heterogeneous modalit

researcharxiv-cs-cv
23 Jul 2026
Research

DuSPiT: Dual-Branch Sub-Patch Pixel Diffusion Transformer

DGX agent

arXiv:2607.18510v1 Announce Type: new Abstract: Diffusion Transformers achieve strong image generation performance, but most operate in compressed latent spaces. Pixel-space diffusion avoids this info

researcharxiv-cs-cv
23 Jul 2026
Research

EAR-Net: Pursuing End-to-End Absolute Rotations from Multi-View Images

DGX agent

arXiv:2310.10051v3 Announce Type: replace Abstract: Absolute rotation estimation is an important topic in 3D computer vision. Existing works in literature generally employ a multi-stage (at least two-

researcharxiv-cs-cv
23 Jul 2026
Model Releases

ECoNGS: Efficient Compressive Neural Gaussian Splats for Volume Visualization

DGX agent

arXiv:2607.18466v1 Announce Type: new Abstract: Recent advances in differentiable Gaussian splatting have highlighted the potential of primitive-based approaches as alternative scene representations f

model-releasesarxiv-cs-cv
23 Jul 2026
Applications

Efficient Tracking and Understanding Object Transformations

DGX agent

arXiv:2607.19743v1 Announce Type: new Abstract: Tracking objects through state transformations is essential for understanding real-world dynamics. However, existing methods are computationally expensi

applicationsarxiv-cs-cv
23 Jul 2026
Safety

EGRNet: A Lightweight Semantic Segmentation Network with Edge-Gated Refinement and Adversarial Sensing

DGX agent

arXiv:2607.19617v1 Announce Type: new Abstract: As autonomous systems and smart cities continue to evolve, the demand for efficient and robust scene understanding becomes increasingly critical. Semant

safetyarxiv-cs-cv
23 Jul 2026
Research

ERank in Latent Space as an Image-Complexity and Richness Measure

DGX agent

arXiv:2607.19315v1 Announce Type: new Abstract: We propose the effective rank (ERank) of the channel covariance of an image's deep feature map as a per-sample, label-free measure of visual richness, c

researcharxiv-cs-cv
23 Jul 2026
Model Releases

ETPDesigner: Multi-Agent Orchestration for Interactive Multimodal Electronic Theater Program

DGX agent

arXiv:2607.19947v1 Announce Type: new Abstract: Electronic Theater Programs (ETPs) serve as critical promotional media in the performing arts, comprising a multi-page collection of heterogeneous visua

model-releasesarxiv-cs-cv
23 Jul 2026
Safety

Evolving Cache Schedules for Fast Diffusion Policy Inference

DGX agent

arXiv:2607.20293v1 Announce Type: new Abstract: Diffusion policies achieve strong visuomotor control by iteratively denoising action chunks, but repeated denoising makes real-time deployment computati

safetyarxiv-cs-cv
23 Jul 2026
Model Releases

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis

DGX agent

arXiv:2607.19341v1 Announce Type: new Abstract: Recent advances in multimodal generative models have enabled instruction-based image generation to move beyond semantic manipulation to knowledge-driven

model-releasesarxiv-cs-cv
23 Jul 2026
Tutorials

Extending a Large View Synthesis Model for Multi-view Panoptic Segmentation

DGX agent

arXiv:2607.19765v1 Announce Type: new Abstract: Large view synthesis models synthesize novel views through cross-view attention without explicit 3D representations, and recent studies have shown that

tutorialsarxiv-cs-cv
23 Jul 2026
Research

Factor-Informed Uncertainty Distillation for Gaze Estimation

DGX agent

arXiv:2607.20072v1 Announce Type: new Abstract: Deep gaze estimation works well in controlled capture but degrades in unconstrained settings, where systems must reject unreliable predictions. Single-p

researcharxiv-cs-cv
23 Jul 2026
Applications

FE-MCFormer: a novel time-frequency interpretable architecture for machinery fault diagnosis under strong noise environments

DGX agent

arXiv:2505.06285v3 Announce Type: replace-cross Abstract: Interpretable fault diagnosis (FD) plays a critical role in industrial manufacturing, as it improves human-machine understanding and operation

applicationsarxiv-cs-cv
23 Jul 2026
Model Releases

FedCC: A Low-Resource Federated Adaptation of Foundation Models for Robust Corpus Callosum localization in Fetal Ultrasound Images

DGX agent

arXiv:2607.18283v1 Announce Type: cross Abstract: Accurate localization of the corpus callosum (CC) in fetal ultrasound (US) images is crucial for the early identification of neurodevelopmental abnorm

model-releasesarxiv-cs-cv
23 Jul 2026
Model Releases

FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling

DGX agent

arXiv:2607.19038v1 Announce Type: new Abstract: Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary prose into long-

model-releasesarxiv-cs-cv
23 Jul 2026
Research

FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility

DGX agent

arXiv:2607.19100v1 Announce Type: new Abstract: Reconstructing animatable 3D human avatars from monocular video is a fundamental problem in computer vision with broad applications in AR/VR and digital

researcharxiv-cs-cv
23 Jul 2026
Model Releases

Fluid-SDF: Ultra-Lightweight and Editable Implicit Shape Representation via Differentiable Primitives

DGX agent

arXiv:2607.18646v1 Announce Type: new Abstract: Implicit Neural Representations (INRs) have become the standard for continuous 2D shape modeling, but they suffer from black-box uneditability, vulnerab

model-releasesarxiv-cs-cv
23 Jul 2026
Applications

Forecasting the Number of Harvest-ready Fruits of Sweet Peppers Using Multimodal Time-Series Data

DGX agent

arXiv:2607.19975v1 Announce Type: new Abstract: Accurate yield forecasting at the individual-plant level is critical for precision agriculture and supply-chain planning, yet public datasets capturing

applicationsarxiv-cs-cv
23 Jul 2026
Model Releases

Frequency-Hierarchical Active k-Space Sampling for Diagnostic MRI

DGX agent

arXiv:2607.19779v1 Announce Type: new Abstract: Active sampling for accelerated MRI must distribute a tight sampling budget across spatial frequencies that carry very different kinds of information. L

model-releasesarxiv-cs-cv
23 Jul 2026
Agents

From Distances to Trajectories: Real-Time Signed Distance Function Mapping and Distance-Accelerated Motion Planning for UAVs

DGX agent

arXiv:2607.19306v1 Announce Type: cross Abstract: Autonomous flight in cluttered environments requires a robot to build a geometric map of its surroundings and plan safe, dynamically feasible trajecto

agentsarxiv-cs-cv
23 Jul 2026
Hardware

From Pixel to Prognosis: Convolutional and GLCM Feature Fusion for Automated Four-Class Cataract Severity Classification

DGX agent

arXiv:2607.18349v1 Announce Type: new Abstract: Objective: To develop a low-cost automated cataract severity classification system operating on standard consumer-grade colour photographs of the eye, w

hardwarearxiv-cs-cv
23 Jul 2026
Model Releases

GATE-3D: Geometry-Aware Test-time Adaptive Reranking for Open-Set 3D Shape Retrieval

DGX agent

arXiv:2607.19111v1 Announce Type: new Abstract: Large pretrained vision models have substantially improved appearance-based 3D shape retrieval, but they still confuse shapes that look similar while di

model-releasesarxiv-cs-cv
23 Jul 2026
Agents

GaussianSeed: Hierarchical Gaussian Seeding for High-Resolution 3D Occupancy Prediction

DGX agent

arXiv:2607.20071v1 Announce Type: new Abstract: Vision-centric 3D occupancy prediction provides dense scene representations essential for autonomous driving and robotic navigation, yet existing method

agentsarxiv-cs-cv
23 Jul 2026
Research

Gaze-DETR: Top-Down Guidance Through Priority Maps for Infrared Weak-Small UAV Detection with DETR

DGX agent

arXiv:2607.19040v1 Announce Type: new Abstract: Infrared small target detection (ISTD) remains challenging because tiny, low-contrast targets are easily overwhelmed by clutter, noise, or occlusion. Co

researcharxiv-cs-cv
23 Jul 2026
Local Ai

Generalized Least Squares Kernelized Tensor Factorization

DGX agent

arXiv:2412.07041v4 Announce Type: replace-cross Abstract: Recovering incomplete multidimensional tensor-structured data is a fundamental task in many real-world applications. Smoothness-constrained lo

local-aiarxiv-cs-cv
23 Jul 2026
Research

Generative World Renderer at the Speed of Play

DGX agent

arXiv:2607.18703v1 Announce Type: new Abstract: Generative world renderer AlayaRenderer receives structured world states exported from physics engines and synthesizes RGB frames. Unlike models that ge

researcharxiv-cs-cv
23 Jul 2026
Local Ai

GeoTrace: Geometry-Aware Trajectory Token Compression for Video Large Language Models

DGX agent

arXiv:2607.09080v2 Announce Type: replace Abstract: Although Video Large Language Models (Video LLMs) have shown strong performance in video understanding, their efficiency is still limited by the lar

local-aiarxiv-cs-cv
23 Jul 2026
Model Releases

GH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision Tasks

DGX agent

arXiv:2512.24592v3 Announce Type: replace Abstract: Systematic failures of vision models on semantically coherent subsets, known as error slices, reveal limitations in robustness and evaluation. Exist

model-releasesarxiv-cs-cv
23 Jul 2026
Local Ai

Glass Surface Detection: Leveraging Reflection Dynamics in Flash/No-flash Imagery

DGX agent

arXiv:2511.16887v5 Announce Type: replace Abstract: Glass surfaces are ubiquitous in daily life, typically appearing colorless, transparent, and lacking distinctive features. These characteristics mak

local-aiarxiv-cs-cv
23 Jul 2026
Model Releases

GLID: Gated Local Intrinsic Dimension Repairs the Blind Spots of Face-Forgery Detectors

DGX agent

arXiv:2607.18770v1 Announce Type: cross Abstract: Fine-tuned foundation-model detectors dominate face-forgery benchmarks, yet they stay blind to generator families absent from training. We present GLI

model-releasesarxiv-cs-cv
23 Jul 2026
Safety

Global Building Area Estimation Products: How Accurate Are They?

DGX agent

arXiv:2607.19766v1 Announce Type: new Abstract: Geo-spatial rasters of building footprint area are useful for a variety of tasks, such as monitoring urbanization, improving energy efficiency, and trac

safetyarxiv-cs-cv
23 Jul 2026
Safety

Great X: A Unified Multi-Modal Simulator Bridging the Sim2Real Gap for 6G

DGX agent

arXiv:2507.08716v4 Announce Type: replace Abstract: Large-scale, precisely synchronized multi-modal datasets are critical for data-driven sixth-generation (6G) wireless research, yet real-world collec

safetyarxiv-cs-cv
23 Jul 2026
Research

Group-of-Latents: Perceptual Video Compression at Extreme Bitrates via Masked Latent Generative Modeling

DGX agent

arXiv:2607.19437v1 Announce Type: cross Abstract: Most existing video compression algorithms follow a paradigm of transformation and quantization, optimizing the trade-off between distortion and bitra

researcharxiv-cs-cv
23 Jul 2026
Safety

Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies

DGX agent

arXiv:2607.18325v1 Announce Type: new Abstract: Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and support decision-making during emergencies. Visi

safetyarxiv-cs-cv
23 Jul 2026
Safety

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement

DGX agent

arXiv:2607.18217v2 Announce Type: replace Abstract: Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two

safetyarxiv-cs-cv
23 Jul 2026
Research

How Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing Inspection

DGX agent

arXiv:2607.20263v1 Announce Type: new Abstract: Housing-level urban physical examination is essential for identifying residential building problems and supporting targeted urban renewal. Existing auto

researcharxiv-cs-cv
23 Jul 2026
Model Releases

IBoxCLA: Towards Robust Box-supervised Segmentation of Polyp via Improved Box-dice and Contrastive Latent-anchors

DGX agent

arXiv:2310.07248v5 Announce Type: replace Abstract: Box-supervised polyp segmentation attracts increasing attention for its cost-effective potential. Existing solutions often rely on learning-free met

model-releasesarxiv-cs-cv
23 Jul 2026
Applications

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

DGX agent

arXiv:2607.19228v1 Announce Type: new Abstract: Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear

applicationsarxiv-cs-cv
23 Jul 2026
Research

Image Editing Models are Numerical Solvers

DGX agent

arXiv:2607.18787v1 Announce Type: new Abstract: We investigate whether a pretrained generative image-editing model can provide a common interface for numerical simulation. Physical inputs and solution

researcharxiv-cs-cv
23 Jul 2026
Local Ai

IMMoE: Incomplete Multi-View Anomaly Detection via Mixture of View Experts Fusion

DGX agent

arXiv:2607.19032v1 Announce Type: new Abstract: Existing Multi-view Anomaly Detection (MAD) methods assume that all views are completely available and model each view separately. However, in real indu

local-aiarxiv-cs-cv
23 Jul 2026
Model Releases

Importance-Aware OBS Pruning for Diffusion Models

DGX agent

arXiv:2607.20048v1 Announce Type: new Abstract: We propose importance-aware pruning for diffusion models, a training-free framework that prioritizes preserving parameters critical to semantically sali

model-releasesarxiv-cs-cv
23 Jul 2026
Model Releases

In-Context Learning for Wound Classification with Small Multimodal Language Models

DGX agent

arXiv:2607.18819v1 Announce Type: new Abstract: Wound image classification is often treated as a task-specific supervised learning problem, requiring substantial amounts of manually labelled data and

model-releasesarxiv-cs-cv
23 Jul 2026
Hardware

InstantSfM: Towards GPU-Native SfM for the Deep Learning Era

DGX agent

arXiv:2510.13310v3 Announce Type: replace Abstract: Structure-from-Motion (SfM) is a fundamental technique for recovering camera poses and scene structure from multi-view imagery, serving as a critica

hardwarearxiv-cs-cv
23 Jul 2026
← Previous
1…4546474849…261
Next →