AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Safety

DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework

DGX agent

arXiv:2506.05199v3 Announce Type: replace Abstract: A core task in embodied intelligence is ego-centric 3D visual grounding. Existing methods typically adopt two-stage, heterogeneous pipelines that pa

safetyarxiv-cs-cv
29 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

DenseScout: Algorithm-System Co-design for Budgeted Tiny Object Selection on Edge Platforms

DGX agent

arXiv:2604.25300v1 Announce Type: new Abstract: Deploying tiny object perception on edge platforms is challenging because practical systems must satisfy both strict compute budgets and end-to-end late

researcharxiv-cs-cv
29 Apr 2026
Model Releases

Detecting Dental Landmarks from Intraoral 3D Scans: the 3DTeethLand challenge

DGX agent

arXiv:2512.08323v2 Announce Type: replace Abstract: Teeth landmark detection is a key task in modern orthodontics, supporting advanced diagnosis, personalized treatment planning, and effective monitor

model-releasesarxiv-cs-cv
29 Apr 2026
Model Releases

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

DGX agent

arXiv:2603.29844v2 Announce Type: replace-cross Abstract: The development of Vision-Language-Action (VLA) models has been significantly accelerated by pre-trained Vision-Language Models (VLMs). Howeve

model-releasesarxiv-cs-cv
29 Apr 2026
Research

Diverse Image Priors for Black-box Data-free Knowledge Distillation

DGX agent

arXiv:2604.25794v1 Announce Type: cross Abstract: Knowledge distillation (KD) represents a vital mechanism to transfer expertise from complex teacher networks to efficient student models. However, in

researcharxiv-cs-cv
29 Apr 2026
Local Ai

DouC: Dual-Branch CLIP for Training-Free Open-Vocabulary Segmentation

DGX agent

arXiv:2604.24997v1 Announce Type: new Abstract: Open-vocabulary semantic segmentation requires assigning pixel-level semantic labels while supporting an open and unrestricted set of categories. Traini

local-aiarxiv-cs-cv
29 Apr 2026
Research

DualGeo: A Dual-View Framework for Worldwide Image Geo-localization

DGX agent

arXiv:2604.25533v1 Announce Type: new Abstract: Worldwide image geo-localization aims to infer the geographic location of an image captured anywhere on Earth, spanning street, city, regional, national

researcharxiv-cs-cv
29 Apr 2026
Research

Edge-Cloud Collaborative Reconstruction via Structure-Aware Latent Diffusion for Downstream Remote Sensing Perception

DGX agent

arXiv:2604.25319v1 Announce Type: new Abstract: The exponential surge in high-resolution remote sensing data faces a severe bottleneck in satellite-to-ground transmission. Limited downlink bandwidth f

researcharxiv-cs-cv
29 Apr 2026
Model Releases

ESICA: A Scalable Framework for Text-Guided 3D Medical Image Segmentation

DGX agent

arXiv:2604.24876v1 Announce Type: new Abstract: Text guided 3D medical image segmentation offers a flexible alternative to class based and spatial prompt based models by allowing users to specify regi

model-releasesarxiv-cs-cv
29 Apr 2026
Model Releases

Evaluating Computational Pathology Foundation Models for Prostate Cancer Grading under Distribution Shifts

DGX agent

arXiv:2410.06723v2 Announce Type: replace-cross Abstract: Pathology foundation models (PFMs) have emerged as powerful pretrained encoders for computational pathology, but their robustness under clinic

model-releasesarxiv-cs-cv
29 Apr 2026
Model Releases

Exploring Remote Photoplethysmography for Neonatal Pain Detection from Facial Videos

DGX agent

arXiv:2604.25680v1 Announce Type: new Abstract: Unaddressed pain in neonates can lead to adverse effects, including delayed development and slower weight gain, emphasising the need for more objective

model-releasesarxiv-cs-cv
29 Apr 2026
Tutorials

Exploring Time Conditioning in Diffusion Generative Models from Disjoint Noisy Data Manifolds

DGX agent

arXiv:2604.25289v1 Announce Type: cross Abstract: Practically, training diffusion models typically requires explicit time conditioning to guide the network through the denoising sampling process. Espe

tutorialsarxiv-cs-cv
29 Apr 2026
Model Releases

FCMBench-Video: Benchmarking Document Video Intelligence

DGX agent

arXiv:2604.25186v1 Announce Type: new Abstract: Document understanding is a critical capability in financial credit review, onboarding, and remote verification, where both decision accuracy and eviden

model-releasesarxiv-cs-cv
29 Apr 2026
Model Releases

Generalizable Human Gaussian Splatting via Multi-view Semantic Consistency

DGX agent

arXiv:2604.25466v1 Announce Type: new Abstract: Recently, generalizable human Gaussian splatting from sparse-view inputs has been actively studied for the photorealistic human rendering. Most existing

model-releasesarxiv-cs-cv
29 Apr 2026
Research

GeoSearch: Augmenting Worldwide Geolocalization with Web-Scale Reverse Image Search and Image Matching

DGX agent

arXiv:2604.25390v1 Announce Type: cross Abstract: Worldwide image geolocalization, which aims to predict the GPS coordinates of any image on Earth, remains challenging due to global visual diversity.

researcharxiv-cs-cv
29 Apr 2026
Model Releases

Golden RPG: Confidence-Adaptive Region-Aware Noise for Compositional Text-to-Image Generation

DGX agent

arXiv:2604.25314v1 Announce Type: new Abstract: Compositional text-to-image (T2I) generation requires a model to honour multiple sub-prompts that describe distinct image regions. Recent work shows tha

model-releasesarxiv-cs-cv
29 Apr 2026
Model Releases

GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment

DGX agent

arXiv:2604.25370v1 Announce Type: new Abstract: The release of GPT-image-2 by OpenAI marks a watershed moment in AI-generated imagery: the boundary between photographic reality and synthetic content h

model-releasesarxiv-cs-cv
29 Apr 2026
Applications

GramSR: Visual Feature Conditioning for Diffusion-Based Super-Resolution

DGX agent

arXiv:2604.25457v1 Announce Type: new Abstract: Despite recent advances, single-image super-resolution (SR) remains challenging, especially in real-world scenarios with complex degradations. Diffusion

applicationsarxiv-cs-cv
29 Apr 2026
Research

High-Precision Dichotomous Image Segmentation via Depth Integrity-Prior and Fine-Grained Patch Strategy

DGX agent

arXiv:2503.06100v5 Announce Type: replace Abstract: High-precision dichotomous image segmentation (DIS) is a task of extracting fine-grained objects from high-resolution images. Existing methods trade

researcharxiv-cs-cv
29 Apr 2026
Model Releases

HuM-Eval: A Coarse-to-Fine Framework for Human-Centric Video Evaluation

DGX agent

arXiv:2604.25361v1 Announce Type: new Abstract: Video generation models have developed rapidly in recent years, where generating natural human motion plays a pivotal role. However, accurately evaluati

model-releasesarxiv-cs-cv
29 Apr 2026
Safety

I-INR: Iterative Implicit Neural Representations

DGX agent

arXiv:2504.17364v4 Announce Type: replace Abstract: Implicit Neural Representations (INRs) have revolutionized signal processing and computer vision by modeling signals as continuous, differentiable f

safetyarxiv-cs-cv
29 Apr 2026
Research

IAM: Identity-Aware Human Motion and Shape Joint Generation

DGX agent

arXiv:2604.25164v1 Announce Type: new Abstract: Recent advances in text-driven human motion generation enable models to synthesize realistic motion sequences from natural language descriptions. Howeve

researcharxiv-cs-cv
29 Apr 2026
Model Releases

Image Classification via Random Dilated Convolution with Multi-Branch Feature Extraction and Context Excitation

DGX agent

arXiv:2604.25188v1 Announce Type: new Abstract: Image classification remains a fundamental yet challenging task in computer vision, particularly when fine-grained feature extraction and background noi

model-releasesarxiv-cs-cv
29 Apr 2026
Research

Image Compression with Bubble-Aware Frame Rate Adaptation for Energy-Efficient Video Capsule Endoscopy

DGX agent

arXiv:2604.25464v1 Announce Type: new Abstract: Video Capsule Endoscopy (VCE) is a promising method for improving the medical examination of the small intestine in the gastrointestinal tract. A key ch

researcharxiv-cs-cv
29 Apr 2026
Research

Improving Diversity in Black-box Few-shot Knowledge Distillation

DGX agent

arXiv:2604.25795v1 Announce Type: new Abstract: Knowledge distillation (KD) is a well-known technique to effectively compress a large network (teacher) to a smaller network (student) with little sacri

researcharxiv-cs-cv
29 Apr 2026
Research

Instruction-Evidence Contrastive Dual-Stream Decoding for Grounded Vision-Language Reasoning

DGX agent

arXiv:2604.25809v1 Announce Type: new Abstract: Vision-Language Models (VLMs) exhibit strong performance in instruction following and open-ended vision-language reasoning, yet they frequently generate

researcharxiv-cs-cv
29 Apr 2026
Safety

Interactive Episodic Memory with User Feedback

DGX agent

arXiv:2604.24893v1 Announce Type: new Abstract: In episodic memory with natural language queries (EM-NLQ), a user may ask a question (e.g., 'Where did I place the mug?') that requires searching a long

safetyarxiv-cs-cv
29 Apr 2026
Applications

Is the Modality Gap a Bug or a Feature? A Robustness Perspective

DGX agent

arXiv:2603.29080v2 Announce Type: replace Abstract: Many modern multi-modal models (e.g. CLIP) seek an embedding space in which the two modalities are aligned. Somewhat surprisingly, almost all existi

applicationsarxiv-cs-cv
29 Apr 2026
Safety

Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization

DGX agent

arXiv:2604.24952v1 Announce Type: new Abstract: Human visual preferences are inherently multi-dimensional, encompassing aesthetics, detail fidelity, and semantic alignment. However, existing datasets

safetyarxiv-cs-cv
29 Apr 2026
Research

Learning Illumination Control in Diffusion Models

DGX agent

arXiv:2604.24877v1 Announce Type: new Abstract: Controlling illumination in images is essential for photography and visual content creation. While closed-source models have demonstrated impressive ill

researcharxiv-cs-cv
29 Apr 2026
Local Ai

Leveraging Previous-Traversal Point Cloud Map Priors for Camera-Based 3D Object Detection and Tracking

DGX agent

arXiv:2604.25405v1 Announce Type: new Abstract: Camera-based 3D object detection and tracking are central to autonomous driving, yet precise 3D object localization remains fundamentally constrained by

local-aiarxiv-cs-cv
29 Apr 2026
Model Releases

Lightweight Real-Time Rendering Parameter Optimization via XGBoost-Driven Lookup Tables

DGX agent

arXiv:2604.25178v1 Announce Type: new Abstract: Achieving a desirable balance between rendering quality and real-time performance is a long-standing challenge in modern game and rendering engines, par

model-releasesarxiv-cs-cv
29 Apr 2026
Model Releases

M^3-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering

DGX agent

arXiv:2604.25122v1 Announce Type: new Abstract: We present M^3-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (ML

model-releasesarxiv-cs-cv
29 Apr 2026
Research

Magnification-Invariant Image Classification via Domain Generalization and Stable Sparse Embedding Signatures

DGX agent

arXiv:2604.25817v1 Announce Type: new Abstract: Magnification shift is a major obstacle to robust histopathology classification, because models trained on one imaging scale often generalize poorly to

researcharxiv-cs-cv
29 Apr 2026
Model Releases

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition

DGX agent

arXiv:2512.07348v2 Announce Type: replace Abstract: In controllable image generation, synthesizing coherent and consistent images from multiple reference inputs, i.e., Multi-Image Composition (MICo),

model-releasesarxiv-cs-cv
29 Apr 2026
Model Releases

MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding

DGX agent

arXiv:2512.17492v2 Announce Type: replace Abstract: Geo-spatial analysis of our world benefits from a multimodal approach, as every single geographic location can be described in numerous ways (images

model-releasesarxiv-cs-cv
29 Apr 2026
Research

MTPano: Multi-Task Panoramic Scene Understanding via Label-Free Integration of Dense Prediction Priors

DGX agent

arXiv:2602.05330v2 Announce Type: replace Abstract: Comprehensive panoramic scene understanding is critical for immersive applications, yet it remains challenging due to the scarcity of high-resolutio

researcharxiv-cs-cv
29 Apr 2026
Research

Multimodal Contextualized Support for Enhancing Video Retrieval System

DGX agent

arXiv:2412.07584v2 Announce Type: replace Abstract: Current video retrieval systems, especially those used in competitions, primarily focus on querying individual keyframes or images rather than encod

researcharxiv-cs-cv
29 Apr 2026
Research

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation

DGX agent

arXiv:2604.25819v1 Announce Type: new Abstract: In this work, we propose Mutual Forcing, a framework for fast autoregressive audio-video generation with long-horizon audio-video synchronization. Our a

researcharxiv-cs-cv
29 Apr 2026
Research

Natural Image Classification via Quasi-Cyclic Graph Ensembles and Random-Bond Ising Models at the Nishimori Temperature

DGX agent

arXiv:2508.18717v3 Announce Type: replace-cross Abstract: Modern multi-class image classification uses high-dimensional CNN features that incur large memory and computational costs and obscure the dat

researcharxiv-cs-cv
29 Apr 2026
Model Releases

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

DGX agent

arXiv:2604.24954v1 Announce Type: cross Abstract: We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, i

model-releasesarxiv-cs-cv
29 Apr 2026
Safety

NimbleReg: A light-weight deep-learning framework for diffeomorphic image registration

DGX agent

arXiv:2503.07768v2 Announce Type: replace Abstract: This paper presents NimbleReg, a light-weight deep-learning (DL) framework for diffeomorphic image registration leveraging surface representation of

safetyarxiv-cs-cv
29 Apr 2026
Safety

No Pedestrian Left Behind: Real-Time Detection and Tracking of Vulnerable Road Users for Adaptive Traffic Signal Control

DGX agent

arXiv:2604.25887v1 Announce Type: new Abstract: Current pedestrian crossing signals operate on fixed timing without adjustment to pedestrian behavior, which can leave vulnerable road users (VRUs) such

safetyarxiv-cs-cv
29 Apr 2026
Research

Novel 3D Binary Indexed Tree for Volume Computation of 3D Reconstructed Models from Volumetric Data

DGX agent

arXiv:2412.10441v2 Announce Type: replace-cross Abstract: In the burgeoning field of medical imaging, precise computation of 3D volume holds a significant importance for subsequent qualitative analysi

researcharxiv-cs-cv
29 Apr 2026
Local Ai

OmniAlpha: Aligning Transparency-Aware Generation via Multi-Task Unified Reinforcement Learning

DGX agent

arXiv:2511.20211v2 Announce Type: replace Abstract: Transparency-aware generation requires modeling not only RGB appearance but also alpha-based opacity and cross-layer composition, which are essentia

local-aiarxiv-cs-cv
29 Apr 2026
Local Ai

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding

DGX agent

arXiv:2604.25276v1 Announce Type: new Abstract: Video Temporal Grounding (VTG), the task of localizing video segments from text queries, struggles in open-world settings due to limited dataset scale a

local-aiarxiv-cs-cv
29 Apr 2026
Model Releases

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations

DGX agent

arXiv:2604.25102v1 Announce Type: new Abstract: Typographic prompt injection exploits vision language models' (VLMs) ability to read text rendered in images, posing a growing threat as VLMs power auto

model-releasesarxiv-cs-cv
29 Apr 2026
Model Releases

OneThinker: All-in-one Reasoning Model for Image and Video

DGX agent

arXiv:2512.03043v3 Announce Type: replace Abstract: Reinforcement learning (RL) has recently achieved remarkable success in eliciting visual reasoning within Multimodal Large Language Models (MLLMs).

model-releasesarxiv-cs-cv
29 Apr 2026
← Previous
1…208209210211212…261
Next →