AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,618 results
Research

Rapid Quantification of Outdoor Object Visibility in Urban Setting Using Connected-Vehicle Fields of View

DGX agent

arXiv:2506.03365v3 Announce Type: replace-cross Abstract: Identifying locations that offer maximum visual exposure to passing vehicular traffic is a core problem in urban analytics, with applications

researcharxiv-cs-cv
23 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation

DGX agent

arXiv:2606.22749v1 Announce Type: new Abstract: Pre-trained Vision Foundation Models (VFMs) have become central to modern computer vision due to their powerful semantic representations and strong gene

researcharxiv-cs-cv
23 Jun 2026
Model Releases

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations

DGX agent

arXiv:2606.22766v1 Announce Type: new Abstract: Audio Description aims to generate concise narrations of essential visual content in audio-visual media for blind and low-vision audiences. Existing met

model-releasesarxiv-cs-cv
23 Jun 2026
Safety

Real-Time Multimodal Activity-Aware Error Detection in Robot-Assisted Surgery

DGX agent

arXiv:2606.23593v1 Announce Type: cross Abstract: Robot-assisted minimally invasive surgery improves surgical precision but introduces complexity, making technical error detection essential for ensuri

safetyarxiv-cs-cv
23 Jun 2026
Hardware

Real-time pedestrian attribute recognition with YOLOv8 and ResNet18

DGX agent

arXiv:2606.21200v1 Announce Type: new Abstract: Pedestrian attribute recognition (PAR) assigns semantic labels to detected pedestrians and is useful in surveillance, video retrieval, and human-centere

hardwarearxiv-cs-cv
23 Jun 2026
Model Releases

Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild

DGX agent

arXiv:2603.04205v2 Announce Type: replace Abstract: While Vision-Language Models (VLMs) achieve near-perfect scores on digital document benchmarks like OmniDocBench, their performance in the unpredict

model-releasesarxiv-cs-cv
23 Jun 2026
Research

ReconMIL: Synergizing Latent Space Reconstruction with Bi-Stream Mamba for Whole Slide Image Analysis

DGX agent

arXiv:2603.19925v2 Announce Type: replace-cross Abstract: Whole slide image (WSI) analysis heavily relies on multiple instance learning (MIL). While recent methods benefit from large-scale foundation

researcharxiv-cs-cv
23 Jun 2026
Model Releases

Region-Specific Calibration Achieves Excellent Inter-Device Reliability for Smartphone Dermatology: A Multi-Device Benchmark on Korean Facial Skin

DGX agent

arXiv:2512.21988v3 Announce Type: replace-cross Abstract: Background: Smartphone-based dermatology requires inter-device colorimetric reliability that holds across calibration regimes, yet quantitativ

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

REKEY: Metadata-Grounded Visual-Key Regeneration for Contamination-Resilient VQA Evaluation

DGX agent

arXiv:2606.20736v1 Announce Type: new Abstract: Static visual question answering (VQA) benchmarks age quickly: Once the items leak into training corpora, scores can reflect memorization rather than ge

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Reliability-Guided Adaptive Ensembling for Robust Test-Time Adaptation

DGX agent

arXiv:2606.22351v1 Announce Type: cross Abstract: Test-time adaptation (TTA) can mitigate domain shift without source data, but it is highly brittle under adversarially contaminated test streams, wher

researcharxiv-cs-cv
23 Jun 2026
Safety

RelightAnyone: A Generalized Relightable 3D Gaussian Head Model

DGX agent

arXiv:2601.03357v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has become a standard approach to reconstruct and render photorealistic 3D head avatars. A major challenge is to religh

safetyarxiv-cs-cv
23 Jun 2026
Research

Render-FM: Feedforward Model for Real-time Photorealistic Volumetric Rendering

DGX agent

arXiv:2505.17338v2 Announce Type: replace Abstract: Photorealistic volumetric rendering of CT scans greatly benefits clinical workflows, yet neural approaches such as Neural Radiance Fields (NeRF) and

researcharxiv-cs-cv
23 Jun 2026
Research

Resolving Multi-Target Association in OFDM-based ISAC via Vision-aided Multi-Modal Learning

DGX agent

arXiv:2606.22195v1 Announce Type: new Abstract: Orthogonal frequency division multiplexing (OFDM)-based integrated sensing and communication (ISAC) systems commonly extract target parameters by peak-s

researcharxiv-cs-cv
23 Jun 2026
Safety

Rethinking Object-Centric Representations for Video Dynamics Modeling

DGX agent

arXiv:2606.23436v1 Announce Type: new Abstract: Unsupervised video object tracking aims to decompose dynamic scenes into persistent, object-centric entities without manual annotations. Many recent app

safetyarxiv-cs-cv
23 Jun 2026
Tutorials

Rethinking Prototype-based Similarity Learning for Few-Shot Object Detection

DGX agent

arXiv:2606.23069v1 Announce Type: new Abstract: Few-shot object detection aims to detect novel object categories from only a few labeled examples, avoiding costly large-scale annotation. Recent protot

tutorialsarxiv-cs-cv
23 Jun 2026
Research

Rethinking the Adaptation of Vision Foundation Models for Efficient Cell Segmentation

DGX agent

arXiv:2606.21913v1 Announce Type: new Abstract: Cell segmentation is critical for computational pathology and biomedical discovery. While recent Vision Foundation Models (VFMs) have demonstrated remar

researcharxiv-cs-cv
23 Jun 2026
Research

Retrieval-Augmented Anatomical Guidance for Text-to-CT Generation

DGX agent

arXiv:2603.08305v2 Announce Type: replace Abstract: Text-conditioned generative models for volumetric medical imaging provide semantic control but lack explicit anatomical guidance, often resulting in

researcharxiv-cs-cv
23 Jun 2026
Safety

Robot Self-Improvement via Human-Video Dynamics Models

DGX agent

arXiv:2606.21406v1 Announce Type: cross Abstract: A central question in robot learning is how to acquire skills from the kinds of data that humans learn from: passive observation, embodied practice, a

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

Robust 3DGS-based SLAM via Adaptive Kernel Smoothing

DGX agent

arXiv:2511.23221v2 Announce Type: replace Abstract: In this paper, we challenge the conventional notion in 3DGS-SLAM that rendering quality is the primary determinant of tracking accuracy. We argue th

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Robust Image-Driven Phenotyping of Ovarian Tumor Cells using Optimized Dynamic Features in Hyperbolic Channels

DGX agent

arXiv:2606.20703v1 Announce Type: new Abstract: Label-free, image-based cellular mechanophenotyping in microfluidic devices provides a high-throughput method for single-cell profiling. However, while

researcharxiv-cs-cv
23 Jun 2026
Safety

Robust Representation Learning in Masked Autoencoders

DGX agent

arXiv:2602.03531v2 Announce Type: replace-cross Abstract: Masked Autoencoders (MAEs) achieve impressive performance in image classification tasks, yet the internal representations they learn remain le

safetyarxiv-cs-cv
23 Jun 2026
Applications

Robust Zero-Shot Generalization for Open-Vocabulary Action Recognition via Task Arithmetic

DGX agent

arXiv:2606.20734v1 Announce Type: new Abstract: Open Vocabulary Action Recognition (OVAR) enables the recognition of novel actions by leveraging vision-language representations, overcoming the limitat

applicationsarxiv-cs-cv
23 Jun 2026
Agents

Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City

DGX agent

arXiv:2606.20980v1 Announce Type: new Abstract: As Self-Driving Cars continue to expand internationally and use multi-modal systems such as VLMs as a cognitive backbone for their Action models; how we

agentsarxiv-cs-cv
23 Jun 2026
Safety

Rotation-Aware Point-Cloud Embeddings for Vision-Based In-Hand Reorientation

DGX agent

arXiv:2606.21788v1 Announce Type: cross Abstract: Point-cloud goals provide a direct way to specify dexterous in-hand reorientation: instead of defining an object-specific pose frame or estimating 6D

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation

DGX agent

arXiv:2606.23221v1 Announce Type: new Abstract: Recent years have witnessed remarkable progress in image generation and editing, particularly regarding instruction following and visual fidelity. Howev

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild

DGX agent

arXiv:2606.23344v1 Announce Type: new Abstract: Accurate document layout analysis remains a critical bottleneck for document parsing systems, due to the intricate coupling among heterogeneous document

model-releasesarxiv-cs-cv
23 Jun 2026
Safety

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation

DGX agent

arXiv:2511.20651v2 Announce Type: replace Abstract: Reinforcement learning (RL) has recently emerged as a promising approach for aligning text-to-image generative models with human preferences. A key

safetyarxiv-cs-cv
23 Jun 2026
Research

S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix

DGX agent

arXiv:2508.08048v2 Announce Type: replace Abstract: While video generation models excel at producing high-quality monocular videos, generating 3D stereoscopic and spatial videos for immersive applicat

researcharxiv-cs-cv
23 Jun 2026
Safety

Safe Few-Step Generation via Velocity Editing

DGX agent

arXiv:2606.23267v1 Announce Type: new Abstract: Flow matching has recently emerged as a strong paradigm for state-of-the-art text-to-image (T2I) generation, enabling high-quality generation with a sma

safetyarxiv-cs-cv
23 Jun 2026
Safety

SAGE: An Expert-Annotated South Asian GI Endoscopy Dataset for Multimodal Learning and Hallucination Analysis

DGX agent

arXiv:2606.22144v1 Announce Type: new Abstract: Gastrointestinal cancers represent a growing health burden in the South Asian region, driven largely by rapid changes in socio-economic conditions & lif

safetyarxiv-cs-cv
23 Jun 2026
Local Ai

SARIF: Segment Anything for Robust Image Forensics

DGX agent

arXiv:2606.21108v1 Announce Type: new Abstract: Image forgery localization remains challenging due to diverse manipulation techniques and distribution shifts. Existing forgery localization models achi

local-aiarxiv-cs-cv
23 Jun 2026
Model Releases

SATURN: Symbolic Spatial Reasoning for Multi-Perspective Grounding

DGX agent

arXiv:2606.22694v1 Announce Type: new Abstract: Vision-Language Models (VLMs) remain unreliable when spatial reasoning requires composing relations whose meanings depend on frames of reference. Existi

model-releasesarxiv-cs-cv
23 Jun 2026
Tutorials

Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology

DGX agent

arXiv:2606.20477v2 Announce Type: replace Abstract: We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations. We introduce RefRad2D, a lar

tutorialsarxiv-cs-cv
23 Jun 2026
Research

ScalePredictor: Instance-aware Scale Learning for Accurate Quantization of Vision Transformers

DGX agent

arXiv:2606.21947v1 Announce Type: new Abstract: Vision Transformers have achieved remarkable success in many fields, yet their deployment on edge devices remains challenging due to their substantial c

researcharxiv-cs-cv
23 Jun 2026
Local Ai

Scaling Diverse Language Generation for 3D Visual Grounding

DGX agent

arXiv:2606.20946v1 Announce Type: cross Abstract: Developing robust models for 3D visual grounding (3DVG), the localization of entities in a 3D scene described in natural language, is important for en

local-aiarxiv-cs-cv
23 Jun 2026
Safety

Scaling Self-Play for End-to-End Driving

DGX agent

arXiv:2606.19641v2 Announce Type: replace-cross Abstract: End-to-end autonomous driving models are typically trained on offline human-demonstration datasets that provide limited state coverage and oft

safetyarxiv-cs-cv
23 Jun 2026
Research

Scaling State-Space Models from Lines to Paragraphs: An Ablation of Mamba-based OCR

DGX agent

arXiv:2606.23524v1 Announce Type: new Abstract: End-to-end OCR increasingly relies on autoregressive sequence models, where the quadratic cost of Transformer attention limits efficient transcription o

researcharxiv-cs-cv
23 Jun 2026
Research

Scaling up fine-grained intracranial vessel annotations in computed tomography angiography

DGX agent

arXiv:2606.21756v1 Announce Type: cross Abstract: In this work, we present SemanticVessel, a dataset for fine-grained brain vessel segmentation in computed tomography angiography scans. Based on the d

researcharxiv-cs-cv
23 Jun 2026
Safety

ScalingAttention: Discovering Intrinsic Sparse Attention Topology for Video Diffusion Transformers

DGX agent

arXiv:2606.23019v1 Announce Type: new Abstract: While Diffusion Transformers (DiTs) have revolutionized high-fidelity video generation, their reliance on 3D full attention creates a quadratic computat

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

Scene-agnostic ALS boresight self-calibration

DGX agent

arXiv:2606.23101v1 Announce Type: new Abstract: ALS boresight calibration has relied for two decades on dedicated flight patterns over structured scenes containing planar surfaces of varied aspect and

model-releasesarxiv-cs-cv
23 Jun 2026
Applications

Scene-Level Heterogeneous Physics Simulation with 3D Gaussian Splats

DGX agent

arXiv:2606.21753v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) has achieved state-of-the-art photorealistic rendering, but the representation gap prevents these assets from being physi

applicationsarxiv-cs-cv
23 Jun 2026
Safety

SCOPE: Scale-Consistent One-Pass Estimation of 3D Geometry

DGX agent

arXiv:2606.21300v1 Announce Type: new Abstract: We present SCOPE (Scale-Consistent One-Pass Estimation of 3D Geometry), a novel approach for estimating 3D geometry from extended monocular video sequen

safetyarxiv-cs-cv
23 Jun 2026
Local Ai

SCRUB-FL: Sanitizing and Cleansing Representations via Unlearning of Backdoors

DGX agent

arXiv:2606.22700v1 Announce Type: cross Abstract: Federated Learning (FL) enables collaborative model training without sharing raw data, making it a promising paradigm for privacy-sensitive applicatio

local-aiarxiv-cs-cv
23 Jun 2026
Local Ai

SEED: Simple ViT and Evolving Harness for Explainable Text Forgery Detection

DGX agent

arXiv:2606.21138v1 Announce Type: new Abstract: AI-assisted image editing threatens trust in financial, legal, and identity records. The GenText-Forensics Challenge at ACM MM 2026 addresses this by re

local-aiarxiv-cs-cv
23 Jun 2026
Safety

Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation

DGX agent

arXiv:2606.19120v2 Announce Type: replace-cross Abstract: On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-level targets conditioned

safetyarxiv-cs-cv
23 Jun 2026
Model Releases

SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion

DGX agent

arXiv:2606.22568v1 Announce Type: new Abstract: Training image generation foundation models consumes substantial resources. Previous methods have attempted to leverage semantic guidance to accelerate

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

SegTME-UNI2: A Foundation Model-Based Framework for Generalisable Multiclass Cell Segmentation and LLM-Driven Tumour Microenvironment Characterisation in Histopathology

DGX agent

arXiv:2606.17702v2 Announce Type: replace Abstract: Characterising the tumour microenvironment (TME) from routine H&E-stained histology images requires simultaneous cell segmentation, feature extracti

model-releasesarxiv-cs-cv
23 Jun 2026
Research

Self-Supervised Dual-Frequency Phase Decomposition for Single-Shot Composite Fringe Projection Profilometry

DGX agent

arXiv:2606.21027v1 Announce Type: new Abstract: Single-shot fringe projection profilometry (FPP) has been actively studied for real-time measurement, dynamic object reconstruction, and motion-sensitiv

researcharxiv-cs-cv
23 Jun 2026
← Previous
1…101102103104105…263
Next →