AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Agents

Geometry-Guided Representations for Coherent Lane and Traffic Topology Reasoning in Driving Scenes

DGX agent

arXiv:2506.13553v4 Announce Type: replace Abstract: Road topology reasoning is fundamental for autonomous driving, requiring both accurate perception of road elements and understanding of their comple

agentsarxiv-cs-cv
27 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

GLI-AL: A Multi-Modal Glioma MRI Label Resource with Unified Anatomy-Lesion Labels

DGX agent

arXiv:2607.22135v1 Announce Type: new Abstract: Existing BraTS-GLI datasets provide a widely used benchmark for adult glioma MRI segmentation, but their task definition focuses on tumor subregions and

model-releasesarxiv-cs-cv
27 Jul 2026
Research

Hash-QNeRF: Multiresolution Hash Encoding for Quantum Neural Radiance Fields

DGX agent

arXiv:2607.21675v1 Announce Type: cross Abstract: Neural Radiance Fields (NeRF) have revolutionized novel view synthesis, yet their classical implementations remain computationally intensive for high-

researcharxiv-cs-cv
27 Jul 2026
Research

Hiding Faces in Plain Sight: Defending DeepFakes by Disrupting Face Detection

DGX agent

arXiv:2412.01101v2 Announce Type: replace Abstract: Face-swapping DeepFakes have become an escalating societal concern, attracting increasing attention in recent years. To counter this, we investigate

researcharxiv-cs-cv
27 Jul 2026
Model Releases

Improving Large Vision-Language Models' Understanding for Flow Field Data

DGX agent

arXiv:2507.18311v3 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) have shown impressive capabilities across a range of tasks that integrate visual and textual understanding, suc

model-releasesarxiv-cs-cv
27 Jul 2026
Research

InnoText: A Unified Model for Visual Text Generation and Editing

DGX agent

arXiv:2607.22101v1 Announce Type: new Abstract: Diffusion models have recently achieved remarkable success in high-fidelity image synthesis, yet their application to visual text generation and editing

researcharxiv-cs-cv
27 Jul 2026
Model Releases

IR275K: A Benchmark for Infrared Multi-Frame Super-Resolution Toward Efficient Remote Sensing

DGX agent

arXiv:2607.22380v1 Announce Type: new Abstract: Efficient processing is becoming increasingly important in infrared remote sensing, where satellite constellations produce large volumes of observations

model-releasesarxiv-cs-cv
27 Jul 2026
Research

ISPCloak: Weaponizing ISP for Optimization-Free Physical Camouflage against Deepfake Detectors

DGX agent

arXiv:2607.21897v1 Announce Type: new Abstract: The rapid advancement of generative models has spurred the critical need to evaluate the worst-case robustness of deepfake detectors. In this paper, we

researcharxiv-cs-cv
27 Jul 2026
Research

Joint Lossless Compression and Steganography for Medical Images via Large Language Models

DGX agent

arXiv:2508.01782v4 Announce Type: replace-cross Abstract: Recently, large language models (LLMs) have driven promising progress in lossless image compression. However, directly adopting existing parad

researcharxiv-cs-cv
27 Jul 2026
Agents

JustDepth: Real-Time Radar-Camera Depth Estimation with Single-Scan LiDAR Supervision

DGX agent

arXiv:2607.22172v1 Announce Type: new Abstract: Accurate yet low-latency depth is essential for radar-camera perception in autonomous systems. Cameras provide rich appearance but lack metric scale, wh

agentsarxiv-cs-cv
27 Jul 2026
Research

Latent Interpolation Learning Using Diffusion Models for Cardiac Volume Reconstruction

DGX agent

arXiv:2508.13826v4 Announce Type: replace-cross Abstract: Cardiac Magnetic Resonance (CMR) imaging is a critical tool for diagnosing and managing cardiovascular disease, yet its utility is often limit

researcharxiv-cs-cv
27 Jul 2026
Safety

LayoutLite: Token-Level Implicit Layout Analysis for Efficient Document OCR

DGX agent

arXiv:2607.22200v1 Announce Type: new Abstract: End-to-end OCR systems based on vision-language models have achieved strong performance in complex document OCR, but their efficiency is limited by the

safetyarxiv-cs-cv
27 Jul 2026
Tutorials

Learning Adaptive Semantic Gaussian Allocation for 3D Occupancy

DGX agent

arXiv:2607.21896v1 Announce Type: new Abstract: Semantic 3D Gaussians provide a compact representation for 3D semantic occupancy prediction by rendering semantic primitives into a voxel volume under v

tutorialsarxiv-cs-cv
27 Jul 2026
Model Releases

LLM-Based Visual Explanation Evaluation Framework for Assessing the Explainability of Facial Skin Disease Classification Models

DGX agent

arXiv:2606.16794v2 Announce Type: replace Abstract: This study proposes a domain-specific LLM-based Visual Explanation Evaluation Framework for assessing visual attention explanations in facial skin d

model-releasesarxiv-cs-cv
27 Jul 2026
Research

Local Synaptic Rules Can Implement a SIGReg Gradient Without Backpropagation

DGX agent

arXiv:2607.21622v1 Announce Type: cross Abstract: We prove that two canonical local synaptic learning rules, the potentiation arm of spike-timing-dependent plasticity (STDP^+) and homeostatic plastici

researcharxiv-cs-cv
27 Jul 2026
Research

Low-Altitude Channel Multipath Prediction via Panoramic Perception and Vision-Language Model

DGX agent

arXiv:2607.21953v1 Announce Type: new Abstract: Unmanned aerial vehicle (UAV) communication is expected to support a wide range of low-altitude applications in 6G mobile networks. However, traditional

researcharxiv-cs-cv
27 Jul 2026
Model Releases

Medical-Checklist: Assessing the Comprehension of Medical Images by Multimodal Models

DGX agent

arXiv:2607.21998v1 Announce Type: new Abstract: This paper introduces a new benchmark test, Medical-Checklist, for assessing medical multimodal models. The recent advancements in multimodal models hav

model-releasesarxiv-cs-cv
27 Jul 2026
Research

Moving Beyond Diversity: Visual Token Pruning as Subspace Reconstruction for Efficient VLMs

DGX agent

arXiv:2606.18681v2 Announce Type: replace Abstract: Despite their remarkable performance, Vision Language Models (VLMs) incur substantial computational overhead due to the large number of visual token

researcharxiv-cs-cv
27 Jul 2026
Model Releases

NWaaS: A Non-Intrusive and Privacy-Preserving Watermarking-as-a-Service System with Adaptive Resource Scheduling

DGX agent

arXiv:2507.18036v2 Announce Type: replace-cross Abstract: Securing intellectual property (IP) in Machine Learning as a Service is critical yet challenging. While deep neural network watermarking serve

model-releasesarxiv-cs-cv
27 Jul 2026
Model Releases

OpenNavMap: Multi-Session Appearance-Based Topometric Mapping for Scalable Visual Navigation

DGX agent

arXiv:2601.12291v2 Announce Type: replace-cross Abstract: Scalable and maintainable maps are fundamental to large-scale navigation and the long-term deployment of robots in real-world environments. Ho

model-releasesarxiv-cs-cv
27 Jul 2026
Model Releases

Optimal Transport Image Representation and Deep Covariance Alignment (CORAL) for Control Valve Stiction Detection

DGX agent

arXiv:2607.22486v1 Announce Type: new Abstract: Control valve stiction is a common cause of unwanted oscillations and poor control-loop performance in industrial processes. Data-driven methods can aut

model-releasesarxiv-cs-cv
27 Jul 2026
Research

Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

DGX agent

arXiv:2607.21694v1 Announce Type: new Abstract: We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is

researcharxiv-cs-cv
27 Jul 2026
Research

Physiological Signals as a Forensic Modality for Talking-Face Deepfake Detection

DGX agent

arXiv:2607.21776v1 Announce Type: cross Abstract: Talking-face (TF) deepfake generation synthesizes photore- alistic facial video from a static source image and an au- dio signal, producing forgeries

researcharxiv-cs-cv
27 Jul 2026
Model Releases

Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images

DGX agent

arXiv:2607.21672v1 Announce Type: cross Abstract: Long source-code contexts consume many text tokens, motivating the proposal to render code as images for vision-language models. Recent work asks whet

model-releasesarxiv-cs-cv
27 Jul 2026
Safety

Projection Pursuit CPCANet for Domain Generalization

DGX agent

arXiv:2607.22117v1 Announce Type: new Abstract: Domain Generalization (DG) aims to learn representations robust to distribution shifts. Recent geometric alignment methods, such as CPCANet, extract dom

safetyarxiv-cs-cv
27 Jul 2026
Tutorials

Quality Text, Robust Vision: The Role of Language in Enhancing Visual Robustness of Vision-Language Models

DGX agent

arXiv:2507.16257v2 Announce Type: replace Abstract: Defending pre-trained vision-language models (VLMs), such as CLIP, against adversarial attacks is crucial, as these models are widely used in divers

tutorialsarxiv-cs-cv
27 Jul 2026
Model Releases

RadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding

DGX agent

arXiv:2607.22293v1 Announce Type: new Abstract: Medical multimodal large language models (MLLMs) are increasingly expected to perform complex image understanding tasks, yet their reliability is often

model-releasesarxiv-cs-cv
27 Jul 2026
Model Releases

ReCowGnition: A Realistic Biometric Benchmark for Cow Face Recognition

DGX agent

arXiv:2607.22071v1 Announce Type: new Abstract: With the development of precision livestock farming and the advances in computer vision, visual animal biometrics has gained attention. Using biometric

model-releasesarxiv-cs-cv
27 Jul 2026
Model Releases

Rethinking Layer-Wise Information Allocation for Vision Foundation Model Adaptation

DGX agent

arXiv:2607.21973v1 Announce Type: new Abstract: Vision foundation models are increasingly reused as frozen backbones for downstream visual recognition, making parameter-efficient adaptation a central

model-releasesarxiv-cs-cv
27 Jul 2026
Research

Rethinking Multi-Branch and Cross-Backbone Fusion for Vehicle Re-Identification in the Foundation-Model Era

DGX agent

arXiv:2607.22068v1 Announce Type: new Abstract: Multi-branch architectures and CNN-Transformer fusion have long been regarded as effective ways to improve vehicle re-identification (Re-ID) by combinin

researcharxiv-cs-cv
27 Jul 2026
Model Releases

Risk-Routed Implicit Boundary Refinement for Robust Ultrasound Image Segmentation

DGX agent

arXiv:2607.21787v1 Announce Type: new Abstract: Medical ultrasound (US) image segmentation faces significant challenges due to speckle noise, low-contrast boundaries, acoustic shadowing, and acquisiti

model-releasesarxiv-cs-cv
27 Jul 2026
Tutorials

Robot-Factored World Models via Robot Rendering

DGX agent

arXiv:2607.22535v1 Announce Type: cross Abstract: Action-conditioned video world models predict future observations from an initial observation and an action signal. In robotics, actions influence fut

tutorialsarxiv-cs-cv
27 Jul 2026
Local Ai

SCALE: Self-Supervised Constraint-Aware Layout GEneration for Local P&R DRV Fixing at Advanced Nodes

DGX agent

arXiv:2607.21850v1 Announce Type: new Abstract: As semiconductor manufacturing advances toward sub-2nm nodes, local place-and-route (P&R) design-rule violation (DRV) fixing is increasingly limited by

local-aiarxiv-cs-cv
27 Jul 2026
Model Releases

SceneActBench: Can Agents Act on the 3D Scenes They See?

DGX agent

arXiv:2607.22393v1 Announce Type: cross Abstract: Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual res

model-releasesarxiv-cs-cv
27 Jul 2026
Research

Self-Poisoning in Adaptive Out-of-Distribution Detection: A Sharp-Threshold Theory and Certified Label-Free Calibration

DGX agent

arXiv:2607.21673v1 Announce Type: cross Abstract: Test-time adaptive out-of-distribution (OOD) detectors update a memory bank from the unlabelled stream. We show this adaptation obeys a provable dynam

researcharxiv-cs-cv
27 Jul 2026
Research

SiPhy: Single-Image Physical Property Reasoning

DGX agent

arXiv:2607.22355v1 Announce Type: new Abstract: Inferring physical properties such as mass, stiffness, and elasticity from a single image is essential for simulation and embodied AI, yet most existing

researcharxiv-cs-cv
27 Jul 2026
Research

SLIP: Segmentation with Low-latency Interactive Prompting for 3D Medical Images

DGX agent

arXiv:2607.22332v1 Announce Type: new Abstract: Interactive deep image segmentation enables efficient medical image annotation by iteratively refining predictions from user prompts, such as positive a

researcharxiv-cs-cv
27 Jul 2026
Applications

SM4RT: Learning Structured Motion Geometry for 4D Reconstruction

DGX agent

arXiv:2607.22534v1 Announce Type: new Abstract: Geometry Foundation Models (GFMs) have substantially advanced monocular 3D reconstruction, yet extending this capability to 4D dynamic understanding rem

applicationsarxiv-cs-cv
27 Jul 2026
Safety

Spectral Prior for Reducing Exposure Bias in Diffusion Models

DGX agent

arXiv:2607.22091v1 Announce Type: new Abstract: Diffusion models typically suffer from error accumulation during iterative sampling, commonly referred to as exposure bias. We reveal systematic frequen

safetyarxiv-cs-cv
27 Jul 2026
Research

TDiR: Transformer based Diffusion for Image Restoration Tasks

DGX agent

arXiv:2506.20302v2 Announce Type: replace Abstract: Images captured in challenging environments often experience various types of degradation, such as noise, color cast, blur, and light scattering. Th

researcharxiv-cs-cv
27 Jul 2026
Safety

TextSLIP: Text Self-Supervised CLIP for Medical Report Generation

DGX agent

arXiv:2607.21970v1 Announce Type: new Abstract: Automating radiology report generation is important for improving reporting consistency and clinical workflows . While Contrastive Language--Image Pretr

safetyarxiv-cs-cv
27 Jul 2026
Model Releases

The 3D Mirage: Probing and Taming 3D Hallucinations

DGX agent

arXiv:2512.15423v2 Announce Type: replace Abstract: Monocular depth foundation models achieve remarkable generalization by learning large-scale semantic priors, but this creates a critical vulnerabili

model-releasesarxiv-cs-cv
27 Jul 2026
Model Releases

The Lift Spectrum: How Measurement-to-Space Adaptivity Shapes Robustness in Image-Free Single-Pixel Sensing

DGX agent

arXiv:2607.22077v1 Announce Type: cross Abstract: Single-pixel sensing encodes a scene as a short sequence of coded measurements, and image-free methods infer the task directly from that sequence. Rem

model-releasesarxiv-cs-cv
27 Jul 2026
Local Ai

The pretraining domain outweighs the training objective in setting the privacy-utility trade-off of differentially private medical image analysis

DGX agent

arXiv:2601.19618v2 Announce Type: replace Abstract: Differential privacy protects the patients whose images train medical imaging models, but it lowers diagnostic accuracy, and the initialization is t

local-aiarxiv-cs-cv
27 Jul 2026
Model Releases

Time-Reversed Imaging: A Multimodal Benchmark and Framework for Reconstructing Past Human-Environment Interactions

DGX agent

arXiv:2607.22352v1 Announce Type: new Abstract: We introduce time-reversed imaging, a new paradigm that infers what just happened in a scene from fading multimodal traces. Instead of extrapolating or

model-releasesarxiv-cs-cv
27 Jul 2026
Local Ai

Toward High-Fidelity 3D Point-Cloud Learning for Brain Folding Morphology Prediction Using Trans-Unet

DGX agent

arXiv:2607.21840v1 Announce Type: new Abstract: Learning high-fidelity point-cloud features in the 3D space poses significant challenges, including permutation invariance, lack of local context, diffi

local-aiarxiv-cs-cv
27 Jul 2026
Local Ai

TRaM-VSR: Importance-Aware Token Routing and Merging for One-Step Diffusion Video Super-Resolution

DGX agent

arXiv:2607.22231v1 Announce Type: new Abstract: Video super-resolution (VSR) using large-scale Diffusion Transformer (DiT) priors achieves exceptional perceptual quality but is often impractical due t

local-aiarxiv-cs-cv
27 Jul 2026
Safety

Twins: Learn to Predict Unified Representations with Focal Loss

DGX agent

arXiv:2607.22531v1 Announce Type: new Abstract: Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the

safetyarxiv-cs-cv
27 Jul 2026
← Previous
1…4142434445…261
Next →