AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,515 results
Safety

RadPRISM: Schema-stratified radiology-report supervision for concept-disentangled image representations and visual grounding

DGX agent

arXiv:2608.00147v1 Announce Type: new Abstract: Vision-language pretraining learns rich medical image representations from radiology reports, but previous model variants commonly operate within a sing

safetyarxiv-cs-cv
4 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Local Ai

RadYOLO: Computationally Efficient 3D Object Detection and Segmentation in CT and MRI

DGX agent

arXiv:2608.00508v1 Announce Type: new Abstract: Object detection and segmentation in three-dimensional medical images is a very active area of research. However, most proposed deep learning models car

local-aiarxiv-cs-cv
4 Aug 2026
Model Releases

Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Metric for Infrared-Visible Fusion Assessment

DGX agent

arXiv:2608.01301v1 Announce Type: new Abstract: Infrared-visible image fusion (IVIF) has no ideal fused reference, so fusion algorithms are routinely ranked by scalar objective metrics that formalize

model-releasesarxiv-cs-cv
4 Aug 2026
Research

ReACT-CLIP: Response-Aware Test-Time Defense for Vision--Language Models

DGX agent

arXiv:2608.01067v1 Announce Type: new Abstract: Training-free test-time defenses offer a practical way to improve the adversarial robustness of CLIP-style vision--language models without modifying the

researcharxiv-cs-cv
4 Aug 2026
Model Releases

Real-Time Visual Obstruction Detection in Surgical Augmented Reality

DGX agent

arXiv:2608.00232v1 Announce Type: new Abstract: Surgical augmented reality (AR) can provide contextual guidance by overlaying virtual annotations, tool cues, and procedural information onto the surgic

model-releasesarxiv-cs-cv
4 Aug 2026
Research

Recompute or Reuse? Diagnosing and Mitigating Textual Shortcuts in VLM Self-Reflection

DGX agent

arXiv:2608.01930v1 Announce Type: new Abstract: Vision-language models (VLMs) are expected to revise their reasoning when visual evidence changes. Failures to do so are often attributed to insufficien

researcharxiv-cs-cv
4 Aug 2026
Local Ai

Reconstruction-Shift Discrimination via Mask-Guided Latent Diffusion for Medical Anomaly Detection

DGX agent

arXiv:2608.00444v1 Announce Type: new Abstract: Unsupervised medical anomaly detection learns normal anatomical patterns from healthy training images and identifies deviations at test time. Reconstruc

local-aiarxiv-cs-cv
4 Aug 2026
Model Releases

Recursive Vision Language Models for General Symbolic Reasoning

DGX agent

arXiv:2608.01534v1 Announce Type: new Abstract: Hard symbolic-reasoning tasks such as Sudoku, maze pathfinding, and ARC remain challenging for LLMs due to their fixed-depth autoregressive reasoning, w

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts

DGX agent

arXiv:2608.00574v1 Announce Type: new Abstract: Vision-language MoE batches contain different numbers of image and text tokens. Image resolution, image count, tiling, and prompt length all change this

model-releasesarxiv-cs-cv
4 Aug 2026
Research

Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning

DGX agent

arXiv:2608.01314v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) increasingly rely on long chain-of-thought reasoning for complex tasks. However, as reasoning sequences lengthe

researcharxiv-cs-cv
4 Aug 2026
Applications

ReMiX-MAE: Learning Missing-Channel Cross-Modal Representations from RGB-Only Clinical Facial Videos for Sympathetic-Mediated Pain Assessment

DGX agent

arXiv:2608.02561v1 Announce Type: new Abstract: Automated pain assessment in real clinics is limited by scarce clinically grounded facial video data with weak labels (often sequence-level self-report)

applicationsarxiv-cs-cv
4 Aug 2026
Research

Representation Transfer of Foundation Models for Ultra-Widefield Retinal Imaging

DGX agent

arXiv:2608.00586v1 Announce Type: new Abstract: Despite the widespread adoption of foundation models as feature extractors for medical imaging, relatively little is understood about how different pret

researcharxiv-cs-cv
4 Aug 2026
Tutorials

Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance

DGX agent

arXiv:2510.21590v3 Announce Type: replace Abstract: Current image super-resolution methods show strong performance on natural images but distort text, creating a fundamental trade-off between image qu

tutorialsarxiv-cs-cv
4 Aug 2026
Research

Rethinking IRSTD: Single-Point Supervision Guided Encoder-only Framework is Enough for Infrared Small Target Detection

DGX agent

arXiv:2604.05363v2 Announce Type: replace Abstract: Infrared small target detection (IRSTD) aims to separate small targets from clutter backgrounds. Extensive research is dedicated to the pixel-level

researcharxiv-cs-cv
4 Aug 2026
Research

Rethinking Pretraining for Specialized Design Data: Evidence from the JONES-19 Cultural Design Dataset

DGX agent

arXiv:2608.00135v1 Announce Type: cross Abstract: Design and architectural archives encode expert human knowledge in graphical formats, providing a critical testbed for design-inspired Machine Learnin

researcharxiv-cs-cv
4 Aug 2026
Research

Rethinking Video Token Compression with a Global Codebook: Learning Once, Compressing Everywhere

DGX agent

arXiv:2608.01271v1 Announce Type: new Abstract: Video large language models (Video-LLMs) represent videos as dense sequences of visual tokens, whose length grows with the temporal and spatial extent o

researcharxiv-cs-cv
4 Aug 2026
Research

Retrieval-Based Cross-Domain Generalization in Optical Networks via Global Features

DGX agent

arXiv:2608.00044v1 Announce Type: cross Abstract: We propose a retrieval-based framework for crossdomain quality-of-transmission (QoT) estimation that leverages transferable feature representations wh

researcharxiv-cs-cv
4 Aug 2026
Model Releases

Robust Watermarks Meet Backdoored Models: Evading Diffusion Semantic Watermarks via Stealthy Backdoor

DGX agent

arXiv:2608.00543v1 Announce Type: cross Abstract: Although semantic watermarking is considered a promising safeguard for images generated by Latent Diffusion Models (LDMs), the reliance of the waterma

model-releasesarxiv-cs-cv
4 Aug 2026
Research

Rolling Shutter Camera Self-Calibration

DGX agent

arXiv:2608.01509v1 Announce Type: new Abstract: Rolling shutter (RS) cameras are widely used in consumer devices, but their row-wise exposure causes distortions under motion, making geometric 3D visio

researcharxiv-cs-cv
4 Aug 2026
Local Ai

Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis

DGX agent

arXiv:2608.01973v1 Announce Type: cross Abstract: Existing indoor layout generators produce globally plausible layouts yet may retain local violations such as collisions, out-of-bounds placements, obs

local-aiarxiv-cs-cv
4 Aug 2026
Applications

RPL-UIE: Reliable Prior Learning for Underwater Image Enhancement

DGX agent

arXiv:2608.00137v1 Announce Type: cross Abstract: Underwater image enhancement (UIE) aims to recover clear images from observations affected by wavelength-dependent absorption, scattering, and spatial

applicationsarxiv-cs-cv
4 Aug 2026
Model Releases

RSC-GestureNet: Reliability-Aware Selective Causal Recognition of Chinese Traffic Police Gestures

DGX agent

arXiv:2608.02200v1 Announce Type: new Abstract: Traffic police gestures are safety-critical perception cues for autonomous driving. A deployable recognizer must infer commands causally from continuous

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

RSRA: Training-Free Probing of Representation Sensitivity for Efficient LoRA Rank Allocation

DGX agent

arXiv:2607.09757v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning enables large language models to adapt to downstream tasks with substantially lower computational and storage cost,

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?

DGX agent

arXiv:2608.02039v1 Announce Type: new Abstract: Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, acti

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

SafeBuild-Bench: A Temporal-Robust Construction Safety Benchmark with Graph-Enhanced Data Mining

DGX agent

arXiv:2608.00068v1 Announce Type: new Abstract: Construction-safety models must handle concrete deployment risks, such as a worker standing near a scaffold edge without guardrails, rather than only re

model-releasesarxiv-cs-cv
4 Aug 2026
Safety

Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression

DGX agent

arXiv:2608.02109v1 Announce Type: new Abstract: Vision-Text Compression (VTC) renders long texts into images and encodes them through the vision encoder (ViT), compressing thousands of text tokens int

safetyarxiv-cs-cv
4 Aug 2026
Local Ai

SARe: Structure-Aware Generative 3D Fragment Reassembly

DGX agent

arXiv:2603.21611v2 Announce Type: replace Abstract: 3D fragment reassembly estimates the rigid pose of each fragment to recover a complete object from unordered point clouds or meshes. The task become

local-aiarxiv-cs-cv
4 Aug 2026
Applications

SCALP: Semi-Supervised Statistical Shape Modeling from Imperfect 3D Photogrammetry via Landmark-Anchored Spectral Warp

DGX agent

arXiv:2608.00187v1 Announce Type: new Abstract: Correspondence-based statistical shape modeling (SSM) is vital for population-level morphometric analysis, but conventional pipelines assume clean, full

applicationsarxiv-cs-cv
4 Aug 2026
Applications

Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds

DGX agent

arXiv:2608.00463v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) turns captured or generated imagery into photorealistic 3D world simulations that users can freely explore, yet these world

applicationsarxiv-cs-cv
4 Aug 2026
Research

Score-Based Turbo Message Passing for Plug-and-Play Compressive Imaging

DGX agent

arXiv:2512.14435v2 Announce Type: replace Abstract: Message-passing algorithms have been adapted for compressive imaging by incorporating various off-the-shelf image denoisers. However, these denoiser

researcharxiv-cs-cv
4 Aug 2026
Research

SecondOpinion: Anatomy-Aware Gated Reasoning for Efficient Medical Image Analysis

DGX agent

arXiv:2608.01808v1 Announce Type: new Abstract: Deep learning models for medical image analysis typically apply a fixed amount of computation to every input, regardless of case difficulty. Anatomy-gui

researcharxiv-cs-cv
4 Aug 2026
Research

Seeing the Unseen: Towards Training-Free Inspection for Wind Turbine Blades Using Knowledge-Augmented Vision Language Models

DGX agent

arXiv:2510.22868v2 Announce Type: replace Abstract: Wind turbine blades operate in harsh environments, making timely damage detection essential for preventing failures and optimizing maintenance. Dron

researcharxiv-cs-cv
4 Aug 2026
Research

Self-supervised DXA representations encode multi-system disease risk, biological aging and heritability

DGX agent

arXiv:2608.02208v1 Announce Type: new Abstract: Whole-body dual-energy X-ray absorptiometry (DXA) scans are routinely acquired to measure bone density and regional body composition, leaving their spat

researcharxiv-cs-cv
4 Aug 2026
Model Releases

Semantic-Guided Cross-Sensor Super Resolution of Remote Sensing Images: A Gated Dual Conditioning Flow Matching Model

DGX agent

arXiv:2510.23816v3 Announce Type: replace Abstract: High spatial resolution satellite imagery is critical for monitoring fine-scale Earth surface processes, but is often limited by cost and revisit ti

model-releasesarxiv-cs-cv
4 Aug 2026
Safety

Semantically Calibrated Evidence Composition for CT Vision-Language Learning

DGX agent

arXiv:2608.00239v1 Announce Type: new Abstract: Learning transferable representations from CT-report pairs requires combining whole-volume context with anatomy-specific evidence. Existing methods typi

safetyarxiv-cs-cv
4 Aug 2026
Safety

Sen-Cap: Sensor-Flexible and Noise-Resilient Human Motion Capture via LiDAR-Camera Integration

DGX agent

arXiv:2608.02285v1 Announce Type: new Abstract: We propose Sen-Cap, a Sensor-Flexible and Noise-Resilient 3D human motion Capture framework that integrates multi-modal data from LiDAR and camera. Whil

safetyarxiv-cs-cv
4 Aug 2026
Safety

SG-Layout: Structured Scene Graph-Guided Layout Generation with LLMs

DGX agent

arXiv:2608.01106v1 Announce Type: new Abstract: Understanding and generating spatially coherent layouts from natural language remains a fundamental yet challenging task for large language models (LLMs

safetyarxiv-cs-cv
4 Aug 2026
Safety

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space

DGX agent

arXiv:2608.01397v1 Announce Type: cross Abstract: World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on whether future dynamics are model

safetyarxiv-cs-cv
4 Aug 2026
Research

Similarity Weighted Aggregation with Global Differential Privacy for Federated Brain Lesion Segmentation

DGX agent

arXiv:2608.00872v1 Announce Type: cross Abstract: Federated Learning (FL) enables collaborative training of machine learning models across multiple institutions without sharing sensitive data, making

researcharxiv-cs-cv
4 Aug 2026
Safety

SPAE: Spectrally Guided Autoencoder for Pretrained Visual Latents

DGX agent

arXiv:2608.01306v1 Announce Type: new Abstract: Latents from vision foundation models (VFMs) are semantically rich and well suited for visual understanding. Recent representation autoencoder methods s

safetyarxiv-cs-cv
4 Aug 2026
Model Releases

SPARC-Rad: A Multimodal Benchmark Dataset and Evaluation Pipeline for Spatial and Anatomical Reasoning in Radiology Vision-Language Models

DGX agent

arXiv:2608.00100v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly being evaluated for medical imaging, but many available benchmarks emphasize disease classification, repo

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

SPARE: Structural Parameter-Free Affinity Regularization for Flow Matching

DGX agent

arXiv:2608.01990v1 Announce Type: new Abstract: Denoising diffusion transformers achieve strong generation quality but converge slowly during training. Regularizing their internal representations has

model-releasesarxiv-cs-cv
4 Aug 2026
Local Ai

SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance

DGX agent

arXiv:2608.00502v1 Announce Type: new Abstract: Affordance grounding aims to localize the functional region for interaction, such as the handle to grasp or the button to press, rather than the whole o

local-aiarxiv-cs-cv
4 Aug 2026
Model Releases

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models

DGX agent

arXiv:2608.01709v1 Announce Type: new Abstract: Vision-language models (VLMs) achieve strong semantic understanding but remain unreliable in metric spatial reasoning, particularly when queries require

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

SPECTRA: Band-Routed Embedding and Stage-Wise LoRA for Cross-Sensor Fine-Tuning of Geospatial Foundation Models

DGX agent

arXiv:2608.01751v1 Announce Type: new Abstract: Geospatial foundation models (GeoFMs), pretrained on large-scale geospatial data such as Earth observation (EO), climate, and weather data, have shown p

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection

DGX agent

arXiv:2608.01334v1 Announce Type: new Abstract: AI-generated video (AIGV) detection aims to distinguish real videos from AI-generated ones. In practice, detectors trained on existing data often fail t

model-releasesarxiv-cs-cv
4 Aug 2026
Research

SpikeRestormer: Towards Energy-Efficient All-in-One Image Restoration via Unified Event Reasoning

DGX agent

arXiv:2608.02290v1 Announce Type: new Abstract: ANN-based All-in-One image restoration (AiOIR) unifies diverse degradation handling but incurs high computational costs, limiting its real-time deployme

researcharxiv-cs-cv
4 Aug 2026
Safety

SPIRIT: Spatio-temporal Pairwise Relational Modeling of Instrument-Tissue Interactions for Surgical Action Triplet Recognition

DGX agent

arXiv:2608.02188v1 Announce Type: new Abstract: Fine-grained understanding of surgical activity is essential for context-aware assistance in the operating room, including safety monitoring, adverse ev

safetyarxiv-cs-cv
4 Aug 2026
← Previous
1…2425262728…261
Next →