AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
12 Aug 2026

When Visual Signals Mislead: A Mechanistic Study of Attribute Hallucination in Vision-Language Models

Model ReleasesDGX agent

arXiv:2608.11024v1 Announce Type: new Abstract: Attribute hallucination---where vision-language models (VLMs) correctly identify an object but mischaracterize its properties---is prevalent yet mechani

Where To Look? : Causal Tracing of Vision Encoders in VLM

ResearchDGX agent

arXiv:2608.10758v1 Announce Type: new Abstract: Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information actua

ZeroPur: Succinct Training-Free Adversarial Purification

ResearchDGX agent

arXiv:2406.03143v4 Announce Type: replace Abstract: Adversarial purification is a kind of defense technique that can defend against various unseen adversarial attacks without modifying the victim clas


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
11 Aug 2026

A Content-Aware Pure Permutation with Intrinsic Avalanche Effect: Breaking the Diffusion-Permutation Dichotomy

ResearchDGX agent

arXiv:2608.09452v1 Announce Type: new Abstract: Pixel permutation is a fundamental tool in image processing, image encryption, and data hiding (including watermarking and steganography) that rearrange

A continually expandable foundation model for brain MRI

ResearchDGX agent

arXiv:2608.08319v1 Announce Type: new Abstract: Brain magnetic resonance imaging (MRI) is central to neuroscience and clinical assessment, but models are commonly developed for individual diseases, po

A Controlled Study of Feature-Based Knowledge Distillation Across Student Designs

ResearchDGX agent

arXiv:2608.08294v1 Announce Type: cross Abstract: Knowledge distillation trains a smaller student to match the outputs of a larger teacher. Feature-based methods also align intermediate representation

A Dynamic-Semantics Framework for Grounding Human Referring Expressions in Visual Perceptual Data

SafetyDGX agent

arXiv:2608.08663v1 Announce Type: cross Abstract: Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment.

A Height-Constrained 2-Point Minimal Solver for Pose Estimation from Active LED Markers with Event Cameras

AgentsDGX agent

arXiv:2608.09520v1 Announce Type: new Abstract: In many autonomous applications requiring real-time localization, active marker-based systems are preferred due to their low latency and ease of deploym

A Hybrid Neural-Microfacet BRDF Model for Real-Time Rendering

ResearchDGX agent

arXiv:2608.09604v1 Announce Type: cross Abstract: Over the past decade, microfacet-based BRDF models have formed the foundation of real-time rendering pipelines. Despite their widespread use, they oft

A Review of Vision-Based Vehicle Detection for UAV-Based Traffic Monitoring: Experimental Insights and Future Directions

ResearchDGX agent

arXiv:2608.07571v1 Announce Type: new Abstract: In Intelligent Transportation System (ITS), unmanned aerial vehicle (UAV)-based surveillance offers an innovative solution to traffic surveillance with

Action- and Language-Conditioned Video Assessment for Embodied Control

SafetyDGX agent

arXiv:2608.08273v1 Announce Type: cross Abstract: Vision-based embodied agents executing multi-step natural language instructions require feedback mechanisms that assess task progress over complete tr

AdaDINO: Pair-Aware In-Backbone Adaptation of Frozen DINO for Efficient Remote Sensing Change Detection

Model ReleasesDGX agent

arXiv:2608.07982v1 Announce Type: new Abstract: Vision foundation models (VFMs) such as DINO are pretrained for single-image representation, whereas remote sensing change detection requires reasoning

AdapterMoE: A Two-Stage Hard-Routing Mixture-of-Experts Architecture for Multi-Crop Disease Recognition with Calibrated Rejection and Incremental Learning

ResearchDGX agent

arXiv:2608.08808v1 Announce Type: new Abstract: Timely crop-disease identification is critical to food security. Multi-crop recognition suits Mixture-of-Experts (MoE), but conventional soft-routing Mo

ADOPD: Reference-Privileged On-Policy Distillation for MLLM-Based Industrial Anomaly Detection

Model ReleasesDGX agent

arXiv:2608.09789v1 Announce Type: new Abstract: Industrial anomaly detection (IAD) requires identifying fine-grained deviations from normal visual patterns. Multimodal large language models (MLLMs) ca

Advantage-Guided Gate: Reshaping Open-Ended Reasoning for Vision-Based Spatial Intelligence

Model ReleasesDGX agent

arXiv:2608.07987v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated significant potential in complex spatial scene understanding and reasoning tasks. However, th

Adversarially Robust Few-Shot Anomaly Detection with Vision Foundation Models

ResearchDGX agent

arXiv:2510.13643v2 Announce Type: replace Abstract: Vision foundation models such as DINOv2 enable strong few-shot anomaly detection (FSAD) through simple non-parametric k-nearest-neighbor (k-NN) scor

AeroReformer2: Spoken-Query Referring Segmentation for Aerial Images

Model ReleasesDGX agent

arXiv:2608.08874v1 Announce Type: new Abstract: Spoken language offers a natural, hands-free interface for specifying an arbitrary target in dense remote-sensing imagery, yet existing referring remote

Agentic AI-powered flexible fiber-bundle endoscopy for high-resolution NIR-II fluorescence imaging in vivo

AgentsDGX agent

arXiv:2608.08402v1 Announce Type: new Abstract: Fiber-bundle endoscopy offers a compact and flexible route for clinical fluorescence imaging through natural human orifices, but since its first report

Agentic Visual Reasoning in Whole-Slide Pathology Images via Active Perception

SafetyDGX agent

arXiv:2608.08648v1 Announce Type: new Abstract: Whole-slide visual reasoning requires identifying sparse diagnostic evidence in gigapixel pathology slides and integrating observations across spatial s

Agreement-Based Audio-Visual Segmentation:Champion Report for the MeViS-Audio Track in the 8th LSVOS Challenge

ResearchDGX agent

arXiv:2608.09475v1 Announce Type: new Abstract: The MeViS-Audio track asks a system to segment the objects described by a spoken motion expression throughout a video and to return empty masks when the

AgriField-40K: Adapting Vision Models to Agriculture With Efficient Continual Pretraining

Model ReleasesDGX agent

arXiv:2608.07984v1 Announce Type: new Abstract: Field-based agricultural computer vision is important for precision agriculture, yet it largely depends on expensive annotations and costly adaptation o

Alpha as an Efficiency Signal: Visibility-Routed RGBA Image-to-Video Generation

ResearchDGX agent

arXiv:2608.09355v1 Announce Type: new Abstract: RGBA videos combine RGB appearance with an alpha channel, enabling animated assets to be applied across arbitrary backgrounds, which are heavily used in

AMD:Anatomical Motion Diffusion with Interpretable Motion Decomposition and Fusion

ResearchDGX agent

arXiv:2312.12763v3 Announce Type: replace Abstract: Generating realistic human motion sequences from text descriptions is a challenging task that requires capturing the rich expressiveness of both nat

Anatomically Consistent Cross-Contrast Super-Resolution of Anisotropic Brain T2w MRI

TutorialsDGX agent

arXiv:2608.08401v1 Announce Type: new Abstract: T2-weighted (T2w) brain MRI provides fluid-sensitive soft-tissue contrast that is important for neuro-oncology and radiotherapy planning. However, T2w s

AQUA20: A Benchmark Dataset for Underwater Species Classification under Challenging Conditions

Model ReleasesDGX agent

arXiv:2506.17455v3 Announce Type: replace Abstract: Robust visual recognition in underwater environments remains a significant challenge due to complex distortions such as turbidity, low illumination,

ARC: Augmented-Rank Conformalization for Changepoint Localization --- Finite-Sample Validity and Distribution-Robust Efficiency

Model ReleasesDGX agent

arXiv:2608.08424v1 Announce Type: cross Abstract: Conformal changepoint localization turns any score into a confidence set for the changepoint with finite-sample coverage. Coverage is universal; effic

Attention-Guided Perturbation Network for Industrial Anomaly Detection

TutorialsDGX agent

arXiv:2408.07490v4 Announce Type: replace Abstract: In unsupervised image anomaly detection, reconstruction-based methods learn normal patterns for data reconstruction, but often undesirably reconstru

Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions

Local AiDGX agent

arXiv:2608.07550v1 Announce Type: new Abstract: Vision-language models return structured chest-radiograph findings through interfaces exposing no confidence score, so a receiving institution cannot re

BAG: Budget-Aware Gating for Diffusion Caching

Model ReleasesDGX agent

arXiv:2608.09231v1 Announce Type: new Abstract: Diffusion caching is a lightweight strategy that accelerates Diffusion Transformers (DiTs) by reusing intermediate features across denoising steps, but

BAP-MOS: Bandit-Based Adaptive Prompting for Boundary-Sensitive Multi-Organ Segmentation

Model ReleasesDGX agent

arXiv:2608.08191v1 Announce Type: new Abstract: Multi-organ ultrasound segmentation remains challenging when anatomically adjacent structures must be delineated jointly, as localized boundary errors c

Beyond Global Editing: Per-Instance Disentangled Subspaces for Training-Free Hallucination Mitigation in LVLMs

ResearchDGX agent

arXiv:2608.09344v1 Announce Type: new Abstract: Recent advances in large vision-language models (LVLMs) have enabled powerful multimodal reasoning by integrating visual encoders with large language mo

Beyond Hazard Resemblance: Contrastive Event Adjudication for Training-Free Video Anomaly Detection

Local AiDGX agent

arXiv:2608.09908v1 Announce Type: new Abstract: Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries f

Beyond Isotropic Assumptions: Continuity-Constrained Segmentation and GPU Morphometry for Nanoscale GBM Analysis

HardwareDGX agent

arXiv:2608.07575v1 Announce Type: new Abstract: Confocal microscopy of optically cleared and swelled tissue resolves complex biological structures in 3D, but such acquisitions are highly anisotropic:

BMDS-Net:Deployment-aware multi-modal brain tumor segmentation with adaptive fusion,decoder regularization,and Bayesian calibration

ResearchDGX agent

arXiv:2601.17504v2 Announce Type: replace Abstract: Multi-modal MRI enables detailed brain tumor sub-region segmentation, but clinical deployment remains affected by missing sequences,boundary errors,

Bootstrapping Vision-Language Model for Hysteroscopic Surgical Scene Segmentation

Local AiDGX agent

arXiv:2608.09302v1 Announce Type: new Abstract: Hysteroscopic surgical scene segmentation plays a pivotal role in understanding the hysteroscopic intraoperative environment as well as computer-assiste

Bridging Object Detection and Segmentation with Polygon Detection Transformers

AgentsDGX agent

arXiv:2603.09245v2 Announce Type: replace Abstract: Box detection and mask segmentation are two dominant paradigms for foreground representation: boxes are efficient but too coarse for object shapes,

Bright-Channel Retinex Enhancement with a Conditional Overdispered-Noise Analysis

Local AiDGX agent

arXiv:2608.09137v1 Announce Type: new Abstract: I present a training-free low-light enhancement method that combines local bright-channel illumination estimation, Retinex division, and edge-preserving

BRUCE: Benchmarking Robustness Under Corruption Escalation for Scientific Vision-Language Reasoning

ApplicationsDGX agent

arXiv:2608.07742v1 Announce Type: new Abstract: Visual-language models (VLMs) frequently struggle with robustness issues in real-world situations due to low- or varying-quality input images. In this p

C^2A: Coupling Spatial Evidence with Clinical Priors via Co-occurrence Aware Class Attention for Multi-Label Chest X-Ray Classification

ResearchDGX agent

arXiv:2608.09774v1 Announce Type: new Abstract: Thoracic pathologies rarely occur in isolation, yet standard multi-label classifiers rely on shared global descriptors, discarding where findings lie an

CableDex: Cable Length Estimation on Industrial Reels Using a Handheld Device

ResearchDGX agent

arXiv:2608.09392v1 Announce Type: new Abstract: CableDex is a computer vision system that addresses the time-consuming and inaccurate manual measurement of cable length on industrial reels from a sing

Can Webcam Gaze Constrain Mesa-Objectives in Driving Models? An Instrument Precision Analysis

AgentsDGX agent

arXiv:2608.08947v1 Announce Type: new Abstract: Current hazard detection systems in autonomous driving may develop mesa objectives, learned internal goals that achieve high training performance throug

CIFA: Contextual-Intersectional Fairness Auditing for Hidden Subgroup Discovery in Face Analysis

SafetyDGX agent

arXiv:2608.09669v1 Announce Type: new Abstract: Fairness evaluation in computer vision commonly relies on aggregate accuracy and demographic subgroup analysis. However, visual models are also sensitiv

Circuit Fine-Tuning for Compute-Efficient Transformer Adaptation

Model ReleasesDGX agent

arXiv:2608.08336v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) has become the de facto standard for adapting Vision Transformers (ViTs) to downstream tasks. While parameter cou

City Sentinel: A Unified AI-Based Smart Surveillance Framework for Real-Time Multi-Threat Detection Using Deep Learning

SafetyDGX agent

arXiv:2608.08887v1 Announce Type: new Abstract: Rapid urbanization has increased the need for surveillance systems that can monitor multiple public safety risks at the same time. Traditional systems o

CodecArena: Codec Quality Assessment via Visual Reinforcement Learning

Model ReleasesDGX agent

arXiv:2608.09139v1 Announce Type: new Abstract: Video coding is advancing into the low and ultra-low bitrate regime, driven by end-to-end codecs that replace the hand-crafted pipeline with jointly opt

CoInS-Net: A Continuous Position-Aware Network for Joint Medical Image Interpolation and Segmentation

ResearchDGX agent

arXiv:2608.09391v1 Announce Type: new Abstract: Accurate medical image interpolation and anatomical structure segmentation are fundamental for computer-aided diagnosis and treatment planning. Anisotro

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making

Model ReleasesDGX agent

arXiv:2608.07584v1 Announce Type: new Abstract: Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images. Many such t

Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation

ResearchDGX agent

arXiv:2608.09101v1 Announce Type: new Abstract: Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misali

ControlRadio: Prompt-Driven Controllable Diffusion for Cross-Modal Radio Map Generation

ApplicationsDGX agent

arXiv:2608.09357v1 Announce Type: new Abstract: Radio maps describe how wireless signals propagate across space and are essential for wireless communication, sensing, and network planning. However, co

CUPA-T2*: Covariance-Aware Uncertainty Propagation and Alignment for T2* Mapping in Accelerated MRI

SafetyDGX agent

arXiv:2608.08693v1 Announce Type: new Abstract: Quantitative T2* maps have strong potential for biomarker discovery but are limited by long scan times, rendering them impractical in clinical settings.

Damage Classification for 3D Point Cloud Data via 3D Data Analysis and Vision Foundation Model-based 2D Projections

ResearchDGX agent

arXiv:2608.08955v1 Announce Type: new Abstract: Fine-grained damage classification of 3D point cloud data (PCD) remains a persistent challenge, constrained by high computational demands and limited la

Dancing Points: Synthesizing Ballroom Dancing with Three-Point Inputs

ResearchDGX agent

arXiv:2601.02096v2 Announce Type: replace-cross Abstract: Ballroom dancing is a structured yet expressive motion category. Its highly diverse movement and complex interactions between leader and follo

Data collection from highways: a geometric, class-agnostic approach to embedded vehicle counting

ResearchDGX agent

arXiv:2608.07643v1 Announce Type: new Abstract: Traffic data collection is dominated today by deep object detectors followed by tracking-by-detection, a pipeline that presupposes what is often missing

DeCo: Zero-Shot Industrial Anomaly Generation through Decoupling and Recoupling

ResearchDGX agent

arXiv:2608.07904v1 Announce Type: new Abstract: Industrial anomaly inspection is severely hindered by the scarcity of real anomalous data.Zero-shot industrial anomaly generation addresses this by gene

Degradation-Guided Underwater Image Restoration with Task-Oriented Latent Control

TutorialsDGX agent

arXiv:2608.08661v1 Announce Type: new Abstract: Degradation information in underwater images plays a dual role: its spatial and spectral cues can guide adaptive restoration, while degradation-entangle

Degraded Infrared Small Object Detection via Degradation-Adapted Physics-Guided Restoration

ResearchDGX agent

arXiv:2608.09311v1 Announce Type: new Abstract: Infrared small object detection has made significant progress in recent years. However, degradations such as fog and nonuniformity can suppress target-b

Dense Point-to-Mask Optimization with Reinforced Point Selection for Crowd Instance Segmentation

SafetyDGX agent

arXiv:2604.01742v2 Announce Type: replace Abstract: Crowd instance segmentation is a crucial task with a wide range of applications, including surveillance and transportation. Currently, point labels

Describe-to-Score: A text-guided framework for image complexity assessment

Model ReleasesDGX agent

arXiv:2509.16609v2 Announce Type: replace Abstract: Accurately assessing image complexity (IC) is essential for many vision tasks, yet existing approaches rely almost exclusively on visual features an

Diagnosing as Cardiologists Do: ECG Agents with Doctor-Grounded Priors for Clinical Reasoning Across Diseases and Populations

Model ReleasesDGX agent

arXiv:2608.09053v1 Announce Type: cross Abstract: Cardiologists interpret electrocardiograms by localizing waveform components, measuring rhythm and interval patterns, and translating these structured

Did the Grid Erase the Event? EndoClock for Auditing Medical World-Model Pipelines

TutorialsDGX agent

arXiv:2608.09266v1 Announce Type: new Abstract: Medical world models commonly learn from multimodal recordings synchronized onto a fixed-rate grid. This preprocessing resamples each native stream onto

← Previous
12345…207
Next →