AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlog
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cv”

GridTimelineEvolution
12,414 results
Research

Signpost Watermarking: Joint Optimization for Visual Watermark Coexistence

DGX agent

arXiv:2608.10091v1 Announce Type: new Abstract: We present a method for training imperceptible visual watermarks to coexist with other such watermarks. Recent work has shown that independently trained

researcharxiv-cs-cv
12 Aug 2026
Research
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

SparSTAR: Sparse Attention for SpaceTime AutoRegressive Video Synthesis

DGX agent

arXiv:2608.10519v1 Announce Type: new Abstract: InfinityStar extends visual autoregressive generation to video through a sequence of image and clip pyramids. Its changing scale and cross-clip context,

researcharxiv-cs-cv
12 Aug 2026
Research

SpecF2M: A Spectral-Aware Multi-task Network Estimating Axial Length and Refractive Error from Pediatric Fundus Photographs

DGX agent

arXiv:2608.09994v1 Announce Type: cross Abstract: Spherical Equivalent Refraction (SER) and Axial Length (AL) are core indicators for pediatric myopia screening, yet their measurements require dedicat

researcharxiv-cs-cv
12 Aug 2026
Model Releases

Static in Frames, Dynamic in Events: Rethinking Features in Event Cameras as Motion Cues

DGX agent

arXiv:2608.11075v1 Announce Type: new Abstract: Event cameras capture intensity changes asynchronously with high temporal resolution, requiring novel preprocessing methods for downstream tasks. Unlike

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

Stream Forcing: Constructing Unified Training Trajectory for Robust Streaming Video Generation

DGX agent

arXiv:2608.10439v1 Announce Type: new Abstract: Streaming video generation holds strong potential for world modeling, where future frames must be inferred online sequentially to form a continuous vide

model-releasesarxiv-cs-cv
12 Aug 2026
Local Ai

Structural Guidance for Unified Joint Demosaicing and Denoising

DGX agent

arXiv:2608.09995v1 Announce Type: cross Abstract: Joint demosaicing and denoising is a fundamental step in camera image signal processing, yet remains challenging because different Bayer-like color fi

local-aiarxiv-cs-cv
12 Aug 2026
Agents

SuperQuadricOcc: Real-Time Self-Supervised Semantic Occupancy Estimation with Superquadric Volume Rendering

DGX agent

arXiv:2511.17361v5 Announce Type: replace Abstract: Self-supervision for semantic occupancy estimation is appealing as it removes the labour-intensive manual annotation, thus allowing one to scale to

agentsarxiv-cs-cv
12 Aug 2026
Model Releases

TEASR: Training-Efficient Any-Step Diffusion Transformer for Real-World Image Super-Resolution

DGX agent

arXiv:2606.16188v2 Announce Type: replace Abstract: Diffusion models excel in Real-World Image Super-Resolution (Real-ISR) due to their powerful generative priors but suffer from slow iterative sampli

model-releasesarxiv-cs-cv
12 Aug 2026
Safety

The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset

DGX agent

arXiv:2608.10839v1 Announce Type: new Abstract: This preprint presents the results of the fourth GENEA Challenge, a large-scale human evaluation of five speech-driven gesture-generation systems traine

safetyarxiv-cs-cv
12 Aug 2026
Local Ai

ThinkAfford: Affordance-Centric Reasoning for Fine-Grained 3D Grounding in Cluttered Scenes

DGX agent

arXiv:2608.10981v1 Announce Type: new Abstract: Task-driven 3D affordance grounding aims to localize the functional region in a cluttered 3D scene that enables an action specified by a natural-languag

local-aiarxiv-cs-cv
12 Aug 2026
Safety

Towards Color-Faithful Low-Light Image Enhancement via Adaptive Color Debiasing and Saturation Rectification

DGX agent

arXiv:2608.10512v1 Announce Type: new Abstract: Low-light imaging often introduces color bias caused by the low signal-to-noise ratio and the image formation process. Although recent low-light image e

safetyarxiv-cs-cv
12 Aug 2026
Research

Towards Geometry-Grounded Dense Semantic Matching with VGGT Priors

DGX agent

arXiv:2509.21263v2 Announce Type: replace Abstract: Semantic matching aims to establish pixel-level correspondences between instances of the same category and represents a fundamental task in computer

researcharxiv-cs-cv
12 Aug 2026
Safety

TRACE-GS: On-Policy Trajectory Distillation with Privileged Geometric Conditioning for Sparse-View 3DGS Restoration

DGX agent

arXiv:2608.10286v1 Announce Type: new Abstract: We present TRACE-GS, an on-policy trajectory distillation framework that leverages privileged geometric conditioning at training time, thereby adapting

safetyarxiv-cs-cv
12 Aug 2026
Tutorials

Transformer Geometry Observatory TGO-IV: Developmental Topology Observatory

DGX agent

arXiv:2608.09997v1 Announce Type: cross Abstract: Transformers have had a profound impact on the world of language processing and computer vision. As efforts to answer the million-dollar question of `

tutorialsarxiv-cs-cv
12 Aug 2026
Safety

UniMod: Enhancing Multi-Modal Medical Diagnosis through Cross-Modality and Within-Modality Alignment

DGX agent

arXiv:2608.10316v1 Announce Type: new Abstract: Multi-modal learning combining medical images and clinical text is promising for disease diagnosis. However, standard multi-modal training leads to shor

safetyarxiv-cs-cv
12 Aug 2026
Local Ai

UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations

DGX agent

arXiv:2608.10835v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) achieve impressive visual reasoning and dialogue capabilities, yet frequently hallucinate content unsupported by th

local-aiarxiv-cs-cv
12 Aug 2026
Local Ai

VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics

DGX agent

arXiv:2608.11201v1 Announce Type: new Abstract: Recent advances in video generation models have significantly improved the realism of synthetic videos, blurring the boundary between generated and auth

local-aiarxiv-cs-cv
12 Aug 2026
Safety

VIDS-Seg: Towards Reliable Uncertainty Quantification in Pediatric Cardiac Ultrasound Segmentation

DGX agent

arXiv:2608.10903v1 Announce Type: new Abstract: Reliable clinical deployment of machine learning requires models that know when they are likely to fail, particularly for subgroups underrepresented in

safetyarxiv-cs-cv
12 Aug 2026
Model Releases

Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction

DGX agent

arXiv:2608.10682v1 Announce Type: new Abstract: Single-frame surround-view reconstruction faces severe geometric instability and rendering artifacts due to minimal inter-camera overlap. While existing

model-releasesarxiv-cs-cv
12 Aug 2026
Safety

Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning

DGX agent

arXiv:2608.11013v1 Announce Type: new Abstract: Text-only training is a popular paradigm in zero-shot video captioning, where the video distribution is not available to the model during training, lead

safetyarxiv-cs-cv
12 Aug 2026
Research

WaveInst: A Frequency-Domain Enhanced Network for Fine-Grained Thin Tree Trunk Extraction in Forest Scenes

DGX agent

arXiv:2505.01656v2 Announce Type: replace Abstract: Analyzing tree morphology, particularly trunk and branch extraction, is valuable for genetic breeding and forestry management. Existing image-based

researcharxiv-cs-cv
12 Aug 2026
Safety

What DINO saw: ALiBi positional encoding reduces positional bias in Vision Transformers

DGX agent

arXiv:2603.16840v2 Announce Type: replace Abstract: Vision transformers (ViTs) - especially feature foundation models like DINOv2 - learn rich representations useful for many downstream tasks. However

safetyarxiv-cs-cv
12 Aug 2026
Applications

When Repository Labels Are Not Image-Level Truth: A Supervision Auditing Framework for Chest Radiograph AI

DGX agent

arXiv:2608.10084v1 Announce Type: cross Abstract: Public chest X-ray repositories are widely used to train medical AI systems, yet their labels are typically extracted from radiology reports rather th

applicationsarxiv-cs-cv
12 Aug 2026
Local Ai

When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs

DGX agent

arXiv:2608.10489v1 Announce Type: new Abstract: Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token p

local-aiarxiv-cs-cv
12 Aug 2026
Model Releases

When Visual Signals Mislead: A Mechanistic Study of Attribute Hallucination in Vision-Language Models

DGX agent

arXiv:2608.11024v1 Announce Type: new Abstract: Attribute hallucination---where vision-language models (VLMs) correctly identify an object but mischaracterize its properties---is prevalent yet mechani

model-releasesarxiv-cs-cv
12 Aug 2026
Research

Where To Look? : Causal Tracing of Vision Encoders in VLM

DGX agent

arXiv:2608.10758v1 Announce Type: new Abstract: Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information actua

researcharxiv-cs-cv
12 Aug 2026
Research

ZeroPur: Succinct Training-Free Adversarial Purification

DGX agent

arXiv:2406.03143v4 Announce Type: replace Abstract: Adversarial purification is a kind of defense technique that can defend against various unseen adversarial attacks without modifying the victim clas

researcharxiv-cs-cv
12 Aug 2026
Research

A Content-Aware Pure Permutation with Intrinsic Avalanche Effect: Breaking the Diffusion-Permutation Dichotomy

DGX agent

arXiv:2608.09452v1 Announce Type: new Abstract: Pixel permutation is a fundamental tool in image processing, image encryption, and data hiding (including watermarking and steganography) that rearrange

researcharxiv-cs-cv
11 Aug 2026
Research

A continually expandable foundation model for brain MRI

DGX agent

arXiv:2608.08319v1 Announce Type: new Abstract: Brain magnetic resonance imaging (MRI) is central to neuroscience and clinical assessment, but models are commonly developed for individual diseases, po

researcharxiv-cs-cv
11 Aug 2026
Research

A Controlled Study of Feature-Based Knowledge Distillation Across Student Designs

DGX agent

arXiv:2608.08294v1 Announce Type: cross Abstract: Knowledge distillation trains a smaller student to match the outputs of a larger teacher. Feature-based methods also align intermediate representation

researcharxiv-cs-cv
11 Aug 2026
Safety

A Dynamic-Semantics Framework for Grounding Human Referring Expressions in Visual Perceptual Data

DGX agent

arXiv:2608.08663v1 Announce Type: cross Abstract: Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment.

safetyarxiv-cs-cv
11 Aug 2026
Agents

A Height-Constrained 2-Point Minimal Solver for Pose Estimation from Active LED Markers with Event Cameras

DGX agent

arXiv:2608.09520v1 Announce Type: new Abstract: In many autonomous applications requiring real-time localization, active marker-based systems are preferred due to their low latency and ease of deploym

agentsarxiv-cs-cv
11 Aug 2026
Research

A Hybrid Neural-Microfacet BRDF Model for Real-Time Rendering

DGX agent

arXiv:2608.09604v1 Announce Type: cross Abstract: Over the past decade, microfacet-based BRDF models have formed the foundation of real-time rendering pipelines. Despite their widespread use, they oft

researcharxiv-cs-cv
11 Aug 2026
Research

A Review of Vision-Based Vehicle Detection for UAV-Based Traffic Monitoring: Experimental Insights and Future Directions

DGX agent

arXiv:2608.07571v1 Announce Type: new Abstract: In Intelligent Transportation System (ITS), unmanned aerial vehicle (UAV)-based surveillance offers an innovative solution to traffic surveillance with

researcharxiv-cs-cv
11 Aug 2026
Safety

Action- and Language-Conditioned Video Assessment for Embodied Control

DGX agent

arXiv:2608.08273v1 Announce Type: cross Abstract: Vision-based embodied agents executing multi-step natural language instructions require feedback mechanisms that assess task progress over complete tr

safetyarxiv-cs-cv
11 Aug 2026
Model Releases

AdaDINO: Pair-Aware In-Backbone Adaptation of Frozen DINO for Efficient Remote Sensing Change Detection

DGX agent

arXiv:2608.07982v1 Announce Type: new Abstract: Vision foundation models (VFMs) such as DINO are pretrained for single-image representation, whereas remote sensing change detection requires reasoning

model-releasesarxiv-cs-cv
11 Aug 2026
Research

AdapterMoE: A Two-Stage Hard-Routing Mixture-of-Experts Architecture for Multi-Crop Disease Recognition with Calibrated Rejection and Incremental Learning

DGX agent

arXiv:2608.08808v1 Announce Type: new Abstract: Timely crop-disease identification is critical to food security. Multi-crop recognition suits Mixture-of-Experts (MoE), but conventional soft-routing Mo

researcharxiv-cs-cv
11 Aug 2026
Model Releases

ADOPD: Reference-Privileged On-Policy Distillation for MLLM-Based Industrial Anomaly Detection

DGX agent

arXiv:2608.09789v1 Announce Type: new Abstract: Industrial anomaly detection (IAD) requires identifying fine-grained deviations from normal visual patterns. Multimodal large language models (MLLMs) ca

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

Advantage-Guided Gate: Reshaping Open-Ended Reasoning for Vision-Based Spatial Intelligence

DGX agent

arXiv:2608.07987v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated significant potential in complex spatial scene understanding and reasoning tasks. However, th

model-releasesarxiv-cs-cv
11 Aug 2026
Research

Adversarially Robust Few-Shot Anomaly Detection with Vision Foundation Models

DGX agent

arXiv:2510.13643v2 Announce Type: replace Abstract: Vision foundation models such as DINOv2 enable strong few-shot anomaly detection (FSAD) through simple non-parametric k-nearest-neighbor (k-NN) scor

researcharxiv-cs-cv
11 Aug 2026
Model Releases

AeroReformer2: Spoken-Query Referring Segmentation for Aerial Images

DGX agent

arXiv:2608.08874v1 Announce Type: new Abstract: Spoken language offers a natural, hands-free interface for specifying an arbitrary target in dense remote-sensing imagery, yet existing referring remote

model-releasesarxiv-cs-cv
11 Aug 2026
Agents

Agentic AI-powered flexible fiber-bundle endoscopy for high-resolution NIR-II fluorescence imaging in vivo

DGX agent

arXiv:2608.08402v1 Announce Type: new Abstract: Fiber-bundle endoscopy offers a compact and flexible route for clinical fluorescence imaging through natural human orifices, but since its first report

agentsarxiv-cs-cv
11 Aug 2026
Safety

Agentic Visual Reasoning in Whole-Slide Pathology Images via Active Perception

DGX agent

arXiv:2608.08648v1 Announce Type: new Abstract: Whole-slide visual reasoning requires identifying sparse diagnostic evidence in gigapixel pathology slides and integrating observations across spatial s

safetyarxiv-cs-cv
11 Aug 2026
Research

Agreement-Based Audio-Visual Segmentation:Champion Report for the MeViS-Audio Track in the 8th LSVOS Challenge

DGX agent

arXiv:2608.09475v1 Announce Type: new Abstract: The MeViS-Audio track asks a system to segment the objects described by a spoken motion expression throughout a video and to return empty masks when the

researcharxiv-cs-cv
11 Aug 2026
Model Releases

AgriField-40K: Adapting Vision Models to Agriculture With Efficient Continual Pretraining

DGX agent

arXiv:2608.07984v1 Announce Type: new Abstract: Field-based agricultural computer vision is important for precision agriculture, yet it largely depends on expensive annotations and costly adaptation o

model-releasesarxiv-cs-cv
11 Aug 2026
Research

Alpha as an Efficiency Signal: Visibility-Routed RGBA Image-to-Video Generation

DGX agent

arXiv:2608.09355v1 Announce Type: new Abstract: RGBA videos combine RGB appearance with an alpha channel, enabling animated assets to be applied across arbitrary backgrounds, which are heavily used in

researcharxiv-cs-cv
11 Aug 2026
Research

AMD:Anatomical Motion Diffusion with Interpretable Motion Decomposition and Fusion

DGX agent

arXiv:2312.12763v3 Announce Type: replace Abstract: Generating realistic human motion sequences from text descriptions is a challenging task that requires capturing the rich expressiveness of both nat

researcharxiv-cs-cv
11 Aug 2026
Tutorials

Anatomically Consistent Cross-Contrast Super-Resolution of Anisotropic Brain T2w MRI

DGX agent

arXiv:2608.08401v1 Announce Type: new Abstract: T2-weighted (T2w) brain MRI provides fluid-sensitive soft-tissue contrast that is important for neuro-oncology and radiotherapy planning. However, T2w s

tutorialsarxiv-cs-cv
11 Aug 2026
← Previous
12345…259
Next →