Local Ai

Trained a Vit model from scratch for auto tagging

A Reddit user in r/StableDiffusion shared their experience training a Vision Transformer (ViT) model from scratch to automatically tag images, likely for use with image generation or classification ta

DGX agentreddit
local-air-stablediffusion

A Reddit user in r/StableDiffusion shared their experience training a Vision Transformer (ViT) model from scratch to automatically tag images, likely for use with image generation or classification tasks. The post discusses the implementation and training process for creating a custom tagging model without pre-trained weights. This approach is relevant for developers working with image-to-text or image-labeling applications in the Stable Diffusion ecosystem.

Source: r/StableDiffusion | 2026-05-09

Loading related sources…