Local Ai
Trained a Vit model from scratch for auto tagging
A Reddit user in r/StableDiffusion shared their experience training a Vision Transformer (ViT) model from scratch to automatically tag images, likely for use with image generation or classification ta
A Reddit user in r/StableDiffusion shared their experience training a Vision Transformer (ViT) model from scratch to automatically tag images, likely for use with image generation or classification tasks. The post discusses the implementation and training process for creating a custom tagging model without pre-trained weights. This approach is relevant for developers working with image-to-text or image-labeling applications in the Stable Diffusion ecosystem.
Source: r/StableDiffusion | 2026-05-09