Local Ai
Microsoft's MIT licensed VibeVoice speech-to-text model (think Whisper with speaker diarization) is really good - my notes on running the 5.…
Microsoft's MIT licensed VibeVoice speech-to-text model (think Whisper with speaker diarization) is really good - my notes on running the 5.71GB 4bit MLX conversion on an M5 MacBook, using about 60GB
Microsoft's MIT licensed VibeVoice speech-to-text model (think Whisper with speaker diarization) is really good - my notes on running the 5.71GB 4bit MLX conversion on an M5 MacBook, using about 60GB of RAM at peak and transcribing 1hr of audio in ~9 mins https://simonwillison.net/2026/Apr/27/vibevoice/
Related
- I built AmicoScript: A local-first Whisper UI with Speaker Diarization and Ollama integration for summaries.
- Got Microsoft's VibeVoice TTS running locally with OpenAI-compatible API — here's my setup
- microsoft/VibeVoice
- OmniVoice: multilingual local TTS with 600+ languages, voice cloning, and an OpenAI-compatible server
Source: Simon Willison (X) | 2026-04-27