Model Releases
Taalas buried the lede for the amazing demo of their first tape out. 15k tokens per second with a llama 8b model, ability to scale that up e…
Taalas buried the lede for the amazing demo of their first tape out. 15k tokens per second with a llama 8b model, ability to scale that up etched onto silicon As models satisfice etching makes sense,
Taalas buried the lede for the amazing demo of their first tape out. 15k tokens per second with a llama 8b model, ability to scale that up etched onto silicon As models satisfice etching makes sense, particularly ternary.. Try it out https://chatjimmy.ai Bullish for $AMD We are pleased to share that Taalas has agreed to join AMD. We built Taalas to rethink AI inference from the ground up: hardware designed around the model, rather than the other way around. The result is the world's fastest and most cost-effective inference silicon. Joining AMD g…
Related
- Ran DS V4-Flash-0731 Locally on 3xMI50 32GB @ ~15 t/s TG
- Cross-Family Speculative Decoding for Polish Language Models on Apple
Silicon: An Empirical Evaluation of Bielik11B with UAG-Extended MLX-LM
Source: Emad Mostaque (X) | 2026-08-06