Has anyone actually benchmarked where the 'big-model orchestrator + local-model worker' split breaks down?
I keep seeing the 'use a big model via API as the architect, run local small/mid models as workers' pattern recommended for people with modest local hardware. I've been running it myself (orchestrator