Local Ai
Most models run on whatever inference engine they ship with. That's rarely the fastest option. We built our own, AGI-RUN. On a phone it beat…
Most models run on whatever inference engine they ship with. That's rarely the fastest option. We built our own, AGI-RUN. On a phone it beat every model's own native engine, and ran up to 9.9x faster
Most models run on whatever inference engine they ship with. That's rarely the fastest option. We built our own, AGI-RUN. On a phone it beat every model's own native engine, and ran up to 9.9x faster than the best SOTA framework out there. Same models, we just use the chip better. We built our own on-device inference engine, AGI-RUN, and ran it against the full field of SOTA frameworks on a Galaxy S25 Ultra. Up to 9.9x faster. It also beat each model’s native engine by up to 6.9x. We didn't get this by shrinking the models. They're identical. We're just mo…
Source: Div Garg (X) | 2026-07-06