Model Releases
Big shoutout to @perplexity_ai for properly benchmarking document understanding in their Portable Computer release 🔥 They took a subset of …
Big shoutout to @perplexity_ai for properly benchmarking document understanding in their Portable Computer release 🔥 They took a subset of our ParseBench benchmark (https://www.parsebench.ai/) and mea
Big shoutout to @perplexity_ai for properly benchmarking document understanding in their Portable Computer release 🔥 They took a subset of our ParseBench benchmark (https://www.parsebench.ai/) and measured across tables, charts, layout, text content, and formatting. Document understanding is the first step towards most knowledge work, and its capabilities are a function of both the model and harness. As companies build new models and agents that push the frontiers of knowledge work, we hope to see document understanding be a core part of any benchmarking effort. New research: Portable Computer is a local-first agent for private and cost-effective work. With an on-device 27B model, our harness scores 82.6% on real knowledge work, beating open-source harnesses Pi and Hermes. Our post-trained PPLX 27B reaches 85.4%.
Related
- We benchmarked GPT-5.5 on document understanding 📄📊 We ran it through ParseBench, our comprehensive OCR benchmark over enterprise document…
- Thanks for the shoutout re: ParseBench! 📑 🙂 Opus 4.7 is definitely a step up from Opus 4.6 on document understanding capabilities. Great t…
Source: Jerry Liu (X) | 2026-08-26