Model Releases
The future of FDE work seems closely related with all work around evals/posttraining/RL envs. FDEs are effectively responsible for the follo…
The future of FDE work seems closely related with all work around evals/posttraining/RL envs. FDEs are effectively responsible for the following: 1. Define the business problem. 2. Codify the business
The future of FDE work seems closely related with all work around evals/posttraining/RL envs. FDEs are effectively responsible for the following: 1. Define the business problem. 2. Codify the business problem into an eval rubric and environment. 3. Hillclimb the environment and output an agent/agentic workflow that solves the business problem. Right now the process of (3) is quite manual - historically FDEs spend hundreds of hours creating bespoke software/workflows that solve the problem. But assuming intelligence is abundant, they can effectively offload (3) to some automated optimization process. This includes RL on the model layer, and using Claude Code/Codex to optimize the harness/workflow. Then the FDE responsibility shifts from implementing the task to defining the right goals and outcomes. In other words, they have access to /goal, and their job is more around making sure the goal, environment, and evals are correct vs. the tactical implementation details.
Related
- great to see more open evals for an important problem
- We need more evals for document understanding. ParseBench is a really great start. I respect @llama_index ‘s work on this. 📈
- There are a lot of coding and reasoning benchmarks for AI agents, but not a lot for document understanding - which is a prerequisite for all…
- Document OCR benchmarks are still an open problem Existing document OCR benchmarks are either too narrowly focused on a specific type (e.g. …
Source: Jerry Liu (X) | 2026-08-09