One of the biggest historical bottlenecks for training specialized models has been curating ground-truth. This is typically a very labor intensive process done with human supervision.
Now frontier models/agents are getting better and better at bootstrapping ground-truth from scratch, especially for verifiable tasks. They can do this either in a single-shot setting for simple tasks or simulate longer rollouts with rewards for longer tasks.
This paves the way for anyone to build mostly automated research loops without needing deep ML expertise or deep access to human labor for data curation. That part can be offloaded to the frontier model/agent; the human's main goal is to define the right evals. This is especially useful for anyone who wants to optimize for a given task at the price-performance frontier. Posttraining open-weight models/harness should be made available to everyone, not just labs/specific startups.
Flow for the average user, who is only responsible for the first step:
1. define goal/evals for a given task
2. run frontier agents to gather ground-truth signals for the task. they can bootstrap annotations directly, and reach out to human review where needed
3. train specialized models (and harness) against this task
