Bullish on this trend of making post-training more accessible.
A new post-training era is upon us.
If you work on agentic RL, long-context tasks (a big focus today) are expensive, inefficient, and don't scale well.
I've been diving into RL envs and evals for long-context tasks, and I can see this being useful.
In agent RL, rollouts use most of the tokens. Every turn re-reads the whole growing context, including tool outputs, files, and earlier turns.
Tinker just cut the price of those tokens. Long-context prefill and sampling now cost the same as short context.
This means that evaluating your trained models on long inputs also gets cheaper. Huge win here.
I believe RL will keep unlocking specialized models that slash the cost of critical agent operations. Cheaper long rollouts make them more practical to build.
Own your intelligence stack!
