Tinker cuts long-context prices up to 70% after efficiency gains
Overview
Tinker, an API for training and fine-tuning models, announced price cuts of up to 70% after making efficiency improvements, and says it is passing the savings on to customers.
The company says savings grow as customers buy more. The cuts apply to long-context work, and the company says that work is driven by users scaling up long-context reinforcement learning.
A follow-up report from Elvis Saravia says the cut covers long-context prefill and sampling, which now cost the same as short context. Saravia says this lowers the cost of agentic RL rollouts, which spend most of their tokens re-reading growing context, and of evaluating trained models on long inputs.
Tinker also made GLM-5.3-Flash and DeepSeek-v4.1-Flash available for cost-efficient long-context work, according to its post on X.
Written by AI from the articles below · updated Oct 9, 7:40 PM ET
Check the sources:
Article timeline
The articles in this story. Times are ET.
elvis@omarsar0XTinker cuts long-context token prices, making agent RL rollouts cheaperAITinker has cut prices up to 70% on long-context prefill and sampling, which now cost the same as short context. The cut lowers the cost of agentic RL rollouts, which spend most of their tokens re-reading growing context, and of evaluating trained models on long inputs. Tinker also added GLM-5.3-Flash and DeepSeek-v4.1-Flash for cost-efficient long-context work.
Soumith Chintala@soumithchintalaXTinker cuts prices up to 70% as efficiency improvesAITinker, an API for training and fine-tuning models, is cutting prices by up to 70% after engineering efficiency gains. The company says the savings are passed on to customers, and that buying more produces greater savings. GLM-5.3-Flash and DeepSeek-v4.1-Flash are also now available on Tinker for long-context work.
Tinker@tinkerapiOfficialTinker adds GLM-5.3-Flash and DeepSeek-v4.1-Flash modelsAITinker adds GLM-5.3-Flash and DeepSeek-v4.1-Flash, both of which natively accept image inputs and use efficient attention architecture. GLM-5.3-Flash costs 4-5 times less on Tinker than GLM-5.3. Long-context options for Qwen3.5-4B and Qwen3.6-35B-A3B are also live.
Tinker@tinkerapiOfficialTinker cuts prices up to 70% and adds two Flash models for long-context RLAITinker says it made efficiency improvements to support scaling long-context reinforcement learning and is passing them on through price cuts of up to 70%. GLM-5.3-Flash and DeepSeek-v4.1-Flash are now live on the platform for cost-efficient long-context work.

Heat trend
Not enough continuous observations to show a trend yet.