Local Qwen Flash inference on consumer GPUs jumps roughly tenfold in a week
Original title (Chinese)
所以,我们的 token 价格能下降10倍吗?
AISummary
The author reports that a dual RTX 5070 Ti setup running Qwen Flash rose from 200 prefill and 10 decode to 2200 prefill and 67 decode, now on a single card, using Strata and a custom PR. The post argues that such consumer-hardware speeds, once limited to top-end machines, could pressure the economics of selling model compute via API.
Source: Orange AI · x.comPublished · added here