Magic's new pretraining recipe matches DeepSeek V4 Pro with 50x less compute
AIMagic says its new pretraining recipe matches DeepSeek V4 Pro's pretraining while using 50x less compute, roughly half the FLOPs used for GPT-3, or about $0.5M on GB200. The post, which congratulates the team, suggests that during recursive self-improvement, automated AI researchers may be less bottlenecked by compute than expected.