LFM2.5-350M trained on 28T tokens, beating Chinchilla scaling
Original title350m model trained in 28T tokens - RIP Chinchilla
AISummary
Awni Hannun says a 350M-parameter model trained on 28T tokens defies Chinchilla's compute-optimal scaling guidance. The quoted Liquid AI post credits scaled RL for LFM2.5-350M, reporting instruction following rising from 18.20 to 40.69, data extraction from 11.67 to 32.45, and tool use from 22.95 to 44.11 over LFM2-350M.
Source: Awni Hannun · x.comPublished · added here