On important coding and agentic benchmarks such as DeepSWE, AutomationBench, and AA-Briefcase, ML4 matches the performance of the best open-weight models.
Original titleOn important coding and agentic benchmarks such as DeepSWE, AutomationBench, and AA-Briefcase, ML4 matches the performance of the best op...
AISummary
It is SOTA on finance and legal workflows, as well as on complex multimodal grounding benchmarks. It can navigate complex terminal workflows, work across spreadsheets, slides, and PDFs, and reason over scientific and multimodal tasks. 4/n
Source: Guillaume Lample · x.com