Thinking Machines' new model matches Inkling on multimodal evals
Original titleLike Inkling, it's natively multimodal. It’s encoder-free, with audio and images processed jointly with text. It nearly matches Inkling a...
AISummary
Thinking Machines' new model is natively multimodal and encoder-free, processing audio and images jointly with text. It nearly matches Inkling across multimodal evaluations and can run Python to crop, zoom, and inspect images while reasoning over documents and charts.
Source: Thinking Machines · x.comPublished · added here