Skip to content
View original post on X: Latent.Space· 28/100AI score28/100

Jev creator argues public AI benchmarks miss what matters

AISummary

CompleteSkeptic, CEO of TypeSafe, argues that public benchmarks such as "Jevbench" miss the point of Jev and can be gamed easily. He says picking the right task matters more than any benchmark score.

Post on XView on X
@latentspacepod

STOP making "Jevbench"es, stop asking for public benchmarks, they completely miss the point of Jev and you won't believe how easy it is to game every benchmark you hold dear

This is @CompleteSkeptic's bitterest lesson of all: picking the right task beats everything

Latent.Space@latentspacepod
Jev and the System One Model: RLCD, intelligence/$, reliable AI, & the end of chat-first AI https://www.latent.space/p/jev @typesafeai CEO @CompleteSkeptic explains why AI can solve extraordinarily hard problems yet still fail to automate basic work, why Jev is built for reliable decisions inside software instead of chat, why TypeSafe rejects public benchmarks and refusals at the API layer, why data and the right task matter more than brute-force compute, how System One Models could reshape coding agents and software, and why even with $1 billion he wouldn’t pre-train a model from scratch.
View quoted post on X

Source: Latent.Space · x.comPublished · added here