Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 7

Oct 7Wed
  1. Google ResearchAI score23

    Google Research invites COLM visitors to ContinuousBench walkthrough on DP synthetic data

    AIGoogle Research is hosting a walkthrough at its COLM booth #107 today at 5:00 PM of ContinuousBench, a standardized benchmark for measuring knowledge transfer in differentially private synthetic data. The session, led by Alex Bie, asks whether DP synthetic data preserve actual information or only style. A paper is linked on arXiv.

    Image from @GoogleResearch's post
  2. IThome · AIAI score72

    Anthropic releases Claude Haiku 5.5, cutting run costs about 75% from Haiku 4.5

    AIAnthropic released Claude Haiku 5.5, which it calls the fastest, cheapest, and most capable Haiku model so far. On average it costs about 75% less to run than Haiku 4.5, with input at $0.10 and output at $0.50 per million tokens for requests up to 100,000 tokens. Anthropic also cut Sonnet 5.5's cache read price from $0.20 to $0.10 per million tokens, which it says lowers run costs by about 20% on many agent tasks.

  3. TypeSafe AIAI score25

    Jev-killer OpenAI Decisions API benchmarked against Jev for HiringCafe

    AIThe main post is a short reply saying reports of a company's death have been greatly exaggerated, with no details about products or figures. The background post from @h_nilforoshan reports that OpenAI's Decisions API, billed as a "Jev-killer," was benchmarked against Jev for HiringCafe, which serves 2.5 million users. On the task of scoring job-description relevance from 1 to 10, the author reports OpenAI costing 2x more and performing 5-10% worse.

  4. Leandro von WerraAI score36

    Snorkel expands Open Benchmarks Grants to $30M for AI evaluation

    AISnorkel AI is expanding its Open Benchmarks Grants tenfold to a $30M commitment to fund more diverse, robust, and continuously updated open AI benchmarks. The program adds an Open Benchmarks Red Team to test and strengthen those benchmarks, plus a Snorkel Research Fellowship for independent researchers developing new evaluation methods. The source says OBG-funded benchmarks have appeared on model cards from every major frontier lab.

  5. Simon WillisonAI score62

    Anthropic releases Claude Haiku 5.5, priced like GPT-6 Luna up to 100,000 tokens

    AIAnthropic has released Claude Haiku 5.5, priced at $0.10 input and $0.50 output per million tokens up to 100,000 tokens, matching GPT-6 Luna. Beyond 100,000 tokens the price rises to $0.50 and $2.50, and the author found the new tokenizer uses about 1.25x as many tokens as Haiku 4.5 on the same long prompt. The model cannot disable reasoning and defaults to medium effort.

  6. MarkTechPostAI score67

    Anthropic releases Claude Haiku 5.5, a small model with 1M context

    AIAnthropic has released Claude Haiku 5.5, its cheapest and fastest small model, priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens. It keeps a 1M token context window, up to 128K output tokens, and is generally available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Anthropic reports 72.4% on OSWorld 2.1 (offline subset) versus 15.7% for Haiku 4.5, and the article notes that non-default temperature, top_p or top_k values return a 400 error.

  7. ChatGPTAI score85

    GPT-6 with Intelligent UI rolls out globally in ChatGPT, Free and Go tiers next day

    AIGPT-6 with Intelligent UI begins rolling out globally in the ChatGPT Chat tab for Plus, Pro, Business, and Enterprise users today. The rollout expands to Free and Go tiers starting tomorrow. Plus, Pro, Business, and Enterprise get GPT-6 Sol, while Free and Go get GPT-6 Luna.

    This story has a top pick“OpenAI rolls out GPT-6 and Intelligent UI to all ChatGPT users”

  8. KhazixAI score60

    Claude Max subscribers get monthly API credits usable across Claude models

    AISubscribers to Claude's Max plan can claim monthly API credits: $100 for the $100 tier and $200 for the $200 tier. The credits work for any Claude model and can be used in the user's own apps and other agents. The author argues that bundling monthly API credits alongside a broad model lineup will make it hard for other model companies to compete.

  9. Wired · AIAI score60

    Researchers Test GPT-6 Astra Driving a Corolla to In-N-Out

    AIThree Axiom engineers had OpenAI's GPT-6 Astra drive a 2024 Toyota Corolla to an In-N-Out drive-thru through a server linked to cameras and power steering, with a safety driver ready to brake. They also built a parking-lot benchmark, DrivingBench, where Astra completed the course slowly, Claude Fable 5.1 finished 45 percent, and Grok finished 11 percent.

  10. Design ArenaAI score44

    Claude Haiku 5.5 is now available on Design Arena

    AIDesign Arena has added Anthropic's Claude Haiku 5.5, which the post describes as the company's fastest and most capable small model yet. It is aimed at high-volume, cost-sensitive work such as coding, classification, summarization, database queries, and speed-sensitive workflows like customer support and browser use. Per the background post from Claude, it costs around 75% less to run than Claude Haiku 4.5.

    Image from @DesignArena's post
  11. MarkTechPostAI score60

    Liquid AI releases open-weight d1-3B and d1-omni-600M decision models

    AILiquid AI released two open-weight multimodal decision models, d1-3B and d1-omni-600M, which return probability answers in one forward pass with zero output tokens. d1-3B scores 48.57 on Decision Index v0.2.1 and answers one question in 8 ms on an RTX 4090, while the models are licensed free for commercial use below $10 million in annual revenue.

  12. 🚨 AI News | TestingCatalogAI score62

    Anthropic releases Claude Haiku 5.5, its fastest and cheapest model

    AIAnthropic has released Claude Haiku 5.5, which the author describes as its fastest and cheapest model to date. The source says it costs about 75% less to run than Claude Haiku 4.5 and is the first Haiku model with an adjustable effort setting. The attached benchmark table reports Haiku 5.5 scores on tasks including computer use (OSWorld 2.1 offline subset, 72.4%) and Terminal-Bench 4.0 (39.2%), compared with Haiku 4.5 and other models.

    Image from @testingcatalog's post
  13. Design ArenaAI score44

    Claude Opus 5.5 tops four Design Arena leaderboards after two weeks

    AIAnthropic's Claude Opus 5.5 has taken first place on four Design Arena leaderboards: Overall Frontend, Data Visualization, 3D Design, and React Native. It also ranks in the top three on the Game Dev and UI Components leaderboards, about two weeks after its release. Design Arena says developers, designers, and casual users have embraced the model.

    Image from @DesignArena's post
  14. Marcus on AIAI score62

    Marcus Says OpenAI's Math Result Lacks Details Needed to Judge Its Generality

    AIGary Marcus argues that OpenAI's math announcement omits the procedure, the model architecture, and the failure rate, so its generalizability cannot be assessed. He says it could be a step toward AGI or a Lean-based verification trick in a verifiable domain, and the initial report cannot distinguish the two. The post includes a quoted Terence Tao post that shares a satirical press release about a fictional film-endings repository.

  15. Liquid AIAI score36

    Liquid AI releases d1-omni-600M, a 600M multimodal model for on-device tasks.

    AILiquid AI has released d1-omni-600M, an experimental 600M-parameter model that handles text plus image or audio input. It combines LFM2.5-Encoder-350M with vision and audio encoders and leads the company's text benchmark comparison on toxicity detection and paraphrase identification. The post suggests uses such as voice-command routing, on-device moderation, and intent classification.

    Image from @liquidai's post
  16. Hugging Face BlogAI score49

    Liquid AI Releases Open d1-3B and d1-omni-600M Edge Decision Models

    AILiquid AI released two open-weight decision models, d1-3B and d1-omni-600M (experimental), built on its Liquid Foundation Models and available on Hugging Face. d1-3B scores 48.57 on the Decision Index 0.2.1, the highest among decision models under 10B parameters, and answers a question in 16 ms on an NVIDIA Jetson AGX Thor and under 50 ms on a Jetson Orin Nano. The models support text and images (d1-3B) or text with image or audio (d1-omni-600M).