Skip to contentSkip to stories

Updated

#Eval/Benchmark

Showing low-relevance items too. Hide low-relevance items

Oct 7

Oct 7Wed
  1. Wired · AIAI score60

    Researchers Test GPT-6 Astra Driving a Corolla to In-N-Out

    AIThree Axiom engineers had OpenAI's GPT-6 Astra drive a 2024 Toyota Corolla to an In-N-Out drive-thru through a server linked to cameras and power steering, with a safety driver ready to brake. They also built a parking-lot benchmark, DrivingBench, where Astra completed the course slowly, Claude Fable 5.1 finished 45 percent, and Grok finished 11 percent.

  2. Design ArenaAI score44

    Claude Haiku 5.5 is now available on Design Arena

    AIDesign Arena has added Anthropic's Claude Haiku 5.5, which the post describes as the company's fastest and most capable small model yet. It is aimed at high-volume, cost-sensitive work such as coding, classification, summarization, database queries, and speed-sensitive workflows like customer support and browser use. Per the background post from Claude, it costs around 75% less to run than Claude Haiku 4.5.

    Image from @DesignArena's post
  3. MarkTechPostAI score60

    Liquid AI releases open-weight d1-3B and d1-omni-600M decision models

    AILiquid AI released two open-weight multimodal decision models, d1-3B and d1-omni-600M, which return probability answers in one forward pass with zero output tokens. d1-3B scores 48.57 on Decision Index v0.2.1 and answers one question in 8 ms on an RTX 4090, while the models are licensed free for commercial use below $10 million in annual revenue.

  4. 🚨 AI News | TestingCatalogAI score62

    Anthropic releases Claude Haiku 5.5, its fastest and cheapest model

    AIAnthropic has released Claude Haiku 5.5, which the author describes as its fastest and cheapest model to date. The source says it costs about 75% less to run than Claude Haiku 4.5 and is the first Haiku model with an adjustable effort setting. The attached benchmark table reports Haiku 5.5 scores on tasks including computer use (OSWorld 2.1 offline subset, 72.4%) and Terminal-Bench 4.0 (39.2%), compared with Haiku 4.5 and other models.

    Image from @testingcatalog's post
  5. Design ArenaAI score44

    Claude Opus 5.5 tops four Design Arena leaderboards after two weeks

    AIAnthropic's Claude Opus 5.5 has taken first place on four Design Arena leaderboards: Overall Frontend, Data Visualization, 3D Design, and React Native. It also ranks in the top three on the Game Dev and UI Components leaderboards, about two weeks after its release. Design Arena says developers, designers, and casual users have embraced the model.

    Image from @DesignArena's post
  6. Marcus on AIAI score62

    Marcus Says OpenAI's Math Result Lacks Details Needed to Judge Its Generality

    AIGary Marcus argues that OpenAI's math announcement omits the procedure, the model architecture, and the failure rate, so its generalizability cannot be assessed. He says it could be a step toward AGI or a Lean-based verification trick in a verifiable domain, and the initial report cannot distinguish the two. The post includes a quoted Terence Tao post that shares a satirical press release about a fictional film-endings repository.

  7. Liquid AIAI score36

    Liquid AI releases d1-omni-600M, a 600M multimodal model for on-device tasks.

    AILiquid AI has released d1-omni-600M, an experimental 600M-parameter model that handles text plus image or audio input. It combines LFM2.5-Encoder-350M with vision and audio encoders and leads the company's text benchmark comparison on toxicity detection and paraphrase identification. The post suggests uses such as voice-command routing, on-device moderation, and intent classification.

    Image from @liquidai's post
  8. Hugging Face BlogAI score49

    Liquid AI Releases Open d1-3B and d1-omni-600M Edge Decision Models

    AILiquid AI released two open-weight decision models, d1-3B and d1-omni-600M (experimental), built on its Liquid Foundation Models and available on Hugging Face. d1-3B scores 48.57 on the Decision Index 0.2.1, the highest among decision models under 10B parameters, and answers a question in 16 ms on an NVIDIA Jetson AGX Thor and under 50 ms on a Jetson Orin Nano. The models support text and images (d1-3B) or text with image or audio (d1-omni-600M).

  9. LlamaIndex 🦙AI score47

    LlamaIndex launches OpenDocRouter, one API for many document parsing models

    AILlamaIndex announced OpenDocRouter, a single API that routes document parsing requests to any of 10 frontier and open-source models at launch, including Claude Opus 5.5, Gemini 3.8 Flash, GPT-6 Luna, MinerU2.5-Pro, and PaddleOCR-VL-1.6. Users can switch models in one line with the same request and markdown output, and each model is scored on ParseBench for quality and cost. Pricing is per-token, failed pages are not charged, and the service costs $0.86 to $48.82 per 1,000 pages depending on the model.

    Video from @llama_index's post
  10. Lucas Beyer (bl16)AI score36

    Reality Check: a public leaderboard for robot manipulation VLA models

    AILucas Beyer praises Reality Check, a new leaderboard for benchmarking VLA and related robot manipulation models. Half of its tasks are fully open, while the other half are held out to detect benchmaxxing by future model versions. The companion post from Nicolas Keller describes the launch as the first public robot manipulation benchmark, built on 14,400 real-world rollouts across four models.

  11. elvisAI score18

    Viktor, a Slack AI employee, reviews overnight agent eval failures

    AIElvis Saravia describes using Viktor, an AI employee in Slack, to review his nightly agent harness evaluation results. Viktor traces tasks that regressed from passing to failing back to the specific harness change that caused them and suggests reverting it, while the human makes the final decision. The post is a sponsored partnership, offering $100 in free credits with no card required.

    Image from @omarsar0's post