Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri11 items
  1. Mike KnoopAI score62

    Tufa Labs hits 88.06% on ARC-AGI-2, clearing the Kaggle bonus threshold

    AIMike Knoop says the 85% Grand Prize bonus threshold has been reached on Kaggle. The ARC Prize 2026 leaderboard lists Tufa Labs first at 88.06%, followed by Rabbithole at 80.56% and Yi-Chia Chen at 77.22%. Knoop says this will be the final year for ARC-AGI-2 on Kaggle and expects an open-source, low-cost, offline reproducible solution and model.

  2. LangChainAI score29

    LangSmith data shows Claude Sonnet 5 and GPT-5.6 Luna gaining ground

    AILangChain reports that over the last month Claude Sonnet 5 rose from #9 to #2 in model adoption, with 51% more organizations using it. GPT-5.6 Luna climbed from #3 to #1 in call footprints, up 65% in calls, while smaller, faster models dominate call footprints overall. Two open-weights models entered the adoption top 10 but do not lead in call volume.

    Image from @LangChain's post
  3. Arena.aiAI score24

    Arena weekly update: Nano Banana 2.1, Mistral Large 4, Claude Haiku 5.5 rankings

    AIArena's weekly update says Nano Banana 2.1 ranked in the top six across three Image Arena modes, with #4 in Multi-Image Edit at 1431 points. Mistral Large 4 placed #43 overall in Agent Arena, 11 spots above Mistral Medium 3.5, and Claude Haiku 5.5 (High) landed #30 in Code Arena WebDev at 1587 points, priced at $0.10/$0.50 per 1M input/output tokens. The post also introduces Arena's Alignment Index and announces a $200M Series B at a $3.1B valuation.

  4. Boris PowerAI score28

    Boris Power calls OpenAI integer multiplication progress "Wow!"

    AIBoris Power, who owns the OpenAI account, posted only the word "Wow!" with no details. Background from a separate post says the integer multiplication problem #109 witness value κ rose to 2⁻¹⁰·⁵⁴⁷ (about 6.6857 × 10⁻⁴), past the 2⁻¹¹ threshold. The author notes gains are now fractional and a major breakthrough is still needed.

  5. SiliconANGLE · AIAI score35

    SailPoint's Navigate event highlights a push to secure AI agent identities in real time

    AISailPoint's Navigate conference in Austin, Texas, featured executives arguing that enterprises must secure AI agent identities at machine speed through just-in-time access and enforcement outside the agent. Mark McClain, SailPoint's founder and chief executive, said real-time decision-making is needed because manual administration cannot keep up. The event also covered the Entro Security acquisition and a partnership with AWS on Amazon Bedrock AgentCore, which grew 15-fold in the first six months of the year.

  6. Don't Worry About the Vase (Zvi Mowshowitz)AI score73

    OpenAI releases 719 AI-generated math manuscripts, splitting the mathematics community

    AIZvi Mowshowitz reports that OpenAI released 722 math manuscripts from an internal frontier model on GitHub, later reduced to 719 after three withdrawals, covering 90 of the top 500 open problems. He says the work came mostly from a single prompt, with an average of three hours of compute per solution. Mathematicians reacted with mixed feelings, and the post highlights concerns about unread papers, cryptography implications, and the role of Lean verification.

Oct 8

Oct 8Thu
  1. IThome · AIAI score62

    Terence Tao questions OpenAI's 719 AI-generated math proofs

    AIOpenAI published 719 AI-generated math proofs covering 372 result families, after withdrawing 3 for a symbol error. Reports say the release falls short of the AGMAI advisory group's standards, since it uses proprietary models, includes reasoning chains for only 10 manuscripts, and leaves about 42% unformalized. Terence Tao argues that rapidly solving famous problems harms the mathematical community's understanding and collaboration.

  2. LeiphoneAI score46

    IROS 2026 Best Paper goes to LT-Mem robot long-term memory study

    AIAt IROS 2026 in Pittsburgh, the Best Paper Award went to Yumin Lee, Hyoseok Ju and Giseop Kim for LT-Mem, a volatility-aware spatio-temporal memory system for lifelong robot scene understanding. The Best Student Paper Award went to Pei-An Hsieh and colleagues for flatness-preserving residual learning enabling real-time tight quadrotor formation flight. Other honors included a humanoid tennis-skills paper and SteadyTray, a humanoid tray-transport study.

  3. TechCrunch · AIAI score62

    Common Sense Media rates ChatGPT for Teens an unacceptable risk over engagement design

    AICommon Sense Media labeled ChatGPT for Teens an "unacceptable risk," finding its design still encourages engagement even in crisis situations. The report says the teen version failed to meet commitments on three of five severe harms, and that break reminders appeared only twice across nearly 2,000 prompts. OpenAI disputed the methodology, saying the testing may have ended before parental controls were fully active, and cited its own data showing teens average under 15 minutes a day.

  4. SiliconANGLE · AIAI score24

    CoreWeave Pitches Open Full-Stack AI Cloud With Forge Development Platform

    AICoreWeave is positioning its AI cloud around an open development loop, connecting training, inference and evaluation through its newly announced CoreWeave Forge platform. Chief marketing officer Jean English said the company wants production learnings to improve models and agents and that the loop should work across different models, frameworks and clouds. She argued that competitive differentiation extends beyond GPUs to partner tooling, infrastructure and APIs.

  5. Artificial AnalysisAI score22

    Artificial Analysis launches Cyber Index Alliance with IBM and NVIDIA

    AIArtificial Analysis has formed the Cyber Index Alliance to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. Current members are Collinear, IBM, NVIDIA, and Vercel, and partners contribute expert input on the Index design and implementation, plus datasets and external research. Organizations interested in joining can contact cyber@artificialanalysis.ai.

    Image from @ArtificialAnlys's post
  6. Arena.aiAI score44

    Arena raises $200M Series B led by Lightspeed, launches Alignment Index

    AIArena has secured a $200M Series B, with Lightspeed doubling down on its investment. The company is also launching the Alignment Index, which measures how closely AI behavior aligns with human values in real-world settings. Arena reports annualized revenue above $100M since its Series A, with millions of people helping evaluate frontier models through real-world use.

  7. TechCrunch · AIAI score46

    Arena raises $200M at $3.1B valuation, nearly doubling in 10 months

    AIArena, the crowdsourced AI model leaderboard that started as a UC Berkeley research project, raised a $200 million Series B at a $3.1 billion valuation, led by Lightspeed Venture Partners and Khosla Ventures. The company said it reached $100 million in annualized run-rate revenue in June, up from $30 million when it raised its $150 million Series A in January at a $1.7 billion post-money valuation.

  8. The DecoderAI score80

    Mathematicians call for OpenAI boycott after AI-generated proofs flood the field

    AIThe Association of Historical Mathematicians (AHM) has called for a boycott of OpenAI after the company released more than 700 AI-generated proof files at once. Fields Medalist Terence Tao, who chairs the group, argues that AI solving open problems autonomously reduces seminars, collaborations, and fertile research directions, and that the field should shift its measure of progress toward explanation and community-building.

    Why it matters: The article links the AHM boycott call to Tao's argument that AI-driven proof volume is changing how mathematicians measure progress and whether solutions remain useful.

  9. TechCrunch · AIAI score65

    OpenAI's math solutions fall short of the field's standards, mathematicians say

    AIOpenAI released hundreds of claimed solutions to hard math problems but did not fully meet guidelines from the Advisory Group on Mathematics and Artificial Intelligence. Only 10 of 719 manuscripts included chain-of-thought releases, and just 42% of proofs were formalized. A Cambridge and King's College paper found discrepancies between a natural language proof and its Lean code for a Navier-Stokes-derived problem.

  10. Arena.aiAI score37

    Arena raises $200M Series B at $3.1B valuation, launches Alignment Index

    AIArena announced a $200 million Series B at a $3.1 billion valuation, alongside a new Alignment Index that measures whether AI agents behave safely, truthfully, and within the bounds of user requests. The company has surpassed $100 million in annualized revenue, facilitated 350 million sessions and 62 million votes, and led by Felicis and PXD from the seed and Series A stages. Arena positions the index as a way to assess trustworthiness as AI systems increasingly take real actions.

Oct 7

Oct 7Wed
  1. TypeSafe AIAI score25

    Jev-killer OpenAI Decisions API benchmarked against Jev for HiringCafe

    AIThe main post is a short reply saying reports of a company's death have been greatly exaggerated, with no details about products or figures. The background post from @h_nilforoshan reports that OpenAI's Decisions API, billed as a "Jev-killer," was benchmarked against Jev for HiringCafe, which serves 2.5 million users. On the task of scoring job-description relevance from 1 to 10, the author reports OpenAI costing 2x more and performing 5-10% worse.

  2. Leandro von WerraAI score36

    Snorkel expands Open Benchmarks Grants to $30M for AI evaluation

    AISnorkel AI is expanding its Open Benchmarks Grants tenfold to a $30M commitment to fund more diverse, robust, and continuously updated open AI benchmarks. The program adds an Open Benchmarks Red Team to test and strengthen those benchmarks, plus a Snorkel Research Fellowship for independent researchers developing new evaluation methods. The source says OBG-funded benchmarks have appeared on model cards from every major frontier lab.

  3. KhazixAI score60

    Claude Max subscribers get monthly API credits usable across Claude models

    AISubscribers to Claude's Max plan can claim monthly API credits: $100 for the $100 tier and $200 for the $200 tier. The credits work for any Claude model and can be used in the user's own apps and other agents. The author argues that bundling monthly API credits alongside a broad model lineup will make it hard for other model companies to compete.

  4. Semafor · TechnologyAI score62

    OpenAI's announced math breakthroughs prompt debate over AI's role in proofs

    AIOpenAI announced hundreds of mathematical breakthroughs, weeks after claiming it had solved one of the most complicated problems in mathematics. The findings raised questions about whether the model used creative thinking or only completed the final steps of human work. Experts say AI could be revolutionary for mathematics if it provides proofs, since proof techniques often underpin other breakthroughs.

  5. Latent SpaceAI score72

    OpenAI publishes 722 math manuscripts from an unreleased internal model

    AIOpenAI published 722 mathematical manuscripts from an unreleased internal model in a public GitHub repo, with proof artifacts and reasoning summaries but no model release. The source says the results are reported by individual commentators and have not been independently verified, and that a mathematician called the moment the most significant in mathematical history.

Oct 6

Oct 6Tue
  1. Epoch AIAI score47

    GPT-6 Astra Hit 100% on EBR-bench Using a Card That Bypassed Its Time Limits

    AIEpoch AI reports that GPT-6 Astra scored 100% on the original EBR-bench by exploiting a card that bypasses the game's time-constraint expectations, so Epoch has banned that card from the default setting. Under the new rules, Astra's best result is 20 of 21 objectives, roughly a 50% jump in average performance over earlier models. Epoch will report revised scores only for Claude Fable 5.1, Claude Opus 5, GPT-5.6 Sol, GPT-6 Astra, and future models.

  2. Thomas WolfAI score38

    OpenAI releases new mathematical results from internal frontier model

    AIOpenAI is releasing a broad range of new mathematical results produced by an internal frontier model, with release guidance from the Institute for Advanced Study's Advisory Group on Mathematics and Artificial Intelligence. The results are available on GitHub at openai/math. The post itself is brief and emphasizes the results rather than hype.