Updated
#Eval/Benchmark
Updated
Items with an AI score under 20 are hidden. Show low-relevance items
Oct 9
Arena.ai@arenaOfficialAI score23
SGLang@sgl_projectOfficialAI score42Cognition first in production on NVIDIA Vera Rubin, 4.8x per-GPU over GB200
AICognition is the first customer running NVIDIA Vera Rubin in production, hosted by CoreWeave, and serving SWE-2 on it with SGLang since September. SGLang says Rubin delivers 4.8x per-GPU throughput over GB200 at matched interactivity.
Ars Technica · AINewsAI score40 Nikon disqualifies AI-tainted winner, names Nguyen Nam Nhat Small World in Motion champion
AINikon disqualified Ning Xu of Tsinghua University from its Small World in Motion competition after an investigation found his entry broke the rules over AI use. Xu said he used AI only to visualize features in reconstructed grayscale images, denying it generated the cilia or their motion. Vietnamese researcher Nguyen Nam Nhat, whose video shows a tiny roundworm and a single-celled organism, is the new winner.
Artificial Analysis@ArtificialAnlysOfficialAI score22Artificial Analysis launches Controlled Voice TTS leaderboard, Speech Arena, and Speech Explorer
AIArtificial Analysis says its Controlled Voice text-to-speech leaderboard now ranks the top models. The post also links a Speech Arena where users vote on models and a Speech Explorer for sample clips.
Mike Knoop@mikeknoopXAI score43Knoop says ARC-AGI-2 is much harder than ARC-AGI-1
AIMike Knoop says ARC-AGI-2 is far harder than ARC-AGI-1, even amid rapid progress on math. He adds that the final open solutions will be useful artifacts to study, and notes that the Kaggle Grand Prize bonus threshold of 85% has been reached this year, the final year for ARC-AGI-2 on Kaggle.
Mike Knoop@mikeknoopXAI score62Tufa Labs hits 88.06% on ARC-AGI-2, clearing the Kaggle bonus threshold
AIMike Knoop says the 85% Grand Prize bonus threshold has been reached on Kaggle. The ARC Prize 2026 leaderboard lists Tufa Labs first at 88.06%, followed by Rabbithole at 80.56% and Yi-Chia Chen at 77.22%. Knoop says this will be the final year for ARC-AGI-2 on Kaggle and expects an open-source, low-cost, offline reproducible solution and model.
François Chollet@fcholletXAI score36ARC-AGI-2 scores on Kaggle reach 88.06% under ARC Prize 2026
AIFrancois Chollet says ARC-AGI-2 scores on Kaggle are now very good. The ARC Prize 2026 high score is 88.06%, set by Tufa Labs, and teams scoring over 85% share a $150K bonus prize.
ARC Prize@arcprizeOfficialAI score29Yi-Chia Chen tops ARC-AGI-3 at 59.17% in ARC Prize 2026
AIARC Prize says Yi-Chia Chen has set a new ARC-AGI-3 high score of 59.17% in ARC Prize 2026. The post says this takes the lead over Tufa Labs.

ARC Prize@arcprizeOfficialAI score42ARC Prize 2026 ARC-AGI-2 high score reaches 88.06%
AITufa Labs posted an 88.06% score on ARC-AGI-2, a new high for the ARC Prize 2026 leaderboard. ARC Prize says a $150K bonus prize, on top of guaranteed prizes, will be split among all teams scoring over 85%.

LangChain@LangChainOfficialAI score29LangSmith data shows Claude Sonnet 5 and GPT-5.6 Luna gaining ground
AILangChain reports that over the last month Claude Sonnet 5 rose from #9 to #2 in model adoption, with 51% more organizations using it. GPT-5.6 Luna climbed from #3 to #1 in call footprints, up 65% in calls, while smaller, faster models dominate call footprints overall. Two open-weights models entered the adoption top 10 but do not lead in call volume.

Arena.ai@arenaOfficialAI score24Arena weekly update: Nano Banana 2.1, Mistral Large 4, Claude Haiku 5.5 rankings
AIArena's weekly update says Nano Banana 2.1 ranked in the top six across three Image Arena modes, with #4 in Multi-Image Edit at 1431 points. Mistral Large 4 placed #43 overall in Agent Arena, 11 spots above Mistral Medium 3.5, and Claude Haiku 5.5 (High) landed #30 in Code Arena WebDev at 1587 points, priced at $0.10/$0.50 per 1M input/output tokens. The post also introduces Arena's Alignment Index and announces a $200M Series B at a $3.1B valuation.
Boris Power@BorisMPowerXAI score28Boris Power calls OpenAI integer multiplication progress "Wow!"
AIBoris Power, who owns the OpenAI account, posted only the word "Wow!" with no details. Background from a separate post says the integer multiplication problem #109 witness value κ rose to 2⁻¹⁰·⁵⁴⁷ (about 6.6857 × 10⁻⁴), past the 2⁻¹¹ threshold. The author notes gains are now fractional and a major breakthrough is still needed.
SiliconANGLE · AINewsAI score35 SailPoint's Navigate event highlights a push to secure AI agent identities in real time
AISailPoint's Navigate conference in Austin, Texas, featured executives arguing that enterprises must secure AI agent identities at machine speed through just-in-time access and enforcement outside the agent. Mark McClain, SailPoint's founder and chief executive, said real-time decision-making is needed because manual administration cannot keep up. The event also covered the Entro Security acquisition and a partnership with AWS on Amazon Bedrock AgentCore, which grew 15-fold in the first six months of the year.
Don't Worry About the Vase (Zvi Mowshowitz)BlogAI score73 OpenAI releases 719 AI-generated math manuscripts, splitting the mathematics community
AIZvi Mowshowitz reports that OpenAI released 722 math manuscripts from an internal frontier model on GitHub, later reduced to 719 after three withdrawals, covering 90 of the top 500 open problems. He says the work came mostly from a single prompt, with an average of three hours of compute per solution. Mathematicians reacted with mixed feelings, and the post highlights concerns about unread papers, cryptography implications, and the role of Lean verification.
QbitAINewsAI score38 Lenovo's TianxiCode Agent Tops SWE-bench-Live Lite Leaderboard at 71%
AILenovo's TianxiCode, paired with DeepSeek-v4.1-Flash, ranked first on the SWE-bench-Live Lite leaderboard with a 71% issue resolution rate and passed official Verified review. The framework combines multi-hop retrieval, autonomous planning with multi-turn tool calling, and test-driven self-correction, and will be applied to Lenovo AI hardware products.
AI SupremacyBlogAI score44 Anthropic Releases Claude Haiku 5.5 as Its Cheapest, Fastest Small Model
AIAnthropic released Claude Haiku 5.5, its cheapest and fastest small model, which it says costs around 75% less to run than Claude Haiku 4.5. The author argues Anthropic is the only frontier model builder, though the piece also covers Nous Research's $90 million Series B and OpenAI's revenue discrepancy.
Oct 8
IThome · AINewsAI score62 Terence Tao questions OpenAI's 719 AI-generated math proofs
AIOpenAI published 719 AI-generated math proofs covering 372 result families, after withdrawing 3 for a symbol error. Reports say the release falls short of the AGMAI advisory group's standards, since it uses proprietary models, includes reasoning chains for only 10 manuscripts, and leaves about 42% unformalized. Terence Tao argues that rapidly solving famous problems harms the mathematical community's understanding and collaboration.
LeiphoneNewsAI score46 IROS 2026 Best Paper goes to LT-Mem robot long-term memory study
AIAt IROS 2026 in Pittsburgh, the Best Paper Award went to Yumin Lee, Hyoseok Ju and Giseop Kim for LT-Mem, a volatility-aware spatio-temporal memory system for lifelong robot scene understanding. The Best Student Paper Award went to Pei-An Hsieh and colleagues for flatness-preserving residual learning enabling real-time tight quadrotor formation flight. Other honors included a humanoid tennis-skills paper and SteadyTray, a humanoid tray-transport study.
The Wall Street Journal · TechNewsAI score60 OpenAI Releases Findings on Over 300 Math Problems After Millennium Prize Solution
AIA month after OpenAI's Millennium Prize solution, the company released findings on more than 300 math problems. The Wall Street Journal says the release was an attempt to win back the mathematics community. Only the excerpt was available, so the source's details on the problems and results are limited.
TechCrunch · AINewsAI score62 Common Sense Media rates ChatGPT for Teens an unacceptable risk over engagement design
AICommon Sense Media labeled ChatGPT for Teens an "unacceptable risk," finding its design still encourages engagement even in crisis situations. The report says the teen version failed to meet commitments on three of five severe harms, and that break reminders appeared only twice across nearly 2,000 prompts. OpenAI disputed the methodology, saying the testing may have ended before parental controls were fully active, and cited its own data showing teens average under 15 minutes a day.
Arena.ai@arenaOfficialAI score44Arena proudly continues working with The House Fund, an investor
AIArena's account of its continued work with The House Fund is a short acknowledgment with no product, figures, or new details. Read alongside the linked background, it sits in the context of Arena's $200M Series B at a $3.1B valuation and its Alignment Index for frontier models.
Arena.ai@arenaOfficialAI score22Salesforce Ventures backs Arena's Series B funding round
AISalesforce Ventures announced an investment in Arena's Series B round, with Arena described as a neutral platform where real users vote on anonymous model outputs across text, coding, vision, image generation, and agentic tasks. Labs and builders use its leaderboards to guide training and purchasing decisions.
Arena.ai@arenaOfficialAI score34Arena raises $200 million Series B, launches AI Alignment Index
AIArena announced a $200 million Series B at a $3.1 billion valuation and launched the Arena AI Alignment Index, a new measure of how AI agents behave in the real world. The main post is a thank-you from Felicis, a backer of Arena since its seed round, and the company's origins as a UC Berkeley research project.
SiliconANGLE · AINewsAI score24 CoreWeave Pitches Open Full-Stack AI Cloud With Forge Development Platform
AICoreWeave is positioning its AI cloud around an open development loop, connecting training, inference and evaluation through its newly announced CoreWeave Forge platform. Chief marketing officer Jean English said the company wants production learnings to improve models and agents and that the loop should work across different models, frameworks and clouds. She argued that competitive differentiation extends beyond GPUs to partner tooling, infrastructure and APIs.
Artificial Analysis@ArtificialAnlysOfficialAI score22Artificial Analysis launches Cyber Index Alliance with IBM and NVIDIA
AIArtificial Analysis has formed the Cyber Index Alliance to set a new standard for evaluating how AI models perform on enterprise cyber defense tasks. Current members are Collinear, IBM, NVIDIA, and Vercel, and partners contribute expert input on the Index design and implementation, plus datasets and external research. Organizations interested in joining can contact cyber@artificialanalysis.ai.

Goodfire@GoodfireAIOfficialAI score25Goodfire launches a challenge to build AI models on a dataset
AIParticipants will use a dataset to build AI models, evaluated through a series of evals ranging from general benchmarks to more complex tasks. Top teams will be shortlisted and have their experimental hypotheses tested in Prima Mente's wet lab.
Goodfire@GoodfireAIOfficialAI score42Goodfire and PrimaMente discover new Alzheimer's biomarkers, launch AI challenge
AIGoodfire and PrimaMente report discovering a new class of Alzheimer's biomarkers. They are now opening a global AI competition to find new treatments, called the Alzheimer's Translation Challenge, led by PrimaMente and AlzData.

Arena.ai@arenaOfficialAI score44Arena raises $200M Series B led by Lightspeed, launches Alignment Index
AIArena has secured a $200M Series B, with Lightspeed doubling down on its investment. The company is also launching the Alignment Index, which measures how closely AI behavior aligns with human values in real-world settings. Arena reports annualized revenue above $100M since its Series A, with millions of people helping evaluate frontier models through real-world use.
TechCrunch · AINewsAI score46 Arena raises $200M at $3.1B valuation, nearly doubling in 10 months
AIArena, the crowdsourced AI model leaderboard that started as a UC Berkeley research project, raised a $200 million Series B at a $3.1 billion valuation, led by Lightspeed Venture Partners and Khosla Ventures. The company said it reached $100 million in annualized run-rate revenue in June, up from $30 million when it raised its $150 million Series A in January at a $1.7 billion post-money valuation.
The DecoderNewsPickAI score80 Mathematicians call for OpenAI boycott after AI-generated proofs flood the field
AIThe Association of Historical Mathematicians (AHM) has called for a boycott of OpenAI after the company released more than 700 AI-generated proof files at once. Fields Medalist Terence Tao, who chairs the group, argues that AI solving open problems autonomously reduces seminars, collaborations, and fertile research directions, and that the field should shift its measure of progress toward explanation and community-building.
Why it matters: The article links the AHM boycott call to Tao's argument that AI-driven proof volume is changing how mathematicians measure progress and whether solutions remain useful.
TechCrunch · AINewsAI score65 OpenAI's math solutions fall short of the field's standards, mathematicians say
AIOpenAI released hundreds of claimed solutions to hard math problems but did not fully meet guidelines from the Advisory Group on Mathematics and Artificial Intelligence. Only 10 of 719 manuscripts included chain-of-thought releases, and just 42% of proofs were formalized. A Cambridge and King's College paper found discrepancies between a natural language proof and its Lean code for a Navier-Stokes-derived problem.
Arena.ai@arenaOfficialAI score37Arena raises $200M Series B at $3.1B valuation, launches Alignment Index
AIArena announced a $200 million Series B at a $3.1 billion valuation, alongside a new Alignment Index that measures whether AI agents behave safely, truthfully, and within the bounds of user requests. The company has surpassed $100 million in annualized revenue, facilitated 350 million sessions and 62 million votes, and led by Felicis and PXD from the seed and Series A stages. Arena positions the index as a way to assess trustworthiness as AI systems increasingly take real actions.
Goodfire@GoodfireAIOfficialAI score25Goodfire monitors cut universal jailbreaks to zero in red-team test
AIGoodfire says a static battery of attacks from FAR AI red-teamed its monitor, and the monitors reduced successful universal jailbreaks to 0. The same testing cut total jailbroken interactions by 97%.

elvis@omarsar0XAI score34Monsoon ASR dataset cuts Bengali Whisper word error rate to 7.65%
AIVoice Arena's Monsoon ASR dataset fine-tuned Whisper Medium on Bengali FLEURS, reducing LLM word error rate from 85.27% to 7.65%. The corpus spans 100,000 hours across 50 languages, and Voice Arena says more than 80 organisations have asked to license it since its launch a week ago.
Cohere@cohereOfficialAI score20Cohere hosts live webinar on future of search and retrieval
AICohere is hosting a live webinar on the future of search and retrieval, covering Embed 5, Parse 5, and its new retrieval methodology, RCP-nDCG. The post is a livestream announcement and does not include details of the methodology or model performance.
Hacker News · AI (150+ points)BlogAI score40 OpenAI withdraws three math papers over a sign error in a proof
AIOpenAI has withdrawn three math manuscripts, including "Algebraicity of Weil classes on split abelian eightfolds," after a sign error invalidated a stabilization-trace cancellation argument. The withdrawals affect two dependent papers, and the withdrawn papers now carry notices linking to archived manuscripts. OpenAI also revised 14 other manuscripts with proof repairs and corrections, and added six formalizations.
Hacker News · AI (150+ points)BlogAI score58 OpenAI withdraws three mathematical results
AIOpenAI has withdrawn three of its mathematical results, according to a Hacker News post linking to a history file in OpenAI's math GitHub repository. The linked page is the only source here, and the feed supplied no further text describing the withdrawn results or the reasons for the withdrawal.
Oct 7
TypeSafe AI@typesafeaiOfficialAI score25Jev-killer OpenAI Decisions API benchmarked against Jev for HiringCafe
AIThe main post is a short reply saying reports of a company's death have been greatly exaggerated, with no details about products or figures. The background post from @h_nilforoshan reports that OpenAI's Decisions API, billed as a "Jev-killer," was benchmarked against Jev for HiringCafe, which serves 2.5 million users. On the task of scoring job-description relevance from 1 to 10, the author reports OpenAI costing 2x more and performing 5-10% worse.
Leandro von Werra@lvwerraXAI score36Snorkel expands Open Benchmarks Grants to $30M for AI evaluation
AISnorkel AI is expanding its Open Benchmarks Grants tenfold to a $30M commitment to fund more diverse, robust, and continuously updated open AI benchmarks. The program adds an Open Benchmarks Red Team to test and strengthen those benchmarks, plus a Snorkel Research Fellowship for independent researchers developing new evaluation methods. The source says OBG-funded benchmarks have appeared on model cards from every major frontier lab.
Khazix@Khazix0918XAI score60Claude Max subscribers get monthly API credits usable across Claude models
AISubscribers to Claude's Max plan can claim monthly API credits: $100 for the $100 tier and $200 for the $200 tier. The credits work for any Claude model and can be used in the user's own apps and other agents. The author argues that bundling monthly API credits alongside a broad model lineup will make it hard for other model companies to compete.