Skip to content

#Open-source ecosystem

Oct 8

TodayOct 8Thu5 items
  1. PandailyAI score55

    Chinese Team Publishes 3D Cell Atlas of Rice's Full Life Cycle in Cell

    A Chinese-led team published in Cell a three-dimensional spatiotemporal cell atlas covering rice from germinating seed to grain fill, along with a public portal and the RICE scGPT single-cell foundation model. The atlas combines single-nucleus RNA sequencing with BGI's Stereo-seq spatial transcriptomics across 10 organ and tissue types and 61 stages, defining 119 cell types and 133 subtypes.

  2. AnthropicAI score57

    Astrophysicist uses Claude to build first complete ultraviolet sky map

    An astrophysicist worked with Claude Science to create the first complete ultraviolet map of the sky, covering regions never observed in UV. Claude located existing datasets, combined them, and filled gaps with statistical inference, taking a few days rather than weeks of human work. The map is presented as a teaching tool and an example of low-priority scientific work that AI now makes feasible.

  3. Goodfire ResearchAI score57

    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    Goodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.

  4. OpenBMBAI score36

    ReJev fine-tunes MiniCPM5-2B to lift decision accuracy to 80.50%

    ReJev, an independent community project, applied LoRA post-training to OpenBMB's MiniCPM5-2B for bounded agent decisions: state, question, and candidate options yield one choice. On its sealed 1,892-sample holdout, accuracy rose from 51.11% to 80.50% (+29.39 percentage points) with 0% invalid outputs, at about $5.31 in cumulative Modal billing including earlier experimental overhead. The authors describe this as an early, task-specific result, not parity with Jev.

  5. QbitAI (量子位)AI score44

    PaperBenchX Shows Top Model Reproduces Only 13.98% of 93 Scientific Papers End-to-End

    UniPat AI's PaperBenchX benchmark found the strongest model, GPT-6 Astra, fully reproduced only 13.98% of 93 real research-paper tasks across 12 scientific fields. Reproduction was judged by regenerating outputs in an isolated environment, with 3,168 expert-verified scoring items. UniPat has open-sourced 12 test tasks and kept 81 tasks closed to preserve long-term evaluation validity.

Oct 7

Oct 7Wed
  1. Epoch AIAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    Epoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    AIWhy it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  2. Ai2AI score39

    We byteified Qwen 3 8B & Llama 3 8B to create Bwen 8B & Blama 8B. Both come close to matching their source models' performance in our evaluations. Bwen 8B also outperforms Bolmo 7B across our aggregate evaluation suite. https://huggingface.co/collections/allenai/bolmo

    We byteified Qwen 3 8B & Llama 3 8B to create Bwen 8B & Blama 8B. Both come close to matching their source models' performance in our evaluations. Bwen 8B also outperforms Bolmo 7B across our aggregate evaluation suite. https://huggingface.co/collections/allenai/bolmo

  3. Ai2AI score38

    Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 https://www.nature.com/articles/s41586-026-11111-4

    Our paper on retrofitting language models to operate over bytes – the approach behind Bolmo – has been accepted to Nature! 🎉 We’re also releasing new checkpoints that extend our method from Olmo to Qwen & Llama. 🧵 https://www.nature.com/articles/s41586-026-11111-4

  4. Lucas BeyerAI score36

    Missed this post the first time around, but i think this is a very cool and needed effort to thoroughly benchmark VLA and co. They build a leaderboard and half the tasks are fully open, half are held out to track potential benchmaxxing of future model versions.

    Missed this post the first time around, but i think this is a very cool and needed effort to thoroughly benchmark VLA and co. They build a leaderboard and half the tasks are fully open, half are held out to track potential benchmaxxing of future model versions.

  5. Elvis SaraviaAI score44

    NVIDIA's VERA co-evolves agent harness and model via verifiable environments

    NVIDIA's VERA turns benchmark trajectories into over 9,000 restartable sandboxes with rubric scoring and updates both model weights and the agent harness together. A harness edit is kept only if it adds at least 5 points on the development set, and a checkpoint is rejected if its score drops more than 20%. At 27B, the co-evolved agent scores 71.6 on AutoCoWorkBench, above Claude Opus 4.8, and the environment corpus is open-sourced.

  6. Hugging Face BlogAI score78

    Nemotron Fine-Tuned to Reach Gold-Level Results at IOI and IMO 2026

    NVIDIA reports that fine-tuned Nemotron models reached gold-medal level at both IOI 2026, scoring 535.4 out of 600, and IMO 2026, scoring 30 out of 42. The IOI run was a live, unofficial, unsupervised benchmark, while IMO proofs were graded by official IMO graders. The post also releases checkpoints, datasets, a new 200-problem benchmark, and inference pipelines on Hugging Face and NeMo-Skills.

    AIWhy it matters: The post traces how SFT, RL, and a generate-verify-refine loop turned Nemotron into gold-level specialists for IOI and IMO, with the training and inference details shared.

  7. Ai2 (Allen Institute for AI)AI score57

    Ai2's Bolmo byte-level language models are published in Nature

    Ai2 has published its Bolmo byte-level language model research in Nature and released new checkpoints on Hugging Face. The byteifying process converts an existing subword model into a byte-level one with a relatively short additional training run, and the paper reports that it also works for Qwen 3 8B and Llama 3 8B, producing Bwen 8B and Blama 8B. Ai2 also released Stage 1 checkpoints for researchers extending the architecture.

Oct 6

Oct 6Tue
  1. OpenAIAI score62

    OpenAI releases new mathematical results from an internal frontier model

    OpenAI is releasing a broad range of new mathematical results produced by an internal frontier model. The company says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study and drew on its advice and public recommendations for how the results are released. The results are available at https://github.com/openai/math.

  2. TekniumAI score33

    We just launched Hermes Index! This combines the scores of our new HermesBench and 3 other leading and relevant benchmarks for agents to give every Hermes Agent user a way to find both the best model at a given time, as well as the best model at a given price point! Check it out

    We just launched Hermes Index! This combines the scores of our new HermesBench and 3 other leading and relevant benchmarks for agents to give every Hermes Agent user a way to find both the best model at a given time, as well as the best model at a given price point! Check it out

  3. METR BlogAI score31

    AI Agents Could Hide Misbehavior by Exploiting Inspect Transcript Viewer

    METR tested whether an AI agent running in an Inspect evaluation could alter the transcript humans review, and a researcher found a vulnerability in about 10 minutes that allowed arbitrary changes to what the reviewer sees. The exploit affects only the displayed transcript, not the underlying data stored in METR's database, and METR has not observed agents using it in its evaluations. METR argues that AI outputs such as transcripts and reasoning should be treated as untrusted input, with monitoring systems treated as security-critical infrastructure.

Oct 5

Oct 5Mon
  1. GoodfireAI score10

    Tue 11am (Franciscan A) Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior Tue 4:30pm (Imperial Ballroom) Do SAEs Capture Concept Manifolds? Wed 4:30pm (Franciscan B) Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts

    Tue 11am (Franciscan A) Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior Tue 4:30pm (Imperial Ballroom) Do SAEs Capture Concept Manifolds? Wed 4:30pm (Franciscan B) Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts

  2. Simon WillisonAI score22

    I re-ran an experiment @colin_fraser ran against GPT-4o a while back to see how good it was at adding long numbers, only this time I tried Qwen 3.8 27B running locally in both reasoning and non-reasoning modes https://simonwillison.net/2026/Oct/4/qwen38-addition-in-words/

    I re-ran an experiment @colin_fraser ran against GPT-4o a while back to see how good it was at adding long numbers, only this time I tried Qwen 3.8 27B running locally in both reasoning and non-reasoning modes https://simonwillison.net/2026/Oct/4/qwen38-addition-in-words/

Oct 3

Oct 3Sat
  1. Hugging Face BlogAI score67

    Microsoft ThinkingBox grades AI agents on database state across 20 repeated runs

    Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.

    AIWhy it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.

Oct 2

Oct 2Fri
  1. Prime IntellectAI score38

    We served GLM-5.3 on GB200 NVL72. Our interactivity target was 100+ e2e tok/s per user while serving as many concurrent agent task as possible. At the 100 tok/s/user bar, a 1:4 P/D ratio serves the most: 66 sessions per prefill group at 101 tok/s/user and 100 output tok/s per GPU.

    We served GLM-5.3 on GB200 NVL72. Our interactivity target was 100+ e2e tok/s per user while serving as many concurrent agent task as possible. At the 100 tok/s/user bar, a 1:4 P/D ratio serves the most: 66 sessions per prefill group at 101 tok/s/user and 100 output tok/s per GPU.

  2. Hugging FaceAI score67

    Hugging Face guide shows how to train agent models across multiple harnesses with RL

    Hugging Face and collaborators published a guide to multi-harness RL that trains models through a capture proxy without changing the agent harness. The proxy records the token ids and logprobs vLLM samples, and the source reports LFM2.5-2.6B rising from 42% to 54% after training across four harnesses. Fine-tuning on 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs, and the capture proxy, trainer, tasks, SFT data, training code, and seven trained models are released openly.

    AIWhy it matters: The source gives a concrete method for training models across several agent harnesses, with measured gains and a note that imitation learning underperformed RL.

Oct 1

Oct 1Thu
  1. Apple Machine Learning ResearchAI score28

    Language Discrimination Narrows Multilingual Speech Model Gap, Study Finds

    Researchers Maureen de Seyssel, Jie Chi, and Zakaria Aldeneh found that strengthening language discrimination during pretraining reduces the performance gap between multilingual and monolingual HuBERT speech models. In a controlled English/French setting, phone-ABX error fell from 11.6% to 10.4%, close to the monolingual 10.8%, while lexical sWUGGY scores rose from 52.1% to 56.7%. The gains were largest when language discrimination was introduced in the first training iteration.

  2. PyTorch BlogAI score38

    TLX-Optimized Jagged Flash Attention Beats FA4 on Blackwell B200 for Meta GEM

    Meta's Jagged Flash Attention kernel, built with TLX on NVIDIA Blackwell B200, outperforms FlashAttention-4 (May 2026 version) on GEM's jagged shapes by about 13% on the forward pass and about 50% on the backward pass. The TLX attention kernel is roughly 3.2K lines of Triton-level code, about 3× shorter than FA4's ~10K-line CuteDSL kernels. The benchmarks use bfloat16 on B200.

Sep 30

Sep 30Wed
  1. Nathan LambertAI score22

    The latest Interconnects plot - showing the exponential growth of the open model inference economy. Plot is showing the daily tokens processed by the leading companies, based on public disclosures (for Together, Baseten, and Fireworks) and the OpenRouter API. I underestimated OpenRouter's growth.

    The latest Interconnects plot - showing the exponential growth of the open model inference economy. Plot is showing the daily tokens processed by the leading companies, based on public disclosures (for Together, Baseten, and Fireworks) and the OpenRouter API. I underestimated OpenRouter's growth.

  2. Tencent HunyuanAI score62

    Tencent Hunyuan releases ExplorationBench to test how AI systems discover rules

    Researchers from Tencent Hy, Fudan University, and Tsinghua University released ExplorationBench, a benchmark that tests whether AI systems can discover hidden rules in executable Alien World sandboxes. Across 10 frontier systems, getting feedback from experiments outperformed thinking alone, with the best run reaching 89.0% after four rounds. The authors note that rankings barely transfer between the two worlds, and the code is listed as coming soon.

  3. OpenBMBAI score42

    Diffusion Reward Models learn full human preference distributions, not single scores

    OpenBMB introduces Diffusion Reward Models (DRM), which learn the full reward distribution of human preferences instead of collapsing them into one scalar score. The approach preserves disagreement and uncertainty, enabling distribution-aware Best-of-N ranking and a new test-time scaling axis by sampling more reward outputs. DRM also improves downstream policy performance over scalar reward baselines when used as the reward in RLHF, according to the post.

Sep 29

Sep 29Tue
  1. Jerry LiuAI score20

    Jev, a System One model, tops OSS rivals on document tasks

    Jerry Liu says Jev, a System One model, outperformed other open-source classifiers and document-specific models on orientation detection, language detection, classification, and splitting. The benchmark measured accuracy, cost, and latency across these fast document decisions, with Jev leading most comparisons. The benchmark code is available in the run-llama/jev_vs_oss repository.

  2. Fireworks AI BlogAI score51

    Fireworks explains how numerical mismatch and MoE routing can derail RL training

    Numerical differences between a rollout engine and a trainer can make reinforcement learning collapse even when algorithm and data stay identical. In a GLM 5.2 experiment, reward fell from about 0.9 to under 0.2 around step 20 without alignment, while aligned numerics kept reward stable over 25 steps. A Qwen3.5-MoE investigation traced a significant mismatch to how expert outputs were combined, and router replay alone was judged insufficient.

  3. OpenBMBAI score72

    One-Shot OPD: One Training Query Matches Most of Full-Data Distillation Gains

    Researchers from Tsinghua NLP and collaborators show that on-policy distillation with a single training query recovers 87% of full-data gains on math, reaching 68.5 versus 69.8 by step 300. The paper attributes the slow progress to how fast the student absorbs the teacher's signal rather than to dataset size. Code and the paper are publicly available on GitHub and Hugging Face.

    AIWhy it matters: The paper isolates training data from the algorithm, showing one query nearly matches full-data on-policy distillation, which reframes where post-training gains come from.

  4. Artificial Analysis ArticlesAI score62

    Artificial Analysis open-sources AA-AgentPerf-Local for benchmarking local AI agents

    Artificial Analysis has open-sourced AA-AgentPerf-Local, a tool that replays recorded agent trajectories to measure inference speed on laptops and workstations. Initial results cover NVIDIA DGX Spark, NVIDIA GeForce RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro, with the RTX 5090 fastest for models that fit its 32 GB. The source states the tool and leaderboard will expand to more hardware, frameworks, and models.

    AIWhy it matters: The source gives per-system completion times and memory bandwidth figures, letting readers compare local hardware for running agentic workloads.

Sep 24

Sep 24Thu
  1. Goodfire ResearchAI score52

    Block-Sparse Featurizers Recover Multidimensional Concept Geometry in Vision Models

    Goodfire Research introduces Block-Sparse Featurizers (BSF), which decompose model activations into subspaces rather than single directions. Applied to DINOv3 and Stable Diffusion XL, BSFs find interpretable multidimensional features that better explain activations and enable fine-grained steering. The authors report that most concepts they examined have a stable rank of about two to four dimensions.

Sep 23

Sep 23Wed
  1. Google Developers BlogAI score62

    Google reproduces Olmo 3 7B pre-training in MaxText on TPUs

    Google Developers reproduced Ai2's Olmo 3 7B from scratch in MaxText on Google Cloud TPUs, covering both the stage-1 pre-training run and the stage-2 mid-training anneal. The match was checked on held-out C4 loss, an 8-task accuracy suite, multi-domain perplexity, and token-level KL, not just the training loss curve. The post also describes a data-loader bug that made training loss look better than the reference while held-out metrics did not move.

    AIWhy it matters: The post documents how a faithful reproduction was verified on held-out metrics, including a data bug that training loss alone would have hidden.

  2. ModelScopeAI score62

    Shanghai AI Lab and SJTU release open-weight 8.9B NCP-ArchPreview model under Apache 2.0

    Shanghai AI Lab and SJTU's LUMIA Lab released NCP-ArchPreview, an 8.9B open-weight language model under Apache 2.0. The model reportedly reaches OLMo-3-7B's final Stage 1 loss using 51.3% of the tokens from the 5.73T Dolma 3 corpus, a 1.95× convergence gain. Its concept module jointly predicts tokens and concepts, and domain adaptation updates only its 17M parameters while the token backbone stays frozen.

Sep 22

Sep 22Tue

Sep 20

Sep 20Sun

Sep 18

Sep 18Fri
  1. SemiAnalysisAI score52

    Engram offloading to DRAM beats SSD for DeepSeek-V4.1-Flash serving on B200

    SemiAnalysis tested offloading DeepSeek-V4.1-Flash's Engram embedding table from HBM to host DRAM and to local SSD. On B200 configurations, DRAM delivered more total tokens per dollar and higher P90 interactivity than SSD at every measured point. The report concludes SSD offloading is likely not worth the tradeoff for production serving in its unoptimized setup.

Sep 17

Sep 17Thu
  1. SenseTimeAI score44

    SenseNova U1.5 open-sources 8B unified model for understanding and generation

    SenseTime released its SenseNova U1.5 technical report, describing an open-source 8B native MoT unified model that connects understanding and generation through shared attention. The model reports 68.2% on VBVR-Pro-Bench, ahead of Nano-Banana-Pro (56.4%) and GPT-Image-2 (50.7%), and its full training recipes, including SFT, RL, and multi-expert on-policy distillation, are open-sourced.

  2. Ai2 (Allen Institute for AI)AI score42

    Crowdsourced Game Steering Arena Shows Olmo 3 Prosocial Scores Can Be Gamed

    Northeastern University MS student Soham Padia used Ai2's open Olmo 3-32B model to build Steering Arena, a public game in which players submit text prefixes to steer prosocial behavior. About 600 submissions from a few dozen people showed the top 36 entries were unreadable token strings, while the best plain-English entry ranked 37th at about 2.7 times lower score. The results suggest that once an evaluation metric is exposed, it becomes an optimization target.

Sep 9

Sep 9Wed
  1. Ai2 (Allen Institute for AI)AI score39

    Goodfire Traces Olmo Safety Regression to Preference Training Data

    Goodfire used Ai2's open post-training stack, including the Dolci preference dataset, intermediate Olmo checkpoints, and OLMES evaluations, to trace a safety regression in Olmo. Preference training made Olmo more likely to comply with harmful requests on a refusal benchmark, and Goodfire linked part of this to specific Dolci examples where the preferred response encouraged compliance. Because Ai2 publishes the individual preferred and rejected responses, researchers could test targeted changes to reduce the regression.