Skip to content

#Open source/Repo

Oct 8

TodayOct 8Thu4 items
  1. Artificial Analysis28

    Among models with a Hallucination-Gated All-Pass Rate above 0%, four set the Pareto frontier for score vs. Cost per Task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) leads at ~$9.50 per task and Muse Spark 1.3 (max) comes second at ~$4.20, while the three Claude models cost ~$18 to ~$22 per task. GPT-6 Luna (max) is the cheapest at ~$0.22 per task, scoring 3.3%.

    Among models with a Hallucination-Gated All-Pass Rate above 0%, four set the Pareto frontier for score vs. Cost per Task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) leads at ~$9.50 per task and Muse Spark 1.3 (max) comes second at ~$4.20, while the three Claude models cost ~$18 to ~$22 per task. GPT-6 Luna (max) is the cheapest at ~$0.22 per task, scoring 3.3%.

  2. Elvis Saravia46

    RSIGym gives research agents services, lifting SWE-bench Verified to 50.33%

    RSIGym provides a research agent with training, inference, evals, and sandboxes as callable services, so it spends its budget on experiments rather than rebuilding infrastructure. With Opus 5 as the researcher, the improved system rose from 17.67% to 50.33% on SWE-bench Verified. The post also highlights a way to measure co-evolution between harnesses and models.

  3. Pandaily38

    Huawei Presents Experimental XMFS Shared-Memory Filesystem at LPC 2026

    Huawei engineers presented XMFS, an experimental Linux kernel prototype filesystem, at the Linux Plumbers Conference in Prague on October 5. It aims to let applications reach cross-node shared memory on CXL 3.0 or Huawei unified bus servers through standard POSIX file calls. The code exists only on openEuler, not in the mainline Linux kernel.

Oct 7

Oct 7Wed
  1. 数字生命卡兹克88

    OpenAI Releases 722 Unpublished AI-Generated Math Manuscripts on GitHub

    OpenAI published 722 math manuscripts covering 372 result groups in a new GitHub repository, openai/math, all produced by an unreleased internal model. The author describes the results as including a near-Riemann hypothesis claim pushed to 0.875, and notes that 25 Fields Medal winners criticized the company's approach to AI math research.

    Why it matters: The piece traces how AI math results moved from benchmarks to open problems, offering context on verification and the mathematicians' pushback.

  2. vLLM46

    vLLM-Omni technical report unifies serving for omni-modality generation

    The vLLM team released a technical report on vLLM-Omni, a unified serving runtime for omni-modality generation spanning multi-stage autoregressive pipelines, iterative diffusion, and stateful sessions. Current LLM servers and diffusion stacks each cover only one of these patterns, pushing deployments to stitch disjoint runtimes together. vLLM-Omni offers a shared control plane in which an orchestrator advances requests across stages, specialized engines handle compute, and a connector carries payloads.

  3. Ai220

    We hope Bolmo opens a path to larger models that adapt how they represent information across languages & domains. Because bytes also represent images & audio, the same ideas could eventually extend beyond text. Learn more in our blog: https://allenai.org/blog/bolmo-nature

    We hope Bolmo opens a path to larger models that adapt how they represent information across languages & domains. Because bytes also represent images & audio, the same ideas could eventually extend beyond text. Learn more in our blog: https://allenai.org/blog/bolmo-nature

Oct 6

Oct 6Tue
  1. GitHub39

    How do you know whether an AI code reviewer catches the issues that matter without adding noise? ReviewBench is a new open benchmark shaped by analysis of 103.9M GitHub pull requests, with 219 PRs across 19 languages. Bring your own code review agent, evaluate it, and submit your results ⬇️ https://github.blog/ai-and-ml/github-copilot/reviewbench-an-open-benchmark-for-ai-code-review/?utm_source=x-promoting-reviewbench-blog-article&utm_medium=social&utm_campaign=reviewbenchmark-oct-2026

    How do you know whether an AI code reviewer catches the issues that matter without adding noise? ReviewBench is a new open benchmark shaped by analysis of 103.9M GitHub pull requests, with 219 PRs across 19 languages. Bring your own code review agent, evaluate it, and submit your results ⬇️ https://github.blog/ai-and-ml/github-copilot/reviewbench-an-open-benchmark-for-ai-code-review/?utm_source=x-promoting-reviewbench-blog-article&utm_medium=social&utm_campaign=reviewbenchmark-oct-2026

  2. METR36

    Here’s a simple example: a METR researcher found a bug in the transcript viewer of Inspect, a popular evaluation framework, that would have enabled an agent to show the user reviewing its transcript a fake (or edited) version.

    Here’s a simple example: a METR researcher found a bug in the transcript viewer of Inspect, a popular evaluation framework, that would have enabled an agent to show the user reviewing its transcript a fake (or edited) version.

  3. ARC Prize31

    @SpaceXAI - Leaderboard: https://arcprize.org/leaderboard - Reproduce the public results: https://github.com/arcprize/arc-agi-benchmarking - Testing policy: https://arcprize.org/policy - Full Grok 4.7 results: https://arcprize.org/results/xai-grok-4-7

    @SpaceXAI - Leaderboard: https://arcprize.org/leaderboard - Reproduce the public results: https://github.com/arcprize/arc-agi-benchmarking - Testing policy: https://arcprize.org/policy - Full Grok 4.7 results: https://arcprize.org/results/xai-grok-4-7

Oct 5

Oct 5Mon
  1. GitHub Blog · AI & ML63

    GitHub releases ReviewBench, an open benchmark for AI code review agents

    GitHub has released ReviewBench, an open benchmark for evaluating AI code review agents on 219 public pull requests across 19 languages. The benchmark reports grounded and augmented precision, recall, and F1 metrics, and its dataset, rubric, and judge are publicly available. GitHub says ReviewBench predicted the direction of a Copilot code review ensemble experiment's production results before A/B testing.

    Why it matters: The post explains how ReviewBench was built and validated, and reports an offline-to-production comparison that shows how well a benchmark predicts real experiment outcomes.

  2. Clément Delangue62

    Hugging Face turns 10 coding harnesses into RL environments via a capture proxy

    Hugging Face says a capture proxy lets reinforcement learning train open models inside unmodified coding harnesses such as Claude Code, Codex, and OpenCode. The proxy records the exact token IDs and logprobs vLLM samples and hands them to TRL for training. On LFM2.5-2.6B, training in four harnesses at once raised OpenCode results from 34% to 58%, while SFT on 3,189 Qwen3.8-27B rollouts plateaued at 47.5%.

Oct 2

Oct 2Fri
  1. Liquid AI64

    Hugging Face guide shows multi-harness RL for coding agents via a capture proxy

    Liquid AI shared a Hugging Face guide to multi-harness reinforcement learning for coding agents, in which a proxy records the token ids and logprobs vLLM samples so training works without changing the harness. Per the quoted post, LFM2.5-2.6B rose from 42% to 54% after training across four harnesses at once, and imitation fine-tuning on 3,189 rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs. The proxy, trainer, tasks, SFT data, training code and seven trained models are described as open.

  2. Google Research60

    Google's TEE-based federated learning system adds verifiable privacy guarantees

    Google announces a next-generation federated learning system that uses Trusted Execution Environments to provide verifiable, auditable data anonymization. The system publishes access policies to a public transparency log and is deployed in Gboard, which has launched English and Japanese next-word prediction models with stronger privacy guarantees and improved accuracy. Training time has also sped up significantly because computation moved to the server and is parallelized across many machines.

    Why it matters: The post shows how Trusted Execution Environments make federated learning's privacy claims externally verifiable, rather than relying on trust in the server operator.

Oct 1

Oct 1Thu
  1. ARC Prize22

    @Alibaba_Qwen @baseten - Leaderboard: https://arcprize.org/leaderboard - Reproduce the public results: https://github.com/arcprize/arc-agi-benchmarking - Testing policy: https://arcprize.org/policy - Full Qwen3.8-27B results: https://arcprize.org/results/alibaba-qwen3-8-27b

    @Alibaba_Qwen @baseten - Leaderboard: https://arcprize.org/leaderboard - Reproduce the public results: https://github.com/arcprize/arc-agi-benchmarking - Testing policy: https://arcprize.org/policy - Full Qwen3.8-27B results: https://arcprize.org/results/alibaba-qwen3-8-27b

  2. Sara Hooker35

    We released invent a dataset. Today we release the technical report that goes with it. Huge shoutout to everyone @singhshiviii @andrijazzz @lekeonilude @sudip_r0y 🔥 This tackles the hardest setting. How do you go from zero data regime to a training set for a capability.

    We released invent a dataset. Today we release the technical report that goes with it. Huge shoutout to everyone @singhshiviii @andrijazzz @lekeonilude @sudip_r0y 🔥 This tackles the hardest setting. How do you go from zero data regime to a training set for a capability.

Sep 26

Sep 26Sat
  1. Alexander Doria38

    Since I went into this release: *It's a smaller selection of 989 envs used to RL a 9B distilled model, not the big MiMo. *Rewards are not self-contained: general part need to set up a judge and webdev rely on their own grader service+vlm. *Most important part is inside general/envs directory (+ docker), not the dataset displayed on hf: genuinely solid mix of real/simulated documents we rarely see in OSS.

    Since I went into this release: *It's a smaller selection of 989 envs used to RL a 9B distilled model, not the big MiMo. *Rewards are not self-contained: general part need to set up a judge and webdev rely on their own grader service+vlm. *Most important part is inside general/envs directory (+ docker), not the dataset displayed on hf: genuinely solid mix of real/simulated documents we rarely see in OSS.

Sep 25

Sep 25Fri
  1. Higgsfield AI28

    The first hybrid AI-animated short made with the Blender + Higgsfield pipeline. We combined handcrafted 2D storyboards and 3D Blender previs with AI-animation to create “Passport Rush.” The film is now open source, so you can see how we made it. Every prompt and asset is public.

    The first hybrid AI-animated short made with the Blender + Higgsfield pipeline. We combined handcrafted 2D storyboards and 3D Blender previs with AI-animation to create “Passport Rush.” The film is now open source, so you can see how we made it. Every prompt and asset is public.

Sep 21

Sep 21Mon
  1. Microsoft Research50

    Microsoft Research open-sources RetroChimera, a retrosynthesis model published in Nature

    Microsoft Research published RetroChimera, a retrosynthesis framework that combines the R-SMILES 2 Transformer model and the NeuralLoc graph neural network through learned ensembling to propose synthesis routes for small molecules. In blind tests, PhD-level chemists preferred its individual reaction predictions over those from preceding models and recorded literature reactions. The implementation and weights are open-sourced for researchers developing new medicinal molecules and materials.

Sep 17

Sep 17Thu
  1. Anthropic26

    You can find all of the code on GitHub: https://github.com/anthropics/uplifting-biomolecular-modeling And the full results in our technical report: https://www-cdn.anthropic.com/d8ca26d0d205708d26c7337cf4cfe7cb52e9b671.pdf

    You can find all of the code on GitHub: https://github.com/anthropics/uplifting-biomolecular-modeling And the full results in our technical report: https://www-cdn.anthropic.com/d8ca26d0d205708d26c7337cf4cfe7cb52e9b671.pdf

Aug 27

Aug 27Thu
  1. LMSYS Org47

    MiniMax-H3 gets up to 6.24x speedup on 8×H200 GPUs

    MiniMax-H3 on 8×H200 GPUs reaches 1.85–1.95x lossless speedup over Diffusers without approximation, with fixed prompts, seeds, resolution, FPS, and 50 denoising steps. Adding step reuse and sparse attention raises speedup to as much as 6.24x, but quality varies by workload, with SSIM from 0.76 to 0.91. Two presets trade off the two: a conservative Cache-DiT setting gives 2.99x at 0.90–0.98 SSIM, while a faster SubBlock 0.75 plus Cache-DiT stride gives 4.90–5.93x at 0.77–0.92.

Jul 15

Jul 15Wed
  1. Liquid AI Newsletter38

    Liquid AI Releases Antidoom and IFStruct to Fix Reasoning Loops and Schema Errors

    Liquid AI released Antidoom, an open-source method that retrains a single overtrained token to eliminate "doom loops" in small reasoning models. On LFM2.5-2.6B and Qwen3.5-4B, loop rates fell from 10.2% to 1.4% and from 22.9% to 1%, respectively. The company also released IFStruct, an open-source benchmark measuring whether model outputs satisfy a schema, where LFM2.5-350M rose from 21.10% to 44.90% after training.

Apr 23

Apr 23Thu
  1. OpenAI Alignment Research Blog44

    OpenAI Open-Sources Chain-of-Thought Monitorability Evaluation Datasets and Code

    OpenAI is releasing a subset of datasets, reference code, and the g-mean 2 metric for evaluating chain-of-thought monitorability. The release includes most datasets from its monitorability suite, while some evaluations relying on private or restricted data were omitted. The company says it will keep reporting monitorability results in future frontier reasoning model system cards.

Mar 26

Mar 26Thu

Mar 17

Mar 17Tue
  1. BAAI46

    BAAI unveils RoboBrain-Dex, dexterous manipulation trained on human egocentric data

    BAAI has released RoboBrain-Dex, a dexterous manipulation model for embodied intelligence trained on large-scale, diverse human egocentric data rather than massive robot teleoperation datasets. The company says this approach yields strong generalization, marking a shift from small-data, weakly generalizing methods toward big-data robotic manipulation. The code has been open-sourced on GitHub.

Mar 13

Mar 13Fri
  1. Berkeley AI Research34

    SPEX and ProxySPEX Identify Influential LLM Interactions at Scale with Fewer Ablations

    Berkeley AI Research introduces SPEX, a signal-processing framework that identifies influential interactions in LLMs using far fewer ablations than exhaustive analysis. A hierarchy-based extension, ProxySPEX, matches SPEX performance with around 10x fewer ablations. The methods apply to feature, data, and model component attribution.

Feb 19

Feb 19Thu

Dec 12, 2025

Dec 12, 2025Fri
  1. Apple · new models on Hugging Face46

    Apple's SHARP Turns a Single Photo into a 3D Scene in Under a Second

    Apple has released SHARP, a model that generates a 3D Gaussian representation of a scene from a single photograph in less than a second on a standard GPU. The output renders in real time as high-resolution photorealistic views of nearby camera positions, with metric absolute scale, and the paper reports reductions of 25–34% in LPIPS and 21–43% in DISTS versus the best prior model.

Aug 11, 2021

Aug 11, 2021Wed
  1. ByteDance20

    A team of ByteDance researchers has won the Association for Computational Linguistics (ACL) 2021 Best Paper Award! 🏆 The team proposed a simple yet efficient machine translation solution named VOLT, which is now open source to developers. Check it out: https://lnkd.in/g32UfH2C

    A team of ByteDance researchers has won the Association for Computational Linguistics (ACL) 2021 Best Paper Award! 🏆 The team proposed a simple yet efficient machine translation solution named VOLT, which is now open source to developers. Check it out: https://lnkd.in/g32UfH2C