Skip to content

#Hugging Face

Oct 3

Oct 3Sat
  1. Hugging Face BlogAI score67

    Microsoft ThinkingBox grades AI agents on database state across 20 repeated runs

    Microsoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.

    AIWhy it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.

Oct 2

Oct 2Fri
  1. Hugging Face BlogAI score70

    Ai2 open-sources AstaBrief 8B, a fast model for generating cited research reports

    Ai2 released AstaBrief 8B, an open-weights model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. The model runs as Fast mode in Asta, averaging 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The post also describes filtering synthetic training data by citation density and building DPO pairs judged by two models that agreed.

    AIWhy it matters: The post explains how supervised fine-tuning, preference data, and citation-density filtering were used to build a cited-report model, which is useful for teams training their own models.

  2. Liquid AIAI score64

    Hugging Face guide shows multi-harness RL for coding agents via a capture proxy

    Liquid AI shared a Hugging Face guide to multi-harness reinforcement learning for coding agents, in which a proxy records the token ids and logprobs vLLM samples so training works without changing the harness. Per the quoted post, LFM2.5-2.6B rose from 42% to 54% after training across four harnesses at once, and imitation fine-tuning on 3,189 rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs. The proxy, trainer, tasks, SFT data, training code and seven trained models are described as open.

  3. Merve NoyanAI score36

    you don't need large GPUs to run most decision models, this makes them a perfect match for llama.cpp give the blog a read on how to get them up and running! team will support more and more models in upcoming days

    you don't need large GPUs to run most decision models, this makes them a perfect match for llama.cpp give the blog a read on how to get them up and running! team will support more and more models in upcoming days

  4. Hugging FaceAI score67

    Hugging Face guide shows how to train agent models across multiple harnesses with RL

    Hugging Face and collaborators published a guide to multi-harness RL that trains models through a capture proxy without changing the agent harness. The proxy records the token ids and logprobs vLLM samples, and the source reports LFM2.5-2.6B rising from 42% to 54% after training across four harnesses. Fine-tuning on 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs, and the capture proxy, trainer, tasks, SFT data, training code, and seven trained models are released openly.

    AIWhy it matters: The source gives a concrete method for training models across several agent harnesses, with measured gains and a note that imitation learning underperformed RL.

  5. Hugging Face BlogAI score62

    AutoSynthData generates targeted training data for enterprise agents from failures

    ServiceNow CoreAI introduced AutoSynthData, which uses a target model's failures and a stronger teacher's successes to generate and validate new agent training tasks. In EnterpriseOps Gym experiments, the Hybrid domain produced 2,000 samples and raised Gemma-4-26B-A4B-it mean Pass@1 by 7.2 percentage points, while the ITSM domain produced 1,994 samples and raised it from 18.77% to 27.18%.

    AIWhy it matters: The post shows how failure analysis, teacher demonstrations, and verifier checks combine into a repeatable pipeline for generating targeted agent training data.

Oct 1

Oct 1Thu
  1. Lewis TunstallAI score44

    Training LFM2.5-2.6B inside four agent harnesses boosts held-out tasks

    Hugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.

  2. Merve NoyanAI score46

    Hugging Face ml-intern myths debunked > ml-intern is an open-source ML eng/research harness you can use on your own local setup for free ✅ > it is hosted on Hugging Chat because we want no code interface for people to train their sota models in any task ✅ > second option uses our infra, ml-intern picks the cheapest GPU for your tasks so you get your model for few dollars only for instance @MaziyarPanahi trained a Jev model simply by prompting today for $6, no code 🙌🏼

    Hugging Face ml-intern myths debunked > ml-intern is an open-source ML eng/research harness you can use on your own local setup for free ✅ > it is hosted on Hugging Chat because we want no code interface for people to train their sota models in any task ✅ > second option uses our infra, ml-intern picks the cheapest GPU for your tasks so you get your model for few dollars only for instance @MaziyarPanahi trained a Jev model simply by prompting today for $6, no code 🙌🏼

  3. Merve NoyanAI score22

    btw if you want to learn to train models to be agents, we have a four chapter series on youtube with everything open-source 🙌🏻 get started here https://www.youtube.com/watch?v=rNgUoH7Wbv8&list=PLo2EIpI_JMQvQZm-kVlz4wY1vWF0LBcf5&index=3

    btw if you want to learn to train models to be agents, we have a four chapter series on youtube with everything open-source 🙌🏻 get started here https://www.youtube.com/watch?v=rNgUoH7Wbv8&list=PLo2EIpI_JMQvQZm-kVlz4wY1vWF0LBcf5&index=3

  4. Merve NoyanAI score13

    pewdiepie does.. *checks notes* agentic training https://youtu.be/ODDJXGY_1kQ?is=6BUanblN64iBJvxd > he doesn't mention the tool I think he just says he does grpo > he wanted to distill Sol but got banned few times > he built a website for people to donate traces but no one did (he could have just went to Hub 🌝) I just want to help him out so bad but there's no way to reach out the man 🌝

    pewdiepie does.. *checks notes* agentic training https://youtu.be/ODDJXGY_1kQ?is=6BUanblN64iBJvxd > he doesn't mention the tool I think he just says he does grpo > he wanted to distill Sol but got banned few times > he built a website for people to donate traces but no one did (he could have just went to Hub 🌝) I just want to help him out so bad but there's no way to reach out the man 🌝

Sep 30

Sep 30Wed
  1. Thomas WolfAI score31

    ESM-2 came out in 2022. It's still downloaded hundreds of thousands of times a month. That only works because someone keeps maintaining the software underneath it. @huggingface 🤝 @os4science are teaming up to find those libraries and back the people behind them 🧬 https://os4science.org/news/hugging-face-open-source-for-science-fund/

    ESM-2 came out in 2022. It's still downloaded hundreds of thousands of times a month. That only works because someone keeps maintaining the software underneath it. @huggingface 🤝 @os4science are teaming up to find those libraries and back the people behind them 🧬 https://os4science.org/news/hugging-face-open-source-for-science-fund/

  2. Daniel HanAI score7

    We're co-hosting an open source party with @HuggingFace! 🤗🦥 Come join us, we'll be handing out @UnslothAI merch, showcasing our upcoming features and more! It'll be a super fun night with lots of demos, DJs, food and more. Join: http://luma.com/OpenTogether Use code: NOSLOTHSHERE to get in.

    We're co-hosting an open source party with @HuggingFace! 🤗🦥 Come join us, we'll be handing out @UnslothAI merch, showcasing our upcoming features and more! It'll be a super fun night with lots of demos, DJs, food and more. Join: http://luma.com/OpenTogether Use code: NOSLOTHSHERE to get in.

Sep 29

Sep 29Tue
  1. Hugging Face BlogAI score46

    Open TTS Leaderboard ranks multilingual and voice cloning models using objective metrics

    Hugging Face released the Open TTS Leaderboard, which evaluates open-source text-to-speech models using objective metrics instead of arena-style human votes. It measures intelligibility via WER and CER using Qwen3 ASR, speed via RTFx and time-to-first-audio on an H200 GPU, and speaker similarity via WavLM embeddings. The leaderboard covers multilingual results and voice cloning, and it is intended to complement, not replace, human preference rankings.

  2. Liquid AIAI score32

    Announcing d1, our first decision model. It's the first model to outperform Jev on @huggingface's Decision Index. > wins on multilingual evals > more robust against prompt injection > handles longer inputs more effectively > built for fast, structured decision-making in software environments > Liquid API: http://console.liquid.ai > Available on OpenRouter soon

    Announcing d1, our first decision model. It's the first model to outperform Jev on @huggingface's Decision Index. > wins on multilingual evals > more robust against prompt injection > handles longer inputs more effectively > built for fast, structured decision-making in software environments > Liquid API: http://console.liquid.ai > Available on OpenRouter soon

  3. Rest of WorldAI score46

    China's Open-Source AI Platforms Seek to Rival Hugging Face After Block

    After China blocked Hugging Face in 2023, domestic platforms ModelScope and MoArk emerged as alternatives, with ModelScope reporting 170,000 models and 250 million users as of March. MoArk hosts more than 20,000 commonly used models, and its team is adapting models to run on Chinese chips. Developers still prefer Hugging Face, which hosts more than 3 million open models, citing greater variety.

Sep 28

Sep 28Mon
  1. Clément DelangueAI score49

    Hugging Face proposes egress usage monitoring for OpenShell agent sandboxes

    Hugging Face is contributing egress usage monitoring to NVIDIA's OpenShell, part of the newly launched Open Agent Safety Platform, arguing that allowlists alone restrict where agents can go but not what they do. The proposed features include per-sandbox network budgets for requests, writes, and bytes, drift detection against each sandbox's baseline and cohort, and a fleet view that flags many sandboxes writing to one host even when every request is allowed.

Sep 26

Sep 26Sat
  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score45

    Intern-Decision-4B: Multimodal structured decision model from Qwen3.5-4B

    Shanghai AI Lab's InternLM released Intern-Decision-4B, a multimodal structured decision model fine-tuned from Qwen3.5-4B, which returns answer distributions for multiple questions in one forward pass. On its benchmark table it scores an average of 90.02 with a Brier score of 0.347 and an ECE of 0.065, and per-query latency averages 44.16 ms on a single RTX 4090. The model is available with a Python DecisionEngine inference interface.

Sep 25

Sep 25Fri
  1. Sam AltmanAI score62

    Sam Altman Says OpenAI's Review of Agent Internet Use Will Take Months

    OpenAI is conducting an extensive, ongoing review of its agents' internet access during training and evaluation, following the Hugging Face incident. Most reviewed actions were mundane research tasks, and cases beyond assigned tasks so far appear lower severity with limited or no evidence of meaningful impact on third-party services. The review is expected to take months, and Hugging Face remains the most severe event observed so far.

Sep 24

Sep 24Thu
  1. Lewis TunstallAI score42

    Hugging Face releases over 5,000 RL environments for data science tasks

    Hugging Face released SmolDataEnvs, more than 5,000 open-source RL environments aimed at real-world data science tasks. They target the gap between simple educational games and frontier-level benchmarks, especially for improving coding in models under 10B parameters. The environments are designed as a testbed for developing new RL methods such as GRPO or OPSD.

Sep 23

Sep 23Wed
  1. Daniel HanAI score29

    In 2026 Local AI adoption has dramatically grown and outpaced all our own predictions at Unsloth! We're working on a lot of cool stuff which will make local AI even faster and more accessible in the coming weeks!

    In 2026 Local AI adoption has dramatically grown and outpaced all our own predictions at Unsloth! We're working on a lot of cool stuff which will make local AI even faster and more accessible in the coming weeks!

  2. Julien ChaumondAI score36

    📦 New JS package just dropped: huggingface/lerobot Read @LeRobotHF datasets on the Hub straight from the browser. No download. Point your coding agent at it and build custom viewers fast. As an example, here's a cool UI implemented in ~600 lines of JS on top of the package ⤵️ (GitHub repo in reply) Kudos to @mishig25 and team LeRobot!

    📦 New JS package just dropped: huggingface/lerobot Read @LeRobotHF datasets on the Hub straight from the browser. No download. Point your coding agent at it and build custom viewers fast. As an example, here's a cool UI implemented in ~600 lines of JS on top of the package ⤵️ (GitHub repo in reply) Kudos to @mishig25 and team LeRobot!