Skip to content
View original post on X: Kala· 34/100AI score34/100

Mistral Large 4 and Reflection Beam promise open weights this month

AISummary

Mistral Large 4 and Reflection Beam are previewed now, with Mistral saying weights drop at the end of October and Reflection promising Apache 2.0 weights this month.

The post argues that these announced future weights should be treated as a conditional migration dependency, not a current self-hosting option.

API previews can be trialed immediately, but they do not prove an unreleased checkpoint will behave the same when downloaded.

Post on XView on X
@kalapowered

Mistral Large 4 and Reflection Beam promise weights this month. We set out what the previews can prove, what to check in the released files, and when a self-hosted migration can proceed. https://x.com/i/article/2107895463051448320

An open model launch should come with the weights

You need to choose a model for a system that has to run on infrastructure you control. A colleague sends you a link to a launch post for a new model that they’re convinced will be a good fit. You skim the post and include it in your migration plan, noting that the weights will be downloadable later this month.

What do you think? Would you treat this as a real plan or as a conditional plan? More broadly, should you treat the announcement of future open weights as a current self-hosting option or a future intent? In this post, we argue for the latter. An API preview can earn a trial immediately. A self-hosted migration still depends on the released files, their terms and your evaluation.

This post is motivated by the announcement this week of Mistral Large 4 and Reflection Beam , both of which have explicit announcements of when the weights will be released.

1. Record what the announcement makes available today

Mistral announced Mistral Large 4 on 6th October . The launch article invites readers to try the preview API on Mistral Studio and says “Weights drop end of this month.” The model doc page lists the model as both “Public Preview” and “Open” .

Reflection announced Beam on 5th October via a post on X which says “Full weights release this month.” The corresponding article page asks you to join a waitlist, says they are releasing an early version of the model for a select number of users and says “This month, we will release the weights under an Apache 2.0 license, along with documentation and the full stack for running, evaluating, and fine-tuning the model.”

What do we think about this? We think that both Mistral and Reflection are doing the right thing by being transparent about when you will be able to download and use the weights. However, we think that it’s a mistake to remove the future tense on these announcements.

If you’re thinking about purchasing either of these models, then we think the right way to take notes is something like: Put the access date and the weight-release date in separate fields, with each source and the checked-at date. Record the promised October release window as a dependency for migration, not a completed prerequisite.

We also think that even once the weights are available, it’s important to keep the distinction between a model whose weights are available and a model that you have evaluated for your needs on your infrastructure.

2. Let the preview answer a question it can answer

Is there value in trying an API preview of a model that will later be open? Yes! If the model is fixing a problem for you then you can learn from a trial. Refusing to use the model until all deployment options are available is a waste of time.

However, we think you need to be very clear about what you’re evaluating when you do this. It’s important to note that you’re evaluating the API version of the model and not the version of the model that you might run locally if the checkpoint you want to use isn’t released. The API trial cannot demonstrate your ability to operate an unreleased checkpoint, recover it from your own storage or maintain a particular local configuration.

Further, you should be aware that the model you’re evaluating may not be the same as the model you download later. In the case of Mistral Large 4, the announcement says “The reinforcement learning run behind this preview is still in flight”. It describes further improvements as expected work. Ask for the relationship between the tested endpoint and the later download when the release arrives.

If you’re trialling the API version of the model, we recommend saving the model ID, the date, all the prompts you sent, the tools you sent (if applicable), the settings you used and the model response. Then you can compute a score for the model on these prompts and re-run the evaluation later when the model is available for download. It’s also important to keep a separate entry for every change you make. For example, if you add, remove or modify a tool, then you should make a new entry in your trial log.

Finally, when the model is released, we recommend re-running this evaluation on the released version of the model to make sure that it is still a good fit for your needs. This is not a comment on either Mistral Large 4 or Reflection Beam, but rather a recommendation for how to evaluate models in general. Include the failures as well as successful cases, and keep the preview result attached to the system that produced it.

3. Read the licence that comes with the release

A final note on licences: it’s worth keeping track of the licence that the model weights will be released under. In the case of Reflection Beam, the article says that the weights will be released under the Apache 2.0 licence. In the case of Mistral Large 4, the launch article just says that the weights will be released but doesn’t mention the licence. It’s important to keep track of these at the appropriate level of fidelity and not conflate them. Do not infer the Large 4 licence from a different Mistral model, or treat Reflection’s promised terms as an already inspected release file.

First: please be sure to contact the person responsible for licensing at your business before deploying any model. They will want to look at the actual license, and discuss with the actual use cases before making a decision. You want to make sure that the person responsible for licensing has the opportunity to make an informed decision, in the context of your business. Part of making an informed decision is having a proposal in hand to review. Keep a record of the decision, in context of the license and release. Specify whether the proposal covers internal inference, a customer service, redistribution of a modified model, or a combination. Save the reviewed licence text and release identifier with the decision.

Second: check separately what you mean by “open source”. The Open Source Initiative’s Open Source AI Definition 1.0 specifies freedoms to use, study, modify and share a system. Its preferred form for making modifications includes data information, code and parameters. Note that merely having the ability to download weights is not enough to establish that a release meets that definition.

Third: you don’t need to resolve any disagreement over names, but you do need to put your requirements in writing, and proceed on that basis. It is possible that you can find a model that is useful for your purposes, even if it doesn’t fit some name. The released terms must permit the particular use your business proposes.

Fourth: If you are making a decision about a license, you should record that decision, and connect it to the decision to deploy. If the deployment requires clarification from the publisher, it should be considered “pending” until you get clarification.

4. A completed launch can supply more than a promise

What is your response to the fact that Mistral did release Mistral 3 on 2 December 2025 ? You can find details about the release at Mistral 3 announcement from 2 December 2025 . The announcement says the models are available that day through named services and Hugging Face collections. It names Apache 2.0 and describes both base and instruction-fine-tuned versions of Mistral Large 3.

The page also contains an example: Mistral says its NVFP4 checkpoint can run with vLLM on Blackwell NVL72 systems and on a single node with eight A100 or eight H100 GPUs. That is the publisher’s documented configuration, not a test we ran.

The availability of Large 3 may not be the reason you want to buy Large 4, but it should be noted that the statement of release for 3 is different from the statement of release for 4.

In the case of a forthcoming release, it would not be unreasonable for the buyer to request a “starting point” from the vendor. In other words, you want the vendor to show you where you can get the release, what you need to download, and how to run it. And if the vendor has published a deployment example, you can ask the vendor to tell you how they set up that example. If you need a compressed version, identify that version in your own evaluation rather than treating files from the same model family as interchangeable.

In general, it is reasonable for a buyer to expect that a release will include a “starting point”, rather than requiring the buyer to guess what to do. Using the vendor’s setup as a “starting point” for experimentation is a good idea, and it makes sense to constrain the scope of the experiment to make it possible to explain the results. Record where your setup differs, including the runtime and numerical format.

5. Retain enough of the release to recover your deployment

Once the release is available for download, the buyer should keep a record of how to retrieve the release. This record should include the location of the release (official source), the version/revision of the release, the list of files that were downloaded, the list of additional files (tokenizer, configuration, etc.) needed to run the model (according to the vendor’s instructions), the version of the runtime needed to run the model, and the command/configuration needed to run the model.

It would be reasonable to enhance the record with file hashes. There is a standard library in Python for this, hashlib , and the documentation includes a section about sha256 , and a section about a helper for computing file digests , and a digest comparison can show whether a file matches a recorded digest . Note that file hashes can help with identification, but they are not sufficient to establish provenance. You will want to keep other evidence to establish provenance. A matching digest also does not, on its own, establish that the model is trustworthy.

You will want to check that you can start the model from your files. You can do this by using a disposable environment, and starting the model from your files using your configuration. For a workflow that must operate without an external model service, test with that service unavailable. Document any other network dependencies the job still needs rather than calling the entire application offline because inference ran locally.

Finally, you will want to have a second engineer start the model from your files, to make sure that you have not missed anything. Record any missing dependency, correct the instructions and repeat the affected step.

Why do this? Because it’s possible that the model file you end up running isn’t the one that’s released. Maybe you or your team convert it to some other format that’s more appropriate for your use. Maybe you make some other changes to the model. At the very least, you should probably store the hash of the model file you’re running, in case of discrepancies between the release and your running copy. For a conversion, retain the source version and conversion settings with the resulting hashes. Distinguish your converted copy from the publisher’s release.

Keep the working release until its replacement passes your acceptance checks. (Obviously you’ll want to make sure that the files are appropriately secured and that you have a deletion policy.) The point is that you need to be able to recover the accepted version when you need it again, so that you can easily reference it later. It’s not good enough to have bookmarked the model’s page, because that may change over time.

6. Test the job that justified self-hosting

This may seem like a lot of work just to switch models, but it’s important to remember that you’re only doing this for a specific workflow. You should have a specific workflow in mind that you want to migrate, or else there’s no point in doing this. Make sure that you’re able to explain the metrics you’ll be using to accept or reject the model. If you’re using the model to extract data from invoices, you need to be able to explain what it means to “extract data from an invoice into your schema correctly”, and what to do about cases where it’s unclear (probably send it to a reviewer). If you’re using it to write code, you need to be able to explain what it means for a patch to be “good” (probably something like “passes the tests for the repository and is accepted after review”).

You should have a fixed set of inputs that your organisation can lawfully use for the evaluation. Obviously these inputs should be representative of the types of inputs you expect to use in the real workflow. You should also know of cases where the current workflow fails, and make sure those are included in the test inputs. You should also have rules for how you’re going to accept or reject a given output, and those rules should be included with the inputs. If the rules require human judgement, you should make sure that the reasons for the judgement are recorded.

You should also make sure to measure the time it takes to run the workflow, as well as the failure rate, with the level of parallelism you expect to use in the real workflow. You should make sure you understand how much hardware you need, and whether it’s cheaper to rent or buy. You should also make sure you understand how much work it is to run the model yourself, vs how much it would cost to have someone else do it. The Beam announcement describes an approximate compute comparison that excludes prompt prefill, context-dependent attention operations and serving overhead. It does not report measured inference cost.

It’s quite possible that a smaller released model meets your needs while a larger preview is still unavailable for the intended deployment. It’s also possible that the larger model doesn’t get released for your intended use for a while, so you end up using the API. The important thing is that you shouldn’t pre-decide what you’re going to do.

What metrics am i going to use? How many errors am i willing to tolerate? How fast does it need to respond? Do i need it to respond quickly, or is it ok to have a batch process that runs overnight? Etc. Write those limits before looking at the new model’s outputs. (Note: these are questions you should answer, not questions that have been answered for Mistral/Beam.)

Approve (and record who approved) the workflow and the configuration you measured. Until you’ve done this, the model should be considered unapproved for your intended use. Name the person who can stop the rollout if it fails. If no evaluation has run, record that the candidate is unevaluated for your deployment.

7. Make future availability an explicit condition

If you’re making a decision this week, you should make sure to separate out the things you can do now from the things you can’t do until the weights are released in October. You can start the process of trialling the preview version of the model, but you should keep decisions about migrating to self-hosted conditional until the weights are actually released. Give the preview trial a budget and an owner, and keep migration conditional on the released files, reviewed terms and acceptance result. Do not retire the working service because a future download has a date.

If you’re leading infrastructure for your company, write the acceptance record and make sure you have a machine to run the model on. If you’re leading procurement, make sure you’re not confusing what you can get now with what you might be able to get later. If you’re the one releasing the model, add links to the model release and license when they’re available, and update the post if the schedule changes.

We’re very happy that models are being released in a usable form, and Mistral and Reflection have given buyers an October release window to watch and previews to consider. The purchasing decision should change when the promised release can be inspected and tested. Until then, keep the self-hosting work conditional and give the available preview credit for the work it can actually do.

Sources

1. Mistral, Introducing Mistral Large 4, 6 October 2026. Checked 6 October 2026.

1. Mistral, Mistral Large 4 documentation, dated 6 October 2026. Checked 6 October 2026.

1. Reflection, Beam announcement, 5 October 2026. Checked 6 October 2026.

1. Reflection, Introducing Beam, checked 6 October 2026.

1. Open Source Initiative, The Open Source AI Definition 1.0, checked 6 October 2026.

1. Mistral, Introducing Mistral 3, 2 December 2025. Checked 6 October 2026.

1. Python documentation, hashlib, checked 6 October 2026.

Source: Kala · x.comPublished · added here