Mistral Large 4 is arriving in two stages. The company opened a public API preview on October 6, while scheduling the model weights for the end of the month. That timing makes the first practical question about access: what can a team evaluate now, and what depends on the later release?

The official announcement describes a multimodal model intended for coding, tool use, and document work. Ars Technica's fresh report brings the preview into the current news cycle. The useful distinction is between running a workload through a provider today and controlling the model's deployment after its artifacts become available.

A preview is an evaluation opportunity

An API lets a developer submit inputs to a model operated by a service. The provider manages the servers and exposes the supported interface. For a team considering Large 4, that is a route to examine responses on its own representative tasks before committing to a deployment design.

Consider a document workflow that extracts a part number from a drawing, finds the corresponding specification, and prepares a table for a reviewer. A persuasive answer is only one part of success. The table must preserve the right identifiers, each claim must point to the relevant evidence, and missing information must remain missing. A model that writes fluently can still attach a correct specification to the wrong part.

That example suggests a concrete comparison: give candidate systems the same inputs, tools, and completion conditions. Record whether each reaches the right result, how much intervention it needs, and whether the reviewer can reconstruct its reasoning from retained evidence. These are proposed evaluation criteria, rather than measured findings about Large 4.

The preview also supplies a useful point of separation between model quality and application design. Retrieval, document parsing, tool permissions, and the presentation of citations affect the final result. Keeping those components steady makes a model comparison easier to interpret. Changing the whole application at once can obscure which component produced an improvement.

Weights change who operates the system

Model weights are the learned numerical parameters used during inference. Having access to them can let an organization operate a model within a deployment it controls, subject to the released license and the practical requirements of running it. That differs from sending each request to a provider's hosted service.

Mistral's planned release remains a future milestone. The preview announcement does not itself supply the final downloadable artifacts or establish their license. Those details determine what a later deployment can actually do, including permitted uses and distribution. Calling the current preview an available open-source release would collapse several different decisions into one phrase.

For an operator, local control also creates work. Someone must provision capacity, manage the inference service, apply access controls, and preserve the version used by an application. A weight release is therefore an additional deployment option, with an operating responsibility attached. Whether that option is useful depends on the workload and the organization's ability to maintain it.

Total size and active computation answer different questions

Mistral's announcement describes roughly one trillion parameters, with 49 billion active. Its Large 4 documentation identifies the mixture-of-experts architecture. In a mixture of experts , a routing mechanism selects among learned components. The researchers' foundational paper explains this conditional computation. The complete parameter set and the components used for an input are different quantities.

The active count helps describe computation within that architecture. The total count still matters when considering storage and how the model is distributed across hardware. Neither number alone supplies a complete serving specification. Numerical representation, the inference implementation, and the workload all influence the eventual system.

For a simple illustration of scale, one trillion values stored at four bits per value would occupy about 500 billion bytes before additional overhead. That is arithmetic, not a claim that Large 4 ships in that representation or that such a system needs exactly that amount of memory. It shows why a smaller active count cannot, by itself, establish an easy single-device deployment.

Read the benchmark figure as a particular comparison

The featured illustration reproduces Mistral's DeepSWE 1.1 comparison. Its orange bar represents Large 4 Preview and rounds the reported result to 62. The figure's footnote says the scores were privately evaluated by Artificial Analysis ahead of the harness's public launch. It is a company-published presentation of a particular coding evaluation, rather than a general statement that every application will perform at that level.

An evaluation harness includes more than the underlying model. Its tasks, tools, execution conditions, and scoring rules shape the outcome. The figure also starts its vertical scale at 40, so bar heights need to be read with the printed values. The relevant next question is how the same configuration behaves on the work a team actually needs completed.

Large 4's immediate contribution is a new preview that can enter those comparisons. The later weight release could add a different operating arrangement. Keeping those stages distinct gives readers a clearer view of the announcement: evaluate the available service now, then assess the released artifacts, license, and deployment requirements when they arrive.