Summary

  • Mistral Large 4 entered public API preview on October 6. The company says weights are expected by month-end; as of October 8, customer-run deployment remains a future option.
  • The model card shows temporary sale rates of $0.68 per million input tokens and $2.09 per million output tokens, against standard prices of $1.36 and $4.18. A cheap trial can measure a buyer’s API bill, not the cost of owning the serving operation.
  • Customers should test a representative workload now, then compare accepted-task cost and control requirements with a self-run design only after files, licence and technical details are published.

Two clocks, one launch

Mistral’s October 6 announcement starts two clocks. One is already running: Mistral Large 4 is available through the company’s Studio API preview. The other is promised: model weights are expected by the end of October. Those are not interchangeable forms of access. In the first, Mistral operates the model and the customer buys inference. In the second, the customer may be able to operate the model itself, subject to the eventual release and its terms.

The sequence matters commercially. An enterprise can begin testing before it knows whether it will ever run Large 4 privately. That reduces the cost of learning whether the model helps with a real workflow. It also means the first practical comparison is between a metered service and the buyer’s current alternative, not between the API and an on-premises system that does not yet exist.

The temporary price is specific enough to make that first test concrete. Mistral’s model card currently displays sale rates of $0.68 per million uncached input tokens and $2.09 per million output tokens; the crossed-out standard rates are $1.36 and $4.18. Cached input is separately priced at $0.07 on sale and $0.14 at standard rates. Mistral’s changelog says the launch offer is 50% off for two weeks. For illustration, 10 million uncached input tokens and one million output tokens would cost $8.89 at the sale rates, before retries, tools, taxes or orchestration; at standard rates the same token mix is $17.78.

That is arithmetic, not a forecast.

The number that matters to a buyer is not dollars per million tokens in isolation. It is cost per accepted result under the buyer’s own prompt lengths, cache reuse, output size, tool calls, retries and human review. A task that needs an agent to run for an hour can consume a different token mix from a short classification request. Quality failures also change the denominator: a cheap attempt that must be repeated is not necessarily a cheap completed task.

A score is not a purchase case

The early evaluation record is mixed in a useful way. Vals AI’s October 7 profile places Large 4 at 33rd of 45 on its Vals Index, with 48.05% accuracy, while listing its best result as sixth of 76 on Harvey’s Legal Agent Benchmark. The index combines finance, coding, legal and tax tasks, weighted by approximate U.S. GDP shares. Vals describes that blend as a proxy and warns that some benchmark runs may use different providers and parameters. It is one structured lens, not a universal quality table.

That distinction argues for testing a buyer’s task, rather than accepting either the launch language or one aggregate rank. Mistral says the preview is still being refined and that further architecture, benchmark and post-training details will accompany the weights. Le Monde reported at launch that independent confirmation of the company’s performance claims was still pending. For now, the buyer can collect its own results through the API, but should label them as preview results from that service and configuration.

The option has no operating price yet

The official model card lists 52 billion active parameters, 1.05 trillion total parameters, a 1.6-billion-parameter vision encoder and a one-million-token context window. None of those figures says how much memory a customer needs to meet a chosen latency and concurrency target. Active parameters per token are not the same as the memory footprint of all model weights. Precision, quantisation, memory bandwidth, routing, context length and the serving stack all matter.

The value of future weights is therefore real but conditional. They may let a customer move inference to a private cloud or its own hardware, alter moderation settings, or retain a fallback if hosted access changes. They may also shift responsibility for patching, security, scaling, uptime and incident response onto that customer. The download itself would not make the model free to run.

A sound purchasing sequence is to use the preview to measure task quality, latency and token spend on a representative set. Preserve the prompts, model version, output checks, cache assumptions and retry counts. When the weights arrive, repeat the same tasks and compare the cost of accepted work, including infrastructure and people. Separately test the licence, hardware, security and service requirements. If the local option cannot meet those constraints, the customer has learned something before treating “open-weight” as a savings plan.

Mistral’s release is thus priced on one side and still unpriced on the other. The API can turn evaluation into a bill today. The weights may later move the control point, but the buyer cannot yet know whether that control is affordable or operationally preferable.

Sources: Mistral’s announcement; model card and pricing; product changelog; Vals model profile; Vals Index methodology; Le Monde launch coverage.