StepFun Step 5 Preview: 600B at $1 per Million Tokens
StepFun opened API access to Step 5 Preview on 20 September, a 600B-parameter sparse mixture-of-experts model priced at a dollar per million input tokens. The interesting number is not 600 billion, it is 27 billion.
StepFun announced Step 5 Preview on 20 September 2026 and opened API access the same day. The specification that gets quoted is 600 billion parameters. The specification that explains the price is 27 billion, which is roughly how many of those parameters actually run on any given token. At $1.00 per million input tokens on a cache miss and $2.70 per million output tokens, it is priced like a mid-size model because, in the only sense that costs money, it behaves like one.
The Numbers, and Which Ones Matter
Specification | Step 5 Preview |
|---|---|
Total parameters | 600B, sparse mixture-of-experts |
Active per token | About 27B, roughly a 4.5% activation ratio |
Context window | 1M tokens |
Input / output | Text and images in, text out |
Input price | $1.00 per million tokens, $0.05 on a cache hit |
Output price | $2.70 per million tokens, reasoning tokens billed as output |
Open weights | Promised for 15 October, API-only until then |
RuntimeWire and AI Weekly both report the same pricing and architecture, and DataStudios records the 15 October open-weights date. Artificial Analysis places the model at 44 on its Intelligence Index, which the launch coverage frames as matching Kimi K3 Max. Treat a single index score with the usual caution: it is a composite, and a benchmark number can be perfectly accurate and still misleading about your workload.
Why 600B Can Cost a Dollar
A dense model runs every parameter for every token. A sparse mixture-of-experts model holds many specialised sub-networks and routes each token to a small subset of them, so the compute per token tracks the active parameters rather than the total. Step 5 Preview stores 600 billion and uses about 27 billion at a time. The 600 billion figure describes how much knowledge the model can hold; the 27 billion figure describes what you are paying to run.
This is why comparing models by total parameter count across architectures tells you very little about price or latency. If the distinction is new, the background is in what mixture of experts means. The short version: the headline number is a capacity claim, not a cost signal, and it is one of the specifics worth checking on any release you are trying to keep up with rather than taking the top-line figure at face value.
The cache pricing is the other detail worth noting. A dollar per million input tokens drops to five cents on a cache hit, a twentyfold difference. Any workload with a large stable prefix, a long system prompt, a document you ask several questions about, will land far closer to the cache price than the list price in practice. That is a general property of how prompt caching works rather than anything specific to StepFun, but the ratio here is steeper than most.
Preview Means Preview
The label is doing real work. Step 5 Preview is API-only until 15 October, when StepFun says full open weights follow. Until then you are building against an endpoint whose behaviour the vendor can change, on a tier that carries no stability commitment. Reasoning tokens are billed as output, which means a model with extended thinking can produce a bill substantially larger than your visible response length suggests, and a change to its default reasoning behaviour moves your costs without moving your code.
None of that is a reason to avoid it. It is a reason to know what preview and GA labels actually commit a vendor to before a preview endpoint ends up on a critical path.
Who This Is Actually For
Long-context work where the input dominates. A 1M window at a dollar per million input tokens, less with caching, changes what is affordable to feed a model in one go.
Teams who want the option to self-host later. If the 15 October open-weights release lands, evaluating on the API now and moving in-house afterwards is a coherent plan. If it slips, you are on a preview endpoint.
Cost-sensitive high-volume pipelines, where the gap between this and a frontier model per million tokens compounds into a real line item.
It is a poor fit if you need stable behaviour under an SLA today, or if your workload is output-heavy, where the $2.70 rate and reasoning-token billing do more of the work than the input price. As always, the comparison that matters is pricing across providers on your own traffic shape, not the list rate.
Frequently Asked Questions
How much does Step 5 Preview cost?
On StepFun's own API, $1.00 per million input tokens on a cache miss, $0.05 per million on a cache hit, and $2.70 per million output tokens. Reasoning tokens count as output.
Is Step 5 Preview open weights?
Not yet. StepFun says full open weights follow on 15 October 2026. Until that date the model is available only through the API.
What does 600B parameters with 27B active mean?
It is a sparse mixture-of-experts design. The model stores 600 billion parameters but routes each token to a subset totalling about 27 billion, so compute and price track the smaller number.
How big is the context window?
One million tokens, with text and image input and text output.
Should I switch my production workload to it?
Not on the strength of a launch and an index score. Run your own evaluation on your own prompts, and be deliberate about putting a preview-tier endpoint anywhere a behaviour change would hurt.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


