StepFun: Step 3.5 Flash
StepFun: Step 3.5 Flash is designed for handling text inputs with a substantial context length of 262,144 tokens and supports tools and reasoning, making it suitable for complex tasks that require extensive data processing. The model does not support structured output, which means its outputs will be in plain text. This model might be a good choice for those who need to work with large volumes of text and value the ability to use external tools and engage in advanced reasoning processes. With a blended benchmark score of 42.4 across one independent benchmark, it performs well but is not at the top tier. Given its pricing at $0.1 per million input tokens and $0.3 per million output tokens, it offers a balanced cost for users who require robust text processing capabilities.
Benchmark results
Independent, published benchmarks. Blended score 42.4 across 1 benchmark, last refreshed 2026-07-23. How scoring works →
| Benchmark | Measures | Score |
|---|---|---|
| AI Index | broad capability composite | 42.9 |
- Model ID
- stepfun/step-3.5-flash
- Vendor
- stepfun
- Released
- January 2026
- Tokenizer
- Other
- Input Modalities
- text
- Output Modalities
- text
- Max Output
- 65,536 tokens
- Tool Calling
- ✓ supported
- Structured Output
- not supported
- Reasoning Mode
- ✓ supported
- Vision
- text only
- Audio
- no
- Moderated
- no
What it costs in practice
Computed from the current $0.10/M input and $0.30/M output rates. Run your own numbers →
| Job | Tokens | Cost |
|---|---|---|
| Summarize a 50-page report | 30k in / 1.5k out | under $0.01 |
| Classify 1,000 customer emails | 500k in / 50k out | $0.07 |
| A month of a busy support chatbot | 5M in / 2M out | $1.10 |
Similar models
Quick answers
- How much does StepFun: Step 3.5 Flash cost?
- $0.10 per million input tokens and $0.30 per million output tokens.
- What is StepFun: Step 3.5 Flash's context window?
- 262,144 tokens, roughly 393 pages of text in a single request.
- Does StepFun: Step 3.5 Flash support tool calling?
- It supports tool calling, a reasoning mode.
- Can StepFun: Step 3.5 Flash process images?
- No, it is text-only on the input side.