head-to-head

Qwen: Qwen3.7 Flash vs StepFun: Step 3.7 Flash

Side-by-side comparison of specs, pricing, benchmark scores, and task rankings. Updated 2026-09-23.

Qwen: Qwen3.7 Flash StepFun: Step 3.7 Flash
Vendorqwenstepfun
Quality Score100100
Benchmark Score-29.1
Input Price$0.03/M$0.20/M
Output Price$0.13/M$1.15/M
Context Window1,000,000262,144
Max Output65,536230,400
Tool Calling
Structured Output
Reasoning Mode
Vision
Audio--
Benchmark Scores
ai_index-32.1

Who wins by task?

TaskQwen: Qwen3.7 FlashStepFun: Step 3.7 Flash
SQL Generation 134 135
Code Review 132 132
Code Completion 132 129
Code Refactoring 136 132
Bug Fixing 136 136
Unit Test Generation 124 125
Code Documentation 132 129
Regex Writing 120 122
CI/CD Pipelines 120 121
Frontend Component Design 122 125
Data Analysis 124 128
CSV / Spreadsheet Cleanup 134 129
ETL Scripting 128 127
JSON Extraction 132 131
Bulk Data Labeling 130 129
OCR / Document Parsing 131 129
Table Extraction from PDFs 131 129
Long-Document Summarization 138 134
Short-Form Summarization 124 125
Blog Post Writing 122 123

Scores reflect capability match + benchmark data + pricing for each task. Methodology →

Related comparisons

Anthropic: Claude Sonnet 5 vs StepFun: Step 3.7 Flash ByteDance Seed: Seed-2.0-Code vs Qwen: Qwen3.7 Flash ByteDance Seed: Seed-2.0-Code vs StepFun: Step 3.7 Flash ByteDance Seed: Seed 2.1 Turbo vs Qwen: Qwen3.7 Flash ByteDance Seed: Seed 2.1 Turbo vs StepFun: Step 3.7 Flash DeepSeek: DeepSeek Flash Latest vs Qwen: Qwen3.7 Flash DeepSeek: DeepSeek V4.1 Flash vs Qwen: Qwen3.7 Flash DeepSeek: DeepSeek V4 Flash Latest vs Qwen: Qwen3.7 Flash

The Model Movers Report

One email every Friday, built from this site's own rankings: the current top five by benchmark score, every model released in the last seven days, and one note worked out from that week's numbers. You can unsubscribe from any issue with one click.