head-to-head

OpenAI: GPT-5.6 Luna Pro vs StepFun: Step 3.7 Flash

Side-by-side comparison of specs, pricing, benchmark scores, and task rankings. Updated 2026-09-23.

OpenAI: GPT-5.6 Luna Pro StepFun: Step 3.7 Flash
Vendoropenaistepfun
Quality Score100100
Benchmark Score-29.1
Input Price$0.20/M$0.20/M
Output Price$1.20/M$1.15/M
Context Window1,050,000262,144
Max Output128,000230,400
Tool Calling
Structured Output
Reasoning Mode
Vision
Audio--
Benchmark Scores
ai_index-32.1

Who wins by task?

TaskOpenAI: GPT-5.6 Luna ProStepFun: Step 3.7 Flash
SQL Generation 133 135
Code Review 132 132
Code Completion 131 129
Code Refactoring 136 132
Bug Fixing 136 136
Unit Test Generation 124 125
Code Documentation 131 129
Regex Writing 119 122
CI/CD Pipelines 120 121
Frontend Component Design 122 125
Data Analysis 124 128
CSV / Spreadsheet Cleanup 133 129
ETL Scripting 128 127
OCR / Document Parsing 131 129
Table Extraction from PDFs 131 129
Long-Document Summarization 137 134
Short-Form Summarization 123 125
Blog Post Writing 121 123

Scores reflect capability match + benchmark data + pricing for each task. Methodology →

Related comparisons

Anthropic: Claude Opus 5 (batch) vs OpenAI: GPT-5.6 Luna Pro Anthropic: Claude Sonnet 5 vs StepFun: Step 3.7 Flash ByteDance Seed: Seed-2.0-Code vs OpenAI: GPT-5.6 Luna Pro ByteDance Seed: Seed-2.0-Code vs StepFun: Step 3.7 Flash ByteDance Seed: Seed 2.1 Turbo vs OpenAI: GPT-5.6 Luna Pro ByteDance Seed: Seed 2.1 Turbo vs StepFun: Step 3.7 Flash DeepSeek: DeepSeek V4.1 Flash vs OpenAI: GPT-5.6 Luna Pro DeepSeek: DeepSeek V4 Flash Vision Exp vs OpenAI: GPT-5.6 Luna Pro

The Model Movers Report

One email every Friday, built from this site's own rankings: the current top five by benchmark score, every model released in the last seven days, and one note worked out from that week's numbers. You can unsubscribe from any issue with one click.