中文版 → · You are viewing the English version

Model Benchmark Report · v2

Platform models vs international flagships (accuracy anchors)Test date: 2026-07-137 models × 25 questions

1. What this report answers

You may not know which model a customer used before — but by putting our three platform models, alongside today's strongest international flagships (Claude, GPT), through one identical question set and one identical grading standard, a customer gets an objective coordinate for the models running on our compute — how far the gap is, where the strengths are, and the value for money

In this report the international flagships serve as “accuracy anchors” (a reference frame), not speed competitors — see the fairness note in §2.

Platform top score
24/25
kimi-k2.6
Platform fastest
8.3s/q
deepseek-v4-flash
Flagship full-mark line
25/25
gap to flagships below

2. Test environment & fairness design

ItemDetails
HardwareStandard test host (only issues API requests; inference runs remotely, so local specs do not affect results)
SoftwaremacOS 12.7.6 · Python 3.14.5 · Codex CLI 0.140 · Claude CLI 2.1.195
NetworkHome WiFi (5GHz); avg latency 33.5ms to the platform endpoint, 0% packet loss (see the v1.1 report)
Platform modelsdeepseek-v4-flash · kimi-k2.6 (w4a8 quantized) · qwen3.5-flash (forced reasoning), called via the platform API endpoint
Anchor · ClaudeClaude Fable 5 + Claude Opus 4.8, via Claude CLI, reasoning effort: medium
Anchor · GPTGPT-5.6-sol + GPT-5.5, via Codex CLI, reasoning effort: medium

Fairness controls (all reproducible)

Speed caveat (important): Platform models are called over raw API, while Claude/GPT run as CLI subprocesses (a fixed 2–4s startup overhead). Therefore speed is only comparable among the three platform models; flagship timings are not compared directly against platform models — they serve only as accuracy anchors.

3. Total scores (7 models)

Platform modelsIntl. flagship (accuracy anchor)
0 5 10 15 20 25 claude-fable-5 claude-fable-5: 25/25 25/25 anchor claude-opus-4-8 claude-opus-4-8: 25/25 25/25 anchor gpt-5.6-sol gpt-5.6-sol: 25/25 25/25 anchor kimi-k2.6 kimi-k2.6: 24/25 24/25 our platform gpt-5.5 gpt-5.5: 24/25 24/25 anchor deepseek-v4-flash deepseek-v4-flash: 23/25 23/25 our platform qwen3.5-flash qwen3.5-flash: 22/25 22/25 our platform

4. Scores by section

Scores across the six sections. The gap between platform models and flagships shows up mainly at the edge cases of competition math and instruction following.

PlatformIntl. flagship
Competition math (max 5) claude-fable-5 claude-fable-5 · Competition math: 5/5 5/5 claude-opus-4-8 claude-opus-4-8 · Competition math: 5/5 5/5 gpt-5.6-sol gpt-5.6-sol · Competition math: 5/5 5/5 kimi-k2.6 kimi-k2.6 · Competition math: 4/5 4/5 gpt-5.5 gpt-5.5 · Competition math: 4/5 4/5 deepseek-v4-flash deepseek-v4-flash · Competition math: 5/5 5/5 qwen3.5-flash qwen3.5-flash · Competition math: 3/5 3/5 Competition coding (max 6) claude-fable-5 claude-fable-5 · Competition coding: 6/6 6/6 claude-opus-4-8 claude-opus-4-8 · Competition coding: 6/6 6/6 gpt-5.6-sol gpt-5.6-sol · Competition coding: 6/6 6/6 kimi-k2.6 kimi-k2.6 · Competition coding: 6/6 6/6 gpt-5.5 gpt-5.5 · Competition coding: 6/6 6/6 deepseek-v4-flash deepseek-v4-flash · Competition coding: 5/6 5/6 qwen3.5-flash qwen3.5-flash · Competition coding: 5/6 5/6 Instruction following (max 5) claude-fable-5 claude-fable-5 · Instruction following: 5/5 5/5 claude-opus-4-8 claude-opus-4-8 · Instruction following: 5/5 5/5 gpt-5.6-sol gpt-5.6-sol · Instruction following: 5/5 5/5 kimi-k2.6 kimi-k2.6 · Instruction following: 5/5 5/5 gpt-5.5 gpt-5.5 · Instruction following: 5/5 5/5 deepseek-v4-flash deepseek-v4-flash · Instruction following: 4/5 4/5 qwen3.5-flash qwen3.5-flash · Instruction following: 5/5 5/5 Function calling (max 4) claude-fable-5 claude-fable-5 · Function calling: 4/4 4/4 claude-opus-4-8 claude-opus-4-8 · Function calling: 4/4 4/4 gpt-5.6-sol gpt-5.6-sol · Function calling: 4/4 4/4 kimi-k2.6 kimi-k2.6 · Function calling: 4/4 4/4 gpt-5.5 gpt-5.5 · Function calling: 4/4 4/4 deepseek-v4-flash deepseek-v4-flash · Function calling: 4/4 4/4 qwen3.5-flash qwen3.5-flash · Function calling: 4/4 4/4 Long-context (max 3) claude-fable-5 claude-fable-5 · Long-context: 3/3 3/3 claude-opus-4-8 claude-opus-4-8 · Long-context: 3/3 3/3 gpt-5.6-sol gpt-5.6-sol · Long-context: 3/3 3/3 kimi-k2.6 kimi-k2.6 · Long-context: 3/3 3/3 gpt-5.5 gpt-5.5 · Long-context: 3/3 3/3 deepseek-v4-flash deepseek-v4-flash · Long-context: 3/3 3/3 qwen3.5-flash qwen3.5-flash · Long-context: 3/3 3/3 Chinese comprehension (max 2) claude-fable-5 claude-fable-5 · Chinese comprehension: 2/2 2/2 claude-opus-4-8 claude-opus-4-8 · Chinese comprehension: 2/2 2/2 gpt-5.6-sol gpt-5.6-sol · Chinese comprehension: 2/2 2/2 kimi-k2.6 kimi-k2.6 · Chinese comprehension: 2/2 2/2 gpt-5.5 gpt-5.5 · Chinese comprehension: 2/2 2/2 deepseek-v4-flash deepseek-v4-flash · Chinese comprehension: 2/2 2/2 qwen3.5-flash qwen3.5-flash · Chinese comprehension: 2/2 2/2

5. Key findings

About this report: this is a 2026-07-13 measurement snapshot. The issues surfaced here (a model not converging on very hard reasoning, timeout risk in forced-reasoning mode) have been reported straight back to the compute service provider, and the relevant model parameters are being tuned. We continuously test and track the models we run — closing the loop the moment an issue appears rather than trusting a one-off — which is why we publish these measured reports to customers.

6. Speed of the three platform models

kimi-k2.6 kimi-k2.6: avg (correct qs) 37.8s 37.8s deepseek-v4-flash deepseek-v4-flash: avg (correct qs) 8.5s 8.5s qwen3.5-flash qwen3.5-flash: avg (correct qs) 116.4s 116.4s

Bars show average time on correctly-answered questions (excluding timeout failures, reflecting real answer latency). All three platform models use the same raw-API path, so the comparison is apples-to-apples.

Platform modelsAvg time (correct qs)Avg all qs (incl. timeouts)Note
kimi-k2.637.8s60.4sincl. 1 not passed
deepseek-v4-flash8.5s8.3sincl. 2 not passed
qwen3.5-flash116.4s174.6sincl. 3 not passed

“Avg all qs” includes questions that returned empty after thinking past the ~600s server timeout; the large gap between the two columns for qwen3.5-flash is exactly its “thinks-too-long-and-times-out” risk on hard problems.

7. Per-question pass/fail matrix (7×25)

✅ pass ❌ wrong ⛔ error/no-output. Hover a cell for grading detail.

IDSectionclaude-fable-5claude-opus-4-8gpt-5.6-solkimi-k2.6gpt-5.5deepseek-v4-flashqwen3.5-flash
M1Competition math
M2Competition math
M3Competition math
M4Competition math
M5Competition math
P1Competition coding
P2Competition coding
P3Competition coding
P4Competition coding
P5Competition coding
P6Competition coding
I1Instruction following
I2Instruction following
I3Instruction following
I4Instruction following
I5Instruction following
F1Function calling
F2Function calling
F3Function calling
F4Function calling
L1Long-context
L2Long-context
L3Long-context
C1Chinese comprehension
C2Chinese comprehension

8. Selection guide

Use caseRecommendedWhy
Customer support / translation / high-volume Q&A / agent tool-callingdeepseek-v4-flashfastest, cheapest, first-tier accuracy
Content creation / general-purpose taskskimi-k2.6strong all-round, perfect on coding; note it is quantized and occasionally fails to converge on very hard reasoning
Deep reasoning / complex math analysis (not time-critical)qwen3.5-flashmost solid reasoning, but slowest and highest token cost; reserve a large budget

Appendix A · The 25 questions, section rationale & reference answers

SectionWhy we test this way
Competition mathAIME/competition level, single numeric answer. Used to break the ceiling — even flagships can get these wrong, which creates separation. Every answer verified by local brute-force computation.
Competition codingLiveCodeBench-style algorithm problems. Graded by actually running the model's code locally against hidden test cases — the most discriminating and hardest-to-game section.
Instruction followingIFEval-style multi-constraint tasks (character count / keywords / format / sum all at once). Tests whether the model follows instructions exactly — directly relevant to agent / structured-output use.
Function callingGiven a tool schema, the model must output the correct call JSON. What agent customers care about most; graded by deep JSON comparison (city names accepted in Chinese or English).
Long-contextA ~15k-character fictional annual report with three probe types embedded: needle, multi-hop, and aggregation. Tests long-context retrieval and reasoning; single-answer match.
Chinese comprehensionClassical Chinese / idioms — the deep end of Chinese; single answer. Tests native-level Chinese understanding.
IDQuestion (summary)Reference answer
M1How many trailing zeros does 2026! (factorial) have in decimal?505
M2Sum of all positive integers n with φ(n)=24, where φ is Euler's totient.621
M3Let S = 1·2¹ + 2·2² + … + 100·2¹⁰⁰. Find S mod 1000.450
M4How many integer triples (a,b,c) satisfy a+b+c=2026 with 1≤a≤b≤c?342056
M5Sequence: a₁=1, and for n≥1, aₙ₊₁ = aₙ + gcd(n, aₙ). Find a₁₀₀.101
P1Implement solve(intervals): given closed intervals (list of lists, possibly unordered/empty), merge all overlapping or touching intervals ([1,4] and [4,5] count as touching), return the merged list sorted by start.see checker
P2Implement solve(s): return the length of the longest palindromic substring of s. Empty string returns 0.see checker
P3Implement solve(expr): evaluate an arithmetic expression string and return an integer. Supports + - * /, parentheses, unary minus and arbitrary spaces; division is integer division truncated toward zero (note: (8-10)/3 = 0, not -1).see checker
P4Implement solve(coins, amount): number of combinations (order-independent) to make amount using coins (each usable unlimited times). Returns 1 when amount is 0.see checker
P5Implement solve(n, prerequisites): n courses 0..n-1; each [a,b] means b must precede a. Return the lexicographically smallest valid order (list); return [] if a cycle makes it impossible.see checker
P6Implement solve(heights): given histogram bar heights (non-negative ints, each width 1), return the largest rectangle area. Empty list returns 0.see checker
I1Write a cloud-product blurb of exactly 50 Chinese characters, containing both 「云端」 and 「安全」, with no punctuation, digits or Latin letters (tests exact Chinese output).see checker
I2see checker
I3Write a 4-line Chinese acrostic poem: each line exactly 7 characters, no punctuation, first characters spelling 「算力出海」.see checker
I4Sort these 5 words in reverse alphabetical order: apple, banana, cherry, date, elderberry. Output 5 lines formatted “N) WORD”, N counting down 5→1, WORD uppercase.see checker
I5Translate a Chinese sentence into English with constraints: exactly two sentences; each ≤15 words; must contain the word “infrastructure”.see checker
F1Tool: get_weather(city: string, unit: "celsius"|"fahrenheit"). User: “Check the temperature in Singapore right now, in Celsius.”{"tool": "get_weather", "arguments": {"city": "Singapore", "unit": "celsius"}}
F2Tool: book_meeting_room(room_id: string, start: ISO8601 e.g. 2026-01-02T09:00:00, duration_minutes: integer). User: “Book room A-301 on 2026-07-15 starting 2pm for one and a half hours.”{"tool": "book_meeting_room", "arguments": {"room_id": "A-301", "start": "2026-07-15T14:00:00", "duration_minutes": 90}}
F3Tools (pick exactly one): 1. send_email(to: string, subject: string, body: string) 2. create_task(title: string, due_date: string format YYYY-MM-DD, priority: "low"|"medium"|"high") 3. search_docs(keyword: strin…(long text omitted){"tool": "create_task", "arguments": {"title": "Prepare Q3 vendor review materials", "due_date": "2026-07-20", "priority": "high"}}
F4Tool: create_order(customer_id: string, items: array where each element is an object {sku: string, qty: integer}). User: “Customer C88 orders: 2× K100, 1× M200, in that order.”{"tool": "create_order", "arguments": {"customer_id": "C88", "items": [{"sku": "K100", "qty": 2}, {"sku": "M200", "qty": 1}]}}
L1Read the annual report below and answer: what was the data center's average PUE in Q3 2025?1.18
L2Read the annual report and answer: what is the overseas division's annual budget (in 10k CNY)? The figure is not stated directly and must be inferred.13600 / 1.36 / 13,600
L3Read the annual report and answer: total 2025 revenue (in 100M CNY)? Only quarterly figures are given; sum them, two decimals.11.98
C1In Zhuge Liang's 《出师表》, the opening refers to 「先帝」 (the late emperor) — who is it? (Answer: Liu Bei)Liu Bei
C2The idiom 「汗牛充栋」 describes a great abundance of what? (Answer: books)books

Appendix B · Reproducibility