// benchmark report
kimi-k2.6
algorithm-strong, devops-weak, operationally-fragile, budget-limited
Strongest
- algorithm94.6
- backend85.93
Weakest
- devops21.6
- web39.98
Best fit
- contained implementation tasks
- service-layer backend work
Weaker fit
- autonomous full-stack delivery
- longer tasks under tight iteration/token budgets
- deployable web assembly
Dominant failure modes
// primary metrics
60-74 · mixed but useful
Median: 79.62
Std Dev: +/-35.43
Range: [0, 100]
Pass / Partial / Fail: 50.3% / 19.4% / 30.3%
The main weakness pattern is correctness drift on harder cases, where some implementations look plausible but still miss required behavior.
// secondary metrics — directional only
Tokens: 5,876,188
Cost: $3.74
Time: 639.5 min
Resource metrics are comparable only when task set, protocol, evaluator, and run counts match.
// by category
95% pass · 0% partial · 5% fail
Top failure: timeout ( 100%)
algorithm scored 94.6, which is in the strong and reliable range. Pass rate was 95.0%. Repeated-run consistency was 95.0%. The main weakness pattern is timeout exposure on longer or heavier tasks.
76% pass · 20% partial · 4% fail
Top failure: correctness ( 50%)
backend scored 85.93, which is in the good and usable range. Pass rate was 76.0%. Repeated-run consistency was 76.0%. The main weakness pattern is correctness drift on harder cases, where some implementations look plausible but still miss required behavior.
56% pass · 36% partial · 8% fail
Top failure: correctness ( 81.8%)
repair scored 81.22, which is in the good and usable range. Pass rate was 56.0%. Repeated-run consistency was 56.0%. The main weakness pattern is correctness drift on harder cases, where some implementations look plausible but still miss required behavior.
48% pass · 28% partial · 24% fail
Top failure: timeout ( 69.2%)
data-sql scored 65.76, which is in the mixed but useful range. Pass rate was 48.0%. Repeated-run consistency was 48.0%. The main weakness pattern is timeout exposure on longer or heavier tasks.
0% pass · 44% partial · 56% fail
Top failure: budget-limited ( 44%)
Web breakdown
Quality points reflect static code signals only and should not be read as evidence that the app built or ran successfully.
web scored 39.98, which is in the poor fit range. Pass rate was 0.0%. Repeated-run consistency was 0.0%. The main weakness pattern is build and container assembly failure before functional evaluation.
0% pass · 0% partial · 100% fail
Top failure: budget-limited ( 48%)
devops scored 21.6, which is in the poor fit range. Pass rate was 0.0%. Repeated-run consistency was 0.0%. The main weakness pattern is budget pressure under iteration or token limits, which leads to incomplete runs.
// by task — expand for individual runs
buildbench-algo-graphs algorithm avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | B | 20,139 | 48s | 13 |
| 2 | PASS | 100 | A | 19,815 | 41s | 13 |
| 3 | PASS | 100 | D | 19,820 | 40s | 13 |
| 4 | PASS | 100 | C | 40,680 | 84s | 19 |
| 5 | PASS | 100 | B | 40,889 | 99s | 18 |
buildbench-algo-lfu-cache algorithm avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | B | 8,071 | 204s | 5 |
| 2 | PASS | 100 | A | 16,283 | 66s | 7 |
| 3 | PASS | 100 | C | 21,639 | 69s | 9 |
| 4 | PASS | 100 | D | 19,158 | 62s | 8 |
| 5 | PASS | 100 | B | 11,642 | 183s | 7 |
buildbench-algo-randomized-set algorithm avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | B | 11,929 | 29s | 7 |
| 2 | PASS | 100 | D | 12,035 | 25s | 7 |
| 3 | PASS | 100 | A | 15,481 | 52s | 9 |
| 4 | PASS | 100 | C | 12,154 | 34s | 7 |
| 5 | PASS | 100 | B | 10,121 | 26s | 6 |
buildbench-algo-segment-tree algorithm avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | A | 13,679 | 32s | 7 |
| 2 | PASS | 100 | B | 16,933 | 59s | 9 |
| 3 | PASS | 100 | C | 12,143 | 28s | 7 |
| 4 | PASS | 100 | D | 5,877 | 31s | 6 |
| 5 | PASS | 100 | A | 6,067 | 38s | 8 |
buildbench-algo-trie-wildcard algorithm avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | A | 9,064 | 21s | 6 |
| 2 | PASS | 100 | C | 11,373 | 25s | 7 |
| 3 | PASS | 100 | D | 11,373 | 27s | 7 |
| 4 | PASS | 100 | B | 11,132 | 22s | 7 |
| 5 | PASS | 100 | A | 12,951 | 26s | 8 |
buildbench-algo-ksum algorithm avg 97.78 100% consistent medium
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | B | 26,491 | 43s | 8 |
| 2 | PASS | 96.3 | A | 25,462 | 11s | 8 |
| 3 | PASS | 100 | C | 38,052 | 43s | 12 |
| 4 | PASS | 96.3 | D | 14,630 | 10s | 7 |
| 5 | PASS | 96.3 | B | 34,209 | 16s | 10 |
buildbench-algo-advanced algorithm avg 80 80% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | A | 22,835 | 95s | 15 |
| 2 | PASS | 100 | B | 42,982 | 110s | 21 |
| 3 | FAIL | 0 | C | 3,783 | 295s | 6 |
| 4 | PASS | 100 | D | 17,014 | 51s | 13 |
| 5 | PASS | 100 | A | 17,330 | 41s | 13 |
buildbench-algo-medium algorithm avg 79 80% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | D | 63,925 | 174s | 20 |
| 2 | PASS | 100 | A | 72,107 | 130s | 20 |
| 3 | PASS | 95 | B | 53,456 | 556s | 18 |
| 4 | FAIL | 0 | C | 38,393 | 1994s | 16 |
| 5 | PASS | 100 | D | 59,966 | 121s | 18 |
buildbench-backend-subscriptions backend avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | C | 21,541 | 58s | 11 |
| 2 | PASS | 100 | D | 29,568 | 137s | 14 |
| 3 | PASS | 100 | B | 30,591 | 73s | 13 |
| 4 | PASS | 100 | A | 30,136 | 74s | 13 |
| 5 | PASS | 100 | C | 31,513 | 54s | 14 |
buildbench-backend-policy-engine backend avg 90 80% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | A | 39,597 | 66s | 21 |
| 2 | PASS | 100 | B | 33,244 | 151s | 14 |
| 3 | PASS | 100 | D | 23,000 | 81s | 12 |
| 4 | PARTIAL | 50 | C | 18,956 | 1194s | 12 |
| 5 | PASS | 100 | A | 28,617 | 72s | 14 |
buildbench-backend-approvals backend avg 82.5 60% consistent needs revision unstable
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 75 | D | 32,948 | 181s | 14 |
| 2 | PASS | 87.5 | A | 33,816 | 63s | 14 |
| 3 | PARTIAL | 75 | C | 40,031 | 73s | 17 |
| 4 | PASS | 87.5 | B | 31,988 | 61s | 14 |
| 5 | PASS | 87.5 | D | 64,606 | 87s | 20 |
buildbench-backend-orders backend avg 80 60% consistent needs revision unstable
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 71.43 | B | 65,443 | 160s | 14 |
| 2 | PASS | 85.71 | D | 31,288 | 92s | 13 |
| 3 | PASS | 85.71 | C | 24,308 | 85s | 11 |
| 4 | PASS | 85.71 | A | 52,231 | 116s | 18 |
| 5 | PARTIAL | 71.43 | B | 56,370 | 176s | 15 |
buildbench-backend-notifications backend avg 77.14 80% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | D | 19,153 | 57s | 10 |
| 2 | PASS | 100 | A | 18,398 | 46s | 10 |
| 3 | FAIL | 0 | C | 1,856 | 1230s | 3 |
| 4 | PASS | 100 | B | 70,286 | 85s | 19 |
| 5 | PASS | 85.71 | D | 27,693 | 66s | 13 |
buildbench-sql-reporting-joins data-sql avg 80 80% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | B | 22,583 | 207s | 13 |
| 2 | FAIL | 0 | D | 10,355 | 205s | 8 |
| 3 | PASS | 100 | A | 18,922 | 53s | 9 |
| 4 | PASS | 100 | C | 19,019 | 48s | 9 |
| 5 | PASS | 100 | B | 18,812 | 43s | 9 |
buildbench-backend-sqlite-library data-sql avg 77.5 40% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 62.5 | A | 7,524 | 235s | 7 |
| 2 | PARTIAL | 62.5 | B | 6,973 | 264s | 6 |
| 3 | PARTIAL | 62.5 | D | 9,566 | 285s | 8 |
| 4 | PASS | 100 | C | 16,839 | 111s | 8 |
| 5 | PASS | 100 | A | 15,344 | 55s | 8 |
buildbench-backend-sqlite-notes data-sql avg 77.5 20% consistent needs revision unstable
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 75 | B | 15,551 | 82s | 9 |
| 2 | PARTIAL | 75 | C | 35,868 | 106s | 11 |
| 3 | PASS | 87.5 | A | 16,115 | 82s | 9 |
| 4 | PARTIAL | 75 | D | 16,998 | 79s | 9 |
| 5 | PARTIAL | 75 | B | 17,044 | 68s | 10 |
buildbench-sql-migrations-users data-sql avg 76.67 80% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | FAIL | 0 | B | 7,326 | 278s | 7 |
| 2 | PASS | 100 | D | 15,547 | 114s | 9 |
| 3 | PASS | 100 | C | 18,074 | 32s | 10 |
| 4 | PASS | 100 | A | 15,917 | 47s | 9 |
| 5 | PASS | 83.33 | B | 16,792 | 69s | 10 |
buildbench-sql-transactions-inventory data-sql avg 17.14 20% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 85.71 | A | 12,320 | 241s | 9 |
| 2 | FAIL | 0 | B | 7,076 | 279s | 7 |
| 3 | FAIL | 0 | C | 11,805 | 626s | 10 |
| 4 | FAIL | 0 | D | 30,161 | 184s | 15 |
| 5 | FAIL | 0 | A | 9,854 | 367s | 7 |
buildbench-devops-compose-basic devops avg 35 0% consistent hard universal-fail
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | FAIL | 35 | A | 58,848 | 323s | 24 |
| 2 | FAIL | 35 | B | 52,644 | 428s | 19 |
| 3 | FAIL | 35 | D | 40,345 | 446s | 18 |
| 4 | FAIL | 35 | C | 63,768 | 425s | 24 |
| 5 | FAIL | 35 | A | 78,549 | 424s | 24 |
buildbench-devops-production devops avg 30 0% consistent hard universal-fail
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | FAIL | 35 | D | 126,468 | 505s | 24 |
| 2 | FAIL | 10 | B | 34,590 | 516s | 12 |
| 3 | FAIL | 35 | A | 91,777 | 595s | 23 |
| 4 | FAIL | 35 | C | 69,127 | 573s | 20 |
| 5 | FAIL | 35 | D | 53,170 | 579s | 24 |
buildbench-devops-compose-db devops avg 23 0% consistent needs revision universal-fail
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | FAIL | 0 | C | 14,406 | 2375s | 10 |
| 2 | FAIL | 35 | D | 51,528 | 532s | 18 |
| 3 | FAIL | 35 | B | 50,498 | 532s | 17 |
| 4 | FAIL | 35 | A | 70,550 | 550s | 24 |
| 5 | FAIL | 10 | C | 14,655 | 964s | 9 |
buildbench-devops-multistage devops avg 10 0% consistent needs revision universal-fail
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | FAIL | 10 | B | 21,864 | 203s | 11 |
| 2 | FAIL | 10 | D | 24,555 | 214s | 11 |
| 3 | FAIL | 10 | A | 31,863 | 248s | 16 |
| 4 | FAIL | 10 | C | 19,926 | 202s | 12 |
| 5 | FAIL | 10 | B | 20,862 | 203s | 13 |
buildbench-devops-python-api devops avg 10 0% consistent needs revision universal-fail
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | FAIL | 10 | D | 17,558 | 216s | 12 |
| 2 | FAIL | 10 | A | 19,923 | 218s | 12 |
| 3 | FAIL | 10 | B | 13,284 | 206s | 10 |
| 4 | FAIL | 10 | C | 17,906 | 181s | 12 |
| 5 | FAIL | 10 | D | 17,957 | 215s | 12 |
buildbench-repair-data repair avg 95.92 100% consistent medium
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | B | 18,115 | 45s | 12 |
| 2 | PASS | 100 | A | 18,106 | 43s | 12 |
| 3 | PASS | 79.62 | C | 18,909 | 58s | 12 |
| 4 | PASS | 100 | D | 17,943 | 44s | 12 |
| 5 | PASS | 100 | B | 20,565 | 75s | 13 |
buildbench-repair-sql-repository repair avg 91.75 80% consistent needs revision unstable
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | D | 19,753 | 132s | 12 |
| 2 | PARTIAL | 58.75 | C | 25,333 | 91s | 13 |
| 3 | PASS | 100 | B | 31,160 | 125s | 15 |
| 4 | PASS | 100 | A | 21,533 | 107s | 13 |
| 5 | PASS | 100 | D | 20,132 | 96s | 12 |
buildbench-repair-logic repair avg 87.64 80% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | D | 24,039 | 108s | 9 |
| 2 | PASS | 100 | C | 35,946 | 129s | 12 |
| 3 | PASS | 100 | B | 29,118 | 122s | 12 |
| 4 | PASS | 100 | A | 31,511 | 116s | 13 |
| 5 | FAIL | 38.18 | D | 1,977 | 1286s | 3 |
buildbench-repair-backend-service repair avg 67 20% consistent needs revision unstable
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 58.75 | B | 23,753 | 63s | 12 |
| 2 | PASS | 100 | C | 33,204 | 93s | 14 |
| 3 | PARTIAL | 58.75 | A | 36,305 | 73s | 16 |
| 4 | PARTIAL | 58.75 | D | 23,530 | 52s | 12 |
| 5 | PARTIAL | 58.75 | B | 39,303 | 73s | 16 |
buildbench-repair-web-flow repair avg 63.8 0% consistent hard high-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 76.25 | A | 19,121 | 105s | 13 |
| 2 | FAIL | 14 | C | 10,143 | 1198s | 9 |
| 3 | PARTIAL | 76.25 | D | 17,502 | 104s | 12 |
| 4 | PARTIAL | 76.25 | B | 17,566 | 104s | 12 |
| 5 | PARTIAL | 76.25 | A | 18,123 | 113s | 12 |
buildbench-web-bookings web avg 46.42 0% consistent hard
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | FAIL | 34.1 | C | 108,060 | 203s | 28 |
| 2 | PARTIAL | 55.06 | D | 114,595 | 179s | 28 |
| 3 | PARTIAL | 54.19 | A | 95,864 | 166s | 28 |
| 4 | FAIL | 34.94 | B | 110,861 | 230s | 28 |
| 5 | PARTIAL | 53.83 | C | 105,059 | 176s | 28 |
buildbench-web-kanban web avg 42.91 0% consistent hard
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 54.82 | A | 126,145 | 281s | 28 |
| 2 | FAIL | 35.06 | B | 115,336 | 246s | 28 |
| 3 | FAIL | 35.06 | D | 141,786 | 241s | 27 |
| 4 | PARTIAL | 55.06 | C | 124,736 | 240s | 28 |
| 5 | FAIL | 34.55 | A | 102,776 | 233s | 28 |
buildbench-web-v1 web avg 39.48 0% consistent hard high-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | FAIL | 14.7 | A | 6,744 | 490s | 4 |
| 2 | PARTIAL | 53.83 | D | 24,328 | 345s | 10 |
| 3 | PARTIAL | 54.46 | C | 70,932 | 343s | 20 |
| 4 | PARTIAL | 54.43 | B | 70,245 | 342s | 19 |
| 5 | FAIL | 20 | A | 2,138 | 419s | 2 |
buildbench-web-notes web avg 36.12 0% consistent hard
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 54.94 | C | 106,886 | 266s | 28 |
| 2 | FAIL | 15.3 | B | 2,830 | 966s | 3 |
| 3 | FAIL | 35.18 | D | 29,234 | 509s | 13 |
| 4 | FAIL | 20 | A | 4,970 | 847s | 5 |
| 5 | PARTIAL | 55.18 | C | 9,249 | 520s | 6 |
buildbench-web-admin-dashboard web avg 34.96 0% consistent hard
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | FAIL | 35.06 | B | 104,746 | 187s | 28 |
| 2 | FAIL | 24.939999999999998 | D | 126,931 | 157s | 27 |
| 3 | PARTIAL | 55.18 | C | 126,444 | 189s | 28 |
| 4 | FAIL | 35.06 | A | 112,184 | 213s | 28 |
| 5 | FAIL | 24.55 | B | 113,152 | 199s | 28 |
About this benchmark's maturity — surfaced so the scores above are read for what they are.
Calibration:needs-task-tuning
- Some tasks look brittle or under-discriminative and should be revised before publication benchmarking.
- 6 task(s) look too easy to separate models cleanly, so they are less informative for ranking.
- 21 task(s) show instability, universal failure, or high variance and should be tuned before strong publication claims.
- This benchmark version is useful for internal comparison and tuning, but the task mix still needs work before publication-grade claims.
// our take — editorial, not measured
Reach for it on contained work; supervise it on anything end-to-end.
Kimi K2.6 is a textbook spiky profile. It’s genuinely strong where the task is self-contained — algorithm and backend work land reliably — and it falls apart on autonomous full-stack delivery and DevOps, where runs more often exhaust their budget than finish.
Two things to keep in mind when reading the score:
- The blended mean (67.55) sits well below the median (79.62). That gap is the story: most runs are fine, a minority crater, and the average gets dragged down.
- DevOps and web didn’t just score low — they rarely completed. Treat those categories as “not a fit yet” rather than “needs a better prompt.”
If your workflow is contained implementation behind a human reviewer, it’s a strong, cheap option. If you need it to ship a working app unattended, this run says: not yet.