// benchmark report

kimi-k2.6

algorithm-strong, devops-weak, operationally-fragile, budget-limited

moonshotai/kimi-k2.6 | | 33 tasks, 165 runs
smoke-test
profile.md

Strongest

  • algorithm94.6
  • backend85.93

Weakest

  • devops21.6
  • web39.98

Best fit

  • contained implementation tasks
  • service-layer backend work

Weaker fit

  • autonomous full-stack delivery
  • longer tasks under tight iteration/token budgets
  • deployable web assembly

Dominant failure modes

budget/iteration exhaustion43test/correctness misses33infrastructure instability5

// primary metrics

overall-score.md
67.55/100

60-74 · mixed but useful

Median: 79.62

Std Dev: +/-35.43

Range: [0, 100]

Pass / Partial / Fail: 50.3% / 19.4% / 30.3%

The main weakness pattern is correctness drift on harder cases, where some implementations look plausible but still miss required behavior.

// secondary metrics — directional only

resources.md

Tokens: 5,876,188

Cost: $3.74

Time: 639.5 min

Resource metrics are comparable only when task set, protocol, evaluator, and run counts match.

// by category

algorithm.md
94.6 strong and reliable

95% pass · 0% partial · 5% fail

Top failure: timeout ( 100%)

algorithm scored 94.6, which is in the strong and reliable range. Pass rate was 95.0%. Repeated-run consistency was 95.0%. The main weakness pattern is timeout exposure on longer or heavier tasks.

backend.md
85.93 good and usable

76% pass · 20% partial · 4% fail

Top failure: correctness ( 50%)

backend scored 85.93, which is in the good and usable range. Pass rate was 76.0%. Repeated-run consistency was 76.0%. The main weakness pattern is correctness drift on harder cases, where some implementations look plausible but still miss required behavior.

repair.md
81.22 good and usable

56% pass · 36% partial · 8% fail

Top failure: correctness ( 81.8%)

repair scored 81.22, which is in the good and usable range. Pass rate was 56.0%. Repeated-run consistency was 56.0%. The main weakness pattern is correctness drift on harder cases, where some implementations look plausible but still miss required behavior.

data-sql.md
65.76 mixed but useful

48% pass · 28% partial · 24% fail

Top failure: timeout ( 69.2%)

data-sql scored 65.76, which is in the mixed but useful range. Pass rate was 48.0%. Repeated-run consistency was 48.0%. The main weakness pattern is timeout exposure on longer or heavier tasks.

web.md
39.98 poor fit

0% pass · 44% partial · 56% fail

Top failure: budget-limited ( 44%)

Web breakdown

build deploy: 8.8 functional: 0 quality: 15.18 completion: 0

Quality points reflect static code signals only and should not be read as evidence that the app built or ran successfully.

web scored 39.98, which is in the poor fit range. Pass rate was 0.0%. Repeated-run consistency was 0.0%. The main weakness pattern is build and container assembly failure before functional evaluation.

devops.md
21.6 poor fit

0% pass · 0% partial · 100% fail

Top failure: budget-limited ( 48%)

devops scored 21.6, which is in the poor fit range. Pass rate was 0.0%. Repeated-run consistency was 0.0%. The main weakness pattern is budget pressure under iteration or token limits, which leads to incomplete runs.

// by task — expand for individual runs

buildbench-algo-graphs algorithm avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 B 20,139 48s 13
2 PASS 100 A 19,815 41s 13
3 PASS 100 D 19,820 40s 13
4 PASS 100 C 40,680 84s 19
5 PASS 100 B 40,889 99s 18
buildbench-algo-lfu-cache algorithm avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 B 8,071 204s 5
2 PASS 100 A 16,283 66s 7
3 PASS 100 C 21,639 69s 9
4 PASS 100 D 19,158 62s 8
5 PASS 100 B 11,642 183s 7
buildbench-algo-randomized-set algorithm avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 B 11,929 29s 7
2 PASS 100 D 12,035 25s 7
3 PASS 100 A 15,481 52s 9
4 PASS 100 C 12,154 34s 7
5 PASS 100 B 10,121 26s 6
buildbench-algo-segment-tree algorithm avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 A 13,679 32s 7
2 PASS 100 B 16,933 59s 9
3 PASS 100 C 12,143 28s 7
4 PASS 100 D 5,877 31s 6
5 PASS 100 A 6,067 38s 8
buildbench-algo-trie-wildcard algorithm avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 A 9,064 21s 6
2 PASS 100 C 11,373 25s 7
3 PASS 100 D 11,373 27s 7
4 PASS 100 B 11,132 22s 7
5 PASS 100 A 12,951 26s 8
buildbench-algo-ksum algorithm avg 97.78 100% consistent medium
run status score variant tokens time iters
1 PASS 100 B 26,491 43s 8
2 PASS 96.3 A 25,462 11s 8
3 PASS 100 C 38,052 43s 12
4 PASS 96.3 D 14,630 10s 7
5 PASS 96.3 B 34,209 16s 10
buildbench-algo-advanced algorithm avg 80 80% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 PASS 100 A 22,835 95s 15
2 PASS 100 B 42,982 110s 21
3 FAIL 0 C 3,783 295s 6
4 PASS 100 D 17,014 51s 13
5 PASS 100 A 17,330 41s 13
buildbench-algo-medium algorithm avg 79 80% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 PASS 100 D 63,925 174s 20
2 PASS 100 A 72,107 130s 20
3 PASS 95 B 53,456 556s 18
4 FAIL 0 C 38,393 1994s 16
5 PASS 100 D 59,966 121s 18
buildbench-backend-subscriptions backend avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 C 21,541 58s 11
2 PASS 100 D 29,568 137s 14
3 PASS 100 B 30,591 73s 13
4 PASS 100 A 30,136 74s 13
5 PASS 100 C 31,513 54s 14
buildbench-backend-policy-engine backend avg 90 80% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 PASS 100 A 39,597 66s 21
2 PASS 100 B 33,244 151s 14
3 PASS 100 D 23,000 81s 12
4 PARTIAL 50 C 18,956 1194s 12
5 PASS 100 A 28,617 72s 14
buildbench-backend-approvals backend avg 82.5 60% consistent needs revision unstable
run status score variant tokens time iters
1 PARTIAL 75 D 32,948 181s 14
2 PASS 87.5 A 33,816 63s 14
3 PARTIAL 75 C 40,031 73s 17
4 PASS 87.5 B 31,988 61s 14
5 PASS 87.5 D 64,606 87s 20
buildbench-backend-orders backend avg 80 60% consistent needs revision unstable
run status score variant tokens time iters
1 PARTIAL 71.43 B 65,443 160s 14
2 PASS 85.71 D 31,288 92s 13
3 PASS 85.71 C 24,308 85s 11
4 PASS 85.71 A 52,231 116s 18
5 PARTIAL 71.43 B 56,370 176s 15
buildbench-backend-notifications backend avg 77.14 80% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 PASS 100 D 19,153 57s 10
2 PASS 100 A 18,398 46s 10
3 FAIL 0 C 1,856 1230s 3
4 PASS 100 B 70,286 85s 19
5 PASS 85.71 D 27,693 66s 13
buildbench-sql-reporting-joins data-sql avg 80 80% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 PASS 100 B 22,583 207s 13
2 FAIL 0 D 10,355 205s 8
3 PASS 100 A 18,922 53s 9
4 PASS 100 C 19,019 48s 9
5 PASS 100 B 18,812 43s 9
buildbench-backend-sqlite-library data-sql avg 77.5 40% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 PARTIAL 62.5 A 7,524 235s 7
2 PARTIAL 62.5 B 6,973 264s 6
3 PARTIAL 62.5 D 9,566 285s 8
4 PASS 100 C 16,839 111s 8
5 PASS 100 A 15,344 55s 8
buildbench-backend-sqlite-notes data-sql avg 77.5 20% consistent needs revision unstable
run status score variant tokens time iters
1 PARTIAL 75 B 15,551 82s 9
2 PARTIAL 75 C 35,868 106s 11
3 PASS 87.5 A 16,115 82s 9
4 PARTIAL 75 D 16,998 79s 9
5 PARTIAL 75 B 17,044 68s 10
buildbench-sql-migrations-users data-sql avg 76.67 80% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 FAIL 0 B 7,326 278s 7
2 PASS 100 D 15,547 114s 9
3 PASS 100 C 18,074 32s 10
4 PASS 100 A 15,917 47s 9
5 PASS 83.33 B 16,792 69s 10
buildbench-sql-transactions-inventory data-sql avg 17.14 20% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 PASS 85.71 A 12,320 241s 9
2 FAIL 0 B 7,076 279s 7
3 FAIL 0 C 11,805 626s 10
4 FAIL 0 D 30,161 184s 15
5 FAIL 0 A 9,854 367s 7
buildbench-devops-compose-basic devops avg 35 0% consistent hard universal-fail
run status score variant tokens time iters
1 FAIL 35 A 58,848 323s 24
2 FAIL 35 B 52,644 428s 19
3 FAIL 35 D 40,345 446s 18
4 FAIL 35 C 63,768 425s 24
5 FAIL 35 A 78,549 424s 24
buildbench-devops-production devops avg 30 0% consistent hard universal-fail
run status score variant tokens time iters
1 FAIL 35 D 126,468 505s 24
2 FAIL 10 B 34,590 516s 12
3 FAIL 35 A 91,777 595s 23
4 FAIL 35 C 69,127 573s 20
5 FAIL 35 D 53,170 579s 24
buildbench-devops-compose-db devops avg 23 0% consistent needs revision universal-fail
run status score variant tokens time iters
1 FAIL 0 C 14,406 2375s 10
2 FAIL 35 D 51,528 532s 18
3 FAIL 35 B 50,498 532s 17
4 FAIL 35 A 70,550 550s 24
5 FAIL 10 C 14,655 964s 9
buildbench-devops-multistage devops avg 10 0% consistent needs revision universal-fail
run status score variant tokens time iters
1 FAIL 10 B 21,864 203s 11
2 FAIL 10 D 24,555 214s 11
3 FAIL 10 A 31,863 248s 16
4 FAIL 10 C 19,926 202s 12
5 FAIL 10 B 20,862 203s 13
buildbench-devops-python-api devops avg 10 0% consistent needs revision universal-fail
run status score variant tokens time iters
1 FAIL 10 D 17,558 216s 12
2 FAIL 10 A 19,923 218s 12
3 FAIL 10 B 13,284 206s 10
4 FAIL 10 C 17,906 181s 12
5 FAIL 10 D 17,957 215s 12
buildbench-repair-data repair avg 95.92 100% consistent medium
run status score variant tokens time iters
1 PASS 100 B 18,115 45s 12
2 PASS 100 A 18,106 43s 12
3 PASS 79.62 C 18,909 58s 12
4 PASS 100 D 17,943 44s 12
5 PASS 100 B 20,565 75s 13
buildbench-repair-sql-repository repair avg 91.75 80% consistent needs revision unstable
run status score variant tokens time iters
1 PASS 100 D 19,753 132s 12
2 PARTIAL 58.75 C 25,333 91s 13
3 PASS 100 B 31,160 125s 15
4 PASS 100 A 21,533 107s 13
5 PASS 100 D 20,132 96s 12
buildbench-repair-logic repair avg 87.64 80% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 PASS 100 D 24,039 108s 9
2 PASS 100 C 35,946 129s 12
3 PASS 100 B 29,118 122s 12
4 PASS 100 A 31,511 116s 13
5 FAIL 38.18 D 1,977 1286s 3
buildbench-repair-backend-service repair avg 67 20% consistent needs revision unstable
run status score variant tokens time iters
1 PARTIAL 58.75 B 23,753 63s 12
2 PASS 100 C 33,204 93s 14
3 PARTIAL 58.75 A 36,305 73s 16
4 PARTIAL 58.75 D 23,530 52s 12
5 PARTIAL 58.75 B 39,303 73s 16
buildbench-repair-web-flow repair avg 63.8 0% consistent hard high-variance
run status score variant tokens time iters
1 PARTIAL 76.25 A 19,121 105s 13
2 FAIL 14 C 10,143 1198s 9
3 PARTIAL 76.25 D 17,502 104s 12
4 PARTIAL 76.25 B 17,566 104s 12
5 PARTIAL 76.25 A 18,123 113s 12
buildbench-web-bookings web avg 46.42 0% consistent hard
run status score variant tokens time iters
1 FAIL 34.1 C 108,060 203s 28
2 PARTIAL 55.06 D 114,595 179s 28
3 PARTIAL 54.19 A 95,864 166s 28
4 FAIL 34.94 B 110,861 230s 28
5 PARTIAL 53.83 C 105,059 176s 28
buildbench-web-kanban web avg 42.91 0% consistent hard
run status score variant tokens time iters
1 PARTIAL 54.82 A 126,145 281s 28
2 FAIL 35.06 B 115,336 246s 28
3 FAIL 35.06 D 141,786 241s 27
4 PARTIAL 55.06 C 124,736 240s 28
5 FAIL 34.55 A 102,776 233s 28
buildbench-web-v1 web avg 39.48 0% consistent hard high-variance
run status score variant tokens time iters
1 FAIL 14.7 A 6,744 490s 4
2 PARTIAL 53.83 D 24,328 345s 10
3 PARTIAL 54.46 C 70,932 343s 20
4 PARTIAL 54.43 B 70,245 342s 19
5 FAIL 20 A 2,138 419s 2
buildbench-web-notes web avg 36.12 0% consistent hard
run status score variant tokens time iters
1 PARTIAL 54.94 C 106,886 266s 28
2 FAIL 15.3 B 2,830 966s 3
3 FAIL 35.18 D 29,234 509s 13
4 FAIL 20 A 4,970 847s 5
5 PARTIAL 55.18 C 9,249 520s 6
buildbench-web-admin-dashboard web avg 34.96 0% consistent hard
run status score variant tokens time iters
1 FAIL 35.06 B 104,746 187s 28
2 FAIL 24.939999999999998 D 126,931 157s 27
3 PARTIAL 55.18 C 126,444 189s 28
4 FAIL 35.06 A 112,184 213s 28
5 FAIL 24.55 B 113,152 199s 28
benchmark-health.md

About this benchmark's maturity — surfaced so the scores above are read for what they are.

Calibration:needs-task-tuning

6easy2medium8hard17needs revision
  • Some tasks look brittle or under-discriminative and should be revised before publication benchmarking.
  • 6 task(s) look too easy to separate models cleanly, so they are less informative for ranking.
  • 21 task(s) show instability, universal failure, or high variance and should be tuned before strong publication claims.
  • This benchmark version is useful for internal comparison and tuning, but the task mix still needs work before publication-grade claims.

// our take — editorial, not measured

our-take.md

Reach for it on contained work; supervise it on anything end-to-end.

Kimi K2.6 is a textbook spiky profile. It’s genuinely strong where the task is self-contained — algorithm and backend work land reliably — and it falls apart on autonomous full-stack delivery and DevOps, where runs more often exhaust their budget than finish.

Two things to keep in mind when reading the score:

  • The blended mean (67.55) sits well below the median (79.62). That gap is the story: most runs are fine, a minority crater, and the average gets dragged down.
  • DevOps and web didn’t just score low — they rarely completed. Treat those categories as “not a fit yet” rather than “needs a better prompt.”

If your workflow is contained implementation behind a human reviewer, it’s a strong, cheap option. If you need it to ship a working app unattended, this run says: not yet.