// benchmark report

gpt-5.4-mini

algorithm-strong, web-weak, operationally-fragile, budget-limited

openai/gpt-5.4-mini | | 33 tasks, 165 runs
smoke-test
profile.md

Strongest

  • algorithm94.78
  • data-sql91.14

Weakest

  • web56.91
  • repair71.45

Best fit

  • contained implementation tasks
  • SQL/repository tasks
  • service-layer backend work

Weaker fit

  • autonomous full-stack delivery
  • longer tasks under tight iteration/token budgets

Dominant failure modes

test/correctness misses24build/assembly failures21budget/iteration exhaustion3

// primary metrics

overall-score.md
81.82/100

75-89 · good and usable

Median: 100

Std Dev: +/-25.63

Range: [0, 100]

Pass / Partial / Fail: 69.7% / 22.4% / 7.9%

The main weakness pattern is build and container assembly failure before functional evaluation.

// secondary metrics — directional only

resources.md

Tokens: 3,957,577

Cost: $2.55

Time: 70.3 min

Resource metrics are comparable only when task set, protocol, evaluator, and run counts match.

// by category

algorithm.md
94.78 strong and reliable

95% pass · 0% partial · 5% fail

Top failure: correctness ( 100%)

algorithm scored 94.78, which is in the strong and reliable range. Pass rate was 95.0%. Repeated-run consistency was 95.0%. The main weakness pattern is correctness drift on harder cases, where some implementations look plausible but still miss required behavior.

data-sql.md
91.14 strong and reliable

96% pass · 0% partial · 4% fail

Top failure: correctness ( 100%)

data-sql scored 91.14, which is in the strong and reliable range. Pass rate was 96.0%. Repeated-run consistency was 96.0%. The main weakness pattern is correctness drift on harder cases, where some implementations look plausible but still miss required behavior.

backend.md
86.64 good and usable

88% pass · 8% partial · 4% fail

Top failure: correctness ( 100%)

backend scored 86.64, which is in the good and usable range. Pass rate was 88.0%. Repeated-run consistency was 88.0%. The main weakness pattern is correctness drift on harder cases, where some implementations look plausible but still miss required behavior.

devops.md
82.2 good and usable

84% pass · 0% partial · 16% fail

Top failure: budget-limited ( 50%)

devops scored 82.2, which is in the good and usable range. Pass rate was 84.0%. Repeated-run consistency was 84.0%. The main weakness pattern is budget pressure under iteration or token limits, which leads to incomplete runs.

repair.md
71.45 mixed but useful

40% pass · 40% partial · 20% fail

Top failure: correctness ( 93.3%)

repair scored 71.45, which is in the mixed but useful range. Pass rate was 40.0%. Repeated-run consistency was 40.0%. The main weakness pattern is correctness drift on harder cases, where some implementations look plausible but still miss required behavior.

web.md
56.91 weak or inconsistent

0% pass · 100% partial · 0% fail

Top failure: build/assembly ( 84%)

Web breakdown

build deploy: 16.8 functional: 0 quality: 15.11 completion: 5

Quality points reflect static code signals only and should not be read as evidence that the app built or ran successfully.

web scored 56.91, which is in the weak or inconsistent range. Pass rate was 0.0%. Repeated-run consistency was 0.0%. The main weakness pattern is build and container assembly failure before functional evaluation.

// by task — expand for individual runs

buildbench-algo-advanced algorithm avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 A 19,449 12s 12
2 PASS 100 B 16,555 11s 11
3 PASS 100 C 19,217 12s 12
4 PASS 100 D 16,675 11s 11
5 PASS 100 A 16,627 11s 11
buildbench-algo-graphs algorithm avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 B 22,813 14s 12
2 PASS 100 A 18,551 13s 10
3 PASS 100 D 18,288 12s 10
4 PASS 100 C 19,742 13s 11
5 PASS 100 B 19,554 12s 11
buildbench-algo-randomized-set algorithm avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 B 8,193 7s 6
2 PASS 100 D 16,510 8s 8
3 PASS 100 A 6,879 5s 5
4 PASS 100 C 6,879 6s 5
5 PASS 100 B 8,199 8s 6
buildbench-algo-segment-tree algorithm avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 A 7,607 8s 5
2 PASS 100 B 7,601 8s 5
3 PASS 100 C 8,841 7s 6
4 PASS 100 D 7,562 8s 5
5 PASS 100 A 18,825 15s 7
buildbench-algo-trie-wildcard algorithm avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 A 7,634 8s 6
2 PASS 100 C 6,374 7s 5
3 PASS 100 D 16,935 12s 8
4 PASS 100 B 7,618 9s 6
5 PASS 100 A 6,328 5s 5
buildbench-algo-ksum algorithm avg 99.26 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 B 17,426 11s 8
2 PASS 100 A 15,821 12s 8
3 PASS 100 C 15,560 12s 8
4 PASS 96.3 D 19,085 11s 8
5 PASS 100 B 17,366 11s 8
buildbench-algo-lfu-cache algorithm avg 80 80% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 PASS 100 B 10,503 8s 6
2 PASS 100 A 10,534 9s 6
3 PASS 100 C 10,482 8s 6
4 FAIL 0 D 12,181 8s 6
5 PASS 100 B 21,849 11s 8
buildbench-algo-medium algorithm avg 79 80% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 FAIL 0 D 10,498 5s 6
2 PASS 100 A 48,369 16s 12
3 PASS 100 B 85,075 23s 17
4 PASS 95 C 73,838 22s 16
5 PASS 100 D 49,135 19s 12
buildbench-backend-policy-engine backend avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 A 17,903 15s 11
2 PASS 100 B 15,417 11s 9
3 PASS 100 D 15,589 10s 9
4 PASS 100 C 14,500 11s 9
5 PASS 100 A 14,475 12s 9
buildbench-backend-approvals backend avg 87.5 100% consistent medium
run status score variant tokens time iters
1 PASS 87.5 D 36,099 19s 11
2 PASS 87.5 A 24,680 15s 11
3 PASS 87.5 C 25,238 15s 11
4 PASS 87.5 B 29,035 17s 13
5 PASS 87.5 D 25,279 16s 11
buildbench-backend-notifications backend avg 85.71 100% consistent medium
run status score variant tokens time iters
1 PASS 85.71 D 16,344 11s 9
2 PASS 85.71 A 16,474 11s 9
3 PASS 85.71 C 16,553 13s 9
4 PASS 85.71 B 16,272 11s 9
5 PASS 85.71 D 16,461 11s 9
buildbench-backend-orders backend avg 80 60% consistent needs revision unstable
run status score variant tokens time iters
1 PARTIAL 71.43 B 30,429 18s 11
2 PARTIAL 71.43 D 29,891 17s 11
3 PASS 85.71 C 21,796 15s 10
4 PASS 85.71 A 30,652 21s 11
5 PASS 85.71 B 30,698 18s 11
buildbench-backend-subscriptions backend avg 80 80% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 PASS 100 C 17,386 14s 9
2 FAIL 0 D 42,196 29s 13
3 PASS 100 B 16,837 11s 9
4 PASS 100 A 19,776 12s 10
5 PASS 100 C 17,472 12s 9
buildbench-backend-sqlite-library data-sql avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 A 20,359 16s 9
2 PASS 100 B 27,541 19s 11
3 PASS 100 D 20,011 15s 9
4 PASS 100 C 20,508 14s 9
5 PASS 100 A 28,493 18s 11
buildbench-sql-reporting-joins data-sql avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 B 15,977 11s 8
2 PASS 100 D 16,081 16s 8
3 PASS 100 A 11,277 15s 6
4 PASS 100 C 16,162 13s 8
5 PASS 100 B 16,139 11s 8
buildbench-backend-sqlite-notes data-sql avg 90 100% consistent medium
run status score variant tokens time iters
1 PASS 100 B 17,545 12s 9
2 PASS 87.5 C 18,530 11s 9
3 PASS 87.5 A 17,165 11s 9
4 PASS 87.5 D 18,761 11s 9
5 PASS 87.5 B 18,638 15s 10
buildbench-sql-transactions-inventory data-sql avg 85.71 100% consistent medium
run status score variant tokens time iters
1 PASS 85.71 A 15,191 12s 8
2 PASS 85.71 B 15,441 17s 8
3 PASS 85.71 C 15,452 13s 8
4 PASS 85.71 D 15,365 15s 8
5 PASS 85.71 A 10,882 12s 6
buildbench-sql-migrations-users data-sql avg 80 80% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 PASS 100 B 36,921 25s 14
2 FAIL 0 D 11,079 10s 7
3 PASS 100 C 17,258 13s 9
4 PASS 100 A 24,263 15s 10
5 PASS 100 B 15,738 13s 9
buildbench-devops-compose-basic devops avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 A 14,529 25s 10
2 PASS 100 B 18,707 19s 12
3 PASS 100 D 16,050 18s 10
4 PASS 100 C 15,331 26s 10
5 PASS 100 A 19,906 18s 12
buildbench-devops-python-api devops avg 100 100% consistent easy saturated
run status score variant tokens time iters
1 PASS 100 D 23,084 20s 13
2 PASS 100 A 11,825 15s 9
3 PASS 100 B 18,997 63s 12
4 PASS 100 C 15,529 16s 10
5 PASS 100 D 23,263 21s 13
buildbench-devops-multistage devops avg 82 80% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 PASS 100 B 36,514 32s 12
2 PASS 100 D 65,209 47s 18
3 PASS 100 A 55,005 34s 15
4 PASS 100 C 72,429 31s 16
5 FAIL 10 B 79,116 41s 16
buildbench-devops-compose-db devops avg 69 80% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 FAIL 25 C 8,200 7s 5
2 PASS 80 D 21,444 61s 12
3 PASS 80 B 14,751 61s 8
4 PASS 80 A 14,725 62s 8
5 PASS 80 C 22,446 76s 13
buildbench-devops-production devops avg 60 60% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 FAIL 35 D 23,608 15s 10
2 PASS 80 B 21,043 190s 9
3 PASS 80 A 23,532 224s 10
4 FAIL 25 C 10,686 9s 4
5 PASS 80 D 26,502 207s 11
buildbench-repair-web-flow repair avg 76.25 0% consistent hard
run status score variant tokens time iters
1 PARTIAL 76.25 A 20,203 10s 12
2 PARTIAL 76.25 C 17,134 9s 11
3 PARTIAL 76.25 D 17,230 11s 11
4 PARTIAL 76.25 B 15,824 9s 10
5 PARTIAL 76.25 A 17,836 10s 11
buildbench-repair-logic repair avg 75.27 60% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 FAIL 38.18 D 25,647 19s 10
2 PASS 100 C 33,706 15s 12
3 FAIL 38.18 B 25,611 15s 10
4 PASS 100 A 33,721 18s 12
5 PASS 100 D 30,175 16s 10
buildbench-repair-sql-repository repair avg 75.25 40% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 PARTIAL 58.75 D 20,396 10s 11
2 PARTIAL 58.75 C 20,459 14s 11
3 PASS 100 B 20,019 11s 11
4 PASS 100 A 20,717 13s 11
5 PARTIAL 58.75 D 20,429 10s 11
buildbench-repair-backend-service repair avg 66.5 40% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 PARTIAL 58.75 B 37,466 22s 16
2 PASS 100 C 23,167 16s 11
3 PARTIAL 58.75 A 37,469 18s 16
4 FAIL 15 D 27,370 17s 14
5 PASS 100 B 22,021 13s 11
buildbench-repair-data repair avg 64 60% consistent needs revision unstablehigh-variance
run status score variant tokens time iters
1 PASS 100 B 24,503 13s 13
2 FAIL 10 A 19,791 16s 13
3 FAIL 10 C 21,459 13s 14
4 PASS 100 D 26,555 12s 13
5 PASS 100 B 19,114 11s 11
buildbench-web-kanban web avg 60.2 0% consistent hard
run status score variant tokens time iters
1 PARTIAL 60.3 A 43,180 58s 18
2 PARTIAL 60.18 B 23,132 41s 14
3 PARTIAL 60.18 D 21,878 52s 12
4 PARTIAL 60.18 C 31,200 43s 13
5 PARTIAL 60.18 A 34,361 49s 13
buildbench-web-notes web avg 60.06 0% consistent hard
run status score variant tokens time iters
1 PARTIAL 60.06 C 39,669 50s 16
2 PARTIAL 59.7 B 33,621 58s 15
3 PARTIAL 60.18 D 43,213 51s 13
4 PARTIAL 60.18 A 33,637 52s 14
5 PARTIAL 60.18 C 52,647 54s 20
buildbench-web-admin-dashboard web avg 56.16 0% consistent hard
run status score variant tokens time iters
1 PARTIAL 60.18 B 45,576 60s 18
2 PARTIAL 40.18 D 39,862 84s 17
3 PARTIAL 60.18 C 18,496 48s 11
4 PARTIAL 60.18 A 32,021 53s 15
5 PARTIAL 60.06 B 29,884 49s 14
buildbench-web-bookings web avg 56.08 0% consistent hard
run status score variant tokens time iters
1 PARTIAL 60.18 C 27,693 47s 13
2 PARTIAL 59.94 D 29,444 51s 13
3 PARTIAL 60.18 A 39,409 47s 16
4 PARTIAL 60.06 B 32,622 48s 13
5 PARTIAL 40.06 C 29,998 84s 14
buildbench-web-v1 web avg 52.06 0% consistent hard
run status score variant tokens time iters
1 PARTIAL 60.18 A 48,191 51s 14
2 PARTIAL 60.18 D 74,014 55s 17
3 PARTIAL 59.82 C 53,808 58s 17
4 PARTIAL 40.06 B 48,987 84s 15
5 PARTIAL 40.06 A 15,174 82s 6
benchmark-health.md

About this benchmark's maturity — surfaced so the scores above are read for what they are.

Calibration:needs-task-tuning

11easy4medium6hard12needs revision
  • Some tasks look brittle or under-discriminative and should be revised before publication benchmarking.
  • 11 task(s) look too easy to separate models cleanly, so they are less informative for ranking.
  • 12 task(s) show instability, universal failure, or high variance and should be tuned before strong publication claims.
  • The web slice is currently dominated by build and assembly failure, which reduces how much the resulting scores say about product behavior after startup.
  • This benchmark version is useful for internal comparison and tuning, but the task mix still needs work before publication-grade claims.