// benchmark report
gpt-5.4-mini
algorithm-strong, web-weak, operationally-fragile, budget-limited
Strongest
- algorithm94.78
- data-sql91.14
Weakest
- web56.91
- repair71.45
Best fit
- contained implementation tasks
- SQL/repository tasks
- service-layer backend work
Weaker fit
- autonomous full-stack delivery
- longer tasks under tight iteration/token budgets
Dominant failure modes
// primary metrics
75-89 · good and usable
Median: 100
Std Dev: +/-25.63
Range: [0, 100]
Pass / Partial / Fail: 69.7% / 22.4% / 7.9%
The main weakness pattern is build and container assembly failure before functional evaluation.
// secondary metrics — directional only
Tokens: 3,957,577
Cost: $2.55
Time: 70.3 min
Resource metrics are comparable only when task set, protocol, evaluator, and run counts match.
// by category
95% pass · 0% partial · 5% fail
Top failure: correctness ( 100%)
algorithm scored 94.78, which is in the strong and reliable range. Pass rate was 95.0%. Repeated-run consistency was 95.0%. The main weakness pattern is correctness drift on harder cases, where some implementations look plausible but still miss required behavior.
96% pass · 0% partial · 4% fail
Top failure: correctness ( 100%)
data-sql scored 91.14, which is in the strong and reliable range. Pass rate was 96.0%. Repeated-run consistency was 96.0%. The main weakness pattern is correctness drift on harder cases, where some implementations look plausible but still miss required behavior.
88% pass · 8% partial · 4% fail
Top failure: correctness ( 100%)
backend scored 86.64, which is in the good and usable range. Pass rate was 88.0%. Repeated-run consistency was 88.0%. The main weakness pattern is correctness drift on harder cases, where some implementations look plausible but still miss required behavior.
84% pass · 0% partial · 16% fail
Top failure: budget-limited ( 50%)
devops scored 82.2, which is in the good and usable range. Pass rate was 84.0%. Repeated-run consistency was 84.0%. The main weakness pattern is budget pressure under iteration or token limits, which leads to incomplete runs.
40% pass · 40% partial · 20% fail
Top failure: correctness ( 93.3%)
repair scored 71.45, which is in the mixed but useful range. Pass rate was 40.0%. Repeated-run consistency was 40.0%. The main weakness pattern is correctness drift on harder cases, where some implementations look plausible but still miss required behavior.
0% pass · 100% partial · 0% fail
Top failure: build/assembly ( 84%)
Web breakdown
Quality points reflect static code signals only and should not be read as evidence that the app built or ran successfully.
web scored 56.91, which is in the weak or inconsistent range. Pass rate was 0.0%. Repeated-run consistency was 0.0%. The main weakness pattern is build and container assembly failure before functional evaluation.
// by task — expand for individual runs
buildbench-algo-advanced algorithm avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | A | 19,449 | 12s | 12 |
| 2 | PASS | 100 | B | 16,555 | 11s | 11 |
| 3 | PASS | 100 | C | 19,217 | 12s | 12 |
| 4 | PASS | 100 | D | 16,675 | 11s | 11 |
| 5 | PASS | 100 | A | 16,627 | 11s | 11 |
buildbench-algo-graphs algorithm avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | B | 22,813 | 14s | 12 |
| 2 | PASS | 100 | A | 18,551 | 13s | 10 |
| 3 | PASS | 100 | D | 18,288 | 12s | 10 |
| 4 | PASS | 100 | C | 19,742 | 13s | 11 |
| 5 | PASS | 100 | B | 19,554 | 12s | 11 |
buildbench-algo-randomized-set algorithm avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | B | 8,193 | 7s | 6 |
| 2 | PASS | 100 | D | 16,510 | 8s | 8 |
| 3 | PASS | 100 | A | 6,879 | 5s | 5 |
| 4 | PASS | 100 | C | 6,879 | 6s | 5 |
| 5 | PASS | 100 | B | 8,199 | 8s | 6 |
buildbench-algo-segment-tree algorithm avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | A | 7,607 | 8s | 5 |
| 2 | PASS | 100 | B | 7,601 | 8s | 5 |
| 3 | PASS | 100 | C | 8,841 | 7s | 6 |
| 4 | PASS | 100 | D | 7,562 | 8s | 5 |
| 5 | PASS | 100 | A | 18,825 | 15s | 7 |
buildbench-algo-trie-wildcard algorithm avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | A | 7,634 | 8s | 6 |
| 2 | PASS | 100 | C | 6,374 | 7s | 5 |
| 3 | PASS | 100 | D | 16,935 | 12s | 8 |
| 4 | PASS | 100 | B | 7,618 | 9s | 6 |
| 5 | PASS | 100 | A | 6,328 | 5s | 5 |
buildbench-algo-ksum algorithm avg 99.26 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | B | 17,426 | 11s | 8 |
| 2 | PASS | 100 | A | 15,821 | 12s | 8 |
| 3 | PASS | 100 | C | 15,560 | 12s | 8 |
| 4 | PASS | 96.3 | D | 19,085 | 11s | 8 |
| 5 | PASS | 100 | B | 17,366 | 11s | 8 |
buildbench-algo-lfu-cache algorithm avg 80 80% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | B | 10,503 | 8s | 6 |
| 2 | PASS | 100 | A | 10,534 | 9s | 6 |
| 3 | PASS | 100 | C | 10,482 | 8s | 6 |
| 4 | FAIL | 0 | D | 12,181 | 8s | 6 |
| 5 | PASS | 100 | B | 21,849 | 11s | 8 |
buildbench-algo-medium algorithm avg 79 80% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | FAIL | 0 | D | 10,498 | 5s | 6 |
| 2 | PASS | 100 | A | 48,369 | 16s | 12 |
| 3 | PASS | 100 | B | 85,075 | 23s | 17 |
| 4 | PASS | 95 | C | 73,838 | 22s | 16 |
| 5 | PASS | 100 | D | 49,135 | 19s | 12 |
buildbench-backend-policy-engine backend avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | A | 17,903 | 15s | 11 |
| 2 | PASS | 100 | B | 15,417 | 11s | 9 |
| 3 | PASS | 100 | D | 15,589 | 10s | 9 |
| 4 | PASS | 100 | C | 14,500 | 11s | 9 |
| 5 | PASS | 100 | A | 14,475 | 12s | 9 |
buildbench-backend-approvals backend avg 87.5 100% consistent medium
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 87.5 | D | 36,099 | 19s | 11 |
| 2 | PASS | 87.5 | A | 24,680 | 15s | 11 |
| 3 | PASS | 87.5 | C | 25,238 | 15s | 11 |
| 4 | PASS | 87.5 | B | 29,035 | 17s | 13 |
| 5 | PASS | 87.5 | D | 25,279 | 16s | 11 |
buildbench-backend-notifications backend avg 85.71 100% consistent medium
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 85.71 | D | 16,344 | 11s | 9 |
| 2 | PASS | 85.71 | A | 16,474 | 11s | 9 |
| 3 | PASS | 85.71 | C | 16,553 | 13s | 9 |
| 4 | PASS | 85.71 | B | 16,272 | 11s | 9 |
| 5 | PASS | 85.71 | D | 16,461 | 11s | 9 |
buildbench-backend-orders backend avg 80 60% consistent needs revision unstable
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 71.43 | B | 30,429 | 18s | 11 |
| 2 | PARTIAL | 71.43 | D | 29,891 | 17s | 11 |
| 3 | PASS | 85.71 | C | 21,796 | 15s | 10 |
| 4 | PASS | 85.71 | A | 30,652 | 21s | 11 |
| 5 | PASS | 85.71 | B | 30,698 | 18s | 11 |
buildbench-backend-subscriptions backend avg 80 80% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | C | 17,386 | 14s | 9 |
| 2 | FAIL | 0 | D | 42,196 | 29s | 13 |
| 3 | PASS | 100 | B | 16,837 | 11s | 9 |
| 4 | PASS | 100 | A | 19,776 | 12s | 10 |
| 5 | PASS | 100 | C | 17,472 | 12s | 9 |
buildbench-backend-sqlite-library data-sql avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | A | 20,359 | 16s | 9 |
| 2 | PASS | 100 | B | 27,541 | 19s | 11 |
| 3 | PASS | 100 | D | 20,011 | 15s | 9 |
| 4 | PASS | 100 | C | 20,508 | 14s | 9 |
| 5 | PASS | 100 | A | 28,493 | 18s | 11 |
buildbench-sql-reporting-joins data-sql avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | B | 15,977 | 11s | 8 |
| 2 | PASS | 100 | D | 16,081 | 16s | 8 |
| 3 | PASS | 100 | A | 11,277 | 15s | 6 |
| 4 | PASS | 100 | C | 16,162 | 13s | 8 |
| 5 | PASS | 100 | B | 16,139 | 11s | 8 |
buildbench-backend-sqlite-notes data-sql avg 90 100% consistent medium
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | B | 17,545 | 12s | 9 |
| 2 | PASS | 87.5 | C | 18,530 | 11s | 9 |
| 3 | PASS | 87.5 | A | 17,165 | 11s | 9 |
| 4 | PASS | 87.5 | D | 18,761 | 11s | 9 |
| 5 | PASS | 87.5 | B | 18,638 | 15s | 10 |
buildbench-sql-transactions-inventory data-sql avg 85.71 100% consistent medium
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 85.71 | A | 15,191 | 12s | 8 |
| 2 | PASS | 85.71 | B | 15,441 | 17s | 8 |
| 3 | PASS | 85.71 | C | 15,452 | 13s | 8 |
| 4 | PASS | 85.71 | D | 15,365 | 15s | 8 |
| 5 | PASS | 85.71 | A | 10,882 | 12s | 6 |
buildbench-sql-migrations-users data-sql avg 80 80% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | B | 36,921 | 25s | 14 |
| 2 | FAIL | 0 | D | 11,079 | 10s | 7 |
| 3 | PASS | 100 | C | 17,258 | 13s | 9 |
| 4 | PASS | 100 | A | 24,263 | 15s | 10 |
| 5 | PASS | 100 | B | 15,738 | 13s | 9 |
buildbench-devops-compose-basic devops avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | A | 14,529 | 25s | 10 |
| 2 | PASS | 100 | B | 18,707 | 19s | 12 |
| 3 | PASS | 100 | D | 16,050 | 18s | 10 |
| 4 | PASS | 100 | C | 15,331 | 26s | 10 |
| 5 | PASS | 100 | A | 19,906 | 18s | 12 |
buildbench-devops-python-api devops avg 100 100% consistent easy saturated
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | D | 23,084 | 20s | 13 |
| 2 | PASS | 100 | A | 11,825 | 15s | 9 |
| 3 | PASS | 100 | B | 18,997 | 63s | 12 |
| 4 | PASS | 100 | C | 15,529 | 16s | 10 |
| 5 | PASS | 100 | D | 23,263 | 21s | 13 |
buildbench-devops-multistage devops avg 82 80% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | B | 36,514 | 32s | 12 |
| 2 | PASS | 100 | D | 65,209 | 47s | 18 |
| 3 | PASS | 100 | A | 55,005 | 34s | 15 |
| 4 | PASS | 100 | C | 72,429 | 31s | 16 |
| 5 | FAIL | 10 | B | 79,116 | 41s | 16 |
buildbench-devops-compose-db devops avg 69 80% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | FAIL | 25 | C | 8,200 | 7s | 5 |
| 2 | PASS | 80 | D | 21,444 | 61s | 12 |
| 3 | PASS | 80 | B | 14,751 | 61s | 8 |
| 4 | PASS | 80 | A | 14,725 | 62s | 8 |
| 5 | PASS | 80 | C | 22,446 | 76s | 13 |
buildbench-devops-production devops avg 60 60% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | FAIL | 35 | D | 23,608 | 15s | 10 |
| 2 | PASS | 80 | B | 21,043 | 190s | 9 |
| 3 | PASS | 80 | A | 23,532 | 224s | 10 |
| 4 | FAIL | 25 | C | 10,686 | 9s | 4 |
| 5 | PASS | 80 | D | 26,502 | 207s | 11 |
buildbench-repair-web-flow repair avg 76.25 0% consistent hard
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 76.25 | A | 20,203 | 10s | 12 |
| 2 | PARTIAL | 76.25 | C | 17,134 | 9s | 11 |
| 3 | PARTIAL | 76.25 | D | 17,230 | 11s | 11 |
| 4 | PARTIAL | 76.25 | B | 15,824 | 9s | 10 |
| 5 | PARTIAL | 76.25 | A | 17,836 | 10s | 11 |
buildbench-repair-logic repair avg 75.27 60% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | FAIL | 38.18 | D | 25,647 | 19s | 10 |
| 2 | PASS | 100 | C | 33,706 | 15s | 12 |
| 3 | FAIL | 38.18 | B | 25,611 | 15s | 10 |
| 4 | PASS | 100 | A | 33,721 | 18s | 12 |
| 5 | PASS | 100 | D | 30,175 | 16s | 10 |
buildbench-repair-sql-repository repair avg 75.25 40% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 58.75 | D | 20,396 | 10s | 11 |
| 2 | PARTIAL | 58.75 | C | 20,459 | 14s | 11 |
| 3 | PASS | 100 | B | 20,019 | 11s | 11 |
| 4 | PASS | 100 | A | 20,717 | 13s | 11 |
| 5 | PARTIAL | 58.75 | D | 20,429 | 10s | 11 |
buildbench-repair-backend-service repair avg 66.5 40% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 58.75 | B | 37,466 | 22s | 16 |
| 2 | PASS | 100 | C | 23,167 | 16s | 11 |
| 3 | PARTIAL | 58.75 | A | 37,469 | 18s | 16 |
| 4 | FAIL | 15 | D | 27,370 | 17s | 14 |
| 5 | PASS | 100 | B | 22,021 | 13s | 11 |
buildbench-repair-data repair avg 64 60% consistent needs revision unstablehigh-variance
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PASS | 100 | B | 24,503 | 13s | 13 |
| 2 | FAIL | 10 | A | 19,791 | 16s | 13 |
| 3 | FAIL | 10 | C | 21,459 | 13s | 14 |
| 4 | PASS | 100 | D | 26,555 | 12s | 13 |
| 5 | PASS | 100 | B | 19,114 | 11s | 11 |
buildbench-web-kanban web avg 60.2 0% consistent hard
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 60.3 | A | 43,180 | 58s | 18 |
| 2 | PARTIAL | 60.18 | B | 23,132 | 41s | 14 |
| 3 | PARTIAL | 60.18 | D | 21,878 | 52s | 12 |
| 4 | PARTIAL | 60.18 | C | 31,200 | 43s | 13 |
| 5 | PARTIAL | 60.18 | A | 34,361 | 49s | 13 |
buildbench-web-notes web avg 60.06 0% consistent hard
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 60.06 | C | 39,669 | 50s | 16 |
| 2 | PARTIAL | 59.7 | B | 33,621 | 58s | 15 |
| 3 | PARTIAL | 60.18 | D | 43,213 | 51s | 13 |
| 4 | PARTIAL | 60.18 | A | 33,637 | 52s | 14 |
| 5 | PARTIAL | 60.18 | C | 52,647 | 54s | 20 |
buildbench-web-admin-dashboard web avg 56.16 0% consistent hard
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 60.18 | B | 45,576 | 60s | 18 |
| 2 | PARTIAL | 40.18 | D | 39,862 | 84s | 17 |
| 3 | PARTIAL | 60.18 | C | 18,496 | 48s | 11 |
| 4 | PARTIAL | 60.18 | A | 32,021 | 53s | 15 |
| 5 | PARTIAL | 60.06 | B | 29,884 | 49s | 14 |
buildbench-web-bookings web avg 56.08 0% consistent hard
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 60.18 | C | 27,693 | 47s | 13 |
| 2 | PARTIAL | 59.94 | D | 29,444 | 51s | 13 |
| 3 | PARTIAL | 60.18 | A | 39,409 | 47s | 16 |
| 4 | PARTIAL | 60.06 | B | 32,622 | 48s | 13 |
| 5 | PARTIAL | 40.06 | C | 29,998 | 84s | 14 |
buildbench-web-v1 web avg 52.06 0% consistent hard
| run | status | score | variant | tokens | time | iters |
|---|---|---|---|---|---|---|
| 1 | PARTIAL | 60.18 | A | 48,191 | 51s | 14 |
| 2 | PARTIAL | 60.18 | D | 74,014 | 55s | 17 |
| 3 | PARTIAL | 59.82 | C | 53,808 | 58s | 17 |
| 4 | PARTIAL | 40.06 | B | 48,987 | 84s | 15 |
| 5 | PARTIAL | 40.06 | A | 15,174 | 82s | 6 |
About this benchmark's maturity — surfaced so the scores above are read for what they are.
Calibration:needs-task-tuning
- Some tasks look brittle or under-discriminative and should be revised before publication benchmarking.
- 11 task(s) look too easy to separate models cleanly, so they are less informative for ranking.
- 12 task(s) show instability, universal failure, or high variance and should be tuned before strong publication claims.
- The web slice is currently dominated by build and assembly failure, which reduces how much the resulting scores say about product behavior after startup.
- This benchmark version is useful for internal comparison and tuning, but the task mix still needs work before publication-grade claims.