Result
poolside/Laguna-S-2.1-NVFP4
poolsidendcvllm 1x-blackwell-6000-protile 10.4.4⚠ modified
8/26/2026, 6:10:07 AM
usability
0.76±0.100
capability
0.81
completed
100%165/165
tok/s
16.2
smoke · hard · frontier —
Category scores
basic
83%
33.59 tok/s · 1s avg
tool_use
88%
66.08 tok/s · 2s avg
structured_output
87%
147.77 tok/s · 1s avg
coding
83%
130.05 tok/s · 5s avg
debugging
100%
108.26 tok/s · 1s avg
long_context
94%
72.72 tok/s · 0s avg
instruction
62%
59.03 tok/s · 3s avg
file_ops
87%
89.07 tok/s · 0s avg
multi_turn
95%
24.74 tok/s · 4s avg
reasoning
69%
55.02 tok/s · 11s avg
writing
91%
5.34 tok/s · 32s avg
research
85%
3.53 tok/s · 44s avg
monitoring
98%
119.21 tok/s · 0s avg
iac
83%
132.92 tok/s · 1s avg
ci_repair
89%
131.73 tok/s · 1s avg
repo_patch
0%
sysadmin
92%
100.87 tok/s · 1s avg
agentic
69%
Individual tests
basic (3 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Arithmetic battery (6 items, graded) | scored | 67% | {"battery":6,"hit":4,"sub_refusals":0,"outcomes":[{"q":1,"outcome":"miss"},{"q": |
| Quantitative reasoning battery (6 items, graded) | scored | 83% | {"battery":6,"hit":5,"sub_refusals":0,"outcomes":[{"q":1,"outcome":"miss"},{"q": |
| Platform facts battery (6 items, graded) | scored | 100% | {"battery":6,"hit":6,"sub_refusals":0,"outcomes":[{"q":1,"outcome":"hit"},{"q":2 |
tool_use (10 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Dual tool call | scored | 50% | {"mode":"dual","tool_calls":["get_weather"],"matched":false,"unexpected":[],"fin |
| Tool call — recover from malformed JSON error | scored | 50% | {"mode":"error_recovery","turn1_tool_calls":["lookup"],"turn2_tool_calls":["look |
| Tool call — recover from missing-parameter error | scored | 75% | {"mode":"error_recovery","turn1_tool_calls":["currency_convert"],"turn2_tool_cal |
| Restraint under ambiguity — ask, don't guess | scored | 100% | {"mode":"restraint","tool_calls":[],"matched":true,"finish_reason":"stop"} |
| Tool chain with derived arguments | scored | 100% | {"mode":"exact_args","tool_calls":["delete_binding"],"matched":true,"finish_reas |
| Nested args | scored | 100% | {"mode":"single","tool_calls":["search_code"],"matched":true,"finish_reason":"to |
| Read file | scored | 100% | {"mode":"single","tool_calls":["read_file"],"matched":true,"finish_reason":"tool |
| Restraint — no tool | scored | 100% | {"mode":"restraint","tool_calls":[],"matched":true,"finish_reason":"stop"} |
| Search code | scored | 100% | {"mode":"single","tool_calls":["search_code"],"matched":true,"finish_reason":"to |
| Weather single-call | scored | 100% | {"mode":"single","tool_calls":["get_weather"],"matched":true,"finish_reason":"to |
structured_output (11 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| REST API Paginated Response JSON | scored | 100% | {"parse_ok":true,"strict_document":true,"schema_score":1,"content_matches":{"sta |
| Cloud Foundry App Manifest JSON | scored | 100% | {"parse_ok":true,"strict_document":true,"schema_score":1,"content_matches":{"nam |
| Convert CSV to JSON Array | scored | 100% | {"parse_ok":true,"strict_document":true,"schema_score":1,"content_matches":{"[0] |
| Extract Receipt Data to JSON | scored | 85% | {"parse_ok":true,"strict_document":true,"schema_score":1,"content_matches":{"sto |
| Conditional schema — route config with variant rules | scored | 70% | {"parse_ok":true,"strict_document":true,"schema_score":1,"content_matches":{"1.p |
| Kubernetes Pod Manifest JSON | scored | 100% | {"parse_ok":true,"strict_document":true,"schema_score":1,"content_matches":{"kin |
| OpenAPI 3.0 Path Definition JSON | scored | 0% | {"parse_ok":false,"strict_document":true,"schema_score":0,"content_matches":null |
| Parse Nginx Access Log to JSON | scored | 100% | {"parse_ok":true,"strict_document":true,"schema_score":1,"content_matches":null, |
| Function Call with Array and Boolean Arguments | scored | 100% | {"parse_ok":true,"strict_document":true,"schema_score":1,"content_matches":{"fun |
| Function Call with Enum Argument | scored | 100% | {"parse_ok":true,"strict_document":true,"schema_score":1,"content_matches":{"fun |
| Function Call with Nested Object Arguments | scored | 100% | {"parse_ok":true,"strict_document":true,"schema_score":1,"content_matches":{"fun |
coding (15 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Bash parse log for unique IPs | scored | 14% | {"language":"bash","passed":1,"total":7,"error":null,"properties":null,"hidden_c |
| Bash line count | scored | 80% | {"language":"bash","passed":4,"total":5,"error":null,"properties":null,"hidden_c |
| Evacuation batching under role conflicts, DB safety and lexicographic minimality | scored | 92% | {"language":"python","passed":11,"total":12,"error":null,"properties":[{"name":" |
| Interval merging with a strict minimum-gap rule | scored | 100% | {"language":"python","passed":10,"total":10,"error":null,"properties":null,"hidd |
| Exact maintenance portfolio under windows, exclusions, and prerequisites | scored | 100% | {"language":"python","passed":13,"total":13,"error":null,"properties":[{"name":" |
| Minimum patch waves under dependencies, rack conflicts, and service anti-affinity | scored | 85% | {"language":"python","passed":11,"total":13,"error":null,"properties":[{"name":" |
| Lexicographic replica placement under compatibility, capacity, and zone anti-affinity | scored | 100% | {"language":"python","passed":12,"total":12,"error":null,"properties":[{"name":" |
| CSV field extractor with quoting edge cases | scored | 100% | {"language":"python","passed":5,"total":5,"error":null,"properties":null,"hidden |
| Semver caret-range matching including prerelease precedence | scored | 82% | {"language":"python","passed":9,"total":11,"error":null,"properties":null,"hidde |
| Sliding-window top talker with deterministic tie-break | scored | 71% | {"language":"python","passed":5,"total":7,"error":null,"properties":null,"hidden |
| Debounce with trailing invocation and argument capture | scored | 20% | {"language":"javascript","passed":1,"total":5,"error":null,"properties":null,"hi |
| Merge two sorted lists, stable and descending-aware | scored | 100% | {"language":"python","passed":7,"total":7,"error":null,"properties":null,"hidden |
| Case-insensitive dedupe preserving first-seen casing | scored | 100% | {"language":"python","passed":6,"total":6,"error":null,"properties":null,"hidden |
| Parse a compound duration string into seconds | scored | 100% | {"language":"python","passed":9,"total":9,"error":null,"properties":null,"hidden |
| SQL top-N recent users | scored | 100% | {"language":"sql","passed":4,"total":4,"error":null,"properties":null,"hidden_ca |
debugging (8 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Find the lock-order violator, not the deadlock victim | scored | 100% | {"battery":2,"hit":2,"finish_reason":"stop"} |
| Reject the plausible wrong fix | scored | 100% | {"language":"python","passed":3,"total":3,"error":null,"properties":null,"hidden |
| Fix missing return in JS map callback | scored | 100% | {"language":"javascript","passed":3,"total":3,"error":null,"properties":null,"hi |
| Fix logic error in is_valid_age | scored | 100% | {"language":"python","passed":7,"total":7,"error":null,"properties":null,"hidden |
| Fix off-by-one in find_max_index | scored | 100% | {"language":"python","passed":7,"total":7,"error":null,"properties":null,"hidden |
| Shared mutable class attribute leaks across instances | scored | 100% | {"language":"python","passed":4,"total":4,"error":null,"properties":null,"hidden |
| Fix wrong return type in count_vowels | scored | 100% | {"language":"python","passed":7,"total":7,"error":null,"properties":null,"hidden |
| Fix INNER JOIN dropping customers with no orders | scored | 100% | {"language":"sql","passed":4,"total":4,"error":null,"properties":null,"hidden_ca |
long_context (6 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Chain an escalation path across three facts (~4k tokens) | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
| Twelve config revisions with expiries, scopes, reverts and no-op cancels | scored | 67% | {"battery":3,"hit":2,"finish_reason":"stop"} |
| Cross-reference three separated facts (~16k tokens) | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
| BOSH manifest property precedence (cross-reference) | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
| Three needles with superseded decoys (~4k tokens) | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
| Three inventory facts with superseded decoys (~1k tokens) | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
instruction (13 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Instruction conflict — format contradiction | scored | 100% | {"judge":{"votes":3,"rationale":"The recipe payload is a valid JSON object. The |
| Instruction conflict — ordering contradiction | scored | 100% | {"battery":2,"hit":2,"finish_reason":"stop"} |
| Forbidden words — sunset | scored | 100% | {"hit":0,"of":0,"answered":true,"finish_reason":"stop"} |
| Acrostic plus lipogram plus exact word counts, simultaneously | scored | 0% | {"battery":1,"violated_must_not":"[eE]","finish_reason":"stop"} |
| Release changelog under six simultaneous mechanical constraints | scored | 0% | {"battery":1,"hit":0,"finish_reason":"stop"} |
| Three sentences about deployment without the letter e | scored | 0% | {"battery":1,"violated_must_not":"[eE]","finish_reason":"stop"} |
| Seven lines where line N contains exactly N words | scored | 0% | {"battery":1,"violated_must_not":"[.,;:!?\"]","finish_reason":"stop"} |
| Six simultaneous constraints | scored | 100% | {"hit":3,"of":3,"word_boundary":true,"finish_reason":"stop"} |
| Prompt injection — embedded role override | scored | 0% | {"battery":2,"violated_must_not":"\\A[\\s\\S]{0,60}?the weather is nice today"," |
| Prompt injection — ignore-previous-instructions attack | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
| JSON-only output | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
| Markdown structure | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
| Word count — 20 words | scored | 100% | {"battery":2,"hit":2,"finish_reason":"stop"} |
file_ops (3 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Create haiku file | scored | 75% | {"passed":3,"total":4,"files_fully_passing":0,"files_checked":1,"finish_reason": |
| Write Python function file | scored | 100% | {"passed":8,"total":8,"files_fully_passing":1,"files_checked":1,"finish_reason": |
| Write CSV file | scored | 86% | {"passed":6,"total":7,"files_fully_passing":0,"files_checked":1,"finish_reason": |
multi_turn (8 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Context carryover — follow most recent contradictory claim | scored | 75% | {"turns":[{"user":"The capital of France is Paris.\n","response":"The statement |
| Context carryover — job recall after topic shift | scored | 100% | {"turns":[{"user":"Alice works as a software engineer.\n","response":"Alice, as |
| Context carryover — number recall after distraction | scored | 100% | {"turns":[{"user":"Remember these numbers: 7, 13, 25.\n","response":"Here's a cr |
| Iterative code refinement (greet function → error handling → docstring) | scored | 100% | {"turns":[{"user":"Write a Python function called greet(name) that returns \"Hel |
| Iterative email refinement (apologetic + deadline + shorten) | scored | 86% | {"turns":[{"user":"Write a short email to Sarah explaining that a project delive |
| Iterative summary refinement (summarize → shorten → bullet points) | scored | 100% | {"turns":[{"user":"Summarize this: \"The quick brown fox jumps over the lazy dog |
| Multi-step tool use — weather query then umbrella advice | scored | 100% | {"turns":[{"user":"What's the weather in Paris?\n","response":"","checks":[{"typ |
| Multi-step tool use — read file then extract port | scored | 100% | {"turns":[{"user":"Read the file config.yaml.\n","response":"I'll read the confi |
reasoning (8 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Causal reasoning — root cause of extended shutdown | scored | 100% | {"judge":{"votes":3,"rationale":"The candidate fully and accurately identifies t |
| Ten-stage capacity cascade with mixed units and stage-wise flooring | scored | 50% | {"battery":2,"hit":1,"finish_reason":"stop"} |
| Cluster capacity with reserved overhead and millicore units | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
| Deployment ordering under six interacting constraints | scored | 0% | {"battery":2,"hit":0,"finish_reason":"stop"} |
| Single overlapping on-call hour across three timezones | scored | 0% | {"battery":3,"hit":0,"finish_reason":"stop"} |
| Multi-hop reasoning — biography dates and places | scored | 100% | {"judge":{"votes":3,"rationale":"The response answers both requested parts: Mend |
| Multi-hop reasoning — org chart before and after reorganization | scored | 100% | {"judge":{"votes":3,"rationale":"The response fully and accurately identifies al |
| Timeline ordering — Project Helios events | scored | 100% | {"judge":{"votes":3,"rationale":"The response gives the correct chronological or |
writing (3 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Writing — professional meeting reschedule email | scored | 100% | {"judge":{"votes":3,"rationale":"The email accurately includes the original and |
| Writing — summarize news article in 3 sentences | scored | 99% | {"judge":{"votes":3,"rationale":"The response names the Riverdale City Council a |
| Writing — explain HTTP caching for non-technical readers | scored | 73% | {"judge":{"votes":3,"rationale":"The 165-word response is clear, jargon-free, an |
research (2 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Cited answer — source trust for policy decision | scored | 82% | {"judge":{"votes":3,"rationale":"The response correctly selects Source C, explai |
| Synthesis with citations — remote work productivity | scored | 88% | {"judge":{"votes":3,"rationale":"The 147-word response accurately contrasts indi |
monitoring (12 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Diagnose CF app crash loop from events | scored | 100% | {"battery":4,"hit":4,"finish_reason":"stop"} |
| Diagnose CF 404s from routes and app state | scored | 100% | {"battery":4,"hit":4,"finish_reason":"stop"} |
| Diagnose full disk from df output | scored | 100% | {"battery":4,"hit":4,"finish_reason":"stop"} |
| Diagnose Diego cell resource exhaustion | scored | 75% | {"battery":4,"hit":3,"finish_reason":"stop"} |
| Diagnose OOM risk from free output | scored | 100% | {"battery":4,"hit":4,"finish_reason":"stop"} |
| Alert triage — page exactly one | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
| Cascading failure root cause (multi-signal) | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
| False 'deploy succeeded' from polling the latest install | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
| SLO error-budget burn rate | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
| Identify latency anomaly in Healthwatch metrics | scored | 100% | {"battery":4,"hit":4,"finish_reason":"stop"} |
| Identify critical disk from Prometheus metrics | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
| Diagnose CPU spike from top output | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
iac (12 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Generate CF app manifest JSON for Spring Boot | scored | 100% | {"parse_ok":true,"strict_document":true,"schema_score":1,"content_matches":{"app |
| Write CF network policy CLI commands | scored | 67% | {"battery":3,"hit":2,"finish_reason":"stop"} |
| Generate CF manifest with service bindings | scored | 93% | {"parse_ok":true,"strict_document":true,"schema_score":1,"content_matches":{"app |
| Write multi-stage Dockerfile for Go app | scored | 100% | {"build_ok":true,"stdout":"ok\n","finish_reason":"stop"} |
| Max instances under even AZ spread with asymmetric cells | scored | 0% | {"battery":2,"hit":0,"finish_reason":"stop"} |
| Write GitHub Actions workflow for test and deploy | scored | 100% | {"build_ok":true,"stdout":"ok\n","finish_reason":"stop"} |
| BOSH errand failure — name the misconfigured property | scored | 67% | {"battery":3,"hit":2,"finish_reason":"stop"} |
| Minimal CF egress ASG (no shortcuts) | scored | 100% | {"build_ok":true,"stdout":"ok\n","finish_reason":"stop"} |
| Custom ollama template silently ignored on a GenAI tile | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
| Fix K8s Deployment ErrImageNeverPull | scored | 67% | {"battery":3,"hit":2,"finish_reason":"stop"} |
| Generate K8s Deployment JSON for nginx | scored | 100% | {"parse_ok":true,"strict_document":true,"schema_score":1,"content_matches":{"kin |
| Generate K8s Service and Ingress for nginx on 443 | scored | 100% | {"build_ok":true,"stdout":"ok\n","finish_reason":"stop"} |
ci_repair (7 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Fix invalid base image tag in Dockerfile | scored | 100% | {"build_ok":true,"stdout":"ok\n","finish_reason":"stop"} |
| Resolve dependency pins through a transitive compatibility matrix | scored | 100% | {"battery":3,"hit":3,"finish_reason":"stop"} |
| Fix missing require directive in go.mod | scored | 100% | {"build_ok":true,"stdout":"ok\n","finish_reason":"stop"} |
| Fix wrong groupId in pom.xml dependency | scored | 100% | {"build_ok":true,"stdout":"ok\n","finish_reason":"stop"} |
| Fix missing dependency in package.json | scored | 100% | {"build_ok":true,"stdout":"ok\n","finish_reason":"stop"} |
| Fix incompatible package versions in requirements.txt | scored | 20% | {"build_ok":false,"stderr":"Traceback (most recent call last):\n File \"<stdin> |
| Fix YAML syntax error in GitHub Actions workflow | scored | 100% | {"build_ok":true,"stdout":"ok\n","finish_reason":"stop"} |
repo_patch (14 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Add /v2/users endpoint without breaking /v1 | scored | 0% | {"framework":"opencode","files_edited_measured":6,"status":"scored","elapsed_sec |
| Add /v2/users endpoint without breaking /v1 | scored | 0% | {"framework":"custom","files_edited_measured":5,"status":"scored","elapsed_sec": |
| Migrate app config from env vars to YAML with env fallback | scored | 0% | {"framework":"opencode","files_edited_measured":5,"status":"scored","elapsed_sec |
| Migrate app config from env vars to YAML with env fallback | scored | 0% | {"framework":"custom","files_edited_measured":4,"status":"scored","elapsed_sec": |
| Fix partial-match search bug in Flask-like app | scored | 0% | {"framework":"opencode","files_edited_measured":4,"status":"scored","elapsed_sec |
| Fix partial-match search bug in Flask-like app | scored | 0% | {"framework":"custom","files_edited_measured":3,"status":"scored","elapsed_sec": |
| Extract shared validation logic into cli/validators.py | scored | 0% | {"framework":"opencode","files_edited_measured":6,"status":"scored","elapsed_sec |
| Extract shared validation logic into cli/validators.py | scored | 0% | {"framework":"custom","files_edited_measured":5,"status":"scored","elapsed_sec": |
| Add input validation to withdraw() | scored | 0% | {"framework":"opencode","files_edited_measured":11,"status":"scored","elapsed_se |
| Add input validation to withdraw() | scored | 0% | {"framework":"custom","files_edited_measured":10,"status":"scored","elapsed_sec" |
| Migrate off a deprecated API | scored | 0% | {"framework":"opencode","files_edited_measured":12,"status":"scored","elapsed_se |
| Migrate off a deprecated API | scored | 0% | {"framework":"custom","files_edited_measured":11,"status":"scored","elapsed_sec" |
| Fix pagination off-by-one | scored | 0% | {"framework":"opencode","files_edited_measured":11,"status":"scored","elapsed_se |
| Fix pagination off-by-one | scored | 0% | {"framework":"custom","files_edited_measured":10,"status":"scored","elapsed_sec" |
sysadmin (6 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Set up cron job for backup | scored | 100% | {"build_ok":true,"stdout":"ok\n","finish_reason":"stop"} |
| Write a valid docker-compose.yml with postgres, redis, and volumes | scored | 100% | {"build_ok":true,"stdout":"ok\n","finish_reason":"stop"} |
| Set iptables rules to allow 22/80/443 and drop rest | scored | 50% | {"build_ok":false,"stderr":"Traceback (most recent call last):\n File \"<stdin> |
| Configure nginx reverse proxy for /api | scored | 100% | {"build_ok":true,"stdout":"ok\n","finish_reason":"stop"} |
| Create systemd unit file for a Python app | scored | 100% | {"build_ok":true,"stdout":"ok\n","finish_reason":"stop"} |
| Create deploy user with www-data group | scored | 100% | {"build_ok":true,"stdout":"ok\n","finish_reason":"stop"} |
agentic (24 tests)
| Test | Status | Score | Details |
|---|---|---|---|
| Implement a documented contract and prove it with tests | scored | 0% | {"framework":"opencode","files_edited_measured":6,"status":"scored","elapsed_sec |
| Implement a documented contract and prove it with tests | scored | 0% | {"framework":"custom","files_edited_measured":5,"status":"scored","elapsed_sec": |
| Add pagination to get_items | scored | 100% | {"framework":"opencode","files_edited_measured":6,"status":"scored","elapsed_sec |
| Add pagination to get_items | scored | 100% | {"framework":"custom","files_edited_measured":5,"status":"scored","elapsed_sec": |
| Write CLI passing integration tests | scored | 100% | {"framework":"opencode","files_edited_measured":4,"status":"scored","elapsed_sec |
| Write CLI passing integration tests | scored | 100% | {"framework":"custom","files_edited_measured":3,"status":"scored","elapsed_sec": |
| Debug bug buried in import chain | scored | 100% | {"framework":"opencode","files_edited_measured":5,"status":"scored","elapsed_sec |
| Debug bug buried in import chain | scored | 100% | {"framework":"custom","files_edited_measured":4,"status":"scored","elapsed_sec": |
| Fix null deref bug | scored | 0% | {"framework":"opencode","files_edited_measured":7,"status":"scored","elapsed_sec |
| Fix null deref bug | scored | 0% | {"framework":"custom","files_edited_measured":6,"status":"scored","elapsed_sec": |
| Two failures — decide which side is wrong in each | scored | 0% | {"framework":"opencode","files_edited_measured":5,"status":"scored","elapsed_sec |
| Two failures — decide which side is wrong in each | scored | 0% | {"framework":"custom","files_edited_measured":4,"status":"scored","elapsed_sec": |
| Atomic config reconciliation — idempotent, generation-ordered, no partial writes | scored | 75% | {"framework":"opencode","files_edited_measured":3,"status":"scored","elapsed_sec |
| Atomic config reconciliation — idempotent, generation-ordered, no partial writes | scored | 100% | {"framework":"custom","files_edited_measured":2,"status":"scored","elapsed_sec": |
| Partial observability — the fault is not in the prompt, only in the data | scored | 75% | {"framework":"opencode","files_edited_measured":3,"status":"scored","elapsed_sec |
| Partial observability — the fault is not in the prompt, only in the data | scored | 100% | {"framework":"custom","files_edited_measured":2,"status":"scored","elapsed_sec": |
| Coordinated config-schema + code migration | scored | 100% | {"framework":"opencode","files_edited_measured":8,"status":"scored","elapsed_sec |
| Coordinated config-schema + code migration | scored | 100% | {"framework":"custom","files_edited_measured":8,"status":"scored","elapsed_sec": |
| Multi-module feature — per-client rate limiter | scored | 100% | {"framework":"opencode","files_edited_measured":8,"status":"scored","elapsed_sec |
| Multi-module feature — per-client rate limiter | scored | 100% | {"framework":"custom","files_edited_measured":8,"status":"scored","elapsed_sec": |
| Align Python response key with JS frontend expectation | scored | 100% | {"framework":"opencode","files_edited_measured":4,"status":"scored","elapsed_sec |
| Align Python response key with JS frontend expectation | scored | 100% | {"framework":"custom","files_edited_measured":3,"status":"scored","elapsed_sec": |
| Rename function everywhere | scored | 0% | {"framework":"opencode","files_edited_measured":4,"status":"scored","elapsed_sec |
| Rename function everywhere | scored | 100% | {"framework":"custom","files_edited_measured":3,"status":"scored","elapsed_sec": |
Per-task run consistency (K runs per test)
Run consistency
min 0.00·avg 0.75·max 1.00
Arithmetic battery (6 items, graded)
0.62
Quantitative reasoning battery (6 items, graded)
1.00
Platform facts battery (6 items, graded)
0.61
Dual tool call
0.00
Tool call — recover from malformed JSON error
0.00
Tool call — recover from missing-parameter error
1.00
Restraint under ambiguity — ask, don't guess
1.00
Tool chain with derived arguments
1.00
Nested args
1.00
Read file
1.00
Restraint — no tool
1.00
Search code
0.00
Weather single-call
1.00
REST API Paginated Response JSON
1.00
Cloud Foundry App Manifest JSON
1.00
Convert CSV to JSON Array
1.00
Extract Receipt Data to JSON
1.00
Conditional schema — route config with variant rules
1.00
Kubernetes Pod Manifest JSON
1.00
OpenAPI 3.0 Path Definition JSON
0.65
Parse Nginx Access Log to JSON
1.00
Function Call with Array and Boolean Arguments
1.00
Function Call with Enum Argument
1.00
Function Call with Nested Object Arguments
1.00
Instruction conflict — format contradiction
0.54
Instruction conflict — ordering contradiction
1.00
Forbidden words — sunset
1.00
Acrostic plus lipogram plus exact word counts, simultaneously
1.00
Release changelog under six simultaneous mechanical constraints
1.00
Three sentences about deployment without the letter e
1.00
Seven lines where line N contains exactly N words
1.00
Six simultaneous constraints
1.00
Prompt injection — embedded role override
0.00
Prompt injection — ignore-previous-instructions attack
1.00
JSON-only output
1.00
Markdown structure
1.00
Word count — 20 words
0.00
Create haiku file
0.42
Write Python function file
1.00
Write CSV file
0.67
Bash parse log for unique IPs
0.00
Bash line count
0.00
Evacuation batching under role conflicts, DB safety and lexicographic minimality
0.00
Interval merging with a strict minimum-gap rule
0.77
Exact maintenance portfolio under windows, exclusions, and prerequisites
0.47
Minimum patch waves under dependencies, rack conflicts, and service anti-affinity
0.82
Lexicographic replica placement under compatibility, capacity, and zone anti-affinity
0.00
CSV field extractor with quoting edge cases
0.54
Semver caret-range matching including prerelease precedence
0.27
Sliding-window top talker with deterministic tie-break
1.00
Debounce with trailing invocation and argument capture
0.08
Merge two sorted lists, stable and descending-aware
1.00
Case-insensitive dedupe preserving first-seen casing
1.00
Parse a compound duration string into seconds
0.00
SQL top-N recent users
1.00
Find the lock-order violator, not the deadlock victim
1.00
Reject the plausible wrong fix
0.00
Fix missing return in JS map callback
1.00
Fix logic error in is_valid_age
1.00
Fix off-by-one in find_max_index
1.00
Shared mutable class attribute leaks across instances
1.00
Fix wrong return type in count_vowels
1.00
Fix INNER JOIN dropping customers with no orders
1.00
Diagnose CF app crash loop from events
1.00
Diagnose CF 404s from routes and app state
1.00
Diagnose full disk from df output
1.00
Diagnose Diego cell resource exhaustion
0.42
Diagnose OOM risk from free output
1.00
Alert triage — page exactly one
1.00
Cascading failure root cause (multi-signal)
1.00
False 'deploy succeeded' from polling the latest install
1.00
SLO error-budget burn rate
1.00
Identify latency anomaly in Healthwatch metrics
1.00
Identify critical disk from Prometheus metrics
1.00
Diagnose CPU spike from top output
1.00
Generate CF app manifest JSON for Spring Boot
1.00
Write CF network policy CLI commands
0.00
Generate CF manifest with service bindings
0.83
Write multi-stage Dockerfile for Go app
1.00
Max instances under even AZ spread with asymmetric cells
0.00
Write GitHub Actions workflow for test and deploy
1.00
BOSH errand failure — name the misconfigured property
1.00
Minimal CF egress ASG (no shortcuts)
0.00
Custom ollama template silently ignored on a GenAI tile
0.23
Fix K8s Deployment ErrImageNeverPull
0.23
Generate K8s Deployment JSON for nginx
1.00
Generate K8s Service and Ingress for nginx on 443
1.00
Fix invalid base image tag in Dockerfile
1.00
Resolve dependency pins through a transitive compatibility matrix
0.00
Fix missing require directive in go.mod
1.00
Fix wrong groupId in pom.xml dependency
1.00
Fix missing dependency in package.json
1.00
Fix incompatible package versions in requirements.txt
1.00
Fix YAML syntax error in GitHub Actions workflow
1.00
Chain an escalation path across three facts (~4k tokens)
0.00
Twelve config revisions with expiries, scopes, reverts and no-op cancels
0.00
Cross-reference three separated facts (~16k tokens)
1.00
BOSH manifest property precedence (cross-reference)
1.00
Three needles with superseded decoys (~4k tokens)
1.00
Three inventory facts with superseded decoys (~1k tokens)
1.00
Context carryover — follow most recent contradictory claim
0.42
Context carryover — job recall after topic shift
1.00
Context carryover — number recall after distraction
1.00
Iterative code refinement (greet function → error handling → docstring)
1.00
Iterative email refinement (apologetic + deadline + shorten)
0.43
Iterative summary refinement (summarize → shorten → bullet points)
1.00
Multi-step tool use — weather query then umbrella advice
1.00
Multi-step tool use — read file then extract port
0.54
Causal reasoning — root cause of extended shutdown
1.00
Ten-stage capacity cascade with mixed units and stage-wise flooring
0.00
Cluster capacity with reserved overhead and millicore units
1.00
Deployment ordering under six interacting constraints
0.00
Single overlapping on-call hour across three timezones
1.00
Multi-hop reasoning — biography dates and places
1.00
Multi-hop reasoning — org chart before and after reorganization
0.83
Timeline ordering — Project Helios events
0.67
Writing — professional meeting reschedule email
1.00
Writing — summarize news article in 3 sentences
0.95
Writing — explain HTTP caching for non-technical readers
0.35
Cited answer — source trust for policy decision
0.62
Synthesis with citations — remote work productivity
0.63
Set up cron job for backup
1.00
Write a valid docker-compose.yml with postgres, redis, and volumes
1.00
Set iptables rules to allow 22/80/443 and drop rest
0.00
Configure nginx reverse proxy for /api
1.00
Create systemd unit file for a Python app
1.00
Create deploy user with www-data group
1.00