A run is a set of scenarios plus a length. Pick a preset, pick what it draws from, and answer.
Afterwards you get percent-correct per domain and per task statement, and the name of the
distractor family behind every wrong option you chose.
No scaled score, ever. The 720 cut score on the
100–1000 scale is fixed by a standard-setting study
against the real item pool. A percentage on original practice items cannot be converted into
it, so this product will not pretend otherwise. What a run tells you is which task statements
you get wrong and which wrong-answer habits you have — which is the part you can act on.
Why the simulation draws its scenarios instead of letting you pick
A real sitting presents 4 of the 6 published scenarios, so there are
15 possible draws. 14 of those
15 exercise all 5 domains
— the exceptions miss D4, which is reachable
only through S5 and S6.
You cannot know which draw you get, so the only rational preparation covers every scenario.
That is why this preset exposes a randomize flag and no scenario picker.
Draws reaching each domain, out of 15
Domain
Blueprint weight
Reachable via
Draws reaching it
D1 Agentic Architecture & Orchestration
27%
S1 S3 S4
15 / 15
D2 Tool Design & MCP Integration
18%
S1 S3 S4
15 / 15
D3 Claude Code Configuration & Workflows
20%
S2 S4 S5
15 / 15
D4 Prompt Engineering & Structured Output
20%
S5 S6
14 / 15
D5 Context Management & Reliability
15%
S1 S2 S3 S6
15 / 15
What the bank can currently supply
120 original items, each written against one of the
30 published task statements. 30 of 30 statements
have at least one item; 30 are at the 3-item floor. A run says exactly
how short it is rather than quietly serving a thinner paper.
5 domains, 6 published scenarios, 4 presented on a real sitting.
Runs on this device
Attempts are kept in this browser only, in the same shape the database uses, so nothing is lost
when they later sync. An unfinished run shows as abandoned — which is itself worth seeing.