MeasureGPT · Research
Exploratory pilot — deprecated
The original MeasureGPT A–E item bank was a small, exploratory pilot used while the corpus and tool surface were still moving. It is not a confirmatory leaderboard and should not be cited as a definitive product benchmark.
What those numbers were
Archival pilot deltas
On a fixed A–E pilot bank, five frontier models were run closed-book and again with MeasureGPT tools. Rough deltas on that bank included citation accuracy about 27% → 87%, computation 63% → 90%, and factual accuracy 74% → 89%. Those figures motivated better evaluation design; they do not survive as the public headline claim for MeasureGPT.
- Small item bank; not pre-registered for a confirmatory claim
- Taxonomy and scoring evolved after the pilot (superseded by ESTIMAND contracts)
- Useful as exploratory evidence of tool-use failure modes — not as a final ranking
What replaces it
ESTIMAND (coming)
ESTIMAND is the next benchmark of grounded psychometric reasoning: whether a model can retrieve a psychometric fact, attribute it, compute with it, and abstain when the fact or the question does not support an answer — scored on accuracy and grounding fidelity as separate axes.
MeasureGPT remains the frozen instrument corpus and the author-built reference-mcp implementation evaluated by that benchmark (conflict of interest disclosed there). The confirmatory ESTIMAND run is not yet complete; no public leaderboard cells should be read as final until that pre-registered evaluation lands.
Project home: github.com/koolkao/estimand. The historical pilot was migrated there as an immutable archive only.