MeasureGPT · Research

Exploratory pilot — deprecated

The original MeasureGPT A–E item bank was a small, exploratory pilot used while the corpus and tool surface were still moving. It is not a confirmatory leaderboard and should not be cited as a definitive product benchmark.

What those numbers were

Archival pilot deltas

On a fixed A–E pilot bank, five frontier models were run closed-book and again with MeasureGPT tools. Rough deltas on that bank included citation accuracy about 27% → 87%, computation 63% → 90%, and factual accuracy 74% → 89%. Those figures motivated better evaluation design; they do not survive as the public headline claim for MeasureGPT.

What replaces it

ESTIMAND (coming)

ESTIMAND is the next benchmark of grounded psychometric reasoning: whether a model can retrieve a psychometric fact, attribute it, compute with it, and abstain when the fact or the question does not support an answer — scored on accuracy and grounding fidelity as separate axes.

MeasureGPT remains the frozen instrument corpus and the author-built reference-mcp implementation evaluated by that benchmark (conflict of interest disclosed there). The confirmatory ESTIMAND run is not yet complete; no public leaderboard cells should be read as final until that pre-registered evaluation lands.

Project home: github.com/koolkao/estimand. The historical pilot was migrated there as an immutable archive only.

← Back to MeasureGPT · Architecture