MeasureGPT
MeasureGPT is a psychometric interpretation service. Your AI system can call it. It gives severity bands, T-scores, percentiles, and citations. It gives a T-score and a percentile only when a percentile is defensible. You can check each citation.
MCP server · stdio + HTTP · no PHI · 317 instrumentsLanguage models invent cutoffs. They make errors in score calculations. They cite PMIDs that do not exist. MeasureGPT is the service behind the model. It calculates scores deterministically from a corpus that humans verified. You can check its calculations and its citations.
Get started
How to install MCP
Ask your ChatGPT, Claude, or Grok to install MeasureGPT as a remote MCP server — paste the line below into the chat:
install this MCP: https://measuregpt.onrender.com/mcpRemote endpoint: https://measuregpt.onrender.com/mcp
Demo
Try it
Do one of these tasks: paste a score table, select instruments, or run a multi-score profile. The demo uses the same engine as the MCP tools. You do not need an account.
Battery of unrelated instruments
Use this form for unrelated instruments in one session. PHQ-9 with GAD-7 is an example. Each row is a separate interpret_score call.
Ad-hoc multi-measure session: enter one score, add several rows, or paste a multi-measure table (PROMIS Outcomes-style: name, score, optional %ile/category). Each matched row calls interpret_score. For one multi-domain instrument (BPI, PROMIS-29, PROMIS Global, PROMIS Peds-25, SF-36, KOOS, HOOS, EQ-5D-5L), use Profile multi-score below. Unmatched instruments stay in the inventory as gaps — no blended global severity.
Profile multi-score
Use this form for one named instrument that has domain slots. The instruments are BPI, PROMIS-29, PROMIS Global, PROMIS Pediatric Profile-25, SF-36 (PCS and MCS, or eight domains), KOOS, HOOS, and EQ-5D-5L (index and VAS). The form uses interpret_profile. Each slot stays separate. The service does not blend the slots into one severity.
Loading multi-score profiles…
Research · exploratory pilot (deprecated)
An early pilot used the A–E item bank. The pilot results suggested large gains when models used the MeasureGPT tools. For example, citation accuracy increased from approximately 27% to 87% on that bank. The pilot is not a confirmatory leaderboard. The project has superseded the pilot methods, the pilot item bank, and the pilot scoring.
The next evaluation is ESTIMAND. ESTIMAND is a pre-registered benchmark of grounded psychometric reasoning on the frozen MeasureGPT corpus. The confirmatory run is still ahead. The pilot numbers below are historical data only.
Background
Why it exists
Language models make three types of error on questionnaire scores. The first type is a fact error: an incorrect cutoff, an incorrect MCID, or a norm that is out of date. The second type is an arithmetic error: an incorrect conversion from a raw score to a z-score, a T-score, or a percentile. The third type is a citation error: a PMID that does not exist. MeasureGPT moves these three tasks to a deterministic engine. The engine uses a corpus that humans verified. Each keyed fact has a quotation from a primary source. The service uses public psychometric facts only. It does not use PHI.
MCP tools
list_measuresThis tool gives the supported measures. Each measure has an id, a construct, a score type, and a range.get_measureThis tool gives the full scoring rubric. The rubric has the bands, the cutoff, the norms, the citations, and the caveats.interpret_scoreThis tool converts a score to a band. It gives the cutoff status with the sensitivity and the specificity. It gives the z-score, the T-score, and the percentile. If a percentile is not defensible, the tool withholds both the T-score and the percentile, and it gives the reason. The tool also gives a summary and the safety caveats.get_validation_studiesThis tool gives the cited validation literature and reference literature for a measure.list_profilesThis tool gives the multi-score profile instruments and their slot ids.interpret_profileThis tool converts the domain scores of one profile to one interpretation for each slot. It does not blend the slots into one severity.Connect
Connect
Configure ChatGPT, Claude, or Grok to use the remote MCP server at https://measuregpt.onrender.com/mcp. The install instructions are at the top of this page. The server uses streamable HTTP at /mcp. Use this host or your own deployment.
On-page assistant
Off by default. The deployed service has no server API key. To enable the assistant on your computer, set OPENROUTER_API_KEY, run npm run dev, and then reload this page.