MeasureGPT

MeasureGPT is a psychometric interpretation service. Your AI system can call it. It gives severity bands, T-scores, percentiles, and citations. It gives a T-score and a percentile only when a percentile is defensible. You can check each citation.

MCP server · stdio + HTTP · no PHI · 317 instruments

Language models invent cutoffs. They make errors in score calculations. They cite PMIDs that do not exist. MeasureGPT is the service behind the model. It calculates scores deterministically from a corpus that humans verified. You can check its calculations and its citations.

Get started

How to install MCP

Ask your ChatGPT, Claude, or Grok to install MeasureGPT as a remote MCP server — paste the line below into the chat:

install this MCP: https://measuregpt.onrender.com/mcp

Remote endpoint: https://measuregpt.onrender.com/mcp

Demo

Try it

Do one of these tasks: paste a score table, select instruments, or run a multi-score profile. The demo uses the same engine as the MCP tools. You do not need an account.

Battery of unrelated instruments

Use this form for unrelated instruments in one session. PHQ-9 with GAD-7 is an example. Each row is a separate interpret_score call.

Ad-hoc multi-measure session: enter one score, add several rows, or paste a multi-measure table (PROMIS Outcomes-style: name, score, optional %ile/category). Each matched row calls interpret_score. For one multi-domain instrument (BPI, PROMIS-29, PROMIS Global, PROMIS Peds-25, SF-36, KOOS, HOOS, EQ-5D-5L), use Profile multi-score below. Unmatched instruments stay in the inventory as gaps — no blended global severity.

Profile multi-score

Use this form for one named instrument that has domain slots. The instruments are BPI, PROMIS-29, PROMIS Global, PROMIS Pediatric Profile-25, SF-36 (PCS and MCS, or eight domains), KOOS, HOOS, and EQ-5D-5L (index and VAS). The form uses interpret_profile. Each slot stays separate. The service does not blend the slots into one severity.

Loading multi-score profiles…

Research · exploratory pilot (deprecated)

An early pilot used the A–E item bank. The pilot results suggested large gains when models used the MeasureGPT tools. For example, citation accuracy increased from approximately 27% to 87% on that bank. The pilot is not a confirmatory leaderboard. The project has superseded the pilot methods, the pilot item bank, and the pilot scoring.

The next evaluation is ESTIMAND. ESTIMAND is a pre-registered benchmark of grounded psychometric reasoning on the frozen MeasureGPT corpus. The confirmatory run is still ahead. The pilot numbers below are historical data only.

27% → 87%Pilot citations
63% → 90%Pilot computation
74% → 89%Pilot facts
Pilot status and ESTIMAND →

Background

Why it exists

Language models make three types of error on questionnaire scores. The first type is a fact error: an incorrect cutoff, an incorrect MCID, or a norm that is out of date. The second type is an arithmetic error: an incorrect conversion from a raw score to a z-score, a T-score, or a percentile. The third type is a citation error: a PMID that does not exist. MeasureGPT moves these three tasks to a deterministic engine. The engine uses a corpus that humans verified. Each keyed fact has a quotation from a primary source. The service uses public psychometric facts only. It does not use PHI.

MCP tools

list_measuresThis tool gives the supported measures. Each measure has an id, a construct, a score type, and a range.
get_measureThis tool gives the full scoring rubric. The rubric has the bands, the cutoff, the norms, the citations, and the caveats.
interpret_scoreThis tool converts a score to a band. It gives the cutoff status with the sensitivity and the specificity. It gives the z-score, the T-score, and the percentile. If a percentile is not defensible, the tool withholds both the T-score and the percentile, and it gives the reason. The tool also gives a summary and the safety caveats.
get_validation_studiesThis tool gives the cited validation literature and reference literature for a measure.
list_profilesThis tool gives the multi-score profile instruments and their slot ids.
interpret_profileThis tool converts the domain scores of one profile to one interpretation for each slot. It does not blend the slots into one severity.

Connect

Connect

Configure ChatGPT, Claude, or Grok to use the remote MCP server at https://measuregpt.onrender.com/mcp. The install instructions are at the top of this page. The server uses streamable HTTP at /mcp. Use this host or your own deployment.

On-page assistant

Off by default. The deployed service has no server API key. To enable the assistant on your computer, set OPENROUTER_API_KEY, run npm run dev, and then reload this page.