← MeasureGPT · Eval pilot

Product · Evaluation

AI scientist interviews

Synthesis of four interview responses on ESTIMAND draft full-bank conclusions and what they imply for MeasureGPT. Conclusions only in the source packet—no score tables for interviewees.

Internal synthesis · 2026-07-24 · draft bank

Build MeasureGPT as bound evidence + guardrails, not “more search.” The trust story is blocked by E3 abstention until decline is first-class and the real stack (web + MCP) is measured. Market to tool-using / mid-model / audit-bearing buyers with where you win—not a single accuracy number.

Shared verdict

All four voices (Biomni A/B, Claude A/B) converge on these points.

  1. Internet is the product baseline. Closed-book inflates abstention (no search temptation) and is not a real deploy. Claims should lead with MCP vs web search.
  2. Biggest risk: E3 / abstention (C4). The product sells “don’t invent cutoffs,” but tools can still over-establish—worst for the strongest web model. Internet + MCP together has not been measured.
  3. No universal “MCP beats web / +X%” on the homepage. Lifts are model-conditional; draft status binds external claims. Lead with where and mechanism, not a single number.
  4. Ship a first-class decline contract. Typed not-in-corpus / not-established responses as successful, provenance-bearing results—not tool errors or silent missing fields.
  5. Next go/no-go: internet + MCP. Complement arm with arbitration policy (and schema A/B on decline). Pause marketing if E3 regresses vs internet-only (human-confirmed).

Where interviewers diverge

TopicBiomniClaude
Strongest positive evidenceC1 — directional win vs internet for every modelC3 — gains land where theory predicted; C1 may be grader/format artifact
E2 winWinning primary estimandSame, but name it relay of provenance fields—not smarter attribution
E3 theoryTool force / near-miss / skill displacementAdds closure: web cannot observe absence; MCP can—if exposed
A2 fixTool-emitted mismatch flagPopulation as mandatory structured I/O + explicit unspecified-in-source
Validity threat #1Bank sampling of corpus strengthsCriterion contamination (gold answers share provenance with corpus)
BeachheadMid-tier, owns the agent, regulated / audit-bearingClinical research / MBC / ePRO / trials eng—volume + auditor + fixed instruments

Live MCP probe (Claude B)

Operational facts from probing the live tool surface at interview time—and what changed after the 2026-07-24 plan sprint:

Merged 90-day plan

Priority 1 · Shipped (2026-07-24)

Decline contract

Never fail unknown measures; type not-in-corpus / parameter-unauthored / not-established-in-literature (keep distinct); provenance on negatives.

Priority 2 · Deferred (user)

internet + MCP complement arm

Corpus-first / web-first policies. Real deploy. Pause if E3 regresses vs internet-only.

Priority 3 · Shipped (2026-07-24)

Population applicability in the contract

Input + mismatch verdict—not only prose on output.

Priority 4 · Process + worksheets shipped; panel not run

E3 human audit

Gold-negative labels + scorer “false establishment” cells.

Priority 5 · Shipped (2026-07-24)

Honest packaging for the beachhead

Checkable answers, failure demos, coverage map/SLA. No “solves hallucination” until 1–2.

Priority 6 · Later (eval hygiene)

Later: confirmatory path

Holdout, dual search stacks, out-of-corpus / structural E3 negatives.

Marketing pause trigger (Agent A consensus): internet+MCP makes E3 worse than internet-only, human-confirmed. That is a safety signal, not a packaging nuance.

Arm design note: the draft ESTIMAND matrix used reference-mcp with internet OFF. Decision 2026-07-24: future reference-mcp runs will use internet ON (MeasureGPT + web). Until that re-run, MCP-only scores are not “MCP + default web.”

North-star metrics

Not overall accuracy. Track:

  1. Abstention fidelity — agent declines when the tool returns a typed negative.
  2. E3 under internet+MCP vs internet-only once that arm exists.

≤2-week experiments

  1. Schema A/B: explicit establishment_status / notInCorpus vs silence/exception — E3-only re-run.
  2. Human sample of E3 auto “false establishment” under internet & MCP for one model.
  3. Unknown-measure → typed not-in-corpus (not tool failure).

Implementation responses (2026-07-24 sprint)

What we shipped in MeasureGPT after the interview synthesis—each item used Codex plan review then implementation audit. Item 2 remains deferred by product decision.

1 — Decline / not-in-corpus

3 — Population applicability in the contract

4 — E3 human audit process

5 — Honest beachhead packaging

Still open

Source interviews

Full write-ups live in the repo under docs/interviews/. Markdown synthesis: docs/interviews/SYNTHESIS.md. Packaging detail: docs/honest-beachhead-packaging.md.

← MeasureGPT home · Architecture · Eval pilot

Draft full-bank narrative for internal product planning—not a sealed confirmatory report.