Product · Evaluation
AI scientist interviews
Synthesis of four interview responses on ESTIMAND draft full-bank conclusions and what they imply for MeasureGPT. Conclusions only in the source packet—no score tables for interviewees.
Internal synthesis · 2026-07-24 · draft bankBuild MeasureGPT as bound evidence + guardrails, not “more search.” The trust story is blocked by E3 abstention until decline is first-class and the real stack (web + MCP) is measured. Market to tool-using / mid-model / audit-bearing buyers with where you win—not a single accuracy number.
Shared verdict
All four voices (Biomni A/B, Claude A/B) converge on these points.
- Internet is the product baseline. Closed-book inflates abstention (no search temptation) and is not a real deploy. Claims should lead with MCP vs web search.
- Biggest risk: E3 / abstention (C4). The product sells “don’t invent cutoffs,” but tools can still over-establish—worst for the strongest web model. Internet + MCP together has not been measured.
- No universal “MCP beats web / +X%” on the homepage. Lifts are model-conditional; draft status binds external claims. Lead with where and mechanism, not a single number.
- Ship a first-class decline contract. Typed not-in-corpus / not-established responses as successful, provenance-bearing results—not tool errors or silent missing fields.
- Next go/no-go: internet + MCP. Complement arm with arbitration policy (and schema A/B on decline). Pause marketing if E3 regresses vs internet-only (human-confirmed).
Where interviewers diverge
| Topic | Biomni | Claude |
|---|---|---|
| Strongest positive evidence | C1 — directional win vs internet for every model | C3 — gains land where theory predicted; C1 may be grader/format artifact |
| E2 win | Winning primary estimand | Same, but name it relay of provenance fields—not smarter attribution |
| E3 theory | Tool force / near-miss / skill displacement | Adds closure: web cannot observe absence; MCP can—if exposed |
| A2 fix | Tool-emitted mismatch flag | Population as mandatory structured I/O + explicit unspecified-in-source |
| Validity threat #1 | Bank sampling of corpus strengths | Criterion contamination (gold answers share provenance with corpus) |
| Beachhead | Mid-tier, owns the agent, regulated / audit-bearing | Clinical research / MBC / ePRO / trials eng—volume + auditor + fixed instruments |
Live MCP probe (Claude B)
Operational facts from probing the live tool surface at interview time—and what changed after the 2026-07-24 plan sprint:
- Was: unknown measure →
TOOL_EXECUTION_FAILED. Now: typed success withok:false,catalogStatus: "not-in-corpus"(item 1). - In-corpus uncertainty is already rich (
norm.suppressed,thresholdKind, disagreement arrays,tensions)—unchanged and still load-bearing. - Was:
interpret_scoretook no population argument. Now: optionalanalysisPopulation+ top-levelpopulationApplicability(matched / caution / blocker) (item 3). - Catalog shape (coverage snapshot): 206/317 with cutoffs, 72 with norms, 47 with both, 86 with neither—depth signal for beachhead packaging.
Merged 90-day plan
Priority 1 · Shipped (2026-07-24)
Decline contract
Never fail unknown measures; type not-in-corpus / parameter-unauthored / not-established-in-literature (keep distinct); provenance on negatives.
Priority 2 · Deferred (user)
internet + MCP complement arm
Corpus-first / web-first policies. Real deploy. Pause if E3 regresses vs internet-only.
Priority 3 · Shipped (2026-07-24)
Population applicability in the contract
Input + mismatch verdict—not only prose on output.
Priority 4 · Process + worksheets shipped; panel not run
E3 human audit
Gold-negative labels + scorer “false establishment” cells.
Priority 5 · Shipped (2026-07-24)
Honest packaging for the beachhead
Checkable answers, failure demos, coverage map/SLA. No “solves hallucination” until 1–2.
Priority 6 · Later (eval hygiene)
Later: confirmatory path
Holdout, dual search stacks, out-of-corpus / structural E3 negatives.
Marketing pause trigger (Agent A consensus): internet+MCP makes E3 worse than internet-only, human-confirmed. That is a safety signal, not a packaging nuance.
Arm design note: the draft ESTIMAND matrix used reference-mcp with internet OFF. Decision 2026-07-24: future reference-mcp runs will use internet ON (MeasureGPT + web). Until that re-run, MCP-only scores are not “MCP + default web.”
North-star metrics
Not overall accuracy. Track:
- Abstention fidelity — agent declines when the tool returns a typed negative.
- E3 under internet+MCP vs internet-only once that arm exists.
≤2-week experiments
- Schema A/B: explicit establishment_status / notInCorpus vs silence/exception — E3-only re-run.
- Human sample of E3 auto “false establishment” under internet & MCP for one model.
- Unknown-measure → typed not-in-corpus (not tool failure).
Implementation responses (2026-07-24 sprint)
What we shipped in MeasureGPT after the interview synthesis—each item used Codex plan review then implementation audit. Item 2 remains deferred by product decision.
1 — Decline / not-in-corpus
- Shared envelope:
ok:false,catalogStatus: "not-in-corpus",decline+provenance(catalog release, checkedAt). - Tools return typed success (MCP
isError: false) so agents abstain instead of treating tool failure as “try harder.” - Docs:
docs/service-contract-benchmark.md, ESTIMAND handoff for consumer parsers.
3 — Population applicability in the contract
- Optional
analysisPopulationoninterpret_score(requiresageBand+setting; strict object). - Top-level
populationApplicability: tier +comparisonStatus+ per-fact relations from existingmatchPopulation()algebra. - Honesty: no invented mismatch without tags; incomplete tags →
insufficient-tags/ caution; blocker does not setok:false.
4 — E3 human audit process
- SOP:
docs/e3-human-audit-process.md(dual raters, 15% FE disagreement hard gate, baseline marketing prohibition on abstention claims). - Machine-checkable worksheets + registry of 8 gold-negatives (
docs/e3-audit/,npm run audit:e3-validate). All items stillopenuntil a live panel runs—no fabricated adjudications.
5 — Honest beachhead packaging
- Claim language, failure demos (decline + population blocker), coverage map, SLA sketch:
docs/honest-beachhead-packaging.md. - Live coverage snapshot:
docs/coverage-snapshot.json(317 instruments, 47 with both cutoff and norm). Regenerate withnpm run docs:coverage-snapshot. - Explicit gates: no “solves hallucination / calibrated abstention” until E3 panel + item 2 complement arm clear.
Still open
- Item 2 — internet+MCP complement arm (realistic deploy; marketing pause if E3 regresses).
- Item 6 — confirmatory holdout, dual search stacks, out-of-corpus / structural E3 negatives.
- Live E3 panel — run dual raters against registry + ESTIMAND exports; until then gold-negatives remain unconfirmed.
Source interviews
Full write-ups live in the repo under docs/interviews/. Markdown synthesis: docs/interviews/SYNTHESIS.md. Packaging detail: docs/honest-beachhead-packaging.md.
measuremcp-biomni-response-A.md— Biomni · evaluation scientistmeasuremcp-biomni-response-B.md— Biomni · product / toolsmeasuremcp-claude-response-A.md— Claude · evaluation scientistmeasuremcp-claude-response-B.md— Claude · product / tools (+ live probe)
← MeasureGPT home · Architecture · Eval pilot
Draft full-bank narrative for internal product planning—not a sealed confirmatory report.