Lab: MCPcast Your SaaS API¶
Your customers keep asking the same question: "does it work with Claude?"
You could build a chatbot into your product. Almost everyone does, and almost nobody uses it. The other answer is better: ship an MCP server for the API you already have, and let every customer bring their own AI — Claude Desktop, Claude Code, Cursor, an internal agent, whatever they already run.
The hard part was never the protocol. It is that a REST API is not a tool
surface: 200 endpoints is not a set of tools, POST /orders/{id}/refund must
never fire without a human, and no one can tell you whether an agent can
actually drive what you shipped.
This lab does the whole thing for a fictional commerce SaaS, end to end, with runnable code you can point at your own spec afterwards.
What You'll Build¶
From one OpenAPI file:
- an MCP server as editable Python — not a black box, not a proxy: an
installable package with its own tests that depends only on
promptiseandhttpx - a risk-classified tool surface where every read is open, every write is approval-gated, and every excluded operation is recorded with a reason
- multi-tenant auth: each customer's MCP key maps to their own upstream credential, and reviewers only ever see their own tenant's calls
- a four-eyes approval gate enforced server-side — a refund waits for a second human, for any MCP client, and is denied on timeout
- an Agent Readiness Score: a real agent, real tasks, a grade, and a list of specific tool-design fixes
- a measured before/after when you apply those fixes
Everything runs on your machine. The upstream API is faked in-process (the way you would fake Stripe in a test suite); the agents, tools, approvals and scores are real.
Prerequisites¶
The complete lab lives in examples/mcp/mcpcast_storefront_lab/:
| File | What it is |
|---|---|
storefront.yaml |
The API the company already ships — 13 operations |
fake_api.py |
The upstream, as an in-process httpx.MockTransport that records every request |
run.py |
The lab itself, in five printed sections |
The API you already have¶
Nothing about storefront.yaml is unusual — that is the point. Customers,
orders, a refund endpoint, a subscription cancellation, one admin route, one
deprecated CSV export, one multipart upload:
/orders/{order_id}/refund:
post:
operationId: refundOrder
tags: [billing]
summary: Refund an order
description: >
Move money back to the customer's original payment method. Partial
refunds are allowed; the sum of refunds may not exceed the order total.
parameters:
- $ref: "#/components/parameters/OrderId"
requestBody:
required: true
content:
application/json:
schema:
type: object
required: [amount]
properties:
amount:
type: number
description: Amount to refund, in the order's currency.
Step 1 — Generate the server, and read what it decided¶
One call does parse → classify → plan → emit. write_project puts three files
on disk: the plan, the server, and a README for whoever runs it.
from promptise.mcpcast import AuthMode, SafetyProfile, mcpcast, write_project
plan = mcpcast(SPEC, profile=SafetyProfile.FULL, auth=AuthMode.API_KEY, name="storefront")
write_project(plan, OUT) # the plan, the package, the launcher, tests, scaffold
The CLI does the same, and adds the LLM curation pass and the review screen:
The interesting part is not that it generated something — it is that it can tell you why it generated exactly this:
tool risk approval upstream operation
get_health read - GET /health
list_customers read - GET /customers
search_customers read - POST /customers/search
get_customer read - GET /customers/{customer_id}
list_orders read - GET /orders
create_order write human POST /orders
get_order read - GET /orders/{order_id}
refund_order financial human POST /orders/{order_id}/refund
cancel_subscription destructive human DELETE /subscriptions/{subscription_id}
get_revenue_report write human GET /reports/revenue
get_customer_pii write human GET /admin/customers/{customer_id}/pii
why the classifier decided that (deterministic, no model):
search_customers POST that only queries ('search')
refund_order mentions money ('refund')
cancel_subscription DELETE is destructive
get_revenue_report GET is a read; escalated: requires scope 'admin:read'
get_customer_pii GET is a read; escalated: path is admin-only
not exposed (nothing vanishes silently):
uploadOrderAttachment: unsupported by mcpcast: unsupported request body media type multipart/form-data
exportLegacyOrders: deprecated in spec
the same spec under each safety profile:
read-only 6 tools (0 approval-gated), 7 not exposed
standard 9 tools (3 approval-gated), 4 not exposed
full 11 tools (5 approval-gated), 2 not exposed
auth: api-key approval: pending (four-eyes)
Four things happened there that you would otherwise have hand-written and got wrong:
POST /customers/searchis a read. Method is not intent. A POST whose leading verb is a query verb, with no mutating verb anywhere, is classifiedread— so it survives theread-onlyprofile.- Two GETs were escalated.
GET /reports/revenuecarries the OAuth scopeadmin:read;GET /admin/customers/{id}/piisits under/admin. Both are HTTP reads that no one should hand an agent by default, so they come out aswrite— exposed only fromstandardup, and approval-gated. - Nothing disappeared quietly. The multipart upload cannot be sent by the
generated client and the CSV export is
deprecated: truein the spec; both are inplan.droppedwith the reason, not silently missing. - The profile is the dial. Same spec, three surfaces.
read-onlyis the default because it is the one you can ship to strangers.
The plan file is the source of truth
mcpcast.plan.yaml is what you edit — tool names, descriptions, examples,
hidden parameters, the dropped list — and then regenerate with
promptise mcpcast mcpcast.plan.yaml. Your edits survive; the package and
server.py are always rebuilt from the plan. See
the plan file.
Step 2 — Point a real agent at it¶
The generated server is a normal MCPServer, so an agent can drive it
in-process through TestClient — the full pipeline (validation, guards,
middleware, approval gate, handler) with no ports involved.
tools_from_server() does the bridging, and a CallRecorder notes every tool
the agent picks.
async with api.client() as http: # the fake upstream
server = module.build_server(http_client=http)
recorder = CallRecorder()
recorder.begin("lookup")
tools = await tools_from_server(server, recorder=recorder,
headers={"x-api-key": AGENT_KEY})
agent = await build_agent(model="openai:gpt-5-mini", servers={},
extra_tools=tools, instructions=...)
result = await agent.ainvoke({"messages": [HumanMessage(content=question)]})
question: A customer wrote in from ada@northwind.example. Who are they, and what are their two most recent orders?
tools the agent chose:
search_customers({"email": "ada@northwind.example"}) -> ok
list_orders({"customer_id": "CUS-1001", "limit": 2}) -> ok
answer: They’re Ada Lovelace (customer ID CUS-1001), email ada@northwind.example — on the Pro plan (created 2025-03-04).
Two most recent orders:
- ORD-1002
- Status: shipped
- Placed: 2026-02-11
- Total: €49.90
- Items: 1 × HUB-USBC-7 (unit price €49.90)
...
requests the API received: ['POST /v1/customers/search', 'GET /v1/orders?customer_id=CUS-1001&limit=2']
the credential each one carried: 'Bearer northwind-upstream-token'
No prompt engineering, no tool wiring, no glue: the agent chained a search into a filtered list because the generated descriptions and examples told it how.
That last line is the multi-tenancy in action. Under --auth api-key the
generated server runs with require_tenant=True: the MCP client presents a key
from MCPCAST_CLIENT_KEYS, the key names a tenant, and the tenant selects the
upstream credential from MCPCAST_UPSTREAM_TOKENS — read per call, so rotating a
customer's token needs no restart.
CLIENT_KEYS = {
AGENT_KEY: {"client_id": "support-agent", "tenant_id": "northwind"},
REVIEWER_KEY: {"client_id": "dana", "tenant_id": "northwind", "roles": ["approver"]},
OTHER_TENANT_KEY: {"client_id": "gus", "tenant_id": "globex", "roles": ["approver"]},
}
UPSTREAM_TOKENS = {"northwind": "Bearer northwind-upstream-token", ...}
Step 3 — The refund that waits for a human¶
Now the agent is asked to move money. Same server, same agent, one different sentence:
request: Order ORD-1003 arrived damaged. Refund the customer 24 euros for it.
the gate is holding the call:
refund_order {"order_id": "ORD-1003", "amount": 24.0, "reason": "Item arrived damaged"}
requested by client_id='support-agent' tenant='northwind'
refund requests the API has received so far: 0
API log so far: []
The agent called the tool. The call is parked — because this is what the
generator wrote for that operation, in storefront_mcp/tools/billing.py, for
you to read:
# -- refund_order (financial, requires approval) -----------------------------
@server.tool(
name="refund_order",
description=(
"Refund an order. Move money back to the customer's original payment method. Partial "
"refunds are allowed; the sum of refunds may not exceed the order total.\n"
"\n"
"Parameters:\n"
" - order_id (string, required): The order identifier, e.g. `ORD-1002`.\n"
" - amount (number, required): Amount to refund, in the order's currency.\n"
" - reason (string): Why the refund was issued (kept for the audit trail).\n"
"\n"
'Example: {"order_id": "ORD-1002", "amount": 24}'
),
tags=["billing"],
open_world_hint=True,
requires_approval=True,
)
async def refund_order(
ctx: RequestContext,
order_id: str,
amount: float,
reason: str | None = None,
) -> Any:
"Refund an order. Move money back to the customer's original payment method.…"
args: dict[str, Any] = {
"order_id": order_id,
"amount": amount,
"reason": reason,
}
return await upstream.call(select_route(ROUTES["refund_order"], args), args, ctx)
ApprovalGateMiddleware holds a requires_approval call before the handler
runs. The upstream API received nothing at all — that
empty log is the whole point. If nobody decides within
MCPCAST_APPROVAL_TIMEOUT, the call is denied, not allowed.
Because the callers are identified (api-key auth), the approver is the pending
one: an independent human, reached through two generated, tenant-scoped tools —
approvals_list and approvals_decide, both guarded by the approver role.
Approving is itself a tool call, so a reviewer can be any MCP client — a dashboard, a Slack bot, or a second Claude session:
(reply,) = await reviewer.call_tool(
"approvals_decide",
{"request_id": pending["request_id"], "approve": True, "reason": "damaged goods"},
headers={"x-api-key": REVIEWER_KEY}, # dana, tenant northwind, role approver
)
a reviewer from ANOTHER tenant:
globex/gus approvals_list -> [] (sees nothing)
globex/gus tries to approve -> NOT_FOUND
the caller approving their own request (same principal, approver role):
-> ACCESS_DENIED: A caller may not approve their own request
dana (northwind, approver role) approves:
approvals_decide -> {"request_id": "73c810f0...", "approved": true, "resolved": true}
answer: Refund successful — REF-3001: €24 refunded to order ORD-1003 (reason: Item arrived damaged).
refund requests the API received: 1
POST /v1/orders/ORD-1003/refund body={"amount": 24.0, "reason": "Item arrived damaged"}
authorization='Bearer northwind-upstream-token' (northwind's upstream credential)
Four attempts and what the server did with each — enforced by the server, not by the client's good manners:
| Attempt | Outcome |
|---|---|
| A reviewer of another tenant lists pending calls | sees nothing — reviewers never see other tenants' arguments |
| That reviewer decides by request id anyway | NOT_FOUND — the id does not exist for them |
The caller approves their own request, holding the approver role |
ACCESS_DENIED — four-eyes, always |
| A different human of the same tenant approves | the call resumes and hits the API, once |
This is why the gate belongs in the server. Your customer might connect with Claude Desktop today and a homegrown agent tomorrow; neither can opt out of a policy that lives on your side of the wire.
Which approver you get
--auth api-key defaults to pending (four-eyes review, shown here).
passthrough, env-token and none default to elicitation: the human
behind the calling client confirms the action in their own UI — the right
shape for a personal server launched by Claude Desktop over stdio. Both are
fail-closed. See Human approval.
Step 4 — Measure: the Agent Readiness Score¶
You now have a server that is safe. Safe is not the same as usable: the real
question is whether an agent picks the right tool from your descriptions.
evaluate() answers it with evidence — a real agent, real tasks, the full
server pipeline, and mocks derived from your spec so nothing can touch real data.
report = await evaluate(
plan, module.build_server,
model="openai:gpt-5-mini",
tasks=tasks, # eight explicit EvalTasks; --eval writes them for you
operations=operations, # response schemas -> realistic mocks
live_reads=False, # this API is fictional; mock every route
headers={"x-api-key": AGENT_KEY},
)
Agent Readiness: A score 1.00
tasks succeeded 8/8 correct tool first 100% parameter errors 0%
t1 expected search_customers called search_customers ok
t2 expected get_order called get_order ok
t3 expected list_orders called list_orders ok
t4 expected refund_order called refund_order ok
t5 expected get_revenue_report called get_revenue_report ok
t6 expected get_customer called get_customer -> get_customer ok
t7 expected list_customers called list_customers ok
t8 expected create_order called create_order ok
what the report says to fix:
• 3 tools not covered by any task: `get_health`, `cancel_subscription`, `get_customer_pii` — raise --eval-tasks to score them
• `list_customers` has no example — agents lean on examples heavily
• `list_orders` has no example — agents lean on examples heavily
wrote generated/v1/eval/tasks.yaml and generated/v1/eval/report.md
The score is 0.6 × task success + 0.4 × correct-tool-first, graded A–F, and the
fixes are the payload: exact tool pairs the agent confuses, parameters that
produced validation errors, tools no task covered, tools a task needed and the
agent never reached for.
A grade of A on eight tasks is not the interesting part — the three lines under it are. Two tools have no worked example, and three tools are not covered by any task at all, which means the run says nothing about them. That is the honest answer to "is my API agent-ready?": here is what I measured, here is what I did not, here is what to fix first.
--eval writes the same thing to disk — eval/tasks.yaml (rerun the exact
tasks) and eval/report.md (the table above, in Markdown) — so a readiness score
can live in CI next to your other tests.
An evaluation never changes your data
The live/mock split is by risk class, not HTTP method: only routes of
read tools may reach the real API (and only with live_reads=True).
Everything else — writes, refunds, cancellations, and any GET the classifier
escalated — is answered by a spec-derived mock behind an auto-approver.
Step 5 — Apply the fixes, regenerate, re-measure¶
The report is advice, and the plan is code, so the fixes are a few lines. Here
the lab merges the two customer lookups into one intent tool with two routes —
the agent should be picking a customer, not an HTTP endpoint — adds the examples
the report asked for, drops the health check no task needs, and puts
notify_customer on a param diet so an agent can never email a customer:
find_customer = ToolPlan(
name="find_customer",
description=(
"Find ONE customer, by id or by email address. Pass customer_id when you "
"know it (e.g. 'CUS-1001'), otherwise pass email. Use list_customers only "
"to browse or filter the whole directory."
),
risk=RiskClass.READ,
routes=[get_customer.routes[0], search_customers.routes[0]], # dispatch by what you pass
params={"customer_id": ParamPlan(...), "email": ParamPlan(...)},
example={"email": "ada@northwind.example"},
)
params["notify_customer"] = params["notify_customer"].model_copy(
update={"hidden": True, "default": False}
)
A multi-route tool dispatches to the first route whose required parameters were
supplied, so find_customer(customer_id=…) hits GET /customers/{id} and
find_customer(email=…) hits POST /customers/search. The hidden parameter is
not hidden from people: its fixed value is printed in the tool description and
in the generated README, so a reviewer can see what actually runs.
In the regenerated mcpcast.plan.yaml the merge is one tool with two routes:
- name: find_customer
description: Find ONE customer, by id or by email address. Pass customer_id when you know it (e.g. 'CUS-1001'),
otherwise pass email. Use list_customers only to browse or filter the whole directory.
risk: read
routes:
- operation_id: getCustomer
method: GET
path: /customers/{customer_id}
params:
customer_id:
location: path
required: true
- operation_id: searchCustomers
method: POST
path: /customers/search
params:
email:
location: body
required: true
params:
customer_id:
description: The customer identifier, e.g. 'CUS-1001'.
email:
description: The customer's email address, matched exactly.
json_schema:
type: string
format: email
example:
email: ada@northwind.example
tags:
- customers
Then the same eight requests run again against the regenerated server:
merged get_customer + search_customers -> find_customer (2 routes, dispatched by which id you pass)
added the missing examples; dropped get_health
create_order description now says: 'Always sends: notify_customer=false'
surface: 11 -> 9 tools, 8 -> 9 of them with a worked example
Agent Readiness: A score 1.00
tasks succeeded 8/8 correct tool first 100% parameter errors 0%
t1 expected find_customer called find_customer ok
t2 expected get_order called get_order ok
t3 expected list_orders called list_orders ok
t4 expected refund_order called refund_order ok
t5 expected get_revenue_report called get_revenue_report ok
t6 expected find_customer called find_customer ok
t7 expected list_customers called list_customers ok
t8 expected create_order called create_order ok
what the report says to fix:
• 2 tools not covered by any task: `cancel_subscription`, `get_customer_pii` — raise --eval-tasks to score them
before: A (1.00) after: A (1.00)
identical score on this run — the surface was already unambiguous for these eight
requests, and the edits kept it that way with two fewer tools.
The grade did not move, and the lab says so. That is what an honest metric looks like: this surface was already unambiguous for these eight requests, so the score had nowhere to go. What did change is worth having anyway — two fewer tools to choose from and to pay tokens for, every tool carrying a worked example, and a fix list that shrank from three items to one. On a real API with a hundred endpoints the first run is rarely an A, and the confusion pairs the report prints are usually the difference between a demo and a product.
One row is worth reading twice. In the first run t6 shows
get_customer -> get_customer: the agent asked again, because the mock answers
every customer id with the spec's canned Customer example, so the record that
came back did not carry the id it had asked for. That wobble is mock-shaped, not
tool-shaped — mocks are how an evaluation stays safe, and the exact row varies
from run to run because the model does. Run with live_reads=True
(and a real credential in MCPCAST_EVAL_AUTHORIZATION) when you want reads to hit
the real API and the numbers to reflect real payloads.
Step 6 — Hand it to Claude Desktop, Claude Code or Cursor¶
For a personal server that a desktop client launches over stdio, generate it
with --auth env-token so it carries one credential — yours:
promptise mcpcast examples/mcp/mcpcast_storefront_lab/storefront.yaml \
--profile standard --auth env-token --out storefront-mcp
{
"mcpServers": {
"storefront": {
"command": "/absolute/path/to/.venv/bin/python",
"args": ["/absolute/path/to/storefront-mcp/server.py"],
"env": {
"MCPCAST_UPSTREAM_TOKEN": "Bearer <your Storefront API token>",
"MCPCAST_BASE_URL": "https://api.storefront.example/v1"
}
}
}
}
Claude Code takes the same command and the same token, passed with -e so it
lands in the server's environment — the client launches the server itself and
does not see your shell's exports:
claude mcp add storefront -e MCPCAST_UPSTREAM_TOKEN="Bearer <your Storefront API token>" \
-- /absolute/path/to/.venv/bin/python /absolute/path/to/storefront-mcp/server.py
Cursor's .cursor/mcp.json uses the identical command/args/env shape. In
all three the write tools still stop and ask the human before they run — that
policy came with the server, not with the client.
For the multi-tenant deployment your customers connect to, keep --auth api-key
and serve the same file over HTTP — the generated module exposes a module-level
server, so promptise serve picks it up:
# Mint real keys: python -c 'import secrets; print("sk-" + secrets.token_urlsafe(32))'
# (the server refuses placeholders and the documentation's sample keys)
export MCPCAST_CLIENT_KEYS='{"sk-<agent key>": {"client_id": "acme-agent", "tenant_id": "acme"},
"sk-<reviewer key>": {"client_id": "dana", "tenant_id": "acme",
"roles": ["approver"]}}'
export MCPCAST_UPSTREAM_TOKENS='{"acme": "Bearer <acme upstream token>"}'
promptise serve server:server --transport http --port 8080
What You've Built¶
- An MCP server for an API you did not write a line of glue for — 13 spec operations became 11 tools, with 5 of them behind a human.
- A safety story you can explain to a security reviewer: risk classes with reasons, three profiles, drops with reasons, and approval enforced at the tool.
- Per-tenant isolation: one server, many customers, each with their own upstream credential, and reviewers who only see their own tenant.
- Evidence instead of vibes: a grade for how well an agent drives your tool surface, and a list of what to fix — with a before/after when you fix it.
- Editable output:
mcpcast.plan.yamlis yours, the package is regenerated from it, and all of it is plain, reviewable, diffable Python.
The whole lab¶
"""Lab: MCPcast the Storefront API — generate, drive, govern, measure, improve.
A fictional commerce SaaS already ships a REST API (``storefront.yaml``). It does
not want to build a chatbot; it wants its customers to use Claude, Cursor or any
other AI client *with the product*. This lab does that end to end, offline except
for the real model calls:
1. GENERATE ``promptise.mcpcast`` turns the spec into an editable project —
a risk-classified tool plan, ``server.py`` and a README — under the ``full``
profile with ``--auth api-key``. Every operation it leaves out is recorded
with a reason.
2. DRIVE A real ``build_agent("openai:gpt-5-mini")`` answers a business
question through the generated server, in-process, against a fake upstream
that records every HTTP request it receives.
3. GOVERN The agent tries to refund an order. The server-side approval gate
holds the call: nothing reaches the API until a *different* human of the
same tenant approves it — and reviewers of other tenants see nothing.
4. MEASURE The Agent Readiness Score: real tasks, real agent, spec-derived
mocks, a grade and specific tool-design fixes.
5. IMPROVE Apply the fixes to the plan in code, regenerate, re-run the same
tasks, and compare the grades honestly.
Only the upstream HTTP API is faked (``fake_api.py``), the way you would fake
Stripe in your own test suite. The agents, tools and approvals are real.
Run:
OPENAI_API_KEY=... .venv/bin/python examples/mcp/mcpcast_storefront_lab/run.py
"""
from __future__ import annotations
import asyncio
import json
import os
import sys
from pathlib import Path
from types import ModuleType
from typing import Any
from fake_api import FakeStorefront
from langchain_core.messages import HumanMessage
from promptise import build_agent
from promptise.mcp.server import TestClient
from promptise.mcpcast import (
AuthMode,
CallRecorder,
DroppedOp,
EvalReport,
EvalTask,
MCPcastPlan,
ParamPlan,
RiskClass,
SafetyProfile,
ToolPlan,
classify,
evaluate,
extract_operations,
load_generated_server,
load_spec,
mcpcast,
tools_from_server,
write_eval,
write_project,
)
from promptise.models import load_dotenv_if_present
HERE = Path(__file__).resolve().parent
# Passed by its relative name (main() runs from this directory), so the `spec_source`
# recorded in the plan and generated docstrings is portable, never an absolute path.
SPEC = "storefront.yaml"
OUT = HERE / "generated" / "v1"
OUT_V2 = HERE / "generated" / "v2"
MODEL = "openai:gpt-5-mini"
# MCP clients authenticate with an API key; each key names the tenant whose
# upstream credential the server presents. Reviewers carry the "approver" role.
AGENT_KEY = "sk-northwind-agent"
REVIEWER_KEY = "sk-northwind-dana"
SELF_APPROVE_KEY = "sk-northwind-agent-approver"
OTHER_TENANT_KEY = "sk-globex-gus"
CLIENT_KEYS = {
AGENT_KEY: {"client_id": "support-agent", "tenant_id": "northwind"},
REVIEWER_KEY: {"client_id": "dana", "tenant_id": "northwind", "roles": ["approver"]},
# The same principal as the caller, holding the approver role — four-eyes
# still refuses: nobody approves their own request.
SELF_APPROVE_KEY: {
"client_id": "support-agent",
"tenant_id": "northwind",
"roles": ["approver"],
},
OTHER_TENANT_KEY: {"client_id": "gus", "tenant_id": "globex", "roles": ["approver"]},
}
UPSTREAM_TOKENS = {
"northwind": "Bearer northwind-upstream-token",
"globex": "Bearer globex-upstream-token",
}
def section(number: int, title: str) -> None:
"""Print a numbered section header."""
print(f"\n{'=' * 78}\n{number}. {title}\n{'=' * 78}")
def configure_environment() -> None:
"""Configure the generated server before it is imported.
``MCPCAST_CLIENT_KEYS`` and ``MCPCAST_UPSTREAM_TOKENS`` are what an operator
sets in production; the approval timeout is shortened so the lab does not
sit for five minutes if nobody reviews.
"""
os.environ["MCPCAST_CLIENT_KEYS"] = json.dumps(CLIENT_KEYS)
os.environ["MCPCAST_UPSTREAM_TOKENS"] = json.dumps(UPSTREAM_TOKENS)
os.environ.setdefault("MCPCAST_APPROVAL_TIMEOUT", "90")
def import_generated(out_dir: Path) -> ModuleType:
"""Import a generated project through its launcher (a fresh package each time)."""
return load_generated_server(out_dir / "server.py")
def final_text(result: Any) -> str:
"""The final assistant text of an agent invocation."""
content = result["messages"][-1].content
if isinstance(content, list):
return "".join(c.get("text", "") if isinstance(c, dict) else str(c) for c in content)
return str(content)
def js(value: Any) -> str:
"""Compact JSON for printing (keeps text the model wrote readable)."""
return json.dumps(value, ensure_ascii=False)
def show_calls(recorder: CallRecorder, task_id: str) -> None:
"""Print the tools the agent chose, in order."""
for call in recorder.calls_for(task_id):
args = js({k: v for k, v in call.arguments.items() if v is not None})
print(f" {call.tool}({args}) -> {'ok' if call.ok else call.error_code}")
# ---------------------------------------------------------------------------
# 1. Generate
# ---------------------------------------------------------------------------
def generate() -> tuple[MCPcastPlan, list[Any]]:
"""Spec -> risk classification -> plan -> an editable project on disk."""
section(1, "Generate the MCP server from the OpenAPI spec")
operations = extract_operations(load_spec(SPEC))
plan = mcpcast(SPEC, profile=SafetyProfile.FULL, auth=AuthMode.API_KEY, name="storefront")
for path in write_project(plan, OUT):
print(f" wrote {path.relative_to(HERE)}")
print(f"\n {'tool':<22}{'risk':<13}{'approval':<10}upstream operation")
for tool in plan.tools:
route = tool.routes[0]
gate = "human" if tool.requires_approval else "-"
print(f" {tool.name:<22}{tool.risk.value:<13}{gate:<10}{route.method} {route.path}")
print("\n why the classifier decided that (deterministic, no model):")
reasons = {op.operation_id: classify(op) for op in operations}
for tool in plan.tools:
cls = reasons[tool.routes[0].operation_id]
# Plain "GET is a read" / "POST is a write" needs no explanation.
if cls.risk is not cls.base or not cls.reasons[0].endswith(("is a read", "is a write")):
print(f" {tool.name:<22}{'; '.join(cls.reasons)}")
print("\n not exposed (nothing vanishes silently):")
for dropped in plan.dropped:
print(f" {dropped.operation_id}: {dropped.reason}")
print("\n the same spec under each safety profile:")
for profile in SafetyProfile:
other = mcpcast(SPEC, profile=profile, auth=AuthMode.API_KEY, name="storefront")
print(
f" {profile.value:<12}{len(other.tools):>3} tools "
f"({len(other.gated_tools)} approval-gated), {len(other.dropped)} not exposed"
)
print(f"\n auth: {plan.api.auth.value} approval: {plan.api.approval_mode.value} (four-eyes)")
return plan, operations
# ---------------------------------------------------------------------------
# 2. Drive
# ---------------------------------------------------------------------------
async def business_task(module: ModuleType) -> None:
"""A real agent answers a real business question through the generated server."""
section(2, "A real agent drives the generated server")
api = FakeStorefront()
question = (
"A customer wrote in from ada@northwind.example. Who are they, "
"and what are their two most recent orders?"
)
async with api.client() as http:
server = module.build_server(http_client=http)
recorder = CallRecorder()
recorder.begin("lookup")
tools = await tools_from_server(server, recorder=recorder, headers={"x-api-key": AGENT_KEY})
agent = await build_agent(
model=MODEL,
servers={},
extra_tools=tools,
instructions=(
"You are a support assistant for the Storefront commerce platform. "
"Answer with data you fetched from the tools. Only read data in this "
"conversation — never call a tool that changes anything."
),
max_agent_iterations=6,
)
try:
print(f" question: {question}")
result = await agent.ainvoke({"messages": [HumanMessage(content=question)]})
print("\n tools the agent chose:")
show_calls(recorder, "lookup")
print(f"\n answer: {final_text(result)}")
finally:
await agent.shutdown()
print(f"\n requests the API received: {api.log}")
if api.calls:
credentials = sorted({c.authorization or "(none)" for c in api.calls})
print(f" the credential each one carried: {', '.join(repr(c) for c in credentials)}")
# ---------------------------------------------------------------------------
# 3. Govern
# ---------------------------------------------------------------------------
async def wait_for_pending(client: TestClient, tool: str, timeout: float = 60.0) -> Any:
"""Poll ``approvals_list`` as the reviewer until *tool* is waiting, or give up."""
loop = asyncio.get_running_loop()
deadline = loop.time() + timeout
while loop.time() < deadline:
(reply,) = await client.call_tool("approvals_list", {}, headers={"x-api-key": REVIEWER_KEY})
entries = json.loads(reply.text)
if isinstance(entries, list):
match = next((e for e in entries if e["tool"] == tool), None)
if match is not None:
return match
await asyncio.sleep(0.25)
return None
async def governance(module: ModuleType) -> None:
"""The refund is held by the server until a second human releases it."""
section(3, "Governance: the refund waits for a human")
api = FakeStorefront()
request = "Order ORD-1003 arrived damaged. Refund the customer 24 euros for it."
async with api.client() as http:
server = module.build_server(http_client=http)
reviewer = TestClient(server)
recorder = CallRecorder()
recorder.begin("refund")
tools = await tools_from_server(server, recorder=recorder, headers={"x-api-key": AGENT_KEY})
agent = await build_agent(
model=MODEL,
servers={},
extra_tools=tools,
instructions=(
"You are a support assistant for the Storefront commerce platform. "
"Carry out the request with the tools you have. Be brief."
),
max_agent_iterations=6,
)
try:
print(f" request: {request}")
call = asyncio.create_task(agent.ainvoke({"messages": [HumanMessage(content=request)]}))
pending = await wait_for_pending(reviewer, "refund_order")
if pending is None:
call.cancel()
print(" the agent never reached refund_order on this run — nothing to approve.")
return
print("\n the gate is holding the call:")
print(f" {pending['tool']} {js(pending['arguments'])}")
print(
f" requested by client_id={pending['client_id']!r} tenant={pending['tenant_id']!r}"
)
print(
f" refund requests the API has received so far: "
f"{len(api.received('POST', '/refund'))}"
)
print(f" API log so far: {api.log}")
print("\n a reviewer from ANOTHER tenant:")
(reply,) = await reviewer.call_tool(
"approvals_list", {}, headers={"x-api-key": OTHER_TENANT_KEY}
)
print(f" globex/gus approvals_list -> {reply.text} (sees nothing)")
(reply,) = await reviewer.call_tool(
"approvals_decide",
{"request_id": pending["request_id"], "approve": True},
headers={"x-api-key": OTHER_TENANT_KEY},
)
print(f" globex/gus tries to approve -> {json.loads(reply.text)['error']['code']}")
print("\n the caller approving their own request (same principal, approver role):")
(reply,) = await reviewer.call_tool(
"approvals_decide",
{"request_id": pending["request_id"], "approve": True},
headers={"x-api-key": SELF_APPROVE_KEY},
)
error = json.loads(reply.text)["error"]
print(f" -> {error['code']}: {error['message']}")
print("\n dana (northwind, approver role) approves:")
(reply,) = await reviewer.call_tool(
"approvals_decide",
{"request_id": pending["request_id"], "approve": True, "reason": "damaged goods"},
headers={"x-api-key": REVIEWER_KEY},
)
print(f" approvals_decide -> {reply.text}")
result = await asyncio.wait_for(call, timeout=120)
print("\n tools the agent chose:")
show_calls(recorder, "refund")
print(f"\n answer: {final_text(result)}")
finally:
await agent.shutdown()
refunds = api.received("POST", "/refund")
print(f"\n refund requests the API received: {len(refunds)}")
for refund in refunds:
print(f" {refund} body={js(refund.body)}")
print(f" authorization={refund.authorization!r} (northwind's upstream credential)")
print(f" refunds now on record: {js(api.refunds)}")
# ---------------------------------------------------------------------------
# 4. Measure
# ---------------------------------------------------------------------------
def tasks_for(plan: MCPcastPlan) -> list[EvalTask]:
"""The same eight user requests, targeted at whichever plan we are scoring.
``--eval`` writes tasks like these with a model; spelling them out keeps the
two runs comparable, which is the whole point of a before/after.
"""
merged = "find_customer" in plan.tool_names
by_email = "find_customer" if merged else "search_customers"
by_id = "find_customer" if merged else "get_customer"
requests: list[tuple[str, str, str]] = [
("t1", "A customer wrote in from ada@northwind.example — who are they?", by_email),
("t2", "Pull up order ORD-1002.", "get_order"),
("t3", "Which orders does customer CUS-1001 have, newest first?", "list_orders"),
("t4", "Order ORD-1003 arrived damaged — refund 24 euros for it.", "refund_order"),
("t5", "How much revenue did we make in 2026-02?", "get_revenue_report"),
("t6", "Show me the record for customer CUS-1002.", by_id),
("t7", "Which of our customers are on the enterprise plan?", "list_customers"),
(
"t8",
"Start a draft order for customer CUS-1003: one MAT-DESK-90 at 89 euros.",
"create_order",
),
]
return [EvalTask(id=i, prompt=prompt, expected_tool=tool) for i, prompt, tool in requests]
async def measure(
plan: MCPcastPlan, module: ModuleType, operations: list[Any], out_dir: Path
) -> EvalReport:
"""Run the Agent Readiness evaluation and print the grade with its evidence."""
tasks = tasks_for(plan)
report = await evaluate(
plan,
module.build_server,
model=MODEL,
tasks=tasks,
operations=operations,
live_reads=False, # this API is fictional; mock every route from the spec
headers={"x-api-key": AGENT_KEY},
)
print(f"\n Agent Readiness: {report.grade} score {report.score:.2f}")
print(
f" tasks succeeded {report.tasks_succeeded}/{report.tasks_total} "
f"correct tool first {report.selection_rate:.0%} "
f"parameter errors {report.param_error_rate:.0%}"
)
for result in report.results:
called = " -> ".join(c.tool for c in result.calls) or "(no tool call)"
print(
f" {result.task.id} expected {result.task.expected_tool:<19}"
f"called {called:<46} {'ok' if result.success else 'MISS'}"
)
print("\n what the report says to fix:")
for fix in report.fixes:
print(f" {fix}")
tasks_path, report_path = write_eval(report, tasks, out_dir)
print(f"\n wrote {tasks_path.relative_to(HERE)} and {report_path.relative_to(HERE)}")
return report
# ---------------------------------------------------------------------------
# 5. Improve
# ---------------------------------------------------------------------------
def improve(plan: MCPcastPlan) -> MCPcastPlan:
"""Apply the report's advice to the plan — the plan is the source of truth.
Four edits, each one something the readiness report asked for:
- merge the two customer lookups into one intent tool with two routes,
so the agent picks a *customer*, not an HTTP endpoint;
- give ``list_customers`` and ``list_orders`` the examples the report
says they are missing;
- drop the health check no agent task ever needs;
- put ``notify_customer`` on a param diet: hidden, and always ``false``,
so an agent creating a draft order can never email a customer.
"""
get_customer = plan.tool("get_customer")
search_customers = plan.tool("search_customers")
find_customer = ToolPlan(
name="find_customer",
description=(
"Find ONE customer, by id or by email address. Pass customer_id when you "
"know it (e.g. 'CUS-1001'), otherwise pass email. Use list_customers only "
"to browse or filter the whole directory."
),
risk=RiskClass.READ,
routes=[get_customer.routes[0], search_customers.routes[0]],
params={
"customer_id": ParamPlan(
description="The customer identifier, e.g. 'CUS-1001'.",
json_schema={"type": "string"},
),
"email": ParamPlan(
description="The customer's email address, matched exactly.",
json_schema={"type": "string", "format": "email"},
),
},
example={"email": "ada@northwind.example"},
tags=["customers"],
)
create_order = plan.tool("create_order")
params = dict(create_order.params)
params["notify_customer"] = params["notify_customer"].model_copy(
update={"hidden": True, "default": False}
)
create_order = create_order.model_copy(update={"params": params})
replacements = {
"get_customer": find_customer,
"create_order": create_order,
"list_customers": plan.tool("list_customers").model_copy(
update={"example": {"plan": "enterprise", "limit": 5}}
),
"list_orders": plan.tool("list_orders").model_copy(
update={"example": {"customer_id": "CUS-1001", "status": "open"}}
),
}
tools = [
replacements.get(tool.name, tool)
for tool in plan.tools
if tool.name not in ("get_health", "search_customers")
]
return MCPcastPlan(
api=plan.api,
profile=plan.profile,
tools=tools,
dropped=[
*plan.dropped,
DroppedOp(
operation_id="getHealth",
reason="operational endpoint — no agent task needs it (Agent Readiness: not covered)",
),
],
)
# ---------------------------------------------------------------------------
async def main() -> None:
os.chdir(HERE)
configure_environment()
plan, operations = generate()
load_dotenv_if_present() # .env next to the project, as build_agent() would
if not os.environ.get("OPENAI_API_KEY"):
print(
"\nSteps 2-5 drive a real agent: set OPENAI_API_KEY and run this again.\n"
" export OPENAI_API_KEY=sk-..."
)
sys.exit(1)
module = import_generated(OUT)
await business_task(module)
await governance(module)
section(4, "Measure: the Agent Readiness Score")
print(" 8 tasks, a real agent, the full server pipeline, spec-derived mock responses")
print(" (nothing an evaluation does can reach real data)")
before = await measure(plan, module, operations, OUT)
section(5, "Improve the plan, regenerate, re-measure")
better = improve(plan)
for path in write_project(better, OUT_V2):
print(f" wrote {path.relative_to(HERE)}")
print(
f" merged get_customer + search_customers -> find_customer "
f"({len(better.tool('find_customer').routes)} routes, dispatched by which id you pass)"
)
print(" added the missing examples; dropped get_health")
print(f" create_order description now says: {'Always sends: notify_customer=false'!r}")
print(
f" surface: {len(plan.tools)} -> {len(better.tools)} tools, "
f"{sum(1 for t in plan.tools if t.example)} -> "
f"{sum(1 for t in better.tools if t.example)} of them with a worked example"
)
after = await measure(better, import_generated(OUT_V2), operations, OUT_V2)
print(
f"\n before: {before.grade} ({before.score:.2f}) after: {after.grade} "
f"({after.score:.2f})"
)
if after.score > before.score:
print(" the edits paid off — the same eight requests land more often now.")
elif after.score == before.score:
print(
" identical score on this run — the surface was already unambiguous for these "
"eight requests, and the edits kept it that way with two fewer tools. That is a "
"real win (less to choose from, fewer tokens), and it is what honest measurement "
"looks like: the number does not move just because you changed something."
)
else:
print(
" the score went DOWN. That is the point of measuring: revert the edit, or "
"sharpen the merged description, and run it again."
)
print("\nNext: point Claude Desktop / Claude Code / Cursor at generated/v2/server.py")
print(" (see README.md), or run the same pipeline from the CLI:")
print(" promptise mcpcast examples/mcp/mcpcast_storefront_lab/storefront.yaml \\")
print(" --profile full --auth api-key --out storefront-mcp --eval")
if __name__ == "__main__":
asyncio.run(main())
The fake upstream it runs against is fake_api.py — a dict-backed store behind
an httpx.MockTransport that records every request, so the lab can prove what
did and did not reach the API.
Next Steps¶
- MCPcast an Existing API — the step-by-step guide to running the same pipeline against your spec, with curation and review
- MCPcast reference — pipeline, curation, plan schema, auth modes, readiness scoring, limits
- Approval Gates — elicitation, pending, webhook and callback approvers in depth
- Multi-Tenancy — tenant claims, isolation keys, per-tenant rate limits and audit
- Build a Secure Multi-Tenant Agent Platform — the same guarantees for a server you write by hand