Skip to content

Make Your Python API MCP-Ready

You are looking at your own codebase — a FastAPI, Django, Flask or Litestar service — and the question is concrete: exactly what do I do so Claude, Cursor and every other AI can use this API?

Not a spec file in hand. An app in hand. This guide is that sequence, step by step, on a small FastAPI app that stands in for yours. Every command was run and every output below is what it printed.

You do not write an MCP server

Your framework already emits an OpenAPI document. promptise mcpcast turns that document into a curated, safe MCP server as editable code. This page is the your own app path; the general walkthrough from a spec file is MCPcast an Existing API, the deep reference is MCPcast, and the recipes show what happens on Stripe, GitHub and legacy Swagger 2 documents.

What You'll Build

  • Step 1 — the URL your framework already serves the spec at
  • Step 2 — a spec that is a good tool surface: names, descriptions and examples a model can act on, with a before/after
  • Step 3 — an MCP server generated from the running app, with reads open, writes gated, and the admin, deprecated and health endpoints kept out
  • Step 4 — your app's authentication mapped onto the right auth mode
  • Step 5 — a real agent using it over MCP stdio, then Claude Desktop, Claude Code and Cursor
  • Step 6 — a write denied fail-closed, then executed once a human approves
  • Step 7 — the server shipped next to your app: CI guard, promptise serve, a Dockerfile, the environment variables
  • Step 8 — "works with Claude" as a feature of your product, with a readiness score as the quality bar

Concepts

The spec you already have is the input. FastAPI, Django Ninja and Litestar generate OpenAPI from your route signatures; Django REST Framework and Flask do it through drf-spectacular and flask-smorest. There is nothing to export by hand and nothing to keep in sync: mcpcast fetches the document from the running app.

The spec's words are the model's only eyes. An agent never reads your code. It sees a tool name, a tool description and parameter descriptions — which are your operation_id, your summary plus docstring, and your Field(description=...). A route without them becomes a tool called search_books_search_post described as "Search", and that is what the model has to choose between. Step 2 is the highest-value twenty minutes on this page.

Exposure is a policy, not a side effect. Every operation is risk-classified by fixed rules (read, write, destructive, financial). The safety profile decides which classes become tools at all, and everything that changes data is requires_approval=True — enforced by the generated server's middleware for every MCP client, not just polite ones. Your DELETE never reaches an agent unless you opt in, and then never runs without a human.

Your app's auth maps onto one of four modes. A personal bearer token, a token per user forwarded over HTTP, or a key per customer — each is one --auth flag, and the generated server carries the credential so the model never sees it.

The plan file is yours. mcpcast.plan.yaml is the reviewable, diffable description of the tool surface; the generated package is regenerated from it. It lives in your repository next to the app, and CI checks it against the app on every commit.


The app

The runnable example at examples/mcp/mcpcast_fastapi_app/ is a Bookshelf API — a deliberately ordinary FastAPI service with the shapes every real app has:

Route operation_id What it is
GET /health health_check liveness probe for the load balancer
GET /books list_books list, optional author filter
GET /books/{book_id} get_book one record
POST /books/search search_books a search that is a POST because it carries a body
POST /books create_book a write
PATCH /books/{book_id} update_book a write
DELETE /books/{book_id} delete_book destructive
GET /books/find find_books_legacy deprecated=True
POST /admin/reset reset_catalogue admin-only, destructive

Everything under /books and /admin needs Authorization: Bearer demo-token, checked by a FastAPI dependency. The store is in memory. run.py starts the app, generates the server, drives it with a real agent and exercises the approval gate — the whole driver is at the end of this page.


Step 1 — Get your spec URL

Start your app and open the OpenAPI document in a browser. You should see a JSON (or YAML) document with "openapi": "3.x.x" and a paths object. Where it lives depends on the framework:

Framework Where the spec is Notes
FastAPI /openapi.json FastAPI(openapi_url=...) moves it; None disables it. Behind a proxy with root_path set, the document gains a relative servers entry
Django Ninja /api/openapi.json NinjaAPI(openapi_url="/openapi.json") by default, relative to where you mounted api.urls — /api/ in the usual path("api/", api.urls)
Django REST Framework + drf-spectacular /api/schema/ if you routed SpectacularAPIView there, as the drf-spectacular quickstart does; YAML by default, ?format=json for JSON — mcpcast reads both. Offline: python manage.py spectacular --file schema.yml
Flask + flask-smorest <OPENAPI_URL_PREFIX>/openapi.json only served when OPENAPI_URL_PREFIX is set — with "/" the spec is at /openapi.json
Litestar /schema/openapi.json OpenAPIConfig(path=...) moves the whole /schema group
Starlette, aiohttp, plain WSGI — no full OpenAPI generator built in (Starlette's SchemaGenerator is docstring-only). Write the document by hand, or skip the spec and build the tools directly with the Server SDK

These are the defaults as of the versions current when this page was written; your project may have moved them, so confirm by opening the URL.

For the Bookshelf app:

cd examples/mcp/mcpcast_fastapi_app
pip install fastapi                    # the app's own web framework (uvicorn ships with Promptise)
uvicorn app:app --port 8011
curl -s http://127.0.0.1:8011/openapi.json | head -c 160
{"openapi":"3.1.0","info":{"title":"Bookshelf API","description":"A small library catalogue: browse, search and maintain books.","version":"1.0.0"},"paths":{"/h

The spec URL is also the API URL

FastAPI emits no servers block. When mcpcast fetches the document over HTTP it resolves the API base against the URL it fetched from — the app's own origin — so nothing else is needed. The same resolution handles a relative servers entry (FastAPI with root_path, DRF's generators). Working from a saved file instead, pass --base-url https://api.yourcompany.com.


Step 2 — Make the spec good enough to be a tool surface

This is where the quality of the result is decided, and it happens in your code. The mapping is exact:

In your app In the spec What the agent sees
operation_id="search_books" operationId the tool name
summary="…" + the docstring summary, description the tool description — the only thing the model reads when choosing
Field(description=…), Query(description=…) parameter description each parameter's description
Field(examples=[…]) examples the worked example in the description
tags=[…] tags tool tags, kept in the plan and on the generated @server.tool(tags=…)
deprecated=True deprecated always dropped, with the reason deprecated in spec
include_in_schema=False absent never seen — the right answer for internal endpoints

Before. A route as most of us write it on day one:

class SearchQuery(BaseModel):
    q: str
    limit: int = 10

@app.post("/books/search")
async def search(request: SearchQuery) -> list[dict]:
    ...

FastAPI derives the operation id from the function, path and method, and the summary from the function name. mcpcast turns that into a tool named search_books_search_post described as Search, with two undescribed parameters. Exactly what the generated tools module registers:

@server.tool(
    name='search_books_search_post',
    description='Search\n\nParameters:\n  - q (string, required)\n  - limit (integer)\n\nExample: {"q": "string"}',
    read_only_hint=True,
    open_world_hint=True,
)

A model choosing between search_books_search_post, get_books_book_id_get and list_books_books_get is guessing.

After. The same route in the Bookshelf app:

class SearchQuery(BaseModel):
    """A search request (POST, because it carries a body)."""

    query: str = Field(
        description="Case-insensitive text matched against title and author.",
        examples=["earthsea"],
    )
    author: str | None = Field(default=None, description="Only books by this exact author.")
    limit: int = Field(default=10, ge=1, le=100, description="Maximum number of results.")


@books.post("/search", operation_id="search_books", summary="Search books by title or author")
async def search_books(request: SearchQuery) -> list[Book]:
    """Full-text search over titles and authors. Read-only despite being a POST."""

which becomes, in the plan and in the server, a tool called search_books (plan excerpt, JSON schemas trimmed):

- name: search_books
  description: Search books by title or author. Full-text search over titles and authors. Read-only despite
    being a POST.
  risk: read
  params:
    query:
      description: Case-insensitive text matched against title and author.
      required: true
    author:
      description: Only books by this exact author.
    limit:
      description: Maximum number of results.
  example:
    query: earthsea

The checklist, in the order it pays off:

  1. operation_id on every route. It is the tool name, and tool names are most of what the model has to go on. Use verbs in your product's language: search_books, not search, not books_post.
  2. A summary and a docstring that say what the operation does and when to use it — "Full-text search … use get_book when you already have an id". Write them for a model that cannot ask follow-up questions.
  3. Field(description=…) on every request-model field and Query(description=…) on every query parameter. A parameter called q with no description is a parameter error waiting to happen.
  4. examples=[…] on required fields. mcpcast builds each tool's worked example from them; without them the example is {"q": "string"}.
  5. Response models (-> list[Book]). They are how the agent knows what comes back, and how the readiness evaluation mocks writes safely.
  6. deprecated=True on routes you are retiring — they are dropped automatically — and include_in_schema=False on anything that was never for API consumers.

Set operation_id before anything else

Renaming tools later in mcpcast.plan.yaml works, but every regeneration from the spec brings the derived names back. Fix the source. In FastAPI you can also set generate_unique_id_function once on the app so every route's id is its function name.

Other frameworks expose the same knobs: Django Ninja's route decorators take operation_id, summary, description, tags and deprecated; Litestar's @get/@post take the same names; drf-spectacular's @extend_schema(operation_id=…, summary=…, description=…) annotates a DRF view; flask-smorest uses the docstring's first line as the summary and @blp.doc(operationId=…) for the id.


Step 3 — Generate the server from the running app

Point mcpcast at the URL. --profile standard opens reads and gated writes, --auth env-token is the mode for a personal server (Step 4 explains why), --review shows the plan before writing, and --no-curate keeps this run deterministic and offline:

promptise mcpcast http://127.0.0.1:8011/openapi.json --no-curate \
    --profile standard --auth env-token --name bookshelf --review -o bookshelf-mcp

Trimmed to the columns that matter:

Parsed 9 operations from http://127.0.0.1:8011/openapi.json; profile=standard auth=env-token
Tools (6) — profile standard
  health_check   read   —         GET /health             —
  list_books     read   —         GET /books              author, limit
  create_book    write  required  POST /books             title, author, year, tags
  search_books   read   —         POST /books/search      query, author, limit
  get_book       read   —         GET /books/{book_id}    book_id
  update_book    write  required  PATCH /books/{book_id}  book_id, notes, tags, year
Not exposed (3)
  find_books_legacy   deprecated in spec
  delete_book         destructive operation excluded by profile 'standard'
  reset_catalogue     destructive operation excluded by profile 'standard'
Write the project? [Y/n]:
╭─────────────────────── mcpcast ────────────────────────╮
│ bookshelf → bookshelf-mcp/                             │
│   tools: 6  (2 require human approval)                 │
│   not exposed: 3 operations (with reasons in the plan) │
│   files: bookshelf_mcp/ (8 modules), tests/ (2), server.py, README.md,           │
│   pyproject.toml, Dockerfile, .env.example, .gitignore                          │
╰──────────────────────────────────────────────────────────────────────────────────╯

Read that against your own endpoints:

  • POST /books/search is read. The classifier looks at the leading verb of the id, summary and last path segment, not just the method: search with no mutating verb anywhere is a read, so the tool is ungated. POST /books (create_book) is a write.
  • create_book and update_book are requires_approval=True. Under standard every write is generated and gated; the gate is middleware in the generated server, so it holds for Claude Desktop, Cursor and a script alike.
  • delete_book is not there. DELETE is destructive, and standard refuses that class. It appears — still gated — only under --profile full.
  • reset_catalogue is not there either, twice over: the verb reset makes it destructive, and a path segment admin escalates whatever class an operation has. An admin-only GET under standard would still be generated as a gated write rather than a free read.
  • find_books_legacy is dropped because it is deprecated: true, under every profile.
  • health_check is a tool, and it should not be: an agent has no task that needs a liveness probe. Deterministic generation keeps every read. Three fixes, pick one: include_in_schema=False in the app (best if no API consumer needs it in the docs either), move it to dropped in the plan (below), or let curation decide — it drops health checks on its own.

The plan file is the artifact. In bookshelf-mcp/mcpcast.plan.yaml the fix is moving one entry:

dropped:
- operation_id: find_books_legacy
  reason: deprecated in spec
- operation_id: delete_book
  reason: destructive operation excluded by profile 'standard'
- operation_id: reset_catalogue
  reason: destructive operation excluded by profile 'standard'
- operation_id: health_check
  reason: operational endpoint — for the load balancer, not an agent

then promptise mcpcast bookshelf-mcp/mcpcast.plan.yaml regenerates the package, server.py, tests/ and README.md from it, leaving the plan (and your comments) untouched. The runnable example makes the same edit in code, so its output shows five tools and four drops.

With curation on

Drop --no-curate and a model designs the surface (it needs OPENAI_API_KEY; --model picks another provider). On this app one run produced two tools: find_books (three routes: by id, search, list — dispatched by which parameters you pass) and save_book (create + update), with health_check, the deprecated route and reset_catalogue dropped with written reasons. It also marked notes on save_book as hidden — a hidden parameter is never shown to the agent, and that is the one field update_book mostly exists for. Curation is a strong first draft that you review, which is what --review is for; every proposal is checked in code and risk is never downgraded. Nine operations is small enough that the deterministic surface is already good; curation earns its keep on the 100-operation apps.


Step 4 — Map your app's auth to an auth mode

The generated server calls your API with a credential. Which one, and where it comes from, is the --auth mode — pick it from how your app authenticates today:

Your app today Who will use the MCP server --auth What the server does
one bearer token / API key per developer or per team you, in Claude Desktop, Claude Code, Cursor (stdio) env-token sends MCPCAST_UPSTREAM_TOKEN as Authorization on every call
a token per user (JWT, OAuth access token) many users, over HTTP, each as themselves passthrough (default) forwards the caller's Authorization header unchanged — your API's own permissions still apply
a key per customer / tenant your customers, on a server you host api-key clients present x-api-key; each key names a tenant whose upstream credential lives on the server, never on the client
nothing (a local sandbox) you, locally none no credentials; the server refuses to bind to anything but loopback

The Bookshelf app takes one bearer token, and the goal is my desktop client using my API, so env-token is right. Desktop clients launch the server over stdio and cannot send request headers, which is why passthrough would fail every call there with UPSTREAM_AUTH_MISSING.

The generated bookshelf-mcp/README.md already carries the client configuration with the credential in the client's own environment block:

{
  "mcpServers": {
    "bookshelf": { "command": "python", "args": ["/absolute/path/to/server.py"], "env": {"MCPCAST_UPSTREAM_TOKEN": "Bearer <your API token>"} }
  }
}

and it says why it is there:

The credential goes in the client's own config below — a desktop client launched from the GUI does not inherit your shell, so export MCPCAST_UPSTREAM_TOKEN=… in a terminal will not reach it.

The value is the whole header — Bearer demo-token, scheme included. Get it wrong and the failure is specific: without the variable the tool returns UPSTREAM_AUTH_MISSING; with the wrong token your API answers and the tool relays it as UPSTREAM_ERROR … HTTP 401 (both verbatim in Troubleshooting).

For the other two shapes: passthrough is a shared HTTP deployment (promptise serve server:server -t http) where every client sends its own Authorization header — see Auth & Security; api-key adds per-tenant credentials and tenant-scoped four-eyes approval — see Multi-Tenancy and the storefront lab, which runs it end to end.


Step 5 — Try it with a real agent

run.py connects build_agent("openai:gpt-5-mini") to the generated server over the real MCP stdio transport — the same way Claude Desktop launches it — with the token in the server's environment, and asks a question the live app has to answer:

agent = await build_agent(
    model="openai:gpt-5-mini",
    servers={
        "bookshelf": StdioServerSpec(
            command=sys.executable,
            args=[str(OUT / "server.py")],          # generated/server.py
            env={"MCPCAST_UPSTREAM_TOKEN": "Bearer demo-token"},
        )
    },
    instructions="You are the Bookshelf assistant. Answer from what the tools return. Be brief.",
)
result = await agent.ainvoke({"messages": [HumanMessage(content=QUESTION)]})
3. A real agent over MCP stdio -> generated/server.py -> the live app
==============================================================================
  question: Which books by Ursula K. Le Guin do we have, and which of them is the oldest?
  tools the agent chose:
    list_books({"author": "Ursula K. Le Guin"})
  answer: We have these Ursula K. Le Guin books:
- A Wizard of Earthsea (1968)
- The Left Hand of Darkness (1969)
- The Dispossessed (1974)

The oldest is A Wizard of Earthsea (1968).

One tool call, the right one, with the author filter the description advertised — because Step 2 gave the model something to read.

The port, the model's wording and the exact arguments it picks (a limit of 10 or 20, say) vary between runs; the tool table, the dropped list and both halves of Step 6's write do not.

Now the clients you actually wanted. Use the interpreter that has promptise installed and the absolute path to server.py:

claude_desktop_config.json:

{
  "mcpServers": {
    "bookshelf": {
      "command": "/absolute/path/to/.venv/bin/python",
      "args": ["/absolute/path/to/bookshelf-mcp/server.py"],
      "env": { "MCPCAST_UPSTREAM_TOKEN": "Bearer demo-token" }
    }
  }
}
claude mcp add bookshelf -e MCPCAST_UPSTREAM_TOKEN="Bearer demo-token" \
    -- /absolute/path/to/.venv/bin/python /absolute/path/to/bookshelf-mcp/server.py

.cursor/mcp.json:

{
  "mcpServers": {
    "bookshelf": {
      "command": "/absolute/path/to/.venv/bin/python",
      "args": ["/absolute/path/to/bookshelf-mcp/server.py"],
      "env": { "MCPCAST_UPSTREAM_TOKEN": "Bearer demo-token" }
    }
  }
}
cd bookshelf-mcp
MCPCAST_UPSTREAM_TOKEN='Bearer demo-token' promptise serve server:server -t http --port 8080
# MCP endpoint: http://127.0.0.1:8080/mcp

Ask "what do we have by Le Guin?" and the client calls list_books. The app must be reachable at the base URL baked into the plan; MCPCAST_BASE_URL overrides it at runtime.


Step 6 — Writes and approval

Reads are half the value. The other half — "add a note to that book" — is where a raw API handed to a model becomes a headline. The generated server gates it, and the gate is not advisory.

Over stdio, with no human to ask. run.py asks the same agent for a write. The Promptise MCP client does not implement elicitation, so the server has nobody to put the question to — and denies:

4. Writes: denied fail-closed over stdio, executed once a human approves
==============================================================================
  a) this MCP client does not support elicitation, so nobody can be asked and the gate denies:
  request: Add the note 'signed first edition' to The Dispossessed.
  tools the agent chose:
    search_books({"query": "The Dispossessed", "limit": 10})
    update_book({"book_id": 3, "notes": "signed first edition"})
      -> APPROVAL_DENIED: Approval denied for tool 'update_book': client declined or returned an invalid elicitation response
  answer: I couldn't add the note — the update_book call was denied by the system (approval denied). The note was not added.

The agent found the book, chose the right tool with the right arguments, and the PATCH never reached the app. The error is structured ({"error": {"code": "APPROVAL_DENIED", …}}), so the model reports it instead of retrying or inventing a result.

With a human who says yes. In a client that implements MCP elicitation (Claude Code, for one) the same call shows you a prompt naming the tool and its arguments, and runs only if you accept. run.py reproduces that in-process: build_server() in the generated module takes an approval_handler, and TestClient runs the full pipeline — validation, middleware, gate, handler — without a port:

def reviewer(request: ApprovalRequest) -> bool:
    print(f"    approval requested: {request.tool_name} {json.dumps(request.arguments)}")
    return True

client = TestClient(import_generated().build_server(approval_handler=reviewer))
(reply,) = await client.call_tool("update_book", {"book_id": 3, "notes": "signed first edition"})
  b) in-process, with a human (here: a callback) who approves:
    get_book(3) before: notes=''  (the denied call changed nothing)
    approval requested: update_book {"book_id": 3, "notes": "signed first edition", "tags": null, "year": null}
    update_book -> {"id": 3, "title": "The Dispossessed", "author": "Ursula K. Le Guin", "year": 1974, "tags": ["science-fiction"], "notes": "signed first edition"}
    get_book(3) after:  notes='signed first edition'

The reviewer sees exactly what will run, the PATCH hits the live app once approved, and the follow-up read proves it. A timeout denies; a handler crash denies; nothing is ever silently allowed. Elicitation is "confirm your own action"; for an independent reviewer — someone other than the person driving the agent — use --auth api-key with --approval pending, covered in Approval Gates.


Step 7 — Ship it next to your app

Keep the project in the repository, beside the code it exposes:

your-service/
├── app/                      # your FastAPI / Django / Flask code
├── bookshelf-mcp/
│   ├── mcpcast.plan.yaml     # the source of truth — reviewed like code
│   ├── server.py             # launcher — regenerated, never hand-edited
│   ├── bookshelf_mcp/        # the server as a package — regenerated
│   ├── tests/                # generated tests + your own files
│   ├── pyproject.toml        # yours after the first write
│   ├── Dockerfile  .env.example  .gitignore
│   └── README.md
└── tests/
    └── test_mcp_surface.py   # below

Guard it in CI. Your API will change; the tool surface should fail loudly when it does, not drift. The pipeline is a Python API, and app.openapi() is the spec — no server needed:

"""CI guard: the shipped MCP tool surface matches the app. Offline, deterministic, fast."""

from pathlib import Path

from promptise.mcpcast import MCPcastPlan, RiskClass, mcpcast

from app import app  # your FastAPI app — app.openapi() is the spec, no server needed

PLAN = MCPcastPlan.load(Path(__file__).parent.parent / "bookshelf-mcp" / "mcpcast.plan.yaml")


def test_every_endpoint_is_a_tool_or_dropped_with_a_reason():
    """A new endpoint must be added to the plan or dropped deliberately — never ignored."""
    fresh = mcpcast(app.openapi(), profile=PLAN.profile, base_url=PLAN.api.base_url)
    in_app = fresh.kept_operations | {d.operation_id for d in fresh.dropped}
    in_plan = PLAN.kept_operations | {d.operation_id for d in PLAN.dropped}
    assert in_app == in_plan, f"unaccounted: {in_app ^ in_plan}"


def test_no_tool_is_less_risky_than_the_classifier_says():
    """Hand edits may raise a tool's risk, never lower it."""
    fresh = mcpcast(app.openapi(), profile=PLAN.profile, base_url=PLAN.api.base_url)
    floor = {op: t.risk for t in fresh.tools for op in t.operations}
    for tool in PLAN.tools:
        for op in tool.operations:
            assert tool.risk.at_least(floor[op]), f"{tool.name} downgrades {op}"


def test_everything_that_writes_is_approval_gated():
    for tool in PLAN.tools:
        assert tool.requires_approval == (tool.risk is not RiskClass.READ), tool.name


def test_the_tools_customers_rely_on_still_exist():
    assert {"list_books", "search_books", "get_book"} <= set(PLAN.tool_names)
$ pytest tests/test_mcp_surface.py -q
4 passed, 2 warnings in 2.26s

When a colleague adds POST /books/{book_id}/lend, the first test fails with unaccounted: {'lend_book'} and the pull request has to say what happens to it: add it to the plan (regenerate from the spec into a scratch directory and copy the new tool's entry over, or write it by hand) or drop it with a reason. Either way promptise mcpcast bookshelf-mcp/mcpcast.plan.yaml regenerates the package, and the diff is reviewed like any other. The nightly readiness test in the general guide adds a model-scored floor on top.

Run it. The launcher exposes server, so the standard runner works from inside the project directory, with the dashboard and hot reload from Deployment — or install the project and run it as a command:

cd bookshelf-mcp
promptise serve server:server --transport http --port 8080
pip install -e ".[dev]" && pytest && bookshelf-mcp --transport http --port 8080
INFO:     Uvicorn running on http://127.0.0.1:8080 (Press CTRL+C to quit)
    Server      bookshelf v0.1.0
    Transport   Streamable HTTP
    Endpoint    http://127.0.0.1:8080/mcp
    Tools       5 registered

Containerise it. The project ships with a Dockerfile (written once — edit it freely): a fully-qualified base image, a non-root user, no baked-in configuration. The server depends only on promptise and httpx and talks to your API over HTTP, so it can run in the same image as your app or in its own:

FROM python:3.12-slim-bookworm

WORKDIR /app
COPY pyproject.toml README.md ./
COPY bookshelf_mcp ./bookshelf_mcp
RUN pip install --no-cache-dir . \
    && useradd --system --create-home --shell /usr/sbin/nologin app
USER app

# The upstream base URL is compiled into bookshelf_mcp/config.py; MCPCAST_BASE_URL at run
# time points the image at another environment. Credentials come from the
# environment at run time too (see .env.example) — never bake them into the image.
EXPOSE 8080

# Auth mode 'env-token' executes every call with the operator's
# MCPCAST_UPSTREAM_TOKEN and has no MCP-level authentication, so the server
# binds loopback only and this image serves stdio. To publish it anyway — only
# behind an authenticating gateway — run it with
# `--transport http --host 0.0.0.0 --public` (or MCPCAST_PUBLIC=1). Regenerate
# with --auth api-key or passthrough for a shared deployment.
CMD ["bookshelf-mcp"]
docker build -t bookshelf-mcp . && docker run -i --env-file .env bookshelf-mcp

This is an env-token project — one credential, no caller authentication — so the image serves stdio and the server refuses a non-loopback bind unless you pass --public behind an authenticating gateway. Regenerate with --auth api-key for a shared HTTP deployment that authenticates its callers.

Everything the server reads at runtime:

Variable Purpose Default
MCPCAST_BASE_URL Where your API is — staging vs production, or the port the app got this time. One clean absolute URL: a query string or fragment, user:password@ or whitespace is refused at start-up api.base_url from the plan
MCPCAST_UPSTREAM_TOKEN env-token: the full Authorization value sent upstream, scheme included — or the raw key in the header / query parameter your app's security scheme names (.env.example says which); a <placeholder> left in place is refused —
MCPCAST_CLIENT_KEYS api-key: JSON map of client key → {client_id, tenant_id, roles}; the server refuses to start with none, with a <placeholder> key or with a documentation key (sk-acme, sk-reviewer) — .env.example ships no working key —
MCPCAST_UPSTREAM_TOKENS api-key: JSON map of tenant → upstream credential, read per call so rotation needs no restart —
MCPCAST_TIMEOUT Total time one upstream call may take, seconds — connecting, sending, waiting and reading the body together 30
MCPCAST_APPROVAL_TIMEOUT How long a gated call waits for a decision before it is denied 300
MCPCAST_MAX_RESPONSE_BYTES, MCPCAST_ERROR_EXCERPT_CHARS, MCPCAST_ALLOW_INSECURE_HTTP, MCPCAST_PUBLIC, MCPCAST_MAX_PENDING, MCPCAST_MAX_PENDING_PER_TENANT, MCPCAST_MAX_PENDING_PER_CLIENT Response cap, error-excerpt length (the credential is scrubbed from the body first), plain-http opt-in (pre-set in .env.example when the plan's upstream is plain http), non-loopback opt-in, pending-approval capacity (server-wide, per tenant, per API key) — see the generated README's configuration table see README

The run.py example bakes the ephemeral port it was given into the plan; for anything that outlives one run, regenerate against the app's stable URL or set MCPCAST_BASE_URL.


Step 8 — Make it a feature

Once the server is in the repo, "works with Claude" is a line in your product's README rather than a roadmap item. The block your users need is three lines:

## Use Bookshelf with Claude, Cursor or any MCP client

pip install promptise
claude mcp add bookshelf -e MCPCAST_UPSTREAM_TOKEN="Bearer <your token>" \
    -- python /path/to/bookshelf-mcp/server.py

Reads work immediately. Anything that changes data asks you to approve it first.

Make the readiness score your quality bar. Tool design is usually a matter of opinion; --eval makes it a measurement. A model writes realistic tasks for your tools, a real agent attempts them against the generated server — reads against your live API, writes against spec-derived mocks behind an auto-approver, so nothing real changes — and you get a grade with named fixes:

export MCPCAST_EVAL_AUTHORIZATION='Bearer demo-token'   # so live reads are authenticated
promptise mcpcast bookshelf-mcp/mcpcast.plan.yaml --eval --eval-tasks 8
Evaluating with openai:gpt-5-mini (8 tasks)…
Agent Readiness: A  (8/8 tasks succeeded)
  ✗ `update_book` vs `get_book` are ambiguous — the agent picked `get_book` in 1/2 runs that needed `update_book` → merge them, or say in each description when NOT to use it
  • `list_books` has no example — agents lean on examples heavily
Report: bookshelf-mcp/eval/report.md  Tasks: bookshelf-mcp/eval/tasks.yaml

(Report paths shortened; the CLI prints them absolute.) eval/report.md records score 0.95, correct tool selected first 88%, parameter error rate 0%, and a per-task table — every task landed, and the one selection miss was get_book called before update_book on a task that said "update book id 1".

Both fixes are small. The generated example only covers required parameters and list_books has none, so its example is one line in the plan: example: {author: Ursula K. Le Guin}. The ambiguity is one sentence in update_book's docstring, in your code — "use get_book to read; this tool only changes fields". Regenerate, re-score. An A before a release is a bar a team can hold, and the score moves a little between runs, so assert a floor rather than a number.

When your API is big. A nine-operation API reads fine as one tool per operation; a two-hundred-operation one does not. That is what curation is for — a model that merges routes serving one intent, renames into your domain language, puts parameters on a diet and drops what no user task needs, with every proposal checked in code. The recipes show it on Stripe and on GitHub's 1,225-operation spec, including how to narrow a spec by tag before curating.


What You've Built

  • The spec URL for your framework, and a spec whose names, descriptions and examples are written for a model
  • An MCP server generated from the running app — reads open, writes requires_approval=True, the delete and the admin reset refused by the profile, the deprecated route and the health check dropped with reasons
  • The right auth mode for your app's authentication, with the credential in the client's config and never in the model's context
  • A real agent using it over MCP stdio, and the Claude Desktop, Claude Code and Cursor configurations
  • A write denied fail-closed when no human could be asked, and executed against the live app once one approved
  • The server shipped next to your app with a CI guard, promptise serve, a Dockerfile and the runtime variables
  • A readiness grade as the quality bar for "works with Claude"

Troubleshooting

Message Cause Fix
could not fetch spec from http://127.0.0.1:8011/: Client error '404 Not Found' The URL is your app's root, not the spec Use the path from the Step 1 table — /openapi.json, /api/openapi.json, /api/schema/, /schema/openapi.json
http://127.0.0.1:8011/docs: not valid JSON or YAML: … You pointed at the Swagger UI page, not the document it renders Same fix — the JSON/YAML URL, not /docs or /redoc
spec has no 'paths' — is this an OpenAPI document? The URL returned JSON that is not an OpenAPI document (an endpoint's payload, an error body) Open the URL in a browser and check for "openapi" and "paths"
base_url must start with http:// or https:// (got '/api'); the spec declares a relative server URL '/api'; pass --base-url https://<api-host>/api A spec read from a file with a relative servers entry (FastAPI root_path, DRF generators) Fetch it over HTTP so it resolves against the spec URL, or pass --base-url
Tools named search_books_search_post, get_books_book_id_get No operation_id on the routes Step 2 — set them in the app; renaming in the plan is undone by the next regeneration from the spec
UPSTREAM_AUTH_MISSING: MCPCAST_UPSTREAM_TOKEN is not set. Auth mode 'env-token' presents that value as the Authorization header on every upstream call … The server has no credential in its environment Put it in the client's env block (or -e for claude mcp add); a shell export does not reach a GUI-launched client
UPSTREAM_AUTH_MISSING on a passthrough server The MCP client sent no Authorization header — stdio clients cannot Regenerate with --auth env-token for a personal server, or run over HTTP with the header
UPSTREAM_ERROR: GET /books/{book_id} returned HTTP 401: {"detail":"invalid token"} Your app rejected the token — MCPCAST_UPSTREAM_TOKEN is wrong Set the token your app accepts, scheme included: Bearer demo-token
UPSTREAM_ERROR: … HTTP 401: {"detail":"Not authenticated"} The variable holds the bare token without Bearer The value is the whole header value
UPSTREAM_UNREACHABLE: GET /books/{book_id} failed: ConnectError The app is not running at api.base_url — for the example, the port changes every run Start the app, or set MCPCAST_BASE_URL
UPSTREAM_TIMEOUT: GET /books did not complete within 30s (MCPCAST_TIMEOUT) The app accepted the connection but did not answer in time — the whole call, headers and body, has a wall-clock deadline Raise MCPCAST_TIMEOUT for slow endpoints; the error is marked retryable
APPROVAL_DENIED: Approval denied for tool 'update_book': client declined or returned an invalid elicitation response A gated tool was called from a client that cannot show an approval prompt Expected. Use a client that implements elicitation, --approval pending with --auth api-key, or an approval_handler in tests
bookshelf-mcp already contains an mcpcast project — pass --force to overwrite it, or regenerate from bookshelf-mcp/mcpcast.plan.yaml to keep your edits Re-running against the spec would discard your plan edits Regenerate from the plan to keep them; --force only if you mean to start over
auth mode 'none' refuses to bind to a non-loopback address --auth none served on 0.0.0.0 none is for local demos; regenerate with a real auth mode

The whole example

examples/mcp/mcpcast_fastapi_app/run.py — start the app, generate, drive with a real agent, deny and approve a write, print the shipping commands. It needs fastapi, uvicorn and OPENAI_API_KEY; the app it MCPcasts is app.py in the same directory.

"""Make your own FastAPI app MCP-ready — end to end, with real calls all the way.

``app.py`` is an ordinary FastAPI service (the Bookshelf API). This driver does
exactly what you would do to your own app, in one run:

  1. START     serve ``app.py`` in-process with uvicorn on a free loopback port
  2. GENERATE  ``promptise.mcpcast`` reads ``http://127.0.0.1:<port>/openapi.json``
               and writes an editable MCP server project to ``generated/``
               (profile ``standard``: reads open, writes approval-gated;
               auth ``env-token``: the server presents one bearer token upstream)
  3. DRIVE     ``build_agent("openai:gpt-5-mini")`` launches ``generated/server.py``
               over the real MCP stdio transport and answers a question by
               calling the live app through the generated tools
  4. GOVERN    the same agent tries a write — denied fail-closed over stdio
               (no human can be asked); then, in-process with an approver, the
               same call runs against the live app
  5. SHIP      what to run once this is your product's MCP server

``app.py`` is a FastAPI app, so ``fastapi`` must be installed
(``.venv/bin/python -m pip install fastapi``; ``promptise[dev]`` includes it).
Beyond that only ``OPENAI_API_KEY`` is needed. Steps 1-2 run offline; the script
stops with a clear message before step 3 if the key is missing.

Run:
    OPENAI_API_KEY=... .venv/bin/python examples/mcp/mcpcast_fastapi_app/run.py
"""

from __future__ import annotations

import asyncio
import json
import os
import socket
import sys
import threading
import time
from pathlib import Path
from types import ModuleType
from typing import Any

import uvicorn
from langchain_core.messages import AIMessage, HumanMessage, ToolMessage

from promptise import StdioServerSpec, build_agent
from promptise.approval import ApprovalRequest
from promptise.mcp.server import TestClient
from promptise.mcpcast import (
    AuthMode,
    DroppedOp,
    MCPcastPlan,
    SafetyProfile,
    load_generated_server,
    mcpcast,
    write_project,
)
from promptise.models import load_dotenv_if_present

try:
    import app as bookshelf  # the Bookshelf API next to this file — a FastAPI app
except ModuleNotFoundError as exc:
    if (exc.name or "").partition(".")[0] != "fastapi":
        raise
    raise SystemExit(
        "app.py is a FastAPI app and fastapi is not a Promptise dependency — "
        ".venv/bin/python -m pip install fastapi"
    ) from None

HERE = Path(__file__).resolve().parent
OUT = HERE / "generated"
MODEL = "openai:gpt-5-mini"
# What the generated server sends upstream as the Authorization header. In a
# desktop client this goes in the client's own config (see generated/README.md).
UPSTREAM_TOKEN = f"Bearer {bookshelf.DEMO_TOKEN}"
QUESTION = "Which books by Ursula K. Le Guin do we have, and which of them is the oldest?"
WRITE_REQUEST = "Add the note 'signed first edition' to The Dispossessed."


def section(number: int, title: str) -> None:
    """Print a numbered section header."""
    print(f"\n{'=' * 78}\n{number}. {title}\n{'=' * 78}")


# ---------------------------------------------------------------------------
# 1. Start your app
# ---------------------------------------------------------------------------


def start_app() -> tuple[uvicorn.Server, str]:
    """Serve ``app.py`` in a background thread and return the server and its origin."""
    with socket.socket() as probe:
        probe.bind(("127.0.0.1", 0))
        port = probe.getsockname()[1]
    config = uvicorn.Config(bookshelf.app, host="127.0.0.1", port=port, log_level="warning")
    server = uvicorn.Server(config)
    thread = threading.Thread(target=server.run, daemon=True)
    thread.start()
    deadline = time.monotonic() + 10
    while not server.started:
        if not thread.is_alive():
            raise SystemExit(f"uvicorn failed to start on port {port} (see the error above)")
        if time.monotonic() > deadline:
            raise SystemExit(f"uvicorn did not start on port {port} within 10 s")
        time.sleep(0.05)
    return server, f"http://127.0.0.1:{port}"


# ---------------------------------------------------------------------------
# 2. Generate the MCP server from the running app's spec URL
# ---------------------------------------------------------------------------


def generate(origin: str) -> MCPcastPlan:
    """``/openapi.json`` -> risk-classified plan -> ``generated/`` project."""
    section(2, "Generate the MCP server from the app's /openapi.json")
    spec_url = f"{origin}/openapi.json"
    plan = mcpcast(
        spec_url,
        profile=SafetyProfile.STANDARD,
        auth=AuthMode.ENV_TOKEN,
        name="bookshelf",
    )
    # FastAPI emits no `servers` block, so the API base was resolved against
    # the URL the spec was fetched from — the app's own origin.
    print(f"  spec:     {spec_url}")
    print(f"  base_url: {plan.api.base_url}   (no servers block -> the spec URL's origin)")

    # Your review pass: a liveness probe is not something an agent should call.
    # `--curate` would drop it for you; deterministic generation keeps every
    # read, so move it to `dropped` — the plan is the file you own.
    health = plan.tool("health_check")
    plan = MCPcastPlan(
        api=plan.api,
        profile=plan.profile,
        tools=[t for t in plan.tools if t.name != "health_check"],
        dropped=[
            *plan.dropped,
            DroppedOp(
                operation_id=health.operations[0],
                reason="operational endpoint — for the load balancer, not an agent",
            ),
        ],
    )
    for path in write_project(plan, OUT):
        print(f"  wrote {path.relative_to(HERE)}")

    print(f"\n  {'tool':<16}{'risk':<8}{'approval':<10}{'upstream operation':<26}params")
    for tool in plan.tools:
        route = tool.routes[0]
        gate = "required" if tool.requires_approval else "-"
        print(
            f"  {tool.name:<16}{tool.risk.value:<8}{gate:<10}"
            f"{route.method + ' ' + route.path:<26}{', '.join(tool.visible_params)}"
        )
    print("\n  not exposed (each with its reason, recorded in the plan):")
    for dropped in plan.dropped:
        print(f"    {dropped.operation_id}: {dropped.reason}")
    return plan


# ---------------------------------------------------------------------------
# 3 + 4a. A real agent over MCP stdio, against the live app
# ---------------------------------------------------------------------------


def final_text(result: Any) -> str:
    """The final assistant text of an agent invocation."""
    content = result["messages"][-1].content
    if isinstance(content, list):
        return "".join(c.get("text", "") if isinstance(c, dict) else str(c) for c in content)
    return str(content)


def show_calls(result: Any) -> None:
    """Print every tool call the agent made, with the server's structured errors."""
    for message in result["messages"]:
        if isinstance(message, AIMessage):
            for call in message.tool_calls:
                print(f"    {call['name']}({json.dumps(call['args'], ensure_ascii=False)})")
        elif isinstance(message, ToolMessage) and '"error"' in str(message.content):
            error = json.loads(str(message.content))["error"]
            print(f"      -> {error['code']}: {error['message']}")


async def drive_over_stdio() -> None:
    """The generated server, launched by the agent exactly as Claude Desktop would."""
    section(3, "A real agent over MCP stdio -> generated/server.py -> the live app")
    agent = await build_agent(
        model=MODEL,
        servers={
            "bookshelf": StdioServerSpec(
                command=sys.executable,
                args=[str(OUT / "server.py")],
                env={"MCPCAST_UPSTREAM_TOKEN": UPSTREAM_TOKEN},
            )
        },
        instructions=(
            "You are the Bookshelf assistant. Answer from what the tools return. "
            "Be brief. If a tool call is denied, say so and stop."
        ),
        max_agent_iterations=6,
    )
    try:
        print(f"  question: {QUESTION}\n  tools the agent chose:")
        result = await agent.ainvoke({"messages": [HumanMessage(content=QUESTION)]})
        show_calls(result)
        print(f"  answer: {final_text(result)}")

        section(4, "Writes: denied fail-closed over stdio, executed once a human approves")
        print(
            "  a) this MCP client does not support elicitation, so nobody can be asked and the gate denies:"
        )
        print(f"  request: {WRITE_REQUEST}\n  tools the agent chose:")
        result = await agent.ainvoke({"messages": [HumanMessage(content=WRITE_REQUEST)]})
        show_calls(result)
        print(f"  answer: {final_text(result)}")
    finally:
        await agent.shutdown()


# ---------------------------------------------------------------------------
# 4b. The same write, in-process, with an approver
# ---------------------------------------------------------------------------


def import_generated() -> ModuleType:
    """Import the generated project — its ``build_server()`` takes an approval handler."""
    return load_generated_server(OUT / "server.py")


async def approve_and_execute() -> None:
    """A reviewer says yes, the PATCH reaches the live app, and the read confirms it."""
    print("\n  b) in-process, with a human (here: a callback) who approves:")
    os.environ["MCPCAST_UPSTREAM_TOKEN"] = UPSTREAM_TOKEN  # env-token mode reads it per call

    def reviewer(request: ApprovalRequest) -> bool:
        print(f"    approval requested: {request.tool_name} {json.dumps(request.arguments)}")
        return True

    client = TestClient(import_generated().build_server(approval_handler=reviewer))
    (reply,) = await client.call_tool("get_book", {"book_id": 3})
    print(
        f"    get_book(3) before: notes={json.loads(reply.text)['notes']!r}  (the denied call changed nothing)"
    )
    (reply,) = await client.call_tool(
        "update_book", {"book_id": 3, "notes": "signed first edition"}
    )
    print(f"    update_book -> {reply.text}")
    (reply,) = await client.call_tool("get_book", {"book_id": 3})
    print(f"    get_book(3) after:  notes={json.loads(reply.text)['notes']!r}")


# ---------------------------------------------------------------------------


async def main() -> None:
    section(1, "Start the Bookshelf app (app.py) with uvicorn")
    app_server, origin = start_app()
    print(f"  serving {origin}  (bearer token: {bookshelf.DEMO_TOKEN!r})")
    try:
        generate(origin)
        load_dotenv_if_present()  # .env next to the project, as build_agent() would
        if not os.environ.get("OPENAI_API_KEY"):
            print("\nSteps 3-4 drive a real agent: set OPENAI_API_KEY and run this again.")
            print("  export OPENAI_API_KEY=sk-...")
            sys.exit(1)
        await drive_over_stdio()
        await approve_and_execute()
    finally:
        app_server.should_exit = True

    section(5, "Ship it")
    print("  generated/ is a real project: edit mcpcast.plan.yaml (never the package), regenerate.")
    print(
        "    cd generated && pytest                    # its own tests: listing, routing, the gate"
    )
    print("    pip install -e generated && bookshelf-mcp --transport http --port 8080")
    print("    cd generated && promptise serve server:server --transport http --port 8080")
    print(
        "    docker build -t bookshelf-mcp generated && docker run -i --env-file generated/.env \\"
    )
    print("      -e MCPCAST_BASE_URL=https://api.yourcompany.com bookshelf-mcp")
    print("  desktop clients (stdio) get the token in their own config: see generated/README.md")
    print(
        "  env-token = one shared credential and no caller authentication, so every HTTP form"
        " above binds loopback only"
    )
    print(
        "  and the image serves stdio; publish only behind an authenticating gateway"
        " (--host 0.0.0.0 --public)"
    )


if __name__ == "__main__":
    asyncio.run(main())

Next Steps

  • MCPcast, end to end — how MCP works, the full real run, and the review checklist

  • MCPcast an Existing API — the general walkthrough: the full profile, four-eyes review, curation and hand edits, acting on a readiness score, the nightly CI job

  • MCPcast Recipes — Stripe, GitHub, Swagger 2, and what to do when the spec fights back
  • MCPcast reference — risk rules, curation post-conditions, the plan schema, auth modes, readiness metrics
  • Approval Gates — elicitation, pending stores, webhooks and custom handlers
  • Deployment — promptise serve, the dashboard, transports, CORS
  • Building Production MCP Servers — for the tools your spec cannot express, hand-written and mounted alongside the generated ones
  • Examples gallery — every runnable example, including this one