Skip to content

Tutorial: MCP & LLM CLIs

AI assistant

Build this with Claude Code, Cursor or Copilot

Copy a Talario-tuned prompt for Tutorial: MCP & LLM CLIs, grounded in 12 real API signatures , into your IDE's AI. No chatbot, just exact context.

Goal: let an MCP-capable LLM CLI (Claude Code, or any MCP client) call your Talario model as a tool — "run the resident model on these features and tell me the class." The LLM talks to a running Talario endpoint; it never needs the source, so this works while your repo stays private.

You need: a running Talario endpoint (from [the PyTorch tutorial(https://github.com/DMJ-CV-91913/talario/blob/main/docs/tutorials/pytorch-to-talario.md) or php/serve.php), and Python with pip install "mcp[cli]" httpx.


Step 1 — have a Talario endpoint running

Any of these exposes GET /health + POST /infer:

# serve an ONNX model you exported from PyTorch:
php -d extension=modules/talario.so php/onnx_serve.php mlp.onnx cpu 127.0.0.1:5055 784
# or the built-in demo server:
php -d extension=modules/talario.so php/serve.php http cpu 127.0.0.1:5055

Step 2 — run the MCP server

The reference server (clients/mcp/talario_mcp_server.py) forwards MCP tool calls to that endpoint:

TALARIO_ENDPOINT=http://127.0.0.1:5055 \
  python clients/mcp/talario_mcp_server.py        # speaks MCP over stdio

Step 3 — register it with your LLM CLI

Claude Code:

claude mcp add talario -- python /abs/path/clients/mcp/talario_mcp_server.py
# if your endpoint isn't the default, set the env in the server config:
#   TALARIO_ENDPOINT=http://host:5055

Any other MCP client: point it at the same command (python .../talario_mcp_server.py, stdio transport). The server exposes two tools:

Tool Input Returns
talario_health — {status, device, in} — the resident model's device + input size
talario_infer features: number[] {output: number[], class: int} (argmax)

Step 4 — use it from a prompt

Once registered, just ask:

"Use talario_health to check the model, then call talario_infer on this 784-length vector […] and tell me the predicted class."

The LLM calls the tools, Talario runs the warm model, and because Talario is CPU==GPU reproducible the answer is stable no matter where the endpoint runs.

Alternative: no MCP, just a prompt (any CLI)

For an LLM CLI without MCP, let it call the endpoint directly — give it this in a prompt or a reusable skill:

curl -s -X POST "$TALARIO_ENDPOINT/infer" \
     -H 'Content-Type: application/json' -d "$FEATURES_JSON"
# -> {"output":[...]}  (or {"logits":[...]} from serve.php)

Deploying the endpoint for remote/agent use

Keep the endpoint on localhost for local agents. To let a hosted agent reach it, front it with a reverse proxy + auth — e.g. Cloudflare Tunnel for a stable hostname and Cloudflare Access to gate /infer — while the GPU box runs elsewhere. The MCP server only needs TALARIO_ENDPOINT to point at the public URL; it never touches the Talario source.

Why this is useful

  • Agents get a reproducible numerical tool. The model's output is deterministic across CPU/GPU, so an agent's tool calls don't drift between environments.
  • Any silicon, no Python ML stack on the agent side. The agent host only needs httpx; the model runs in the Talario engine wherever you placed it.
  • Private by default. The source never leaves your machine; the agent only sees the live endpoint.

Edit this page on GitHub