Tutorial: MCP & LLM CLIs
AI assistant
Build this with Claude Code, Cursor or Copilot
Copy a Talario-tuned prompt for Tutorial: MCP & LLM CLIs, grounded in 12 real API signatures , into your IDE's AI. No chatbot, just exact context.
Goal: let an MCP-capable LLM CLI (Claude Code, or any MCP client) call your Talario model as a tool — "run the resident model on these features and tell me the class." The LLM talks to a running Talario endpoint; it never needs the source, so this works while your repo stays private.
You need: a running Talario endpoint (from [the PyTorch tutorial(https://github.com/DMJ-CV-91913/talario/blob/main/docs/tutorials/pytorch-to-talario.md) or
php/serve.php), and Python with pip install "mcp[cli]" httpx.
Step 1 — have a Talario endpoint running
Any of these exposes GET /health + POST /infer:
# serve an ONNX model you exported from PyTorch: php -d extension=modules/talario.so php/onnx_serve.php mlp.onnx cpu 127.0.0.1:5055 784 # or the built-in demo server: php -d extension=modules/talario.so php/serve.php http cpu 127.0.0.1:5055
Step 2 — run the MCP server
The reference server (clients/mcp/talario_mcp_server.py) forwards MCP tool calls to that
endpoint:
TALARIO_ENDPOINT=http://127.0.0.1:5055 \ python clients/mcp/talario_mcp_server.py # speaks MCP over stdio
Step 3 — register it with your LLM CLI
Claude Code:
claude mcp add talario -- python /abs/path/clients/mcp/talario_mcp_server.py # if your endpoint isn't the default, set the env in the server config: # TALARIO_ENDPOINT=http://host:5055
Any other MCP client: point it at the same command (python .../talario_mcp_server.py, stdio
transport). The server exposes two tools:
| Tool | Input | Returns |
|---|---|---|
talario_health |
— | {status, device, in} — the resident model's device + input size |
talario_infer |
features: number[] |
{output: number[], class: int} (argmax) |
Step 4 — use it from a prompt
Once registered, just ask:
"Use talario_health to check the model, then call talario_infer on this 784-length vector
[…]and tell me the predicted class."
The LLM calls the tools, Talario runs the warm model, and because Talario is CPU==GPU reproducible the answer is stable no matter where the endpoint runs.
Alternative: no MCP, just a prompt (any CLI)
For an LLM CLI without MCP, let it call the endpoint directly — give it this in a prompt or a reusable skill:
curl -s -X POST "$TALARIO_ENDPOINT/infer" \ -H 'Content-Type: application/json' -d "$FEATURES_JSON" # -> {"output":[...]} (or {"logits":[...]} from serve.php)
Deploying the endpoint for remote/agent use
Keep the endpoint on localhost for local agents. To let a hosted agent reach it, front it with
a reverse proxy + auth — e.g. Cloudflare Tunnel for a stable hostname and Cloudflare
Access to gate /infer — while the GPU box runs elsewhere. The MCP server only needs
TALARIO_ENDPOINT to point at the public URL; it never touches the Talario source.
Why this is useful
- Agents get a reproducible numerical tool. The model's output is deterministic across CPU/GPU, so an agent's tool calls don't drift between environments.
- Any silicon, no Python ML stack on the agent side. The agent host only needs
httpx; the model runs in the Talario engine wherever you placed it. - Private by default. The source never leaves your machine; the agent only sees the live endpoint.