For PyTorch users
AI assistant
Build this with Claude Code, Cursor or Copilot
Copy a Talario-tuned prompt for For PyTorch users, grounded in 12 real API signatures , into your IDE's AI. No chatbot, just exact context.
Talario is not a faster PyTorch, and it isn't trying to be. PyTorch is where you research and train. Talario is where you ship — the deployment layer that runs your model on any GPU, gives the same answer on CPU and GPU, and lives inside your web app with no Python serving sidecar.
If your pain is serving — CUDA lock-in, a separate TorchServe/Triton/FastAPI tier, results that drift between environments — Talario is the alternative to that slice of your stack. You keep PyTorch for everything upstream.
What you get
- Any silicon, one artifact. Vendor-neutral SPIR-V on Vulkan runs on NVIDIA, AMD, Apple, Intel, and ARM — chosen at runtime, with a CPU fallback. No per-vendor kernels, no CUDA lock-in.
- Reproducible to the bit. Execution is atomic-free and CPU==GPU bit-identical for
exact ops (and ~1 ULP where
exp/sqrtare involved). A model trained on a cloud GPU gives the same result on a laptop CPU or a phone — auditable and debuggable. - In-process, zero-IPC serving. As a native extension the warm model lives in your request process (FrankenPHP/Swoole) — inference is one submission, no sidecar, no serialization hop.
The path (no rewrite)
You do not port your model to a new language. Keep training in PyTorch, then hand the trained graph to Talario via ONNX:
# 1) train in PyTorch, then export (ONNX import matches PyTorch to ~1e-7 in Talario) import torch torch.onnx.export(model, example_input, "model.onnx", input_names=["x"], output_names=["logits"], opset_version=17)
// 2) run it in Talario — portable, reproducible, in-process $model = Talario\Onnx\onnx_import($dev, 'model.onnx', [1, 784]); // fn(float[]) => float[] $logits = $model($features);
# 3) or serve it behind an HTTP endpoint and call it from anything (Python included) php -d extension=modules/talario.so php/serve.php http gpu 127.0.0.1:5055 curl -s -X POST http://127.0.0.1:5055/infer -d '[/* features */]' # {"logits":[...]}
A small Python client (pip install -e clients/python) wraps that /infer call so you call
the portable Talario endpoint from Python exactly like a local function, and
export_for_talario() / run_via_talario() cover the ONNX path — training stays in PyTorch,
serving becomes portable and reproducible. It is NumPy-interoperable (not a NumPy-native API);
a deep Python binding is planned (see docs/project/python-binding-design.md). See
[clients/python/README.md(https://github.com/DMJ-CV-91913/talario/blob/main/clients/python/README.md).
When to reach for Talario — and when not
Reach for it when you need to: deploy the same model across mixed GPUs without CUDA lock-in; guarantee reproducible/auditable inference (regulated, scientific, finance); or serve in-process inside a PHP/web app with no separate model service.
Stay on PyTorch when your bottleneck is large-model training throughput — Talario is f32 and does not use Tensor Cores today, so for frontier training on A100/H100-class hardware, PyTorch+CUDA wins. Talario's bet is portability + reproducibility + web-native deployment, not peak FLOPs. Use both: PyTorch to train, Talario to ship.
Native authoring (optional)
If you are in the PHP/web world, Talario is a full define-by-run framework in its own right
— tape autograd, Linear/Conv2d/LayerNorm/RMSNorm/Embedding/attention/transformer
blocks, GNNs, SGD/Adam, save/load. See the [User Guide(https://github.com/DMJ-CV-91913/talario/blob/main/docs/USER-GUIDE.md). But you never have
to leave PyTorch to adopt the deployment layer.