Skip to content

ONNX import/export

AI assistant

Build this with Claude Code, Cursor or Copilot

Copy a Talario-tuned prompt for ONNX import/export, grounded in 12 real API signatures , into your IDE's AI. No chatbot, just exact context.

Talario imports ONNX models and runs them natively in PHP: no onnxruntime, no libtorch, no Python. You train anywhere (PyTorch/TF → ONNX) and deploy in your PHP backend, on any GPU. This is the ecosystem play: borrow the industry's model zoo rather than rebuild it.

Status

  • Working: a minimal ONNX protobuf reader + op-mapper imports a model and reproduces its output to < 1e-4 of a reference, on CPU and GPU (onnx import gate in php/suite.php). Proven self-contained; the test fixture is a real ONNX file emitted by php/onnx/fixture_mlp.php (valid wire format; a conforming onnxruntime could read it).
  • Real PyTorch exports imported: a genuine torch.onnx.export ReLU-MLP matches PyTorch to ~1e-8 (php/onnx_import_real.php), and a real pre-norm Transformer block (php/onnx/fixtures/block.onnx, exported at opset 17) matches PyTorch to <1e-4 on CPU and GPU (suite: onnx transformer block). No onnxruntime/PyTorch at run time. Fixtures via tools/export_onnx_fixture.py.
  • Full nn.TransformerEncoderLayer imported: the default PyTorch module, exported by the dynamo exporter (43 nodes, external-data weights), imports and matches PyTorch to ~1e-7 on CPU and GPU (suite onnx transformer (nn.TransformerEncoderLayer)).
  • Supported ops today: Gemm, MatMul (incl. batched/N-D), Add (bias + scalar), Mul, Div, Relu, Gelu, Erf, Sigmoid, Tanh, LayerNormalization, Softmax, Transpose (any perm), Reshape, Squeeze, Unsqueeze, Gather, Constant, Identity, plus external-data (.onnx.data) weights and N-D tensors. Unsupported ops throw a clear error at import time.
  • Foundation in place: the .tal graph-plan snapshot (CompiledGraph::save / Talario\load_compiled, bit-identical round-trip), so an imported graph can be compiled once and shipped as a .tal (the zero-overhead production path; Phase 2).

How it works

model.onnx ──▶ protobuf decoder ──▶ op-mapper ──▶ Talario Tape ──▶ run (any GPU / CPU)
 (no onnxruntime; a small wire-format reader)   (ONNX node → Tape op)
  • php/onnx/import.php: a ~100-line protobuf decoder (_decode: varint / length-delimited / f32) + parse_onnx() + the op-mapper.
  • Inference-first: initializers (weights) are baked in as constants; the graph input is the fed tensor; the mapper walks nodes in order, emitting Tape ops.

Usage

require 'php/talario.php';
require 'php/onnx/import.php';

use Talario\Device;
use function Talario\Onnx\onnx_import;

$dev = new Device('gpu');                               // or 'cpu'
$run = onnx_import($dev, '/path/to/model.onnx', [1, $IN]);  // bytes or a file path
$logits = $run($features);                              // float[] -> float[]

Run the showcase

bin/kphp php/onnx_demo.php cpu      # or gpu

Emits a valid ONNX MLP, imports it, runs it, and checks the output against an analytic reference.

Fixtures (two routes)

  • Self-contained (no PyTorch): Talario\Onnx\make_onnx_mlp($w1, $in_h, $w2, $h_out) emits real ONNX bytes for relu(x@W1)@W2. Used by the parity gate.
  • Real exports: export a model from PyTorch/TF to .onnx and import it directly (fully exercises the real exporter, the faithful validation route).

Roadmap

  • Widen the op-mapper: Gemm/Add (bias MLPs, the common PyTorch export), then LayerNormalization/Softmax/Gelu/Transpose/MatMul (→ transformer blocks, Talario already has every op these need) and Conv/MaxPool (NCHW→NHWC).
  • Phase 2: import → .tal emit: compile the mapped graph once and serialize it (reusing the .tal snapshot) → warm, zero-overhead InferenceService::fromArtifact.
  • Shape specialization / batch buckets (onnx-import-design.md §5).

Related: docs/project/onnx-import-design.md (design), onnx-import-build-plan.md (phases), tal-format.md (the .tal artifact), transformers.md (the ops it maps to).

Edit this page on GitHub