Skip to content

Transformers

AI assistant

Build this with Claude Code, Cursor or Copilot

Copy a Talario-tuned prompt for Transformers, grounded in 12 real API signatures , into your IDE's AI. No chatbot, just exact context.

Talario trains a pre-norm Transformer encoder block natively in PHP, on any GPU (Vulkan/SPIR-V) or CPU, with no Python, no serving sidecar, no new hand-written attention kernel. The block is composed from gradcheck-verified primitives, and the whole block is itself gradchecked end-to-end on CPU and GPU.

Status

  • Proven: forward + backward of a single-head, batch-1 pre-norm encoder block, with a passing finite-difference gradcheck on CPU and GPU (transformer_block in php/suite.php), and a training demo that overfits a sequence pair (loss falls sharply) in php/transformer_demo.php.
  • Not yet: batched / multi-head attention (needs a batched matmul bmm + a head permute, on the roadmap) and Dropout. MultiHeadAttention throws for nhead > 1.

The primitives it's built from

Added this round, each with its own CPU↔GPU gradcheck:

Op PHP Role in a Transformer
GELU gelu() / Activation('gelu') FFN activation (alt to ReLU)
Softmax Tape::softmax, softmax() attention weights (row-wise)
LayerNorm LayerNorm layer, Tape::layerNorm norm1 / norm2
Transpose Tape::transpose, transpose() Kᵀ for QKᵀ
Scale Tape::scale, scale() /√d attention scaling

plus the pre-existing matmul, Linear, relu, and residual add.

The classes

use Talario\{Device, Tape, TransformerEncoderLayer};

$dev   = new Device('gpu');                 // or 'cpu'
$block = new TransformerEncoderLayer($dev, dim: 64, dff: 256, nhead: 1, seed: 1);

$t   = new Tape($dev);
$x   = $t->tensor($features, [$L, 64], false);   // [seq_len, dim], batch=1
$out = $block($t, $x);                            // [seq_len, dim]
  • MultiHeadAttention($dev, $dim, $nhead = 1, $seed): single-head scaled dot-product self-attention: Q,K,V,O are Linear projections, and out = Wo( softmax( scale(Q·Kᵀ, 1/√d) ) · V ).
  • TransformerEncoderLayer($dev, $dim, $dff, $nhead = 1, $seed): pre-norm: x = x + Attn(LN1(x)); x = x + FFN(LN2(x)), FFN = Linear→ReLU→Linear. Mirrors a PyTorch TransformerEncoderLayer(norm_first=True).

Both expose params() / state() / loadState(), so they drop into the normal training loop (Adam/SGD) and weight save/load.

Mapping from PyTorch

self.norm1 = nn.LayerNorm(d_model)          # -> LayerNorm
self.self_attn = nn.MultiheadAttention(d_model, nhead)   # -> MultiHeadAttention (nhead=1 today)
self.feed_forward = nn.Sequential(          # -> Sequential([Linear, relu(), Linear])
    nn.Linear(d_model, d_ff), nn.ReLU(), nn.Linear(d_ff, d_model))
self.norm2 = nn.LayerNorm(d_model)          # -> LayerNorm
# forward: x = x + attn(norm1(x)); x = x + ff(norm2(x))   # -> TransformerEncoderLayer

Run the demo

bin/kphp php/transformer_demo.php cpu 300     # or gpu

Overfits one (input → target) sequence pair; the loss should drop >10×, proving the block trains end-to-end (forward + backward + Adam) through LayerNorm, attention, and the FFN.

How correctness is guaranteed

Every primitive has a finite-difference gradcheck (CPU and GPU) in php/suite.php, and so does the assembled block (transformer_block). A passing gradcheck means the gradients are numerically correct; the block genuinely trains, it doesn't just run.

Roadmap: bmm + head-permute → batched multi-head attention; Dropout; stacked N-layer encoder. See docs/project/outstanding-issues.md.

Edit this page on GitHub