Transformers
AI assistant
Build this with Claude Code, Cursor or Copilot
Copy a Talario-tuned prompt for Transformers, grounded in 12 real API signatures , into your IDE's AI. No chatbot, just exact context.
Talario trains a pre-norm Transformer encoder block natively in PHP, on any GPU (Vulkan/SPIR-V) or CPU, with no Python, no serving sidecar, no new hand-written attention kernel. The block is composed from gradcheck-verified primitives, and the whole block is itself gradchecked end-to-end on CPU and GPU.
Status
- Proven: forward + backward of a single-head, batch-1 pre-norm encoder block, with
a passing finite-difference gradcheck on CPU and GPU (
transformer_blockinphp/suite.php), and a training demo that overfits a sequence pair (loss falls sharply) inphp/transformer_demo.php. - Not yet: batched / multi-head attention (needs a batched matmul
bmm+ a head permute, on the roadmap) and Dropout.MultiHeadAttentionthrows fornhead > 1.
The primitives it's built from
Added this round, each with its own CPU↔GPU gradcheck:
| Op | PHP | Role in a Transformer |
|---|---|---|
| GELU | gelu() / Activation('gelu') |
FFN activation (alt to ReLU) |
| Softmax | Tape::softmax, softmax() |
attention weights (row-wise) |
| LayerNorm | LayerNorm layer, Tape::layerNorm |
norm1 / norm2 |
| Transpose | Tape::transpose, transpose() |
Kᵀ for QKᵀ |
| Scale | Tape::scale, scale() |
/√d attention scaling |
plus the pre-existing matmul, Linear, relu, and residual add.
The classes
use Talario\{Device, Tape, TransformerEncoderLayer}; $dev = new Device('gpu'); // or 'cpu' $block = new TransformerEncoderLayer($dev, dim: 64, dff: 256, nhead: 1, seed: 1); $t = new Tape($dev); $x = $t->tensor($features, [$L, 64], false); // [seq_len, dim], batch=1 $out = $block($t, $x); // [seq_len, dim]
MultiHeadAttention($dev, $dim, $nhead = 1, $seed): single-head scaled dot-product self-attention:Q,K,V,OareLinearprojections, andout = Wo( softmax( scale(Q·Kᵀ, 1/√d) ) · V ).TransformerEncoderLayer($dev, $dim, $dff, $nhead = 1, $seed): pre-norm:x = x + Attn(LN1(x)); x = x + FFN(LN2(x)),FFN = Linear→ReLU→Linear. Mirrors a PyTorchTransformerEncoderLayer(norm_first=True).
Both expose params() / state() / loadState(), so they drop into the normal
training loop (Adam/SGD) and weight save/load.
Mapping from PyTorch
self.norm1 = nn.LayerNorm(d_model) # -> LayerNorm self.self_attn = nn.MultiheadAttention(d_model, nhead) # -> MultiHeadAttention (nhead=1 today) self.feed_forward = nn.Sequential( # -> Sequential([Linear, relu(), Linear]) nn.Linear(d_model, d_ff), nn.ReLU(), nn.Linear(d_ff, d_model)) self.norm2 = nn.LayerNorm(d_model) # -> LayerNorm # forward: x = x + attn(norm1(x)); x = x + ff(norm2(x)) # -> TransformerEncoderLayer
Run the demo
bin/kphp php/transformer_demo.php cpu 300 # or gpu
Overfits one (input → target) sequence pair; the loss should drop >10×, proving the block trains end-to-end (forward + backward + Adam) through LayerNorm, attention, and the FFN.
How correctness is guaranteed
Every primitive has a finite-difference gradcheck (CPU and GPU) in php/suite.php, and
so does the assembled block (transformer_block). A passing gradcheck means the
gradients are numerically correct; the block genuinely trains, it doesn't just run.
Roadmap: bmm + head-permute → batched multi-head attention; Dropout; stacked
N-layer encoder. See docs/project/outstanding-issues.md.