Going beyond mere token input embeddings to the attention, self-attention and multi-head attention process that lies at the heart of transformer architecture and generative prompt/completion models. The central section contains minor additions to the first and third.