Set Difference, Part 2
A two-layer, single-head-per-layer transformer reads a set of symbols followed by a permutation of that set, one symbol at a time, and predicts which symbols could still come next. Your goal is to explain the algorithm it implements, end to end.
Choose K distinct symbols to form X (2 ≤ K ≤ 8),
and let Y be an independent random permutation of X.
The model sees
[BOS] x1 … xK [SEP] y1 … yK [EOS].
At SEP and after each yi, it is trained to output the uniform
distribution over the symbols of X that have not yet appeared in Y;
once all of Y has been read, it predicts EOS.
| Vocabulary | 16 symbols, a–p, plus BOS, SEP, and EOS |
| Set size | K = 2–8 distinct symbols, mixed during training |
| Architecture | Two causal attention-only layers, one head each; RMSNorm before each attention layer and before the unembedding; no MLPs or biases |
| Dimensions | d_model=64, d_head=64 |
| Positions | Rotary position embeddings (RoPE) |
| Parameters | 35,392 |
| Accuracy | Every valid next token outranks every invalid one, at every position, on all held-out sets tested |
We are looking for a human-understandable computational account of how the model obtains its answer. A complete solution should connect the entire path from embeddings to attention to head outputs to the final readout.
Prioritize clarity and understanding over exhaustiveness. A short, well-explained account of what the model actually computes is worth more than a long collection of plots.
The starter notebook loads the pre-trained weights from HuggingFace and walks you through basic inference and attention visualization. Open it in Colab, save a copy to your own Drive, and start exploring.
Submit a link to a clean Colab notebook that explains your findings. Include well-labeled figures, state claims plainly, and distinguish observations from causal evidence. Think of the notebook as a presentation of the algorithm you found, not a transcript of exploratory work.
Deadline: October 31, 2026 (anywhere on Earth)