Neural Nets & Transformers quiz
24 questions from the Neural Nets & Transformers track. Score 80% or higher to unlock a certificate.
Tokens
1. Why do modern language models use subword tokens instead of whole words?
Byte Pair Encoding
2. What does one BPE merge do?
Special Tokens
3. What is a chat template for?
The Embedding Table
4. What does an embedding table do?
Positions
5. Why do transformers need a position signal?
Attention
6. In causal self-attention, why are some scores set to a huge negative number?
Multi-Head Attention
7. What problem do multiple attention heads address?
The Transformer Block
8. What do residual connections do in a transformer block?
Residuals and LayerNorm
9. Why add the layer output back to the input (a residual)?
Encoder vs Decoder
10. Why must a GPT-style decoder hide future tokens?
Pretraining
11. What is the standard pretraining task for GPT-style models?
Logits and Unembedding
12. What is a logit in a language model?
Decoding
13. When should an agent decode tool names greedily?
The KV Cache
14. What does the KV cache store?
Context Windows
15. What is a common dangerous truncation bug?
Prefill and Decode
16. Why is time-to-first-token mostly prefill?
Fine-Tuning
17. What does SFT mainly copy?
LoRA
18. What is LoRA for?
Limitations
19. Why do transformers hallucinate facts?
Lost in the Middle
20. Where should the latest tool observation usually sit in a long prompt?
Scaling
21. What do scaling laws mainly predict?
Padding and Masks
22. What is an attention mask for padding doing?
Why Transformers
23. What is a main training advantage of transformers over classic RNNs?
The Stack an Agent Actually Runs
24. What is the main thing you control in a hosted transformer agent?