Type a sentence. A real Transformer (BERT) runs in your browser and shows its self-attention — which words each word looks at — across all 6 layers × 12 heads.
Loading the Transformer (~91 MB, first time only)…
Layer
Head
View
Each row is a word doing the looking; brighter cells are the words it attends to. Hover a cell for the exact weight, or hover a token to isolate its row/column.
lowhighattention weight (each row scaled to its own max)