Beats me how it works, honestly can't wrap my head around it.
From what I understand, at position Aha in each layer it's constructing a query based on the current activation and looking at the key of each other token position for that layer, in order to decide how much attention to pay to the value.
In this way it attends to the previous values, such as perhaps the incorrect assumption and plausible explanation.