You might be right. I have not measured it on typical CL benchmarks yet. I will try to do that after the reading of the whole corpus. And don't forget there still is a normal slow forgetting happening, ofc. But my main evidence of the network doing something right is an ability to learn from a continuous stream of data. There is no other architecture I know of that can do it on a similar scale.
I think as an experiment your setup is interesting. But in realistic systems even standard evals, i.e. making sure things work well for normal use cases, are already enormous, let alone thinking about catastrophic forgetting.