Yes I read that one, and that’s why I said the explanation is not clear. Anyway, catastrophic forgetting is theoretically not possible. “ Existing capability must substantially learn new information, so its representation must change, while retaining its still-valid old knowledge”. I think your setup is demonstrating something weaker.
You might be right. I have not measured it on typical CL benchmarks yet. I will try to do that after the reading of the whole corpus. And don't forget there still is a normal slow forgetting happening, ofc. But my main evidence of the network doing something right is an ability to learn from a continuous stream of data. There is no other architecture I know of that can do it on a similar scale.
I think as an experiment your setup is interesting. But in realistic systems even standard evals, i.e. making sure things work well for normal use cases, are already enormous, let alone thinking about catastrophic forgetting.