Is this architecture actually able to generalize or is it mostly based on memorization? Have you tried some basic tasks that require generalization? e.g. number addition etc?
The model is way too small and undertrained to make any generalization claims. I want to wait until it reads the whole corpus I gave and then test it on some simple established benchmarks to see how it will behave.
Seems a bit premature to make an HN post about then, imho.
It's an interesting idea, but it doesn't really do anything interesting yet. I looked at the output in the training run and it is a far, far cry from intelligence. Worse than GPT-2 as it stands.
I do hope it will perform well when scaled and trained, though; best of luck.
And I'm not sure video, images and motion are actually 3 different things. Images are just still motion and videos capture motion, so it's really just "video and audio" of which one is visual, the other is not, thus my confused/surprised comment.
reply