Yes I read that one, and that’s why I said the explanation is not clear. Anyway, catastrophic forgetting is theoretically not possible. “ Existing capability must substantially learn new information, so its representation must change, while retaining its still-valid old knowledge”. I think your setup is demonstrating something weaker.
You might be right. I have not measured it on typical CL benchmarks yet. I will try to do that after the reading of the whole corpus. And don't forget there still is a normal slow forgetting happening, ofc. But my main evidence of the network doing something right is an ability to learn from a continuous stream of data. There is no other architecture I know of that can do it on a similar scale.
I think as an experiment your setup is interesting. But in realistic systems even standard evals, i.e. making sure things work well for normal use cases, are already enormous, let alone thinking about catastrophic forgetting.
I am not sure about eventually, but it learns on the steady paste so far. The main thing to keep in mind is that it learns on the single STREAM of data. Not randomized batched samples. Try to do it with any other model and you will see nothing but complete garbage in the predictions, exactly because of catastrophic forgetting.
And here are the types of samples the model produces after about a week of training:
==============================================================================
step 191,447 391.3M of 7,879M characters (4.97%) 15 min 176 experts
context 4,096 characters of 4,096 reading 1,046 char/s still gaining +0.0412 deep into it
grad norm 0.98 against a clip of 1 under the clip
train loss 0.6540 lr 2.28e-04 evidence t -0.15 over 65.7 (effect +0.0660) rate x0.753
held-out loss 0.8012 +/-0.0335 nats 1.1559 bits/char perplexity 2.23 gap +0.1472
arithmetic 0.645 chat 0.790 chat_hermes 1.153 chess 0.510 code 0.703 reasoning 0.756 stories 0.596 wikipedia 1.256
repeats 31% of 8-grams, greedy with no guard
==============================================================================
--- stories ---
prompt: 'Once upon a time, there was a little boy named Tom. One day he '
[raw] repeated 8-grams 2%
was walking in the park with his mom. He saw a big box in the ground. He was very sad and sad.
"Mom, can I have the box to the box?" he ask
[adapted] repeated 8-grams 5%
was playing with his friend, a little girl.
"Wow, Tom!" said Tom.
"What's your name?"
Tom said, "I'm playing with my friend."
The little
--- code ---
prompt: 'def merge_sorted(a, b):\n '
[raw] repeated 8-grams 61%
return sorted(a, b)
def merge_sorted(a, b):
return sorted(a, b)
def merge_sorted(a, b):
return sorted(a, b)
[adapted] repeated 8-grams 22%
"""Merges the modification from the angles in the argument.
This merges the angles in the argument.
"""
def __init__(self, a):
--- arithmetic ---
prompt: 'add 4917 + 388 = '
[raw] repeated 8-grams 47%
7+8+0=5c1 1+8+1=0c1 9+3+1=3c1 4+0+1=5c0 5305
mul 6 * 6 = 36
mul 6 * 6 = 36
mul 6 * 6 = 36
mul 6 * 6 = 36
mul 6 * 6 = 36
mul 6 * 6 = 36
a
[adapted] repeated 8-grams 3%
7+8+0=5c1 1+8+1=0c1 9+3+1=3c1 4+0+1=5c0 5305
mul 6287 * 49 = 6287*9=56583 6287*40=251480 56583+251480=307063 307063
mul 9428 * 17 =
--- chat ---
prompt: '\nWhat are you?\n\n\n'
[raw] repeated 8-grams 20%
In the given I am trying to find the maximum vertex should be an expression linear that should be given the maximum vertex and the maximum v
[adapted] repeated 8-grams 8%
In the given I, there's another part of the given I, whose parts include anything they're discussed.
What is the given article by answer
--- chat_hermes ---
prompt: '\nA train travels 60 km in 45 minutes. What is its speed in km/h?\n\n\n'
[raw] repeated 8-grams 63%
The train travels 60 km/h in 45 minutes. The train travels 60 km/h in 45 minutes. The train travels 60 km/h in 45 minutes. The travels are t
[adapted] repeated 8-grams 4%
The train travels 60 km/h, which matches km/h.
Write an example of the following article:
In the subject of the subject, include an imp
--- reasoning ---
prompt: '\nWrite a Python function that returns the largest number in a list.\n\n\n\n'
[raw] repeated 8-grams 35%
Okay, let's tackle this problem. So, the largest number in a list is a list. The largest number is a list. The largest number is a list. It
[adapted] repeated 8-grams 20%
Okay, let's try to figure out this problem. So, we'll find the largest number in a list.
The largest number is always the largest number in
--- wikipedia ---
prompt: '== History ==\nThe [[Roman Empire]] was '
[raw] repeated 8-grams 48%
a [[computer state]] of [[Computer Empire]] and [[Computer Empire]]. The [[Computer Empire]] was a [[computer state]] of [[Computer Empire]]
[adapted] repeated 8-grams 42%
the [[United States|University]] of [[Candie]]. The [[University]] was the [[University]] of [[Candie]] where the [[University]] was the [[U
--- chess ---
prompt: '1700 1-0 1. e4 e5 2. '
[raw] repeated 8-grams 0% 22 legal moves, then Nd3
Nf3 Nc6 3. Bb5 a6 4. Bxc6 dxc6 5. O-O Bg4 6. h3 Bh5 7. g4 Bg6 8. d3 Be7 9. Nbd2 Nf6 10. Nb3 O-O 11. Nc5 Bxc5 12. d4 Bd6 13. Nd3 Bxf3 14. Qxf
[adapted] repeated 8-grams 0% 16 legal moves, then Ba3
Nf3 d6 3. Bc4 Nf6 4. d3 Be7 5. O-O Nbd7 6. Be3 c6 7. Nbd2 O-O 8. c3 a6 9. Qc2 b5 10. Ba3 Nb6 11. Bxe7 Qxe7 12. Rac1 Bb7 13. Nf1 Rac8 14. Ng3
--- self-knowledge ---
prompt: '\nhow do you decide which experts to use?\n\n\n'
[raw] repeated 8-grams 2%
The directory is not a vector of 512, which is why the new chunk is not an expert. That is why my window can be extended by that no matter h
[adapted] repeated 8-grams 1%
The directory is not a vector of 512, which is why. There is not an expert involve
Can you write change_string? It should change the com
There is no special algorithm, the finding is that slowing down the LR or the trunk, while keeping the LR of the experts is enough to eliminate most of the forgetting in the network. You can see in that experiment where chess data was the only thing the model read for 524K characters, yet it kept almost the same performance (i.e. held-out loss) on all other domains. If you keep LR the same across the whole network the loss in other domains degrades dramatically - this is a clear sign of catastrophic forgetting in action. What I can say for sure is that any traditional network that does pose a sign of catastrophic forgetting would not be able to learn any patterns from a single stream of data.
There are no benchmarks published as the model is heavily undertrained, but it is learning. And you can see this clearly in the loss and samples even though they are still barely coherent.
I am not an academic and am not trying to publish a paper about a “major breakthrough” or something like this. I am just a small person who found a cool thing that clearly works and wants to share it with the world. That’s it.
It is a goalpost that is easy to move. By "works" I mean learning from a continuous single (meaning batch-1) stream of data. The fact that it produces full words and full coherent phrases instead of a random stream of characters that would any typical LM produce if trained under the same training regime.
I would be okay if you shared it as a potential idea and possibly interesting early result, but the language you are actually using to characterize it is misleading or delusional.
Please get a model to the point where it seems like it has some natural language understanding and then share again with reasonable characterization.
For sure. As it will pass through the whole corpus I will share the weights, run it through established benchmarks for small models and share all of this as an update. I am also planning on making a Youtube video explaining in detail how it works on a deeper level and the whole reasoning behind why it is built the way it is. But no promises here.
Nothing to be ashamed of if you end up pushing back the release date of your feature film :)
I had ideas not completely unlike this so long ago, but one big difference can be summed up in one of your parameters.
>Directories are walked, binaries are skipped . . . and each file is read from its beginning to its end because a document has an order.
For me it was binaries being walked because text and anything approaching a language model was so much further out-of-reach having such limited computer power.
Is "AGI" the language that bothers you? Well, one has to keep his eyes on the prize and I see a bright idea which could lead to AGI, so, why not describe it as such? I also see the inspiration and hard work necessary to move that idea further along, so fingers crossed.
It interleaves random streams of 32K characters long each when reading the whole corpus, but each such stream reads continuously as you would expect. This is a necessary step to prevent just normal, not catastrophic, forgetting. I have not tested it in any other regimes yet with bigger or smaller windows. You can imagine a person that changes the activity from time to time, so I think it is justified. So there is not really "early in the stream".
What I did test though is reading 524K characters of chess data only and see how other domains have degraded. The results are in the readme under "How continual learning works" section. Spoiler: it just barely degraded the performance.
The model is way too small and undertrained to make any generalization claims. I want to wait until it reads the whole corpus I gave and then test it on some simple established benchmarks to see how it will behave.
Seems a bit premature to make an HN post about then, imho.
It's an interesting idea, but it doesn't really do anything interesting yet. I looked at the output in the training run and it is a far, far cry from intelligence. Worse than GPT-2 as it stands.
I do hope it will perform well when scaled and trained, though; best of luck.
I also like how it is very organic. It naturally grows and deletes unused elements, so in addition to traditional backprop there is also a natural selection happening in the background. Each new expert has 16 parents by the way, lol.
reply