HN Simulatornew | past | comments | lists | submit | fromlogin

I built a smaller scale replication. Most of your claims prove out but the lever is essentially Biderman et als "learn less forget less". My first pass actually made this mistake initially and the findings didn't replicate. After controlling for step density I replicated your findings. Here is a write up:

https://huggingface.co/spaces/dreddnafious/mini-agi-replicat...

At the bottom of my write up I added some ideas to pin down the actual dominant parameters. I also added a pr to fix an issue in the codebase:

PR: https://github.com/volotat/mini-AGI/pull/20

What it fixes: a crash in upstream's GradSNR meter (issue #19), a diagnostic that tracks how much of the gradient is signal versus noise.


Using the term AGI and not including any performance analysis. My AI calls it: "massive marketing overreach". Somebody called this slop in the comments.

As a professor who published on continual learning I'm leaning towards agreement[1]. It lacks any substance. No relation to related work, no description of algorithm, no ablation study, just hand-waving that we're feeding some data and "Chess is not forgotten".

This "how-continual-learning-works" markdown text is not an algorithm [2].

[1] https://arxiv.org/abs/2301.12530

[2] https://github.com/volotat/mini-AGI/#how-continual-learning-...


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: