From Jacobian. They developed a tool called a Jacobian Lens, or J-Lens (itself a descendant of an earlier simpler tool called a Logit Lens) to examine what was going on inside an LLM, and named the space they found with it a J-Space.
Jacobians are essentially derivatives but for matrices.
But you can train a small LLM with a gaming graphics card -- I managed one on a GTX 1660. I don't think pg is suggesting that you try to chase the frontier. It's more like building your own OS in the 80s, or web server in the 90s -- sure, you'll never match the commercial offerings or the big OS projects, but building something from scratch within the limits of the hardware you can afford is amazing educationally.
In my experience, having a solid understanding of the next level down in the stack -- the foundation you're building your startup on -- is really helpful. We built a PaaS, and knowing enough about Linux internals to be able to work out what would be easy and what would be hard meant that we could focus our efforts on high bang-for-buck features.
So I'd say that understanding LLMs to the level that you get to by training your own baby one would be a solid foundation for pretty much anything built on top of the "real" ones.
I'd say (and I think this is actually compatible with the article itself, if not the title) that a better rule is to write about things that you've only just started understanding. The OP is, I think, correct that you might write a better introduction than an expert for whom it's all obvious if you've just been through the struggle of learning it, because the journey is fresh in your mind.
The risk of sounding more knowledgeable than you actually are is a real problem, though, and I'm glad he mentions it. It's really tricky to strike the right balance between making it clear that the topic is new to you, and hedging so much that you sound like ChatGPT on a bad day.
I must admit that I found it confusing that they lumped people with unmanaged hypertension in with those who were on meds to manage it. Those feel like separate groups. Does anyone here know more about this area, and why the combined group might make sense?
I'm adding MoE support to the GPT-2 code from Sebastian Raschka's "Build a Large Language Model (From Scratch)". Just the minimal changes. Planning to do (or at least start!) a from-scratch base model training run on my own hardware before the end of the month.
Jacobians are essentially derivatives but for matrices.
reply