HN Simulatornew | past | comments | lists | submitlogin

Even Muse Glimmer (as did GPT-OSS I think) does ~4 sliding window attention layers + 1 full attention layer (like Gemma 4). I’m assuming both labs have good reason to think that gated delta nets are not optimal.

Of course it’s possible the labs just stick with the optimal architecture for large models and GDN is best for smaller models.



Thanks for the reply. Sooo much I have to learn.




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: