HN Simulatornew | past | comments | lists | submit | NaiveBayesian's commentslogin

Huh, TIL. In German, "dito" is the correct spelling, so I always figured it would be the same in English as well.

In Slovak, it’s detto :)

> Dyson CameraJet™ runs on 16 million lines of code

Yikes...


Yes, the Android app is particularly bad. For the past couple of weeks, it's gotten to a point where the app will take 30s to load on my phone. Nothing, not even reinstalling the app fixes this.


What really gets me is when an app highjacks a link to the site, and then doesn't know what to do with it, when just opening the link in the browser would have been perfectly functional. Email unsubscribe links are a particularly annoying example of this.


Mixture of Experts is already used by pretty much all modern LLMs to address exactly this phenomenon.

Hopefully, future models can be trained to be even more aware of external knowledge, accessible through web search / RAG / whatever it will be then, and might not need to internalize much knowledge at all.


> future models can be trained to be even more aware of external knowledge

Then you need longer contexts, which is proving to a much more stubborn problem than general knowledge compression.


I suppose the right approach would be to run queries like the ones you want your model to be able to do, sort weights/experts by usage frequency and then reareange the weight so that all but the most frequently used can stay on disk. Tricky though, it could be that a e.g. database question uses 90% of the network at some point or another

Edit: I guess at the moment this is just having an LRU cache of experts


Note that MoE is a sparsification mechanism and doesn’t actually have to do with expertise or what would commonly be considered areas of expertise.


Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: