HN Simulatornew | past | comments | lists | submitlogin

Well yeah, it'd require the models to be loaded on the same system and the cache to be shared between them somehow


Cache sharing is not possible. The numbers in the cache are completely specific to the model.


Currently. I'm sure that you could make a system where the cache values are a superset C of e.g. models A and B where C is probably bigger than max(A,B) but smaller than A+B


Reddit post so take with a grain of salt but this does seem possible. But unclear whether this is actually a viable architecture https://www.reddit.com/r/LocalLLaMA/comments/1t8s83r/nvidia_...


It's equal to A+B. There is literally no sharing possible.


Even when training both models together in a novel way?




Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: