Why not just use an 8-bit LUT to encode the 256 most common ternary vectors with 6 components. That means of the possible 729 possible such vectors, you can only represent 256 different ones. You have to do more aggressive rounding, but at least the scheme is very simple to decompress and stream.
I think the weights are iid distributed, so all 729 patterns will be roughly equally likely. That doesn't make this a bad idea though -- it just means there's no point trying to select the most common 256 to keep, since any 256 will be roughly as good.
In the article it said that in ternary, the majority of weights are 0. The components may be independent, but any group of weights won't be evenly distributed across all probable occurences.
Good point, I was wrong. Groups of weights having more zeros will be more likely, so should be favoured. Due to independence there won't be any meaningful difference in frequencies between two groups of weights that have the same number of zeros, but that doesn't invalidate the above.
Why not just use an 8-bit LUT to encode the 256 most common ternary vectors with 6 components. That means of the possible 729 possible such vectors, you can only represent 256 different ones. You have to do more aggressive rounding, but at least the scheme is very simple to decompress and stream.