This suggests inference margins in the ballpark of 98% if we assume 5.6 Sol is about as efficient to serve as Kimi K3.
We also do not know what efficiency improvements have been made with GPT 6 Sol and Luna.
There is some speculation that 6 Sol could be a smaller model comparable in size to 5.6 Terra, and that this is why the improvement in intelligence is modest over 5.6 Sol.
This would line up with a faster serving speed and benchmarks that show a small improvement in coding tasks with regressions in knowledge tasks.