I was very excited to try the new GPT-6 Luna on our internal evals for a cybersecurity application. The 50% cost savings over 5.6-Luna sounded like an easy win.
Unfortunately in our benchmarks 6-Luna uses way more tokens to produce the same result at a given thinking effort, and takes longer to do so. The additional tokens wipe out the cost savings for us. Has anyone else seen this?
Unfortunately in our benchmarks 6-Luna uses way more tokens to produce the same result at a given thinking effort, and takes longer to do so. The additional tokens wipe out the cost savings for us. Has anyone else seen this?