AMD bought Talaas specifically to make AI accelerator pieces to be embedded into generalized chips.
At work (we're a medium sized manufacturing firm), we bought our own inference server for $107k and run Kimi 2.8 for nearly all of our use cases (and dropped our cloud AI spend to $0).
You dropped your cloud AI spend down to the price of capital plus the cost of electricity and maintenance on that server. When/if tokens become a commodity, the price of tokens would be the marginal cost, AKA about the same. Big when/if, though.
At work (we're a medium sized manufacturing firm), we bought our own inference server for $107k and run Kimi 2.8 for nearly all of our use cases (and dropped our cloud AI spend to $0).