IBM’s $240m Bet on Cheap, Open-Source AI Inference
IBM has committed $240 million to a multi-year deal with the startup Together AI to establish a large-scale AI inference cluster on its cloud platform. The move is aimed at enterprises seeking to reduce their artificial intelligence (AI) operational costs, challenging the dominance of hyperscalers in this domain.
The company believes that the value in AI is shifting from building sophisticated models to efficiently running them, particularly as enterprise demand for AI continues to grow. To achieve this, IBM will utilize Nvidia HGX B300 systems built on the Blackwell architecture, known for its optimization in inference tasks. This infrastructure will be further enhanced by Nvidia’s Spectrum-X Ethernet networking technology.
Together AI, valued at $8.3 billion as of July, offers a platform that enables companies to train and execute workloads using open-source models like DeepSeek, MiniMax, and Kimi. Their solution is marketed as a more cost-effective and flexible alternative to proprietary, closed systems. With around 400 trillion tokens processed monthly, Together AI represents a significant share of the growing inference traffic.
The focus on AI inference has sparked intense competition, with various players investing heavily in optimization and infrastructure. Nebius, for instance, acquired a team of 20 experts specializing in inference optimization for $643 million. This trend reflects the enterprise shift towards open-source models and the desire to control sensitive data within their own infrastructures.
IBM’s strategy leverages the appeal of cost-effective, open-source AI infrastructure, positioning its cloud as a competitive alternative to dominant players like Amazon, Microsoft, and Google. This move aligns with Europe’s efforts to establish sovereignty in AI inference, as demonstrated by TensorX, which secured €8 million to build such capacity on Nvidia Blackwell.
The race for AI inference dominance is on, and IBM aims to capitalize on the growing demand for efficient, controlled AI deployment.