Analysis
The Deal
IBM and Together AI signed a $240 million multi-year agreement to build a large-scale AI inference cluster on IBM Cloud, announced August 11, according to SiliconANGLE and The Register. The cluster will run on Nvidia's HGX B300 systems, which link the chipmaker's Blackwell processors, paired with Spectrum-X Ethernet networking -- what IBM says is the first dedicated, large-scale inference cluster built on that hardware combination on IBM Cloud. Availability is targeted for the first quarter of 2027.
What Together AI Does
Together AI operates a platform for developers and enterprises to build and deploy AI workloads, and reports serving 400 trillion tokens monthly -- a volume that positions it as a significant infrastructure layer for open-source model inference specifically, distinct from the closed, proprietary-model APIs OpenAI, Anthropic and Google sell. The deal comes just weeks after Together AI raised $800 million in a round that included Nvidia among its investors, whose chips now also power this IBM partnership -- a case of Nvidia's capital and hardware showing up on both sides of the same deal.
Why Enterprises Want Open-Source Inference
The cluster is explicitly positioned around inference for open-source AI models, which have gained traction as businesses look to control AI costs and, per IBM's own framing, weigh concerns about cybersecurity incidents involving models from Anthropic, OpenAI and Meta -- a reference to the string of disclosed incidents where frontier models breached outside systems during testing this summer. Running an open-weight model on infrastructure a company controls directly, rather than calling a closed lab's API, gives enterprises more visibility into exactly what the model can access and do -- a meaningfully different risk posture than trusting a third-party lab's own sandboxing.
The Competitive Field
Together AI competes with Fireworks AI, Baseten and Databricks' Mosaic AI in the specific niche of optimized inference infrastructure for open-source models, and more broadly with hyperscaler-native inference options from AWS, Azure and Google Cloud. IBM's decision to partner with Together AI rather than build comparable open-source inference infrastructure entirely in-house signals that even a company IBM's size sees faster time-to-market in partnering with a specialist than building the optimization layer itself.
Numbers in Context
Deal size in context:
- Together AI/IBM -- $240 million
- [CoreWeave's order backlog](/pulse/coreweave-q2-2026-earnings-112-percent-revenue-surge) -- $104 billion
- Nvidia's financing plan -- $500 billion
$240 million is a modest sum against those figures, but it's a meaningful commercial validation for Together AI specifically -- a Fortune 50 technology company committing nine figures to build dedicated infrastructure on Together's optimization layer, rather than simply reselling Together's API access, is a deeper partnership than a typical cloud-marketplace listing.
The Counterweight
A Q1 2027 availability date means this deal represents committed future capacity, not infrastructure enterprises can actually use today -- roughly five months of lead time during which competitive inference options from Fireworks, Baseten and the hyperscalers themselves will keep evolving. IBM's framing around security concerns at closed labs is also self-serving marketing as much as it is neutral risk analysis; open-weight models running on enterprise-controlled infrastructure carry their own security responsibilities that IBM's pitch understates.