Illustration for: Clockwork.io Raises $31M To Stop Wasted GPU-Hours

Clockwork.io Raises $31M To Stop Wasted GPU-Hours

Clockwork.io raised $31 million to expand its GPU fault-tolerance software after LinkedIn, Together AI and WhiteFiber adopted it to stop losing GPU-hours to hardware failures.

TC
Early-stage VC & angel · Founder, New York Venture Partners · Value Add Pulse Funding Desk
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

GPU idle-time from hardware failures is an invisible tax on every AI lab's compute spend -- Clockwork.io's pitch is recovering that waste without the customer having to rewrite training code, a low-friction sell that explains fast enterprise adoption.

2

LinkedIn, Together AI and WhiteFiber all adopting the same resilience layer suggests infrastructure reliability is becoming a shared utility across large enterprise AI fleets, neoclouds and GPU-as-a-service providers alike, not a one-off internal tool each builds itself.

3

Together AI planning to resell TorchPass as a service shows how fast infrastructure startups can become embedded inside larger GPU-cloud platforms rather than selling directly to end customers -- a distribution path worth tracking for infra-layer founders.

4

At $73M in total funding against a single infrastructure-reliability niche, Clockwork.io's backers are betting GPU failure-recovery becomes as standard a line item as networking or storage -- a bet on category durability, not a one-time trend.

TC

The VC Read · Trace's Take

Trace Cohen

LinkedIn, Together AI and WhiteFiber adopting the same tool in one cycle is the diligence signal, not the $31M check size. Worth asking Clockwork.io directly what percentage of GPU-hours saved translates into renewed contracts versus one-time pilot deployments -- infra tools die fast when the savings don't show up in a renewal conversation.

Analysis

Clockwork.io has raised $31 million to expand its fault-tolerance software for AI infrastructure, led by Premji Invest, Wing Venture Capital and Seligman Ventures, with participation from NEA and e& Capital, bringing its total funding to $73 million, according to The SaaS News and Converge Digest.

What the software actually does

The Palo Alto-based company, founded in 2021, builds two core products: LinkPass, which reroutes network traffic around failed links, optics, cables or NICs, and TorchPass, which migrates work from a failing GPU to healthy capacity without forcing a full restart from an earlier checkpoint. Clockwork.io is extending TorchPass with multi-node platform snapshots that capture an entire distributed job's execution state without requiring changes to training code, plus asynchronous checkpoints that run while a workload keeps training.

“At the scale frontier labs now train at, that idle time compounds into real money.”

The problem Clockwork.io is solving is a real and expensive one: a single failed GPU, NIC, optical link or switch port in a large distributed training run can force recovery from an earlier checkpoint, leaving the rest of an otherwise-healthy GPU cluster idle while it waits. At the scale frontier labs now train at, that idle time compounds into real money.

Early customers signal fast adoption

LinkedIn has deployed LinkPass across its AI infrastructure fleet and reports the software prevents tens of thousands of GPU-hours of downtime every month. Together AI, a GPU-cloud provider that competes with names like CoreWeave and Crusoe, plans to offer TorchPass as a service directly to its own customers, while WhiteFiber is expanding Clockwork.io across its global GPU-as-a-service infrastructure. That combination -- a large enterprise AI fleet, a neocloud, and a GPU-as-a-service provider all adopting the same resilience layer in the same funding cycle -- is a stronger adoption signal than the dollar amount of the round itself.

What the headline misses

Clockwork.io has not disclosed revenue, a valuation, or customer count beyond the three named deployments, which makes it hard to size how much of the GPU-cloud market has actually standardized on its tools versus built in-house alternatives. Hyperscalers like Google and Amazon have their own internal failure-recovery systems for their own fleets, and nothing here suggests Clockwork.io has displaced those -- its traction so far is concentrated among neoclouds and GPU-as-a-service providers that lack hyperscaler-scale infrastructure teams.

Watch whether Clockwork.io discloses a hyperscaler-scale customer win next, which would be the real test of whether this becomes infrastructure-layer standard or stays a neocloud-tier tool.

ShareXLinkedInEmail

Key Sources

3 sources

Reported by The SaaS News · First reported by Converge Digest · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take, a few times a week. Free to subscribe, no spam.