Skip to content

Clockwork Raises $31M To Keep AI Training Running

Premji Invest, Wing Venture Capital and Seligman Ventures co-led the round, with LinkedIn, Together AI and WhiteFiber running the software across the computers that train their AI models.

Clockwork Raises $31M To Keep AI Training Running
Image courtesy: Unsplash

Training a large AI model is a single job split across thousands of chips working together, and the whole job stops the moment one of those chips, or the cable connecting it, gives out. Clockwork.io, a Palo Alto company that sells software to keep those runs going when a part fails, has raised $31M.

The round takes the eight-year-old firm's total funding to $73M. Premji Invest, Wing Venture Capital and Seligman Ventures co-led it, and the existing backers NEA and e& Capital also took part, the company announced on 5 October.

What makes that expensive is the recovery rather than the breakage. The standard fix is to go back to the last saved copy of the work, called a checkpoint, and redo everything done since, which can take up to 90 minutes on hardware that is rented by the hour.

Clockwork's answer is to avoid the restart altogether. Its software either reroutes traffic around a broken network link or lifts the work off a failing chip and puts it on a healthy one, so the run carries on instead of going back to a save point. Three products do that work. LinkPass handles the network side, TorchPass moves a running job from a failing chip to a working one without the customer rewriting any code, and FleetLens watches the whole estate and reports which parts are degrading.

LinkedIn is the largest named deployment. Raghu Hiremagalur, the social network's infrastructure chief technology officer, said that before the software "one InfiniBand NIC flap could remove an eight-GPU server from service," meaning a single network card dropping its connection for a moment took eight chips offline. He said LinkPass now saves tens of thousands of chip-hours of downtime every month.

Together AI, which rents out clusters to companies training their own models, is putting TorchPass into that service, and WhiteFiber, a Nasdaq-listed provider of the same kind of capacity, is extending the software across its estate. Wells Fargo, Nebius, Nscale and DCAI are also listed as customers.

The company's chief executive, Suresh Vasudevan, summarised the pitch as treating failure as routine. "Failures are inevitable at AI scale," he said. "Losing hours of useful work to them should not be."

A Failure Every Three Hours

The reference point everyone in this field cites comes from Meta, and it is unusually detailed. During a 54-day run training Llama 3 on 16,384 Nvidia H100 processors, Meta recorded 419 unexpected interruptions, an average of one every three hours.

The causes were overwhelmingly physical. Processors and NVLink accounted for 148 of them and HBM3 memory for another 72, which puts the two together at 52.5% of the total, while the central processors failed twice in the whole period.

Meta still achieved more than 90% effective training time, and only three of the 419 incidents needed significant manual intervention. The rest were absorbed by automation the company had built for the purpose.

That detail matters to the commercial case in both directions. It establishes the failure rate as real and constant at scale, and it also shows that the largest operators have already written their own tooling for it. The companies without that engineering capacity are the ones buying, because a GPU rental business or an enterprise running a few hundred processors has the same failure physics and none of Meta's infrastructure team.

The Benchmark Is The Company's Own

Clockwork published a comparison in March that sets out what its software does against the alternatives, and the figures are worth reading alongside their source. The test ran on 64 Nvidia H200 processors across eight nodes, training a 109B-parameter model over 3,000 steps with faults injected at random intervals.

TorchPass finished in 404.6 minutes with no steps lost and about 10 minutes of total downtime across six failures. A standard checkpoint restart took 817.5 minutes and had to recompute 869 steps, spending 307 minutes recovering, while Meta's own fault-tolerance library, TorchFT, lost no steps either but took 930 minutes, slowed by moving coordination traffic through the central processors rather than the network fabric.

Clockwork ran that benchmark itself and published it on its own site. The research firm SemiAnalysis has separately reported that the software cuts training goodput loss from 14% to under 3%, which is the figure the company leads with.

The scale gap is the part the numbers do not cover. Sixty-four processors is two orders of magnitude below the clusters where interruptions arrive every three hours, and the behaviour of a migration system at that size is not established by a test at this one.

Nvidia And Meta Give This Away

The competitive position is unusual for an infrastructure startup, because the two obvious alternatives cost nothing. Nvidia publishes a resiliency extension for Python that lets framework developers build fault tolerance into their own training code, aimed at the same goal of reducing downtime from failures, and Meta's TorchFT is open source and sits inside PyTorch, where most of this work happens.

Clockwork's argument against both is where it operates. The free tools work inside the training framework, reconfiguring the job after something breaks, while LinkPass sits in the network and TorchPass moves a live process between machines without the job noticing. Whether that distinction holds commercially depends on customers who have tried the free option first, and LinkedIn, Together AI and WhiteFiber all run large engineering teams, which makes their adoption the more useful signal in the announcement.

Dylan Patel of SemiAnalysis pointed at where the demand is heading. "Cluster fault tolerance used to be a training problem," he said. "It is now an inference problem too." That shift changes the economics, because a failed training step costs recomputation while a failed inference node costs a request a customer is waiting on, and inference now consumes a growing share of installed capacity.

Built For A Different Problem

Clockwork came out of Stanford in 2018 to solve something unrelated to artificial intelligence. Balaji Prabhakar, a professor of computer science there, founded it with Yilong Geng and Deepak Merugu, and Mendel Rosenblum, who co-founded VMware, joined as chief scientist, with the technology synchronising server clocks to within nanoseconds using software rather than dedicated hardware.

The first customers were in finance. Nasdaq, Wells Fargo and the Royal Bank of Canada were named when the company raised $21M in 2022 in a round led by NEA, with John Hennessy, Ram Shriram and Jerry Yang investing as individuals.

That earlier work is still underneath the current product. Measuring where a packet went and when it arrived, to nanosecond precision, is what allows FleetLens to tell a customer which link in a fabric is degrading before a job falls over. Vasudevan, who previously ran Sysdig and Nimble Storage, became chief executive as the company moved towards AI infrastructure, and Wells Fargo appears on both customer lists, seven years apart.

What the round establishes is that three operators of substantial GPU estates have put a startup's software into the path of their training jobs, which is not a casual purchase. What it leaves open is whether that holds as clusters grow and as the free alternatives, maintained by the companies that sell the hardware and the framework, continue to improve.

Add Morning Tick on Google