✉ The Friday AI Brief: the week's 5 best AI stories, tools & comparisons — in your inbox every Friday morning.

Clockwork.io — $31M raised for GPU fault-tolerance software used by LinkedIn, Together AI, and WhiteFiber

Clockwork.io Raises $31M to Stop GPU Failures From Wasting AI Training Time

Key takeaway

Clockwork.io raised $31 million (bringing total funding to $73 million) for software that keeps AI training and inference jobs running through hardware failures — and it announced production deployments at LinkedIn and Together AI, plus expanded adoption by WhiteFiber. LinkedIn says the software prevents tens of thousands of GPU-hours of downtime every month across its fleet.

What happened

Clockwork.io announced a $31 million funding round co-led by Premji Invest, Wing Venture Capital, and Seligman Ventures, with participation from existing investors NEA and e& Capital. The round brings the company’s total funding to $73 million, which it will use to roll out its fault-tolerance suite across training, inference, and reinforcement learning, and to expand enterprise and cloud adoption. Alongside the raise, the company announced two new capabilities for its TorchPass product and named production customers including LinkedIn, Together AI, WhiteFiber, Wells Fargo, Nebius, and NScale.

The problem it’s solving

Large AI workloads can span thousands of GPUs that must stay in sync — one failed GPU, one dropped network link, or one frozen server can stall the entire job. Meta reported unexpected interruptions averaging roughly one every three hours during a 54-day stretch of Llama 3 training on 16,384 GPUs. The usual fix is reloading a checkpoint, but recovery can take up to 90 minutes, healthy GPUs sit idle waiting, and the job has to repeat work done since that checkpoint. Customers pay for idle GPUs and repeated computation; models take longer to finish.

What the software does

Clockwork.io’s tools deploy as a layer between the hardware and the workload. LinkPass reroutes traffic around a failed network link so the job never notices the fault. TorchPass migrates work off a failing GPU onto a healthy one so training continues instead of rolling back. The two new capabilities announced today: multi-node platform snapshots, an industry first that saves an entire running distributed job across every node with no changes to the training code, and fast, asynchronous application checkpoints taken in the background while the job runs, which also speed up reinforcement learning by getting updated weights to inference replicas sooner.

Why it matters

This is the unglamorous infrastructure layer the AI boom runs on — and it’s where the money leaks. A single training run on thousands of GPUs costs a fortune, and losing hours of it to routine failures is the kind of waste that compounds at scale. LinkedIn’s numbers are the strongest signal here: tens of thousands of GPU-hours saved monthly, from one InfiniBand flap no longer being able to knock an eight-GPU server out of service. SemiAnalysis’s independent benchmarking found TorchPass cuts training “goodput” loss from 14% to under 3% for a gold-rated neocloud. The broader point: as training jobs get bigger, fault tolerance stops being an optimization and becomes a requirement. Clockwork’s pitch is that failures should be treated as the normal state, not the exception. If its snapshot approach really works with no code changes, it could become standard plumbing for anyone operating a GPU fleet.

FAQ

How much did Clockwork.io raise, and from whom?

The $31 million round was co-led by Premji Invest, Wing Venture Capital, and Seligman Ventures, with NEA and e& Capital participating. Total funding is now $73 million.

What is the difference between TorchPass and LinkPass?

LinkPass handles network failures: it reroutes traffic around failed links so jobs keep running. TorchPass handles GPU failures: it migrates work from a failing GPU to a healthy one, and now can snapshot an entire distributed job with no code changes.

Which companies are using Clockwork.io in production?

LinkedIn has deployed LinkPass across its AI infrastructure fleet; Together AI is bringing TorchPass to market as a service on its GPU clusters; WhiteFiber is expanding its use across its GPU-as-a-service footprint. The company also names Wells Fargo, Nebius, NScale, and DCAI as customers.

What is “goodput”?

Goodput is the share of GPU-hours that actually move the model forward — the opposite of time lost to failures, idle waiting, and repeated computation. Clockwork.io’s CEO calls fault tolerance a “goodput multiplier.”

Sources: PR Newswire (Clockwork.io company announcement)

Leave a Comment

Your email address will not be published. Required fields are marked *

Get the 5 best AI tools every week

Top AI news, tools, and prompts — one short email. Free, unsubscribe anytime.

Scroll to Top