Key takeaway
Clockwork.io raised $31 million (bringing total funding to $73 million) for software that keeps AI training and inference jobs running through hardware failures — and it announced production deployments at LinkedIn and Together AI, plus expanded adoption by WhiteFiber. LinkedIn says the software prevents tens of thousands of GPU-hours of downtime every month across its fleet.
PALO ALTO, Calif. — October 5, 2026
What happened
Clockwork.io announced a $31 million funding round co-led by Premji Invest, Wing Venture Capital, and Seligman Ventures, with participation from existing investors NEA and e& Capital. The round brings the company’s total funding to $73 million, which it will use to roll out its fault-tolerance suite across training, inference, and reinforcement learning, and to expand enterprise and cloud adoption. Alongside the raise, the company announced two new capabilities for its TorchPass product and named production customers including LinkedIn, Together AI, WhiteFiber, Wells Fargo, Nebius, and NScale.
The problem it’s solving
Large AI workloads can span thousands of GPUs that must stay in sync — one failed GPU, one dropped network link, or one frozen server can stall the entire job. Meta reported unexpected interruptions averaging roughly one every three hours during a 54-day stretch of Llama 3 training on 16,384 GPUs. The usual fix is reloading a checkpoint, but recovery can take up to 90 minutes, healthy GPUs sit idle waiting, and the job has to repeat work done since that checkpoint. Customers pay for idle GPUs and repeated computation; models take longer to finish.
What the software does
Clockwork.io’s tools deploy as a layer between the hardware and the workload. LinkPass reroutes traffic around a failed network link so the job never notices the fault. TorchPass migrates work off a failing GPU onto a healthy one so training continues instead of rolling back. The two new capabilities announced today: multi-node platform snapshots, an industry first that saves an entire running distributed job across every node with no changes to the training code, and fast, asynchronous application checkpoints taken in the background while the job runs, which also speed up reinforcement learning by getting updated weights to inference replicas sooner.
Why it matters
This is the unglamorous infrastructure layer the AI boom runs on — and it’s where the money leaks. A single training run on thousands of GPUs costs a fortune, and losing hours of it to routine failures is the kind of waste that compounds at scale. LinkedIn’s numbers are the strongest signal here: tens of thousands of GPU-hours saved monthly, from one InfiniBand flap no longer being able to knock an eight-GPU server out of service. SemiAnalysis’s independent benchmarking found TorchPass cuts training “goodput” loss from 14% to under 3% for a gold-rated neocloud. The broader point: as training jobs get bigger, fault tolerance stops being an optimization and becomes a requirement. Clockwork’s pitch is that failures should be treated as the normal state, not the exception. If its snapshot approach really works with no code changes, it could become standard plumbing for anyone operating a GPU fleet.
FAQ
How much did Clockwork.io raise, and from whom?
The $31 million round was co-led by Premji Invest, Wing Venture Capital, and Seligman Ventures, with NEA and e& Capital participating. Total funding is now $73 million.
What is the difference between TorchPass and LinkPass?
LinkPass handles network failures: it reroutes traffic around failed links so jobs keep running. TorchPass handles GPU failures: it migrates work from a failing GPU to a healthy one, and now can snapshot an entire distributed job with no code changes.
Which companies are using Clockwork.io in production?
LinkedIn has deployed LinkPass across its AI infrastructure fleet; Together AI is bringing TorchPass to market as a service on its GPU clusters; WhiteFiber is expanding its use across its GPU-as-a-service footprint. The company also names Wells Fargo, Nebius, NScale, and DCAI as customers.
What is “goodput”?
Goodput is the share of GPU-hours that actually move the model forward — the opposite of time lost to failures, idle waiting, and repeated computation. Clockwork.io’s CEO calls fault tolerance a “goodput multiplier.”
Sources: PR Newswire (Clockwork.io company announcement)

