Clockwork.io, a provider of fault-tolerance software for AI infrastructure, has raised $31 million in a funding round co-led by Premji Invest, Wing Venture Capital, and Seligman Ventures. Existing investors NEA and e& Capital also participated, bringing the company's total capital raised to US$73 million. The company also announced new TorchPass capabilities, production deployments at LinkedIn and Together AI, and expanded adoption by WhiteFiber.
The Cost of Failure in Large AI Clusters
Large distributed AI workloads can span thousands of GPUs that must stay synchronized, and a single failed GPU, network link, or server can stall an entire job. Meta reported unexpected interruptions averaging roughly one every three hours during a 54-day Llama 3 training run on 16,384 GPUs. Conventional checkpoint-and-restart recovery can take up to 90 minutes, leaving healthy GPUs idle and forcing repeated work.
Production Adoption Across Enterprises and Neoclouds
LinkedIn has deployed Clockwork.io's LinkPass network fault tolerance across its AI infrastructure fleet, preventing tens of thousands of GPU-hours of downtime each month. Together AI is bringing TorchPass to market as a service on its GPU Clusters, with a live demonstration planned at the PyTorch Conference. WhiteFiber is expanding its use of the software across its global GPU-as-a-service footprint to validate new capacity before production.
New Capabilities for Training and Reinforcement Learning
Clockwork.io launched TorchSnap, a checkpointing technology designed to protect distributed AI workloads without code changes, and extended TorchPass with multi-node platform snapshots. Fast asynchronous application checkpoints operate in the background and accelerate reinforcement learning by delivering updated model weights to inference replicas sooner. These features allow platform teams to protect workloads they do not control and reduce the amount of progress lost after failures.
Executive and Investor Perspectives
CEO Suresh Vasudevan said failures are inevitable at AI scale, but losing hours of useful work to them should not be. He described fault tolerance as a goodput multiplier that keeps GPUs doing useful work instead of waiting for recovery or repeating completed computation. NEA Venture Partner Greg Papadopoulos added that Clockwork.io treats failure as the normal state and is defining the performance layer of the AI cluster.
A Broader AI Infrastructure Market
The opportunity extends beyond fault tolerance as global AI infrastructure investment accelerates and GPU clusters become larger and more expensive. Together AI raised US$800 million at a US$8.3 billion valuation, while Crusoe recently raised more than US$3 billion at a US$30 billion valuation in the same infrastructure race. Clockwork.io is attacking a different layer by ensuring that the enormous amount of compute being deployed actually gets used.
What the Funding Will Support
The new capital will help Clockwork.io accelerate the rollout of its fault-tolerance suite across AI training, inference, and reinforcement learning. The company also plans to expand enterprise adoption and scale delivery through cloud partners. With AI clusters becoming larger and harder to keep synchronized, Clockwork.io is betting that resilience will become as important as raw GPU performance.
Clockwork.io's US$31 million round underscores growing demand for infrastructure that keeps AI workloads running through inevitable hardware and network failures. Its production traction at LinkedIn, Together AI, and WhiteFiber demonstrates that large-scale GPU operators now view fault tolerance as a practical operating requirement. By combining live migration, network rerouting, and state capture, the company aims to make resilience a core performance layer for the next generation of AI infrastructure.