AI News

Clockwork.io Raises $31 Million to Keep AI Workloads Running When Hardware Fails

AI’s Expensive Problem: Starting Over

Imagine spending hours training an AI model across hundreds of powerful chips. The work progresses nicely. Then a network connection fails.

Suddenly, healthy processors are waiting, engineers are investigating, and part of the completed work may need repeating. The hardware bill keeps ticking.

Clockwork.io wants to make those interruptions less disruptive.

On October 5, the company announced $31 million in new funding, alongside production deployments at LinkedIn and Together AI and expanded adoption by WhiteFiber. Its software aims to protect AI workloads against infrastructure failures.

The announcement also introduced new capabilities for preserving running jobs and saving progress more efficiently.

That makes this an AI story about the machinery behind the magic.

Models receive the applause. Infrastructure teams handle the awkward moment when the magic stops halfway through.

Clockwork’s pitch is that a broken component should not automatically erase useful progress across an expensive computing cluster.

Who Is Backing Clockwork’s Next Phase?

The funding round was co-led by Premji Invest, Wing Venture Capital, and Seligman Ventures, with participation from existing investors NEA and e& Capital.

It brings Clockwork.io’s total funding to $73 million.

The company says the capital will support wider deployment of its fault-tolerance software across training, inference, and reinforcement learning. It also plans to expand enterprise adoption and distribution through cloud partners.

Those three workload categories cover different activities.

Training develops a model. Inference runs it to produce outputs. Reinforcement learning involves a feedback-driven process that can combine training and inference.

Clockwork is therefore pursuing several related infrastructure needs.

The funding gives it resources to expand that effort. It does not establish how much revenue the company generates or what return customers will receive.

Those outcomes will depend on deployment, compatibility, operating costs, and how effectively the software protects useful work under real conditions.

Why One Failure Can Affect Many GPUs

A graphics processing unit, or GPU, performs calculations used extensively in AI.

Large workloads distribute those calculations across multiple GPUs. The processors must exchange information and coordinate their work.

That coordination creates dependencies.

Clockwork’s technical explanation describes how a stalled network connection can leave a distributed operation waiting, even when the GPUs themselves remain healthy.

Think of an orchestra waiting for one section to reach the next bar.

Everyone may have a working instrument. The performance still depends on coordination.

The comparison has limits, but it captures the central problem: individual components cannot always make useful progress independently.

Consequently, buying additional GPUs does not automatically solve the issue.

A cluster also needs dependable communication and a way to manage failures. Otherwise, a small fault can interrupt a much larger investment.

Clockwork targets that connection between component reliability and the progress of the overall job.

GPU-Hours Explain the Financial Stakes

A GPU-hour represents one GPU used for one hour.

The unit makes the scale of an interruption easier to understand.

Consider a hypothetical job running on 1,000 GPUs. If all of them spend 30 minutes waiting, that amounts to 500 GPU-hours without forward progress.

That is an illustration, not a measurement from a Clockwork customer.

It also excludes any work that must be repeated afterwards.

The financial impact depends on the price of the computing capacity and the surrounding operation. Delayed experiments and engineering time can matter, too.

This explains why infrastructure operators care about more than whether their machines are powered on.

A running server can still be contributing little to the task.

The practical goal is to convert more paid computing time into completed work. Recovery software becomes commercially interesting when it improves that conversion enough to justify its own cost and complexity.

LinkPass Gives Network Traffic Another Route

Clockwork’s LinkPass addresses network interruptions.

The company describes it as a network plugin that redirects traffic through healthy network interfaces when a connection fails. It is designed to let supported distributed workloads continue instead of restarting.

Its technical documentation also acknowledges that available bandwidth can fall during this process. The original path can be used again after it recovers.

That detail matters.

Continuing at reduced speed can still be preferable to stopping, restoring saved progress, and repeating calculations. But uninterrupted execution does not necessarily mean unchanged performance.

Picture a delivery vehicle encountering a closed road.

An alternative route may take longer. It still gets the shipment moving.

The infrastructure version requires careful coordination so communication remains correct through the transition.

The failed connection also needs repair. LinkPass’s proposed benefit is to reduce the disruption while operators address the underlying problem.

TorchPass Moves Work to Healthy Resources

Clockwork’s TorchPass adds another recovery approach: moving affected work onto replacement resources.

The company’s technical material describes support for planned maintenance, warning signs of impending failure, and certain sudden failures.

The mechanism depends on the situation.

Before a planned move, software can capture state from a working component. After a component has failed, that same opportunity may no longer exist.

Clockwork explains that some unplanned recovery paths reconstruct the required state from healthy workers in supported distributed configurations.

This is more specific than a universal promise that software can rescue anything.

The practical result depends on available replacement capacity, supported workloads, and the recovery path being used.

The company also describes brief pauses during migration.

Its aim is to preserve progress and resume useful computation with less disruption than a broad restart.

That distinction keeps the technology understandable: resilience manages failure; it does not make hardware incapable of failing.

Snapshots Add a Different Safety Net

Clockwork.io AI infrastructure resilience

The new TorchPass snapshots preserve the state of supported distributed training jobs.

Clockwork says platform teams can capture running process and GPU state without changing the training code, then restore the job when compatible capacity is available.

This offers a different option from immediate migration.

Migration needs somewhere suitable to move the work. A saved snapshot can preserve progress while replacement capacity becomes available later.

Consider a hypothetical maintenance operation.

An operator needs to release several machines, but no spare resources are immediately ready. Capturing the job could allow the machines to be serviced without requiring the workload to begin again from scratch.

The feature’s scope matters.

Clockwork’s documentation specifies compatibility requirements involving GPU configurations, drivers, kernels, and other execution conditions.

“Without training-code changes” describes an integration advantage for supported jobs. It does not mean every workload can be captured and restored in every environment.

Saving Progress Has Its Own Trade-Offs

Saving state is useful, but the process consumes resources and time.

Clockwork reports internal snapshot testing on an eight-node H200 configuration using Megatron-LM. In that test, the total pause was under 20 seconds.

That is a company-reported result for a specific setup.

The broader operational question is how frequently to save progress.

More frequent snapshots can reduce the amount of recent work exposed to an unexpected failure. They can also increase the cumulative overhead of saving.

Less frequent snapshots reduce that overhead but leave a longer gap between recovery points.

There is no single ideal interval for every job.

Operators need to consider workload size, storage performance, failure patterns, and recovery requirements.

Clockwork’s documentation also says application checkpoints remain valuable, including when moving to an incompatible runtime.

The new approach therefore adds another recovery tool. Its usefulness depends on choosing the right mechanism for the situation.

LinkedIn Provides a Concrete Adoption Signal

LinkedIn is one of the named production users.

In Clockwork’s announcement, LinkedIn infrastructure executive Raghu Hiremagalur says the company deployed LinkPass across its AI infrastructure fleet.

He attributes the prevention of tens of thousands of GPU-hours of downtime per month to the technology.

That figure comes from a customer statement included in company communications. It should not be treated as an independently audited result.

Nevertheless, the account describes a concrete operational use.

Network faults can become maintenance events while workloads continue using healthy paths. For an infrastructure team, that could change both the urgency and the consequences of an incident.

The important question extends beyond how quickly an alert appears.

Does the running job survive? How much performance changes? What work remains for engineers?

LinkedIn’s reported experience gives the announcement substance, while leaving room for further evidence about deployment conditions and measurement.

Together AI Brings Resilience to Cloud Customers

Together AI is bringing TorchPass to market as a service on its GPU Clusters, according to the announcement.

Its product lead describes fault tolerance as an additional layer beyond detecting problems and provisioning replacement capacity.

That distinction is useful.

Replacing a broken machine addresses infrastructure health. Preserving the work that depended on it addresses workload continuity.

A provider needs both.

For customers, integrating protection into a cloud service could make resilience easier to adopt. They would not necessarily have to assemble every recovery component themselves.

However, the announcement does not establish identical protection across all customer workloads.

Availability, integration requirements, and supported configurations remain relevant.

The commercial opportunity is to make reliability part of the service customers purchase.

A cluster’s value comes from what users accomplish on it. Faster repairs help, but keeping useful progress intact may be an additional reason to choose one provider over another.

WhiteFiber Focuses on Problems Before Production

WhiteFiber, an existing customer, is expanding its use of Clockwork software across its GPU service operations.

Its chief technology officer, Tom Sanfilippo, describes another application: validating clusters before customers begin production work.

He says Clockwork’s automated fleet audit helps identify faulty links, problematic network interfaces, and other infrastructure issues.

These are customer claims presented in company communications.

The preventative angle complements recovery.

A cluster can be powered on and still contain weaknesses that emerge under demanding workloads. Finding those problems earlier could reduce the chance that a customer discovers them through a failed job.

For providers, this creates two related objectives.

First, deliver capacity that has been tested meaningfully. Second, manage the failures that still occur afterwards.

Neither objective eliminates the other.

Reliable operations require preparation and recovery together. WhiteFiber’s account illustrates why a resilience company may have value before a model’s first training step even begins.

Inference and Reinforcement Learning Broaden the Need

The announcement also highlights inference and reinforcement learning.

When a model operates across multiple servers, network communication remains important during serving. An interruption can affect a running response or task.

Clockwork positions LinkPass as protection against relevant link failures in those supported configurations.

Reinforcement learning adds another connection.

Copies of a model generate examples, often called rollouts. Training uses those examples, and updated model weights must return to the generating systems.

Clockwork says its fast asynchronous application checkpoints can help move updated weights through that cycle sooner.

These are the company’s stated benefits, rather than universal performance results.

The analytical implication is that resilience can affect more than time spent recovering from a crash.

It can also influence how smoothly interconnected stages exchange work.

However, the size of any improvement depends on the pipeline. A system constrained somewhere else may gain less from faster recovery or weight transfer.

Useful Progress Is the Metric That Matters

Clockwork and its customers use the term goodput for computing time that actually moves a workload forward.

It is a helpful distinction.

A GPU can be busy repeating calculations lost during an interruption. That activity consumes capacity, but it does not produce the same net progress as successful new work.

Clockwork’s website reproduces a SemiAnalysis statement describing lower training goodput loss in a particular cluster scenario.

Such a result needs its configuration and assumptions attached. It should not become a promise of identical gains everywhere.

For prospective customers, a relevant evaluation would compare the software with their current recovery approach.

They could examine job completion time, overhead during healthy operation, behaviour under different faults, and the replacement capacity required.

The question is practical: does the full deployment produce more useful work?

That is stronger evidence than a recovery demonstration alone, however impressive the demonstration looks.

What This Funding Could Change

Clockwork.io AI infrastructure resilience

Clockwork’s announcement combines fresh funding, named production users, and additional recovery capabilities.

The promising idea is straightforward: preserve more of the work organisations already pay to perform.

If deployments succeed, customers could complete jobs with fewer disruptive restarts and less repeated computation. Platform teams could also gain more flexibility when maintaining or reallocating resources.

Those are potential benefits, not outcomes established for every cluster.

The next stage should provide more evidence about compatibility, overhead, recovery behaviour, and sustained customer results. Financial savings would need to include the cost of the software and its supporting infrastructure.

Still, the direction is meaningful.

AI progress depends on capable models and the systems that keep their work moving. Protecting that progress can make existing hardware more productive.

The glamorous announcement may belong to the next model.

The infrastructure victory is quieter: a component fails, engineers repair it, and the expensive job keeps making progress.

Sources