The Fort Worth Press - Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training

USD -
AED 3.672502
AFN 65.000288
ALL 79.92783
AMD 363.200558
ANG 1.790365
AOA 916.99996
ARS 1512.724296
AUD 1.405896
AWG 1.8
AZN 1.701759
BAM 1.70475
BBD 2.015105
BDT 122.866121
BGN 1.683441
BHD 0.377158
BIF 3009.975725
BMD 1
BND 1.276579
BOB 10.179685
BRL 5.132305
BSD 1.000449
BTN 95.931629
BWP 13.567684
BYN 3.041441
BYR 19600
BZD 2.012246
CAD 1.39962
CDF 2310.0002
CHF 0.82448
CLF 0.024171
CLP 954.419682
CNY 6.706402
CNH 6.70556
COP 3134.56
CRC 447.578413
CUC 1
CUP 26.5
CVE 96.109511
CZK 21.1723
DJF 178.159742
DKK 6.51085
DOP 59.260873
DZD 133.876605
EGP 52.161402
ERN 15
ETB 163.423038
EUR 0.87102
FJD 2.24075
FKP 0.743677
GBP 0.747565
GEL 2.604997
GGP 0.743677
GHS 11.515883
GIP 0.743677
GMD 73.999749
GNF 8795.70108
GTQ 7.633807
GYD 209.2854
HKD 7.845245
HNL 26.85343
HRK 6.5669
HTG 130.759174
HUF 316.173502
IDR 17733
ILS 3.03294
IMP 0.743677
INR 95.77355
IQD 1310.636666
IRR 1374575.000327
ISK 121.769736
JEP 0.743677
JMD 157.885791
JOD 0.70903
JPY 155.678496
KES 129.590798
KGS 87.450162
KHR 4074.837551
KMF 428.000211
KPW 900.000318
KRW 1382.815037
KWD 0.30861
KYD 0.833795
KZT 445.677328
LAK 22408.751171
LBP 89594.128747
LKR 331.841379
LRD 173.581452
LSL 16.298676
LTL 2.95274
LVL 0.60489
LYD 6.355114
MAD 9.52989
MDL 17.564075
MGA 4345.078556
MKD 53.623182
MMK 2099.664815
MNT 3596.851201
MOP 8.084808
MRU 40.080539
MUR 47.659514
MVR 15.409852
MWK 1734.87843
MXN 17.19097
MYR 4.0986
MZN 63.903695
NAD 16.298676
NGN 1330.980258
NIO 36.817998
NOK 9.42809
NPR 153.487582
NZD 1.74142
OMR 0.384505
PAB 1.000458
PEN 3.377685
PGK 4.451379
PHP 62.716015
PKR 277.316313
PLN 3.79595
PYG 5918.423381
QAR 3.647021
RON 4.583802
RSD 102.225971
RUB 84.725008
RWF 1476.267885
SAR 3.713538
SBD 8.000512
SCR 13.793554
SDG 601.483424
SEK 9.816345
SGD 1.27541
SHP 0.746965
SLE 24.639878
SLL 20969.491881
SOS 571.784692
SRD 37.752499
STD 20697.981008
STN 21.354764
SVC 8.754581
SYP 13002.000254
SZL 16.292377
THB 33.304993
TJS 9.229529
TMT 3.5
TND 2.943174
TOP 2.40776
TRY 48.67542
TTD 6.792558
TWD 31.893029
TZS 2653.280206
UAH 44.708507
UGX 3931.959922
UYU 40.225402
UZS 11781.061917
VES 845.45495
VND 26012
VUV 118.307083
WST 2.742913
XAF 571.351358
XAG 0.015409
XAU 0.00023
XCD 2.70255
XCG 1.803158
XDR 0.707052
XOF 571.351358
XPF 103.951572
YER 236.650164
ZAR 16.254503
ZMK 9001.190528
ZMW 19.635051
ZWL 321.999592
SSP 5658.175743
MXV 1.949099
  • RBGPF

    0.0000

    69.99

    0%

  • NGG

    1.1400

    77.17

    +1.48%

  • GSK

    0.4800

    50.77

    +0.95%

  • CMSC

    0.2400

    20.72

    +1.16%

  • RIO

    2.1000

    97.89

    +2.15%

  • RYCEF

    0.4200

    20

    +2.1%

  • AZN

    2.3300

    165.22

    +1.41%

  • RELX

    0.0600

    34.33

    +0.17%

  • BCE

    0.1250

    22.625

    +0.55%

  • VOD

    0.0450

    17.505

    +0.26%

  • BTI

    -0.3300

    55.76

    -0.59%

  • JRI

    0.0650

    11.565

    +0.56%

  • BP

    -0.1550

    45.225

    -0.34%

  • BCC

    -0.3800

    74.93

    -0.51%

  • CMSD

    0.1700

    20.46

    +0.83%

Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training
Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training

Clockwork.io Introduces A New Class of Fault Tolerance to End Failure-Driven GPU Waste in AI Training

New TorchPass solution addresses a multi-million dollar challenge with AI infrastructure; uses Live GPU Migration to keep large-scale AI training running through hardware failures instead of forcing costly restarts

Text size:

PALO ALTO, CA / ACCESS Newswire / March 11, 2026 / Clockwork.io, the leader in Software-Driven AI Fabrics- a programmable, vendor-neutral software layer that optimizes large-scale GPU clusters for real-time observability, fault tolerance, and deterministic performance-today announced the general availability of TorchPass Workload Fault Tolerance. This new class of software-driven fault-tolerance eliminates one of the most costly failure modes in large-scale AI training: catastrophic job restarts caused by infrastructure faults.

Delivered as a core capability of the Clockwork.io FleetIQ platform, TorchPass applies the principles of Software-Driven AI Fabrics to distributed training, using Live GPU Migration to allow workloads to continue running through GPU failures, network disruptions, driver bugs, and even full node crashes-without checkpoint restarts or lost progress.

"Companies are investing billions in next-gen chips, yet the costs of running distributed AI jobs remains grossly inflated because the ecosystem has accepted failure as a constant," said Suresh Vasudevan, CEO of Clockwork.io. "We built TorchPass to fundamentally reject that premise. Instead of treating failure as inevitable and restarting after the fact, TorchPass makes infrastructure faults invisible to the workload-training continues through failures transparently, in software. For a typical 2,048-GPU deployment, that translates into over $6 million a year in recovered compute. This is what our Software-Driven AI Fabric approach was designed to deliver: fault-tolerant AI infrastructure."

Dylan Patel, Founder and CEO of SemiAnalysis agreed that large-scale training jobs are limited by interruptions.

"As Blackwell clusters roll out with an NVL72 domain, and we look to the future with Rubin Ultra's NVL576 domain, the idea that a single GPU error or network link flap can take down an entire run is totally unacceptable," said Patel. "TorchPass solves a huge challenge with cluster reliability: it provides transparent failover and live workload migration that keeps MFU high, which in turn drives better GPU economics."

Why AI Training Fails at Scale

Distributed AI training remains one of the most failure-prone workloads in modern infrastructure. As cluster sizes grow, fragility increases sharply. Research from Meta FAIR shows that mean time to failure drops to 7.9 hours in a 1,024-GPU cluster and to just 1.8 hours at 16,384 GPUs. This means that for most large, AI-focused enterprises or AI clouds, failure-driven restarts are completely inevitable - making this a major barrier to scaling AI's impact.

Each failure forces training jobs to roll back to the most recent checkpoint, discarding minutes or hours of completed work and wasting additional time on manual intervention, reprovisioning resources and restarting training. These restarts silently cap GPU utilization, making reliability one of the largest hidden costs in AI infrastructure.

TorchPass addresses this problem by proactively addressing costly AI workload failures, solving them before the job stops or needs to restart. Vital for enterprises running large AI workloads and AI clouds alike, TorchPass dramatically improves the reliability of workloads and cluster utilization. For AI clouds, who can now address impacted GPUs while preserving the training run as planned, this translates into better customer SLAs and overall AI cloud economics, improving their ability to protect margin and deliver new models sooner.

"Managing compute output across large-scale GPU clusters is vital to ensuring we're delivering reliable capacity to our customers. By using TorchPass we have the support of a company that focuses on resilience like it is a core business function: it replaces any specific failing GPU and keeps the rest of the job moving, rather than making one small problem impact our large-scale operations," said David Power, CTO of Nscale. "In our evaluation, Live GPU Migration preserved both run continuity and throughput under real fault conditions, which is exactly what you need to deliver predictable time-to-train and a better customer experience at scale."

How Live GPU Migration Works: Reliability Without Restart

TorchPass performs transparent, in-flight migration of impacted training ranks to spare resources when failures occur. TorchPass typically completes recovery in approximately three minutes while the training process continues uninterrupted.

It supports resilience across three failure scenarios:

  • Unplanned migration, handling sudden events such as kernel crashes, power failures, or GPU faults by reconstructing state from healthy replicas

  • Pre-emptive migration, triggered by early warning signals such as rising temperatures or ECC memory errors, enabling controlled migration before a hard failure

  • Planned migration, enabling maintenance, patching, and workload rebalancing without interrupting training

This approach reduces wasted training progress by 95%, cutting lost time from approximately three hours per day to under ten minutes in a 1,024-GPU cluster.

Jordan Nanos, Member of Technical Staff and lead author of ClusterMAX-SemiAnalysis' independent benchmark for large-scale AI training-stress tested Clockwork.io TorchPass and found it delivered leading performance and efficiency for large-scale distributed training, enabling users to reduce checkpointing overhead in training. He shared the following results:

"In our testing, Clockwork.io TorchPass delivered the fastest and most efficient fault-tolerant performance for a gpt-oss-120B training run. We used TorchTitan on a Kubernetes cluster with 64x H200 GPUs. During our testing we measured job completion time (JCT) and Model FLOPs Utilization (MFU) against a standard approach (checkpoint-restart) and the leading open-source fault-tolerant training framework (TorchFT). We simulated multiple hardware failures on the cluster in order to stress test the fault-tolerant training frameworks.

When compared to checkpoint-restart, TorchPass was significantly faster to recover from failures. This reduced overall JCT and maintained high MFU. And when compared to TorchFT, TorchPass had a significantly higher MFU. This reduced overall JCT while also maintaining an equal time to recover from failures.

Using TorchPass also has a downstream effect where it provides users with an opportunity to reduce or even remove checkpointing from their training code. This means larger effective batch sizes, lower risk of out of memory errors (OOMs), and less time spent thinking about storage. For a research organization, this can ultimately mean a faster time to reach their training objective," concluded Nanos.

Measurable Business Impact from Software-Driven Fault-Tolerance

For customers operating large AI clusters, the impact is immediate and measurable. In a typical 2,048-GPU H200 deployment, TorchPass Workload Fault Tolerance delivers over $6 million in annual savings by preventing wasted compute.

These savings come from eliminating hundreds of thousands of GPU-hours that would otherwise be lost to failure-driven restarts, cascading retries, and idle recovery time. By keeping training jobs running through infrastructure faults instead of restarting them, TorchPass converts lost GPU time into productive training, significantly improving the return on GPU investments that today often operate at just 30-50% of theoretical performance.

Enabling the Next Generation of AI Infrastructure

By making reliability a software-defined capability rather than a hardware constraint, TorchPass provides the operational confidence required to deploy next-generation, tightly coupled systems such as NVIDIA GB200 and GB300 NVL72 and future rack-scale systems, where dense architectures amplify the cost of even small failures.

TorchPass builds on Clockwork.io's prior release of Network Fault Tolerance, which applies the same Software-Driven AI Fabric principles to network resilience by transparently rerouting traffic around link failures.

Together, these capabilities form Clockwork.io's Software-Driven AI Fabric, a vendor-neutral software layer spanning network, compute, and storage. As modern AI workloads run on tightly coupled clusters where hundreds or thousands of processors must operate in coordinated lockstep, infrastructure behaves as a single system, where reliability and performance directly determine overall efficiency. By managing this complexity in software, Clockwork.io enables operators to run heterogeneous AI infrastructure as a unified platform-maintaining high utilization, predictable performance, and resilience while preserving the flexibility to evolve hardware and improve the economics of large-scale AI deployments.

To learn more about the launch of TorchPass, visit the Clockwork.io team in-person at NVIDIA GTC from March 16-19, Booth #205, or visit https://clockwork.io.

About Clockwork.io
Clockwork.io pioneers Software-Driven AI Fabrics™, delivering a programmable software layer that makes large-scale AI clusters observable, deterministic, and resilient by design to drive continuous workload progress and peak cluster utilization. Its FleetIQ platform enables enterprises to train, deploy, and serve the world's most demanding AI workloads faster, more reliably, and at lower cost. Companies including Uber, Wells Fargo, DCAI, Nebius, Nscale, and White Fiber trust Clockwork.io to power their AI infrastructure. Learn more at www.clockwork.io.

Media Contact
Dana Trismen
[email protected]
650-269-7478

SOURCE: Clockwork



View the original press release on ACCESS Newswire

W.Lane--TFWP