The Fort Worth Press - Clockwork.io Launches The Industry's First Contractual Commitment to End GPU Waste in AI Training

USD -
AED 3.672497
AFN 65.000299
ALL 79.569641
AMD 363.409858
ANG 1.790365
AOA 918.000006
ARS 1507.786404
AUD 1.402003
AWG 1.80125
AZN 1.700088
BAM 1.695609
BBD 2.01466
BDT 123.160418
BGN 1.683441
BHD 0.377118
BIF 2991.820156
BMD 1
BND 1.273738
BOB 10.998357
BRL 5.144898
BSD 1.000308
BTN 95.98391
BWP 13.567066
BYN 3.037656
BYR 19600
BZD 2.011747
CAD 1.39366
CDF 2312.503214
CHF 0.818501
CLF 0.024154
CLP 953.749848
CNY 6.71145
CNH 6.70591
COP 3127.62
CRC 447.438348
CUC 1
CUP 26.5
CVE 95.595821
CZK 21.07295
DJF 178.127262
DKK 6.47764
DOP 59.048684
DZD 133.877982
EGP 52.198901
ERN 15
ETB 163.385076
EUR 0.86655
FJD 2.21295
FKP 0.741937
GBP 0.743145
GEL 2.597171
GGP 0.741937
GHS 11.488194
GIP 0.741937
GMD 73.503157
GNF 8795.314945
GTQ 7.636255
GYD 209.277425
HKD 7.844699
HNL 26.847168
HRK 6.5298
HTG 130.738787
HUF 315.586989
IDR 17687
ILS 3.027498
IMP 0.741937
INR 95.907349
IQD 1310.410896
IRR 1374599.999914
ISK 121.140073
JEP 0.741937
JMD 157.856864
JOD 0.709008
JPY 155.045019
KES 129.501832
KGS 87.449895
KHR 4050.422208
KMF 427.000011
KPW 900.000318
KRW 1367.304993
KWD 0.30854
KYD 0.833583
KZT 444.773881
LAK 22392.336201
LBP 89575.662472
LKR 331.544497
LRD 173.548166
LSL 16.285389
LTL 2.95274
LVL 0.60489
LYD 6.353115
MAD 9.432659
MDL 17.439779
MGA 4325.19127
MKD 53.344315
MMK 2099.62457
MNT 3595.075141
MOP 8.082447
MRU 40.051151
MUR 47.290362
MVR 15.389513
MWK 1734.543043
MXN 17.13729
MYR 4.044802
MZN 63.909789
NAD 16.285247
NGN 1326.92021
NIO 36.811146
NOK 9.352695
NPR 153.579928
NZD 1.735345
OMR 0.384499
PAB 1.000282
PEN 3.355918
PGK 4.522953
PHP 62.718023
PKR 277.25311
PLN 3.775165
PYG 5918.765443
QAR 3.636395
RON 4.561303
RSD 101.694019
RUB 84.330137
RWF 1475.913668
SAR 3.753075
SBD 8.03625
SCR 13.676962
SDG 601.498266
SEK 9.76976
SGD 1.27315
SHP 0.742225
SLE 24.639846
SLL 20969.491881
SOS 571.673797
SRD 37.752501
STD 20697.981008
STN 21.240534
SVC 8.752834
SYP 13002.000254
SZL 16.283395
THB 33.229499
TJS 9.226022
TMT 3.51
TND 2.928519
TOP 2.40776
TRY 48.657935
TTD 6.782056
TWD 31.7655
TZS 2645.002972
UAH 44.609389
UGX 3916.077853
UYU 40.244484
UZS 11793.315705
VES 841.184029
VND 25998.5
VUV 118.157011
WST 2.736734
XAF 568.419469
XAG 0.01544
XAU 0.00023
XCD 2.70255
XCG 1.802758
XDR 0.707052
XOF 568.419469
XPF 103.393717
YER 236.497564
ZAR 16.25115
ZMK 9001.19823
ZMW 19.630588
ZWL 321.999592
SSP 5655.283496
MXV 1.943135
  • BCC

    0.2600

    76.19

    +0.34%

  • JRI

    0.0400

    11.66

    +0.34%

  • CMSC

    0.1550

    20.475

    +0.76%

  • RIO

    -0.0800

    97.18

    -0.08%

  • GSK

    0.2350

    50.245

    +0.47%

  • AZN

    1.4250

    163.275

    +0.87%

  • BCE

    -0.3750

    22.525

    -1.66%

  • NGG

    1.4200

    76.35

    +1.86%

  • BP

    -1.1500

    45.81

    -2.51%

  • RYCEF

    0.2600

    19.3

    +1.35%

  • CMSD

    0.1900

    20.26

    +0.94%

  • RBGPF

    0.0000

    69.99

    0%

  • VOD

    -0.1300

    17.55

    -0.74%

  • RELX

    0.1550

    34.375

    +0.45%

  • BTI

    -0.1600

    56.36

    -0.28%

Clockwork.io Launches The Industry's First Contractual Commitment to End GPU Waste in AI Training
Clockwork.io Launches The Industry's First Contractual Commitment to End GPU Waste in AI Training

Clockwork.io Launches The Industry's First Contractual Commitment to End GPU Waste in AI Training

"You Only Compute Once" (YOCO) guarantees to resolve 90% of AI training failures with no lost progress, or customers get credit

Text size:

PALO ALTO, CA / ACCESS Newswire / July 1, 2026 / Clockwork.io, pioneer of Software-Driven AI Fabrics™ and the company behind TorchPass AI fault tolerance, today announced the YOCO Guarantee - the industry's first contractual commitment to dramatically reduce the hidden, compounding cost of training failure in large-scale AI infrastructure. The announcement marks a turning point in how the AI industry measures infrastructure reliability - moving beyond uptime metrics designed for a previous era towards goals AI teams value most: whether the job finishes on time, without losing work.

Under the YOCO (You Only Compute Once) Guarantee, Clockwork.io commits that at least 90% of training failures on supported TorchPass workloads will be resolved through live GPU migration, with no lost training progress, no checkpoint rollback, and no recompute. If Clockwork.io falls short of that commitment in any contract year, customers receive a 25% credit against their next TorchPass renewal or expansion.

"We built TorchPass to make training failure irrelevant," said Suresh Vasudevan, CEO of Clockwork.io. "The YOCO Guarantee is a line in the contract. We're putting skin in the game because we know TorchPass delivers, and we want our customers to know it too."

The Hidden Tax on AI Progress

Every AI organization training at scale faces the same brutal math: GPU clusters fail constantly, and every failure triggers an expensive restart cycle. According to research published by Meta FAIR at HPCA 2025, a 1,024-GPU cluster experiences a mean time to failure of just 7.9 hours - and at 16,384 GPUs, that drops to 1.8 hours. Each failure forces teams to provision replacement nodes, restore from the last checkpoint, and recompute every training step since that checkpoint was taken. That recomputed work costs full GPU dollars - compute you already paid for, run again from scratch. The cycle typically costs three or more hours of progress per failure event, with losses accumulating daily.

The consequence is that current GPU clusters effectively operate at 30-50% of their theoretical performance - not because the hardware is slow, but because the reliability framework governing it was never designed for workloads of this nature, duration, or scale.

"AI teams need their models to be done, not their nodes to be up. The industry has been measuring node uptime and calling it reliability. YOCO holds us accountable for the only thing that matters - your model, done," said Vasudevan.

The financial toll is severe. In a typical 2,048-GPU H200 deployment, failure-driven restarts drain over $6 million per year in wasted compute - hundreds of thousands of GPU-hours lost to cascading retries, idle recovery time, and recomputed training steps. For AI builders, the real unit of value is not GPU uptime but time to trained model - yet the infrastructure contract they've been buying guarantees node availability, not job continuity. For AI operators, the gap is equally costly: when a customer's training job fails, restarts, and loses days of progress, the experience is one of unreliability - regardless of what the SLA technically said.

"Recompute and restart is the hidden tax of large-scale training," said Vasudevan. "Most teams treat it as a fact of life. It isn't."

The YOCO Guarantee changes that contract.

TorchPass: Reliability Redefined in Software

Clockwork.io's answer is to make reliability a software-defined property rather than a function of hardware uptime - a fundamental architectural rethink that decouples job continuity from the failure rate of any individual component.

TorchPass addresses failure at its root through live GPU migration - when a fault occurs, TorchPass transfers the training job's full in-memory state, including model weights, gradients, and optimizer state, to a healthy spare node. Training continues from exactly where it stopped, typically completing recovery in approximately three minutes. No checkpoint restore. No recompute. No lost progress.

TorchPass handles three classes of failure: unplanned migration for sudden, catastrophic faults - kernel crashes, power failures, GPU failures - where state is reconstructed from healthy replicas; pre-emptive migration triggered by early warning signals like rising ECC error rates or thermal thresholds, enabling a controlled handoff before failure occurs; and planned migration for proactive maintenance, security patching, and firmware updates, allowing infrastructure hygiene without interrupting training. Across all three scenarios, the job never stops.

This approach reduces wasted training progress by 90%, cutting lost time from approximately three hours per day to under ten minutes in a 1,024-GPU cluster - meaning research teams no longer discover hours of progress silently erased, and model release timelines become predictable rather than probabilistic.

In independent testing conducted by SemiAnalysis, a leading AI infrastructure research firm, TorchPass outperformed every competing fault-tolerance framework - the only solution that "maintains the same training performance as jobs without fault tolerance."

TorchPass is 100% software-based, runs in cloud and on-premises environments, and supports popular training frameworks including TorchTitan, Megatron-LM, and DeepSpeed, on schedulers including Kubernetes and Slurm. It works across NVIDIA and AMD hardware, and across InfiniBand, RoCE, and Ethernet fabrics - with no hardware lock-in of any kind.

Why the Guarantee Changes the Market

For AI builders, it redefines the SLA they should demand. The question is no longer "what is your node uptime?" but "what percentage of my training failures will be resolved without losing progress?" - a metric tied directly to GPU ROI, not an availability percentage that has historically had little relationship to whether models get trained on time. The YOCO Guarantee makes that question answerable and auditable.

For AI operators, it raises the competitive bar. AI Cloud operators and infrastructure providers who can offer job-level continuity guarantees - backed by contractual credits - will command premium pricing, win customers burned by restart-driven losses, and protect their margins by dramatically reducing their GPU idle time. Those who cannot will find themselves competing only on raw GPU price in a commoditizing market.

And for the industry as a whole, it establishes a new accountability standard. The AI infrastructure market has long accepted vendor claims about fault tolerance at face value, with no contractual obligation behind them. The YOCO Guarantee - measurable and contractually backed - introduces a standard the market will increasingly expect others to match or explain why they cannot.

"There's a big difference between a vendor making a slide that says their product works and them writing it into a contract," said Jordan Nanos, Member of Technical Staff and lead author of ClusterMAX at SemiAnalysis. "In our testing, TorchPass delivered the fastest and most efficient fault-tolerant performance for a GPT-OSS-120B training run on a 64x H200 cluster when compared to checkpoint-restart on job completion time. TorchPass also outperformed TorchFT (in terms of MFU and tokens/sec/GPU) for this job, while matching its recovery time. The YOCO Guarantee just reflects what we saw in testing, and makes it contractual."

"Every enterprise running large-scale AI training knows the cost of a failed job: hours of progress lost, recomputes billed, model timelines slipping. Every product decision we make at Scaleway comes back to one question: are we making our customers' outcomes more predictable? Node uptime answers a different question entirely. The YOCO Guarantee is the first infrastructure commitment we've seen built around the right metric - whether progress is protected and the jobs keep running to completion, not whether the hardware stays up. That's the accountability model the AI infrastructure market has been missing," said Fred Bardolle, Head of Products and AI at Scaleway.

Availability

The YOCO Guarantee is available to new and renewing TorchPass customers effective August 3, 2026. Existing TorchPass customers should contact their Clockwork.io account team to discuss adding the guarantee to their current agreement. To learn more or get started, visit clockwork.io/yoco.

Clockwork.io will be at RAISE Summit in Paris, France, July 8-9, Booth #27A. Suresh Vasudevan, CEO of Clockwork.io, will also take part in the panel "Infrastructure as Destiny: The Compute-Capital-Cloud Trinity" on July 8th at 10:40 a.m. local time on the Main Stage.

About Clockwork.io

Clockwork.io pioneers Software-Driven AI Fabrics™ - a programmable layer between hardware and workload that delivers nanosecond-accurate telemetry, AI fault tolerance, and performance optimization across any accelerator, network, or deployment model. Modern AI workloads need the whole cluster to act as one machine, but failures and infrastructure bottlenecks severely compromise efficiency. Clockwork.io's FleetIQ platform recovers that lost capacity, letting enterprises train, deploy, and serve the world's most demanding AI workloads faster, more reliably, and at lower cost - across any Ethernet, RoCE, or InfiniBand fabric, without hardware lock-in. TorchPass, Clockwork.io's AI fault tolerance product, is independently benchmarked by SemiAnalysis as the only solution that maintains full training throughput during failures, outperforming checkpoint-restart and leading open-source frameworks. Uber, Wells Fargo, DCAI, Nebius, NScale, and White Fiber trust Clockwork.io to power their AI infrastructure. Learn more at www.clockwork.io

© 2026 Clockwork Systems Inc. TorchPass and YOCO Guarantee are trademarks of Clockwork Systems Inc. All other trademarks are the property of their respective owners.

Media Contact

Dana Trismen
[email protected]
650-269-7478

SOURCE: Clockwork



View the original press release on ACCESS Newswire

A.Maldonado--TFWP