The Fort Worth Press - TrustScale Launches ArgusRL as Automated AI Evaluation Surpasses Human Performance

USD -
AED 3.673026
AFN 63.000366
ALL 80.15939
AMD 363.440085
ANG 1.790365
AOA 916.99972
ARS 1515.990901
AUD 1.420031
AWG 1.8
AZN 1.683254
BAM 1.712962
BBD 2.013821
BDT 122.932477
BGN 1.683441
BHD 0.37695
BIF 3000
BMD 1
BND 1.277595
BOB 10.867156
BRL 5.162599
BSD 0.99986
BTN 95.678145
BWP 13.595481
BYN 3.027431
BYR 19600
BZD 2.01093
CAD 1.409435
CDF 2310.000279
CHF 0.824015
CLF 0.024404
CLP 963.620227
CNY 6.711349
CNH 6.711075
COP 3292.13
CRC 453.685538
CUC 1
CUP 26.5
CVE 97.250035
CZK 21.44175
DJF 177.720446
DKK 6.563599
DOP 59.450034
DZD 133.903899
EGP 51.420199
ERN 15
ETB 163.319699
EUR 0.87798
FJD 2.237196
FKP 0.750528
GBP 0.754775
GEL 2.59503
GGP 0.750528
GHS 11.585003
GIP 0.750528
GMD 74.00031
GNF 8750.000302
GTQ 7.63211
GYD 209.186734
HKD 7.843395
HNL 26.837396
HRK 6.6152
HTG 130.683635
HUF 320.695497
IDR 17896
ILS 3.03393
IMP 0.750528
INR 95.84645
IQD 1309.814191
IRR 1374575.000286
ISK 121.160089
JEP 0.750528
JMD 157.538176
JOD 0.70903
JPY 158.193974
KES 129.439858
KGS 87.449797
KHR 4054.823313
KMF 430.000184
KPW 900.000318
KRW 1366.610269
KWD 0.309199
KYD 0.833209
KZT 444.620586
LAK 22404.007848
LBP 89537.860574
LKR 328.955608
LRD 172.465735
LSL 16.310417
LTL 2.95274
LVL 0.60489
LYD 6.383128
MAD 9.57123
MDL 17.695844
MGA 4393.217489
MKD 53.88149
MMK 2099.577364
MNT 3597.923561
MOP 8.077068
MRU 40.034859
MUR 47.879632
MVR 15.460192
MWK 1733.73447
MXN 17.48236
MYR 4.079899
MZN 63.910346
NAD 16.310274
NGN 1323.820162
NIO 36.793911
NOK 9.49116
NPR 153.084711
NZD 1.761825
OMR 0.384507
PAB 0.99986
PEN 3.375476
PGK 4.454546
PHP 62.84201
PKR 277.084952
PLN 3.84022
PYG 5925.912951
QAR 3.645178
RON 4.630977
RSD 103.134956
RUB 84.772762
RWF 1475.675439
SAR 3.756768
SBD 8.026013
SCR 13.866488
SDG 601.498687
SEK 9.915065
SGD 1.279779
SHP 0.74758
SLE 24.650091
SLL 20969.491881
SOS 571.398543
SRD 37.822504
STD 20697.981008
STN 21.4581
SVC 8.748774
SYP 13002.000254
SZL 16.306476
THB 33.395008
TJS 9.253737
TMT 3.5
TND 2.95341
TOP 2.40776
TRY 48.829803
TTD 6.795649
TWD 31.791102
TZS 2649.997977
UAH 44.878367
UGX 3889.609025
UYU 40.078651
UZS 11798.503182
VES 851.341302
VND 26010.5
VUV 118.348377
WST 2.752527
XAF 575.916998
XAG 0.015477
XAU 0.000233
XCD 2.70255
XCG 1.801955
XDR 0.707052
XOF 575.916998
XPF 104.452317
YER 236.549962
ZAR 16.35505
ZMK 9001.19823
ZMW 19.472654
ZWL 321.999592
SSP 5712.590503
MXV 1.981384
  • RIO

    -0.3300

    97.04

    -0.34%

  • CMSC

    -0.0400

    20.67

    -0.19%

  • RELX

    0.0000

    33.41

    0%

  • RBGPF

    0.0000

    67.95

    0%

  • BCE

    -0.0700

    21.99

    -0.32%

  • BCC

    -0.0300

    75.66

    -0.04%

  • GSK

    0.8600

    51.08

    +1.68%

  • JRI

    -0.0300

    11.52

    -0.26%

  • NGG

    -0.1200

    76.68

    -0.16%

  • CMSD

    0.0900

    20.54

    +0.44%

  • BTI

    -0.0800

    55.75

    -0.14%

  • AZN

    2.0200

    168.1

    +1.2%

  • BP

    -1.4200

    43.16

    -3.29%

  • RYCEF

    0.4600

    19.7

    +2.34%

  • VOD

    0.0700

    17.02

    +0.41%

TrustScale Launches ArgusRL as Automated AI Evaluation Surpasses Human Performance
TrustScale Launches ArgusRL as Automated AI Evaluation Surpasses Human Performance

TrustScale Launches ArgusRL as Automated AI Evaluation Surpasses Human Performance

Customer calls it a "singularity moment" for reinforcement learning with human feedback, as evidence-grounded automation crosses the human-quality threshold at scale

Text size:

LOS ALTOS, CA / ACCESS Newswire / September 23, 2026 / TrustScale, an AI training, evaluation and assurance company, today announced the launch of ArgusRL, an automated evaluation and reinforcement feedback platform that outperformed a customer's highest qualified human evaluators in a production deployment, including identifying errors human reviewers missed. In testing, more than 95% of ArgusRL's automated evaluations were accepted by the customer without correction.

"One of our early customers, a leading global technology company, told us we've reached a singularity moment for reinforcement learning with human feedback, where evidence-grounded automation crosses the human-quality threshold at scale," said Lawrence Snapp, CEO of TrustScale. "This fundamentally changes the economics of AI training and reinforcement learning. AI makers and deployers no longer have to choose between the scale of automation and the quality of human evaluation. They can have both, grounded in deterministic evidence rather than another probabilistic AI opinion."

Human feedback has long been the gold standard for evaluating and improving AI models through reinforcement learning. As AI development accelerates, model makers are increasingly automating that process with AI judges and other model-based evaluation systems. ArgusRL takes a different approach, using empirical evidence and deterministic verification to evaluate AI-generated responses and generate structured feedback that can be used to continuously improve model performance.

In a production deployment with a leading global technology company, ArgusRL's automated evaluation delivered better results than the customer's human annotators and identified mistakes the human reviewers had missed.

"We've worked with a range of partners and approaches to improve the quality of reinforcement learning and model evaluation, and ArgusRL has consistently stood out for the quality and accuracy of its prompt and response review," Former Apple and Amazon AGI Leader "Its ability to identify errors missed during human review is particularly compelling, demonstrating the potential for deterministic automation to improve both the quality and scale of AI evaluation."

Unlike AI-Judge approaches that rely solely on probabilistic AI to evaluate another probabilistic system, TrustScale's patent-pending ArgusRL technology grounds its evaluations in retrieved external evidence. The platform analyzes each prompt and response, breaks responses into individual claims and searches multiple data sources for supporting or contradictory evidence. It then returns structured deterministic verdicts with citations and confidence scores.

ArgusRL also evaluates the quality of the original query and overall response and identifies cases that warrant human review. With ArgusRL, contradicted claims, claims without sufficient evidence and other flagged responses can be routed to human annotators, allowing people to focus on the cases where human judgment adds the greatest value rather than manually evaluating every response.

Because ArgusRL continuously evaluates outputs after deployment, its reinforcement feedback can incorporate current evidence and information that may not have been available during a model's initial training.

"The implications go well beyond accuracy," said Snapp. "AI companies spend billions of dollars each year on the data, human evaluation and infrastructure required to train and improve models. If AI makers can automate more of the reinforcement feedback process without sacrificing quality, they can improve models faster and at lower cost while reserving human expertise for the cases that actually require it."

ArgusRL is built on the same evidence-based TrustScale Engine that powers Argus, the company's AI assurance platform for detecting and correcting hallucinations at the point of use. ArgusRL takes that evidence-based approach upstream, giving AI makers and developers structured feedback they can incorporate into model training, fine-tuning pipeline, and continuous improvement.

ArgusRL operates as an API-backed evaluation service and supports multiple languages, locales and input formats. It can be integrated with existing model development, evaluation, and annotation workflows and returns claim-level verdicts, supporting evidence, citations, and structured results for downstream use.

ArgusRL is available today through the AWS Marketplace and directly through TrustScale. To learn more or request a demonstration, visit https://trustscale.ai/en/argusrl.

About TrustScale

TrustScale is an AI training, evaluation and assurance company helping organizations create, shape and use artificial intelligence with greater confidence and control. Built on more than 20 years of experience in AI data across 200-plus languages, TrustScale develops independent technologies that detect AI mistakes, evaluate claims against deterministic empirical evidence and keep people at the center of consequential decisions. Its Argus suite spans the AI lifecycle, from real-time hallucination detection and evidence-based correction at the point of use, to automated evaluation and reinforced feedback for model training and continuous improvement. Learn more at TrustScale.ai.

Media contact:

Songue PR for TrustScale
[email protected]

SOURCE: TrustScale



View the original press release on ACCESS Newswire

L.Holland--TFWP