The Fort Worth Press - LLM Consensus Matches or Outperforms the Best AI Models in Expert Evaluation Without Performance Degradation

USD -
AED 3.672498
AFN 64.99967
ALL 80.055475
AMD 365.038314
ANG 1.790365
AOA 916.999686
ARS 1512.769098
AUD 1.405432
AWG 1.8
AZN 1.692558
BAM 1.705954
BBD 2.026944
BDT 123.912408
BGN 1.683441
BHD 0.379431
BIF 3010.06145
BMD 1
BND 1.281516
BOB 11.065415
BRL 5.148994
BSD 1.006407
BTN 96.572478
BWP 13.649903
BYN 3.056204
BYR 19600
BZD 2.024065
CAD 1.398695
CDF 2310.000083
CHF 0.825565
CLF 0.024159
CLP 953.919973
CNY 6.706403
CNH 6.707855
COP 3134.19
CRC 450.170308
CUC 1
CUP 26.5
CVE 96.179091
CZK 21.20565
DJF 179.213314
DKK 6.513475
DOP 59.409222
DZD 134.033032
EGP 52.038498
ERN 15
ETB 164.385519
EUR 0.87133
FJD 2.24075
FKP 0.743677
GBP 0.746544
GEL 2.605021
GGP 0.743677
GHS 11.558238
GIP 0.743677
GMD 74.00016
GNF 8849.017188
GTQ 7.682814
GYD 210.553402
HKD 7.844945
HNL 27.011091
HRK 6.5665
HTG 131.535909
HUF 317.228009
IDR 17742
ILS 3.035498
IMP 0.743677
INR 95.925501
IQD 1318.411974
IRR 1374575.000204
ISK 121.819979
JEP 0.743677
JMD 158.820015
JOD 0.709037
JPY 155.747498
KES 129.601128
KGS 87.449902
KHR 4075.170853
KMF 428.000022
KPW 900.000318
KRW 1384.149724
KWD 0.30872
KYD 0.838672
KZT 447.50121
LAK 22527.582748
LBP 90114.343029
LKR 333.53975
LRD 174.607813
LSL 16.383324
LTL 2.95274
LVL 0.60489
LYD 6.391878
MAD 9.490211
MDL 17.544808
MGA 4351.675382
MKD 53.66543
MMK 2099.664815
MNT 3596.851201
MOP 8.131937
MRU 40.29552
MUR 47.659857
MVR 15.410109
MWK 1745.118648
MXN 17.21102
MYR 4.099499
MZN 63.901804
NAD 16.384681
NGN 1330.92999
NIO 37.035906
NOK 9.41832
NPR 154.516313
NZD 1.74333
OMR 0.384497
PAB 1.006407
PEN 3.376408
PGK 4.550549
PHP 62.750497
PKR 278.944747
PLN 3.80184
PYG 5954.878299
QAR 3.658582
RON 4.582903
RSD 102.270153
RUB 84.425401
RWF 1484.92527
SAR 3.701375
SBD 8.000512
SCR 13.800587
SDG 601.504736
SEK 9.817025
SGD 1.275805
SHP 0.746965
SLE 24.640284
SLL 20969.491881
SOS 575.164311
SRD 37.752501
STD 20697.981008
STN 21.370224
SVC 8.806277
SYP 13002.000254
SZL 16.382675
THB 33.349023
TJS 9.283923
TMT 3.5
TND 2.946387
TOP 2.40776
TRY 48.675301
TTD 6.823643
TWD 31.899397
TZS 2657.177986
UAH 44.88157
UGX 3939.971477
UYU 40.489856
UZS 11865.271642
VES 845.45495
VND 26013
VUV 118.307083
WST 2.742913
XAF 571.55478
XAG 0.015633
XAU 0.000232
XCD 2.70255
XCG 1.813765
XDR 0.707052
XOF 571.55478
XPF 104.025016
YER 236.650369
ZAR 16.28741
ZMK 9001.197507
ZMW 19.750448
ZWL 321.999592
SSP 5658.175743
MXV 1.951372
  • RBGPF

    0.0000

    69.99

    0%

  • CMSD

    0.1785

    20.29

    +0.88%

  • RYCEF

    -0.0100

    19.29

    -0.05%

  • CMSC

    0.1400

    20.48

    +0.68%

  • RELX

    0.0500

    34.27

    +0.15%

  • NGG

    1.1000

    76.03

    +1.45%

  • GSK

    0.2800

    50.29

    +0.56%

  • RIO

    -1.4700

    95.79

    -1.53%

  • VOD

    -0.2200

    17.46

    -1.26%

  • BCE

    -0.4000

    22.5

    -1.78%

  • JRI

    -0.1200

    11.5

    -1.04%

  • AZN

    1.0400

    162.89

    +0.64%

  • BTI

    -0.4300

    56.09

    -0.77%

  • BCC

    -0.6200

    75.31

    -0.82%

  • BP

    -1.5800

    45.38

    -3.48%

LLM Consensus Matches or Outperforms the Best AI Models in Expert Evaluation Without Performance Degradation
LLM Consensus Matches or Outperforms the Best AI Models in Expert Evaluation Without Performance Degradation

LLM Consensus Matches or Outperforms the Best AI Models in Expert Evaluation Without Performance Degradation

A multi-model consensus system matches or outperforms GPT-5.4, Claude Opus 4.6 and Gemini 3.1 Pro across 100 expert-level questions infinance, law, medicine and technology, with no performance degradation.

Text size:

SHERIDAN, WY / ACCESS Newswire / April 2, 2026 / LLM Consensus has released the results of its Expert-Domain Evaluation Benchmark v1.0, an independent study analyzing the performance of its multi-model consensus technology across 100 high-complexity questions in areas such as financial regulation, law, clinical medicine and technical architecture.

According to the results, the system matches or outperforms the best individual AI model across all evaluated questions, achieving measurable improvement in 44.9% of cases and with no instances of performance loss.

Key findings

In nearly half of the questions (45%), responses generated by the consensus system clearly outperformed those of the best individual model. The system was able to identify regulatory details that other models missed, resolve contradictions across sources, and deliver more complete answers.

In the remaining 55%, performance matched that of the best available model, ensuring a consistent baseline of quality without requiring users to choose between different models.

Additionally, in none of the 100 questions analyzed did the system produce a worse result than an individual model.

Performance by domain

The analysis focused on complex questions typical of regulated industries:

  • Clinical medicine (59% improvement): stronger performance in complex drug interactions, comorbidities, and application of clinical guidelines.

  • Financial regulation (50% improvement): advantages in scenarios combining multiple European regulatory frameworks such as DORA, PSD2, GDPR, and NIS2.

  • Legal analysis (44% improvement): greater precision in multi-jurisdictional and cross-regulatory compliance questions.

  • Technical architecture (30% improvement, 70% match): consistent results in system design decisions under regulatory and technical constraints.

Why it matters

The use of artificial intelligence in regulated industries continues to grow, yet no single model consistently excels across all domains. A system may perform well in financial regulation but fall short in clinical medicine, or vice versa.

LLM Consensus addresses this challenge by combining multiple leading models into a single response. It integrates technologies from OpenAI, Anthropic, Google, Mistral, and Meta, applying a synthesis process with cross-verification that leverages each model's strengths while reducing their weaknesses.

"Reliability is the core value proposition," the company said. "Users no longer have to decide which model to use. They get a single answer that consistently matches or outperforms the best available model for each case."

Evaluation methodology

The benchmark was specifically designed to assess tasks that require combining multiple sources of knowledge. Each question was evaluated by three independent reviewers from different AI providers, who scored responses blindly based on accuracy and quality.

Responses - from both the consensus system and individual models - were presented anonymously and in random order. Cases where sufficient agreement was not reached were classified as inconclusive and excluded from the final results.

The full dataset has been published to enable independent verification.

About LLM Consensus

LLM Consensus is an AI orchestration API that combines multiple advanced models into a single optimized response using patent-pending consensus technology.

The solution is available via REST API with different operating modes and is designed for developers and organizations in regulated sectors such as finance, healthcare, legal, and technology.

Press contact

Francisco Javier Nunez
Email: [email protected]
Web: llmconsensus.io

Patent pending: US 19/215,933 | EU EP25176020.3

This press release contains forward-looking statements based on current benchmark results. The evaluation was conducted using specific model versions as of March 2026; performance may vary with model updates. LLM Consensus is a system benchmark evaluating multi-model orchestration on expert synthesis tasks and should not be interpreted as a general-purpose comparison of individual AI models.

SOURCE: LLM Consensus



View the original press release on ACCESS Newswire

T.Dixon--TFWP