The Fort Worth Press - Modulate Launches Velma Transcribe: High-Performance Transcription For Real World Conversations at 90% Lower Cost

USD -
AED 3.672502
AFN 65.000288
ALL 79.92783
AMD 363.200558
ANG 1.790365
AOA 916.99996
ARS 1512.724296
AUD 1.405896
AWG 1.8
AZN 1.701759
BAM 1.70475
BBD 2.015105
BDT 122.866121
BGN 1.683441
BHD 0.377158
BIF 3009.975725
BMD 1
BND 1.276579
BOB 10.179685
BRL 5.132305
BSD 1.000449
BTN 95.931629
BWP 13.567684
BYN 3.041441
BYR 19600
BZD 2.012246
CAD 1.39962
CDF 2310.0002
CHF 0.82448
CLF 0.024171
CLP 954.419682
CNY 6.706402
CNH 6.70556
COP 3134.56
CRC 447.578413
CUC 1
CUP 26.5
CVE 96.109511
CZK 21.1723
DJF 178.159742
DKK 6.51085
DOP 59.260873
DZD 133.876605
EGP 52.161402
ERN 15
ETB 163.423038
EUR 0.87102
FJD 2.24075
FKP 0.743677
GBP 0.747565
GEL 2.604997
GGP 0.743677
GHS 11.515883
GIP 0.743677
GMD 73.999749
GNF 8795.70108
GTQ 7.633807
GYD 209.2854
HKD 7.845245
HNL 26.85343
HRK 6.5669
HTG 130.759174
HUF 316.173502
IDR 17733
ILS 3.03294
IMP 0.743677
INR 95.77355
IQD 1310.636666
IRR 1374575.000327
ISK 121.769736
JEP 0.743677
JMD 157.885791
JOD 0.70903
JPY 155.678496
KES 129.590798
KGS 87.450162
KHR 4074.837551
KMF 428.000211
KPW 900.000318
KRW 1382.815037
KWD 0.30861
KYD 0.833795
KZT 445.677328
LAK 22408.751171
LBP 89594.128747
LKR 331.841379
LRD 173.581452
LSL 16.298676
LTL 2.95274
LVL 0.60489
LYD 6.355114
MAD 9.52989
MDL 17.564075
MGA 4345.078556
MKD 53.623182
MMK 2099.664815
MNT 3596.851201
MOP 8.084808
MRU 40.080539
MUR 47.659514
MVR 15.409852
MWK 1734.87843
MXN 17.19097
MYR 4.0986
MZN 63.903695
NAD 16.298676
NGN 1330.980258
NIO 36.817998
NOK 9.42809
NPR 153.487582
NZD 1.74142
OMR 0.384505
PAB 1.000458
PEN 3.377685
PGK 4.451379
PHP 62.716015
PKR 277.316313
PLN 3.79595
PYG 5918.423381
QAR 3.647021
RON 4.583802
RSD 102.225971
RUB 84.725008
RWF 1476.267885
SAR 3.713538
SBD 8.000512
SCR 13.793554
SDG 601.483424
SEK 9.816345
SGD 1.27541
SHP 0.746965
SLE 24.639878
SLL 20969.491881
SOS 571.784692
SRD 37.752499
STD 20697.981008
STN 21.354764
SVC 8.754581
SYP 13002.000254
SZL 16.292377
THB 33.304993
TJS 9.229529
TMT 3.5
TND 2.943174
TOP 2.40776
TRY 48.67542
TTD 6.792558
TWD 31.893029
TZS 2653.280206
UAH 44.708507
UGX 3931.959922
UYU 40.225402
UZS 11781.061917
VES 845.45495
VND 26012
VUV 118.307083
WST 2.742913
XAF 571.351358
XAG 0.015409
XAU 0.00023
XCD 2.70255
XCG 1.803158
XDR 0.707052
XOF 571.351358
XPF 103.951572
YER 236.650164
ZAR 16.254503
ZMK 9001.190528
ZMW 19.635051
ZWL 321.999592
SSP 5658.175743
MXV 1.949099
  • RBGPF

    0.0000

    69.99

    0%

  • JRI

    -0.1200

    11.5

    -1.04%

  • AZN

    1.0400

    162.89

    +0.64%

  • GSK

    0.2800

    50.29

    +0.56%

  • NGG

    1.1000

    76.03

    +1.45%

  • BCE

    -0.4000

    22.5

    -1.78%

  • RIO

    -1.4700

    95.79

    -1.53%

  • BCC

    -0.6200

    75.31

    -0.82%

  • CMSC

    0.1400

    20.48

    +0.68%

  • BTI

    -0.4300

    56.09

    -0.77%

  • RELX

    0.0500

    34.27

    +0.15%

  • RYCEF

    -0.0100

    19.29

    -0.05%

  • BP

    -1.5800

    45.38

    -3.48%

  • VOD

    -0.2200

    17.46

    -1.26%

  • CMSD

    0.1785

    20.29

    +0.88%

Modulate Launches Velma Transcribe: High-Performance Transcription For Real World Conversations at 90% Lower Cost
Modulate Launches Velma Transcribe: High-Performance Transcription For Real World Conversations at 90% Lower Cost

Modulate Launches Velma Transcribe: High-Performance Transcription For Real World Conversations at 90% Lower Cost

Modulate's ELM model architecture unlocks transcription for the masses, cutting costs by 10x while achieving industry-leading accuracy.

Text size:

BOSTON, MA / ACCESS Newswire / March 18, 2026 / Modulate, the frontier conversational voice intelligence company, today announced Velma Transcribe, a speech-to-text API delivering high-accuracy, low-latency transcription at 90% lower cost per hour than other leading transcription providers. This significantly lower price point represents a fundamental shift in the economics of transcription. For a fraction of the cost, Modulate unlocks affordable speech-to-text transcription for every audio conversation in the world, empowering real-time voice agents, call center platforms, social apps, and more with industry-leading transcription tools at a global scale.

Built using Modulate's industry-leading Ensemble Listening Model (ELM) research, Velma Transcribe orchestrates an ensemble of specialized transcription models to improve accuracy, latency, and cost efficiency compared to any single model. In addition to the outstanding unit economics, Velma Transcribe achieves industry-leading results on widely used datasets, including Earnings-22 and the AMI Meeting Corpus. The result is a new standard for conversational audio transcription, combining strong accuracy on complex multi-speaker audio with dramatically improved unit economics for processing voice data at scale.

"Modulate is the world leader in using voice understanding AI, and our goal is to make the tools to understand audio available to anyone, at any scale," said Carter Huffman, CTO and Cofounder of Modulate. "Our full ensemble for conversation understanding, Velma, already outperforms LLMs in recognizing key behaviors, and now Velma Transcribe makes one of our core underlying capabilities available directly to developers who simply need accurate transcripts, not behavioral insights."

In addition, Velma Transcribe offers features built for Enterprise use cases:

  • Emotion detection (20+ emotions)

  • Accent detection (20+ accents)

  • Multilingual (70+ languages)

  • PII redaction, diarization, streaming support, and more

Lower Transcription Costs By up to 10X

Velma Transcribe reduces transcription costs to approximately $0.03 per hour of audio, more than 90% lower than leading providers. These economics make it far more cost-effective for enterprise organizations to analyze and monetize their voice data.

  • $0.03 - Modulate Velma Transcribe

  • $0.40 - ElevenLabs Scribe v2

  • $0.31 - Deepgram Nova-3

  • $0.26 - Deepgram Nova-2

  • $0.21 - AssemblyAI Universal-3 Pro

*Based on publicly listed pricing as of March 18, 2026

Compare the leading speech-to-text transcription companies on cost and accuracy at Speechtxt.com.

Top Marks for Conversational Audio Accuracy at Scale

Velma Transcribe is engineered for real-world conversations that challenge traditional systems, including overlapping speakers, interruptions, accents, and background noise. On the AMI Meeting Corpus dataset, a widely used benchmark for complex multi-speaker conversational audio, Velma avoids over 40% of the errors made by Eleven Labs and over 70% of the errors made by OpenAI GPT-4o-transcribe.

Huffman explains the top marks, "We've tuned Velma for conversational audio, including emotion and accent detection, leading to materially lower error rates on meeting and call data while delivering dramatic cost savings versus incumbent providers. That combination makes high-quality transcription practical at scale."

Built for Secure Enterprise Voice Production

Velma Transcribe includes all the capabilities developers expect and enterprise operations need, including:

  • Batch and streaming transcription endpoints with structured output and segment timestamps

  • Zero data stored, ensuring privacy-safe workflows

  • Sub-second streaming latency with partial transcripts for live applications and agent pipelines

  • Robust formatting optimized for conversational speech and long recordings

  • Broad language coverage in 70 of the world's most commonly spoken languages

  • Personally Identifiable Information (PII) detection and redaction

  • Advanced transcription enrichments, including speaker diarization, emotion detection, and accent identification

Backed by Modulate's security practices and ISO 27001 certification, these capabilities allow developers to build secure, voice-enabled applications and help organizations extract insights from large volumes of conversational data.

Models that Listen and Understand

Velma Transcribe is part of Modulate's growing family of Velma 2.0 voice analytics models built to deliver a new, context-rich listening layer for AI systems. It represents the first step in Modulate's expanding developer API strategy, with additional capabilities planned across synthetic voice detection, emotion analysis, and deeper conversational intelligence. Together, these capabilities allow developers and enterprises to move beyond transcription to understand how conversations unfold, enabling applications such as fraud detection, customer sentiment analysis, compliance monitoring, and real-time decision support.

"The industry has spent years teaching AI how to generate and respond. The next frontier is teaching it how to listen," said Mike Pappas, CEO and Cofounder of Modulate. "Most systems today rely on transcription, reducing rich conversations to flat text and losing the signals humans naturally understand. Velma is the listening layer for AI, giving developers and enterprises the 'ears' needed to build voice-native applications that can capture the nuance and intent within spoken dialogue."

Availability and Pricing

Velma Transcribe is available today with batch and sub-second streaming transcription. Modulate pricing is usage-based and optimized for high-volume workloads: https://www.modulate.ai/pricing

About Modulate

Modulate is a voice intelligence company building AI models and APIs designed to understand real-world conversational audio at scale. Its technology combines speech recognition, acoustic analysis, and conversational context to deliver reliable, explainable, and cost-effective voice intelligence for developers and enterprises.

For more information or to get started, visit modulate.ai.

Media Contact

Megan Fasy
Grithaus Agency
(e) [email protected]
(m) +1 (617) 480-3674

###

SOURCE: Modulate



View the original press release on ACCESS Newswire

L.Rodriguez--TFWP