NEWS · SEPTEMBER 25, 2026 · AI

Patientdesk.ai introduced two speech models built for Turkish: Alania and Duyu

Webrazzi reported on September 25, 2026 that voice AI startup Patientdesk.ai had introduced two Turkish-focused speech models: Alania for text to speech and Duyu for speech to text. The company reports Duyu's word error rate below OpenAI Whisper large-v3 on the Turkish FLEURS test and below ElevenLabs Scribe v2 on its own phone-quality test set.

01 · WHAT HAPPENED?

One model speaks, the other writes it down

Webrazzi reported the launch on September 25, 2026 under a headline naming both models, Alania and Duyu. Alania runs in the text to speech direction and supports tone variations such as appointment, question and calming delivery. It ships a single voice output. According to Webrazzi, Alania can stream audio without waiting for the sentence to finish, which is why the company points it at low-latency situations such as phone calls.

Duyu runs the other way, from speech to text. It transcribes pre-recorded audio files and also converts live speech into text in real time. The report lists the hard cases the model is meant to cover: phone-quality audio, and conversations where an email address or a phone number is spoken aloud. Between them the two models hold both ends of a voice session, one producing audio, the other turning incoming audio into text.

02 · DETAILS

The word error rate numbers, and open source Antalia-1

Duyu's claim rests on two comparisons. On the Turkish FLEURS test Duyu is reported at a 4.71 percent word error rate against 5.04 percent for OpenAI Whisper large-v3. The second comparison was run on the company's own phone-quality audio test set, where Duyu scores 9.87 percent and ElevenLabs Scribe v2 scores 10.4 percent. The gap between the two sets is informative on its own: Duyu's reported rate goes from 4.71 percent on FLEURS to 9.87 percent on the phone-quality set.

Two details matter on the access side. Both models are compatible with the widely used OpenAI API format, and as part of the launch they are free for one month from the sign-up date. The company also published Antalia-1, an experimental open source Turkish text to speech model, on Hugging Face together with its weights, code and technical report. The model card lists Sezgin Saygılı, Emre Kaplaner, Öncel Özgül and Fikri San Köktaş as authors, with Patientdesk.ai as the affiliation. The accompanying antalia-voice-corpus consists of 1,073 segments and 5.008 hours of audio from a single consenting voice actor.

03 · WHY IT MATTERS

A measured number for Turkish, and latency for the phone line

What counts here is not only that the models exist but the ground the comparison was run on. Duyu's figures come from a Turkish evaluation set and from a second set built out of phone-quality audio, and the models placed opposite it are OpenAI Whisper large-v3 and ElevenLabs Scribe v2. The contents of that second set also reveal the target, because phone-quality audio plus spoken email addresses and phone numbers sit squarely inside the phone-call scenario the company describes for Alania.

The second point is integration. Both models are said to be compatible with the OpenAI API format, and our reading at UNALSOFT is that existing clients written against that format would not need to be rewritten. We looked at why standard interfaces get so much attention in agent architectures in our September 11 piece on the OpenAI Agents API. The funding behind the company appears in the report as well: Webrazzi recalls that Patientdesk.ai announced a 1 million dollar pre-seed round led by Y Combinator in February.

04 · TÜRKİYE

Turkish is the target language here, but open ends remain

There is no Türkiye gap in this story; the whole announcement is built on Turkish. Both models are presented as Turkish-focused, the speech recognition comparison uses the Turkish portion of FLEURS, and the open source release, Antalia-1, is a Turkish text to speech model published alongside a Turkish voice corpus. The funding side is in the report too: a 1 million dollar pre-seed round led by Y Combinator in February.

The things the sources do not contain deserve the same clarity, and this paragraph is UNALSOFT commentary. Neither the report nor the model card says: what pricing looks like once the one free month ends, what the measured latency actually is, whether Duyu's word error rates have been verified by an independent third party, whether additional voices are planned for Alania, whether the models cover any language other than Turkish, which regions the service can be called from, where voice data is processed and how that is handled under KVKK, or how many customers and production deployments exist. Those questions need answers before a Turkish speech model goes into a live business flow.

The UNALSOFT view

For us this is not a product recommendation but a reminder about measurement. If you are building a Turkish-speaking voice flow, the deciding question is not which model is described best, it is which word error rate each one produces on your own recordings. Patientdesk.ai publishing two separate test sets makes exactly that point, because the distance between the two test sets shows up directly in the numbers. That is why our agentic AI work starts by pulling a small test set out of the client's own call recordings and comparing models on it. Since pricing, latency and data processing terms are absent from the sources, this article makes no assessment of them.

Ready to measure your Turkish voice flow on your own recordings?

A short conversation is enough to plan which speech model actually holds up in your call and appointment flows.

Message on WhatsApp