Toward Real-Time Medical IVR in Latvian: Evaluation Framework and Comparative Analysis of STT and TTS Models
DOI:
https://doi.org/10.7250/csimq.2026-47.05Keywords:
Speech-to-Text, Text-to-Speech, Interactive Voice Response, Medical Domain, Latvian Language, Low-Resource LanguageAbstract
This article investigates state-of-the-art speech-to-text (STT) and text-to-speech (TTS) technologies for the development of a conversational interactive voice response (IVR) solution in the Latvian medical domain. The study addresses challenges associated with low-resource languages, domain-specific medical terminology, real-time interaction, and sensitive data processing. A requirement-driven evaluation framework is proposed that integrates linguistic and domain accuracy, architectural and deployment suitability, real-time performance, scalability, and regulatory considerations relevant to healthcare environments. Based on evidence reported in scientific literature, model documentation, and publicly available repositories, candidate STT and TTS models with Latvian support or adaptation potential are comparatively assessed against these requirements. Rather than providing a standardized experimental benchmark, the analysis identifies architectures and models whose documented characteristics appear most closely aligned with the operational requirements of real-time Latvian medical IVR deployment. The study provides a structured basis for candidate model selection and defines the metrics and deployment conditions to be addressed in subsequent common-condition empirical evaluation and clinical validation.
References
Y. Chen, C. Zhang, R. Bai, T. Sun, W. Ding, and R. Wang, “A Review of Medical Text Analysis: Theory and Practice,” Information Fusion, vol. 119, article 103024, 2025.
N. Kondai, A. Devapangu, S. K. Hosur, P. Raj, N. Tiwari, and K. S. Nataraj, “Enhancing Medical ASR Accuracy through the Integration of T5-Based Small Language Models and Error Correction Mechanisms,” in Proceedings of the 2024 IEEE Conference on Engineering Informatics (ICEI), pp. 1–6, 2024.
S. J. Adams, J. N. Acosta, and P. Rajpurkar, “How Generative AI Voice Agents Will Transform Medicine,” npj Digital Medicine, vol. 8, article 353, 2025.
L. Liu et al., “A Survey on Medical Large Language Models: Technology, Application, Trustworthiness, and Future Directions,” arXiv preprint arXiv:2406.03712, 2024.
H. Ahlawat, N. Aggarwal, and D. Gupta, “Automatic Speech Recognition: A Survey of Deep Learning Techniques and Approaches,” International Journal of Cognitive Computing in Engineering, vol. 6, pp. 201–237, 2025.
A. M. Alkalbani et al., “A Systematic Review of Large Language Models in Medical Specialties: Applications, Challenges and Future Directions,” Information, vol. 16, no. 6, article 489, 2025.
R. Safarik and L. Mateju, “Automatic Development of ASR System for an Under-Resourced Language,” in Proceedings of the 41st International Conference Telecommunications and Signal Processing (TSP), pp. 100–103, 2018.
R. Dargis et al., “BalsuTalka.lv – Boosting the Common Voice Corpus for Low-Resource Languages,” in Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pp. 2080–2085, 2024.
M. Pinnis, A. Salimbajevs, and I. Auzina, “Designing a Speech Corpus for the Development and Evaluation of Dictation Systems in Latvian,” in Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016), pp. 775–780, 2016.
I. Auzina et al., “Recent Latvian Speech Corpora for Linguistic Research and Technology Development,” Baltic Journal of Modern Computing, vol. 12, no. 4, pp. 646–658, 2024.
A. Znotins, D. Gosko, and N. Gruzitis, “LATE: Open Source Toolkit for Latvian and Latgalian Speech Transcription,” in Proceedings INTERSPEECH 2025, pp. 306–307, 2025. Available: https://www.isca-archive.org/interspeech_2025/znotins25_interspeech.pdf
A. Salimbajevs and J. Kapočiūtė-Dzikienė, “Automatic Speech Recognition Model Adaptation to Medical Domain Using Untranscribed Audio,” in Digital Business and Intelligent Systems. Baltic DB&IS 2022. Communications in Computer and Information Science, vol. 1598, Springer, pp. 65–79, 2022.
J. Huh, S. Park, J. E. Lee, and J. C. Ye, “Improving Medical Speech-to-Text Accuracy using Vision-Language Pre-training Models,” IEEE Journal of Biomedical and Health Informatics, vol. 28, no. 3, pp. 1692–1703, 2024.
A. Znotins, N. Gruzitis, and R. Dargis, “From Conversational Speech to Readable Text: Post-Processing Noisy Transcripts in a Low-Resource Setting,” in Proceedings of the 10th Workshop on Noisy and User-generated Text (WNUT), pp. 143–148, 2025.
R. Dargis and I. Auzina, “Towards a Modern Text-to-Speech System for Latvian,” in Human Language Technologies – The Baltic Perspective, vol. 307, IOS Press, pp. 26–29, 2018.
I. Auzina, et al., “Development of a Specialized Latvian Speech Corpus and Pronunciation Dictionary for the Linguistic Analysis and Systematic Transcription of Visual Diagnostic Examinations,” Letonica, vol. 47, pp. 244–262, 2022 (in Latvian).
R. Dargis, N. Gruzitis, I. Auzina, and K. Stepanovs, “Creation of Language Resources for the Development of a Medical Speech Recognition System for Latvian,” in Human Language Technologies – The Baltic Perspective, vol. 328, IOS Press, pp. 135–141, 2020.
M. Kronis, A. Salimbajevs, and M. Pinnis, “Code-Mixed Text Augmentation for Latvian ASR,” in Proceedings of the 2024 Joint International Conference Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pp. 3469–3479, 2024.
N. Gruzitis, R. Dargis, V. J. Lasmanis, G. Garkaje, and D. Gosko, “Adapting Automatic Speech Recognition to the Radiology Domain for a Less-Resourced Language: The Case of Latvian,” in Intelligent Sustainable Systems, Lecture Notes in Networks and Systems, vol. 333, Springer, pp. 267–276, 2022.
I. Skadina, et al., “Latvian Language in the Digital Age: The Main Achievements in the Last Decade,” Baltic Journal of Modern Computing, vol. 10, no. 3, pp. 490–503, 2022.
Children’s Clinical University Hospital, “Automated Voice Communication Solutions for the Healthcare Industry,” 2024 (in Latvian). Available: https://www.bkus.lv/lv/automatizeti-balss-komunikacijas-risinajumi-veselibas-nozarei. Accessed on Feb. 16, 2026.
Tilde, “TILDE Develops New Voice-Enabled AI Solution to Support Hospital Visitors and Doctors,” Tilde News, 2024. Available: https://tilde.ai/news/tilde-develops-new-voice-enabled-ai-solution-to-support-hospital-visitors-and-doctors/. Accessed on Feb. 16, 2026.
SoftMedico, “Viola,” n.d. Available: https://viola.softmedico.eu/. Accessed on July 22, 2026.
Labs of Latvia, “An AI-Powered Transcription Tool for the Medical Field in the Latvian Language Has Been Developed,” 2025 (in Latvian). Available: https://labsoflatvia.com/aktuali/izstradats-maksliga-intelekta-transkripcijas-riks-medicinai-latviesu-valoda. Accessed on July 22, 2026.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust Speech Recognition via Large-Scale Weak Supervision,” in Proceedings of the 40th International Conference on Machine Learning (ICML), vol. 202, article 1182, pp. 28492–28518, 2023.
Hugging Face, “Whisper,” OpenAI Models, 2024. Available: https://huggingface.co/openai/whisper-large-v3. Accessed on Jan. 15, 2026.
Hugging Face, “General-purpose Latvian ASR model,” AI Lab at IMCS (University of Latvia) Models, 2024. Available: https://huggingface.co/AiLab-IMCS-UL/whisper-large-v3-lv-late-cv19. Accessed on Jan. 15, 2026.
Hugging Face, “Latvian Whisper Small Speech Recognition Model,” Raivis Dejus Models, 2024. Available: https://huggingface.co/RaivisDejus/whisper-small-lv/blob/main/README.md. Accessed on Jan. 15, 2026.
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “Wav2Vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” in Proceedings of the 34th International Conference on Neural Information Processing Systems (NeurIPS), article 1044, pp. 12449–12460, 2020.
Hugging Face, “Wav2Vec2,” Transformers Documentation, 2025. Available: https://huggingface.co/docs/transformers/main/model_doc/wav2vec2. Accessed on Jan. 17, 2026.
Hugging Face, “Wav2Vec2-Large-XLSR-Latvian,” Jim O'Regan Models, 2025. Available: https://huggingface.co/jimregan/wav2vec2-large-xlsr-latvian-cv. Accessed on Jan. 13, 2026.
V. Pratap et al., “Scaling Speech Technology to 1,000+ Languages,” The Journal of Machine Learning Research, vol. 25, no. 1, article 97, pp. 4798–4849, 2024.
Hugging Face, “Massively Multilingual Speech (MMS) – Finetuned ASR - ALL,” AI at Meta Models, 2023. Available: https://huggingface.co/facebook/mms-1b-all. Accessed on Jan. 17, 2026.
G. Keren et al., “Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages,” arXiv preprint arXiv:2511.09690, 2025.
Hugging Face, “Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages,” AI at Meta Models, 2025. Available: https://huggingface.co/facebook/omniASR-LLM-1B. Accessed on Jan. 16, 2026.
M. Sekoyan et al., “Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST,” arXiv preprint arXiv:2509.14128, 2025.
Hugging Face, “Canary 1B v2: Multitask Speech Transcription and Translation Model,” NVIDIA Models, 2025. Available: https://huggingface.co/nvidia/canary-1b-v2. Accessed on Jan. 17, 2026.
Hugging Face, “parakeet-tdt-0.6b-v3: Multilingual Speech-to-Text Model,” NVIDIA Models, 2025. Available: https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3. Accessed on Jan. 17, 2026.
M. Z. Boito, V. Iyer, N. Lagos, L. Besacier, and I. Calapodescu, “mHuBERT-147: A Compact Multilingual HuBERT Model,” arXiv preprint arXiv:2406.06371, 2024.
GitHub, “mHuBERT-147: A Compact Multilingual HuBERT Model,” UTTER Project Repository, 2024. Available: https://github.com/utter-project/fairseq/tree/main/examples/mHuBERT-147. Accessed on Jan. 18, 2026.
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli, “Unsupervised Cross-Lingual Representation Learning for Speech Recognition,” in Proceedings INTERSPEECH 2021, pp. 2426–2430, 2021.
J. Shen et al., “Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions,” in Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4779–4783, 2018.
GitHub, “Tacotron 2,” NVIDIA Corporation Repository, 2024. Available: https://github.com/NVIDIA/tacotron2. Accessed on Jan. 13, 2026.
Y. Ren et al., “FastSpeech 2: Fast and High-Quality End-to-End Text to Speech,” arXiv preprint arXiv:2006.04558, 2020.
GitHub, “FastSpeech 2,” Chung-Ming Chien Repository, 2023. Available: https://github.com/ming024/FastSpeech2. Accessed on Jan. 14, 2026.
J. Kim, J. Kong, and J. Son, “Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech,” in Proceedings of Machine Learning Research, vol. 139, pp. 5530–5540, 2021. Available: https://proceedings.mlr.press/v139/kim21f.html
Hugging Face, “VITS,” Transformers Documentation, 2025. Available: https://huggingface.co/docs/transformers/model_doc/vits. Accessed on Jan. 15, 2026.
W. Ping et al., “Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning,” in Proceedings of the 6th International Conference on Learning Representations (ICLR), pp. 1-16, 2018. Available: https://openreview.net/pdf?id=HJtEm4p6Z
GitHub, “Deep Voice 3 (PyTorch),” Ryuichi Yamamoto Repository. Available: https://github.com/r9y9/deepvoice3_pytorch. Accessed on Jan. 14, 2026.
Z. Zhang et al., “Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling,” arXiv preprint arXiv:2303.03926, 2023.
GitHub, “VALL-E-X,” Songting Repository, 2024. Available: https://github.com/Plachtaa/VALL-E-X. Accessed on Jan. 14, 2026.
Hugging Face, “OuteTTS Version 1.0,” OuteAI Models, 2025. Available: https://huggingface.co/OuteAI/Llama-OuteTTS-1.0-1B. Accessed on Jan. 13, 2026.
AI Models, “Latvian Female TTS Model VITS Encoding Trained on CV Dataset at 22050Hz,” AI Models for Text-to-Speech Synthesis, 2022. Available: https://aimodels.org/ai-models/text-to-speech-synthesis/latvian-female-tts-model-vits-encoding-trained-on-cv-dataset-at-22050hz/. Accessed on Jan. 16, 2026.
GitHub, “tts_models--lv--cv--vits.zip,” Coqui Repository, 2022. Available: https://github.com/coqui-ai/TTS/releases/tag/v0.8.0_models. Accessed on Jan. 20, 2026.
R. Dargis and I. Auzina, “Ilvars – Latvian Male VITS Text-to-Speech Model (vers. 2023),” CLARIN-LV Digital Library at IMCS, University of Latvia, 2023. Available: http://hdl.handle.net/20.500.12574/89
Hugging Face, “Latvian Piper TTS Voice ‘Aivars’,” Raivis Dejus Models, 2025. Available: https://huggingface.co/RaivisDejus/Piper-lv_LV-Aivars-medium. Accessed on Jan. 14, 2026.
Hugging Face, “Massively Multilingual Speech (MMS): Latvian Text-to-Speech,” AI at Meta Models, 2023. Available: https://huggingface.co/facebook/mms-tts-lav. Accessed on Jan. 19, 2026.
R. Dargis, P. Paikens, N. Gruzitis, I. Auzina, and A. Akmane, “Development and Evaluation of Speech Synthesis Corpora for Latvian,” in Proceedings of the 12th Conference on Language Resources and Evaluation (LREC), pp. 6633–6637, 2020. Available: https://aclanthology.org/2020.lrec-1.818.pdf
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Sintija Petrovica-Klavina, Ingars Erins, Roberts Kirsteins, Ints Meijers, Biruta Eliza Aunina-Kirmuska (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.