Nordic speech models, built from the ground up.
Nordic speechmodels, built fromthe ground up.
K—Speech is our proprietary speech program. We build the models, training data and evaluation loop ourselves, keeping the capability in-house.
Nordic speech, reviewed by people.
The workspace below is part of K—Speech. Every take passes a human before it counts: play it, mark the exact second, leave a note, approve it or send it back.
Four passes before a take counts. Nothing ships on a model's word alone.
One model family. Every function of speech.
Delayed Streams Modeling is symmetric: the same architecture, codec and aligner run both directions. Swap the delay and you swap the task. Everything below shares one Norwegian/Swedish adaptation recipe.
Speech to text ASR
Text to speech TTS
Same architecture, both directions. Given text and audio streams, ASR corresponds to the text stream being delayed, while the opposite gives a TTS model.Zeghidour et al. · arXiv:2509.08753
Everything a reviewer needs, and nothing else.
Speech is judged, not scored.
We open the review console to partners who record with us. If your organisation has Nordic audio and a standard to hold, we would like to hear from you.
Become a reviewer → hello@kapllan.ai →