Godfrey Njoroge · · 4 min read
Your Video Has English, Swahili and Kikuyu Sections - Here's How to Get It Right
Contents
A documentary that's mostly English with a Kikuyu-language testimonial. An NGO training video in Swahili with an English-language expert interview cut in. A news package with an English anchor intro and a Swahili interview clip. None of that is unusual - and it quietly breaks most automated video transcription and translation tools, including a lot of dedicated Swahili-to-English translators built for exactly this language pair.
Key takeaways
- Most transcription pipelines - generic tools and Kauli's own AI-drafting step alike - assume one language for the whole file, not a different language per section.
- A section that doesn't match the assumed language gets transcribed as if it were that language, producing garbled or nonsensical text for exactly that part.
- Kauli's real fix isn't automated per-segment detection (we don't have that yet, and say so honestly) - it's that a human editor listens to the actual audio and corrects whatever language each part really is, on every order, regardless.
Why this trips up automated tools specifically
Automatic speech recognition models are usually run with a single language parameter for an entire audio file - it's how Whisper and most other ASR engines are actually called in a typical pipeline, Kauli's included.
That single parameter tells the model what language to expect for every second of the recording, not just the dominant one.
The industry term for a recording that doesn't hold to one language is code-switching - most often used for a speaker mixing languages mid-sentence, but the same underlying problem shows up at the larger scale of a whole section in a different language too.
Either way, a pipeline that assumes one language throughout will mishandle whatever doesn't match it.
What actually goes wrong if nobody accounts for it
If you submit a mostly-English video and tell the tool it's English, the English portions transcribe fine.
The Swahili or Kikuyu section gets forced through the same English-language decoder anyway.
The result is usually garbled, nonsensical text for exactly that section - not a clear error message, just quietly wrong output that looks like a transcript until you actually read it against the audio.
A translation or dub built on top of that mistranscription inherits the same error, now harder to catch because nobody's comparing it back to the source anymore.
What most video-translation tools actually offer today
Searching for how other services handle this turns up mostly single-language-pair tools - an "English to Swahili" translator, a "Swahili to English" one, built the same way, one fixed language assumption per file.
Some larger transcription platforms now advertise dedicated code-switching detection, usually aimed at the mid-sentence version of the problem rather than a whole separate section in a different language.
None of this is a knock on those tools specifically - it's a real, structural limitation of how automatic speech recognition is normally deployed, not a bug unique to any one vendor.
How Kauli actually handles this today - honestly
We don't have automated per-section language detection either, and we're not going to pretend otherwise.
What we do have, on every single order regardless of plan: a real human editor who listens to the actual source audio and corrects the transcript against what's really being said, in whatever language it's really in.
If the AI's first pass mishandled a Swahili or Kikuyu section because the order was set to English, that's exactly the kind of mistake the human review step exists to catch - not a special case, just the same real review every order already gets.
The one thing worth doing on your end
Tell us. When you submit an order with a section in a different language, add a note - "there's a Swahili section around the 4-minute mark" is enough.
The notes field on every order goes straight to the editor working on it, visible on the same brief they check before they start.
It doesn't change what we do - a human is already checking the whole file against the audio either way - but it means they know to expect the switch instead of discovering it mid-review, which usually means a faster, cleaner first pass.
What you actually get back
One transcript, one translation, one set of captions or a dub - covering the whole video, each section in the language it's actually in, checked by a person against the real audio before anything ships.
Not a fully automated per-segment miracle - a real editor doing the part that automation still genuinely can't do reliably yet.
See our formatting standards page for how a transcript is formatted by default, or start with a real clip that has this exact problem - the first few minutes are free to try.