Godfrey Njoroge ·
AI-Only Dubbing vs. Human-Verified: What Actually Matters for Swahili and Kikuyu
Contents
Short version: AI-only dubbing is fast and cheap, and it genuinely works well for high-resource languages with huge training datasets. Swahili and Kikuyu aren't those languages yet - which is exactly the gap a human-in-the-loop process exists to close.
What AI-only dubbing actually gets wrong
Automatic speech recognition and machine translation are trained on however much real-world audio and text exists for a language. English, Spanish, and Mandarin have enormous datasets behind them. Swahili has meaningfully less, and it shows up in specific, predictable ways: code-switching (a speaker moving between Swahili and English mid-sentence), regional accent variation, and idioms that translate literally into nonsense.
Kikuyu is a harder case again. It's genuinely under-resourced compared to Swahili - though that's changing faster than it used to. Google's WAXAL dataset and Microsoft's Paza research project have both published real 2026 work specifically targeting Kikuyu speech recognition, alongside smaller open-source efforts. None of that means a production-ready, reliably accurate Kikuyu ASR model is a drop-in replacement for a human transcriber today - research datasets and benchmark papers are a different thing from a service you can point real client audio at and trust. We route Kikuyu to a human transcriber for exactly that reason: not because nothing exists at all, but because "exists in a research paper" and "accurate enough for a client deliverable" aren't the same claim, and we're not willing to guess which side of that line a given clip lands on.
Where AI genuinely helps
None of this is an argument against AI - it's an argument against AI *alone*. Speech recognition and machine translation are excellent at producing a fast first draft: getting words on a page in minutes instead of hours, catching the easy 90% correctly so a human editor's time goes toward the hard 10% that actually needs judgment. That's the whole design of Kauli's pipeline - AI drafts, a real editor checks every word before anything ships, for every order regardless of service level.
What to actually check before trusting an AI-only dub or transcript
- Play a sample against the source audio. Don't just read the transcript - listen while you read it, specifically around names, numbers, and anywhere the speaker switches between languages.
- Check for confident-sounding wrong answers. AI models rarely say "I'm not sure" - they produce fluent, plausible-looking output whether or not it's accurate, which makes errors harder to spot than they should be.
- Ask what happens when the model doesn't know a language well. If a vendor can't tell you their real accuracy for Swahili or Kikuyu specifically (not just "we support 100+ languages"), that's a real signal.
FAQ
Is AI-only dubbing ever good enough?
For high-resource languages and low-stakes content, often yes. For Swahili and Kikuyu, or for anything where accuracy has real consequences (compliance, an NGO's public messaging, a broadcast), we don't think it's there yet - which is a statement about the current state of the underlying models, not a permanent one.
Does Kauli use any AI at all?
Yes, on every order - AI drafts the first pass because it's genuinely faster than starting from nothing. The point isn't "no AI," it's "AI alone isn't the deliverable."