Godfrey Njoroge ·
Sheng to Standard Swahili: Why Machine Translation Gets East African Ads Wrong
Contents
Short version: Sheng and Standard Swahili aren't the same language, and treating them as interchangeable is exactly how a generic AI translation tool makes an ad sound stiff, dated, or accidentally hilarious to the audience it's trying to reach.
Where Standard Swahili actually comes from
What's taught in schools and used on the news - Kiswahili Sanifu - isn't a natural, organically-evolved dialect. It's a constructed standard: in 1928 the Inter-Territorial Language Committee met in Mombasa and formally chose Kiunguja, the Swahili dialect of Zanzibar Town, as the basis for standardization across East Africa. That's why Standard Swahili sounds a certain way - it was deliberately built from one coastal dialect and spread through schools and broadcast media, not picked up naturally in Nairobi's streets.
Sheng is a different, real thing
Sheng grew out of Nairobi itself, starting with rural-to-urban migration bringing together Kikuyu, Luo, Kamba, Swahili and English speakers into the same low-income neighborhoods from the late 1970s onward. Linguists describe it as sitting on a continuum - Standard Swahili at one end, Sheng at the other, with everyday Nairobi vernacular Swahili in between. Sheng's grammar stays recognizably Swahili underneath, but it borrows vocabulary heavily from English and other Kenyan languages, and it innovates - new verb forms, shifted noun-class agreement, slang that changes fast enough that a term from five years ago can already sound dated.
Why this actually matters for a translation
A generic machine translation model, trained mostly on Standard Swahili text (news articles, government documents, formal writing), will confidently produce grammatically correct Standard Swahili for content that was never meant to sound that formal. For an ad, a social campaign, or anything targeting a younger urban audience, that's not a small stylistic miss - it's the difference between sounding like you understand your audience and sounding like a government pamphlet. The reverse mistake happens too: rendering something into heavy Sheng when the actual audience is older, rural, or expects a more formal register.
Machine translation also has no reliable way to know which register a given piece of source content calls for - it just produces its most statistically likely output, which skews toward whatever dialect dominates its training data.
What we actually do about it
This is a judgment call, not a lookup-table problem - it's exactly the kind of thing our human review step exists for. A human editor reading the actual context (who's the audience, what's the tone, is this a government notice or a youth-targeted ad) can make that call in a way no model can yet. If you're localizing marketing or social content specifically for a younger Kenyan audience, say so when you submit the order - it changes the register we translate into.
FAQ
Can Kauli translate into Sheng specifically?
We translate into Standard Swahili by default, since that's the safe, broadly-understood register for most content. For marketing or social content aimed at a youth audience where Sheng or a more casual register is genuinely the right call, tell us in your order notes and a human editor will make that judgment - it's not something we'd want a model guessing at unsupervised.
Isn't this the kind of thing AI will just get better at over time?
Possibly, for register detection specifically. But knowing an audience well enough to pick the right register isn't purely a language-modeling problem - it's a judgment call about who you're actually talking to, which is exactly why we keep a human in that loop rather than waiting for a model to get confident about it.