AI for brokers

Your AI reads isiZulu less accurately than English: test the inbox before you automate

Sentiment models score lower on isiZulu and Sesotho than on English. How a brokerage should test its WhatsApp classifier before letting it reply on its own.

Published on 6 min readFCB.ai
Contents
  1. What the research actually reports
  2. What a wrong label costs in a brokerage
  3. A test any brokerage can run in an afternoon
  4. Rules that hold up once you switch it on
  5. Frequently asked questions

A client in Pietermaritzburg writes: "Sawubona, ngicela ukwazi about my policy, angikwazi ukukhokha this month." Half isiZulu, half English, one sentence, and three things a brokerage needs to catch — a greeting, a question about cover, and a payment problem. Your inbox is full of messages like this. If an AI layer is reading them to score sentiment, flag churn risk or draft a reply, the honest question is not whether it works. It is how well it works in each language your clients actually use, and what happens in the cases where it does not.

What the research actually reports

There is now solid published work on this, and it is not reassuring for anyone planning to switch on automation across a multilingual book. A 2025 study in PLOS ONE by Mabokela, Primus and Celik benchmarked pre-trained language models on sentiment in five South African languages — Sepedi, Sesotho, Setswana, isiXhosa and isiZulu. The best monolingual fine-tuned setups averaged roughly 71% weighted F1 across the set. The spread matters more than the average: the Nguni languages, isiZulu and isiXhosa, exceeded 77%, while the Sotho-Tswana group sat around 63%. The same corpus was full of code-switching with English, which the authors had to work around during annotation.

The picture across the continent is similar. The AfriSenti shared task at SemEval-2023, which the Masakhane community contributed to, built sentiment datasets across a dozen-plus African languages precisely because the resources did not exist. Purpose-built Afro-centric models outperformed general multilingual ones — which is another way of saying that a general-purpose assistant, handed an isiZulu message cold, is working further outside its training distribution than most vendors admit.

Two caveats before you quote these numbers at a supplier. They are benchmarks on social-media text, not on your inbox, and short transactional insurance messages may behave better or worse. And a system does not have to classify perfectly to be useful. The point is the gap between languages, and the fact that nobody — including your provider — knows what that gap looks like on your book until someone measures it.

What a wrong label costs in a brokerage

An error rate is abstract until you map it onto the things your inbox actually does with a label.

What the client sentLikely misreadWhat the system doesWhat it costs
A bereavement notice in SesothoNeutral sentimentSends a cheerful automated replyA family that never comes back, and a complaint you deserve
Polite cancellation intent in isiZuluNo churn signalNo flag, no follow-upA lapse discovered at renewal
An urgent motor claim mixed with EnglishGeneral enquiryOrdinary queue, no escalationA late notification and a hard conversation with the insurer
Sarcasm or an idiom about a declined claimPositive sentimentAuto-reply instead of a humanAn escalation that starts one level angrier

There is a regulatory edge to this too. The FAIS General Code of Conduct requires that what a provider tells a client is in plain language, is not misleading, and does not create confusion — a standard that is hard to meet with a machine-translated reply about cover terms. Treating Customers Fairly says the same thing in outcome language. Nothing in either framework prohibits AI in the inbox; both make you responsible for what leaves it.

A test any brokerage can run in an afternoon

You do not need a data scientist. You need one hundred real messages and two colleagues who speak the languages.

  1. Pull a sample. One hundred inbound messages from the last quarter, weighted to match your real language mix — not a hundred English ones plus five for flavour.
  2. Label them by hand first. Two colleagues label each message independently for the things your system acts on: sentiment, urgency, and whether it signals a claim, a payment problem or an intention to leave. Where the two humans disagree, note it; that disagreement is your ceiling, and no model will beat it.
  3. Run the same hundred through the system without showing it the labels. Most tools, ORIS included, have a way for an administrator to test the classifier on sample text.
  4. Score by language, never in aggregate. A single accuracy figure will be carried by your English majority and will hide exactly the problem you are looking for.
  5. Look at the errors, not the score. Sort the misses by cost. Ten missed upsell hints are an annoyance; one auto-reply to a death notice is an incident.
  6. Set the escalation rule from the evidence and re-run the test each quarter, or whenever the model behind the feature changes.

Rules that hold up once you switch it on

The point of the test is not a certificate. It is to decide, with evidence, where the machine acts and where a person does.

  • Automate intents, escalate emotions. Confirming receipt of a document or acknowledging a payment is safe in any language. Anything carrying sentiment — a complaint, a claim, a bereavement, a cancellation — goes to a human. In ORIS, negative sentiment always produces a draft and a notification rather than an automatic send, and auto-reply rules carry a cooldown and a cap per client; keep both conservative in your lower-confidence languages.
  • Reply in the language the client wrote in. Obvious, routinely broken, and easy to check in your own audit of sent messages.
  • Approve templates per language. Meta reviews each language version of a message template separately and does not translate for you — you supply the wording. Budget for the translation and the review cycle before you promise a multilingual campaign.
  • Keep a house glossary of cover terms. Excess, waiting period, cession, beneficiary, lapse: agree the wording in each language you serve, have a person who knows insurance sign it off, and use it in templates rather than letting a model improvise. Nothing generated should introduce a term your glossary does not contain.
  • Log what was AI-drafted. Your compliance file should be able to show, per message, whether a person wrote it or approved it. Our note on what a compliance officer should sign off sets out the minimum, and the piece on risk scores under POPIA section 71 covers the automated-decision question that sits behind it.

Worth saying plainly: the ORIS interface is in English. Your clients' conversations are not, and that asymmetry is the thing to design around — an English-speaking principal reviewing a queue of isiZulu drafts is not doing quality control, they are approving something they cannot read. Put the reviewer who speaks the language on the queue that needs them. If you are still weighing scripted automation against a supervised assistant, our comparison of chatbots and supervised AI lays out where each one belongs.

Frequently asked questions

Should we simply not use AI on non-English conversations?

That is over-correcting. Classification is useful even when imperfect, provided it routes rather than decides: use it to surface, prioritise and draft, and keep the send button with a person for anything sentiment-bearing. The risk lies in automatic action, not in analysis.

How many messages do we need to test properly?

A hundred per language is enough to see a serious gap, though not enough for a precise figure. If a language makes up a large share of your book and you cannot gather a hundred real messages in it, that itself tells you something about how well you are serving those clients.

Does POPIA stop us from using client messages to test a classifier?

Testing on your own client conversations is processing that needs a lawful basis and should sit within the purpose you told clients about. Keep the sample internal, do not send it to a third party without checking your operator agreements, and de-identify it where the test does not need names.

Our clients mostly write in English anyway. Does this matter?

Check that assumption against your own inbox rather than against the impression of your English-speaking staff. Clients often switch to English with a brokerage while writing in another language everywhere else — and the messages that matter most, about a death or a claim, are the ones most likely to come in a home language.

Who is accountable if an AI-drafted message misleads a client?

Your FSP is. Regulators treat the tool as a means of delivery, not a defence: the plain-language and fair-treatment duties attach to the licensed provider whose name appears at the top of the chat.

See ORIS in action

Shared WhatsApp inbox, client records, follow-ups and opportunities for the whole brokerage. 15-minute demo.

Book a demo
Book a demo