Arabic & English AI support engine

Multilingual AIRiyadh & London
Customer support team with headset working

The results

89%

Auto-resolved

4.9★

Satisfaction score

The challenge: two languages, two support cultures, one queue

A client operating support desks in both Riyadh and London faced a problem that generic chatbot vendors kept underestimating: Arabic and English support tickets aren't just the same content in two languages. They arrive with different phrasing conventions, different levels of formality, different assumptions about what "resolved" means, and different regulatory context, since Saudi consumer-protection and data-handling rules don't map one-to-one onto UK equivalents. Every off-the-shelf chatbot the client had trialed handled Arabic as an afterthought: a translation layer bolted onto an English-first model, which produced replies that were grammatically passable but culturally off in ways that Arabic-speaking customers noticed immediately and complained about.

The support team was handling over 12,000 tickets a month across both languages, split roughly evenly, with response times stretching during any volume spike because Arabic-fluent agents were a smaller pool than English-fluent ones, creating an uneven bottleneck: English queues moved fine, Arabic queues backed up. The client needed a system that treated Arabic as a first-class language, not a translation of one, and that understood enough about both regulatory environments to know when a ticket needed to escalate rather than auto-resolve.

The approach: fine-tuned models per language, not one model translated twice

Rather than building a single English model and running it through translation in both directions, we fine-tuned separate model behavior for Arabic and English on the client's own historical ticket data, so idiom, tone, and common request patterns were learned natively in each language rather than approximated through translation. The Arabic NLP layer was tuned specifically for the dialectal mix the client's actual customer base uses, which blends Modern Standard Arabic in written form with regional phrasing patterns from both Gulf and broader MENA customers, something a generic Arabic language model trained primarily on MSA text handles poorly.

Cortex AI underpins the reasoning and retrieval layer that grounds every response in the client's actual policy documentation rather than letting the model improvise on regulatory or refund questions, and the system runs on AWS, chosen for the regional infrastructure options that let ticket data for Saudi customers stay within the data-residency boundaries the client's compliance team required. A shared compliance-and-escalation layer sits behind both language models, so a ticket that touches a refund threshold, a data-access request, or anything with real regulatory weight routes to a human agent in the appropriate market regardless of which language it arrived in.

Implementation: earning trust on low-risk tickets first

The rollout deliberately started with the lowest-risk ticket categories, order-status questions, account-access resets, basic product FAQs, in both languages, and only expanded auto-resolution scope to higher-stakes categories once accuracy and customer satisfaction on the low-risk categories had been measured over several weeks of live traffic. This was a direct decision made after early internal testing showed the Arabic model performing extremely well on straightforward requests but needing more guardrails on anything touching money or personal data, exactly the kind of ticket where a culturally fluent but overconfident reply does real damage.

A constraint worked around during implementation was tone calibration: what reads as appropriately polite and deferential in Arabic customer service correspondence would read as stilted or overly formal if the same phrasing structure were applied to English tickets, and the reverse was also true. Rather than sharing a single tone template across languages, each language's model was tuned against its own set of agent-approved historical replies, which took longer to prepare than a shared template would have, but avoided the uncanny, translated feel that made competitor tools unpopular with the client's Arabic-speaking customers.

Results: 89% auto-resolved, satisfaction actually improved

At full scale, the system auto-resolves 89% of the 12,000+ monthly tickets across both languages without human intervention, with the remaining 11% escalating cleanly to a human agent already carrying the full ticket context and the reason for escalation, rather than starting the conversation over. Critically, that 89% figure held roughly equally across Arabic and English tickets, closing the gap that used to mean Arabic-speaking customers waited longer purely because of agent-pool imbalance.

Customer satisfaction, measured through the client's existing post-ticket survey, sits at 4.9 out of 5 across the combined ticket volume, a number that matters more than the resolution rate on its own, since it confirms the auto-resolved tickets aren't just closed quickly, they're closed in a way customers actually consider genuinely resolved rather than deflected.

Tech stack

Fine-Tuned LLMsArabic NLPCortex AIAWS

Want results like this for your business?

Tell us about your project. We respond within 24 hours.

Book a consultation