Voho Saudi Chat
4B
An Arabic assistant that answers in spoken Saudi, not newsreader Arabic.
Sounds more Saudi than Alibaba’s Qwen. Says it in a fifth of the words.
Replies that read as Saudi. Higher wins.
27.5 points ahead of Alibaba Qwen
vs Qwen3-4B-Instruct, 400 held-out questions; 6.4 words a reply vs 30.2.
- 89.8%
- Replies read as Gulf, from 62.3%
- 6.4
- Words per reply, from 30.2
- 2.5 GB
- Q4 download, runs on a laptop
- Apache 2.0
- Commercial use allowed
What Voho Saudi Chat 4B is for
Ask a general 4B model a question in Saudi and the reply comes back as Gulf Arabic only 62% of the time, with Levantine and Egyptian leaking into the rest, and at 30 words: three times what a person says in one turn on the phone.
Voho Saudi Chat 4B replies the way a person does: consistently Najdi, at phone-call length. It sits between speech-to-text and text-to-speech in a voice agent, deciding what to say.
Why teams pick it
Each reason is a number from the evaluation below, against the models named there.
It sounds Saudi
89.8% of its replies are classified Gulf by an independent dialect classifier, against 62.3% for the model it started from and 94.5% for the reference replies themselves.
It answers like a phone call
6.4 words a reply instead of 30.2. A voice agent that speaks paragraphs loses the caller; this one takes a turn and stops.
You can use it commercially
Apache 2.0 end to end: the base model and every dataset in the mix. Nothing about the licence stops a business adopting it.
It runs where you need it
4 billion parameters. The Q4 build is 2.5 GB and runs locally with Ollama or llama.cpp; the full weights run on one GPU.
What it is, in one table
- Parameters
- 4.0B
- Base model
- Qwen3-4B-Instruct-2507 (Apache 2.0)
- Input / output
- Arabic text in, spoken Najdi Arabic out
- Reply length
- About 6 words, one spoken turn
- Full weights
- 8.0 GB, safetensors
- Smallest build
- 2.5 GB, GGUF Q4_K_M
- Licence
- Apache 2.0, commercial use allowed
Measured on held-out data it never trained on
Share of replies an independent dialect classifier (MARBERTv2) assigns to each dialect, on 400 held-out questions. Higher Gulf and lower Modern Standard Arabic is the goal.
| Classified as | Held-out reference replies | Qwen3-4B-Instruct-2507 | Voho Saudi Chat 4B |
|---|---|---|---|
| Gulf (Saudi) | 94.5% | 62.3% | 89.8% |
| Modern Standard Arabic | 0.2% | 1.5% | 0.8% |
| Egyptian | 0.2% | 8.2% | 0.0% |
| Levantine | 4.0% | 19.8% | 7.8% |
| Maghrebi | 1.0% | 8.2% | 1.8% |
The reference column is the ceiling: what the held-out replies themselves score. The classifier does not call every Saudi sentence Gulf, so a perfect model would not reach 100%.
| Reference | Base model | Voho Saudi Chat 4B | |
|---|---|---|---|
| chrF++ against held-out replies | — | 13.0 | 12.0 |
| Mean reply length (words) | 10.7 | 30.2 | 6.4 |
chrF++ is flat against the base: the model says the right kind of thing in the right register, not the same words as the reference, which a dialect classifier measures and a string-overlap metric cannot.
Real outputs
Questions from the held-out set.
| Asked | Voho Saudi Chat 4B |
|---|---|
| هلا والله يا أبو ناصر، تدري أنا أبي أسجل حقوق الملكية؟ | هلا بك، أبشر. عندك ورقة من المكتب؟ |
| وش رايك في ذا التلفزيون؟ ودنا نشتري واحد عشان نتابع النهائي. | والله فكرة، بس أنا ما أحب أشتري شي جديد. |
| الطالب: يبه، ودي اروح محل الحيوانات اللي في التحلية ونشتري قطه. | الله يعينك، بس وين تبي تروح؟ |
| وش رايك في التصميم الجديد للمسجد؟ | حلو، بس أخاف يزعل الناس. |
Every file, its size, and where it runs
| Build | File | Download | Runs on |
|---|---|---|---|
| Full weights | model.safetensors | 8.0 GB | Transformers on one GPU |
| GGUF Q8_0 | voho-saudi-chat-4b-Q8_0.gguf | 4.3 GB | llama.cpp, Ollama; near-lossless |
| GGUF Q6_K | voho-saudi-chat-4b-Q6_K.gguf | 3.3 GB | llama.cpp, Ollama |
| GGUF Q5_K_M | voho-saudi-chat-4b-Q5_K_M.gguf | 2.9 GB | llama.cpp, Ollama |
| GGUF Q4_K_M | voho-saudi-chat-4b-Q4_K_M.gguf | 2.5 GB | llama.cpp, Ollama; laptop CPU or GPU |
Sizes from Hugging Face's file listing. Speed depends on your hardware; no benchmark is published yet.
Run it in a few lines
Run it locally with Ollama
ollama run hf.co/VohoAI/voho-saudi-chat-4b-GGUF:Q4_K_MPython, with Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "VohoAI/voho-saudi-chat-4b"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")
SYSTEM = "أنت مساعد صوتي سعودي. رد باللهجة النجدية كما يتكلم الناس في الرياض، بجمل قصيرة مثل المكالمة الهاتفية، بدون رموز ولا تنسيق ولا شرح زائد."
messages = [{"role": "system", "content": SYSTEM},
{"role": "user", "content": "أبي أحجز موعد بكرة الصبح، فيه وقت فاضي؟"}]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=128, do_sample=False)
print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))How it was trained
- Base model: Qwen3-4B-Instruct-2507 (Apache 2.0), LoRA r=32 on all attention and MLP projections, 2 epochs, one NVIDIA L4. Loss on assistant turns only.
- Data: 7,601 Voho service-call dialogues across eight enterprise verticals (oil and gas, utilities, telecom, banking, government, healthcare, logistics, facilities), 5,551 Voho everyday conversations, 202 rows from 2A2I/Arabic_Aya and 1,181 from arbml/CIDAR, all Apache 2.0.
- Every dialogue had to pass a Najdi lexicon filter and the MARBERTv2 dialect classifier, which took no part in training. Anything with markdown, tables or a reply longer than a spoken turn was dropped. The dialogue data is published at VohoAI/voho-saudi-dialogues.
- Keep the system prompt from the example: it is in every training example, and the dialect is noticeably weaker without it.
Where you can use it
Open weights on Hugging Face
Apache 2.0, commercial use allowed
Commercial use
Allowed under Apache 2.0
CPU and GPU
Transformers, llama.cpp and Ollama
Offline, on your own servers
Nothing calls out once it is downloaded
Limitations
- Najdi, mostly. Hijazi and Khaleeji replies drift toward Najdi or Modern Standard Arabic.
- The classifier’s Gulf class covers Saudi, the UAE and Kuwait: a high score means "reads as Gulf", not "reads as Riyadh".
- Not a knowledge model. It is tuned for register and voice; ground it with retrieval for facts.
- Writes without diacritics.
For production Saudi Arabic voice, including in-Kingdom and on-premise deployment, use the Voho API.
Hear it, decide what to say, say it the Saudi way.
The three Voho models are one voice agent: speech to text, the reply, and the reply rewritten the way a Saudi would say it.
- 01 · HearVoho Saudi STT SmallSpeech to text
- 02 · DecideVoho Saudi Chat 4BArabic assistant, spoken Saudi
- 03 · SayVoho Saudi Speak 0.6BFormal Arabic to spoken Saudi
Other models: Voho Saudi STT Small · Voho Saudi Speak 0.6B · All models