Hada News
Abu Dhabi's TII launches Falcon-Emirati dialect model alongside Arabic speech-to-text and OCR models
The Technology Innovation Institute on 6 October launched Falcon-Emirati, a 7-billion-parameter model built for Emirati Arabic, together with Falcon-ASR, which transcribes speech in six languages with word-level timing, and Falcon-OCR-Arabic for pulling text from images and documents.
Abu Dhabi's Technology Innovation Institute (TII), the applied research arm of the Advanced Technology Research Council (ATRC), launched three AI models on 6 October focused on Arabic, and on the Emirati dialect in particular.
Falcon-Emirati is a 7-billion-parameter language model built on TII's Falcon-H1-Arabic. TII says it was trained on native Emirati content, Modern Standard Arabic material about Emirati culture and heritage, and synthetic data. It scored 84.83% on Alyah, a native Emirati Arabic benchmark covering everyday language, figurative expressions, heritage knowledge and poetry. TII says that beats every Arabic and multilingual open-source model it tested. The model will be offered through a Falcon Chat platform.
Falcon-ASR is a compact 1.6-billion-parameter speech-recognition model. It turns spoken Emirati Arabic, Modern Standard Arabic, English, French, Spanish and Portuguese into text and records when each word is spoken. On Emirati speech, TII says, it outperformed a 30-billion-parameter multimodal model. Falcon-OCR-Arabic is a lightweight model that extracts Arabic text, tables and formulas from images and documents.
TII lists subtitling, meeting and interview transcription, accessibility tools, archive digitisation and multilingual customer service as likely uses. ATRC secretary general Faisal Al Bannai framed the launch as keeping Emirati language part of the country's technology future, and TII chief executive Najwa Aaraj said sovereign AI has to reflect how people actually speak.
Why it matters: most Arabic AI is trained mainly on Modern Standard Arabic and often misses the meaning of dialect. Gulf newsrooms, broadcasters, podcasters and public services have long needed tools that handle everyday speech. A small speech model with word-level timing is directly useful for subtitling and transcribing Arabic video and audio, and an Arabic OCR model helps digitise print archives.
Sources: TII announcement, 6 Oct 2026; TII technical blog on Hugging Face, 6 Oct 2026.
Watch next: whether TII publishes model weights and licence terms, independent tests on the Alyah benchmark, and take-up by UAE broadcasters and government services.