Saudi Arabian artificial intelligence company HUMAIN has introduced humain-m3, an Arabic-focused large language model developed by Chinese AI company MiniMax. Announced at the LEAP technology conference, the model is initially being offered as a research and evaluation preview through HUMAIN Node, the company’s platform for accessing and testing AI models.
Built on the MiniMax-M3 model family, humain-m3 has approximately 428 billion parameters and uses a mixture-of-experts architecture. Such systems contain multiple specialized components but activate only part of the network when processing an input. MiniMax’s documentation for the underlying M3 model says it activates about 23 billion of its 428 billion parameters at a time, an approach intended to reduce the computing resources required for each response.
HUMAIN said the model underwent additional pretraining using more than one trillion tokens of “Arabic-native content.” A token is a unit of text processed by an AI system and can represent a word, part of a word or a punctuation mark.
Tareq Amin, CEO of HUMAIN, said, “Arabic is spoken by hundreds of millions of people, yet it remains significantly underrepresented at the frontier of artificial intelligence. With humain-m3, we are investing in changing that. And through HUMAIN Node, we are making that intelligence accessible so developers, researchers and innovators can experiment with it, build on it and create the next generation of Arabic AI experiences.”
Arabic Benchmark Performance
Humain-m3 recorded the highest average score among the frontier models evaluated across seven public Arabic benchmarks. The company said the assessments covered Arabic-language understanding and reasoning.
Public benchmarks provide researchers with a consistent way to compare how language models perform across defined tasks. Results can help identify strengths in areas such as knowledge, comprehension and reasoning while giving developers a reference point for further evaluation during the model’s research-preview period.
Arabic-language evaluation continues to expand through initiatives such as the Open Arabic LLM Leaderboard. One of its components, Native Arabic MMLU, contains nearly 15,000 multiple-choice questions spanning 40 subjects. The AlGhafa benchmark was also developed specifically to evaluate Arabic language models using a collection of multiple-choice tasks.
Access, Licensing and Saudi Arabia’s AI Strategy
Researchers and developers can currently reach humain-m3 through HUMAIN Node, rather than downloading and operating the model independently. HUMAIN said it expects to release the model weights under the MiniMax Community License after completing additional safety training and alignment work, with the release targeted for October 2026.
An open-weight release would allow qualified users to inspect, adapt and deploy the model on their own infrastructure, subject to the license. Open weights are not necessarily equivalent to fully open-source development: training data, source code and detailed training methods may remain unavailable. Running a model with hundreds of billions of total parameters would also require substantial computing and memory capacity, even though its mixture-of-experts design activates a much smaller portion of the network for each token.
The launch broadens HUMAIN’s model portfolio beyond ALLAM, its existing Arabic-language model family. Saudi Arabia’s Public Investment Fund launched HUMAIN in May 2025 to operate across data centers, cloud platforms, AI models and applications.
Commercial Relevance Extends Beyond Chatbots
Improved Arabic-language processing could have practical uses in banking, insurance and other regulated industries, where errors involving terminology, customer instructions or legal documents may carry significant consequences. Potential applications range from document search and customer-service assistance to internal knowledge management. Actual adoption would still depend on accuracy, privacy, cybersecurity and regulatory assessments conducted for each use case.
Regional organizations also have to account for the linguistic complexity of the Arabic-speaking market. Modern Standard Arabic is widely used in formal communication, while customers may speak or write in dialects that vary considerably across the Gulf, North Africa and the Levant. A model that performs strongly on academic multiple-choice questions may not show the same consistency when interpreting informal messages, mixed Arabic-English text or specialized financial language.
humain-m3’s research preview provides developers with an opportunity to test those gaps before a wider release. More meaningful conclusions will require independent benchmark reproduction, real-world evaluations and clearer information about the training corpus and safety procedures.




