Home / News / MBZUAI and G42 Introduce K2-Think: Redefining AI Reasoning Efficiency Standards with 32 Billion Parameters
MBZUAI

MBZUAI and G42 Introduce K2-Think: Redefining AI Reasoning Efficiency Standards with 32 Billion Parameters

Sep 14, 20251 min read
MBZUAI and G42 Introduce K2-Think: Redefining AI Reasoning Efficiency Standards with 32 Billion Parameters

Abstract

The Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) Foundational Models Institute, in collaboration with G42, jointly released K2-Think, an open-source AI inference system, on September 9. Despite using only 32 billion parameters, the model's performance in mathematics, code, and scientific reasoning rivals that of large models with up to 20 times more parameters. Across various mathematical benchmarks, K2-Think scored 90.8 points on AIME 2024, 81.2 points on AIME 2025, and 73.8 points on HMMT 2025, achieving an inference speed of up to 2,000 tokens per second on Cerebras hardware.

Technical Breakthrough: Small Model, Big Capabilities

K2-Think is built upon six key technical pillars: long-chain-of-thought supervised fine-tuning, verifiable reward reinforcement learning, agentic planning, test-time augmentation, speculative decoding, and inference-optimized hardware. Challenging the traditional "bigger is better" paradigm, K2-Think demonstrates a significant breakthrough in parameter efficiency through intelligent design.

The system is built on the Qwen2.5 foundational model, leveraging advanced post-training and test-time computation techniques to enable smaller models to compete at the highest levels. During training, the model's performance on mathematical benchmarks rapidly improved within the initial one-third of the training phase (approximately 0.5 epochs), achieving a 79.3% pass rate on AIME 2024 and a 72.1% pass rate on AIME 2025.

Outstanding Performance

K2-Think demonstrates exceptional performance across multiple domains:

  • Mathematical Reasoning: Scored 60.7 points on OMNI-MATH-HARD, leading all open-source models in mathematical benchmarks.
  • Code Generation: Scored 64.0 points on LiveCodeBench v5.
  • Scientific Reasoning: Scored 71.1 points on GPQA-Diamond.

True Open-Source Transparency

Unlike many "open-source" AI models that only release model weights, K2-Think achieves full open-source transparency, including training data, parameter weights, deployment software code, and test-time optimization techniques. This transparency ensures that every step of the model's learning and reasoning can be studied, reproduced, and extended by the global research community.

Ultra-High-Speed Inference Capability

K2-Think achieves an unprecedented throughput of 2,000 tokens per second on the Cerebras wafer-scale inference-optimized computing platform. In comparison, Google's Gemini 2.5 Flash performed at 258 tokens per second on the third-party AI performance measurement website Artificial Analysis, significantly lower than K2-Think's speed.

Leadership Statements

His Excellency Khaldoon Khalifa Al Mubarak, Chairman of the MBZUAI Board of Trustees and member of the Advanced Technology Research Council, stated: "The new global benchmark set by K2-Think underscores the pioneering excellence of the MBZUAI Foundational Models Institute's initiative, serving as a fast track for global collaboration and cutting-edge research."

Peng Xiao, Group CEO of G42, commented: "K2-Think has shifted the AI inference paradigm from 'bigger is better' to 'smarter is better'."

Professor Eric Xing, President and University Professor at MBZUAI, said: "K2-Think represents a significant advancement for the global AI research and development community. By making these advancements available within a fully transparent framework, we are ushering in a new era of cost-effective, reproducible, and responsible AI."

Model Lineage and Future Outlook

K2-Think builds upon a growing family of UAE-developed open-source models, including Jais (the world's most advanced Arabic LLM), NANDA (Hindi), and SHERKALA (Kazakh), and continues the pioneering legacy of K2-65B, the world's first fully reproducible open-source foundational model released in 2024.

The model is currently available at https://www.k2think.ai/ and on the Hugging Face platform. MBZUAI plans to build future models based on the K2-Think recipe, including adapted versions for healthcare and genomics.

Technical Specifications

  • Parameter Size: 32 billion parameters
  • Base Model: Qwen2.5-32B
  • License: Apache 2.0
  • Inference Speed: Up to 2,000 tokens per second (Cerebras hardware)
  • Deployment Platforms: k2think.ai, Hugging Face

Conclusion

This release marks a significant milestone for the UAE in the global AI landscape, demonstrating a new path to achieving AI breakthroughs through intelligent engineering design rather than solely relying on computational scale.

MBZUAI