3T-Parameter Kimi 3 Narrows the MMLU Gap With Opus 4.8
Moonshot AI is preparing to launch Kimi 3, a 3-trillion parameter open-weights model targeting Anthropic's Opus 4.8 performance levels.
Moonshot AI is finalizing the release of Kimi 3, a massive open-weights model expected to scale between 2 trillion and 3 trillion parameters. According to the latest industry reports, the model performs within 3 to 5 percent of Anthropic’s Opus 4.8 on core reasoning benchmarks. This release positions Moonshot to deliver the largest open-weights model to emerge from China to date.
Hardware and Architecture
The training phase for Kimi 3 utilized a large cluster of H800 and H20 GPUs combined with domestically produced accelerators. This hybrid hardware strategy allowed the company to bypass ongoing export restrictions while scaling the model to unprecedented sizes for the region.
Moonshot continues to build on the Long Context architecture established in previous Kimi iterations. The new model expands the context window significantly to process multi-modal inputs of extreme length. Prior versions supported up to 2 million Chinese characters. Kimi 3 pushes this boundary further to accommodate complex, long-running multi-agent systems operating across lengthy documents.
Benchmark Performance
Internal testing data positions Kimi 3 as a direct competitor to Western frontier models. The primary target for this release is Anthropic’s Opus 4.8.
Leaked internal benchmarks highlight the performance profile:
- MMLU: Within 3 to 5 percent of Opus 4.8.
- GSM8K: Within 3 to 5 percent of Opus 4.8.
- Language Specifics: Outperforms Opus 4.8 on Chinese-language coding tasks and domestic legal and financial document analysis.
The regional performance advantage stems from superior native linguistic tokenization. The model dedicates a larger portion of its vocabulary to Chinese tokens, reducing the compute overhead required to process localized inputs.
Release Timeline and Compliance
Moonshot AI, now valued at over $2.5 billion, is rolling out Kimi 3 in two distinct phases. An initial closed beta for enterprise partners begins in late July 2026. A public API release and an open-weight distribution under a restrictive license will follow in mid-August 2026.
Before the public rollout, the model must clear regulatory hurdles. As of July 16, Moonshot is undergoing the final rounds of security assessment by the Cyberspace Administration of China to verify output alignment with domestic content regulations.
For developers evaluating open models for enterprise deployments, Kimi 3 presents a viable high-parameter alternative to restricted Western APIs. If you are building applications that require deep reasoning over massive Chinese-language datasets, you should prepare your infrastructure to host 3-trillion parameter weights by mid-August.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
How to Benchmark Custom AI Agent Tools via Hugging Face
Learn how to evaluate open-weights models against your proprietary APIs using Hugging Face's private benchmarking framework and sandboxed environments.
32B Inkling Open Model Hits 88.4% on GSM8K via Dynamic Sparsity
Thinking Machines has released Inkling, an open-weights model family optimized for local inference, edge deployment, and task-specific reasoning.
2.4T Qwen3.8-Max Beats GPT-5.6 Sol Max in Agentic Benchmarks
Alibaba's 2.4-trillion-parameter multimodal model delivers top-tier agentic performance and aggressive $2 pricing ahead of a scheduled open weights release.
Google Vertex Adds Mirendil's Self-Evolving RFT Architecture
Mirendil secured a $100 million Google Cloud partnership to scale its self-improving RFT architecture to 50 trillion tokens and integrate with Vertex AI.
Ai2 TutorMoments Benchmark Targets AI Over-Scaffolding
The Allen Institute for AI has released TutorMoments, a 520-scenario dataset designed to evaluate whether educational LLMs provide too much help to students.