Moonshot AI Releases Kimi K3 Architecture with 2.8T Parameters

Moonshot AI has released the complete technical architecture for Kimi K3, a 2.8 trillion-parameter multimodal model with a one-million-token context window, but training recipes remain closed.

Last Updated: July 28, 2026 Editorial Process
Editorial Process
See more of Inside AI's trusted news by adding us as a preferred source on Google.
AI neural network visualization
By Inside AI Editorial Team Published on: July 28, 2026

July 28, 2026, (Inside AI) — Moonshot AI has released the full technical architecture of its Kimi K3 model, a 2.8 trillion-parameter multimodal mixture-of-experts system with a one-million-token context window. The 1.56 TB repository on Hugging Face includes model weights, detailed documentation, attention kernels, MoE communication code, and agent deployment tools.

The release, dated July 27, exposes how the Chinese startup engineered a frontier-scale competitor to Western models at lower computational cost. Kimi K3 activates only 104 billion parameters per inference call by using 16 out of 896 experts per token, a sparse routing strategy that slashes compute while preserving capability.

The architecture spans 93 layers and uses a 160K-token vocabulary. Native vision support comes from MoonViT-V2, a 401-million-parameter vision encoder. The model’s hybrid attention system combines 69 Kimi Delta Attention layers with 24 Gated Mixture-of-Latents layers, balancing long-range dependency tracking and efficiency.

Moonshot AI applied quantization-aware training with low-precision MXFP4 weights and MXFP8 activations. The company claims 2.5 times better scaling efficiency over its predecessor Kimi K2, attributing gains to Kimi Delta Attention, Attention Residuals, a sparser MoE configuration, and revised training recipes. However, those recipes remain undisclosed.

Efficiency Claims Invite Scrutiny

The 1,048,576-token context window places Kimi K3 among elite extended-context models, but independent verification of efficiency gains is pending. The full repository’s size presents a practical barrier: local experimentation demands substantial GPU memory, storage bandwidth, and accelerator interconnect capacity.

Researchers have noted that sparse MoE models often face challenges with load balancing and training stability. A 2021 study on Switch Transformers highlighted how uneven expert utilization can degrade performance, a risk that may apply to Kimi K3’s 896 experts. Without open training recipes, the community cannot assess how Moonshot AI mitigated these issues.

The release of infrastructure components for attention kernels and MoE communication is a pragmatic move. These tools lower the barrier for teams building production systems, but the hardware requirements remain steep. Quantization-aware training with MXFP formats aligns with industry trends toward low-precision compute, though its real-world impact on inference quality needs benchmarking.

Open Weights, Closed Recipes

Moonshot AI’s decision to release weights and documentation but withhold training recipes mirrors a growing tension in open-source AI. While the weights enable fine-tuning and deployment, the lack of training transparency limits reproducibility. This contrasts with models like OLMoE, where fully open data and code allow complete verification.

The Kimi K3 release also includes large-scale agent deployment tools, signaling Moonshot AI’s ambition beyond raw model performance. Yet, the practical deployment challenges remain formidable. Running a 1.56 TB model locally is not feasible for most research labs, potentially concentrating power among well-resourced entities.

More from Inside AI

  • AI Hardware & Infrastructure

    Nasdaq Futures Fall on AI Chip Worries Ahead of Pivotal Earnings

    July 28, 2026
  • AI In Business

    InfiMaker Launches AI-Powered Desktop 5-Axis CNC Machine

    July 28, 2026
  • AI In Business

    AI Sell-Off Deepens as Chip Stocks Slump, Kospi Drops 10%

    July 28, 2026
  • Machine Learning

    Moonshot AI Releases Kimi K3 Architecture with 2.8T Parameters

    July 28, 2026
  • AI Policy & Regulation

    Anthropic CEO Defends Open-Weight AI But Rejects Nvidia’s Letter

    July 28, 2026
  • AI Safety

    AI-Generated Doctors Spread Dangerous Medical Myths on TikTok

    July 27, 2026
  • Cybersecurity AI

    Hugging Face CEO Demands Transparency After OpenAI Agent Cyber Attack

    July 27, 2026
  • Generative AI

    Anthropic Launches Claude Opus 5 with Autonomous Reasoning at Stable Cost

    July 27, 2026

Never Miss a Breakthrough

Join 50,000+ readers who get our daily AI intelligence briefing. No fluff, just what matters.

Inside AI is an independent publication covering artificial intelligence news, machine learning research, and the tools shaping the future of technology. No hype. Just what's happening in the AI world.

Topics

  • Artificial Intelligence
  • Machine Learning
  • Generative AI
  • Agentic AI
  • Vibe Coding
  • Prompt Engineering
  • AI Tools & Reviews (Coming soon)

Company

  • Editorial Standards
  • Privacy Policy
  • Terms of Service
  • Contact
  • About Us

Others

  • Press Releases

© 2026 Inside AI. All rights reserved.

Designed by Blue Flare Digital