Thinking Machines Lab released Inkling, a 975B-parameter Mixture-of-Experts model with open weights and a 1M-token context window. The model activates 41B parameters per token and was pretrained on 45 trillion tokens across text, images, audio, and video. Its MoE architecture largely follows DeepSeek-V3. A smaller variant, Inkling-Small, matches the larger model on many benchmarks and will release after testing. Inkling supports fine-tuning on Tinker and is deployable via multiple runtimes.
No score is assigned. Sources and their independence are shown in the citation chain below.