You could actually use the model's MTP head to make a ~decent prediction on what experts would be activate in future tokens and preload them
Will look into tightening this up
You could actually use the model's MTP head to make a ~decent prediction on what experts would be activate in future tokens and preload them