You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Note:i9 14900、1+13 8ge4 use 4 threads,others use the number of threads that can achieve the maximum speed
35
+
Note:
36
+
- i9 14900, 1+13 8ge4 use 4 threads, others use the number of threads that can achieve the maximum speed
37
+
- sparse: refers to leveraging the sparsity induced by the ReLU activation function to skip certain computations during the UP/DOWN calculation of each expert based on the GATE output, as well as using a predictor to perform sparse computation when calculating the lm_head
Note:lm_head sparsity is not included. If needed, please merge model_lm_head.pt into the safetensors file before executing the above commands, or directly download the GGUF file we provide.
0 commit comments