https://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-models 1b to 4b active parameters, full weight does not load to the DRAM.