Microsoft’s own technical report on Phi-3 shows a 3.8 billion parameter model scoring 69 percent on the MMLU benchmark, competitive with GPT-3.5, while being small enough to run directly on a phone. That single result is why small language model fine-tuning…