Microsoft extends the Phi-4 line into multimodal reasoning: a 15B vision-language model built on Phi-4-reasoning (14B dense, 40 layers, hidden 5120) with a SigLIP vision encoder (phi4-siglip), 32K context, released on HuggingFace under MIT on August 31, 2026. The card targets chain-of-thought visual reasoning, math, OCR, and notably GUI grounding / computer use.

Model-card benchmarks: AI2D 84.8, ChartQA 83.3, MathVista (mini) 75.2, MMMU 54.3, OCRBench 76.0. No accompanying technical report at release (the Phi small-model reports typically follow); not yet scored on the Artificial Analysis Intelligence Index as of 2026-09-01.

Model Details

Architecture DENSE
Parameters 15B
Context window 32,768
License MIT
Base model phi-4

Benchmark Scores

Benchmark Score Mode
AI2D 84.8 —
ChartQA 83.3 —
MathVista (mini) 75.2 —
MMMU 54.3 —
OCRBench 76.0 —
multimodalreasoningopen-weightsmall-model

Related