Extrinsic evaluation: Evaluate the performance of a language model by embedding it in an application and measure how much the application improves.
Intrinsic evaluation: Measures the quality of a model independent of any application.
Perplexity is the inverse probability of the test set, normalized by the number of words: