Pearson correlation - prerequisites

Опубликовано: 15 Июнь 2026
на канале: Statistik am PC
75,834
468

// Pearson Correlation - Prerequisites //

A bivariate Pearson correlation requires certain prerequisites that are all too often ignored. Admittedly, they are usually met anyway, but it definitely doesn't hurt to have a feel for the prerequisites when calculating and subsequently interpreting a Pearson correlation coefficient.

In particular, a metric scale level for both variables to be correlated is a basic prerequisite for Pearson. In addition, the relationship should be linear or quasi-linear and not follow a U or inverted U shape. Furthermore, there must be no extreme outliers. And last but not least, a bivariate normal distribution is required.

I show and explain the prerequisites in detail in the video.

#####
🎥 Other important videos I mention:
🎥 Bivariate Correlation (Scale Levels):    • Bivariate Korrelation in SPSS - Skalennive...  
🎥 Graphical Diagnosis of Outliers:    • Ausreißer in SPSS grafisch diagnostizieren...  
🎥 Interpreting Boxplots:    • Boxplot interpretieren (Kastendiagramm int...  

Note:
========
In an earlier version of the video, the central limit theorem (CLT) was not explained precisely enough. As the sample size increases, the distribution of sample parameters, e.g., means, increasingly approaches a normal distribution—regardless of the shape of the population distribution.

Here is an excerpt from Field, A. (2018) Discovering Statistics Using SPSS, p. 235 f.:
“1. For confidence intervals around a parameter estimate (e.g., the mean, or a b in equation (2.4)) to be accurate, that estimate must come from a normal sampling
distribution. The central limit theorem tells us that in large samples, the estimate will have come from a normal distribution regardless of what the sample or population data look like. Therefore, if we are interested in computing confidence intervals then we don’t need to worry about the assumption of normality if our sample is large enough.
2. For significance tests of models to be accurate the sampling distribution of what’s being tested must be normal. Again, the central limit theorem tells us that in large samples this will be true no matter what the shape of the population. Therefore, the shape of our data shouldn’t affect significance tests provided our sample is large enough. However, the extent to which test statistics perform as they should do in large samples varies across different test statistics, and we will deal with these idiosyncratic issues in the appropriate chapter.
3. For the estimates of model parameters (the b's in equation (2.4)) to be optimal (using the method of least squares) the residuals in the population must be normally distributed. The method of least squares will always give you an estimate of the model parameters that minimizes errors, so in that sense you don’t need to assume normality of anything to fit a linear model and estimate the parameters that define it (Gelman & Hill, 2007). However, there are other methods for estimating model parameters, and if you happen to have normally distributed errors then the estimates that you obtained using the method of least squares will have less error than the estimates you would have got using any of these other methods.”

Additionally, I highly recommend the article on this topic: Lumley, T., Diehr, P., Emerson, S., & Chen, L. (2002). The importance of the normality assumption in large public health data sets. Annual review of public health, 23(1), 151-169.

If you have any questions or suggestions regarding the assumptions of the Pearson correlation coefficient, please use the comments function. You can indicate whether you found the video helpful by giving it a thumbs up or down. #statistikampc

⭐Become a channel member⭐:
======================
   / @statistikampc_bjoernwalther  

Support the channel? 🙌🏼
===================
Paypal donation: https://www.paypal.com/paypalme/Bjoer...
Amazon affiliate link: http://amzn.to/2iBFeG9

Thank you for your support! ♥