Most individuals know that a correlation implies “moving together.” While this definition is correct, it is rather futile without additional information. The correlation of 0.6 between two variables might reflect a robust, dependable relationship, or the actual pattern might have nothing to do with what the value suggests. It is only by distinguishing the various types of correlation that one can recognize which one applies in a particular situation.
This comprehensive guide will cover all main types of correlation, interpret the coefficient properly, identify the proper measurement technique for your variables, and understand what lies behind the value. Jump right in wherever you like since each section stands alone.
What Will I Learn?
What Is Correlation?
The correlation is the degree to which there exists a statistical relationship between two variables. You express the correlation as a numerical value, known as the correlation coefficient, ranging from -1 to +1.
Here is what these endpoints signify:
- +1: perfect positive relationship, such that when one variable increases, the second increases proportionally as well, always
- -1: perfect negative relationship, meaning when one variable increases, the other decreases proportionally
- 0: zero correlation, meaning that there is no relationship at all
It would be rare to get a perfect correlation in real-world research. Coefficients are usually distributed around the 0 point between -0.8 and +0.8, and the interpretation of these mid-point coefficients depends on knowledge of the subject matter as well as the data type.
It was Francis Galton who first came up with the idea of correlation, having published it in a paper he wrote on the topic in 1888. Galton called his method of analysis “correlation.” Later, Karl Pearson provided the mathematical definition of correlation and the Pearson coefficient has been the standard ever since.
Types of Correlation: Quick Reference
Before the detailed breakdown, here’s every major type at a glance:
| Type | Based On | Direction | Example |
| Positive | Direction of change | Both variables rise or fall together | Study hours up, exam score up |
| Negative | Direction of change | One rises as the other falls | Price up, demand down |
| Zero | Direction of change | No linear relationship | Shoe size vs. IQ |
| Linear | Shape of relationship | Change is proportional and constant | Temperature in Fahrenheit vs. Celsius |
| Non-linear (curvilinear) | Shape of relationship | Change is real but not constant | Anxiety up then performance peaks then drops |
| Simple | Number of variables | Two variables only | Rain vs. umbrella sales |
| Partial | Number of variables | Two variables, others held constant | Coffee vs. heart rate with stress controlled |
| Multiple | Number of variables | One outcome, several predictors | House price from size, location, and age |
| Perfect | Degree | Exact +1 or -1 coefficient | Extremely rare in real data |
Types of Correlation Based on Direction of Change
Positive Correlation
A positive correlation occurs when both the variables rise and fall together.
The value of this coefficient lies between 0 and +1. The nearer it gets to +1, the stronger the correlation; the closer it is to 0, the weaker the connection between the two.
Some real-life examples that are logical:
- Number of hours spent studying and exam results – the more hours you devote to studying, the better your result will be
- Spending on advertisements and the number of clients – there might be exceptions, but in general, the two variables depend on each other
- Exercises and calories consumed – those who exercise regularly consume more calories
Negative Correlation
In negative correlation, one of the variables increases while the other decreases. The range of the coefficient is between -1 and 0.
- Price and demand for a commodity: The standard example of a negative correlation, which applies to elastic commodities only
- Amount of exercise done and the resting pulse rate: More frequent exercise results in lower resting pulse rate
- Amount of sleep and workplace error rates: People who lack sleep tend to make more errors
The negative correlation does not imply inferiority to positive correlations. They are simply opposites. Therefore, a coefficient of -0.85 implies a stronger correlation than +0.4.
Zero Correlation (No Correlation)
The concept of zero correlation means that there is no linear relation existing between the two sets of data. When plotting a graph, one would see nothing but scattered data points without a trend.
What the articles generally fail to highlight is that a zero correlation means only a lack of linear correlation. It does not imply that there is no correlation whatsoever between the two sets of data points. There might be a very real, strong correlation between the two that is missed entirely by a correlation coefficient of 0.
In terms of applied statistics, this may be regarded as the most dangerous misconception. In Anscombe’s Quartet, first published back in 1973, Francis Anscombe presented four graphs with identical coefficients of correlation equal to 0.816, while showing different scatter plots.
One represents a clear linear correlation. One represents a curved relationship where Pearson r is clearly misleading. And one includes an outlier distorting the coefficient upward.
The takeaway from the Anscombe example: always plot your data.
Types of Correlation Based on Shape
Linear Correlation
Linear correlation is a correlation where the relation between two variables is in a straight line. An increase in X by one unit gives an almost equal increase or decrease in Y.
This is the basis of Pearson’s correlation coefficient. You can see this correlation by plotting the variables on a graph showing a straight line with some scatter around it.
The standard example is the conversion of Fahrenheit to Celsius. This gives an absolute linear correlation, where 1°F equals 0.556°C. That’s a perfect linear relationship (r = 1). Real data won’t be that clean, but the shape will still be roughly linear.
Non-Linear (Curvilinear) Correlation
Non-linear correlation or curvilinear refers to the situation whereby a genuine relationship exists between two variables but is not linear in nature.
One of the best examples is the Yerkes-Dodson Law. As arousal or anxiety increases, performance on tasks improves until a certain optimal point is reached and then falls sharply. In such a case, it is an inverted-U relationship, where there is a connection between the two variables, but if we use Pearson’s r, we may obtain a result close to zero.
In fact, this does not imply a lack of correlation but a poor choice of a question. If we were to conduct Pearson’s r in such cases, the result obtained would be incorrect since the question posed does not reflect the true relationship.
Types of Correlation Based on Number of Variables
Simple Correlation
Simple correlation examines only two variables at a time. What most of you have learned till now about correlation from this article belongs to simple correlation. You measure one thing against another thing and then find one coefficient.
It should be used to determine whether or not these two variables move together.
Partial Correlation
Partial correlation is a statistical technique used to find the correlation between two variables by eliminating the effect of other variables. In simpler terms, it is the method to find out the exact effect of one thing without the interference of other variables.
For example, imagine you wish to see whether drinking coffee causes a high heart rate in people. The issue here is that stress increases the drinking of coffee as well as the heart rate. Running a correlation between the two variables will give you a value; however, you will not be able to tell whether this is the actual effect caused by coffee or stress causing the increase in both of these things.
Multiple Correlation
The analysis of the multiple correlation is done by observing the relationship between one dependent variable and two or more independent variables. This type of analysis doesn’t seek to find whether X predicts Y. Instead, it tries to find out “do X, Z, and W predict Y, and how good?”
This type of relationship occurs all the time in machine learning. Simple correlation would be where house prices are predicted based on square feet alone. However, when predicting the price of a house depending on the square feet, neighborhood rating, year built, and the proximity of public transport, you are dealing with multiple correlations.
This kind of correlation is measured using the coefficient of multiple correlation (R), and not the ordinary r.
Strong vs. Weak Correlation: Reading the Coefficient Correctly
What about a correlation of 0.5—is that considered strong or weak? And again, the real truth is: it all depends on your discipline.
The most common criteria for strength of effect, put forth by Jacob Cohen in 1988, are as follows:
- Weak effect: r = 0.1
- Moderate effect: r = 0.3
- Strong effect: r = 0.5
However, they were developed with social psychology experiments in mind, in which human behavior is unpredictable and coefficients higher than 0.5 are simply unrealistic. A coefficient of 0.5 in physics or engineering would be deemed very low. An epidemiological study involving a risk factor for disease could see a coefficient of 0.2 between behavioral and clinical data as highly clinically relevant due to the sheer size of the population under study.
Thus, it is important to compare any coefficient with other, similar statistics from the same field to define its strength.
Degree of Correlation: Perfect, Zero, and Everything Between
Perfect correlation (r = +1 or -1) implies that all the points are on a straight line. The value of one variable will perfectly predict the value of the other variable without any error. This rarely happens with real data, and when it does, it is wise to ask if the variables measure the same thing.
Zero correlation (r = 0) indicates no linear correlation between the two variables. This means there is no movement in any direction of both variables relative to each other. Just to reiterate, it does not mean that they have no connection at all.
A limited degree of correlation includes everything in between, and it is here that almost all data points actually fall. A correlation coefficient of 0.65 indicates there is a fairly positive correlation between two variables, and thus, they do tend to move together to an extent.
The threshold ranges cited most often: low (0 to 0.25), moderate (0.25 to 0.75), high (0.75 to 1). Treat those as starting points, not law.
Types of Correlation Coefficients: Which Method to Use
This is where all other writings about the subject fail. Understanding that there are positive, negative, and zero correlation is great; but having the knowledge of which equation to use when dealing with real data – that’s the key.
As far as I understand, all books on statistics begin with the presentation of several correlation methods in an equal manner, which means you can choose whichever you like. The choice of an incorrect method will render all your calculations void.
So first, look at your data type:
If your data is measured on an interval or ratio scale (scores, heights, temperatures, sales figures) and it’s roughly normally distributed: use Pearson’s r. It’s the most common, most widely reported, and easiest to interpret.
If your data is ordinal (survey ratings like 1 to 5, rankings, order-based measurements) or your distribution is skewed: use Spearman’s rho. It converts values to ranks first, then measures how well those ranks correspond. This makes it more resistant to outliers and non-normal distributions.
If you have a small sample or many tied values: Kendall’s tau is often more reliable than Spearman in these cases. It counts concordant and discordant pairs rather than ranking, a subtly different approach that performs better when your dataset is small.
If one of your variables is binary (yes/no, pass/fail, clicked/didn’t click): use Point-Biserial correlation, which is a special case of Pearson applied to a continuous variable paired with a dichotomous one.
Pearson Correlation
Pearson’s r measures the linear relationship between two continuous variables. The formula:
The calculation itself will not be done by hand when actually working with the formula. However, an understanding of what is being measured will certainly come in handy: in essence, it measures the variance of the two variables against each other compared to the variance of each on its own (the quotient of the covariance of the variables and the product of their standard deviations).
Assumptions for Pearson’s r to apply to your situation: both variables must have an interval or ratio scale, linear relation, and no outliers skewing results.
Spearman Rank Correlation
The rho, otherwise known as Spearman’s correlation coefficient, was named in honor of Charles Spearman, a British psychologist and statistician who pioneered the use of the correlation in a 1904 journal article in The American Journal of Psychology. Spearman’s rho is based on rankings instead of the actual numerical data. So, instead of giving a rank of 95 to one person and 72 to another, Spearman’s would rank them at number one and two.
This makes the Spearman correlation coefficient an appropriate tool to measure relationships in survey answers, among other types of ordinal data, and even continuous data that contains extreme values or is heavily skewed since the extreme value will merely be ranked.
The formula: ρ = 1 − (6Σd²) / (n(n² − 1)), where d is the difference between each pair of ranks and n is the sample size.
Kendall’s Tau
The use of Kendall’s tau started in 1938 by Maurice Kendall. Although it is not as widely used as the others, it is more stable than Spearman’s with smaller sample sizes and data containing numerous ties.
Point-Biserial Correlation
What if you wanted to see if there was a relationship between taking a training course (pass/fail) and how well someone would perform at work later on (score)? This is an example where you would employ a point-biserial correlation; mathematically, it’s identical to Pearson’s r for this situation.
Correlation vs. Causation: The Mistake That Keeps Getting Made
Two variables being correlated says nothing about whether one caused the other. But that lesson apparently needs repeating, because people get it wrong constantly.
There are three possible explanations for any correlation you observe:
1. Direct causation — X genuinely causes Y. Smoking causes lung cancer. The correlation between them reflects a real mechanism.
2. Reverse causation: Y causes X, and not vice versa. In research suggesting that individuals suffering from depression eat more sugar, it is possible to interpret this as “sugar causes depression,” but just as likely that depression leads to an increased craving for sugar.
3. A confounding variable: A third, unobserved variable Z causes both X and Y, creating a correlation that is purely coincidental. Ice cream sales and drowning deaths increase during summer months. They are correlated because summer causes both. Ice cream does not cause drowning.
Spurious correlations, statistically real correlations between two unrelated variables, keep appearing all the time whenever one tries to find correlations in big data sets. Per capita consumption of cheese correlates with accidental deaths by getting tangled in bed sheets over several years, according to the Spurious Correlations site by Tyler Vigen, where many of such statistically real but logically meaningless correlations are documented.
Real-World Applications of Correlation
Business and Marketing
Correlation is used to see if your marketing expenditure makes any difference. Doing a correlation between weekly expenditures and weekly sign-ups would not show causality since there are so many factors that could change from one week to another – but a positive correlation coefficient consistently month after month will at least make for an interesting investigation using a controlled experiment.
When developing a model to predict churn within SaaS companies, the starting point might be to run correlations among variables associated with user behavior (logins per week, use of certain features, volume of support tickets) and then correlate those variables with churn.
Machine Learning: Feature Selection and Multicollinearity
High correlation between variables in machine learning poses a challenge rather than a desirable situation. Suppose two variables in your feature set have very high correlation, for instance, “number of rooms” and “total square footage” in case of real estate data. In that case, those two variables are providing almost similar data information, leading to multicollinearity, thus increasing the variance inflation factor (VIF) of the regression model, making it less effective.
Rule of thumb: perform a correlation analysis of all variables in your feature set before fitting a model. If you find any variable pair whose r-value is greater than 0.9, then one of them should be omitted.
[VISUAL: Example 4×4 correlation matrix heatmap — variables: income, education_years, credit_score, debt_ratio. Color scale from deep blue (-1) to deep red (+1). Shows that income and credit_score are highly correlated at r = 0.81, flagging a multicollinearity concern.]
Finance: Portfolio Diversification
Negative correlation between assets in investing is a great thing to have. When there is a correlation of -0.4 between stock A and bond B, it means that as stocks decline, bonds will go up. This reduces the volatility of the portfolio, which makes the basic principle of diversification.
The portfolio whose all assets have positive correlations provides no security at all. It can easily be understood that constructing a portfolio is all about finding assets with low or negative correlations.
Medicine and Epidemiology
Risk factors are studied primarily through correlation studies. The first correlation study regarding the link between cigarettes and lung cancer showed that those who smoked at a greater rate had more incidences of lung cancer in their country and individually. The link between smoking and lung cancer holds true in dozens of different countries and study methods.
How to Calculate Correlation in the Tools You Actually Use
Python (Pandas)
python
import pandas as pd
df = pd.DataFrame({
'study_hours': [2, 4, 6, 8, 10, 3, 7],
'exam_score': [55, 65, 72, 80, 91, 60, 78]
})
# Pearson (default)
print(df.corr())
# Spearman
print(df.corr(method='spearman'))
The .corr() method returns a correlation matrix: every variable correlated against every other variable. The diagonal is always 1.0, since a variable perfectly correlates with itself.
Excel
Use the CORREL function: =CORREL(A2:A20, B2:B20). Select your first column of data, select your second, and you get a coefficient. No setup, no additional libraries.
R
r
# Pearson
cor(df$study_hours, df$exam_score)
# Spearman
cor(df$study_hours, df$exam_score, method = "spearman")
For a full correlation matrix across multiple columns: cor(df).
Limitations of Correlation Analysis
Correlation is genuinely useful. But it has real limits, and knowing them is more valuable than knowing how to calculate it.
It only measures linear relationships. Anscombe’s Quartet proves this: the same coefficient can come from a straight-line relationship, a perfect curve, or an outlier-driven distortion. Always check the scatter plot — this can’t be said enough.
Outliers can badly distort Pearson’s r. A single extreme data point can pull the coefficient toward 1 or -1 and make a weak relationship look strong. If your data has outliers, Spearman is more reliable.
Correlation doesn’t tell you the size of the effect. Knowing that X and Y correlate at r = 0.4 doesn’t tell you how much Y changes for a given change in X. That’s what regression gives you; correlation alone won’t get you there.
Restricted range compresses correlation. If you study marathon finishers and try to correlate training miles per week with finish time, you’ll find a weaker correlation than you’d get from the general running population — because your sample only includes people who already train heavily. Cutting off the range of a variable compresses the relationship artificially.
More data doesn’t automatically mean more accuracy. With a very large sample, you can find statistically significant correlations that are completely trivial in practice. An r = 0.05 might reach p < 0.001 with 10,000 observations; but a relationship explaining 0.25% of the variance matters for almost nothing.
Frequently Asked Questions
Q1. What are the main types of correlation?
Ans. Correlation can be classified using three methods. One way is through the direction of correlation, which could be positive correlation, where two variables increase or decrease together; negative correlation, where one variable increases as the other decreases; or zero correlation, where there is no linear correlation between two variables. Another way is by means of a form of correlation, which can either be linear or non-linear/curvilinear.
Q2. What does a correlation coefficient of 0.7 mean?
Ans. A value of 0.7 suggests a moderately strong correlation: the two variables move together relatively well, but not in absolute agreement. Using Jacob Cohen’s criteria, anything over 0.5 is considered a large effect size; 0.7 is definitely one. This aside, context matters: 0.7 would be very good for social sciences but possibly too small for engineering.
Q3. What is the difference between positive and negative correlation?
Ans. In positive correlation, both variables will rise or fall simultaneously. In other words, if one variable increases in value, the second variable will tend to increase in value. If one decreases, the second one will follow suit. On the other hand, in negative correlation, both variables move in opposite directions.
Q4. Can correlation prove causation?
Ans. No. Correlation tells us that there’s a connection between two variables, but doesn’t indicate which one causes the other or if there’s another cause that affects both of them. Causation can be established through an experiment or using some particular method of causal inference.
Q5. What is the difference between Pearson and Spearman correlation?
Ans. Pearson correlation reflects the linear correlation between two variables that are continuous and normally distributed. The Spearman rank correlation method ranks the data points and then assesses whether there is a monotonic relationship, which means that it consistently increases or decreases but does not need to follow a linear pattern. Use Pearson for clean, continuous data. Use Spearman if there are outliers or skewness.
Q6. What is zero correlation?
Ans. Zero correlation (r=0) implies that there is no linear relationship between the two variables. However, this does not imply that there is no relationship at all; in a case where there is a curvilinear relationship, zero correlation (r=0) can still result despite high interdependency. Graph your data.
Q7. What is partial correlation?
Ans. Partial correlation is an analysis that seeks the correlation between two variables after accounting for the effect(s) of another variable(s). It is used when one needs to establish the direct correlation between two variables when there are interfering variables present.
The Number Is the Starting Point, Not the Answer
The value of the correlation coefficient shows us two very important things: its direction and magnitude. That’s pretty neat stuff, indeed. However, this is only the beginning of your analysis; don’t end here!
It turns out that those researchers who blindly believed in r = 0.816 in all of Anscombe’s data sets without even plotting the number. Well, they made a serious blunder which may be fixed in around thirty seconds. Compute the correlation. Plot the data. Think about the graph’s pattern. Consider other possible factors that might influence the data.
Step by step.