Internal Validity vs. External Validity: What’s the Difference?

We often hear that this study shows X therapy works. Or that research finds Y therapy helps with anxiety. Headlines often sensationalize research, sometimes going far past what the researchers actually found or said themselves. Understanding internal and external validity can help us understand the study more clearly.

We might investigate if the study was designed well, if the treatment actually caused the outcome, if there were other variables or explanations, and if the findings apply to populations outside the study. The two pieces of internal and external validity help us clarify these questions.

  • Internal Validity - How confident are we that the study’s findings are truly what they say? Are the results actually caused by what researchers believe caused them?

  • External Validity - Do these findings apply to populations outside the study? Do the results translate to other people, settings, and situations outside the study parameters?

internal validity vs external validity

What is Internal Validity?

The central question of internal validity is whether or not the study proved a cause-and-effect relationship without bias, outside interference, or other variables. For example, there’s the famous bee and ice cream example of causation. Bee stings and ice cream purchases are heavily correlated. Does this prove cause-and-effect? No. There’s a third variable in the weather and season.

Internal validity tells us whether or not a study really is sound. With internal validity, we can confidently say a treatment caused the measured effect. Without internal validity, we are left making claims that aren’t necessarily rooted in evidence.

Common Threats to Internal Validity

There are a few common threats to internal validity. This is by no means an exhaustive list. Many aspects of a study can change the validity. We’ll use the example of a hypothetical study that found CBT can reduce anxiety.

Confounding Variables

This is the bee and ice cream example. Sometimes, outside variables affect a study. For example, a study finds that CBT improves symptoms of anxiety. But, participants also were encouraged to start exercising during treatment. Can we say with certainty that the improvement in anxiety was from CBT, or might it be from the exercise?

Selection Bias

There are many forms of selection bias, including sampling bias, self-selection bias, survivorship bias, and more. A selection bias is an error in which the participants are selected in a way that decreases the internal validity of the study. For example, if you look for volunteers for the CBT and anxiety study, you have the self-selection bias. The group you have for your study is perhaps more motivated than the general population, skewing results.

History

History threat is when something happens during the study that is often out of the researchers’ control. Let’s say you are doing a study on CBT and anxiety. During that time, there is a worldwide pandemic. Anxiety levels may increase. The outside circumstances impact results in a significant way.

Maturation

Maturation, as the name suggests, has to do with the biological, physchological, and physical changes people experience over time. For example, you might do a long-term study of anxiety. Those in their 20’s have high levels. After 15 years, you return to the participants and find they (now in their 30’s) have much less anxiety. This may not be from CBT; it may be from simply aging.

Regression to the Mean

Some people may begin a study while symptoms are unusually severe (or absent). Over time, there is a regression to the mean, or typical for them. The person may have more or less anxiety because of this regression to the mean, and not the CBT treatment.

Attrition

The rate of attrition is simply how many people drop out of the study or don’t complete the treatment protocol. This is somewhat similar to a selection bias. With attrition, you are losing important data points and getting an incomplete picture.

Measurement Effects

Finally, measurement effects are changes in results due to testing. There are various ways this can happen. It may be poor question structure, early questions influencing later questions, repeated testing, or the Hawthorne Effect.

What is External Validity?

External validity, on the other hand, is whether or not the study can be applied outside the context of the study. For example, if you find CBT is effective for anxiety, but you only test 18-22 year old male college students at an elite university, you may not find these results useful in the general population.

Populations

The people of a study are a major factor in external validity. A significant amount of psychology research is done on college students, as they are available to researchers at universities. Furthermore, psychology research has historically been heavily male-biased. For decades, research really focused on white, college-educated males.

Obviously this presents a problem in external validity. CBT may help with anxiety in this population, but can we say for certain it’s likely to help a middle-aged Asian woman? A study needs a sample representative of the population it is hoping to treat. If we want to show CBT is useful for anxiety in general, we need a study population representative of the general population.

Settings

The setting is exactly as it sounds. Let’s take the example of a university health clinic. A therapy intervention may work well in this setting, but does that mean it is going to work in a busy clinic in a city? What about in a hospital? What works in a lab or controlled environment may not work in general.

A good example of this is the Oxytocin and Trust studies. In 2005, researchers found that intranasal oxytocin increased trust of others more than those who were given placebo in lab experiments. Later studies tried to replicate, but outside the setting of the lab, the effect was nonexistent or highly context-dependent. This shows a lack of external validity.

Circumstances

Circumstances can include many things including time, task/situation, and implementation. For this example, let’s go back to the oxytocin study and speak hypothetically. During that study, participants were asked to send money to an anonymous partner in a game of trust. But as it was fake money, the circumstances are different than when we deal with real money. These are different circumstances, and can decrease external validity.

Internal Validity vs. External Validity

Below is a simply table showing internal and external validity. You can see the differences between the two, and how both are important in understanding what exactly a study says.

Internal Validity External Validity
Main question Did the study accurately establish what caused the outcome? Do the findings apply outside the study?
Focus Cause and effect Generalizability
Concern Alternative explanations Different people, settings, and situations
Example Did CBT cause the reduction in anxiety? Will CBT have similar effects for people outside this study?

Why Both Types of Validity Matter

We’ll cover the tradeoffs in a moment, but both internal validity and external validity matter. Internal validity helps us determine whether a treatment actually works the way we believe it does, whether the improvements or effects are attributable to the intervention, and ultimately whether the study supports a causal claim.

On the other hand, external validity helps determine who the treatment may work for (and who it may not work for), whether the findings apply in general clinical settings, whether the results can be safely generalized to other populations, and whether the research suggests we can use the treatment in everyday practice.

Good research needs both internal and external validity, but they do answer different questions. Many studies are quite interesting and have strong internal validity, but not much external validity. That’s okay. We just have to analyze and understand the research and what exactly it is finding.

The Tradeoff Between Internal and External Validity

Now, on to the tradeoff. Many highly controlled experiments increase internal validity. But too much control can decrease external validity. This is a tradeoff, and something researchers consider when designing their studies.

For example, a study on CBT and anxiety might control for how the therapy is delivered (in-person vs. online therapy, in groups or individuals, etc.), who participates (only college-aged women at a university), and how outcomes are measured (specific testing criteria and collection of data on outside variables). This can provide a solid claim that the CBT helps anxiety. But it may really only be confidently said about that population.

On the other hand, research conducted in ordinary clinical settings often represents typical treatment but may not control for outside variables. This may have more external validity, but lacks the same internal validity. Some research is designed specifically with one or the other in mind.

Improving Internal and External Validity

There are many ways researchers improve both internal and external validity. Some of these may be plausible for some studies but not others, and improvements in one may impact the other.

Improving internal validity:

  • Randomization

  • Control groups

  • Blinding when possible

  • Standardized procedures

  • Reliable and valid measurements

  • Controlling confounding variables

  • Adequate sample sizes

  • Addressing attrition

Improving external validity:

  • Recruiting diverse participants (age, gender, race, socioeconomic status, severity of symptoms, location, etc.)

  • Using representative samples if applicable

  • Replicating the study with different populations

  • Conducting research in real-world settings

  • Using real-world clinical trials

How to Evaluate a Psychology Study

Often, someone in my life sends me a study or asks me to comment on it. When I read a study, I ask a few questions in regards to both internal and external validity. You might ask these same questions when evaluating a study yourself.

  • Who participated in the study? Is the population or sample representative of who we are talking about?

  • How was the study conducted? Is there a control group? Were participants randomized? What measures were used?

  • What else could explain the results? Are there confounding variables? Is there another possible explanation for the findings?

  • Where was the study conducted? Was it at a lab, a hospital, or a clinic?

  • How large was the study?

  • Has the finding been replicated?

  • Are researchers (or the article about the research) making a claim that goes beyond what the study actually showed?

Which is More Important?

This is a question that can’t really be answered one way or the other. It really depends on what you want to know. If you want to know if the treatment did cause the outcome, then internal validity perhaps is more important. If you want to know if the treatment will work for people like a specific client, then external validity is perhaps more important.

Often, some research will offer internal validity while other studies offer external validity. Although we may not find both in one study, we can combine studies to get a better picture of internal and external validity.

Elizabeth Sockolov, LMFT

Elizabeth Sockolov is the founder of One Mind Therapy, and offers CBT, mindfulness-based therapy, and EMDR to clients in Petaluma and around California.

https://OneMindTherapy.com
Next
Next

How to Deal with Intrusive Thoughts: 10 Ways to Respond Without Getting Stuck