Why group disparities persist: A new way to break down the causes
Ang Yu and Felix Elwert introduce award-winning research on a groundbreaking methodological framework that explains why group disparities persist.

Why do some groups consistently fare better than others? Researchers and policymakers ask this about income, education, health, and other outcomes. But identifying a gap is only the beginning. If we want to reduce it, we need to know what generates it. Is the gap driven by unequal access to a helpful resource? By different returns to that resource? Or by who within each group actually receives it?
That turns out to be harder than it sounds. Two groups might have very different outcomes for several different reasons at once, and if we cannot tell those reasons apart, we risk reaching for the wrong solution. Our paper ‘Nonparametric causal decomposition of group disparities’ (read in full here) introduces a new way to answer those questions. We propose a causal decomposition of group disparities (a formal framework that breaks a total observed gap or difference into distinct, specific components). Causal decomposition analysis (CDA) is relatively new, with foundational papers having been published in the late 2010s and early 2020s.
Unlike many earlier approaches, it is designed for the observed group disparity itself, rather than the causal effect of group membership. It is also nonparametric, meaning it is not tied to any specific identifying or functional form assumptions.
In our paper, we consider the example of why people from wealthier family backgrounds tend to earn more as adults than those from poorer backgrounds. We know that going to college plays some role. But how, exactly? We show that there are three ‘mechanisms’ at play here, three distinct ways college could be contributing to that gap.
The first is straightforward, people from wealthier families are simply more likely to go to college. If access is unequal, outcomes will be too. The second is that a degree might pay off differently depending on your background, perhaps because of the institutions you can access, the networks you build, or the opportunities available to you afterwards.
But there is also a third factor, and it is one that researchers have previously overlooked. We pinpoint this mechanism as being differential selection based on individual-level effects, or in other words, it comes down to who, within each group, actually ends up going to college.
If the people who complete college in one group are especially likely to benefit from it, while those who complete college in another group are less strongly selected on gains, that also shapes the disparity. This can happen even if the two groups had the same graduation rate and even if the average return to college were identical.
Isolating this selection mechanism is our paper’s key conceptual contribution, because none of the earlier decomposition methods could capture it separately.
Distinguishing between mechanisms matters because each one points to a different policy lever. If prevalence is the issue, the answer may be to expand access. If effects differ, we need to ask why the same treatment helps one group more. If selection matters, attention shifts to sorting: who is able to effectively act on their own treatment effect. In that sense, the framework helps policymakers think more clearly about what kind of intervention might reduce a disparity.
To estimate the decomposition, this paper develops machine-learning-based procedures, which make the approach flexible and robust in functional form. This matters because real-world relationships can be highly complex. At the same time, we provide rigorous statistical theory showing that researchers can still make valid statistical inferences even when machine learning is used.
We illustrate the approach by studying intergenerational income persistence in the United States: the gap in adult income between people from higher-income and lower-income family backgrounds. Using the National Longitudinal Survey of Youth 1979, we examine how much college graduation contributes to that gap. The observed disparity is substantial: individuals from lower-income origins end up, on average, about 21 income percentiles lower in adulthood.
We find that higher education plays contradictory roles in the reproduction of income inequality. Unequal college completion increases income persistence, while differential selection into college offsets part of that persistence. Differences in average effects across groups, by contrast, are comparatively modest. In other words, college matters not only because some groups complete it more often, but also because the people who complete it are selected differently across groups.
For applied researchers, we provide an R package, `cdgd`, available on CRAN. The package implements the causal decompositions developed in the paper and offers both parametric and machine-learning-based options for estimation, making the framework easier to use in practice. We believe the paper offers a clearer way to think about disparities: not just whether a treatment matters, but how it matters and through which mechanism. That makes explanations more precise for researchers and gives policymakers a better map of where inequality is coming from and which levers might be most promising for reducing it.
Read our full paper here.
By Ang Yu (Assistant Professor, Hong Kong University of Science and Technology) and Felix Elwert (Professor of Sociology, University of Wisconsin-Madison)