EPITalk: Behind the Paper
This stimulating podcast series from the Annals of Epidemiology takes you behind the scenes of groundbreaking articles recently published in the journal. Join Editor-in-Chief, Patrick Sullivan, and journal authors for thought-provoking conversations on the latest findings and developments in epidemiologic and methodologic research.
EPITalk: Behind the Paper
Predicting Depression Beyond Race
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Dr. Patrick Sullivan, Editor-in-Chief, is joined by co-guest and recent graduate from the University of Texas, Dr. Priya Thomas, to discuss her paper, "Disentangling Race and Ethnicity in Predicting Symptoms of Depression among Young Adults: A Machine Learning Approach.” In this episode, Dr. Thomas describes how categorizing race and ethnicity could impact how well we could predict an outcome. Dr. Thomas’ article is published in the June 2026 Issue (Vol. 118) of Annals of Epidemiology.
Read the full article here:
https://www.sciencedirect.com/science/article/pii/S1047279726000918
Episode Credits:
- Executive Producer: Sabrina Debas (Episodes 1-18) and Sofina Tran (19-)
- Technical Producer: Paula Burrows
- Annals of Epidemiology is published by Elsevier
Hello, you're listening to EpiTalk Behind the Paper, a monthly podcast from the Annals of Epidemiology. I'm Patrick Sullivan, editor-in-chief of the journal, and in this series, we take you behind the scenes of some of the latest epidemiologic research featured in our journal. Today, we're here with Dr. Priya Thomas to discuss her article, Disentangling Race and Ethnicity in Predicting Symptoms of Depression Among Young Adults, a Machine Learning Approach. You can read the full article online in volume 118, the June 2026 issue of the journal at www.annals of epidemiology.org. So first let me introduce our guest, Dr. Priya Thomas, is a recent graduate of the University of Texas School of Public Health. Her research centers on depression, anxiety, suicidal thoughts and behaviors, and mental health treatment utilization with a focus on sociodemographic differences, especially on race and ethnicity among young adult college and student populations. Dr. Thomas, congratulations on your recent graduation and thanks so much for joining us today.
Dr. Priya ThomasThank you. I'm so honored and excited to be here and discuss this work and get this started.
Patrick SullivanWonderful. So before we dive into the details of your paper, can you just give us a little overview of what the mental health landscape looks like for young adults today and where race and ethnicity fit into that conversation?
Dr. Priya ThomasSure. So overall, uh young adults have some of the highest burden of mental illness and higher than other adult age groups. So the prevalence usually hovers around 20 to 25%. But what really makes us a vulnerable group is that onset of mental illness tends to peak at around 15 years of age, with 75% of onset usually occurring by the age of 25. So this is based on a variety of factors because this age range is in a transition period socially, economically, biologically. So in the last few years, prevalence and incidence of mental illness has increased among the young adult population as they've been experiencing rising inequalities and intergenerational inequity. So that means like not hitting life milestones at the rates that previous generations were able to, lack of independence, and financial instability. Further to that, we also saw these acute increases among young adults during the COVID-19 pandemic. So this manifested through school disruption, unemployment, social isolation. So while mental health has improved some since the peak of COVID-19, it largely has not returned to pre-COVID levels. So in other words, these acute pandemic effects have not entirely gone away and has fundamentally shaped the experiences of this young adult age group. And then further to that, actually the largest loss of GDP and other economic loss is based in mental illness. And this is explainable mainly by this early onset and the subsequent impairment in your daily functioning. So if left untreated, a lot of mental illness can lead to a worse prognosis. So what does that mean? That means more frequent episodes, longer episodes, a higher risk of relapse, and then more chronically experienced symptoms in general. So then, kind of to tie into that second half of your question. So how does race and ethnicity tie into this? So generally speaking, mental health disparities by race and ethnicity are well established and their extent in the literature. Racial and ethnic minority groups compared to non-Hispanic white individuals experience more chronic stressors that negatively impact their mental health. And a lot of these are race-related stressors like discrimination and social disadvantage. And then further to that, you know, in 2021, the Surgeon General had put out this advisory about the youth mental health crisis, and it outlined the necessity to urgently address this. But as part of that, it really included aspects of compounded risk for racial and ethnic groups. So, for example, like black youth experience higher suicide rates, American Indian and Alaska Native youth are at higher risk of social isolation and less access to treatment. So that's kind of how the landscape looks currently in terms of the youth mental health epidemic today.
Patrick SullivanGreat. Thanks for that background and so many important thoughts about how we handle race, ethnicity in epi analyses.
Dr. Priya ThomasYep.
Patrick SullivanThey're almost never really about race and ethnicity. They're about how our society engages with people in the context of their race and ethnicity and you sort of disambiguating that. So, with that background, tell us a little bit about the question you were trying to answer, the study design and the methods that you use to do this analysis.
Dr. Priya ThomasRight. So our primary aim was to see how various categorizations of race and ethnicity could impact prediction of depression symptoms in young adults. So for our exposure, the primary exposure was racial and ethnic categories. And what really drove our schema for our model building was the Office of Management and Budget, the OMB, is a federal organization that sets the standards of how we as researchers and throughout are supposed to, or maybe categorize race and ethnicity in applying that to analyses. So what does that standard really entail? A lot of that, and it's changed a little bit, they tried to update it, but a lot of those standards analytically really look at both separating the Hispanic ethnicity from racial identity. So, you know, if you choose in a survey, whatever racial category you put, Asian, black, white, whatever, if you also choose Hispanic ethnicity, you are automatically put into the Hispanic ethnicity category, and that supersedes whatever racial identity you may have chosen. So it really means that whatever racial category that you're inputting is completely masked by that analytic separation. So conceptually and analytically, there's inherent heterogeneity that's forced by this exclusion criteria. So that really drove part of our operationalization for race and ethnicity. And then the second part of it is this disaggregation of what's usually collapsed into the other race group. This is usually comprised of racial and really racial groups that just aren't sampled enough or are not represented enough in a lot of surveys, national surveys, local surveys, however you call it, usually this comprises of Asian, multiracial, American India, Alaska Native, just to name a few. So we really were aiming to disaggregate and separate some of those categories as well, while also considering this overlapping Hispanic ethnicity and racial identity as part of our operationalization in the models that we tested. So that was really the focus for the exposure. And then for our outcome, which is symptoms of depression, we were fortunate in our survey to have the PhQ9, the Patient Health Questionnaire 9, which is a widely used validated clinical measure of depression to use in our analyses. I mean, there are limitations to its utility because it's not always validated in larger populations and across diverse groups, which is discussed in the paper itself. But it's good that we were able to use a clinical measure that's mapped to the DSM, the diagnostical and statistical manual, rather than rely on a single item space validity. So we kind of have these two exposure outcomes that we were really lucky with these data in particular, able to drive.
Patrick SullivanSo I wonder if you can just start out about telling us in your analysis what were your exposures and what were your outcomes? And then we want to get a little bit into the machine learning approach, but big picture, you know, what are the exposures and what are the outcomes in this analysis?
Dr. Priya ThomasRight. So just to kind of summarize some of the key elements of the paper, we were really, you know, the title says disentangling race and ethnicity. So we really wanted to see how the heterogeneity of how we categorize race and ethnicity could impact how well we could predict an outcome. So just to summarize some of those elements, so this is a cross-sectional study design using Wave 14 from fall 2021 of this larger Texas-based cohort study, which was the original study was actually about adolescent tobacco and marketing surveillance, but these were some of the key variables we were able to extract from this Texas-based cohort in choosing to do this analysis because we have so many race variables and validated depression measures as well. But then the second kind of half of that design and the methods was really implementing a machine learning approach to this. So the first part of that was choosing which machine learning model out of a handful of models to use. So that's sort of the first iteration of well, what can best be, you know, what do our data fit best? What does it fit our needs best? And then how are they performing when we're applying them to an outcome? And then kind of based within that classifier, how to use that to answer our question. So ultimately, you know, we went with a random forest model for a few reasons, because it's a pretty robust model, especially in terms of handling collinear features, which so we don't even have to worry about that. And more importantly, is we can get a little bit of the peak behind that black box and we can derive some feature importance and a little bit of directionality with partial dependence plots, for example. So we can look at how these various racial and ethnic categorizations can be applied to an outcome and see which of those variables can impact that prediction. So those are sort of the key elements in choosing how to come to our statistical analysis plan.
Patrick SullivanSo we've talked a lot about the methods. So there are a couple of important points here. One is that for OM B reasons, we consider race and ethnicity separately and for important reasons, but in the same way we might talk about people as being black, non-Hispanic.
Dr. Priya ThomasYeah.
Speaker 1I think the the question is always in even in big surveys, whether there are sufficient numbers of smaller groups to be able to make confident sort of inference.
Dr. Priya ThomasYeah.
Patrick SullivanAnd I think that a lot of the folks, and myself included and the things that I see come in, take different approaches to this, which is like have enough for a stable estimate, or one that really is is often problematic, is like so we combine everybody else and just Yeah, and I'm not immune to this.
Dr. Priya ThomasI've done this too. Like, you know, it's not like blaming anybody for anything. I I totally get analytically why you're kind of restricted if you just don't have the sample to make a meaningful discussion about some disparity or some within group or between group difference. That's really but then that goes back to sampling because if these are just the data you have, then that's just the data you have.
Patrick SullivanIt is just a problem that I think all of us deal with, which is the choice of saying we don't have enough data to make a confident estimate, but also not wanting to present data where we're excluding data that was provided in the end from a subject's point of view, you have people who consent and provide data. There is a sort of understanding that those data are going to be used to improve health. So, what do we do with those? So that's a little bit of a sidetrack and a problem for the field, but you actually uh use some great methods to try to tease these things out. So talk a little bit about what some of your main findings were, and to you, what is sort of the headline finding of this that would help people understand this?
Dr. Priya ThomasYeah. So this headline finding, sort of the like, you know, what I would put on a neon sign, this is part of what made this, I think, so difficult to get out there, is the fact that regardless of the heterogeneity that we were able to parse out, you know, in this just preliminary analysis, overall race and ethnicity are not strong determinants or strong predictors of mental health, of depression in young adults. Frankly, if I were building a larger model, I would drop these variables because they're not doing much to impact how well we can predict our outcome of interest, depression. But on the flip side of that, some of the findings that did emerge showed that socioeconomic status and specifically lower socioeconomic status was positively predictive and predicted depression in young adults as well. So kind of those two findings together were really just the overall sort of that flagship takeaway from these analyses, I would say.
Patrick SullivanAnd I think this raises such an important point about race ethnicity, which I think many authors, many analysts sort of feel like those are just a priori variables that we want to look at in their associations with outcomes. But as you found, you know, most often these aren't really about race ethnicity, they're about social disadvantage and inequities in society. So they become essentially confounded with socioeconomic factors. And I think this is something, you know, just in teaching epidemiology, even, I always just put a strong caution to the knee-jerk to say, like, we have to look at age, we have to look at race. And yes, those are exposures that are associated with outcomes, but we shouldn't be inferring that they're like causative of those outcomes. It's really societal structures, which are about inequitable opportunity on the basis of race and ethnicity and poverty and family of origin, like all these things that create unequal opportunity. There's actually some nice articles that have these sort of unequal opportunity tags in them. Yeah, so I think it's the way you're looking at it, both analytically and substantively, is sort of exactly, you know, gets to a different question, which is what is really predictive of you know symptoms of depression. So you have this, you know, data set that you use. One thing I always like to think about, like to ask people about is how generalizable are these findings and conclusions to broader populations, either of people in Texas or more broadly. So, what do you think about the um sort of generalizability of what you found?
Dr. Priya ThomasYeah, sure. So I'm a proud Texan born and raised. So, you know, because Texas is so diverse, we are uniquely positioned to have several non-white racial categories. And for the sample in specific, it this was a primarily Hispanic lower income sample. So I think these findings are generalizable to regions as diverse across the United States, especially given how the sample, how the distribution of the sample. So I think, you know, the findings of race and ethnicity being poorly predictive of depression are made more robust given this diversity. But we also, you know, we sampled from five major metropolitan regions of Texas. So that would be the Dallas-Fort Worth area, Houston, San Antonio, Austin. So those counties, it's where the population density is. So some of this may not be as generalizable to more rural areas, especially in Texas, but also to less diverse regions in the country, which, you know, is a lot of the United States doesn't have as much diversity or, you know, mixing kind of in that way. So I think to take that and how generalizable these findings are. When you have a more diverse sample, you know, maybe race and ethnicity aren't really predictive of depression, but that could change, you know, it's part of science's replication. So finding this, how do these apply to another area would be, you know, kind of the next step in terms of inquiry for this.
Behind the Paper
Patrick SullivanGreat. Okay, so now we're gonna move to a part of the podcast that's called Behind the Paper. And I think it's always just so interesting to understand, you know, how colleagues come to ideas, like what struggles come up. And in the end, you you have this very nice, like technically right, organized paper, but often the process of getting there is uh also quite interesting. So I just want to start by asking you how this research question came up for you. How did you decide this was a worthwhile thing to invest your academic power into?
Dr. Priya ThomasYeah, so I love this question because this really was like based on a confluence of factors of two phenomena happening in parallel that converged to this research question and idea. So the first part of it was I had just published a paper on which this paper loosely builds investigating the independent and joint effects of race, ethnicity, and socioeconomic status on increasing severity of depression and anxiety among young adults. So those findings showed that there were no significant joint effects, but strong effects of low SES on both uh depression and anxiety outcomes. But the thing that kind of made me have a little bit of okay, what's going on? Was the results also show that some racial and ethnic groups, specifically non-Hispanic black, actually had a protective effect of the outcome compared to non-Hispanic white. And so this kind of led me to reading and really for the first time understanding this concept called the mental health paradox, wherein minorities experience more chronic and severe stressors that impact their mental health, but have better measurable mental health outcomes compared to their white counterparts. So this kind of made me think of a few things like what are we really measuring with race? Are we categorizing race correctly? Maybe some racial and ethnic heterogeneity is involved, and this explains the low effect sizes or the protective effects. So maybe that explains these quote unquote conflicting findings. And so, was this like a categorization and measurement issue? So that was like the first half of it. At the same time, so when I had started my PhD, I had no vested interest in machine learning or AI. But after a couple of years, I just I felt like I was just hearing about it everywhere, uh, how to use it, pros, cons, what it can and can't do, and so forth. So I was like, I'm the kind of person who's like, I want to see what the fuss is about. So I just enrolled in a few data analytics and machine learning courses to gain an understanding of what this methodology was, and just to add that to my statistical toolbox in case it's something that I could use to answer a question in the future. So then, you know, taken together, I thought, well, so both of these are happening at the same time. So I'm like, oh, maybe I can use machine learning to think about how to disaggregate some of these, like this racial and ethnic heterogeneity, and I can subvert some of these like strict necessities from traditional regression analyses. Maybe I can use machine learning to subvert some of that so I can have multiple categorizations of race and ethnicity, have those in stable, robust models, and then apply that to see how they can affect how well we can predict a depression outcome. So how those various combinations can do that in the first place. So yeah, it's kind of these two parallel phenomena.
Patrick SullivanYeah, and and the idea that you have this sort of substantive question that that co-evolves with the methodologic question is a nice place, uh nice place to land. You've alluded to this a little bit, but um, in terms of what was challenging about the analysis, I mean, I hear just some like methodologic questions, you know, already coming out, but anything else that, and you know that that where the challenges come from when we're planning to do something often are from unexpected places, like availability of the data or like ability to sort of combine different data sets that bring in different elements. So anything else for you that became a challenging part of doing the analysis that you hadn't um anticipated or that surprised you?
Dr. Priya ThomasYeah, so the you know, there were a few challenges, but the most challenging part wasn't actually the coding or the writing itself. Although, I mean, coding-wise, it did pose challenges because I'm learning about this while I'm doing this. And, you know, for example, I was like, oh, there's over under-sampling techniques. That's why none of these models worked in the first place. I'm like, oh, okay, then I have to go redo everything because I didn't know about that before. So I'm learning about this as I'm going through. Because again, coursework, as good as it is, it's very broad and it doesn't cover your specific issues that you have to deal with in your like publishable analyses. You're covering the basics. So that was sort of the more technical challenge. But the, you know, challenges coming from a surprising place was more the challenge in getting this published. That was the real challenge for me because I had a lot of polarizing views on the takeaway from these data and from these findings. Because, you know, we all I expect rejection. That's it's tough to get published in general. That is par for the course, and I expect to feel that. But, you know, I wasn't expecting some reviewer comments, you know, over the last couple of years that I did get that were like I've never experienced this kind of pushback before. So, you know, specifically what that means was like I was told that I'm using machine learning because it's fashionable and not relevant. Uh, I was told that my results were superficial, that my coding was wrong. And then also the one that really, and it kind of comes back to the like takeaway from the findings, was I was still uh there were people who wanted me to focus on the minimal effects of race and ethnicity and predicting the outcome and not to talk about income really at all. Uh, they just wanted me to still try to say that race and ethnicity was important in predicting a depression outcome within the sample when that's just not what our findings were. Uh, so it was it was really challenging kind of sifting through all of these uh polarizing opinions about a paper that I was really excited about, you know, and it was it was kind of ironic because I would also be told, you know, race is a social proxy. And that was kind of a major takeaway initially, you know, because for epic studies, it's well known that race is a social construct and it's a way that we investigate population differences. But our results gave some preliminary evidence that really aligns with a larger body of work that race is not a causal determinant, but rather that social implications, perception. So in our study, it was lower socioeconomic status, but others would be like racism, discrimination, and social disadvantage. So, you know, what I learned, I learned to be resilient, I learned to have a tough skin and keep moving. But I also learned that some disparities-focused research can be more polarizing than I had initially experienced or that I could imagine. Because in trying to shift research focus away from something so identity embedded, such as race and ethnicity, as a construct itself, and more towards a measurable determinants of health, such as you know, economic condition in this case, it can be it can be really tough to get somebody to want to read that if they're not in a space to view that objectively.
Patrick SullivanYeah, and it's such an important just concept that race and ethnicity for most things are markers of other societal inequities. And so when they get treated as elemental, I mean, there are some things, you know, like sickle cell disease, or there's some things that really are hard-coded uh and represent biological phenomena. But I think the central understanding in terms of social epidemiology is that these are markers for like a set of history and current experiences. And the process of teasing out, you know, the question you're asking, and then my other question always is like, what's the actionability of these things? And unfortunately, in the case of a lot of the things that are confounded with race, the action is societal level um you know changes that are needed. So I think the other thing that can feel unsatisfying about these things is that like we can propose, and and there are governmental you know programs, there are methods historically that have been meant to try to minimize the inequities that that just sort of co-occur. But it's also always a always a touchy topic. I mean, I think race and ethnicity are hard topics to to talk about uh sometimes in our society. And so I I really appreciate that when we can have these conversations and and disambiguate what we're talking about, because so often that we can superficially talk about exposures based on race and ethnicity, but we have to get down further into you know what's confounded with those. And those stories, unfortunately, are pretty consistent across health conditions and it's access, it's where people historically have been, communities have been, you know, live in the space of cities and what that means for access, what that means for transportation recruits, and like all those things that are associated. So I have a couple more questions. One is, you know, I appreciate you representing this transition in your own career of some work that you did as a student and some work as a colleague. What advice would you have for students who are interested, and early career folks who are interested in doing analytic work or methodologic work? Your paper had some component components of both about the this issue of race and mental health outcomes, which you dealt with. Like, what advice do you have about like how do we need to be going about this so that we're creating understanding and pathways to make things better, as opposed to to describing inequities that may occur around race ethnicity, which is not which there's no need to intervene on, right? Right, there's no need to intervene on race ethnicity, there's need to intervene on inequities that um arise. So, so what advice do you have? Like, what are important questions to ask? You know, what like where do we go next with trying to advance this line of thinking?
Dr. Priya ThomasNo, right, absolutely. So first and foremost, I in my experience, methodological and analytic work at these intersections of race, ethnicity, mental health, and then also socioeconomic status is specifically challenging because when you're conducting like measurement-based or methods-related research, you're challenging how someone is conceptualizing an assessment or a construct itself rather than an exposure outcome. This exists in the world. Well, how are we measuring the exposure? Someone's someone's already gone through their own heuristics, both from you know extant literature or lived experience, about how to conceptualize a specific variable. So when you're going upstream and thinking about how to reassess that or reevaluate its use, that's I it's a little bit more challenging to convince somebody about that. So to me, in moving forward with that, at least for me, I'm always like, be bold and go for it. Uh, just be aware of these specific challenges. Um, because sometimes it feels like you're challenging identity itself. And and that's not the purpose. But, you know, when we're reinforcing things like racial reasoning, so like for me, for example, to say, oh, I have some outcome because I'm Asian, when you say that out loud, that doesn't really make sense. So really to think about uh, especially for race and for mental health, to think about why you're including the variable that you're including and what you are actually needing for it to represent. I know that, you know, in data sets, sometimes you don't have a discrimination variable. Even if you do, a lot of people don't answer it. So you're limited in that capacity. But just to say we have racial and ethnic disparities is also a limiting way to talk about your research. So that's, I guess, one part of it. But my other part of it is learn as much as you can, like methodologically, when you are more flexible with your analysis and your analytic tools, you can be more creative with what you maybe can do. Like that's what happened with me. I learned about machine learning, and that's the only way I could even write this paper. So, you know, just to really take it upon yourself to learn and combine a bunch of different methods. But, you know, just like any other statistical methods, models again are only as good as we operationalize or consider the variables we use. You know, it's the same mentality of garbage in, garbage out. So, you know, really think about if including race in an analysis benefits the people that you are trying to, you know, what is the health outcome? How are you going to intervene on that? Well, if you want to talk about like social class instead, that's, you know, that's like one step away from that proxy. So that's kind of how I think about that intersection is really to be aware of what these variables are trying to say and the purpose of including them in an analysis rather than, you know, an epi, we just kind of throw in our regular sociodemographic confounders and covariates and call it a day, but to be more mindful of that, I think.
Patrick SullivanYeah, and I think it works um both ways, which is uh just coming from a surveillance background in a public health sense, it may be really important to say that you know, the Latinx people or the black people have a higher occurrence of a disease because we need to get more supplies or more screening to a place. And so, in that sense, um descriptive Epi uh may have a different take on the importance of those designations and just demographically and where we need services. I'm gonna literally send you a coffee mug. If, like if you tell me, and it's gonna say on it, why are you collecting race ethnicity and what do you want it to represent in your analysis? I'm gonna make one coffee mug for you and make one coffee mug for me. Have another Zoom meeting and we'll um coffee mugs.
Dr. Priya ThomasCheers. Yep.
Patrick SullivanThis is really worthy of coffee mug status, which is when you decide to use it, what do you want it to represent? And so in descriptive Epi for the purpose of screening programs or other things, maybe it is uh elemental. But in most of the analyses where we're putting it, you know, if you're putting in a table in a model, I hope you've asked that question and you've um inspired me to a new line of Epi coffee mugs. We'll make them available for now on Epi website, and you'll get a royalty for each one. So that won't actually. I I've really enjoyed our discussion, and I appreciate especially that you've talked about um some things that all of us experience, which is how we put our our work out in the world, how journals see it, how they receive it, and how frustrating it can be when the conversation between you as an author and a journal is these sort of like emails back and forth about and and sometimes you just want to say, like, you're missing the point. Oh my god. Or is there another way to um to convey that? But but I really uh appreciate you bringing your work to Annals. I appreciate the focus of the work, and I really appreciate you making time to talk today about just your process and how you got there and best of luck in your next steps and phases of career.
Dr. Priya ThomasI appreciate that. Thank you so much.
Patrick SullivanI'm your host, Patrick Sullivan. Thanks for tuning in to this episode and see you next time on Epitalk. Brought to you by Annals of Epidemiology, the official journal of the American College of Epidemiology. For a transcript of this podcast or to read the article featured on this episode and more from the journal, you can visit us online at www.annals of epidemiology.org.