Alyssa Lawson
Does ChatGPT improve student learning?
Today we explore the research base that tries to answer if Generative AI improves student learning. My guest, Alyssa Lawson, recently co-wrote an article that critically analysed two meta analyses of ChatGPT that found it did indeed improve learning. Alyssa and her colleagues aren’t so sure.
Alyssa Lawson is a lecturer in the department of psychological and brain sciences at Johns Hopkins University. Together with Amedee Martella, Joshua Weidlick, Miriam Mulders and Josef Buchner, Alyssa co-wrote the new article “Color me confounded: a critical analysis of media comparisons on ChaptGPT in education” which was published in Computers & Education.
Will Brehm 1:15
Alyssa Lawson, welcome to FreshEd.
Alyssa Lawson 1:18
Thank you so much for having me. I’m very excited to talk with you today.
Will Brehm 1:22
So there’s a lot of excitement right now — perhaps like a bubble, let’s say, a bubble of excitement — about ChatGPT or just generative AI in particular in education. And there have been a few meta-analyses that have come out basically claiming how ChatGPT or generative AI more generally improves student learning. I remember seeing this hype all over social media. What did these studies actually find, and why do you think these studies spread so quickly?
Alyssa Lawson 1:50
So one of the major claims made in both the meta-analyses that we analyzed for this paper was that ChatGPT enhances academic performance for learners. And both meta-analyses found this with pretty large effect sizes. What they were looking at specifically was research that compared a condition where participants learned with ChatGPT to a condition where students learned without ChatGPT. And so from looking across all of these different studies, they were able to make this claim.
One thing I do want to note is that one of the papers that we looked at in our critical analysis did end up being retracted. The Wang and Fon paper was retracted back in April due to concerns related to discrepancies in the meta-analysis. So one of them is no longer being used as evidence. But that being said, the research that they looked at is still in the zeitgeist.
Through our investigation, it seems like these findings were based on studies that compared conditions that differed in more than just whether ChatGPT was being implemented, which makes it very difficult to know whether ChatGPT is actually useful or beneficial in learning. It’s hard to really credit any benefits seen to ChatGPT itself or generative AI itself.
Will Brehm 2:58
And we’ll dive into some of those methodological issues shortly. But before that, I guess — why do you think these two papers spread so quickly? So it’s the Wang and Fon paper, which was retracted, and then the Deng et al. paper, which is the other meta-analysis that you critique. These papers received hundreds or thousands of citations in a very short period of time. Why did they spread so quickly, both in terms of citations but also on social media, in the news, in the press? These pieces had so much power in shaping perception.
Alyssa Lawson 3:32
That’s a really great question, and I don’t have a certain answer, but I do have a hunch that the spread was likely due to both the excitement and the fear surrounding the integration of generative AI technology into learning spaces. My thinking is that people look to meta-analyses to help us make decisions about implementing technology into learning environments. And so having these meta-analyses say, hey, ChatGPT is really beneficial for learning — it tells us that across a whole bunch of different studies, this is what we’re finding, so we should trust this. I think that’s probably what sparked a lot of interest in these meta-analyses and similar ones that have come out since.
Will Brehm 4:10
So let’s get into the core methodological problem. Walk us through the critique that you’re making — the methodological issues that both of these studies were committing, let’s say.
Alyssa Lawson 4:20
So what we were actually looking at were confounds in methodological design. What we mean by that is that the conditions being compared are different in ways that go beyond just the technology being used. So if you’re asking the question, is ChatGPT effective for learning, then you want to know if that technology is having a causal impact on learning. But if you start adding extra things into one of the conditions that the comparison condition doesn’t get, then it makes it really difficult to say — if you do find effects — whether it is the technology causing the differences, or if it’s those other differences across conditions, or maybe an interaction.
Will Brehm 4:58
So can you give me an example of what that might look like? If ChatGPT was given in one group but not the other, what are these other things that might also have appeared?
Alyssa Lawson 5:08
This looked different across a lot of different studies, but for an example, a confounded study could look like a researcher comparing a condition where students are solving problems with absolutely no external feedback to a condition where students are problem-solving and receiving feedback simultaneously from ChatGPT. That difference in learning outcomes between the two conditions could be because ChatGPT is causing benefits for learning. But it could also be because there’s a benefit of feedback in the ChatGPT condition that isn’t present in the comparison condition.
Will Brehm 5:42
And since these papers were meta-analyses looking at multiple studies and trying to find some effect across studies — was this issue of confounding a problem within individual studies that were being done? Like, were researchers of these individual studies not having like-for-like comparisons? Or was this an issue with the meta-analysis itself, in the way it was being conducted?
Alyssa Lawson 6:05
So what we were looking at was the articles and the comparisons that were included within the meta-analyses. We were actually looking at the — you could say the raw data — the actual studies that had been published, that the meta-analyses then combined and found effect sizes for. So we were looking at the design of the original research studies.
Will Brehm 6:25
If you give an experimental group and a control group — ChatGPT in one and not the other — but there’s also different instructional design, is there a way to manage that difference and still say something about ChatGPT?
Alyssa Lawson 6:38
Yes. And I think one of the important things that we touch on in this paper — and something I think is really important — is to make sure that when we’re doing research, we have alignment between our questions, our methods, and our conclusions. There’s nothing wrong with doing research that looks at using ChatGPT with feedback compared to a condition where someone’s just problem-solving with no external resources at all. That’s fine if your question is, how does this instructional package impact learning compared to this other situation? That is a totally valid research question. But it’s not the same as asking, is ChatGPT beneficial for learning? It sounds very minor in the wording, but the meaning is really different. And that’s what’s really important.
Will Brehm 7:20
I want to dig into some of these meta-analyses that you looked at, because surely thinking about how comparable the different conditions are when you bring in the studies must be a common part of doing a meta-analysis. But in the Deng et al. piece, you found that only one comparison out of 27 was actually well-controlled. How did that happen? Shouldn’t it be 27 out of 27?
Alyssa Lawson 7:45
That’s the goal. But when it comes to educational technology — especially new technology coming out — it oftentimes feels like a race to the finish of who can get their work out first, how can we make a name for ourselves in this field. And to do that, you can’t quite achieve strong methodological control because it takes time and effort to think through all of those issues. If you don’t do that, it’s a lot easier to get to the finish line. That’s my guess at what the problem is.
I do also want to note — you say one in 27 comparisons in the Deng et al. analysis were well-controlled, and that is what we found, but it can be a little misleading. There were also comparisons that we identified as being difficult to determine because in at least one of the control categories we were looking at, there wasn’t enough information to make a determination about whether that variable was controlled for. So there is a possibility that there are more well-controlled comparisons among those 27 — we just couldn’t determine that from the information provided in the articles. But even without that information, it’s still hard to make strong conclusions, because we simply don’t know.
Will Brehm 8:45
One of the practical issues with these meta-analyses is that they sort of become the quote-unquote evidence to show that ChatGPT and generative AI actually increase learning. Whereas when you don’t account for all these confounding factors, you miss out on perhaps what is actually increasing learning.
Alyssa Lawson 9:02
Yes. And I think that’s a problem across a lot of different educational technology research. A lot of this technology does incorporate methods that we know from educational psychology research to be effective — like feedback and scaffolding. But if that’s not being taken into account in the research, then we really can’t say: is it the technology that’s making that happen, or is it just the method being used? If I were to take the technology out completely, would I still find strong effects between giving students feedback versus not giving students feedback, regardless of whether ChatGPT is there or not?
Will Brehm 9:38
Historically, has this been an issue in studies on ed tech? Is your critique of these two meta-analyses around ChatGPT just the latest in a long line of research on educational technology?
Alyssa Lawson 9:50
I would say yes. I think this is not an uncommon problem. And in fact, this paper was inspired by another critical analysis that part of this team — Ahmadi Martella and myself and a couple of others who weren’t on this team — did looking at the research in immersive virtual reality. We found very similar findings: a lot of the research published on immersive virtual reality environments looking at effects on learning was not very well controlled. I can’t say with 100% certainty beyond that, but my guess is that this is a prevalent problem across a lot of different educational technology. Another member of our team, Joseph, looked at the research on augmented reality and learning and found very similar results. So a lot of people are doing similar work with a lot of different educational technologies, and we’re finding very similar patterns.
Will Brehm 10:42
Why do you think that’s the case? Why is it that ed tech is being pushed as a solution, there’s a lot of research claiming effects, and then when researchers dig into the methodology, they find a whole bunch of methodological issues — repeatedly?
Alyssa Lawson 10:57
This is a very interdisciplinary field of study. There are people studying this who don’t have backgrounds in how to conduct human research. Without that background, it would be really difficult to understand: what is validity? What is reliability? Why is comparability so important? Maybe if you don’t have an education background, you might think, oh, well, lecture is the normal comparison. But I haven’t been in a classroom in the last 20 years where it’s just been straight lecture. So there are so many different people with so many different backgrounds — which is great — but if you’re trying to conduct research with humans, you need to know these foundational pieces.
Will Brehm 11:42
A question I have is around how we can do this well. Ed tech is so prevalent in education systems around the world and has been for a very long time. So it seems important to try and figure out the impact that it’s having on student learning. How can we do this successfully, rigorously, and with validity?
Alyssa Lawson 11:58
That is definitely something I think about a lot. But it goes back to what I was talking about before — it’s all about alignment between what questions you’re trying to answer, how you design the study, and what your conclusions are. So if you really want to understand whether this technology causes learning to happen, then think about: are the conditions I’m comparing comparable on everything that isn’t inherent to the technology itself? If you can’t say yes to that question, then you can’t answer the overall question. But maybe your question isn’t that. Maybe it’s: does this technology work in this specific way? That question changes how you design your study completely. Just making sure that there is alignment in those three components is the most important thing we can do in educational technology.
That being said, not everyone agrees with me on this. There is a lot of back and forth on whether it even makes sense to do media comparison — can we even do strong media comparison work? I’m taking a more pragmatic view, where this kind of research is going to be done regardless of what we say about whether it should or shouldn’t be. So if we’re going to do it, let’s do it well, or at least make sure there’s alignment. That being said, there are other people who say there’s no way to do media comparison well, and they would say just stop doing it.
Will Brehm 13:02
And when you say media comparison, you mean trying to figure out if one technology works better than another?
Alyssa Lawson 13:08
Trying to understand if, typically, a new technology is better than another technology, or better than a more traditional learning environment, like a lecture.
Will Brehm 13:17
Right. So some people say that’s an impossible task. And why do they say that?
Alyssa Lawson 13:22
There are a lot of reasons, but one thing that comes up a lot is that it becomes impossible to control for everything once you start breaking things down further and further. And so if you can’t possibly control for every single thing, there’s always going to be some sort of confound that makes it difficult to draw those conclusions. I’m not on that side of the spectrum, but there are people who argue for it.
Will Brehm 13:43
Right. And you’d describe yourself as more of a pragmatist — where you can identify the different confounds and design studies that more reliably try to understand the effect of a new medium.
Alyssa Lawson 13:53
Exactly. And I think it’s really important — I recognize that it is very difficult to control for everything. But being very explicit about this is what we weren’t able to control for, this is what we were able to control for, and this is how that impacts our conclusion — I think that is one of the more important things we can do when it comes to this kind of research.
Will Brehm 14:12
And what about meta-analyses? Because these are becoming quite popular in the published academic literature, trying to look across studies. What should we be aware of? What makes a good meta-analysis?
Alyssa Lawson 14:23
That’s a question I’ve been thinking about a lot since starting to do these critical analyses. I’ve been going back and forth with myself about whether there is a benefit to doing a meta-analysis. My co-authors and I like to joke and say “garbage in, garbage out.” If you’re not evaluating the literature that goes into your meta-analysis, then it makes it really difficult to know whether the meta-analysis is giving you good information. And that’s not the meta-analysis’s fault — it can’t control the research that is in the field. But I think it is important to start thinking more carefully about how good the research is that we’re putting into meta-analyses.
I do know that there’s some push for that, and there are different measures people use to assess how well-designed the research is. But to my knowledge, it’s not really getting down to that very specific, methodological, controlled aspect. So even when meta-analyses report that the research they’re using is strong, it’s not always being evaluated in the same way that we’ve looked at it here.
Will Brehm 15:18
I have a practical question. I guess the meta-analysis takes on so much power because it’s supposedly looking at multiple studies around the same phenomenon — so instead of having to read 27 articles, I can read one article and get the broad understanding. And it begins to cut across contexts and be seen as more generalizable. But if it’s garbage in, garbage out, and if I’m not going to dig into all 27 studies the way you have done, how should I be a skeptical reader of these? A lot of people listening to this podcast are probably students in education, or people who work in schools and don’t have the time to read a lot of research or are still learning how to read research. What tips would you have for those people who are going to be impacted by some of these studies?
Alyssa Lawson 16:05
Well, one of the things I do is look for the conversation that’s happening in more informal spaces. When these ChatGPT meta-analyses came out, there was quite a bit of discussion about them on LinkedIn from a lot of different educational researchers. And they very quickly pointed out that these meta-analyses were made up of really poor studies — that if you just look at two of them, they are not particularly strong. So I think that is a really great way to see what kind of conversations are happening around a meta-analysis in a more informal space — not in academic journals, but researchers communicating with one another.
I also think it’s really hard to know exactly what to do in this case, because you want to trust the research. But it really does depend on looking at least at a couple of the underlying articles and asking: do we see common flaws happening across these papers? And does that seem to be a problem across more than just a couple?
I would just be very hesitant to believe any really big, sweeping claims made in meta-analyses, especially with new technology. That’s probably me being a little skeptical about how impactful media can actually be in learning. But I think we do need to be cautious, particularly in the educational technology world at this current moment.
Will Brehm 17:12
I think a healthy skepticism is probably the right approach. And we have to recognize the sensationalism that news media will generate based on some of these research articles. Even really good research, when it gets taken up into the BBC or similar outlets, gets twisted a little — it doesn’t say exactly what the article is saying. So there are multiple levels at which things get reinterpreted and reimagined. And being skeptical at all of those levels is something a consumer of research needs to be aware of.
What about policymakers? So many systems are trying to figure out what to do about AI. It seems as if there’s a consensus that we have to cut deals with Claude and ChatGPT and embrace this new general technology, and if we don’t, we’ll be left behind. And they often cite research to justify that. What tips would you have for policymakers on how to approach this evidence?
Alyssa Lawson 17:58
This research doesn’t necessarily allow us to say whether or not ChatGPT or other generative AI tools are effective in learning. That being said, what we do know is that a lot of the techniques and methods that ChatGPT can use and make more prevalent for students are things that we’ve investigated very deeply in the literature — such as feedback, scaffolding, and executive function support. We know through decades of research that these things help learning. And ChatGPT can be used to support these approaches — to help students give themselves feedback, to help students scaffold how they’re thinking.
So I think trying to rely more on that, in this moment, is where we can have a little more certainty. Yes, this is probably going to help — not necessarily because of the technology itself, but because of the methods being used for learning. Now, if ChatGPT does support learning on its own, then great. But if not, at least students are getting those specific kinds of methods put to use, hopefully.
Will Brehm 18:55
And finally, what advice would you give to academics who are in this moment — this new zeitgeist — with all this pressure to focus on AI? You’ve looked at two meta-analyses that are seemingly pretty problematic, and one so problematic that it’s been retracted. The authors of that retracted article have probably suffered some serious consequences. So if garbage in, garbage out is the standard line — even as a joke — it raises some serious questions for academics. What tips would you have around how we should do media comparisons and meta-analyses going forward?
Alyssa Lawson 19:32
I know I keep saying it and probably sound like a broken record, but I think alignment is really, really important. Study what you are trying to answer, ask questions that you can answer, and use what you find to address those questions in your conclusions. Don’t just kind of hop around and say, oh yeah, we looked at this package compared to nothing at all, and we found this was really helpful, so we can say ChatGPT is effective. That’s not very helpful.
What is helpful is being very methodical about: what question do I actually have? How am I actually going to address this question? And what kind of conclusions can I draw from this?
And this is moving a little bit away from the question, but something I’ve been thinking about through this work is: maybe the question of whether ChatGPT is effective for learning isn’t a very interesting question. The reason for that is because ChatGPT can do so many different things — generative AI can do so many different things. So if you try to compress that into just saying, this is what ChatGPT is, this is what generative AI is, it’s not very informative. To me, it’s much more interesting to ask: if I use ChatGPT in this specific way, is that more beneficial than using that same strategy without ChatGPT? That’s a much more informative question and allows people to actually use the result, rather than just making a very broad statement.
Because if you say ChatGPT is effective for learning, that should mean any teacher can go and say, hey student, use ChatGPT, and it should be effective. The likelihood of that happening is very low. But saying, hey, use it in this specific way — maybe that’s more effective.
Will Brehm 20:58
So what you’re saying is: compare a traditional approach to scaffolding versus an approach to scaffolding that uses ChatGPT, and see if there’s a difference.
Alyssa Lawson 21:07
Yes. I think that would be more interesting and more informative than just asking, generally, is ChatGPT beneficial?
Will Brehm 21:14
And I guess finally — can I ask how, if at all, do you use generative AI yourself?
Alyssa Lawson 21:20
My husband is actually in machine learning and AI, so I kind of see it from a lot of different lenses. But in my own work and in my own classroom, I try to make sure that it’s a useful resource. One thing that I really try to emphasize is that generative AI is really good for executive function support. So you don’t need it to write your papers, but you can say, hey, help me break down this assignment, or help me understand what I need to add in this assignment. That’s a really good use of it to support your learning. You’re still doing the hard part — you’re still doing the learning — but you’re taking away some of that extraneous processing that isn’t as important to the learning process.
Do my students take all of that advice? Probably not. But I think I get through to students a little bit. And that’s what I use it for myself as well — things like, okay, I have these ingredients in my fridge, what can I make? So I dabble in that way. But I’m still trying to wrap my own mind around what I want to use it for and how I can use it. And because we don’t know how effective it is for learning, I’m really trying to lean into: how can we use it to deploy the methods that we know are effective?
Will Brehm 22:18
Well, Alyssa Lawson, thank you so much for joining FreshEd. It’s really been a pleasure to talk. Thank you for your work on these meta-analyses. I think it’s really helpful and points to new ways for researchers, the public, and policymakers to critically interrogate this deluge of articles coming out about the importance and effectiveness of ChatGPT in learning. Thank you so much.
Alyssa Lawson 22:40
Thank you so much. I had a great time chatting with you today.
Translate this transcript into any language. Before you do, we'd love to know a little about who this is for — it helps us understand which language communities find FreshEd useful.
Translations are limited to 10 per day per visitor to keep this free for everyone.
Color Me Confounded: A Critical Analysis of Media Comparisons on ChatGPT in Education
Reconsidering Research on Learning from Media
Cognitive Load During Problem Solving: Effects on Learning
Cochrane Handbook for Systematic Reviews of Interventions
Have any useful resources related to this show? Please send them to info@freshedpodcast.com



