Jack Rossiter
Problems with standardised effects sizes in education evaluations
Today we unpack various sources of distortion when calculating standardised effect sizes in education evaluations. This might seem technical, but understanding how measures of effectiveness are calculated can help us temper the confidence we place in claims about “what works.” My guest is Jack Rossiter.
Jack Rossiter is an independent researcher and non-resident fellow at the Centre for Global Development. Together with David Evans, Susannah Hares, and Catherine Henny, he recently co-wrote the working paper entitled “The Illusion of Comparability Among Standardised Effect Sizes: Why Education Evaluations Should Report Raw Effects.”
Will Brehm 1:40
Jack Rossiter, welcome to FreshEd.
Jack Rossiter 1:44
Thank you, Will. Nice to be here.
Will Brehm 1:46
Policymakers and donors and researchers alike often rely on standardized effect sizes to compare education programs, and policymakers in particular use them to decide where to put money, or are influenced by this research as to where they should put money. But before we get into some of the problems that you and your colleagues have uncovered about some of this, can you just explain in simple terms — what is a standardized effect size?
Jack Rossiter 2:10
I’ll try to do it with an example. If I’m interested in understanding what the difference between two groups of children is at the end of some sort of intervention — so say I’ve provided additional hours of instruction for one group, and the others had a regular class experience — I have a test. I can choose that test from wherever I want, or I could write it myself. I then put that test to the children, and at the end I find that the children who had the additional hours of instruction get seven points more on this test. That’s in the units of whatever the test was that I put together. That’s a raw effect size in the terminology that we’ve used in the paper.
Jack Rossiter 2:52
And a standardized effect is then saying, okay, I want to try and put that effect — and all of these other effects — onto a common scale so that I can compare them more directly, because the test that I used isn’t going to be the test that the next person uses, or the next person, and the next person, and so forth. And so in order to get from the raw effect to a standardized effect, we need to divide it by a measure of the spread of skills of the children who sat the test. So we say, okay, six points — what was the spread of skills? Let’s say it was 12 points of variation among the children, and we say, okay, that was half a standard deviation of effect. And then I can say, okay, I’ve got half, and then someone else has done the same process and they’ve got a quarter, and someone else may have one. And I can put these on a scale and start to understand what the difference between these interventions is.
Jack Rossiter 3:42
That sounds quite technical, but I think the main reason this grew was in order to support comparison. In the 1960s and 70s, there was work to say, okay, I’ve got this measure of — it was often psychology measures — a measure of something which is not comparable with another measure of some other trait, and I want to be able to put these into a meta-analysis and think about what we’re learning across all of the knowledge. And so they said, well, we need to get rid of the units of the test, let’s do this. But also I think it helps, within an assessment, to understand how meaningful a change is. So six points — if the difference among all the children in my test, in my intervention, was 100 points, well, six points, you’re not going to be able to see it very easily. And if the difference was only five points between some kids who were doing really well and some who were doing badly, then that intervention is quite powerful. So it gives you a sense of, within the assessment, how meaningful it is, but also across studies, how important one intervention might be in comparison with another.
Will Brehm 4:48
In education policy today — and perhaps historically — a lot of conversations are around getting children to read. Fluency is a huge effort by school systems worldwide, and there’s a lot of effort to understand whether children are learning to read with different programs being implemented. One way this is being measured is how many additional words per minute of oral reading fluency corresponds to some effect. And you looked at a whole bunch of programs and showed that there was a range between 0.03 standard deviations and 0.55 standard deviations for one additional word per minute — roughly a 20-fold difference. What does that range tell you? What do you read into that range of effect sizes?
Jack Rossiter 5:32
It doesn’t tell us how much children learned in these different assessments — that was my interpretation. It was quite an interesting exercise to do because you don’t usually see this metric. “What is a word worth?” is the way that we were thinking of it when we were trying to think about how to communicate the problem. Instead, I think it tells us how unusual the gain was in this particular study relative to the spread of scores. And the main thing driving this difference is that the spread of scores is just so different across all of these studies.
Jack Rossiter 6:05
So if you imagine one study where it’s relatively high literacy — and all of these are coming from low and middle income countries, so we’re generally talking about lower literacy in primary grades — but relatively higher literacy. So let’s say a grade four class in South Africa compared to a grade one class in Malawi. The spread of scores is quite large. Some kids can read with fluency and some kids still can’t read. And so there’s a wide range of scores. And in another environment, you’ll have almost everyone who cannot read — 95, almost 100% of kids can’t get any score. And so this range of a word being worth half a standard deviation or a word being worth one twentieth of a standard deviation is really reflecting something else about the spread of scores in the population, not how important or how meaningful the intervention itself was.
Will Brehm 6:58
Your paper ends up arguing that there are three main sources of this distortion. The first one is about how much scores vary within the comparison group. Walk me through that concept. What does it actually mean to have variation within the comparison group?
Jack Rossiter 7:15
Yeah, and I should flag that I remember reading a blog post from around 2015 that Abhijit Singh wrote, and his was the first time I came across this problem. His was a very similar argument to ours — he was just trying to make the point that standardized effects in education interventions of this type don’t convey the meaning that we often think. And others have written on this, typically within a single study, trying to compare and say, hold on, this measure is giving me a very large standardized effect, whereas that measure is giving me a relatively small standardized effect with the same intervention on the same population. So what we were trying to do was say, okay, how can we reiterate the point, but also make it a bit more general across the literature and see what it means.
Jack Rossiter 7:58
So this is probably the most familiar of the three distortions that we refer to. The example I gave a moment ago is probably as simple as they are. In a relatively high literacy environment, some kids can read, some kids still cannot read — there’s a very wide range of scores in this oral reading fluency measure. If you picked a kid out, it might be that they can read 10 words, or it might be that they could read 55 words; you just don’t know. Whereas in a very low literacy environment, particularly in the earlier grades, and also in languages that are orthographically more complicated or that construct sentences differently — so fewer words to convey the same amount of meaning — you may find that almost all children can’t read. And so the variability of scores is very narrow.
Jack Rossiter 8:44
If you think of it as: that is the denominator in this function that turns a raw into a standardized effect — dividing by a big number gives you something smaller with the same thing on top; dividing by a small number gives you something bigger with the same thing on top. So each of these studies could move scores by two words per minute, but because I’ve got a much wider range of scores in one group than in another, the standardized effect is completely different.
Jack Rossiter 9:10
There is something real about that — it does help to compensate a little bit for differences in language systems. I think it would be wrong to assert that we should expect to see the same standardized effect per word across all languages and grades. That’s incorrect, for sure. But 20-fold? I don’t think we could believe that is representing just something about differences in language. That also reflects problems with differences in samples and everything else.
Will Brehm 9:35
And you bring some of these problems to light when you look at the examples in Uganda and Kenya. Can you walk me through that — what did you find there that was so surprising?
Jack Rossiter 9:46
It was actually a lot of fun to do this, because I have so many small files on my computer with odd file names saying, wow, Kenya, what’s going on here? Or Nigeria? Just going through some of the data from different studies and surprising myself. The file names — very odd file names. And then I tried to rediscover what was what and tried to systematize it, and I couldn’t. But as I was starting to explore the issue — and I think this Uganda-Kenya example is there not because it’s necessarily the strongest, but partly to acknowledge the authors, because the data are available. Not always are the data available, so I’m very grateful for us to be able to access the replication data and look at these more carefully.
Jack Rossiter 10:25
Also, the Uganda study was one we looked at a bit more closely because the authors were making a similar sort of point to ours — that you’re measuring reading using oral reading fluency, or you’re measuring written comprehension, or reading comprehension, or the ability to sound out letters, and whichever one of these different outcome measures I choose creates a different standardized effect. So how I put these together matters. That was just to clarify why we came to this example.
Jack Rossiter 10:55
But I found this one fairly easy to use to demonstrate the problem because Uganda and Kenya are neighbors. They use the same type of intervention — very similar, a structured pedagogy approach, providing materials and support to teachers. They use the same instrument, obviously adapted to the relevant languages, but it’s the same EGRA, the Early Grade Reading Assessment, that a lot of USAID programs and others have used. And so we felt it was a fair basis for comparison.
Jack Rossiter 11:22
What we see is that in Kenya, lots of children at the beginning and at the end of the program — in both the control and the treated groups — can’t read. It’s still something like 50% of the children who cannot read in the control group at end line. But there is still a range; some children can. And so there is a distribution of scores that you could imagine as a reasonably stable denominator for this standardization. Whereas in Uganda, there were 473 kids in the control group at end line, and only 21 of those could read anything. So almost everyone got zero on the assessment. And among those 21, very few could read more than five or 10 words.
Jack Rossiter 12:00
If you visualize this — as we do in the paper — if you think of a distribution of scores, most people think of a bell curve, a fairly smooth, normal-like Gaussian curve. But this was completely different. This is basically just a vertical line with a few little dots next to it. And that’s what is, from my point of view, a pretty big problem, because you divide a small effect by something that’s very narrow, and the standardized number that comes out the back is very large. Whereas in Kenya, a large effect — I think it was 14 words per minute that children gained — divided by a fairly wide spread of scores, doesn’t look nearly as impressive, despite the fact that that intervention helped kids to read a lot more words per minute than the intervention in Uganda helped kids learn to read words per minute.
Jack Rossiter 12:48
That’s not — and I want to be very clear — a judgment of how good those programs were in comparison with each other, because it might well be that in Uganda, kids were at a level where they could have made progress in identifying letters or sounds, but weren’t really ready to demonstrate performance in reading fluency. But that’s a different type of problem that we could discuss as we go through.
Will Brehm 13:10
Right — so you’re not actually testing the kids on the things that matter to them in that context. But given that it was the same program, the effect sizes that get produced end up creating a distortion: the program that didn’t look like it did as well ends up looking like it did really well. So the Uganda effect size is bigger than the Kenyan effect size, even though what actually happened in practice was kind of the reverse.
Jack Rossiter 13:32
Exactly. So it’s like a 14 words per minute gain in Kenya translated as a 0.4 or 0.5 standard deviation effect, whereas a two word per minute gain in Uganda was worth about the same. So there’s an enormous difference in how we then interpret the evidence, depending on which of those effect sizes we present.
Will Brehm 13:52
Exactly. And how this evidence gets used in policy and in research — it takes on a life of its own afterwards. I want to move to the second source of distortion that you and your co-authors highlight, and this is about which group’s spread of scores you measure the gain against — the control group, the treatment group, some pool of both, or some combination. What does this mean, and why is it an important source of distortion in how we understand effect sizes across different programs?
Jack Rossiter 14:22
This particular distortion — what the reference group is — think of it again as the denominator in your sum. And I probably wasn’t fully aware of this. My earlier training was in development economics, and within that discipline we effectively become familiar with the idea of standardizing things divided by the control group. So we say, well, at the end of that program, we did something for half the kids. They were in the treated group, and we monitored the control group as well. The control group is our reflection of what would have happened — it’s our counterfactual. And so let’s look at the distribution of scores in that control group, and that’s going to determine our spread that we use to create the standardized effect.
Jack Rossiter 14:58
What I hadn’t appreciated is that there are lots of other choices one could make. It seemed that, qualitatively, going through these hundreds of papers, those coming from an economics background tended to follow that idea — take the control group. But those coming from a more psychology or educational discipline, still quantitative but within those streams, would take the spread of scores in all of the children at the end of the intervention, both those in the treated group and the control group. And that’s a very well-known standardized effect called Cohen’s D, named after Jacob Cohen. So we say, what was the spread of scores in all the kids in the control and treatment group, and that’s our denominator.
Jack Rossiter 15:42
But if you imagine a program that’s influenced the treatment group kids — it’s really effective and it’s moved them up, now they can read a lot more in that group — well, the spread of scores is presumably quite a bit wider if I include those in my denominator. And there are other questions as to whether you choose the denominator at the beginning of the assessment or at the end. If I took a baseline test, I could use that as my denominator, or I could do it at end line. So there’s no regulation or clearly best version of this. That discretion is pretty wide. You can choose.
Jack Rossiter 16:15
I certainly didn’t get the impression that people were gaming it — it was more that there are just disciplinary norms that seem to be being followed. But that means those who are using the pooled effect size probably, more often than not, underestimate the effect in standardized terms compared to what they would have gotten using just the control group, in a low literacy environment. We use one example — I think a dozen or so different programs — and try to show how much those effects would change depending on which choice you make. And they change enormously. You might go from an effect size of 0.4 standard deviations in Yemen to over one standard deviation. Nothing has changed — it’s exactly the same data. All you’ve done is change the way that you’ve standardized the effect.
Will Brehm 17:00
It’s so interesting that a researcher’s disciplinary background and their own norms of what is the right way to do research can, on its own, create huge distortion when you’re comparing different programs or different effect sizes. And that very technical choice distorts what we know about a program dramatically.
Jack Rossiter 17:18
Absolutely. And it’s hidden as well — that was probably what I found most frustrating about this. To work out how something had been standardized usually meant poring through the footnotes of tables or reading through the text, because all you see is “0.4 standard deviations” — okay, but a standard deviation of what? In the better examples, it would simply be described as “0.4 standard deviations of the control group scores at end line.” That’s nice and easy. But in some cases, there isn’t a single mention in the whole paper of how these things were standardized.
Will Brehm 17:52
In those cases, how did you find out?
Jack Rossiter 17:54
We didn’t. There’s a category for which we just have to report it as unspecified. Of 119 papers that report standardized effects, there were 20 for which it wasn’t specified — we didn’t know if it was the baseline or the end line, the control group or the joint pooled sample. So it’s not a small share either: 20 out of 119, that’s around 17%.
Will Brehm 18:15
That’s a fascinating insight, and it sounds like something researchers can absolutely address — just specify what denominator they’re using.
Jack Rossiter 18:23
Exactly. And I think that’s the first step. It also helps to discipline us as researchers. My takeaway was that we’re not spending enough time explaining what it is that we’ve measured, and why that’s going to be meaningful for the children who were part of this intervention — and perhaps others in the country this is relevant to. The first step would be, instead of just saying “standard deviations of some arbitrary measure,” let’s discuss what that means: what is a standard deviation of what, and what does that look like? But we can talk about that a bit more.
Will Brehm 18:55
And you get into this issue of raw effect sizes — like what an actual increase of seven units means, as you said at the beginning. But before we get there, I want to touch on the third source of distortion, which is really about whether the test measures what we actually want it to measure and what children should be learning. This happens in contexts where many children, as you’ve said, basically score zero on a reading assessment. How does having such a low floor in an assessment distort effect sizes and what we can tell about the power of a particular program?
Jack Rossiter 19:28
There are other papers from the psychometrics literature on the suitability of assessments for a given group. One of the big issues in any test development is floor effects and ceiling effects. What we want for all of these programs is a fairly smooth distribution of scores — I want to be able to separate all of the children, to discriminate and place them on a line, so that I can say this kid has the highest level of performance in my test and this kid has the lowest. If we have a floor effect, a lot of kids score zero on the test — let’s say half of the kids scored zero, which would actually be quite common among many of these oral reading fluency assessments. I can’t distinguish any of those children from one another. They all just look like they can’t read with any fluency.
Jack Rossiter 20:10
And a ceiling effect is the other end — if the test is just way too easy. So let’s say I measure only letter sound identification, and in this environment children are very good at that. Seventy percent of them score 100%, but I can’t separate them — they all look like they’ve done the same. When we develop and design tests, we want to make sure we avoid these sorts of phenomena because they don’t help us and they create exactly this sort of distortion that we’ve demonstrated.
Jack Rossiter 20:38
The biggest problem is that you can’t see progress underneath the floor or above the ceiling. It just appears that a child started unable to read and finished unable to read. We used one example in the paper from Uttar Pradesh where, quite exceptionally, the same children were assessed at both baseline and end line. So we can see how children did in terms of their ability to identify letters, words, and read a short sentence at both the beginning and the end. And we asked, okay, imagine we only looked at the reading fluency part of this — what would we not be able to see?
Jack Rossiter 21:12
We would have missed 1,500 children — about a third of the sample — who were able to climb that ladder. They had definitely moved up from identifying letters to identifying words, but they hadn’t managed to reach the point of fluency. And so we wouldn’t see it if we only focused on the reading fluency measure. Similarly, on a different type of assessment just looking at reading comprehension — maybe the intervention does help children start to read, identify words, decode, and so forth, but it doesn’t mean that they can all of a sudden read with comprehension. We miss it.
Jack Rossiter 21:48
And this is the problem of the assessment itself — how the test is designed and how suitably it is targeted to the children being measured. I think this is an area that we’ve let go for a very long time. We’ve focused a lot on identification and getting better at measuring causal effects — being very confident that this intervention is helping kids to read better — but not much better in terms of the way we then measure the outcomes and understand what that means for the children’s progress towards reading with meaning.
Will Brehm 22:18
One of the reasons this might be happening is because certain assessments have gone kind of global in the Global South — USAID’s Early Grade Reading Assessment, for example, which you mentioned earlier. Is that part of the reason this is happening?
Jack Rossiter 22:32
We thought that might be the case, but I’m not too sure, because we tried to codify the assessment being used across all the papers. For 170 papers we asked: what was used here? Was it a government test — something like a district exam? Was it the EGRA? Was it an assessment provided by another vendor? Or was it a researcher-designed test? And that last category was actually dominant — 60% of assessments are researcher-designed. So there’s a lot of discretion on that side as well. Again, that’s not inherently bad — it might be someone saying, I don’t think that test is quite suitable for my purposes, so I’m going to design my own. If that’s done well, it can help us and mean that we have better quality evidence. But if it’s done poorly — which is sometimes, maybe often, the case — it can create exactly these sorts of problems.
Jack Rossiter 23:20
So I think it’s less about the globalization of particular tests and more that we haven’t pushed people to explain what was in the test, to demonstrate some basic properties of that test in terms of reliability and other things. It’s a lot of hard work and it comes at quite a cost. It seems we can pass that up so long as we’ve done a very good job of explaining how we randomized and how we structured an intervention that didn’t create undesirable spillovers and other things. The measure itself gets a little bit less attention.
Will Brehm 23:55
And in the same vein, it seems like one of the things you highlight in the paper is a troubling trend in research where before 2010 in particular, it was more common for people to report the raw effect size — this group of students performed seven points better on the test, for instance — rather than just reporting standard deviations. More recently, the number of researchers reporting raw effect sizes has decreased. Why do you think that is, and why would you like to see more people reporting raw effect sizes?
Jack Rossiter 24:28
I was actually surprised by this. I probably didn’t expect it to show that pattern. And also a caution: when we were putting together these papers, they aren’t built out of a strict systematic search review. We explain exactly how we put them together. We’ve not cherry-picked by any means, but it’s possible that with different search criteria we might see something different. We built them from systematic reviews in education — we thought, well, these are the sorts of reviews that influence policy and practice and that summarize what the group of readers likely interested in this paper refer to. So just a caveat that this number might not carry to a different set of papers.
Jack Rossiter 25:05
But nevertheless, it becomes clear that more often than not we see standardized effects reported without their raw effect counterpart — sometimes both, but not always. And that trend over time suggests that fewer papers are explaining what it is in raw terms that children could do differently. I don’t have a single answer for why, but reading the papers, my impression is it has something to do with how we want to communicate the effect.
Jack Rossiter 25:32
There’s a great paper from some of the early TAAL interventions in India which reports the effect and then has a short paragraph that says, just to put this into more tangible terms, half of the children have managed to move from being able to identify letters to now being able to sound out words. It’s just a much more tangible example of what changed. But there’s nothing in there about this being comparable to all of the other standardized effect sizes in the literature — sitting in the top 5%, or whatever — because that sort of reporting just didn’t exist when that paper was written.
Jack Rossiter 26:08
And my impression is that you now read the abstracts and introductions of these papers, and quite quickly it goes to: we did this intervention, the effect size was X, this is comparable to all of the other papers in the literature, sitting at the 6th percentile or in the top 15% of papers, we’ve also got cost data, so when we divide by every $100 we spent this is one of the most cost-effective interventions in the literature, and so on. It seems like the space of the paper has been gobbled up by these conventions — now we want to compare and contrast with others — whereas the former version had a little bit of space to talk about what it meant locally and in terms of the children that were part of it. So that has gotten a bit squeezed out.
Jack Rossiter 26:52
And maybe also there’s this issue of identification. There’s a lot more time spent now on making sure that we’ve created an intervention and a design that is watertight, so that we’re confident our effect is causal at the end. So there’s a lot more space devoted to that, and maybe that means less space for talking about the raw effects. I also think — and maybe this isn’t fair — but often when you talk about those raw effects, it’s not particularly impressive. And I think maybe we need to be a bit more modest on that front and just accept that when we write it up and say “this is equivalent to helping kids read three additional words per minute,” instead of being nervous that that’s going to undermine the intervention’s quality, we just accept that that’s not uncommon for these sorts of interventions, which are really hard to do and which often don’t have transformational effects on their own.
Will Brehm 27:45
So by way of conclusion — given that foundational literacy and numeracy is such a major zeitgeist in education development policy right now, given that you see things like “best buys” and cost-effectiveness rankings, and there seems to be so much pressure to generate the kind of results that will get picked up and used — what advice would you have for researchers working in this area? There’s a lot of pressure to emphasize effect sizes from the institutions funding the research and involved in this ecosystem. What would you say to researchers trying to navigate this space and be more transparent, even if the result isn’t so amazing in the end?
Jack Rossiter 28:30
I think transparency is the word I would lean toward rather than honesty. I don’t think there’s a huge amount of dishonesty here — I think it’s people responding to their professional incentives, and sometimes not thinking too hard about it. I certainly hadn’t thought as hard about the standardized effects I’ve worked with as I have in the last year or so working on this paper. And I think the transparency steps are pretty easy.
Jack Rossiter 28:55
The first step is something we can do with the data we already have — we don’t need to collect more. In addition to reporting the standardized effect, report the raw effect and present something about the distribution of scores in that raw effect and how you went through the standardization process. That’s very simple to do. I just need a second number alongside the standardized effect that expresses what it means: three words per minute, six additional items on the test, whatever it happens to be. And then something that shows visually what the distribution of scores looks like that informed the standardized effects. Because if you go back to that Kenya-Uganda example, you could quickly then say, okay, there’s something different about these two things, and I’m going to use that to think about how to interpret these two pieces of evidence alongside each other.
Jack Rossiter 29:42
I think those steps are pretty easy to do. And the bonus is that it will discipline us as authors to confront what one point means and how we can communicate that better. Some of the really good papers that I saw going through here — where there’s not necessarily even a very easy way of explaining what six points means on my test, because that’s still somewhat arbitrary — they’ll just have a simple table showing some example items or breaking items down into the competencies that were assessed. So maybe there was some basic operations with numbers or some sort of simple geometry questions, and just show what percentage correct the children achieved in these types of skills or skill domains. That would, I think, mean that any paper we turn to, we can do a much better job of interpreting what that evidence means for the sorts of decisions we might make, or how we convert it into useful communication to policymakers, and how relevant we think these studies are to different settings.
Jack Rossiter 30:35
And then all of the work we’ve done on evidence synthesis — things like the GEEP, the Global Education Evidence Advisory Panel — I think that’s helped us a lot to think about what’s in the literature. But I think it’s almost like the first phase was doing it with standardized effects, and I think we have to stick with that, but it would be quite nice to see something that runs alongside it. Look at some subset of these interventions that use a common type of instrument — what do we learn from those? — or some other subsets and start to think about whether we’re learning the same things or different things by looking at those raw effects.
Jack Rossiter 31:10
One thing we didn’t want to do is go too far, because it’s easy to say we should do this more and that more and the other thing more. But when we shared this with a reviewer, he mentioned that he now routinely, when he’s reviewing a journal paper, asks to see what the raw effects were. I don’t think we can imagine that becoming a norm standardized across the whole field, but those sorts of small steps — providing a bit more awareness — help others in that role to think, okay, I can see that it’s a big effect in standardized terms, but there’s nothing here that explains what it means. Ask for that. It doesn’t need to take up a lot of space and time, but it would move us forward in understanding what the different pieces of evidence mean.
Will Brehm 31:58
I think it’s such a fascinating topic, because it shows you the evolution of a field and the methodology being used, and you’re pointing to the next evolution that might make the work better. Jack Rossiter, thank you so much for joining FreshEd. Congratulations on the paper — it was really fascinating to dive into some of these technical statistical elements that maybe a lot of people on the podcast don’t think too much about, but are obviously hugely important for researchers and policymakers. Thanks again for joining.
Jack Rossiter 32:26
Thank you very much, Will.
Translate this transcript into any language. Before you do, we'd love to know a little about who this is for — it helps us understand which language communities find FreshEd useful.
Translations are limited to 10 per day per visitor to keep this free for everyone.
A Standard Deviation of What? Rethinking How We Report Learning Gains
How Big Are Effect Sizes in International Education Studies?
Interpreting Effect Sizes of Education Interventions
The Early Grade Reading Assessment: Applications and Interventions to Improve Basic Literacy
Have any useful resources related to this show? Please send them to info@freshedpodcast.com


