Zero Months: 58 EEF Studies On Setting And Streaming
Zero months. Not a small gain, not "modest for some" — the EEF Toolkit's own average for setting and streaming, across 58 studies, is nothing. A setting coordinator and a form tutor autopsy the promise (teach to the top, protect the bottom, everyone wins) against what actually happened: a small negative effect on low-attaining pupils, and a structural risk of lower expectations, misallocation and less experienced staffing following disadvantaged pupils into lower sets.
What's covered:
• The Number: 0 months average impact, 58 studies, evidence security rated "Very limited" — the weakest band the Toolkit uses
• The promise vs the autopsy: teach-to-the-top-protect-the-bottom vs a small negative effect specifically on low-attaining pupils
• The structural risk: lower teacher expectations, misallocation of disadvantaged pupils to lower sets, less experienced staffing in those sets
• Try it today: check your bottom set's staffing against your best timetable slot, and ask one pupil what they believe the set says about them
Actions for consideration:
• Check whether your bottom set's staffing is a deliberate choice or a default outcome of timetabling — don't assume it's neutral.
• Ask a bottom-set pupil, directly, what they believe the set says about their ability — not just what mark placed them there.
• Don't treat "setting helps everyone" as self-evidently true — the Toolkit's own average is zero, and the harm is specific and named.
• If you set, own the disadvantage-compounding risks explicitly (expectations, misallocation, staffing) rather than assuming the policy is neutral by default.
Printable one-pager for this episode, free: https://theeducationcommute.co.uk/#newsletter-staff
Two free newsletters, one for staff and one for parents and carers, same evidence written for the person reading it.
Sources:
EEF Teaching & Learning Toolkit — Setting or Streaming: https://educationendowmentfoundation.org.uk/education-evidence/teaching-learning-toolkit/setting-or-streaming
Disclosure: Both voices on this show are synthesised (NotebookLM, Google); the research, reading and editorial judgement are done by a serving practitioner.
Strand: Teachers
Free printable one-pagers for this episode, one for staff and one for parents and carers.
Get the free newsletters · Listen on RSS.com · Watch on YouTube
Full transcript
Read the transcript
This is the education commute. Both voices on this show are synthesized. The judgment isn't. So today, we are unpacking a really fascinating stack of data from the Education Endowment Foundation. Yeah, we're trying to uncover the truth behind one of the most common structural practices in our schools. Exactly. We are evaluating the traditional logic of grouping pupils by attainment. And, you know, we're comparing that against what the evidence actually tells us about academic progress. Which is always a slightly tense conversation. Oh, absolutely. So I want you to picture the scene.
It is early September. Right. You are looking at the timetable. or perhaps you were sitting across from a family at parents' evening. Yeah, we've all been there. Right. And you were delivering the classic time -honored staff room pitch for how we organize our classrooms. I mean, we all know the script. We do. You group by attainment, which allows you to teach a narrow range of ability in each room. So the top set moves faster because nobody's waiting to catch up. Yeah. And then the bottom set gets more targeted support.
Right. Exactly, because the gap between the highest and lowest performer in that room is, well, it's so much smaller. It is an incredibly compelling pitch. It really is. Teach at the top, support the bottom, everyone moves faster, everyone wins. Yeah, it's a beautiful piece of logic. I mean, it makes intuitive sense to almost everyone who hews it. Right. Because it sounds like this really elegant solution to the very real problem of You know, managing 30 different cognitive baselines in a single room. And I know that logic intimately.
Because for this conversation, I am putting myself firmly in the setting coordinator's chair. Ah, right, the person building the grid. Exactly. I am the person who builds those math sets. every single September. I'm the one who sits with the spreadsheets, balancing the columns, making the actual structural decisions. And defending it. Oh, completely. Yeah. I have defended this exact policy using that exact logic at every parent's evening for years. It is my job to make sure this machinery works on a practical level. Well, I will be taking the form tutor's chair for this one.
OK, right. I am the one who sees every set, including the bottom one, walk through the door for registration every single morning. Right. So a very different view. Yeah, completely. Because while you are looking at the structural spreadsheet, I am seeing the reality of how that policy lands on the pupils themselves, like day in and day out, right there in the classroom. So what we're doing today is taking a hard look at the gap between that theoretical promise and the lived reality. Yeah. We are looking at the findings.
from the Education Endowment Foundation, specifically their teaching and learning toolkit strand on this exact topic. And the findings are stark. They really are. I'm going to deliver the headline finding exactly as it is written in their research, because it is a finding that completely stops you in your tracks if you are the one building these timetables. I will stretch yourself. According to the data, setting and streaming shows zero months of additional progress on average. Zero. Zero. Not, you know, a modest gain, not slightly positive, just zero.
Exactly. I want you, the listener, to really hold that in your head. Zero months. It's hard to wrap your head around, honestly. It is. We are going to autopsy why the evidence shows the exact opposite of that classic staffroom pitch. We will look at the promise, the document itself, what survived the scrutiny, and, well, what absolutely didn't. Yeah, so starting with the foundation we are standing on is crucial here. Right. This data is pulled from the Teaching and Learning Toolkit, and the research was completed on the 26 Aug 2026.
OK. First off, this policy is listed as having a very low cost. Which is key. Massively key. We have to acknowledge that this low cost is a huge driver of why the practice is so ubiquitous. Yeah, because it doesn't cost a school extra money from the budget to simply sort the pupils they already have into different rooms. Exactly. You're not buying new software or hiring external intervention tutors. You are essentially just moving names around on a timetable grid. It is an administrative task, which makes it feel like a free intervention.
Right. And the evidence base they are drawing from to evaluate this free intervention is actually substantial in terms of sheer volume. How big are we talking? They are looking at 58 different studies. Wow. OK. Yeah. And across those 58 studies, observing all these different contexts, the average impact is exactly zero months of additional progress. So the grand promise of the top moving faster and the bottom being supported. It results in a net average of absolutely nothing. Right. But see, as the As the setting coordinator, my immediate instinct when you hand me a zero average is to scrutinize the data itself.
Naturally. We need to look closely at the evidence security rating for that zero. For those who haven't obsessed over these reports, they use a padlock rating system. Right, from zero to five padlocks. Yes, to show school leaders how secure and reliable their evidence is. So five padlocks means you can take the finding to the bank. But if you look at the toolkit's own wording for this specific page, the evidence security is rated as very limited. Very limited. That is the absolute weakest security ban they use.
Exactly. There is no numeric padlock count published for this page at all. They say very limited explicitly. And the reasoning behind that very limited rating is important. There are too few recent studies and there is a large percentage of non -randomized controlled trial studies in that mix of 58. Right, which matters a lot in educational research. It does because a randomized controlled trial is the goal standard for proving that a specific intervention caused a specific result. Without a high volume of those, the waters get muddy.
Yeah they really do. So I am looking at this and thinking, yes the number is zero. But the zero itself is built on incredibly shaky ground. We shouldn't let this headline number carry more weight than the evidence behind it actually warrants. That pushback is totally necessary. And I think it leads directly into the crucial caveat of this data, which is zero months is an average. It is a statistical middle ground. It is not a universal law of physics dictating the outcome for every single school in the country.
The bell curve of those 58 studies means that some schools specific setting arrangements will do better than that. average. So generating some positive progress. Exactly, and some will do worse, generating negative progress. The evidence does not claim that setting never works, or that every set in school automatically gets a zero. Okay, so if I am a head teacher listening to this, I shouldn't just assume my own school is hitting a flat zero. No, you shouldn't. But even taking that caveat into account, an average of zero across 58 studies is still a devastating blow to the everyone wins pitch.
It really is. If zero is still the average outcome, we have to figure out where the logic of that pitch actually breaks down. If everyone isn't winning, who is losing? Well, the breakdown happens at the bottom of the structure. OK. When you examine the anatomy of this null result, it is not simply a case of every pupil in the building staying completely static and making zero progress. Right. It's not a uniform zero across the board. Exactly. The toolkit explicitly names a, and I'll quote, small negative impact on low attaining learners.
Wait, negative? Yes. And we need to be highly precise with this wording, that negative effect is specific to low attaining pupils. So it is not a broad generalization. that setting harms all pupils equally. Right. The specific harm, the mathematical drag pulling that average down to zero, is landing squarely on the lowest detainers. Which is honestly the most uncomfortable truth to face for the person building the sets. Yeah, I imagine so. The bottom set was the cornerstone of the moral defense for this policy. It was supposed to be the place where they got the most help.
The promise was targeted support. Yes, targeted support. But instead, the evidence points to a small negative impact on their progress. You know, if we step away from the spreadsheets for a moment and look at the lived experience of an 11 -year -old walking into a secondary school. Yeah, that's where you really see it. The mechanics of that negative impact become painfully obvious. Oh, absolutely. If you consider the long -term effects on attitudes and engagement, which the toolkit also explicitly warns about, you start to see the psychological mechanics at play.
Because setting is often treated like a purely academic sorting hat, right? But it is actually a profound exercise in identity formation. That's a great way to put it. Imagine a pupil placed in the bottom set at age 11. They look at their timetable in the first week. They learn something fundamental about their own perceived value and capability before they learn a single thing about the actual subject material. Yeah, the internal takeaway for that child isn't, oh, good, I get to move at my own pace now in a highly supportive environment.
No, definitely not. The takeaway is, I am the bottom. You are assigning them an academic identity on day one, and they internalize that ceiling immediately. And sitting in the form tutor's chair, do you actually see that play out? Oh, every day. I see that identity formation solidify in morning registration. You witness the promise of targeted support failing in real time. How does it actually look, though? It manifests as the way they talk about their own potential. You have a conversation with a pupil about their effort and they just shrug and say, why bother?
I'm in the low set anyway. Wow. You are watching the grouping mechanism itself erode their engagement. And the toolkit points that out, doesn't it? That grouping pupils on the basis of attainment may have longer -term negative effects on the attitude and engagement of those low -attaining pupils. Yeah, it's a slow drain on motivation. It isn't just about the test scores at the end of the first term. It is about how they view their entire academic trajectory over five years. The system has essentially codified their limits.
Right. But that psychological damage doesn't happen in a vacuum, though. No, it doesn't. It is the direct result of how the adults in the building are deploying their resources and managing their own biases. Which brings us to the structural half of the policy. And this is where the systemic flaws really compound the issue. Yes. Because based on the evidence in the toolkit, There are specifically named risks for disadvantaged pupils when it comes to setting. This is a crucial point. First, disadvantaged pupils may suffer from lower teacher expectations.
Right. The implication here is that the adults have lower expectations before the teaching even begins. Yeah, exactly. And because of this bias, these pupils are more likely to be misallocated to lower sets than their actual academic attainment would justify. Which completely undermines the whole concept. The entire premise of setting relies on it being a pure objective sorting mechanism based strictly on attainment data. Right. It's supposed to be just the numbers. But the evidence names the risk that bias heavily influences that sorting process. If a teacher assumes a disadvantaged pupil can not handle complex concepts, they recommend them for a lower tier, regardless of what that child might actually be capable of achieving with the right support.
So it disproportionately pushes disadvantaged pupils down the tiers. It does. And once they are pushed down into those lower groups, we hit the mechanism that completely shatters the targeted support promise. The staffing issue. Exactly. The toolkit names the risk that pupils and lower groups are more likely to be taught by less experienced and less qualified teachers. This is the staffing paradox, and it is entirely driven by the mechanics of how a timetable is constructed. As the setting coordinator, I can tell you exactly how this happens.
If we look at the literal building of a timetable grid, the constraints dictate the resources. Right. The year 11 top set dictates the grid. So the head of department, or the most qualified specialist, gets locked into that slot to secure the highest exam grades. Which makes sense on paper. It does, but that pattern cascades downwards. By the time you get to the year seven or year eight bottom set, you're often filling gaps on the grid with whoever happens to have a free period. Yeah, and that might be a non -specialist or a newly qualified teacher finding their feet.
Exactly. So setting is theoretically pitched as a triage system. You group the pupils with the highest need together so you can give them the most specialized help. But the data shows the reality is often the exact opposite. Yes. We are putting our most vulnerable academic cases in a single room. And instead of sending in our most experienced practitioners, we are sending in the staff with the least experience. The resourcing runs entirely backwards to the educational need. It's staggering when you lay it out like that. Lower expectations lead to misallocation.
Misallocation puts vulnerable pupils in the lowest tier. That tier is then staffed by the least experienced teachers. It's a chain reaction. It is. This structural chain reaction perfectly explains why we see a small negative impact on progress and long -term damage to attitudes and engagement. We do need to establish clear boundaries around what the toolkit is actually claiming here though to ensure we are representing the data impartially. That is a very good point. The listener needs to be clear that these disadvantaged compounding effects, the lower expectations, the misallocation, the less experienced staffing, these are explicitly named risks in the evidence.
Exactly. They are not universal facts asserting that this catastrophic chain of events happens in every single set at school across the country. No, of course not. It is what the evidence flags as a significant danger inherent to the system. But it doesn't mean your specific school's leadership team is definitely making these errors. Furthermore, we must be incredibly precise about what this autopsy of setting means for alternative models. The toolkit does not claim that mixed attainment teaching is proven superior. Right. The data we are unpacking today is strictly an evaluation of setting and streaming itself.
It is looking purely at the mechanics and outcomes of sorting by attainment. It is not a head -to -head trial against mixed attainment grouping. No, it's not. Mixed attainment has its own separate toolkit entry. with its own evidence base. We cannot use this zero -month average finding to suddenly declare that mixed attainment is the undisputed flawless champion of school organization. That is simply not what this specific evidence is doing. No. We are solely looking at the machinery of setting and discovering that on average it produces zero additional progress while carrying highly specific structural risks for the lowest attaining and most disadvantaged pupils.
Okay, so if I am sitting in a school right now commuting in and listening to this conversation, the theoretical debate have to give way to practical application. Yeah, exactly. If the honest takeaway is an abolished setting by Monday... Because... Again, the evidence doesn't say it never works, and the evidence security is very limited. Right. So if that's not the takeaway, what should a school actually do? School leaders need to know how to audit their own machinery. Well, there are two highly practical checks a school can do to test whether their specific system is falling into the traps, the toolkit names.
Yeah, these aren't long -term policy reviews. They are immediate audits of the reality on the ground. So check one. If I'm a head teacher, I'm pulling the staffing list for my bottom sets before the day is over. OK. You look at that list and you have to be brutally honest with yourself about how it was built. Is that bottom set staffed by your least experienced teachers by default because they just got whatever was left on the timetable grid? Or was it a deliberate strategic choice to put your most capable experienced staff in front of the pupils who actually need the targeted support you promised at parents evening?
Exactly. Because if it is the former, Your school is living the exact named risk in the evidence, and your resourcing does not match your rhetoric. Yeah, that first check exposes the structural reality. Right, and what's the second check? The second check exposes the psychological reality. You need to test that attitude and engagement risk in your own building. How do you do that? Find one cupel in a bottom set. Sit down and have a proper, meaningful conversation with them about their learning. Do not look at their September marks or their baseline data.
Ask them if they think their set reflects what they can actually do. Ask them what they believe about their own academic potential now compared to before they were placed in that group. And listen closely to the language they use to describe themselves. Exactly. Because that single conversation will illuminate the reality of the attitude risk. far more effectively than staring at a spreadsheet ever will. It takes the abstract phrase, negative effects on attitudes and engagement, and tests it against the human beings navigating your system. But both of those checks require a willingness to look at the unvarnished reality of your own structural decisions.
Yeah, rather than relying on the comfort of the traditional staff room pitch, Right, so we are left with this highly fragile reality regarding how we organize our schools. We have a foundational policy built on a very limited evidence base. It's lacking in numeric padlock count entirely, yet deployed almost universally due to its very low cost. And despite the widespread reliance on this policy, it yields exactly zero months average impact across 58 studies. The traditional assumption that the top moves faster and the bottom is protected is not reliably supported by the data.
Instead, we see a specific small negative effect on low -attaining pupils' progress, operating alongside a named risk to their long -term attitudes and engagement. And the overarching warning is how those risks compound structurally. The machinery itself facilitates lower teacher expectations, the misallocation of disadvantaged pupils, and the tendency for lower sets to be staffed by less experienced teachers. Which means you have to actively audit your own school against those specific structural flaws rather than just assuming the zero average doesn't apply to you. It requires a fundamental shift in how we view the data we use to sort these peoples.
It really does. So I want to leave you with final thought to mull over as you finish your journey today. We have dissected the data, the timetabling structures, and the psychology of the form room. But consider the sequence of events at play here. Okay. If a pupil's placement in a lower set is driven by lower teacher expectations before they even step into the classroom, well how much of that subsequent small negative impact on their progress is actually a self -fulfilling prophecy created by the adults rather than a true measure of the child's academic ceiling?
Wow. That is the question that demands an answer from anyone holding a timetable. Are we measuring their limits or are we building them? Everything's linked in the show notes where the trailer, not the film. Both voices on this show are synthesized. The judgment isn't. Safe trip in. See you at the gates.
More episodes
Newer: No 'Plus X Months': Five Meta-Analyses On Maths Manipulatives
Older: +6 Months, 145 Studies: But Not If It's Just 'Kids Teaching Kids'
All episodes