'One Set Is Likely Sufficient': Ofqual's Mock Exam Warning
Autumn mocks, spring mocks, a third round before study leave — many schools run more assessment cycles than Ofqual's own guidance says are needed. An assessment lead and a class teacher work through what "one set is likely to be sufficient" actually means, and what to do with results once you have them.
Sources:
- Ofqual guide for schools and colleges 2026 (pub. 15 Jan 2026, updated 3 Feb 2026), gov.uk
- EEF blog: New case studies, making effective use of diagnostic assessment
- Printable one-pager for this episode: theeducationcommute.co.uk
What's covered:
- The honest finding: no statutory date or format for mock exams exists anywhere in DfE or Ofqual guidance, January is a scheduling convention
- Ofqual's real, current, named purpose for mocks: contingency evidence for Teacher Assessed Grades, should exams be disrupted
- The same guide's warning against over-assessment, verbatim: one set of mocks likely to be sufficient for evidence purposes
- What makes a mock cycle worth the disruption: sorting scripts by gap, not just rank, and reteaching before the next set, per EEF's diagnostic-assessment evidence
Actions for consideration:
- Before scheduling a second or third mock cycle this year, check whether it's gathering genuinely new evidence or just repeating what the last one already showed.
- Sort marked scripts by topic gap, not just by score, before any report goes home.
- Build a reteaching slot into the calendar straight after each mock cycle, the evidence only pays off if something happens with it before the next paper.
- Be ready to show a governor or LA reviewer one clear evidence-gathering cycle rather than several disconnected mock folders.
Disclosure: Both voices on this show are synthesised (NotebookLM, Google); the research, reading and editorial judgement are done by a serving practitioner.
Strand: Leaders
Free printable one-pagers for this episode, one for staff and one for parents and carers.
Get the free newsletters · Listen on RSS.com · Watch on YouTube
Full transcript
Read the transcript
This is the education commute. Both voices on this show are synthesized. The judgment isn't. Irm, I want to paint a picture for you. And if you are listening to this on your way into school, well, it's a scene you probably know far too well. Imagine it's half term. Oh, yeah, the classic half term scene. Exactly. You are sitting at your kitchen table. The coffee in your mug went cold. about, I mean, probably two hours ago, and staring back at you, just dominating the entire table, is this mountain of 92 mock exam scripts waiting to be marked.
Just looming over you. Yeah, looming. You've got the mark scheme open on your laptop, a red pen in hand, and you're basically trying to figure out how to physically get through this pile before Monday morning. And, you know, the physical weight of those papers, it's usually just the start of it. There's this immense mental load that comes with mock marking. Absolutely. Largely because of the speed at which schools demand the data to be turned around. I mean, you are often expected to mark, moderate, input the grades into the school system and then write a summary report for your head of department all within a matter of days.
Relentless. It is. And this pressure is sort of compounded by the fact that the students themselves are deeply anxious about the numbers you're gonna write on the front of those booklets. That anxiety is just a constant undercurrent, isn't it? And for a lot of you listening right now, this isn't even a one -off event. This is likely what, the third mock cycle since September? Very likely, yeah. Your department probably ran a baseline set in November. Now you are buried under the February half -term pile. Yeah.
And... If you look at the staff room whiteboard, the leadership team is probably already penciled in a fourth cycle to run just before Easter. Yeah, it's just the water we swim in at this point. Right. The common staff room orthodoxy is that doing three, maybe even four mock cycles a year is just, well, it's just common sense. The prevailing mantra is that practice makes perfect. That more practice under exam conditions inevitably equals better preparation on the actual day. That's the assumption Yeah, but I have to push back a bit on the idea that this is inherently flawed I mean surely exposing students to the friction of the exam hall, you know the strict timings the absolute silence Surely that pressure cooker environment is necessary.
You're saying the entire staff room consensus is genuinely wrong here Well, when we test that consensus against what Ofqual actually requires right now, the foundation of that assumption looks incredibly fragile. Really? In what way? Erm, the issue isn't that practice is bad. The issue is that we have confused the act of measuring a student with the act of teaching a student. We really need to move away from the idea of mocks merely as these sort of practice runs and look at what their actual official government purpose is within the system.
Okay, so what does Ofqual say their purpose actually is? Well, if you read the Ofqual Guide for Schools and Colleges, which by the way was published on the 15th of January, 2026, and then updated on the 3rd of February, 2026, you will find that mock exams have a very specific named purpose. Right. Ofqual refers to this purpose as long -term resilience arrangements. Long -term resilience arrangements. Wow. I mean, that phrase sounds like it was drafted in a Whitehall committee room, doesn't it? It really does. Very bureaucratic.
But in practical terms, for the teacher actually looking at those 92 scripts, this is fundamentally about contingency planning, right? Spot on, yes. Like, if you look back to the disruption of the pandemic, the system was caught without a safety net. So now, OCL explicitly requires schools to have an established evidence base of student performance. Exactly. The logic being that... In the unlikely event that summer exams cannot go ahead as planned, grades can still be determined fairly and robustly. So we're essentially talking about the evidence base for teacher -assessed grades.
That is the core function. But notice the specific wording Okul uses there. These are long -term resilience arrangements. Right, long -term. Yeah, this is not a temporary one -off pandemic fix that is going to quietly vanish from the statute books next year. It is an ongoing structural part of the education system. Ofqual requires schools and colleges to have this evidence gathered, moderated, and ready on standby every single academic year. Okay, but... If Ofqual requires this evidence to be gathered every year, it raises a pretty massive logistical question.
Does Ofqual actually dictate when we have to gather it? Ah, now that is the question. Because, I mean, every school I know treats the massive January mock season like it is a non -negotiable law of physics. We freeze the entire curriculum, we clear out the sports hall, and we run three weeks of intensive exams. Yeah, and that is perhaps the most pervasive myth we need to address today. There is absolutely no statutory date dictating when mock exams must happen. Wait, none at all. Nothing. Nothing in the off -school guide sets a month or a specific seasonal window.
The whole January mock season that basically takes over the entire school calendar, it is purely a timetabling convention. Just a convention. Exactly. It is not a government compliance rule. Schools naturally gravitate towards January because, well, it sits nicely after the Christmas break, a substantial amount of the curriculum has been covered, and there's theoretically enough time left in the academic year to act on the results. but Ofqual does not say you must run your mocks in January. That distinction is so crucial. If Ofqual sees mocks essentially as an insurance policy kept on standby, rather than like a mandatory training montage, it really changes the conversation.
It absolutely does. But following that logic, If it is an insurance policy, wouldn't buying three policies running three distinct mock cycles across the year mean we are extra covered? If I am a head teacher wanting to ensure my teacher -assessed grades are bulletproof, wouldn't Ofqual prefer me to have a mountain of evidence rather than a molehill? It's a very logical assumption, but it is the exact opposite of what Ofqual advises in their current guidance. The exact opposite. Yes. And this brings us directly to why having three or four of those insurance policies is actually a massive systemic problem.
To answer your point about building a mountain of evidence, I am going to quote the Ofqual guide directly. OK, let's hear it. This is verbatim from Ofqual. Schools and colleges should avoid over -assessment, with one set of mocks likely to be sufficient for evidence purposes. Wait, I just wanna pause there. if Ofqual is explicitly framing one set of mocks as likely to be sufficient. It really calls into question why school leadership teams are still willing to sacrifice weeks of teaching time for sets two and three.
Exactly. Because we aren't just talking about an hour in a classroom. We're talking about paying external invigilators, displacing physical education classes because the sports hall is full of desks, and throwing the entire school timetable into chaos. If one set satisfies Ofqual, the institutional inertia driving these extra mock cycles is staggering. Staggering and costly. The fallout is severe on multiple fronts. You are losing weeks of instructional teaching time, which is arguably the most valuable resource a school has. You are generating immense stress for students who are already feeling the pressure of their final year.
And going back to the teacher at the kitchen table, running a third mock cycle in February isn't gathering any new evidence that Ofqual actually asked for. Right. You satisfy the resilience arrangement with the first set. The subsequent sets are just generating extra unmandated marking. Okay, but if we strip away the extra mock cycles, the pressure is entirely on that single paper to actually do some serious heavy lifting. Yeah. I mean, that lands incredibly hard for anyone marking their 90 -second script right now. If one set is genuinely sufficient for off goal, then the disruption of running them is huge.
How do we justify that one set? What makes the disruption actually worth it? That's the pivotal question. If we only run one mock, it cannot just be an exercise in compliance. The Education Endowment Foundation's recent work on diagnostic assessment completely changes the framework here. Okay, the EEF. Yes. If we are reducing the quantity of exams, we have to radically increase the quality of how we use the data from the single exam we do run. The Education Endowment Foundation, or EEF, has published a series of new case studies looking exactly at this making effective use of diagnostic assessment.
I feel like whenever the EEF comes up, everyone immediately wants to know the impact score. Right, and we need to be radically honest about the numbers here. We often look to the EEF toolkit for a very specific quantified impact metric, but there is no headline plus X month's impact figure for diagnostic assessment in this context. Got it. So no magic number to quote. Exactly. It is a highly well -evidenced practice, but it is not a numbered toolkit strand, and we are not going to sit here and invent a metric for it.
The value lies in the methodology, not a shiny impact number you can pin to a staff room notice board. Which makes sense. So without a shiny impact number, the discipline here is purely functional. The methodology requires a shift in how we view the marked script itself, right? A mock exam paper has to do two very specific jobs once the red pen is put down. It cannot just be a cumulative grade on a cover sheet that gets inputted into a spreadsheet and then filed away. No, definitely not.
The EEF highlights that diagnostic assessment from these scripts must identify learning gaps on two distinct levels. Yes, macro and micro. Right. So the first is the macro view, the class level. This involves looking at the data to identify whole topics, concepts, or skills that the majority of the class fundamentally misunderstood. Yeah. Like if 60 % of the room failed a specific question, that topic requires explicit whole class reteaching before the teacher can even think about moving on. to the next unit on the syllabus. That's the macro.
And then the second job is at the micro view, the individual level. How does that look in practice? Well, this means drilling down past the overall grade to identify specific pupils who have very localized idiosyncratic gaps. I mean, two students might both score 40 % on the overall paper, but the actual knowledge they are missing could be completely different. Oh, right. One might have failed algebra and the other failed geometry. Exactly. Identifying those individual gaps allows for highly targeted intervention rather than just putting all the 40 % students in the exact same generic revision session.
It functions exactly like medical triage in a busy A &E department. You wouldn't walk into a waiting room full of patients and just hand everyone with an injury a generic bandage, regardless of whether they have a broken ankle or a migraine. No, you'd be a terrible doctor if you did. Exactly. You have to use the diagnostic data to sort the specific injuries before you can treat them. Handing a student a script with a grade 5 written on the front tells them they are injured, but it offers absolutely zero information on how to heal.
Such a good analogy. The true value of the mock exam is the rigorous triage process that happens after the marking is finished. And grounding this in Monday morning reality is where we see the culture actually shift. The EEF case studies provide excellent blueprints for what this looks like across different phases and subjects. For instance, if we look at how a secondary math department might restructure around this, Walton High School and Staffordshire provides a brilliant example. What did they do? They tied diagnostic assessment directly to retrieval practice.
In a secondary math setting, I imagine dealing with hundreds of students across an entire year group. Criage requires deep organization. Very deep. This usually takes the form of a detailed topic grid, doesn't it? Every single question on the mock exam is tagged to a specific specification point. When the department sits down for their post -mock meeting, they aren't looking at generic averages. No, averages hide the detail. Right. They are looking at the grid and identifying that, say, question 14 on algebraic fractions was answered incorrectly by 40 % of the entire cohort.
And this is where the retrieval practice comes in. Walton High doesn't just note that algebraic fractions are a weakness and then plan a generic revision lecture for April. They act on it immediately. Yes. They immediately pull that specific gap out of the mock data and inject it straight into the short -term curriculum plan. Algebraic fractions become the mandatory starter activity, the retrieval practice, for the next three weeks of lessons. Wow. The mock data actively dictates the daily teaching. The gap is identified, and the reteaching is scheduled immediately, spaced out over time, to ensure it actually embeds in the student's long -term memory.
That's incredibly powerful. Let's apply that same discipline to a primary setting then. The EEF highlighted Burton and Primary Academy in Suffolk, looking at low -stakes quizzing, and South Shore Academy in Blackpool, which focused heavily on reading. Yeah, two great examples. If I am a year six teacher who has just finished marking a stack of reading mock papers, following this diagnostic model means I am completely abandoning how I physically handle the papers. I'm not sorting them alphabetically for the mark book, and I am certainly not sorting them from highest score to lowest.
Definitely not. You are sorting those physical papers into literal piles on your desk based on the specific misconceptions they reveal. Just literally putting them in different piles. Yeah. If you are assessing reading comprehension, Pile A might be the pupils who fundamentally lack the background vocabulary to access the text. Pile B might be the pupils who can decode the words perfectly. They completely missed the inference questions. They cannot read between the lines. And Pile C. Pile C might be the pupils whose answers suggest they simply ran out of time and started guessing on the final page.
That is brilliant. So before the children even walk into the classroom on Monday morning, the teacher has three distinctly different intervention groups ready to go, built entirely from the mock data. Ready to treat the specific injury. Right. You avoid the trap of giving a whole class lecture on time management when only a third of the class actually struggled with pacing. Exactly. And this approach is universally applicable, especially when we look at specialist settings or students with special educational needs and disabilities. How does it translate there?
Well, Burton and Primary Academy specifically focus their case study on pupils with education, health and care plans, or EHCPs. The execution of the assessment might adapt, but the core diagnostic discipline remains absolute. Yeah, because in a specialist setting, a student might be taking an adapted mock paper. They might require symbol support, a reader, a scribe, or 25 % extra time. Right. The crucial point here is that an adapted paper is still robust diagnostic evidence. If a pupil with an EHCP takes an adapted paper and the results reveal a persistent gap in their phonics knowledge, that gap still requires a highly targeted reteaching plan.
We cannot fall into the trap of lowering our expectations just because the paper was adapted. The evidence is the evidence, and the professional response is targeted intervention, not a sympathetic lowering of the bar. Spot on. It requires treating the adapted assessment with the exact same diagnostic respect as any other paper in the school. The focus remains squarely on identifying the barrier to learning and systematically dismantling it through teaching. But this entire methodology, I mean, the topic grids, the physical sorting of reading papers, the targeted EHCP interventions, it all relies on one fundamental commodity, doesn't it, time?
Time, yes. Which brings us right back to Ofqual's warning against over -assessment. Because if a school immediately jumps into a second or third mock cycle... As many have penciled in for February or Easter, they are testing the students again before the teachers have actually had the time to execute the reteaching plan from the first set. It is literally like running a diagnostic scan on a car. seeing a flashing engine light for transmission failure, printing the error code, clearing the dashboard, and then just driving the car for another month, hoping it fixes itself before you run the scanner again.
That's hilarious, but so true. They were just generating the exact same error codes. It is entirely counterproductive. You are measuring the exact same gaps. If you run a mock cycle in November, spend all of December marking it, and then shove the students into another mock cycle in February without spending six solid weeks ex - explicitly reteaching the gaps you found, while the entire exercise is futile. Completely futile. You are generating immense stress, burying your staff and marking, and you are not gathering any new useful data for your off -goal or resilience arrangements.
It completely dismantles the More is Better orthodoxy. It forces schools to shift their focus from the logistical act of sitting the exam to the pedagogical act of responding to the exam. If you use the first set of mocks properly as a diagnostic tool, you will have more than enough reteaching work to keep you busy until the actual summer exams arrive. And that response is the only thing that actually moves the needle on student learning. It is the targeted reteaching, not the repetitive testing, that builds the genuine academic resilience Ofqual is looking for in a student, and the systemic resilience they require from a school.
We have covered a vast amount of ground today. navigating between official guidance and classroom reality. I want to make sure we distill this down into clear, actionable points for everyone listening on their commute. Let's recap the four key takeaways we have pulled from the sources. Sounds good. First, mock exams have a current, named official purpose in the system. They are the contingency evidence for teacher -assessed grades mandated by Ofqual's 2026 guide. They are a structural requirement, not a temporary pandemic measure. Second, That exact same Ofqual guide actively warns schools against the culture of over -assessment.
They state verbatim that one set of mocks is likely to be sufficient for evidence purposes. Ofqual is setting a ceiling to protect instructional time, not a floor that schools need to scramble over. Third, despite what the rhythm of the school calendar might suggest, January is just a timetabling convention. There is absolutely no statutory compliance requirement to run mocks in January. And fourth, the true value of any mock exam does not come from the conditions under which it is sat or how many times you repeat the process.
The value comes entirely from using the marked scripts for rigorous diagnostic assessment, sorting the data by gap, and executing a detailed reteaching plan at both the whole class and individual level. It's entirely about the quality of the pedagogical response over the sheer quantity of exam practice. Before we wrap up today, I want to leave you with a final thought to mull over as you walk through the school gates and head to your department office. It's a challenging one. Next time you were handed a timetabling spreadsheet with three separate mock cycles mapped out across the year, ask yourself, are you genuinely building student resilience or are you just building a heavier folder of data?
If a governor or local authority reviewer asked to see your Ofqual mandated evidence gathering arrangements tomorrow, could you show them one clear set of mocks connected to a well -documented reteaching plan? Or would you just point to three separate spreadsheets of grades with absolutely no teaching plan connecting them? If it is the latter, it is deeply worth fixing that disconnect before you clear the sports hall for the next set of exams. The evidence strongly suggests that having one meaningful diagnostic conversation about the data is far more powerful for student outcomes than subjecting them to three superficial practice rounds.
Everything we have discussed today, from the specifics of the 2026 of Sokol guide to the operational details of the EEF case studies at Walton High and Burton End, is linked in the show notes. We also have a printable one -pager summarizing all of this available at theeducationcommute .co .uk. It is perfect for leaving on the staffroom table or sharing with your head of department who might be wondering why you are suddenly questioning the mock calendar. Have a fantastic day and I wish you a safe trip in.
See you at the gates.
More episodes
Newer: Every Universal Credit Household Now Qualifies For FSM
Older: No 'Plus X Months': Five Meta-Analyses On Maths Manipulatives
All episodes