Where a Real Person Must Listen to You Speak
One regulator refuses any test whose speaking section is computer marked. Others require two examiners, or a second independent rating. Here is who insists on a human.
Speaking is the only skill in a language test that requires somebody to be there. That makes how it is marked a live question, and one regulator has taken a side.
The UK nursing regulator will not accept any test whose speaking section is marked by a computer.
SpeakShark is free for the practice before the person listens, three AI conversation sessions a day with no card.
In this guide: the criterion · what it excludes · two examiners · a second rating · no doubt · a role play · our position · limits · method · FAQ
Key takeaways
- The NMC requires the speaking element to be tested by a person.
- Linguaskill uses hybrid marking; MET grades speaking with certified raters.
- Cambridge speaking tests use an interlocutor and an assessor.
- CELBAN records the speaking test for an independent second rating.
- Aviation requires an examiner to be in no doubt, or record nothing.
The criterion
Among six criteria the Nursing and Midwifery Council publishes for accepting a language test:
The speaking element is tested by a person and not via a computer test.
Six words of it do the work. Not marked by a person, but tested by a person, which reaches the conduct of the assessment as well as the scoring.
It sits alongside another criterion in the same list: that the test must assess English in a healthcare or academic context and must not be a general test. Both point in the same direction, at an assessment that resembles the situation the language will be used in.
We set out all six, and what they exclude, in the NMC wants speaking marked by a person.
What it excludes
The criterion has teeth because part-automated speaking assessment is now normal.
| Test | How speaking is handled |
|---|---|
| Cambridge Linguaskill | hybrid marking, human examiners combined with auto-marking |
| Michigan MET | certified raters grade speaking and writing; other sections automatic |
| Oxford Test of English | adaptive delivery, invigilated and marked by real people |
| OET | a live role play with an interlocutor |
| CELBAN | conducted over video, recorded for a second independent rating |
Those descriptions come from the providers themselves, gathered across this series in Linguaskill Speaking, MET for Australian visas and Oxford Test of English.
A test can be rigorous, secure and four-skilled and still fall outside the NMC's list on the marking question alone.
Two examiners in one room
Cambridge runs its speaking tests with two people in different roles: an interlocutor who conducts the conversation and manages the tasks, and an assessor who observes and marks against the criteria.
That separation is deliberate. Somebody managing a conversation cannot simultaneously give it full analytical attention, and somebody marking is not distracted by keeping it moving.
It also produces a specific benefit for the candidate that machine marking cannot: the person you are talking to is not the person judging you, so the conversation can be a conversation.
We work through how those criteria are applied, using real scored candidate performances, in Cambridge speaking criteria by level and Interactive Communication.
Recorded and rated twice
CELBAN, the Canadian nursing test, takes a different route to the same reliability problem.
Its speaking test runs for 35 minutes over video, with a role play against a visible prompt, and it is recorded so it can be rated independently a second time. Two ratings, one performance, covered in CELBAN speaking for Canadian nurses.
Double rating is the standard psychometric answer to rater variation, and it is expensive, which is why not every test does it.
Ontario and Nova Scotia both put CELBAN's listening requirement at 9 and speaking at 8, the two highest components on their tables, covered in Ontario ranks speaking above reading on every test.
The examiner must be in no doubt
Aviation takes human judgement furthest, and states the consequence.
The UK Civil Aviation Authority requires a candidate to demonstrate all aspects of the expert level descriptors, and says examiners must be in no doubt that a candidate is an expert speaker. If the examiner has doubts about any element, no language proficiency level is recorded and the candidate is referred for formal assessment.
Not a lower level. No level.
It also directs examiners to recorded speech samples at every level, designed to promote rating standardisation between different raters, test providers and regions of the world.
That combination, a human decision plus a calibration library, is the most explicit treatment of rater consistency in this series, covered in the examiner must have no doubt.
A person playing a patient
OET's speaking sub-test is a role play. An interlocutor plays a patient, a relative or a carer, and the candidate has to explain, reassure, check understanding and handle whatever comes back.
Nine bodies accept OET and every one that distinguishes between the sub-tests sets speaking at grade B or its numeric equivalent, covered in one OET grade, nine different passing scores.
What that design assesses is not vocabulary. It is whether the other person understood, which is the same thing aviation's top level measures from a completely different direction.
Our own position on this
We build an AI speaking practice app, so we should say where we stand rather than leaving it implied.
We think the NMC is right.
When a decision is being made about whether somebody can nurse, fly, dispense medicines or treat a patient, a person should be listening to them speak. That is not a hard call, and it does not hurt us, because we are not a test and we never claim to be.
What an AI partner is for is the hundred conversations before the one that counts. Nobody can give a candidate unlimited hours with a human examiner. That is the gap, and it is the only gap we are trying to fill.
Anything we said that blurred the line between practice and assessment would make us exactly the kind of product the NMC's criterion exists to keep off its list.
What to do with this
- Check whether your regulator has a position on how speaking is marked.
- Do not assume a hybrid marked test is refused. Several bodies accept them.
- Do not assume it is accepted either. One publishes a criterion that excludes it.
- Expect a live interlocutor on OET, CELBAN and Cambridge speaking tests.
- Practise against a partner, since every design above assesses interaction.
None of these formats rewards a rehearsed monologue, and none of them improves by reading quietly. Role play scenarios and a daily speaking partner are the practical routes, speaking practice for nurses covers the clinical register, and testing your speaking level free shows where you stand. CEFR levels for speaking explains the scale, speaking is the highest bar in four countries covers why it matters most, and English certificates that expire covers validity.
What we could not verify
We did not observe any of these assessments. Every description is the provider's or the regulator's own.
The NMC does not publish a reason for the human marking criterion.
We did not check how OET's role play is scored beyond the fact that it is conducted live.
We did not verify Cambridge's two examiner model against a current specification in this batch; it is described from the provider's published account of its speaking tests.
Marking arrangements change, and providers revise them without announcement.
We have an interest in this topic. We build AI speaking practice, and our position is stated openly above rather than buried.
This is not registration advice. It is a comparison of published descriptions, all linked in the guides above.
How we researched this guide
The criterion came first and the comparison came from it. Reading a body that publishes six conditions a test must meet, rather than a list of tests, raises an obvious question: what do those conditions exclude.
Answering it meant going back through marking arrangements already gathered one provider at a time earlier in the batch, which is the only reason the table exists. None of those descriptions was collected with this question in mind.
The aviation material was the last piece and the most striking, because it is the only regime that states what happens when the human is unsure. Every other system converts uncertainty into a lower score. That one converts it into no score and a referral.
We wrote our own position into the guide rather than leaving it to be inferred. A company selling AI speaking practice, writing about a rule that excludes computer marked speaking, has an obvious conflict, and the honest handling is to state the view plainly and let a reader weigh it.
Practise speaking, from SpeakShark
SpeakShark is an AI English speaking practice app, and everything in this guide describes a moment we do not participate in.
A person will listen to you. An interlocutor will play a patient, an examiner will decide, a second rater may review the recording. None of that is something software should be doing, and none of it is something we offer.
What we offer is the practice beforehand, in volume, at a time you choose, without a fee per attempt.
You talk, the AI answers what you actually said, and you get speaking feedback while the conversation is still running. The free tier gives basic feedback; the detailed pronunciation and grammar breakdown is on Premium.
Being straight about the limits: a free session runs five minutes with four turns, we issue no certificates, and by the NMC's own criterion we would not qualify as a test, which is correct.
Start free with three sessions a day and no card. Paid sessions run ten minutes with unlimited turns. Limits are on the pricing page, and how it works walks through a session. Sign up here for the conversations before the one that counts.
We are a speaking improvement tool. We are not an exam preparation provider, and we are not affiliated with Cambridge English, the IELTS partners, Pearson or any exam board. For test format, booking and official practice material, go to the exam body directly. Use SpeakShark to make your spoken English stronger, and use official material to learn the test.
- Site: speakshark.com
- Support: speaksharksupport@gmail.com
- Free tier: 3 AI conversation sessions per day, no card
Sources
- NMC, accepted English language tests: the six criteria a test must meet, including that the speaking element is tested by a person and not via a computer test and that the test must not be a general test
- UK Civil Aviation Authority, language proficiency: the requirement that examiners be in no doubt, the rule that doubt in any element means no level is recorded, and the direction to the ICAO Rated Speech Samples Training Aid for rating standardisation
- Cambridge English, Linguaskill: the hybrid marking model combining human examiners with auto-marking
- Oxford Test of English: the statement that the tests are invigilated and marked by real people while using adaptive technology
FAQ
Which regulator requires speaking to be marked by a person?
Are any speaking tests computer marked?
Do any tests use two examiners?
Why would a regulator insist on human marking?
Does the aviation system use human judgement?
Should I avoid computer marked speaking tests?
Curious how it works? Explore SpeakShark's features or see plans and pricing.
Keep reading
Where the Real Rule Hides on a Regulator's Page
Ten places the decisive sentence turned up while reading twenty English requirements, and none of them is where the score is. A method rather than a fact.
One French Route Asks B1 Where the Others Ask C1
Ontario's teaching regulator accepts four French qualifications. Three require C1 and one requires B1, and the two diplomas never expire while the two tests do.
The NMC Wants Speaking Marked by a Person, Not a Computer
The UK nursing regulator publishes six criteria a language test must meet. One of them rules out computer-marked speaking outright, and it decides the list.