OPIc AI: Who Grades You, and Can AI Prepare You?
The OPIc is computer delivered but rated by certified humans, with blind double rating on Official tests. What that means for using AI to prepare.
Two questions hide inside the phrase OPIc AI, and they need separate answers. The first is whether the thing grading you is a machine or a person, because that decides whether the test can be gamed at all. The second is whether an AI conversation partner is a real way to prepare.
One line of housekeeping before that. This page is about the ACTFL Oral Proficiency Interview Computer. It is not about opic.ai the photo sharing service, the legaltech company with a similar name, or Opik, the LLM monitoring tool from Comet.
If you want to practise unscripted speaking while you read, SpeakShark is free with three AI conversation sessions a day and no card.
In this guide: who grades you · what the machine does · how human rating works · ACTFL and machine scoring · why AI practice is legitimate · where it crosses the line · unpredictability · auditing a predicted grade · AI during the test · a practice method · how we researched this guide · FAQ
Key takeaways
- A computer program selects and delivers your questions. That is the full extent of the machine's role in the test itself, and rating is done by ACTFL certified humans.
- On an Official or Certified OPIc, at least two certified raters must independently agree before a rating is released, and a third rater arbitrates blind if they do not. Commercial OPIcs are single rated.
- ACTFL does run machine scoring, on the AAPPL. Spanish presentational writing came first in 2023, then Spanish interpersonal listening and speaking on 20 January 2026, running alongside human raters rather than replacing them.
- The opening "tell me something about yourself" prompt is a warm up and is explicitly not rated, so one of the most heavily rehearsed answers in OPIc preparation is scored zero times.
- Running an AI assistant during the live test is not a lower grade risk. The Examinee Handbook says that if the rater detects you are receiving assistance, the test will not receive a rating.
Is the OPIc graded by AI or by a person
By a person, and this is not a close call.
The OPIc feels like a machine test from start to finish. You sit alone, an avatar speaks to you, a microphone records you, and nobody in the room reacts to anything you say. It is completely reasonable to assume a machine is on the other end. That assumption is wrong at every stage after the recording stops.
ACTFL's own Familiarization Guide describes the purpose of the test as obtaining a speech sample that a rater can evaluate against the ACTFL Proficiency Guidelines in order to assign a rating, and states that the recordings of the test taker's responses are made available electronically through a secure internet site to ACTFL certified OPIc raters.
Read that sentence again with the search intent in mind. The document that explains the test to candidates says the destination of your audio is a certified human being with a login, not a scoring engine.
The ACE report is more explicit still. A certified OPIc rater listens to the sample holistically, reaches a preliminary rating, compares the sample to the descriptions in the Proficiency Guidelines, selects the best match between the sample and the descriptors, and enters that rating into the online system.
That is a criterion referenced judgement, not a computable score. There is no point total, no deduction per error and no threshold you can reverse engineer, which is why the "how many mistakes am I allowed" question has no answer.
What the machine actually does
The machine's job is real, and it ends before anyone forms an opinion of your English.
| Stage | Machine or human | What actually happens |
|---|---|---|
| Background Survey and Self Assessment | Machine | Your answers set the topic pool and generate one of five forms |
| Question selection | Machine | A carefully designed computer program picks prompts from an item bank |
| Question delivery | Machine | Pre-recorded audio, played by a virtual avatar |
| Your responses | You | Recorded and uploaded to a secure internet site |
| Rating | Human | A certified rater listens holistically and matches the descriptors |
| Second rating, Official and Certified tests | Human | A second certified rater rates blind, and must agree |
| Disagreement between the two | Human | A third certified rater performs a blind arbitration |
Three details from that table are worth pulling out, because they are where the AI assumption comes from.
The avatar generates nothing. The ACE report states that prompts are selected from an item bank of pre-recorded prompts organised into testlets by topic and proficiency level, and that each OPIc explores four to five topics depending on the form. Ava is a playback surface. She is not listening to you and she cannot follow up on anything you said.
The order is random. The Familiarization Guide says the avatar asks randomly selected questions from within the predetermined pool of prompts. Your topic areas are predictable from your own survey answers. The sequence is not.
Identical inputs do not produce identical tests. ACTFL states that even if two test takers select the same combination of Background Survey and Self Assessment responses, the resulting test would not be the same, because of the size of the item bank and the selection algorithm.
You will still see it claimed that Ava adapts her follow ups to what you just said, the way a chatbot would. Nothing in the primary documents supports that, and the two documents that describe prompt selection both describe a fixed pool. If it matters to your preparation, ask the exam body directly rather than trusting a forum post.
How the human rating works
The rating pipeline is not one person's opinion, and it is also not always two.
Language Testing International, the organisation that sells and administers the test, draws a distinction that decides how many people listen to you. Commercial OPIcs are single rated, with one ACTFL certified rater identifying the proficiency level met. On an Official or Certified OPIc, the sample is rated autonomously and independently by at least two ACTFL certified raters, and their independent ratings must agree before an Official rating is released.
The ACE report describes what happens when they do not agree. The OPIc is blindly second rated by another certified rater following the same protocol. If the two ratings agree exactly, the rating is finalised. If they differ, the sample goes to a third rater for a blind arbitration, and ACTFL Quality Assurance monitors interrater reliability closely.
Notice what is missing from that description: any automated tie break. When two trained humans disagree about your English, ACTFL's answer is a third trained human.
You will find Korean coaching blogs stating flatly that three raters score every OPIc. That is not what the source says. The third rater is an arbitrator invoked only on disagreement, and the base count depends on whether your sitting is commercial or Official. Which bucket a given Korean corporate administration falls into is not something we could establish from any primary source, so ask the administrator rather than assuming.
What the rater weighs is also not countable. The Familiarization Guide describes proficiency as defined by four criteria ACTFL labels FACT, evaluated holistically based on overall performance, with the rating awarded on how evidence of all the criteria contributes to a description of global proficiency. That is the structural reason no transcript based tool can reproduce the judgement. The tool is counting things the rater is not scoring.
ACTFL does use machine scoring, just not on the OPIc
Here is the part that reframes the whole question.
The comfortable story is that ACTFL is a traditional body that has not got around to AI yet. It is not true. ACTFL has already shipped automated scoring. Its announcement, dated 16 January 2026, says the automated scoring system was first introduced in 2023 for the AAPPL Spanish presentational writing component, and that it extended to live Spanish interpersonal listening and speaking tests on 20 January 2026.
The design of that launch is the interesting bit. ACTFL states that the integration of machine scoring alongside human ratings means all Spanish interpersonal listening and speaking tests will now be double rated. The machine was added as a second opinion, not as a replacement.
The OPI and OPIc are not mentioned in that announcement at all.
So the correct sentence is not that ACTFL has avoided automation. It is that ACTFL has deployed automation on one assessment, deliberately paired it with human raters, and has not put it anywhere near the OPIc. Language Testing International has argued publicly in the same direction, suggesting that rather than sidelining human interaction, AI makes the gold standard Oral Proficiency Interview more significant as a measure of true oral proficiency involving negotiation of meaning.
For you, that changes the preparation question completely. You stop asking how to beat an algorithm and start asking what a trained listener is listening for. Our guides to Intermediate High and Advanced Low go through those descriptors band by band, and OPIc IM2 and who requires it covers a grade that published Korean recruitment notices name and restate each hiring cycle, such as Hyundai Motor asking for OPIc IM2 or TOEIC Speaking 130, and Lotte Chemical asking IM2 for research and production roles but Intermediate High for product planning and sales. Those are examples from individual notices, not a national standard, which is why the published picture is a grid rather than a single corporate cutoff.
Why an AI partner is legitimate preparation
The case for practising with an AI is not our marketing claim. It is ACTFL's own advice, and it is unusually blunt.
The Examinee Handbook says the best advice for doing well on the OPIc is practice, and repeats the word three times. It then rules out the thing most candidates do in the final week: last minute language learning, grammar review or vocabulary practice will most likely not improve your final results. Knowing more about the language will not affect your rating unless it reflects on what you can do.
The handbook closes with the clearest statement available on this: practising your listening and speaking skills as much as you can in the target language is the best preparation for a successful OPIc assessment.
Volume of unscripted listening and speaking is the prescription. That is exactly what an AI partner is good at supplying, for reasons that have nothing to do with intelligence and everything to do with availability. It is there at midnight, it does not get bored of your third attempt at the same story, and it costs nothing to fail in front of. Talking to an AI in English free and the benefits of an AI English tutor cover that trade in general, and AI tutor versus human tutor covers what you give up.
Be honest about the ceiling, though. An AI partner supplies practice hours. It cannot supply a rating, and neither can we. No AI tool, ours included, can issue an ACTFL rating, and no third party preparation app has any published endorsement or partnership from ACTFL or Language Testing International that we could find. If an app implies otherwise, that is a claim to verify with the exam body, not a credential.
Where AI preparation crosses the line
The same handbook that recommends practice draws a hard boundary, and the penalty on the wrong side of it is worse than most candidates realise.
Your responses must be authentic. The handbook states that you should not try to memorise responses before taking the OPIc, that preparing a response or using responses from online sources or books means you will not receive an accurate rating, and that because raters are experienced at identifying rehearsed responses, using them may leave you with no rating for your test.
Not a lower grade. No rating at all. ACTFL's separate tips page reinforces it from the other side: an abundance of rehearsed language or memorised speech may prevent a rating beyond the Novice proficiency level, and cramming will not improve a proficiency rating because proficiency develops over time.
So the line runs through what you ask the AI to do, not through whether you use one:
- Asking an AI to talk with you produces practice hours, which is what ACTFL recommends.
- Asking an AI to write your answers produces exactly the rehearsed material ACTFL trains raters to detect.
There is a specific and expensive version of the second one. The self introduction is one of the most heavily rehearsed answers in OPIc preparation, and the Familiarization Guide states that when the avatar opens with "Tell me something about yourself", this serves as a warm up and an opportunity to begin using the language, and this warm up activity is not rated.
Every hour spent polishing that answer with a chatbot is an hour spent on a question the rater does not score. Why scripted dialogues do not make you fluent is the longer version of this argument.
The unpredictability argument, in one paragraph
There is a second reason scripts fail, which is that you cannot predict what you will be asked. Your prompts are drawn from a pool your own Background Survey creates, the order is randomised, and two candidates with identical survey answers still receive different tests. We have covered this in full in OPIc questions and what to expect, including the five question shapes that cover every prompt and how to practise without a question list. If that is what you came for, read that one instead of this section.
How to audit an AI tool that promises a grade
Plenty of apps will show you an estimated OPIc level after a mock session. Some are honest tools with an optimistic label. Others are a number generator with a progress bar. Here is how to tell, in order of how much it tells you.
- Ask which form it assumed. Your rating range is set by which of five forms your Self Assessment triggers, except where a requesting client fixes the form instead, which the ACE report notes happens in some administrations and covers exactly the corporate sittings this post discusses. Form 1 cannot award above Intermediate Low, and Form 4 returns nothing below Intermediate High. A tool that never asked which form you sat is predicting a band without knowing which bands are available to you. Our guide to which self assessment level to pick has the full table and why it decides your ceiling.
- Ask what it scored. ACTFL raters weigh four criteria holistically across the whole sample. A tool reporting error counts, filler word tallies or a words per minute figure is measuring things that are visible in a transcript rather than the thing being rated.
- Ask for the validation data. Any real predictor would have correlation figures against actual ACTFL ratings. We could not find published validation for any vendor's predicted OPIc grade.
- Check the endorsement claim. No third party AI preparation tool that we could find has a published ACTFL or Language Testing International endorsement. Implied official alignment is a marketing choice, not an accreditation.
- Watch what the number does to your behaviour. A predicted grade that makes you practise more is doing something useful. A predicted grade that makes you stop practising because the app said IH is doing harm with a fabricated input.
Our own position, stated plainly so you can hold us to it: SpeakShark gives you feedback on pronunciation, grammar and fluency during a conversation. It does not issue OPIc levels, because we cannot, and neither can anyone else outside ACTFL's rater pool. If you want a rough starting point for your own planning rather than a prediction, test your English speaking level free is the honest version of that exercise.
Using AI during the test voids your result
This one is short because the rule is short.
The Examinee Handbook states that during the interview you are not permitted to review documents or dictionaries, or to ask for help, and that you must rely exclusively on what you can do in the language on your own. If the rater detects that you are receiving assistance, the test will not receive a rating.
A second screen running an AI assistant is receiving assistance. Text on a monitor being read aloud is exactly the kind of delivery a rater trained on rehearsed speech listens for, and the cost is not a lower band. It is nothing, plus the wait before you can sit again.
The reason it gets caught is worth understanding. Read speech has a different rhythm from produced speech: the pauses fall in the wrong places, the syntax is too tidy for the fluency around it, and the register does not drift the way spontaneous speech does. Raters listen to samples for a living.
A practice method built from the testlet
If a human is listening, imitate the thing that human hears. The ACE report describes the internal structure of each topic block, and it is a good template for a practice session.
Within a testlet of two or three prompts, you answer at least one level check, a prompt that elicits functions at the floor or major level you can sustain, followed by either more level checks or a probe, a prompt that targets a function at the next major level, or the ceiling where you cannot sustain performance.
The test deliberately pushes you one level above where you are steady. Practice that only stays comfortable never touches the part of your range the rating actually turns on.
A session that mirrors it:
- Pick one topic and stay on it for ten minutes. The OPIc explores four to five topics in blocks, not one question each. Topic hopping trains a skill the test does not test.
- Start at your level check. Describe the topic plainly. This is the floor, and it should feel easy.
- Force the probe yourself. Push into the next function up: compare the topic across time, then handle something going wrong inside it. Role play scenarios are the fastest way to reach the complication reliably.
- Notice where you stop being able to sustain it. That boundary is the useful data from the session, and it is the only thing about your level a practice tool can tell you honestly.
- Repeat the boundary, not the floor. Rehearsing what you already do well is the most comfortable way to waste a week.
- Do it daily and unscripted. Practising English speaking alone covers building tolerance for speaking into silence, which the OPIc requires and ordinary conversation does not, and SpeakShark's free tier gives you a partner that answers back.
Underneath all of it, clarity outranks range. Six pronunciation fixes in fourteen days targets the sounds that block comprehension, and how to improve English pronunciation is the longer route.
If you are still deciding whether the OPIc is even the test you need, OPIc versus TOEIC Speaking covers the alternative and the TOEIC Speaking test guide covers that route on its own. If you are choosing practice tools more broadly, the best English speaking apps and apps for conversation practice compare the field, and using ChatGPT voice mode for English practice covers the free general purpose option.
How we researched this guide
The rating chain comes from three primary documents read directly rather than from summaries. The delivery mechanism, the destination of your recordings, the not rated warm up, the random prompt selection and the four holistic criteria come from the ACTFL OPIc Familiarization Guide, 2024 edition. The holistic rating protocol, the blind second rating, the third rater arbitration, the item bank of pre-recorded prompts and the level check and probe structure come from ACTFL's ACE report on the OPIc, Part A General Test Information. The commercial versus Official and Certified rating distinction comes from Language Testing International's own product page. The rehearsed response warning, the practice advice and the outside assistance rule come from the ACTFL OPIc Examinee Handbook. The AAPPL machine scoring launch comes from ACTFL's news announcement.
Three things could not be verified, and we have not stated any of them as fact. We could not establish whether the Korean mass administration is single rated or double rated, because Language Testing International draws the distinction without naming which bucket that administration falls into, so the post says to ask the administrator.
The other two concern the AI preparation market. We could not find any published validation data, correlation figure or ACTFL recognition for any vendor's predicted OPIc grade. And we could not find any statement from ACTFL or Language Testing International endorsing or partnering with a third party AI preparation app, which is why this post states plainly that no AI tool, ours included, can issue an ACTFL rating.
A further set of claims had no measurement behind them, so they are no longer stated here. No survey of competing OPIc pages sits behind this guide, so we do not say what other pages do or do not mention. No source ranks English tests by how computer shaped they are, so the opening now describes how the OPIc feels rather than where it sits in a league table. And no published measurement identifies which OPIc answer candidates rehearse most, so the self introduction is described here as one of the heavily rehearsed answers rather than the most rehearsed one. On the Korean employer side there is no count of recruitment notices anywhere either, so the grades named in this post are examples from individual published notices rather than a national requirement.
One inconsistency worth flagging for anyone reading the sources themselves. The 2024 Familiarization Guide refers to the ACTFL Proficiency Guidelines 2024 for Speaking, while the ACE report carries publication number AAR/OPIc/ACE/I 2020 002 on its cover and benchmarks to the superseded 2012 edition, and it gives both four to five topics and four to six topic areas in different places. The 2023 label often attached to that report is only the file name it is published under. Where the two documents disagree we have followed the 2024 guide and said so rather than smoothing it over.
Practise speaking, from SpeakShark
SpeakShark is an AI English speaking practice app, and it is our pick for OPIc preparation because it supplies the thing the Examinee Handbook puts first, which is volume of unscripted listening and speaking. It will not give you a predicted band, because a certified human rater is the only source of an OPIc rating and we would rather say so than sell you a number. What it does give you is a partner that answers unpredictably, at any hour, so the hours ACTFL says you need are actually available. You talk, the AI answers, and you get feedback on pronunciation, grammar and fluency while the conversation is still going.
Start free with three sessions a day and no card. Limits are on the pricing page, and how it works walks through a full session.
We are a speaking improvement tool. We are not an exam preparation provider and we are not affiliated with ACTFL, Language Testing International, ETS or any exam board. For test format, rules, booking and official practice material, go to the exam body directly. Use SpeakShark to make your spoken English stronger, and use official material to learn the test.
- Site: speakshark.com
- Support: speaksharksupport@gmail.com
- Free tier: 3 AI conversation sessions per day, no card
Sources
- ACTFL OPIc Familiarization Guide: computer selected and delivered questions, recordings sent to certified raters, the unrated warm up, randomly selected prompts, the four FACT criteria
- ACTFL ACE report, OPIc general information: holistic rating protocol, blind second rating and third rater arbitration, item bank of pre-recorded prompts, level check and probe structure
- ACTFL OPIc Examinee Handbook: rehearsed response warning and the no rating consequence, the practice advice, the outside assistance rule
- Language Testing International, what the OPIc is: commercial tests are single rated, Official and Certified tests require at least two agreeing raters
- ACTFL, tips for OPI and OPIc test takers: rehearsed speech may prevent a rating beyond Novice, cramming does not move a proficiency rating
- ACTFL news, automated scoring for AAPPL Spanish ILS: machine scoring launched alongside human raters, making those tests double rated
- ACTFL, 2025 to 2026 AAPPL release and educator resources: context for where ACTFL's automation work sits
- ACTFL, Oral Proficiency Interview Computer overview: official description of the assessment
- ACTFL, OPIc rater certification: confirmation that OPIc raters are certified people rather than systems
- Language Testing International blog, leveraging AI in language education: the administrator's public position that AI raises rather than lowers the value of human rated oral proficiency testing
- Language Testing International, OPIc product page: commercial description of the test and its administrations
- ACTFL OPIc Examinee Handbook, university hosted copy: mirror of the handbook for readers who cannot reach the primary link
- OPIc Korea: the Korean administrator, which we could not read programmatically and therefore could not use to settle whether Korean corporate sittings are single or double rated
FAQ
Is the OPIc scored by AI?
Does ACTFL use machine scoring anywhere?
Can an AI app predict my OPIc grade?
Is using AI to practise for the OPIc allowed?
Should I use AI to write my OPIc self introduction?
Does the OPIc avatar respond to what I say?
Curious how it works? Explore SpeakShark's features or see plans and pricing.
Keep reading
OPIc AL: How to Reach Advanced Low in 2026
Three different OPIc forms can return Advanced Low, and they carry completely different risk. Which one to pick, and what AL actually requires.
OPIc Demo: What the Official Sample Test Shows You
The official OPIc demo is a system check with a short test attached. What it shows you, why it cannot score you, and how to tell it from the fakes.
OPIc Exam: What to Expect on the Day, Start to Finish
What the OPIc exam actually involves: booking, cost, ID, equipment, the screens in order, rating, results, and a retake clock that changes by country.