Word vs. Sentence Transformation Drills: Why the Difference Matters for Grading, Not Just Design
August 3, 2026 · Writing, no kidding
"Transformation Task" is one exercise type in the builder, with a toggle between Word Formation and Key Word Transformations. It's easy to treat the two as cosmetic variants of the same thing — swap a root-word gap for a keyword-and-blank format, same idea. They're not graded the same way underneath, and that difference isn't trivia. It changes what you need to do when you write the answer key, especially if you're using Autopilot to publish scores without reviewing them first.
What each format is actually asking a student to do
Word transformation (Cambridge B2 First "Use of English" Part 3) gives a passage with numbered gaps. Each gap shows a root word in capitals — say, CHALLENGE — and the student has to reshape it to fit the sentence: a different part of speech, a prefix or suffix, a plural, a tense change. There's one blank, one root word, and (in the intended design) one correctly transformed answer.
Sentence transformation, or Key Word Transformation (Part 4), works differently. The student sees a complete original sentence, a keyword they must use unchanged, and a second sentence with part of it blanked out. Filling the blank correctly means the second sentence ends up meaning exactly the same as the first — which usually means restructuring the whole clause around the keyword, not just swapping one word for another.
Both formats test whether a student's grammar is flexible enough to reshape a thought instead of just recognizing one. That's where the similarity ends.
Why word transformation gets graded strictly, on purpose
A word-transformation gap has a single defensible correct form. "Challenging" is right; "challengeful" or "challenge" isn't, no matter how close it looks. Because of that, grading checks the student's answer against the correct answer exactly — not the more forgiving fuzzy match used for quiz short-answers, which tolerates small typos and near-misses. The exactness is the point here, not a limitation: a small-edit-distance mistake in a word-transformation answer is often the exact mistake the gap is testing for (the wrong suffix, the wrong word class), so a grading pass that quietly forgave it would be checking less than it looks like it's checking.
Why sentence transformation doesn't work the same way — until it has to
A sentence-transformation answer can't be checked by exact string match, because there's rarely only one correct phrasing. "I haven't seen her for a long time" and "It's a long time since I last saw her" can both satisfy the same keyword-and-meaning requirement, worded differently. So by default, sentence-transformation answers are evaluated for meaning: does the keyword appear unchanged, is the sentence grammatically sound, does it mean the same as the original. That's a judgment call, and it's handled by AI evaluation when you review a submission yourself.
It stops being handled that way the moment Autopilot is switched on for an assignment. Autopilot publishes scores straight to students without a teacher looking at them first — which means an AI-judged score can't be part of what it publishes; a human has to be the one applying judgment to anything that isn't a deterministic check. So for any assignment where Autopilot will auto-publish, sentence-transformation grading falls back to matching the student's answer against exactly the model answers you supplied — the same kind of strict, literal check as word transformation, just applied to a list of acceptable phrasings instead of a single form. The in-app description on the Autopilot toggle says this outright: scores publish without a review step, and sentence-transformation answers are matched strictly against your model answers, with AI evaluation running only when you grade manually.
What this means for how you actually build the exercise
If you're grading by hand — reviewing each submission before it's returned — the AI evaluation step does the paraphrase tolerance for you, and the exact wording of your model answers matters less. Generate the exercise, skim the draft for a keyword that doesn't force real restructuring or a sentence pair that reads awkwardly, and the meaning check handles the rest.
If you're running the assignment on Autopilot, that safety net isn't there for sentence transformation. The model-answers list generated for each item is what gets checked against, in full, with nothing softening a legitimately correct answer phrased slightly differently. Take the "long time" example above: if the model-answer list only holds "a long time" and a student writes "such a long time," that's a legitimate paraphrase a human reviewer would accept without thinking twice — but under Autopilot, with no AI evaluation running, it gets marked against the list as it stands. Before turning Autopilot on for an assignment with a sentence-transformation task, it's worth actually reading the generated model answers for each item rather than trusting the first draft, and regenerating an item if the phrasing options look thin. Word transformation doesn't need this same pre-Autopilot check, since it's graded the same exact-match way regardless of whether a teacher reviews it first or not.
None of this makes either format worse — it's just a reason to know which grading path is running before you decide how closely you need to check the draft. The mechanics for building both, including the Word Formation / Key Word Transformations toggle and where the CEFR level and gap-count settings live, are covered step by step in how to create word and sentence transformation exercises. For how transformation drills compare to gapped text and matching as a design choice rather than a grading one, see gapped text, matching, and transformation drills; for the reading-comprehension side of the same "which exercise actually tests what" question, see designing ESL reading comprehension passages and questions.
Build a transformation drill with full model answers. Start Free — no credit card required.
Ready to try this in your own classroom?
More from the blog
Gapped Text, Matching, and Transformation Drills: When to Use Each ESL Exercise Type
What gapped text, matching, and transformation drills each actually test, which CEFR levels suit them, and common design mistakes to avoid.
Designing ESL Reading Comprehension Passages and Questions That Actually Test Understanding
Most comprehension quizzes test scanning, not understanding. What good questions test by CEFR level, common mistakes, and how to generate a levelled set.