AI homework help that does not do the homework
Students already have AI homework help. The question is no longer whether to allow it, but what shape of help you are willing to defend at a parent meeting.
Routines that work in a room with 31 students, one cart of Chromebooks and four minutes of passing time.
Almost every AI classroom idea published in the last two years assumes a room that does not exist: one device per student, all charged, a full period, reliable wifi and a teacher with prep time. Here are six routines built for the actual room. Thirty-one students, a shared cart, four minutes of passing time, and a lesson that has to work whether or not the network holds.
Each routine below has the same shape: what it is, the setup cost, the teacher move that makes it work, and the failure mode to watch for. None of them require a new pacing guide.
What it is. Students arrive, open one problem from yesterday, attempt it, and get hints rather than answers when stuck. Four minutes, no grading, no scores.
Setup cost. One problem, chosen the previous afternoon. If the tool groups by misconception, zero additional prep, because the grouping is the output.
The teacher move. Do not circulate helping individuals. Watch the aggregate. By minute three you know whether today's lesson can proceed as planned or whether six students need five minutes of something else first.
Failure mode. Letting it run to ten minutes. It is a diagnostic, not a lesson. If the problem takes longer than four minutes it was the wrong problem.
What it is. One or two devices on a table at the side. Any student who did not follow the explanation goes there for three minutes and gets the same idea a different way, then returns.
Setup cost. Two devices and a chair. This routine exists precisely because a full class set is not available.
The teacher move. Normalize it out loud, repeatedly, in the first two weeks. "If that did not land, go to the station, come back, we carry on." If the station reads as remedial it will be empty, and the students who needed it most will be the ones who avoid it.
Failure mode. The station becomes an escape from the lesson. Three-minute limit, timer on the table, and the student returns to the same task rather than a different one.
What it is. Homework is done with hints available but answers switched off. The student's session produces a per-skill record; the teacher reads the class summary before the next lesson.
Setup cost. Zero, if the assignment is one you already set. The change is what the student does when they get stuck and what you receive afterwards.
The teacher move. Read the summary before you plan, not after you teach. The routine's whole value is that tomorrow's lesson starts from real misconceptions rather than from what you expected. The mechanics of that loop are on how it works, and the policy framing is in AI homework help that does not do the homework.
Failure mode. Treating the record as a grade book. It is a diagnostic. The moment students believe their at-home stumbling is being graded, they stop stumbling honestly and the signal dies.
What it is. Ten minutes on one anonymized wrong answer from the class, projected. Not "what is the right answer", but "what was this person thinking, and where exactly does it stop working?"
Setup cost. One wrong answer, which you already have. If the tool names the misconception you have the teaching point ready-made.
The teacher move. Anonymize properly and say that you have. Then insist on precision: not "they did it wrong" but "they applied the second percentage to the original number". The whole value is in the precision, and students find the exercise genuinely interesting when it is not about catching anyone out.
Failure mode. A student recognizing their own work in front of the class. Retype it, change the numbers, or use a wrong answer from a different class.
What it is. The three students who always finish first work a harder version with hints available, while the teacher works with the group that needs it. No printing, no separate worksheet.
Setup cost. One line: "harder version of the same idea". The point is that it removes prep rather than adding it.
The teacher move. Make the extension genuinely harder rather than longer. Fast finishers punished with more of the same stop finishing fast, which is a real and common own goal.
Failure mode. Extension becomes a permanent second track. Rotate who does it, and make it available to anyone who wants it.
What it is. Five minutes, one problem, no devices, on paper, once a week. It is not assessed and it does not go in the grade book.
Setup cost. One problem and a stack of paper.
The teacher move. Explain it as calibration, not as a trap. "This tells me whether the help you are getting is working." Students accept that framing. They do not accept a surprise test.
Failure mode. Letting it become an assessment. The moment it is graded it becomes a test, students prepare for it, and it stops telling you anything about ordinary weeks. Related: what students are actually using, most of which is invisible without a checkpoint like this one.
Three constraints determine whether any of the above survives past two weeks.
Three widely recommended activities that consistently disappoint in a real schedule.
Abstract routines are easy to agree with and hard to run, so here is Routine 4, the whole class evaluating one generated answer, written out as a 50-minute grade 7 lesson on percent change. To be clear about what this is: a lesson plan and the reasoning behind each block of time, not a case study. Coeducate is new, and no school has run this exact lesson with it. The numbers below are the ones this topic reliably produces, not measurements from a classroom we watched.
The objective is successive percent change: a $40 shirt goes up 20%, then goes down 20%, and students have to say what it costs. Assume single percentages have already been taught. The predictable class-wide error is that students treat the two 20s as cancelling, which gets them back to $40 instead of $38.40.
No devices out. Students write their answer and one sentence of reasoning underneath it. The sentence is the part that matters: an answer alone tells you nothing about which of four different faults produced it. Expect a majority of the room to write $40, and expect those sentences to split into at least two groups, along the lines of "the 20s cancel out" and "up 20 then down 20 is the same number". Those two look identical in a grade book and are not the same misconception. That split is the whole reason for asking for the sentence.
Project a single AI-generated worked solution that contains the exact error the class has just made. Not a corrected answer: a wrong one, deliberately. The task is: find the first line that is wrong, and say why it is wrong rather than what the right number is.
This is the whole trick of Routine 4. Thirty students prompting freely produces thirty unrelated outputs and no shared reference. One shared output produces one argument the whole room can have. The moment to steer towards is a student asking "20% of what, though", because that question is the entire lesson objective, and it lands differently when it comes from the room rather than from the front of it.
Devices from the shared cart, students working in pairs on four more successive-change problems. The tutor runs in hints-only mode, so a stuck pair gets a question back rather than a solution. Pairs without a device do the same four problems on paper and swap with a device pair halfway through.
What the teacher does during those sixteen minutes is the point. Not circulating and guessing, but reading the aggregated skill view as it fills in. On this topic the shape it tends to take is predictable: "percent of a quantity" secure across most of the class, "choosing the base of a percentage" not yet for a large group, and "checking a result for reasonableness" developing for almost everyone. Whatever the split turns out to be in your room, it is three facts you cannot get from a set of crosses, and they point at three different lessons.
The reteach comes out of the evidence rather than out of a hunch: put $40 and $48 side by side, ask for 20% of each, and make students say out loud which number the percentage is "of". Five minutes, no preparation, aimed at the one skill the class view flagged. The remaining ten minutes are two more problems with the base written under every step.
Two things, and both are worth planning for rather than discovering. Logins will eat several minutes at the start of the device block if the class is on individual accounts rather than SSO. That is exactly the ninety-seconds-per-student tax described above, and it is what stops a routine like this reaching a second lesson. And some pairs will work out within a minute that hints-only mode will not hand over an answer, get bored, and go back to paper. That is the correct outcome and it should be expected: a tutor that refuses to do the work is less popular than one that does it, and that is the trade the routine is making on purpose.
Total teacher preparation for the whole lesson: generating the one wrong worked answer for the board, which is a few minutes the evening before. Everything else is produced by the lesson itself.
Two questions determine whether any of these routines outlive the first month, and both have their own page rather than a paragraph here.
The first is subject-specific. Math is where the misconception vocabulary has to be precise, because "struggled with percentages" is not a diagnosis and cannot be reteaching. What a per-skill record looks like for math specifically, covering base confusion, additive thinking, notation slips and plain arithmetic error, all of which produce the same cross in a grade book, is set out on the AI math tutor page, along with the fault lines the diagnosis is built to separate across fractions, ratio, algebra and measure.
The second is the one that ends pilots. Somebody will ask where minors' work is going, whether it trains a model, and what happens to it at the end of the year. If you cannot answer that in one sentence, the routines never get past the first staff meeting no matter how well they work in the room. The student data privacy and COPPA compliance page states the model plainly, including which formal attestations exist and which do not, which is the version worth taking into a review rather than a marketing claim.
The test for any of these is not whether it looked impressive in a lesson observation. It is whether it still happens in week nine when you are covering someone else's class and the wifi is bad. Routines that need a good day are not routines. If you want to see what the teacher-facing output of Routines 1, 3 and 4 actually looks like, the Mastery Check demo produces one from a piece of student work in a few seconds.
Where this comes from
Socratic hints instead of answers, evidence the teacher can act on, and student work that is never used to train models. Run the Mastery Check on the homepage and read the output before you decide whether any of this is worth your time.
Students already have AI homework help. The question is no longer whether to allow it, but what shape of help you are willing to defend at a parent meeting.
The phrase AI teacher does a lot of damage in a staff meeting. Here is the split that makes the conversation productive again.
Flashcard generators, summarizers, solvers, voice tutors. A survey of what is actually in use, and what each one hides from you.
Run the Mastery Check on the sample work, or paste your own practice text. Nothing you type is stored.
Student work is never used to train models. No card required to try the demo.
Mastery evidence after every session.
Get started
One field. We email you a 6-digit code to confirm the address, then the next steps for getting your classroom, school or tutoring program set up.
No card required. Student work is never used to train models. See our privacy policy.