AI That Knows What It Doesn't Know: Why Grounded Tutoring Is the Feature That Matters
By EduGears AI Team

Picture a student at eleven at night, stuck on the third question of a problem set that is due in the morning. They are not going to email their instructor, and they are not going to wait for office hours. They are going to type the question into an AI, because that is what students do now, in every subject and at every level. The debate about whether students should use AI to learn has quietly been settled by the students. The question that is still open — and the one that actually decides whether AI helps or harms a course — is which AI they are talking to, and what it knows.
The AI most students reach for first is an internet-scale chatbot. It has read an enormous amount, and it will answer anything you ask it, fluently and immediately. For general curiosity that is a genuine strength. For a course, it is precisely the wrong design. A course is a deliberate narrowing of the world: this textbook, this notation, this definition, this method, in this order, by this week. A chatbot drawing on everything will cheerfully answer with a different convention, a technique the class has not been taught yet, a definition from a different jurisdiction or syllabus, or a fact from an edition the department retired years ago.
What makes this dangerous is not that the chatbot is often wrong. It is that a wrong-for-this-course answer looks exactly like a right one. It is written in the same confident register, with the same tidy structure. The student cannot tell the difference — they asked because they do not know — and the instructor never sees the exchange at all. Confidence without scope is the failure mode, and it is invisible to everyone involved until the exam.
There is a better design, and it starts by giving up something. The tutor answers from the course's own materials — the documents the instructor uploaded, the course's lessons and resources — and politely declines the rest. It sounds like a limitation. It is the feature. A tutor that stays inside the course speaks the course's language, uses the course's examples, and cannot drift into a method the class has not met, because it is not permitted to go looking for one.
What a good decline looks like matters. Ours reads, in full: "That's outside this course's materials — ask your instructor if you think it should be covered." Notice what that sentence does. It is honest about the boundary instead of bluffing past it. It is short, so it costs the student a few seconds rather than a wrong turn. And it points the student to a person, which is exactly where a question the course does not answer should go.
A decline is also information — arguably the most useful information a tutor produces. Every one is a record that a student wanted something the course does not currently cover. Collect them, group them, rank them by how often students asked, and you have a list of your course's content gaps, in order of demand, ready to fill. An AI that answers everything destroys that signal: it papers over each gap with an answer from somewhere else, and the instructor never learns the gap was there. An AI that knows what it doesn't know hands the instructor a map of it.
The signal works in advance, too. If the tutor answers only from course materials, you can look at those materials before term starts and see where there is simply nothing — a module with no documents and no lesson content behind it is a module where the tutor will decline. A coverage view cannot tell you whether a document covers every topic a student might raise, and it should say so plainly. What it can show is where there is nothing at all, so you can add it before a student asks.
And declines say something about students as well as about courses. A learner whose questions are mostly being declined may be lost, working from the wrong course, or trying to get ahead of material that is not there yet. Any of those is worth a teacher's attention, early. That pattern only exists if the tutor is willing to say no in the first place.
How strict the tutor should be is a teaching decision, so it belongs to the teacher. Some courses benefit from a tutor that prefers the course materials but may also draw on general knowledge; some need a tutor that stays strictly inside them; some need explicit topic boundaries — stay within the chemistry of this unit, treat everything else as out of scope. Our view is that strict should be the default, because the costs are lopsided. A decline costs a student a minute and a question to their instructor. A confident wrong answer can cost them the exam, and cost the teacher their students' trust in every answer that came before it.
Transparency is the second half of trust. A student should always know when they are talking to an AI, not only on the first day. That means a badge that never goes away — ours says "AI tutor — answers are AI-generated" on every tutor surface — and a first conversation that opens with a plain notice that answers can be wrong, with a pointer back to the course materials and to the instructor. None of this is legal boilerplate. It is AI literacy delivered at the one moment it is guaranteed to be read.
Then there is pedagogy. Every tutor has a teaching style, whether anyone designed it or not. A Socratic tutor answers a question with a question and makes the student do the reasoning. A scaffolded tutor breaks the problem into steps and hands over one rung at a time. A direct tutor explains. All three are legitimate, and which one fits depends on the subject, the level, and what the course is trying to build. The question is who chooses. With a general chatbot the answer is the student's mood, and at eleven at night the mood is usually "just tell me". We think the instructor should choose, per course — and when they do, the student sees "Style set by your instructor: Socratic" in place of their own picker. Where it makes sense to let learners pick for themselves, that is a choice the instructor can make too.
Graded work draws the sharpest line of all. A tutor that will complete an assessment for a student is not a tutor; it is a shortcut with a friendly voice. On graded work, the tutor should scaffold rather than complete: help the student work toward an answer, not produce it for them. The distinction between help and cheating is not mysterious, and a well-built tutor should know exactly where it sits.
The same principle — know what you don't know, and say so — applies with even more force to marks. AI is genuinely useful for marking short written answers, because it can recognise a correct paraphrase that an exact-match check would reject. But a mark is a consequential thing, and an AI mark has to be one a teacher can see, question and correct. Seen means every short answer records how its mark was decided — by exact match, by the AI, or without the AI because it was unavailable — and the submissions view says so on the attempt. Questioned means the instructor decides how strictly the AI marks, course by course: lenient marking that accepts paraphrases and reasonable variants, or strict marking for terminology-heavy subjects, with the honest warning that strict may mark more answers wrong than a teacher would. Corrected means any mark can be overridden, with a note if the teacher wants one; the attempt is rescored and sent to the gradebook again, and the student sees the corrected mark with the teacher's name on it. The AI drafts; the teacher owns the mark.
The most important marking rule is the least glamorous. When the AI cannot check an answer — because the service was unavailable at the moment the student submitted — the easy fallback is to mark the answer incorrect. That quietly converts a system failure into a student's failure, and nobody finds out unless the student complains. The right behaviour is to wait. The student sees "Awaiting teacher review" instead of a wrong mark, the gradebook shows the attempt as waiting, and the teacher sees each waiting answer highlighted, ready to mark with the same override they already use. It is the marking equivalent of a tutor that declines: an AI that knows it could not check is worth far more than one that guesses and keeps quiet.
None of this is happening in a vacuum. U.S. federal guidance on AI in education asks for tools that are educator-centered and transparent, and a growing number of states are turning those principles into rules. We wrote a plain-language guide to what the guidelines require at www.edugears.ai/en/blog/us-ai-education-guidelines-what-they-mean-for-schools. Grounded tutoring is what those principles look like once they are built into software. Educator-centered means the teacher sets the strictness, the topic boundaries and the guidance style, and has the last word on every mark. Transparent means the student always knows it is AI, every question put to the tutor — answered or declined — is visible to the instructor in a report, and every mark carries a record of how it was decided.
If you are evaluating an AI tutor for your school, college or training programme, a handful of questions will separate the designs quickly. What exactly does it answer from? What happens when the answer is not in the course — does it decline, or does it improvise? Can I see every question my students asked, including the ones it would not answer? Who chooses how it teaches: me, or the student at eleven at night? What does a student see when the AI could not mark their answer? And when I override a mark, does the correction reach the gradebook? Vendors with a real design answer these in a sentence each. Where the answers come back as adjectives, keep asking.
This is how EduGears AI now works, across the EduGears AI LMS and our Moodle and Canvas toolkit alike: grounded by default. The tutor answers from each course's own materials and politely declines what falls outside them; instructors choose, per course, how strictly it grounds, which topics it stays within, and how it guides; every tutor surface says plainly that it is AI; and AI marking is something teachers can see, set and correct, with answers the AI could not check waiting for a teacher instead of being marked wrong. Existing courses have moved to the strict default, and an instructor who prefers the broader behaviour can set any course to Relaxed. If you would like to see it on a real course, bring your hardest off-topic question — we would be glad to show you the decline.


