The Feedback That Ships Without the Grade
Google Classroom’s Gemini integration shipped on time — and that was the constraint that mattered. This piece examines what the product team chose to leave out: multilingual feedback, audit logs, automated grading — each absent feature a resourcing decision, not a capability gap. The bill arrives when schools discover the tool live, teachers unprepared, and no safeguarding audit layer in place.
Google expanded Gemini AI study help in Google Classroom to younger schoolchildren in August 2026, and the Classroom tab defaulted to on for students whose schools already enabled the AI chat or notebook service. The announcement emphasised personalised learning support grounded in actual class materials. Three months later, Plymouth High School principal Michael Martin told Education Week he learned about the change through a news story and disabled Gemini because teachers had not been trained collectively to use it. That gap between release and readiness is a planning problem, but it sits downstream of a different question: what did the product team have to leave out so the feature could ship in August?
I spent eight years building content moderation and safety systems inside platforms before Real AI, and another six advising teams that sell software into education. The patterns are consistent. A vendor announces a capability, schools adopt it or find it turned on by default, and six months later an engineer somewhere is holding the ticket for the feature that should have been in version one. The ticket is always something procedural: audit logs the institution can export, a review queue a department head can access without elevation, or the ability to generate feedback in the language half the district actually speaks. It was cut because launch dates do not move and roadmaps do not defend themselves in budget meetings.
Google launched AI-suggested feedback for written assignments in Google Classroom on 19 February 2026, enabling educators to generate draft comments using Gemini tailored to a student's work, grade level, and selected focus areas. At launch, access was limited to English-speaking educators aged 18 and older. The August expansion added the Classroom tab for students, but the core feedback tool remains monolingual and sits behind a paywall for the use case schools actually need. Educators can now create an audio lesson or convert a file into a Classroom-ready rubric, or soon get suggested feedback for writing assignments, but suggested feedback on written assignments requires Education Plus or the Teaching and Learning add-on. A free tier exists; the marquee feature does not live there.
The feature matrix nobody sees
Google Classroom has some built-in AI grading features through Gemini, but teachers can only use them to leave feedback for students, and it is unable to grade student work and produce rubric-aligned scores that can go into gradebooks. That is a design choice, not a technical limit. Every grading tool comparison published in the last four months lists automated rubric scoring as table stakes. AI auto-grading for objective assessments grades MCQ and short-answer questions at 95%+ agreement with human markers on objective items and 80-88% on well-scoped short answers, and the models Google has access to can do that work. The feature was scoped out.
The decision makes sense if you are the product manager writing the PRD. Automated grading is a different liability surface than suggested feedback. A draft comment that a teacher reviews and edits before posting is an assistive tool; a score that populates a gradebook without human confirmation is an automated decision that implicates FERPA, accuracy disputes, and the operational question of what happens when a parent contests a grade the system assigned. Deferring that feature pushes the hardest compliance and trust conversations into a future release, which means this release can ship in February instead of June.
What that leaves in the market is a feedback drafter that works in one language, requires a subscription most districts have not bought, and stops short of the workflow step teachers actually need to automate. AI-generated comments may still require careful refinement to avoid generic or overly broad feedback, per Google's own documentation. The verb there is "may". In practice it is "will", because the model has no persistent memory of how an individual teacher typically frames developmental feedback, and first-draft suggestions revert to the statistical average of praise-concern-suggestion unless the teacher writes a custom prompt every time.
What the engineer could see in March
If you were the ML engineer who shipped the February feedback feature, you knew in March what the support load would look like in September. Teachers in multilingual districts would submit feedback requests in Spanish, Mandarin, or Arabic and receive either an error or an English-language response. Schools that had enabled Gemini for general use would discover the feedback capability was not included and would open billing enquiries. Administrators would ask why student Gemini conversations were not logged in a dashboard they could audit, and support would route the ticket to a product gap that was marked "future consideration" in the backlog.
Administrators have minimal ability to monitor conversations or receive alerts about concerning content, and teachers cannot review student conversations unless they are Workspace administrators. That is not an oversight. Building an audit layer that preserves student privacy, satisfies institutional oversight requirements, and does not generate ten thousand low-value alerts per school per week is a product in itself. It was not in scope for the August student release, so it did not ship. The consequence is that the tool went live in a configuration where a student can use Gemini to discuss personal or mental health topics, and the district's safeguarding team will not know unless the student or a peer reports it outside the platform.
The technical capacity to build these features exists. Turnitin released an update to its AI writing detection model for Spanish submissions in August 2026 that improves detection of AI-generated content, and the AI writing indicator shows an overall percentage of the text that may have been AI-generated. Multilingual model deployment is a solved problem in 2026; Turnitin added Arabic detection the same month. The Google Translate API has been in production for fifteen years. The fact that Classroom feedback shipped English-only in February and stayed that way through August is a resourcing decision, not a capability gap. Someone prioritised other features, or the internationalisation work was estimated at eight engineering weeks and the release window was six.
The adoption curve nobody modelled
MIT education-technology scholar Justin Reich argued that the default expansion compressed schools' usual time to vet a tool. The rollout decision makes commercial sense and terrible operational sense. Turning a feature on by default increases activation rates, which moves the product metric the leadership review cares about. It also means thousands of schools went into the autumn term with a generative AI tool enabled for students in an environment where most teachers had not been trained to use it, did not know it was active, and had no shared framework for when student use of the assistant constituted acceptable support versus policy violation.
The alternative was an opt-in release, which would have reduced Week One adoption by 60% and required a months-long enablement campaign to hit the same MAU target by December. Product teams do not get rewarded for slow, careful rollouts. They get rewarded for shipping, for activation numbers, and for feature parity with competitors who already have AI in the classroom. McGraw Hill completed its acquisition of TeachFX, an AI-native coaching platform that analyses recorded lessons and gives feedback on practices, in September 2024, and the deal followed McGraw Hill's earlier purchase of Teachally. The sector is moving fast enough that a six-month delay is a market position you do not get back.
If I were still inside a product org and this landed on my desk, I would have flagged three risks in the decision document: districts discovering the feature live without prior communication, support load from schools asking how to disable something they did not enable, and reputational exposure if a student used the tool in a way that triggered a safeguarding incident the school could not have detected. All three happened. None of them were enough to delay the launch, because the calculus is that most schools will absorb the adjustment cost, 95% of student interactions will be benign, and the PR risk of one bad outcome is lower than the strategic risk of ceding the education AI market to OpenAI and Anthropic while Google spends another quarter in review.
The engineer who built the feedback model knows it works. The engineer who scoped the feature matrix knows what was cut. The person who made the default-on call knows the trade. What is left on the desk is a tool that does one thing well, in one language, for the educators who already paid for the tier that includes it, in a rollout that assumed schools could adapt faster than their professional development cycles actually run. That is not a bad product. It is a product that shipped on time, and time was the constraint that mattered.
Tarry Singh is the founder and CEO of Real AI (realai.eu), an enterprise AI advisory and deployment firm working with global enterprises on production agent systems, model risk, and AI sovereignty strategy. He also leads Earthscan (earthscan.io) for Energy AI, and is a founding contributor to the EU-funded HCAIM and PANORAIMA programmes for responsible AI education across European universities. He writes at tarrysingh.com.