New Tutoring Benchmarks Continue to Reveal Gaps in AI Models for Teaching and Learning
There are two new interesting papers about education benchmarks for LLMs:
EduClaw-Bench evaluates AI-agent configurations over a 30-day learning horizon using simulated learners based on real knowledge-tracing data. Their study revealed that almost none of the model/agent combinations they studied (including GPT5.5 and Alibaba’s Qwen models) maintained strong tutoring performance over time; sustained tutoring demands ongoing capabilities like remembering learner progress, spacing practice, and maintaining strategy, suggesting the need for learner modeling, orchestration, and pedagogical memory. (arXiv)
The ELBench research benchmark evaluated nine models (seven general-purpose frontier systems and two education-specialized models) across general capability, safety/trustworthiness, basic education, and higher-order educational development. Specialized Models Underperformed: The two education-specialized models studied failed to lead in either of the education-specific modules. (arXiv)
Why It Matters
Edtech still lacks solid, universally accepted benchmarks to define what actually makes an AI product or tutor “good” at teaching and learning… but it’s not for lack of trying.
Learning Commons has been publicly building a suite of ‘evaluators’ of Edtech tool outputs
Google has developed a thoughtful internal rubric to evaluate the pedagogical value of its LLM outputs
Digital Promise recently announced the recipients of the first grants from its K-12 Infrastructure program which include funding for benchmarks
…and last, but not at all least, research teams around the world have been suggesting lots of clever ways to benchmark the outputs of LLM models and their agents/harnesses/product delivery systems.
Our Take
As always, the gap between research and practice is vast and most Edtech companies are not even attempting to benchmark their own products against any of these research-lab-released benchmarks.
That said, maybe they should… at least one lead VC in the field is paying close attention. Here’s a quote from Jennifer Carolan of Reach Capital on Alison Dulin Salisbury’s excellent Humanist Substack:
ALLISON: What have you changed your mind about in the past year?
JENNIFER: I changed my mind about how useful AI tutors are becoming. The underlying technology is improving more quickly than I had ever imagined which is rapidly improving the quality of the AI tutors. Combined with the changes happening in society writ large, these AI tutors may become one of the primary drivers of change of our current formal education system.
I teach a class at Stanford with Steve Blank called Lean Launchpad, and we’ve been so blown away by how quickly students built working prototypes this year—like, Week Two. We’ve had a couple of education teams in this cohort, and it’s been fascinating to see how fast they created compelling AI tutors. In just a few weeks, one team made the leaderboard for Scale AI’s “TutorBench” evaluation—a benchmark for assessing AI tutor quality—so we’ve watched the speed at which these tutors improve in almost real time. Tutors are meeting a pivotal moment, as higher ed is getting pushed hard on ROI and K-12 is under tremendous pressure.
Thanks for reading! Support our work by becoming a paid subscriber, or email info@edtechinsiders.org to learn more about partnerships and sponsorships.


