AI Killed the Essay, and Grade Inflation Killed the GPA

‍I began thinking about this while having a conversation with a colleague during a break in a lecture recently delivered to faculty and staff at my institution. Although the lecture was not about our new teacher education program, some of the issues it raised found their way into a conversation we had already been having about how we should assess students in a new program preparing its first cohort of teacher candidates. In that conversation, I found myself returning to two issues, artificial intelligence and grade inflation in teacher education. This was because the lecture seemed to raise a much larger question about what we are actually trying to assess when we prepare teachers.

This question feels particularly important to me because teacher education is not just another university program. Whatever form it takes, whether education is studied alongside another undergraduate discipline or a student completes a first degree and then enters a consecutive BEd, the purpose is ultimately professional formation. We are preparing people to enter classrooms and take responsibility for the education of other human beings. That means the assessment practices we use in teacher education deserve more attention than asking whether they are familiar or convenient. They should tell us something meaningful about whether a person is becoming capable of doing the work of teaching.

The structure of teacher education varies across Canada. Some students study education alongside another discipline, while others complete an undergraduate degree first and then enter a post-degree BEd. Programs also vary considerably in length and organization. Ontario has moved toward shorter consecutive teacher education programs, and Manitoba has its own arrangements for preparing teacher candidates. My own institution is now beginning this work with its first cohort, which makes some of these questions feel less abstract. When you are building a program, you have an opportunity to ask not only how things have traditionally been done, but also whether the traditional way still makes sense.

‍Artificial intelligence (AI) is probably the most obvious place to begin. I have struggled with AI myself. I am still trying to understand what it means for education, writing, learning, and the formation of students. But I have reached a point where I no longer think the most useful response is to pretend that AI will disappear. It is here, and I suspect it will remain here. We can regulate it, restrict it, teach students about its dangers, establish circumstances in which it should not be used, and develop institutional policies around it. All of those things matter. But even if an institution decided to prohibit AI completely, the underlying question would remain. What kinds of learning are we trying to cultivate, and what kinds of assessment can actually provide evidence of that learning?

I do not think AI has literally killed the essay. Nor do I think grade inflation has made the GPA entirely meaningless. But both have exposed something important about the limits of the ways we have traditionally assessed learning. AI has made that question harder to ignore because it can now do things that universities have traditionally asked students to do on their own. It can produce a lesson plan, explain a theory, generate a reflection, organize an argument, revise a piece of writing, suggest classroom activities, and produce polished prose almost immediately. Anyone who has worked with these systems seriously knows that their capabilities are not trivial. They can be useful precisely because they can help a person think through a problem, generate possibilities, notice connections, and revise an idea. That is part of what makes them educationally interesting. It is also what makes assessment difficult.

‍For years, we have often treated a take-home assignment, let's say essay, as evidence that a student understands something. There is nothing inherently wrong with this. Writing remains an important intellectual practice. I still want my students to write. I want them to struggle with ideas, organize arguments, interpret texts, make judgments, and find language for things they are trying to understand. But when a student submits a polished essay, I increasingly have to ask myself what exactly that essay establishes. Does it show what the student knows? Does it show what the student can think through independently? Does it show how they respond when an argument is challenged? Does it show whether they can recognize a weakness in their own position? Or does it primarily show that they can produce, with or without technological assistance, a polished piece of writing?

These questions reveal a weakness that was already present in our assessment practices. We have sometimes asked an assignment to provide more evidence than it was capable of providing. AI has made that problem much harder to ignore. The same kind of difficulty appears when we think about grades. Grade inflation is not just a matter of students receiving too many high grades. There are complicated reasons why grades change over time, and an increase in grades does not automatically mean that students are learning less. But there is a point at which a grade begins to lose some of its ability to distinguish among levels of achievement. If an A becomes increasingly common, the difference between an A and an A-minus becomes less meaningful. Students become anxious about very small differences in percentages, professors feel pressure to respond to those anxieties, and institutions continue to use GPA as though it were a precise representation of intellectual ability.

‍I have watched students become concerned about whether an assignment will move them from an A-minus to an A. It is not difficult to understand why. Grades matter for scholarships, graduate school, employment, awards, and sometimes for how students understand themselves in a corporatized milieu. But the more heavily we attach consequences to small numerical differences, the more tempting it becomes to treat the number itself as the achievement. That is where AI and grade inflation begin to intersect. A student who is already capable of producing good work can use AI to make that work more polished. If the difference between a B and an A carries significant consequences, there is an obvious incentive to use whatever tools are available to gain that advantage. We can tell students not to do it, and sometimes we should. But it seems to me that we should also examine the assessment environment that makes the difference between an 84 and an 89 feel so consequential in the first place.

This brings me back to teacher education. I am not persuaded that teacher education should reproduce the traditional grading practices of the university because that is what universities and faculties of education have always done. The purpose of teacher education is to prepare teachers, and teaching is a profession in which competence cannot be adequately represented by a GPA. Think about what we actually ask teachers to do. They need knowledge, certainly. They need to understand curriculum, pedagogy, assessment, child development, subject matter, educational policy, and the philosophical, historical, and social contexts in which schooling takes place. But knowledge alone does not make someone a good teacher. Teachers also need judgment. They need to be able to read a classroom, notice when students are confused, respond to unexpected situations, adapt an explanation, work with colleagues, communicate with families, recognize their own mistakes, and make decisions when there is no obvious answer available.

These, I argue, are what make one intellectually virtuous, connecting what a person knows with what a person must be able to do. This is why I find the comparison with professional education useful, particularly medicine. We do not ordinarily tell the public that a doctor is competent because the doctor accumulated an impressive GPA. What matters much more is whether the person has demonstrated the knowledge, judgment, practical ability, and professional competence necessary to care for patients. Medical education has therefore developed forms of assessment that go well beyond conventional course grades. Harvard Medical School, for example, uses Satisfactory/Unsatisfactory grading in significant parts of its curriculum while still maintaining examinations and other forms of assessment. Yale Law School similarly uses a grading structure that includes Honors, Pass, Low Pass, and Fail rather than calculating a conventional GPA. I do not think this means that teacher education should copy medical school or law school. The professions are different, and their assessment needs are different. What interests me is the principle underneath these examples.

Professional competence does not necessarily require a numerical ranking of every student. Perhaps teacher education should take that possibility more seriously. I have, therefore, begun to wonder whether a competency-based Pass/Fail model might make more sense for teacher education than the traditional grading system. I do not mean that everyone would simply receive a Pass for completing assignments or attending class. Nor do I mean that standards would become lower. In fact, I think the opposite would have to happen. A Pass would need to represent demonstrated competence. A teacher candidate would still have to read difficult material, write, participate in discussions, engage with educational theory, complete assignments, collaborate with colleagues, and demonstrate knowledge. They would still have to complete a demanding practicum(s). They could still fail. What would change is the role of the final number. Instead of treating an 84 and an 89 as evidence of meaningfully different levels of professional readiness, the program could establish a clear threshold of competence and use multiple forms of evidence to determine whether the candidate had reached it.

That evidence could include essays, but the essay would become one piece of evidence rather than the entire case. It could include lesson planning, classroom presentations, teaching demonstrations, reflective writing, collaborative projects, curriculum analysis, practicum evaluations, professional portfolios, action research, and opportunities for candidates to explain and defend their decisions. This is important in an AI-mediated environment because the assessment can move closer to the person. Suppose a teacher candidate submits a lesson plan. Instead of assuming that the document itself tells us everything we need to know, we might ask the candidate to teach part of it. We might introduce an unexpected classroom scenario and ask what they would do. We might ask why they selected one instructional strategy rather than another. We might ask them to identify weaknesses in their own plan. We might even give them an AI-generated lesson plan and ask them to evaluate it. The candidate does not necessarily have to demonstrate that they can produce something without technological assistance. They have to demonstrate that they can exercise judgment about what the technology produces.

That seems increasingly important for future teachers. Teachers will encounter AI in schools whether we want them to or not. They will have students using it. They will have colleagues debating its use. They will have administrators developing policies around it. They will have to decide when it is appropriate, when it is inappropriate, and how it changes teaching and learning. Preparing them to avoid AI seems inadequate. They need to understand it well enough to use it responsibly and critically. This is where I think the idea of intellectual virtues becomes important. I have written elsewhere about the importance of habits such as intellectual humility, open-mindedness, and attentiveness. AI makes these dispositions more important, not less. A teacher who knows how to use an AI system but cannot recognize when its answer is wrong is not necessarily well prepared. A teacher who can generate hundreds of lesson ideas but cannot judge which ones are appropriate for a particular group of students has not solved the pedagogical problem. The technology may expand our possibilities, but it does not remove the need for judgment.

The same principle applies to grading. If teacher education adopted Pass/Fail, I would not want it to become a system in which a candidate does the minimum amount of work necessary to cross a 70 percent threshold and then stops. That would be a failure of the proposal. The point would be to change what the program recognizes and how it recognizes it. A candidate who demonstrates exceptional ability should have that excellence visible somewhere. It simply does not have to appear as a 98 percent on a transcript. A strong practicum evaluation could demonstrate it. A sophisticated capstone project could demonstrate it. A portfolio could demonstrate it. A professor’s detailed assessment could demonstrate it. A candidate’s ability to defend their pedagogical decisions could demonstrate it. There are many ways of making excellence visible without turning every difference in performance into a difference in GPA.

‍This also raises an obvious concern about students who later want to pursue graduate education. Suppose someone completes a BA in History and then enters a Pass/Fail BEd. If that person later applies to an MEd, the admissions committee would still need evidence of academic ability. Their first degree would provide some of that evidence through the traditional transcript and GPA. The BEd would provide a different kind of evidence through practicum evaluations, a capstone or research project, references, and a professional portfolio. In some respects, this could make admissions more demanding rather than less demanding. Instead of looking at an A on a transcript and assuming that the student performed well, an admissions committee would have to look at the work itself. What did the student actually produce? How did they think about educational problems? What did their practicum supervisors observe? Can they conduct research? Can they write an argument? Can they connect educational theory to practice?

There is something attractive about that. Universities have become very good at converting complicated forms of human achievement into numbers. Numbers are convenient because they allow us to rank, sort, compare, and make decisions quickly. But convenience does not necessarily mean that the number is the best evidence. A GPA can tell us something. It should not be expected to tell us everything. The same applies to narrative assessment. I would not want to replace numbers with unstructured opinions and assume that the problem had been solved. A principal’s evaluation of a student teacher can be affected by expectations, personality, cultural assumptions, or unconscious bias. A Pass/Fail system that relies entirely on narrative evaluations could create its own problems, particularly for students who already experience inequities within educational institutions. This means that a competency-based system would need to be carefully designed. Practicum assessments would need clear criteria. Multiple evaluators could be used where appropriate. Candidates would need opportunities to respond to feedback. Programs would need to monitor whether particular groups of students were being evaluated differently.

The criteria should focus on observable practices rather than vague impressions of whether someone “seems like a good teacher.” In other words, moving away from grades does not mean moving away from rigour. It may actually require us to become more rigorous about what we mean by competence. This is why I am not particularly interested in the argument that Pass/Fail would make university easier. If implemented properly, it might make some aspects of teacher education harder. A student cannot simply produce an excellent essay and hope that the grade carries them through. They would have to demonstrate their competence in several different ways. They would have to teach. They would have to explain. They would have to reflect. They would have to respond. They would have to make judgments. They would have to show that they can move between theory and practice. That seems much closer to the work they are preparing to do.

‍There is also something worth considering about the kind of teacher we want our teacher education programs to produce. If we train teachers in an environment where every assignment is understood primarily as an opportunity to earn points, we may unintentionally teach them to think about assessment in the same way. They may come to see learning as something that is divided into percentages and converted into a final number. But teachers themselves will later be responsible for assessing children. Perhaps they should experience an educational milieu in which assessment is understood differently. Assessment can be about identifying growth, diagnosing difficulties, providing feedback, making professional judgments, and determining whether a learner has demonstrated a particular competency. It does not have to be primarily about ranking one person against another. This does not mean that competition disappears. Nor should it necessarily. Universities serve many purposes, and grades can be useful for some of them.

I am not making a case for abolishing grades across higher education. A physics degree, a history degree, and a teacher education program may have different assessment needs. Even within teacher education, there may be good reasons to retain numerical assessment in some circumstances. My argument is more modest. Teacher education should be one of the places where we seriously reconsider what grades are for. The emergence of AI gives us a reason to do this now. So does the growing concern about the meaning of GPA. But the deeper reason has little to do with technology or grade inflation. It has to do with the nature of teaching itself. A teacher is not hired because they can produce the highest-scoring essay in a cohort. A teacher is trusted with a classroom because they are expected to possess knowledge, judgment, professional responsibility, and the ability to act well in situations involving other human beings.

‍If AI has made it harder to know whether a piece of writing represents a student’s own thinking, perhaps we should not respond just by trying to make the old assignment harder to complete with AI. If grades have become increasingly difficult to interpret, perhaps we should not respond by finding increasingly precise ways to distinguish between an 88 and a 90. Perhaps we should ask whether we have been asking these instruments to do too much. The essay still matters. Writing still matters. Grades may still matter. But neither should be expected to carry the entire burden of demonstrating that someone is ready to become a teacher. I keep returning to the conversation with my colleague because our discussion was ultimately about something much larger than whether students should receive grades. We were trying to think about what kind of teacher education program we want to build. For a new program, that is an opportunity. We do not have to assume that every inherited practice is necessarily the best practice.

Perhaps the most important thing we can give teacher candidates is not another way to compete for the highest grade, but an education that helps them develop the intellectual, moral, and professional capacities that teaching requires. AI has not made thinking unnecessary. If anything, it has made our ability to recognize genuine thinking, and what I have encountered elsewhere as “meditative thinking,” more important. And grade inflation has not made achievement meaningless. It has made us more aware of how difficult it can be to represent achievement with a single number. Teacher education therefore has a choice. We can continue asking whether students earned an A+, an A, an A-minus, or a B-plus and hope that those distinctions tell us enough. Or we can become more ambitious about assessment and ask whether the evidence we gather actually helps us determine whether someone is ready to teach. I think we should choose the latter. Not because grades have no value, and not because AI has destroyed education, but because preparing teachers is too important to leave the question of assessment unquestioned.

Next
Next

Must You Call Me by My Title?