Content and Authority for AI Answers

AI Has Exposed Higher Education’s Assessment Problem

Westcliff University’s Anthony Lee offers insights on how AI has exposed higher education’s assessment problem. This article originally appeared in Insight Jam, an enterprise IT community that enables human conversation on AI.

Higher education has spent a great deal of time worrying about whether students are using artificial intelligence to complete assignments. That concern is understandable, but I think we are focusing too much on the wrong problem.

If a university can no longer look at a paper, discussion post, or presentation and feel confident that the work reflects what the student actually understands, then the problem is larger than academic integrity. It becomes a problem of assessment.

For a long time, we were able to make a reasonable assumption that the work a student submitted bore a close relationship to the work the student had done intellectually. Generative AI has weakened that assumption very quickly. A polished piece of academic work can now be produced with far less evidence of mastery behind it. That changes what a university can reasonably infer from the finished product.

I do not regard this as an argument against AI. Students will graduate into organizations where AI will be part of daily work, and universities have a responsibility to prepare them for that reality. The more difficult question is how we continue to know what the student knows when technology can produce increasingly convincing evidence on the student’s behalf.

When a university awards a degree, it is making a claim about the graduate. We are saying that this individual has developed a level of knowledge and judgment that can be relied upon beyond the classroom. Employers take that claim seriously. So do families and students. If we become less certain about the connection between submitted work and actual understanding, then universities have to find better ways to establish that connection.

The workplace already does this constantly. People are asked to explain their reasoning in meetings. A recommendation gets challenged. A supervisor asks a follow-up question that was never anticipated. New information appears, and the person has to adjust. These moments reveal very quickly whether someone understands the subject or has simply produced a competent-looking answer.

Oral examination has been part of universities for centuries. We continue to use it in dissertation defenses and in disciplines where direct questioning reveals far more than a written response alone. There is a reason for that. Conversation makes understanding harder to imitate. A good question followed by an unexpected follow-up can reveal whether a student has command of an idea, where the limits of that understanding are, and whether the student can think when the path has not been prepared in advance.

The obstacle has always been practical. Individual oral assessment takes faculty time, and faculty time is finite. A professor can have a serious conversation with a student. Doing that consistently across several hundred students is another matter.

Ironically, AI may also help solve the problem it helped create. The technology can support individualized oral assessment at a scale that would have been unrealistic only a few years ago. A student can respond to an open-ended question orally, and the next question can be shaped by what the student actually said. The conversation can move deeper when an answer is vague or move in another direction when the student demonstrates understanding.

At Westcliff University, our work with AI-enabled oral assessment has forced us to think much more carefully about what actually constitutes evidence that a student has learned. One lesson has become clear to me: a finished answer tells us less than it once did. The reasoning behind the answer is becoming much more important.

But we should be careful about assuming that technology automatically makes assessment better. An automated oral assessment can be badly designed just as easily as a written exam can be badly designed. Institutions have to know what they are trying to measure before they introduce technology into the process. Faculty judgment remains essential. Questions have to be aligned with what was actually taught, and students need to understand how their responses will be evaluated. Accessibility, privacy, and fairness also have to be designed into the process from the beginning.

I would be equally cautious about turning oral assessment into another form of policing. The most useful purpose is educational. We should want students to explain what they know because explanation itself is part of learning. Having to respond to a thoughtful follow-up question requires students to take ownership of what they know in a way that is increasingly difficult to establish through take-home work alone.

This is especially important as universities prepare students for an AI-enabled workforce. Employers will certainly value people who know how to use generative tools well. They will also continue to value people who can walk into a room, understand a difficult problem, and explain what they think should happen next. Technology may assist that person throughout the day, but at some point the judgment still belongs to the individual.

Universities cannot keep making confident claims about what graduates know if we are becoming less confident in the evidence behind those claims. Universities therefore need to become more demanding about what we mean when we say a student has learned something. Written work will remain important. So will projects, research, and the intelligent use of AI. But there should also be moments during a student’s education when the institution can hear the student think.

Generative AI has made this question impossible to avoid. We can produce better-looking academic work than ever before, faster than ever before. That tells us surprisingly little about whether understanding has kept pace.

A degree should still mean that it has.

Share This

Related Posts