AI Assessment in Higher Ed: What Proves a Student Learned?
Solutions Review’s Tim King explores how AI assessment in higher education is changing as educators rethink what actually proves student learning in the age of GenAI.
AI assessment in higher education is becoming a much larger challenge than determining whether students used ChatGPT.
When generative AI can produce increasingly polished essays, research, analysis, and other academic work, educators face a more fundamental question: What actually proves that a student learned?
For decades, higher education has relied heavily on final artifacts as evidence of learning. Students write papers, complete projects, answer questions, and submit assignments. Educators evaluate those outputs and assign grades based, at least in part, on the assumption that the finished work reflects the intellectual process that produced it.
Generative AI complicates that relationship.
A student can now produce an impressive final artifact without necessarily performing all of the thinking traditionally associated with creating it. That doesn’t make the artifact worthless, nor does it mean every use of AI undermines learning. It does mean the finished product can no longer tell educators everything they need to know.
That challenge was at the center of “Navigating Student and AI Work in Higher Ed,” a panel from the recent Q3, 2026 Mini Jam on Insight Jam, which explored education in an age of abundant intelligence. Higher education leaders, faculty members, and AI experts debated how educators can make learning visible when AI increasingly mediates the work students produce.
Their discussion revealed something important: AI assessment in higher education isn’t simply an academic integrity problem.
It’s an evidence-of-learning problem.
Why AI is Changing Assessment in Higher Education
AI did not create the difficulty of determining whether students actually understand something.
Educators have always worked with imperfect proxies for learning. An essay demonstrates some things. An exam demonstrates others. A presentation, project, discussion, or lab can provide additional evidence.
None provides a perfect window into cognition.
Generative AI makes that limitation considerably harder to ignore.
As the panel discussed, education has traditionally placed significant value on the polished final product. But students can increasingly use AI to improve—or potentially produce—exactly what educators are evaluating.
One panelist described encountering students who did not want an instructor to see their writing before it had been processed through an LLM. The instinct is understandable: students have learned that better-looking outputs tend to receive better evaluations.
But educators may now need precisely what students have been conditioned to hide: the messy work.
Drafts. Mistakes. Decisions. Revisions. Questions. False starts. Interactions with AI. Explanations of why one approach was chosen over another.
The educational value may increasingly exist in the path toward the final product rather than the polish of the product alone.
What Counts as Evidence of Learning Now?
That leads to one of the defining questions for AI assessment strategies:
If the finished assignment is no longer sufficient evidence of learning, what should replace it?
There was no universal answer from the panel—and that’s important.
One approach is greater process documentation. Instead of evaluating only what students ultimately submit, educators can capture the decisions students make along the way, including how and why they use AI.
Another panelist described an approach where students document their AI interactions throughout the semester and eventually review the entire arc of that work. The objective isn’t simply to create an audit trail. Looking back across months of interactions can help students recognize how their own relationship with AI and their thinking changed over time.
Reflection and metacognition could therefore become increasingly important components of assessment.
So could demonstrations, conversations, iterative work, explanations of decisions, and other forms of evidence that make it harder to separate an artifact from the person who supposedly learned by creating it.
But there is a danger in treating any of these approaches as a simple replacement for the traditional assignment.
Can Educators Really Make Student Thinking Visible?
“Make thinking visible” sounds like an obvious response to generative AI.
The panel complicated that assumption.
Students can document what they did, but that does not necessarily mean educators have captured what they thought.
One participant challenged the entire premise that cognition can reliably be surfaced through assessment. Another pointed out the difficulty people can have articulating the intuitive processes behind their decisions, particularly when they have not yet developed sufficient expertise.
That distinction matters.
Requiring students to explain every decision could produce additional evidence of learning. It could also produce additional academic theater: students reconstructing plausible explanations for decisions after the fact because they know an instructor expects them.
The appropriate approach may also vary considerably depending on the student, discipline, assignment, and level of expertise.
That means higher education should be careful not to replace one imperfect proxy for learning with another.
A five-page essay isn’t automatically proof of cognition.
Neither is a five-page reflection explaining how the essay was written.
AI Reflection Can Become Performative, Too
This is where the discussion becomes particularly valuable for institutions redesigning assessment around AI.
Reflection is often proposed as a way to preserve human thinking. Ask students what they learned. Have them explain their process. Require them to document how AI was used. Ask them to critique the output.
But the panel raised a difficult question:
- Who is the reflection really for?
A private reflection designed purely for the learner is different from a reflection submitted to an instructor who controls the student’s grade.
Students understand that power relationship.
Once reflection becomes an assessed activity, students may begin performing reflection in much the same way they learned to perform traditional academic work. They tell the instructor what they believe the instructor wants to hear.
The panel did not resolve this tension. Instead, the disagreement exposed an important principle for AI assessment in higher education: simply adding reflection to an assignment does not automatically make the assessment authentic.
Assessment design has to consider incentives as well as activities.
The Goal Shouldn’t Be Proving That Students Didn’t Use AI
This distinction could also help higher education move beyond an increasingly narrow debate over AI detection.
The most useful question isn’t always whether AI touched an assignment.
It’s whether the student developed and demonstrated the capabilities the course was supposed to teach.
In some circumstances, AI use may interfere with that objective. In others, using AI effectively may itself be part of the learning objective.
What matters is the relationship between the tool, the learner, and the capability being developed.
One panelist framed a particularly significant AI-era challenge around early experiential learning. If AI performs tasks traditionally completed by students and early-career professionals, educators have to consider where people will develop the intuition, judgment, and meta-skills previously built through doing that work themselves.
This creates a much more useful framework for AI assessment than simply categorizing AI as allowed or prohibited.
Educators can instead ask:
What knowledge does the student need to acquire?
What skill does the student need to practice?
What thinking should the student be capable of doing independently?
Where can AI deepen that learning?
And where would AI remove the experience responsible for developing the capability in the first place?
Those questions connect AI policy directly to learning outcomes.
AI Assessment Strategies May Need to Become More Adaptive
Another conclusion emerged as the panelists compared their approaches: there may not be one post-AI assessment model.
The appropriate strategy depends on context.
A first-year composition student, graduate public health student, engineering major, and experienced professional are not developing the same capabilities. The value of reflection, documentation, AI assistance, independent work, or process-based assessment will consequently differ.
The nature of the task matters, too.
Panelists discussed moving toward a more adaptive conception of curriculum and pedagogy in which educators select approaches according to the learner, the task, and where the student currently sits in the learning process.
That is a significant departure from searching for a universal “AI-proof assignment.”
Higher education may instead need a portfolio of assessment strategies capable of producing different forms of evidence.
Sometimes the final product will remain valuable.
Sometimes the process will matter more.
Sometimes students may need to defend their reasoning.
Sometimes AI collaboration itself should be evaluated.
And sometimes students may need to demonstrate that they can perform important cognitive work without AI at all.
The common denominator isn’t the assessment format.
It’s whether the assessment produces credible evidence that the intended learning actually occurred.
Faculty Need Time to Redesign Assessment for AI
There is also an institutional problem hiding beneath the pedagogical one.
Redesigning assessment is work.
Faculty members have to understand what generative AI can do, reconsider existing assignments, experiment with new approaches, observe what happens, compare results, and continue adapting as both students and AI systems change.
The panel repeatedly returned to the importance of collaboration in that process.
Many higher education instructors were never formally trained as instructional designers. Some institutions have teaching and learning centers capable of supporting experimentation, while others have significantly fewer resources.
Meanwhile, faculty members are already being asked to fulfill demanding teaching, research, administrative, and service responsibilities.
Expecting every instructor to independently reinvent assessment on top of those responsibilities is unlikely to produce systemic change.
Panelists consequently emphasized providing faculty with time, support, training, and communities of practice where educators can compare approaches and challenge one another’s assumptions.
That may prove just as important as any individual AI assessment strategy.


