Jul 15, 2026 | 14 min read
What a 96% Midterm and a 48.6% Final Exam Reveal About Unproctored Assessments
A welfare economics class at Brown University took a take-home midterm this spring. The average score was 96 percent, well above the course’s historical range of 65 to 80 percent. When the professor moved the final exam in person, the average dropped to 48.6 percent, a result he called a historic low for the course. Nineteen students ultimately failed.
That gap is the story. It is not really about one professor, one class, or one university. It is about what happens when an assessment has no layer of oversight between the moment a student submits an answer and the moment a grade gets issued.
What happened in the Brown University case
According to Inside Higher Ed, economics professor Roberto Serrano gave a take-home midterm for the first time in nearly two decades of teaching the course, largely to accommodate students who did not want to sit in a classroom exam after a shooting on campus in December. Enrollment in the class had also grown to 86 students, up from a typical 30, which he attributed to the promised take-home format.
When the midterm scores came back unusually high, he and his graders ran the exam through ChatGPT. Separate reporting from Fox News Digital puts a finer point on the scale of it: 40 students earned a perfect score on the midterm. The AI produced answers that closely resembled what his students had submitted, using a proof method that technically worked but read as unnatural for the level of the course. He told students he suspected widespread AI use and switched the final exam to an in-person, closed-book format. Eighteen students dropped the class outright, and nine more stayed enrolled but skipped the final exam entirely, meaning 27 students in total did not sit for it. Of those 27, 22 had scored a perfect 100 on the midterm. The final exam, taken by the 59 students who remained, averaged 48.6 percent, compared with a historical low of 65 percent for the course. The midterm was voided.
The university’s response has been a point of dispute. As of the initial reporting on July 8, Serrano described the process as slow, saying Brown’s Standing Committee on the Academic Code had not acknowledged evidence he first shared in late May. Serrano later wrote in an op-ed for The Free Press that he does not believe the university would have acted at all had his account not first been published by El Pais in late June and then picked up by Inside Higher Ed. Brown disputes that framing. In a statement reported by Fox News Digital on July 13, the university said its academic leaders had been in contact with Serrano as early as May, that he provided the documentation the Standing Committee needed on July 8, and that the committee is now moving the case forward under its normal procedures.
Whichever account of the timeline is more complete, the underlying pattern holds: resolving this case required a professor to independently investigate 86 exam responses, go public across three separate news outlets, and press the matter for roughly two months before a formal review process began in earnest.
Why this is a detection problem, not a discipline problem
It is worth separating two different failures here, because they call for different fixes.
The first failure was upstream. There was no oversight mechanism in place during the exam itself. A take-home format with unlimited time and no monitoring gives a student every opportunity to use an AI tool, and gives the institution no record of how the answer was actually produced. The professor’s suspicion, however well founded, arrived only after the fact, based on score patterns and stylistic analysis he did personally.
The second failure is what happens next. Brown’s own newly published committee report on generative AI in teaching and learning found that three-quarters of surveyed faculty are concerned about AI-enabled cheating, a figure that lines up with a 2025 national faculty survey from AAC&U and Elon University, which found that the large majority of faculty nationwide are worried about AI’s effect on academic integrity and student learning. The Brown report recommends that the university avoid over-reliance on detection tools with known false-positive and false-negative rates, and that it move institutional codes toward addressing the specific mechanics of generative AI misuse rather than relying on generic academic dishonesty language.
Both of these are structural gaps. Neither is solved by asking one professor to build 86 individual cases from memory and suspicion after the semester has ended.
What a human-reviewed process would have looked like
This is where the distinction between automated and human-backed oversight actually matters, and it is worth being precise about it rather than making a sweeping claim.
An AI detection tool run after the fact, as the professor described doing informally with ChatGPT, can surface a pattern. It cannot make a defensible determination. Independent research on AI-detection accuracy backs this up: one analysis of 14 detection tools found false-positive rates as high as 50 percent and false-negative rates as high as 100 percent depending on the tool, with accuracy dropping further once AI-generated text was lightly edited or paraphrased. A tool like this has no way to distinguish a genuinely unusual but honest answer from an AI-assisted one, and it cannot document its reasoning in a way that holds up if a student challenges the finding, which is exactly the dispute now unfolding at Brown.
A proctored, human-reviewed exam changes the sequence entirely. Instead of a professor discovering an anomaly in aggregate score data weeks later, a trained reviewer looks at the flagged session at the time it happens, with the actual exam conditions in view, not just the output. That review becomes a documented record: what was flagged, who looked at it, and why a decision was made. When a result is challenged, whether by one student or by dozens at once, the institution has something concrete to point to instead of reconstructing intent from a grade curve.
This does not mean every exam needs to become high-stakes and adversarial. Brown’s own committee is right that de-emphasizing punishment and building AI-aware academic codes matters too. But policy language only works if there is a way to apply it consistently and fairly at the point of assessment, not just after a problem has already reached this scale.
The scaling problem administrators are underestimating
The academic integrity director quoted in the article makes an important point: faculty are not compensated or incentivized to build cheating cases against dozens of students at once, and most institutions do not staff for it. That is a resourcing problem as much as a policy one. Tricia Bertram Gallant, who directs the Academic Integrity Office and Triton Testing Center at the University of California San Diego, made this point directly in the original reporting.
It is also precisely the moment where AI-assisted cheating changes the math. A single suspicious paper is manageable for a professor to investigate. Eighty-six unusually strong exam answers, submitted in an unmonitored take-home format, is not something any single instructor can reasonably adjudicate alone, especially without institutional support. Programs that rely entirely on faculty judgment after the fact, with no monitoring layer during the assessment itself, are the ones most exposed as AI tools become more capable and harder to detect through scoring patterns alone.
The takeaway for programs issuing results that need to hold up
Every program that issues a grade, a certificate, or a credential is implicitly telling the people who rely on that result that it means something. When an assessment has no oversight during the exam and no documented review process afterward, that claim becomes hard to defend the moment it is questioned, which is exactly what is playing out at Brown right now.
The fix is not more automation layered on top of automation. It is a human reviewer in the loop before a result is finalized, producing a record that can actually be defended. That is a different design decision than choosing a take-home format for convenience, and it is one institutions can make before their own version of this story runs in the news.
Don’t wait for a semester-end surprise.
Human-reviewed proctoring catches what automated tools and after-the-fact suspicion miss, before a result is issued, not after.
