Content and Authority for AI Answers

What You Actually Buy When You Buy an AI Detector

What You Actually Buy When You Buy an AI Detector
Ahn Mai offers this commentary on what you actually buy when you buy an AI detector. This article originally appeared in Insight Jam, an enterprise IT community that enables human conversation on AI.

Every AI-detection vendor will give you an accuracy figure. Very few evaluation processes ask the second question, which is the one that determines what the product does to your students: not how often it is wrong, but who it is wrong about. Those are two different products, and only one of them appears in the procurement pack.

The Evidence Buyers Should be Reading

A widely cited study tested seven off-the-shelf GPT detectors against 91 TOEFL essays written by people who do not speak English as a first language, alongside 88 essays by US eighth-graders (Liang et al., Patterns, 2023). On the eighth-grade essays the detectors were near-perfect. On the TOEFL essays the average false-positive rate was 61.3 per cent. All seven detectors unanimously flagged 19.8 per cent of them, and 97.8 per cent were flagged by at least one.

Two caveats matter for procurement and are usually dropped. The TOEFL essays came from a Chinese educational forum, so this is one population of second-language writers rather than all of them. And Turnitin was not among the seven tools tested, a detail that gets misreported in both directions.

The Mechanism is the Product, Not a Defect

The researchers then changed the writing instead of the tools. Prompting a model to enrich the vocabulary of those same TOEFL essays cut the false-positive rate from 61.3 per cent to 11.6 per cent. Running the reverse on the US essays, simplifying word choice to imitate a second-language writer, produced a substantial rise in misclassification; the preprint version of the same study puts that rise at 5.19 per cent to 56.65 per cent.

After that single intervention, only one of the 91 essays was still flagged unanimously by all seven detectors.

Read those numbers together. The output moved by roughly fifty points while the authorship stayed exactly the same. What the classifier measures is linguistic range, and it reports range as authorship. That is not a configuration problem you tune out during onboarding. It is what the product does.

The Counter-Evidence, Which Buyers Also Need

The record is not one-sided. A separate evaluation found that a widely used generative-AI detector produced zero false positives across 160 genuine student assignments (Gosling et al., 2024). That result is real, and it should stop anyone treating detection as uniformly broken.

It also hands buyers the correct question. Both findings can hold at once, because they test different tools, on different writing, at different thresholds. So the question is never whether detection is accurate. It is: measured on which corpus, at which threshold, and how closely does that corpus resemble the submissions my institution actually receives?

One per cent is a number about your own cohort

Vanderbilt University worked this through in public. Turnitin reported a false-positive rate of roughly one per cent at launch. Vanderbilt had submitted around 75,000 papers through the system in 2022, which implies about 750 pieces of genuine student work incorrectly flagged. Vanderbilt disabled the tool.

A headline error rate is not a property of the software. It is a figure you have to multiply by your own volume before it tells you anything about your own risk.

The Workflow Nobody Bought

Detection budgets are not falling, and the reason is legitimate. HEPI’s 2026 student survey found that 94 per cent of UK undergraduates now use generative AI to help with assessed work, and that the share of students putting AI-generated text directly into a piece of assessed work rose from 3 per cent in 2024 to 8 per cent in 2025 and 12 per cent in 2026. Real use is climbing, so the pressure to buy something is real.

The problem is what got bought. Detection products were procured with a workflow for flagging. Very few were procured with a workflow for appeal, and that asymmetry is now visible in the field.

ABC News reported that Australian Catholic University logged close to 6,000 alleged misconduct referrals across its nine campuses in 2024, around 90 per cent of them AI-related. The university states that approximately half of confirmed breaches involved undisclosed AI use, that around a quarter of all referrals were dismissed after investigation, and that any case resting solely on a Turnitin AI score was dropped immediately. ACU also disputes the ABC figures as substantially overstated. That disagreement is the part buyers should notice: an institution running detection at scale could not produce an uncontested account of its own error rate.

ACU switched its detector off in March 2025. Curtin University has disabled Turnitin AI writing detection across all campuses and study periods from 1 January 2026, while keeping text matching in place. At least a dozen of Australia’s 43 universities still run detection. The regulator, TEQSA, told the ABC that detection tools are permitted and simply not recommended on their own, and that it holds no data on which universities are using them.

3 Questions for the Evaluation Pack

Ask the vendor for the assumed false-positive rate, the corpus it was measured on, and the linguistic composition of that corpus. If the answer stops at the first part, you have been handed a marketing number rather than an error profile.

Then multiply that rate by your annual submission volume and by your share of international students, and ask whether your appeals capacity can absorb the resulting number of contested cases.

Finally, apply one operational test before you sign. With the score removed, could a panel still make its case? If it could not, the product has not given you evidence. It has given you a reason to look more closely, which is a legitimate thing to buy, but only if you also buy the process that has to follow.

Share This

Related Posts