I’m pausing our series on the history of American education to discuss something that we’ve begun to encounter in our work with client districts: the use of AI to grade student essays and other work. AI — artificial intelligence — is cropping up everywhere, even as the ethics of its use have yet to be defined or even thoroughly debated. Teachers and professors appear united in barring students from using AI to generate essays and other written work. But there is a growing trend of teachers themselves using AI to grade that same work and that poses a thorny question of whether using AI in this way is ethical or even advisable. And AI isn’t quite the miracle some of its proponents claim.
AI operates from enormous databases of factual* information. The algorithm is trained on these massive data sets to analyze patterns and relationships between words and phrases. It uses that training to generate new content — that’s the main reason these engines are called “generative AI.” To some extent it works like predictive text in Microsoft Word or on an iPhone: it’s trying to predict the next word, the next idea you’re looking for accurately — not necessarily correct, factual information, just something that conforms to what you are searching for. In fact, AI isn’t capable of determining truthfulness; to it, all digital information is a valid source for generating a match with the request. If you ask ChatGPT to write a newscast in the style of Lincoln’s Gettysburg Address, it draws on examples of newscasts and Lincoln’s speech to generate a new text. That kind of exercise is harmless and entertaining, but other uses (like chatbots that provide information to customers) have become demonstrably problematic. The reason for this is that AI is not — and never has been –infallible, and in fact, appears to be getting factually shakier with each subsequent generation.**
AI engines like ChatGPT, DeepSeek, and others are prone to what their creators call “hallucinations” — demonstrably incorrect or fabricated information, even fabricated articles and URLs and blends of truth and fiction that become very difficult to pick apart. In 2023, leading AI developers told the New York Times that these hallucinations would be solved. Instead, as of 2025, they have only gotten worse and no one knows why. That New York Times article cites independent research that shows hallucination rates are rising. Systems that rely on a reasoning model can have hallucination rates as high as 27%, meaning that almost 1/3 of returned results are entirely made up by the AI engine. This is complicated by the fact that AI generates this bad information so confidently that people fail to recognize the error and use it to complete tasks and/or pass it along to others. A different New York Times article cited a case where a chatbot used to produce legal research for a court case both invented wholly fictional cases and confidently asserted that they were available in major legal databases. They weren’t. While researchers and developers of AI systems know this is happening, they still don’t know how to prevent it.
And that’s just the stuff AI made up. Factual accuracy rates are also highly variable, depending on the task and the AI engine used to complete it. The BBC published research earlier this year (2025) on tests it conducted using 4 four prominent, publicly available AI assistants – OpenAI’s ChatGPT, Microsoft’s Copilot, Google’s Gemini, and Perplexity. They narrowly constructed the test to be solely questions about news using the BBC’s own news database as the source. The results are striking: 51% of all AI answers to questions about the news had significant issues, 19% of AI answers which cited BBC content introduced factual errors such as incorrect factual statements, numbers and dates, and 13% of the quotes sourced from BBC articles were either altered or didn’t actually exist in that article (i.e. were hallucinations). Some of these errors were really basic, like asserting that a prime minister was still in office when he was not or inaccurately citing health information from the NHS. Others ascribed adjectives like aggressive and restrained to countries involved in an armed conflict, even though those adjectives had nowhere appeared in BBC coverage. These results are disturbing in such a narrowly constructed test.
That’s an appropriate segue to yet another fundamental problem plaguing AI: because the databases AI’s algorithms rely on are produced and curated by humans, they are prone to human biases. In 2022, researchers at USC Viterbi examined two of these AI databases to determine whether there was any bias — and what types — in the factual information they generated. While one database showed just 3.6% bias, the other clocked in at 38.6% biased information. Those biases — involving gender, race, nationality, religion, and occupation — were truly vile and I won’t include them here, but they were bad enough that the researchers debated whether to include them in an academic paper. Another study from 2023 found that images generated by AI perpetuated gender and racial stereotypes. For example, asking AI for a picture of a CEO produced white men 97% of the time, even though only 88% of CEOs are white men. Other positions of power, like Director, also mostly produced pictures of white men. AI also associated people of color with positions like taxi driver and maid, and women with positions like dental assistant and receptionist. On the flip side, descriptive requests like compassionate returned pictures of women while intellectual and unreasonable returned images of men.***
I wonder though, if these biases should have been anticipated: way back in 2016, ProPublica conducted a similar study of an AI system used to evaluate the risk of reoffending for people convicted of crimes. This information was then provided to judges prior to sentencing for a current crime. The study found that the AI tool incorrectly predicted much higher rates of potential recidivism for Black people, and substantially lower rates for White people. Because they were generated by an algorithm, these predictions were assumed to be neutral when they weren’t neutral at all: they were demonstrably biased against Black people and toward White people, even when the subject’s prior criminal history (or lack thereof) didn’t support that bias. Although this was noted in 2016 it’s still relevant today because AI’s own developers and researchers admit that they don’t know how to prevent biased, false information from proliferating. And by “proliferating” we mean this: the USC Viterbi study found that when AI used biased data to provide information, the information it produced was even more biased than the original data.**** Rather than controlling for biases, AI amplified them, sometimes significantly.
So: AI is a tool, based on information that is biased to an unknown degree, that can and often does generate incorrect and/or entirely false information in a highly confident way that sounds completely plausible (O’Brien, 2023). Those are severe limitations that require us to approach anything produced by AI with extreme caution and careful vetting for accuracy and biases.
And it should leave us asking if AI should — or even could — play a role in grading student work. To answer that, we need to examine how it’s being used already and with what results. We’ll get to that in Part 2.
_________________________________________________
*Hold that thought, because whether those databases are strictly factual is a major part of the problem with AI usage.
**The University of Maryland has a great set of articles explaining the ways AI creates content that appears reliable and factual but is, in fact, made up.
***And this is really problematic because some of these AI imaging tools are being marketed to police forces to auto-generate sketches of suspects from witness statements. If negative descriptors are associated with particular ethnicities or religious groups or genders, the potential for harm here is terrifying.
****Remember, generative AI takes all that data and creates new texts, new formats, new applications. Because of this, it actually doubles down on whatever biases are already inherent in the data.


