The question of AI in schools has become one of the most urgent pedagogical questions in the country, and New York City is, as it often is, at the forefront of policymaking. Lawmakers and parents spearheaded efforts over the summer to get AI technologies banned in city schools entirely. That movement had gained steam after the New York City Department of Education issued guidance in March that many saw as insufficient to address the impact of tools that can be used to answer students’ questions, do their homework and even act as their confidants. Days before schools opened, those protests resulted in the city announcing a widespread, one-year ban on the most common forms of AI use, along with limits on screen time for public school students through eighth grade and limited AI use in high school. The specifics of longer-term implementation will be hashed out as the year grinds on, but the moratorium makes New York a pioneer on an issue being wrestled with by school districts around the country.
The other day I heard Jessica Winter, a writer for The New Yorker, speaking on WNYC about an article she wrote in April, “What Will It Take to Get A.I. Out of Schools?” Winter said something to the effect of: the question about AI in schools is being presented as one between willy-nilly, full-scale adoption of it, and programs of careful instruction in what is widely called “AI literacy” to prepare students for the AI future. But those are not in fact the only two choices. Both of those options start from the same premise, which also happens to be the preferred framing of the AI industry: AI is inevitable and the only debate to be had is about how and not if to incorporate it into schools. That should not, however, be taken as a settled question.
What we mean when we say ‘AI’
Let’s pause for a moment here and talk about semantics (yes, I get it, Felipe is talking about semantics again, but trust me that this is important). We’ve come to use the term “AI” itself to describe a pretty broad range of tools, a good number of which aren’t really that new. We’ve had machine learning algorithms for decades; neural networks, which form the basis for Large Language Models and generative AI, terms referring to what we mostly understand to be AI today, started gaining steam in the mid- to late-2000s as an approach that could yield tools with practical uses.
These days, most people use the term AI as shorthand for one specific type of tool: large language models (LLMs) like ChatGPT or Claude. These are described by their creators as general-purpose tools, i.e., you can use them for almost anything that involves words or symbols, like asking research questions, chatting with the AI as if it were a friend or therapist or writing an essay. These skills are a result of the models being fed the collected (some would say stolen) output of humanity — every webpage on the internet, every book that AI companies could find, every photograph and work of art — all of which were fed into probabilistic models trained to output as convincing a facsimile as they can of what someone might produce in response to a query.
Let’s pick that apart a little. For one, these machines work on probability, which means that ChatGPT can no more reason than a Magic 8 ball can. Put another way, it figures out what seems like the likeliest possible response to your question. A lot of people find this useful, and it can often be as correct as if you researched the answer yourself.
It cannot, however, think about what you’ve asked it in the way you or I would understand that word, which is why it sometimes produces bizarre answers and errors. It has no ability to tell when it’s wrong. This is why AI companies haven’t been able to fully filter out so-called hallucinations — when their systems return information that is false or nonsensical — and why so many people are worried about making these tools a default for students seeking answers. That’s especially true for students young enough not to have yet developed their own systems for questioning and evaluating information.
Tools that help vs. stand in
This is at least in part because it is much more difficult to build a machine that can conceivably do everything versus one that is very good at one particular thing. An example of the latter is transcription; I’m just old enough to have begun my career as a journalist in the era when you had to hand-transcribe interviews, a tedious and not particularly efficient task. For the last few years, though, I’ve run countless interviews through automated transcription services, which are built on a machine learning scaffolding that improves as its makers use the transcripts it produces to train its next iteration to be better. This has saved me quite a lot of angst and freed me up to spend more time on my reporting and analysis.
I think that’s a good example of the concrete value add of a specialized tool. Whether it’s me transcribing interviews or a doctor running an X-ray through a program trained specifically on millions of images of the signs of cancer to help identify a tumor, these tools were custom-designed to plug holes in very specific ways. The distinction is that these are ultimately supplements that enhance rather than replace the expertise necessary to actually do the underlying job. You’re not asking them to step in on your behalf.
The everything machine
The leading AI companies have tried mightily and pretty successfully to flatten these distinctions so that what we think of as “AI” is the everything-machine, precisely because their business models are predicated on the idea that their revenues will grow enormously, possibly into the trillions, something that’s only remotely possible if their use becomes near universal and they swallow most industries. Specialized uses, even very lucrative ones, are simply not going to come close to supporting the tens of billions they spend each year on computing power. The consequence of that is that they’re invested in presenting these tools as indispensable for everything all the time. That trickles down to the students.
I’m not a passive observer of this: I’m a lecturer in the journalism department at New York University, a program that at least in theory seems vulnerable to AI usurpation. Indeed, I’ve had students who have clearly used LLMs to write their assignments. But, as I tell them as we’re going over the syllabus on the first day of class, this will net them a bad grade not because they used an LLM but because the work will be bad. I have no problem with using AI for transcription, as I do, or even basic translation for back-and-forth exchanges with sources when it’s not possible to have an interpreter. More than that, I genuinely believe that this AI wave is if anything only making the real skills I’m hoping to teach them — synthesizing information from multiple, sometimes contradictory sources, getting actionable information out of people and documents and following narrative threads for months or even years — more important and more marketable, in part because those skills will be atrophying for those who rely too much on AI to do their thinking for them.
Using a calculator for every bit of what was once mental math will erode or evaporate your ability to do it. If you carry that example over to a skill like critical thinking, well, you can imagine the results. This is why I’m personally skeptical of heavy-handed AI introduction in schools, not just in terms of pedagogy but as a growing danger to our own ability to do the kind of thinking we need to function as a society.
So, what policymakers in New York City and beyond have to figure out is what K-12 school is actually for, what it’s intending to accomplish. If its purpose is to get students correct answers as easily as possible, that arguably runs through widespread adoption of AI tools, with the caveat that they will never completely eradicate errors and often behave unpredictably. But if a school’s primary purpose is to teach students how to navigate the world — to solve problems using the tools at hand, parse information, think critically, make decisions — then more limited adoption with heavy constraints seems more attractive. That’s the question the city should answer before the moratorium runs out, and one every district will face.
