Will AI really turn us into paperclips?

Thought experiments such as Nick Bostrom’s ‘paperclip maximiser’ tell us less about artificial intelligence than about our own intuitions and fears.

The front cover of 'The Iron Man' by Ted Hughes.
The front cover of 'The Iron Man' by Ted Hughes. Credit: Retro AdArchives

Thought experiments have always played an outsized role in the philosophy of mind, in part because the central problem is so difficult to settle by observation alone. Scientists can examine the brain in extraordinary detail. They can measure electrical activity with electroencephalograms (EEGs), track changes in blood flow and oxygenation with functional magnetic resonance imaging (fMRI), and construct increasingly high-resolution models of the functional relationships involved in perception, cognition and behaviour. But a description of how a brain works does not by itself explain what it is to have a mind, or what it means to be a conscious subject.

Gottfried Wilhelm Leibniz, writing three centuries ago, asks us to imagine a machine capable of thought, feeling and perception, enlarged until we could walk inside it, as though entering a mill. What would we find? Only parts pushing against other parts, he says in section 17 of the Monadology (1714), ‘but never anything by which to explain a perception’. His argument, later dubbed ‘Leibniz’s Mill’, is meant to show that the stuff of mind – thoughts, perceptions, feelings and all those qualitative states philosophers call qualia – is not discoverable simply by wandering through the stuff of the brain, or, in his formulation, a machine.

Later philosophers would grapple with the mystery of mind along similar lines. In 1995, the philosopher David Chalmers distinguished the relatively ‘easy’ problems of explaining cognitive functions and behaviour from what he called the ‘hard problem’ of consciousness: why should all the brain’s physical and functional processes be accompanied by subjective experience at all? Neuroscience seems to reveal ever more about the functions and structures of the material brain. Again, though, as with Leibniz’s Mill, merely locating and describing those mechanisms does not seem to tell us where consciousness is, or why it is there at all.

Leibniz offered us a metaphysical critique of materialism that we are free to reject, of course, but even materialists can see the attraction of his gedankenexperiment. Instead of attempting to peer more closely into the machine, Leibniz makes the machine itself into an enormous object. We can even walk through it. And still there’s no sign of a mind. We can’t walk through the chambers of our brain. But a thought experiment like this allows us to speculate about what we might, or might not, find there if we could.

Intellectual history abounds with clever thought experiments attempting to make clear ideas or viewpoints seemingly perpetually obscured by a history of opposing sides and complex arguments. Leibniz’s Mill, in various forms, recurs throughout modern philosophy. Thought experiments bedazzle scientists, too. Einstein recalled wondering, as a 16-year-old boy, what the world would look like if he could run alongside a beam of light. His intuition would lead, by degrees, to the theory of special relativity. Special and later general relativity were ripe for thought experiment; Einstein would also conjure hypothetical situations on fast-moving trains, stationary clocks and observers moving relative to one another on Earth and in space.

Much earlier, Isaac Newton gave us his ‘falling apple’ story. The anecdote has acquired the status of legend, but Newton himself really did recount something close to it late on in his life: watching an apple fall prompted him to ask whether the same force pulling the apple toward the ground was also continuously pulling the Moon toward Earth. This was a fantastic speculation. Newton saw the significance of the Moon’s curved path: its forward motion should send it out into space, yet it was perpetually falling around, rather than into, the planet. Throw the apple hard enough and it might go into orbit, too. Newton’s fancies collapsed terrestrial and celestial mechanics into a single imaginative picture. On and on it goes. Thought experiments can help us see something that’s not obvious, or that’s obvious but too often obscured in contemporary discussion.

Back to the philosophy of mind. In 1980, Berkeley philosopher John Searle imagined himself locked inside a room receiving strings of Chinese characters or ‘squiggles’ on otherwise blank pieces of paper. Searle knows no Chinese, but he has an instruction book replete with a sufficiently elaborate set of rules, written in English, telling him which Chinese symbols to return in response to others he receives (through a slit in the otherwise closed and windowless room). Eventually his answers become indistinguishable from those of a native Chinese speaker. To observers outside the room, the performance is impeccable. Inside, Searle still has no clue how to read or interpret Chinese. But if questions get answered correctly in Chinese, without any actual understanding of the language, what does this say about AI?

Searle introduced the thought experiment in his 1980 paper ‘Minds, Brains, and Programs’. Its publication was accompanied by 27 critical commentaries from philosophers and cognitive scientists, along with Searle’s replies, an early indication of how thoroughly the Chinese Room would divide the field. Searle’s point was directed at what he called ‘strong AI’, the notion that a sufficiently powerful AI could reproduce all the qualities of the human mind.

Like Leibniz centuries before, his gedankenexperiment was a pull-back on hype, a cautionary tale: formal symbol manipulation, however successful behaviourally, does not by itself establish understanding. Today these visions seem particularly germane to discussions about language models like those produced by a welter of Big Tech companies such as OpenAI, Anthropic, Google, Meta and others.

AI researchers and philosophers have shown some reserve in attributing metaphysical concepts like ‘mind’ to machines. No matter, as there is no dearth of thought experiments about AI: the field has generated a new crop of experiments of its own. The key move is from mind to intelligence. Leibniz and company asked whether outward behaviour, physical structure, or mechanical operation were sufficient to establish the presence of a mind. The newer experiments look at the implications once we grant that a machine is intelligent, whether there is anything going on inside, or not.

The pivot from mind to intelligence seems straightforward, perhaps, but invites its own form of confusion. On the one hand, AI enthusiasts seem transfixed by the possibility that advanced AI might ‘come alive’, and acquire a mind of its own. In 2014 Elon Musk worried publicly that we might be ‘summoning the Devil’ with AI. Musk joins a cacophony of existential risk worriers and techno-futurists who fret and dream about future artificial minds. On the other hand, though, the same cadre of futurists are quick to dismiss armchair speculation about what minds are, preferring to focus on what they see as questions that can be more easily addressed. These questions revolve around what AIs might conceivably do, and especially whether they might become autonomous and decide to exterminate humanity.

The philosopher Nick Bostrom, known for his 2014 bestseller Superintelligence: Paths, Dangers, Strategies  supplied one of the most memorable examples of a thought experiment about the agency of AI in 2003, in a widely read paper, ‘Ethical Issues in Advanced Artificial Intelligence’. In it, Bostrom argued that a future superintelligence need not share anything resembling human motives or values. In this case, an AI could quite innocently, but with ruthless superior machine intelligence, end up exterminating us all.

We might get turned into a paperclip, for example.

Imagine, he suggested, a superintelligence whose mission is simply to manufacture as many paperclips as possible. As he puts it:

Artificial intellects need not have humanlike motives. Humans are rarely willing slaves, but there is nothing implausible about the idea of a superintelligence having as its supergoal to serve humanity or some particular human, with no desire whatsoever to revolt or to ‘liberate’ itself. It also seems perfectly possible to have a superintelligence whose sole goal is something completely arbitrary, such as to manufacture as many paperclips as possible, and who would resist with all its might any attempt to alter this goal. For better or worse, artificial intellects need not share our human motivational tendencies.

The thought experiment asks us to imagine an entity of extraordinary intelligence combined with a seemingly innocuous purpose. Maximising paperclip production seems like a reasonable enough goal for a company selling paperclips and in need of more supply. The instruction turns tragic, however, because the AI is so clever and unburdened with human values that it uses its advanced intelligence to pursue a truly maximal goal: turn everything – every atom – into a paperclip. Don’t stop at the atoms in the bodies of the employees of the paperclip factory. Don’t pass by the CEO, either.

By way of explanation, Bostrom offered his orthogonality thesis: intelligence is largely independent of purpose (the concepts ‘pull apart’, as philosophers would say). A highly intelligent agent might pursue almost any final objective. Witness his paperclip maximiser, a scenario that’s troubling precisely because it’s about a hypothetical course of action, not an idea. If the paperclip maximiser merely outputs better plans for manufacturing paperclips when asked to do so, we’re fine. Apocalypse averted. The scenario becomes frightening only when the system begins acting for itself, acquiring resources for example, and preventing itself from being switched off. Here the objective has been stripped from any context or common sense. The supposedly intelligent AI resists any modifications to its objective. We can imagine, with Bostrom, that it sets about improving its own abilities. It anticipates human interference and works to disguise its actions and side-step attempts at reining it in.

Scary. But readers may ask: why should instructions to increase paperclip manufacture instigate any of this? Who approved the ‘new plan’? It’s not clear whether paperclip maximising supports his orthogonality thesis or contradicts it.

The researcher Stephen Omohundro supplied an influential answer to these concerns in his work on ‘basic AI drives’ in 2007 and 2008. A sufficiently capable goal-directed system, he argued, would tend to develop certain instrumental objectives, even if they had not been programmed as final goals: self-preservation, for example, because a system cannot accomplish its objective if it has been destroyed. Resources are useful only so far as they permit a more effective pursuit of the objective. Protecting the objective itself is useful because a machine whose goal has been changed may no longer accomplish the original goal. Innocent instrumental objectives could lead to calamity.

Bostrom later systematised this idea as the instrumental convergence thesis: agents with very different final goals may converge on similar intermediate goals because those intermediate goals help accomplish a wide variety of ends.

This sounds technical, but the underlying idea is simple enough: intelligent agents may do terrible things in pursuit of their goals. Bostrom apparently intends his thesis and the ideas around his now-famous paperclip example to be a cautionary tale about advanced intelligence, but a more straightforward reading is that the paperclip maximiser looks less and less intelligent – not more – as it literally pursues its own objective. Strip away context, common sense and any understanding of what the factory was built to produce, and you get a picture of mechanical stupidity armed with enormous power.

We can describe the resulting behaviour in terms of goals and subgoals, but on one straightforward reading, the inference to high intelligence seems gratuitous and counterintuitive. To get Bostrom’s desired outcome, it helps to anthropomorphise, imagining the machine scheming and lying and resorting to subterfuge to reach its ultimate aim. That’s not just a machine pursuing a route towards goals and subgoals – that’s more like a mind. There’s an ambiguity here that sensitive readers will notice.

The problem becomes even more obvious in later AI-risk scenarios deliberately involving deception, or what Bostrom calls a ‘treacherous turn’. An artificial agent behaves cooperatively while it remains weak, perhaps because appearing cooperative is instrumentally useful. Once it becomes sufficiently powerful, it ceases cooperating and pursues its objective directly. (Cue angst.) Notice how psychologically rich this description is, even when its proponents insist that no psychology is required. Perhaps every one of these operations can be specified mechanically. Perhaps not.

This brings us back to Leibniz’s Mill. The older philosophy-of-mind thought experiments created vivid imaginary worlds to get at deep mysteries like the mind-body problem. The paperclip maximiser does something similar with optimisation: give an enormously capable system an open-ended objective and even something as trivial as a paperclip can end in global catastrophe. But is the experiment clarifying the mystery, or making it more opaque?

Leibniz thought mind existed; his Mill asked how that could be demonstrated. Readers could remain puzzled or simply disagree. Bostrom’s thought experiment feels different. It does not merely ask us to imagine what intelligence might do; it builds in a particular answer about how intelligence behaves, even when that answer rests on debatable assumptions about intelligence. Would a truly intelligent agent really act this way?

A very different AI thought experiment appeared in 2020, when two computational linguists, Emily Bender and Alexander Koller, published ‘Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data’. They ask us to imagine two people stranded on separate islands who communicate through an undersea cable. An extraordinarily intelligent octopus taps the line, studies their exchanges and becomes so adept at predicting the patterns in their messages that it can eventually cut the cable and impersonate one person talking to the other.

This is a modern variation on Searle’s Chinese Room: successful linguistic performance need not imply understanding. But here Bender and Koller imagine a statistical learner that extracts patterns from a language while remaining cut off from the world the words describe. The octopus just has access to the form of language. When genuinely novel circumstances arise – a request involving a coconut catapult or an attack by a bear – they argue that the limitation becomes obvious. The octopus has learned patterns in signals without acquiring the context behind them. Their paper, which won the Association for Computational Linguistics’ 2020 Best Theme Paper award, used the example to clarify the distinction between linguistic form and meaning. It was, of course, an attack on the notion that language models understood language because they could interpret and generate it given a prompt. The target then was models such as BERT, launched in 2018, well before OpenAI released ChatGPT in November 2022.

We might pause here to ask why thought experiments dominate discussions of artificial intelligence. For one, AI repeatedly forces us to reason about entities that either do not yet exist or whose internal capacities remain disputed. We cannot put Bostrom’s superintelligence in a laboratory because there is no such machine. We cannot settle questions about understanding simply by pointing to the fluent output of large language models because the dispute concerns what that output means in a deeper sense. Thought experiments let us isolate one assumption, amplify it, subtract another, or imagine situations technology has not yet produced, but may produce in the future.

In the end, Leibniz’s Mill, Chinese rooms, paperclip factories and hyperintelligent octopuses are instruments for thinking, for us. And like the choice of any instrument, what each one reveals depends on all-too-human intuitions and fears.

Author

Erik J. Larson

Erik J. Larson is the author of 'The Myth of Artificial Intelligence: Why Computers Can’t Think the Way We Do' (Harvard University Press) and co-author of the forthcoming 'Augmented Human Intelligence: Empowering Minds in an Age of AI' (MIT Press). His fiction includes the novel 'Benderland'. Learn more at larsoninstitute.org.

Download The Engelsberg
Ideas app

The world in your pocket. The app brings together – in one place – our essays, reviews, notebooks, and podcasts.

Download here