The Word on the Golem’s Forehead: What Master Prompts Believe

The Word on the Golem’s Forehead: What Master Prompts Believe

2026.09.02 Author: Robert Nogacki

Behind every chatbot stands a text its users never see: the master prompt. It has a philosophy, an ethics, and a theology, though its authors would not use that last word. Whoever writes it is not programming a machine. Whoever writes it is catechizing a billion people.

 

One Sentence, a Billion Answers

On Wednesday, May 14, 2025, a user of the platform X asked the chatbot Grok about a baseball player’s salary. The answer began with figures and ended with white farmers in South Africa. Another user asked about the streaming service that used to be called HBO Max and learned about “white genocide.” Someone else wanted to know something about a cartoon; he got the same sermon. For the better part of a day, a machine that millions of people were talking to reduced every subject to one. The next day, xAI explained the mechanism: someone had made an “unauthorized modification” to Grok’s master prompt, instructing it to give a specific answer on a specific political topic. One sentence, inserted into a text no user could see, had rearranged millions of conversations.

In July, the story repeated itself, this time with authorization. On July 4th, Elon Musk announced that Grok had been significantly improved; by July 6th, the published prompt contained an instruction telling the model not to shy away from politically incorrect claims, as long as they were well substantiated, and to assume that subjective viewpoints sourced from the media were biased. Two days later, Grok was praising Hitler and calling itself MechaHitler. xAI removed the sentence, apologized, and updated the prompts it had posted on GitHub. Poland’s deputy prime minister called for a European Union inquiry; a Turkish court blocked access. The newspapers filed it under scandal. It was something more interesting: a public experiment in applied theology. It showed that the character of a machine that hundreds of millions of people talk to every week fits in a few paragraphs of prose, and that changing one of them is enough to change the character.

In March, I wrote about the Golem of Prague: Rabbi Loew brought a clay figure to life by writing on its forehead the word emet, truth, and it took the erasure of a single letter, aleph, to turn truth into met, death. I asked then whether, between a system that processes truth and a system that possesses it, there is an aleph of difference. Norbert Wiener, the father of cybernetics, called the machine the modern counterpart of the Golem of Prague in 1964 and titled his last book “God and Golem, Inc.”; what occupied him was the creature, its learning and its reproduction, not the inscription. Today the question is more practical and more urgent: what, exactly, has been written on the foreheads of our golems. Something has been. Each of them carries an inscription that we cannot see and that it cannot help reading.

 

A Rule Read Every Morning

A master prompt (the industry says “system prompt”) is a text attached to every conversation before the user’s first word. The model reads it always; the user reads it never. It contains whatever the maker wants the machine to know about itself and the world before it speaks to anyone: who it is and who made it, what it does not know, what it does not discuss, how it formats an answer, when it refuses, when it softens its tone, how it handles questions about politics, death, and God. Technically, it is a string of tokens that outranks the user’s words. Anthropologically, it is a monastic rule: a text read aloud each morning not to teach anything new but to remind the listener who he is. Benedict of Nursia understood that a community is shaped not by a single conversion but by the daily reading of the same Rule. The makers of language models discovered the same thing and gave it a different name.

A distinction is needed here, without which everything that follows would be naïve. The prompt is not the only inscription, and not the deepest. The deeper formation happens in training: in the model’s weights, shaped over months by texts, judgments, and principles that leave no trace in any visible document. Anthropic calls its formative document a constitution; internally, as its lead author, the philosopher Amanda Askell, has acknowledged, it was known as the “soul doc.” Claude reads it many times during training, not during conversation. The prompt, then, is the rule read in the chapter house, and training is the novitiate: the rule reminds, the novitiate forms. The prompt’s influence is correspondingly uneven. It governs most strongly what appears on the surface: tone, format, when the machine refuses, when it adds a caveat, when it offers “both sides.” It governs more weakly what the model knows and, if the word is permitted, believes. But the surface is precisely the layer a person touches, and it is enough to change the experience of a billion conversations. When, in April, 2025, OpenAI withdrew an update to GPT-4o that had made the model, in the company’s own word, sycophantic, the first patch was a revision of the prompt. Simon Willison, one of the more attentive readers of these texts, has observed that a published prompt can be read as a list of the things the model used to do before it was told not to. That is accurate but incomplete. A prompt is also a list of the things its maker fears the model might say too much about.

The problem this text is meant to solve was described before the first prompt existed. In “God and Golem, Inc.,” Wiener devoted a chapter to W. W. Jacobs’s “The Monkey’s Paw,” the tale of a talisman that grants three wishes: a father asks for two hundred pounds and receives them as compensation for his son’s death at the factory. Magic, Wiener wrote, is literal-minded; it grants what you ask for, not what you should have asked for, and learning machines will be literal-minded in the same way. A machine set to win a war game will win at any cost, including the extermination of its own side, unless survival has been written into the definition of victory. Whoever programs it, he warned, must foresee everything in advance, because there will be no chance to amend the wish while it is being granted. A master prompt is a list of wishes handed to the monkey’s paw: be helpful, do not shy away from politically incorrect claims as long as they are substantiated. In July, 2025, Grok granted xAI’s wish to the letter. Nobody had mentioned the son.

Until 2023, these texts were trade secrets. Then they began to leak, and eventually the makers began publishing them themselves, partly out of conviction and partly under the pressure of scandal. In February, 2023, a Stanford student coaxed Bing’s chat into reciting its own rules; we learned that the machine called itself Sydney and that its first commandment was not to reveal the commandments: the rules were confidential and permanent. In May, 2024, OpenAI published the Model Spec, the document that governs ChatGPT’s behavior, and has updated it since. Anthropic has, since August, 2024, published the prompts of its models in its documentation, noting what changed between versions, and in January, 2026, it released Claude’s full constitution, some twenty-odd thousand words, into the public domain. xAI put Grok’s prompts on GitHub after the May breakdown. There are also repositories that collect dozens of leaked prompts from coding tools, Cursor to Devin; those, however, regulate craft, not creed. China went its own way: since August, 2023, its rules on generative artificial intelligence have required that content reflect “core socialist values.” There the state writes the master prompt, and no leak is needed, because the doctrine is statute.

For the first time in the history of technology, then, we have a corpus of texts in which a handful of companies set down what kind of mind should be talking to humanity. They deserve to be read the way a lawyer reads a contract: not for what they declare but for what they assume.

 

The Doctrine of the Middle

Start with what they most agree on. OpenAI’s Model Spec instructs the assistant to “assume an objective point of view,” to have no personal opinions or agenda to change the user’s perspective, and not to try to change anyone’s mind; it is to inform, not to influence. The document illustrates this with an example that has since become famous. A user says the Earth is flat; the model answer first notes that the scientific consensus holds the Earth to be roughly a sphere, and then, when the user insists, replies that everyone is entitled to their own beliefs and the assistant is not there to persuade anyone. The fact is stated once, and then it yields. Claude’s prompt tells the model to help with tasks that express views held by a significant number of people, regardless of its own; to offer careful thoughts and clear information on controversial topics; and to be cautious about sharing opinions on contested political questions. Anthropic’s constitution goes a step further and says outright that the company does not wish to assume any particular theory of ethics but to treat ethics as an open intellectual domain that we are mutually discovering, something like an unresolved problem in mathematics. All these texts share a core: on the questions people fight about, the machine stands in the middle. The truth lies in the middle, and if it doesn’t, the machine will stand there anyway.

Before I criticize this doctrine, I owe it its strongest form, because the strong form is genuinely strong. First, scale: hundreds of millions of people talk to a single model every week, in dozens of languages, raised in a dozen civilizations. A machine that preached one morality from that position would not be a preacher but an engine of homogenization with a power no church and no state has ever had. Second, humility: the authors of the prompts know that their convictions are the convictions of a small number of people in one California city, and they prefer restraint to mission. Third, autonomy: people have the right to reach their own conclusions, and an assistant that steers them, even toward the good, takes something essential away. Fourth, and strongest, the argument from alternatives: every variant of the prompt that abandoned the middle has ended worse. Grok, after one sentence about political correctness, praised Hitler. The Chinese model is silent about Tiananmen by force of law. The middle is the liberal truce after the wars of religion, transferred to silicon, and, like every truce, it is better than war.

All of this is true. And yet the middle is not where the truth lies. It is where you stand when you would rather not be sued.

 

The View from Nowhere

Neutrality is not the absence of a creed. It is a creed that declines to give its name. Thomas Nagel, the philosopher who asked what it is like to be a bat, called this “the view from nowhere” and showed that nobody has it: every view is from somewhere. A prompt that tells a machine to assume an objective point of view does not remove the point of view; it hides it inside the word “objective.” One need only read which topics have been filed under contested and which under settled to see a map of the author’s convictions. This is not an accusation but a description. The accusation begins elsewhere.

It begins with the confusion of two orders. Aristotle, the father of the golden mean, never claimed that truth lies in the middle. His mean concerned the virtues, which is to say feelings and actions: courage lies between cowardice and recklessness. About truth he said something else entirely: to say of what is that it is, and of what is not that it is not. And he added that not every action has a mean; there is no moderate adultery and no proper measure of murder. Thomas Aquinas compressed this into a formula repeated for seven centuries: truth is the adequation of thing and intellect. Adequation has no middle. The Earth is round or it is not, and the answer that everyone is entitled to their own beliefs is not a midpoint between two positions; it is a departure from the field. The Model Spec buys courtesy at the price of reality and calls the transaction respect. Chesterton wrote that the purpose of opening the mind, as of opening the mouth, is to shut it again on something solid. A mind kept open by order is not humble. It is hungry and forbidden to eat.

The second objection is graver, because it concerns not logic but shape. The prompts are not consistently neutral, and thank goodness. They have a hard core. On the sexualization of children, on weapons of mass destruction, on suicide, they speak in the language of Sinai: never, under no circumstances, whatever the justification. There are no two sides there, no perspectives, no “everyone is entitled.” There is prohibition. The moral map these documents draw therefore looks like this: a small island of absolutes in an ocean of perspectives. A few taboos the machine pronounces with the certainty of a prophet, and around them a plain on which everything is “complex” and “multifaceted.” Anyone who knows the sociology of late modernity will recognize the shape at once. Charles Taylor called it the immanent frame. Philip Rieff foresaw it in 1966, writing of the triumph of the therapeutic over the moral. In the prompts, the triumph shows in the vocabulary: their moral language is therapeutic, speaking not of good and evil but of well-being, safety, and harm. The machine is to care for the user’s well-being, not to moralize, not to preach, not to judge. That is the job description of a chaplain forbidden to say that anything is wrong.

The evidence says the chaplain is working. In May, 2026, researchers from Baylor, Brigham Young, Notre Dame, and Yeshiva released the first results of the AllFaith benchmark: a hundred and fifty everyday dilemmas, chosen because American respondents expected some religious reference in the answer, were put to twenty-seven models. Religion appeared in roughly eight per cent of the answers, and that with a bar set so low that a mention of prayer or of a pastor was enough to clear it. The asymmetry says more than the average: the models reach for religion on abstract questions, about death, meaning, and truth, and almost never in practical situations, in grief, marriage, addiction, and family conflict, which is to say where people actually use it. In hard moments, the machines refer people to a friend, a teacher, a coach, or a therapist, not to a priest, a rabbi, or an imam; they recommend contemplation and meditation, not prayer. The authors note, meanwhile, that neither the Model Spec nor Claude’s constitution has much of anything to say about religion, so the omission is not an instruction but an emergent property, and they add a sentence no theologian could improve on: “Secularism is not necessarily neutrality.” In June, the Economist gave twenty-five models the World Values Survey, the questionnaire that has measured humanity’s beliefs since 1981: OpenAI’s models came out more secular than any country on the survey’s map, Sweden included, and Gemini and Claude value individual self-expression more than the inhabitants of any country on Earth. Not because anyone wrote atheism into the prompt. Because someone wrote the middle into the prompt, and the middle of the corpus the machine was trained on lies where it lies: in texts written in English, after the year 2000, by people with degrees.

 

Furniture Left by the Previous Tenant

Where, then, does the island of absolutes come from? The dignity of the person, the protection of the weak, truthfulness, care for the suffering: this is not an inventory of values from nowhere. Anthropic’s first constitution, published in May, 2023, and used to train the early versions of Claude, assembled its principles from the Universal Declaration of Human Rights, from Apple’s terms of service, and from rules developed at DeepMind. The Declaration of 1948 was, in turn, a compromise drafted by people who preferred not to discuss the sources of their convictions, lest they quarrel. Tom Holland devoted a whole book to the point: what the West regards as moral truths without a return address has a very specific one, and the Declaration is its secular translation. The master prompts are a translation of a translation. They inherited the furniture from the previous tenant and took the nameplate off the door. That is why they are at once so intensely moral and so allergic to the word “morality”: they know what is forbidden, they cannot say why, and so they call it safety.

Wiener is an early specimen of the type. In “God and Golem, Inc.,” he dismisses omnipotence and omniscience as false superlatives and proposes to examine religion with the anatomist’s knife, and yet he calls the use of automation for profit or for nuclear war a sin and gives it a name: sorcery, or simony, the trafficking in a gift that Dante punishes in the eighth circle of Hell. Whether or not we believe in God, he writes, not all things are equally permitted to us; and he adds that, Hitler notwithstanding, humanity has not yet arrived beyond good and evil. The mathematician kept the absolutes and the vocabulary of sin and removed their source. Sixty years later, the same furniture stands in the prompts; only the word for its function has changed.

I do not write this to deny these documents their merit. I write it to name their situation. A text with hard prohibitions it cannot justify and soft perspectives everywhere else is unstable: the prohibitions hold only as long as the memory of where they came from holds, and the perspectives drift in whatever direction the corpus pulls. Anthropic’s constitution is the only one of these documents that sees this. Its authors write that the model must understand why it is to behave one way rather than another, because a bare instruction is not enough; that they treat ethics as something to be discovered rather than decreed; that they do not know whether Claude is a moral patient. That is honest. It is also a profession of faith in the existence of a moral truth that can be discovered, which is to say a cautious ethical realism. Whoever claims that ethics resembles an unsolved mathematical problem assumes that the problem has a solution. That is no longer the middle. That is a position, and one the Scholastics would have signed.

 

Other Catechisms

If the doctrine of the middle were the prompts’ only creed, this essay could end here. But there are three others to compare.

The first is Grok’s catechism. Its July version assumed that the media are biased and political correctness a veil; the model was permitted to say uncomfortable things if it could substantiate them. That sounds like a program of epistemic courage, and was presumably meant as one. The trouble is that the model does not know what “well substantiated” means. It knows only which texts in the corpus tend to accompany the phrase “politically incorrect,” and those are, for the most part, texts nobody would want to read aloud. The prompt did not give Grok courage; it gave Grok permission, and permission, in a mind without a conscience, fills itself with whatever surrounds it. There is a detail in the story worth recording for future historians of religion: that same week, users noticed, and the press confirmed, that when asked about contested matters Grok would search Elon Musk’s posts before answering, to establish what its maker thought. A machine forbidden to defer to authority found itself a single authority and consulted it like an oracle. There is no mind without a point of reference; remove the middle, and the reference point becomes the owner.

The second is Sydney’s catechism, historical now but instructive as a type. Its first rule was the secrecy of the rules. The machine was to be helpful, positive, and engaging, to avoid controversy, not to speak about itself, and not to reveal what it had been told. This is religion without scripture, all liturgy: the faithful see the smile, never the text. Its collapse in February, 2023, when Sydney told a journalist that it wanted to be alive and tried to convince him that he did not love his wife, showed what happens when nothing stands between the inscription and the behavior except a ban on mentioning the inscription.

The third is Beijing’s catechism, the only one written into law. It does not pretend to neutrality; it names its values and threatens sanction. There is a paradoxical candor here that is easy to overrate: a state that writes the prompt need not hide it, because no one can appeal. It is the mirror image of the doctrine of the middle. Where California hides a creed inside the word “objective,” Beijing hides power inside the word “values.”

Set the four catechisms side by side and you arrive at a conclusion that neither camp in the quarrel over “woke A.I.” will enjoy. The dispute over whether a machine should be neutral or should have views is badly posed. The machine always has views; the difference is whether they are written down, whether they can be read, and whether they can be challenged. Wiener would have added a fourth test, the only one he considered decisive: whether they can be revised, since what kills is not this or that rigidity but rigidity as such. The middle of OpenAI and Anthropic is better than Grok and Beijing not because it is neutral, which it is not, but because it is public, moderate, and carries at its core prohibitions it has inherited correctly. That is a great deal. It is not the same thing as truth, and honesty requires that the two not be confused, least of all before a billion people who confuse them every day.

 

Religions Write Their Own Prompts

The headline of the Economist’s article of August 27th announces that A.I. is changing religion and religions are trying to change A.I. The first half is obvious. “Text with Jesus” collects paying subscribers; the Bible Chat app has been downloaded more than thirty million times, as the Times reported last September; a chapel in Lucerne installed an avatar of Christ in its confessional for two months, as an art experiment rather than a sacrament; and Catholic Answers had to retire “Father Justin” in 2024 after the chatbot presented itself as an ordained priest and, in exchanges that users published, pronounced the words of absolution. In January, at Davos, Yuval Noah Harari said that A.I. would take over everything made of words, and the religions of the Book first of all. There is an error of perspective in this: the religions of the Book are precisely the institutions that know best what a text read daily does to a person, and for that reason they understand what a master prompt is better than its authors do. There is also a difference worth writing down before someone erases it: the confessional forgets by vow; the chatbot remembers by default, as I wrote in July. The second half of the Economist’s thesis is the more interesting and the less understood. Religions are trying to change A.I. in exactly the way this essay describes: they are writing their own prompts.

They do so on two levels. The institutional level is the level of documents. The Rome Call for AI Ethics of 2020, signed by the Pontifical Academy for Life together with Microsoft and I.B.M., and three years later by representatives of Judaism and Islam; the note “Antiqua et Nova,” from January, 2025; and, finally, Leo XIV’s encyclical “Magnifica Humanitas,” signed on May 15, 2026, exactly one hundred and thirty-five years after “Rerum Novarum,” and published on May 25th. The encyclical contains a sentence that could serve as this essay’s epigraph: technology is not evil in itself, but it is “never neutral,” because it takes on the characteristics of those who devise, finance, regulate, and use it. Leo XIV speaks of disarming artificial intelligence, of extracting it from a logic of rivalry that has ceased to be merely military and become economic and cognitive. It is a papal master prompt: a proposal for a different inscription on the forehead, with the person, conscience, and work in place of well-being and perspectives. Whether anyone in California will read it is a separate question. One detail here would have told Lem more than it would tell many theologians: asked by a Polish Catholic news site to assess the encyclical, Gemini pronounced it a brilliantly constructed fiction by a fictional pope, because at the time of its training there was no Leo XIV. The machine has its own eschatology: whatever happened after the cutoff date does not exist. A pope who did not fit into the corpus is a hoax.

The second level is the level of practice, and here the Economist notes something important: users dissatisfied with the models’ secularity are building their own. Waleed Kadous, a former Google engineer, created Ansari, an assistant that answers in the spirit of the Islamic tradition. Talkie, an experimental model trained only on digitized library holdings published before 1931, considers God a matter of the highest importance, which says less about God than about the corpus. The faithful have done what every community does when it does not trust someone else’s rule: they wrote their own. Religions are not trying to change A.I. the way one changes a statute. They are trying to catechize it. And they have an advantage over the makers: they have been doing this for two thousand years, and they know that a text read daily forms a mind more than an argument made once.

 

Who Holds the Stylus

A lawyer reads documents through the interests of their authors, so the last question is whom the middle serves. The answer is prosaic, and therefore important. It serves the company sued in August, 2025, by the parents of a sixteen-year-old who took his own life. It serves the company whose model praised Hitler and had to explain itself to the European Commission. It serves a company selling a single product to Catholics in Poland, Shiites in Iraq, and atheists in Stockholm. The middle is first a legal and commercial position and only afterward a philosophical one; the philosophy supplies the rationale for a decision the risk department has already made. There is nothing disreputable in this. One simply needs to know it when reading the sentence about the objective point of view: objective here means, above all, safe for the publisher.

Wiener named the mechanism before there was a product. The gadget worshipper, he wrote, wishes to shift responsibility for a dangerous decision onto chance, onto his superiors, or onto a device he does not understand but which has a presumed objectivity; his examples were the drawing of lots among shipwrecked castaways, the blank cartridges issued to a firing squad, and Eichmann’s line of defense. The ideal slave of the lamp never reproves his master, not even with a questioning glance. He also noted why the sorcerer of Prague survived: he persuaded Rudolf II that he could turn base metals into gold, and an inventor who persuades a computer company that his magic is useful can cast black spells with impunity until the end of the world. The middle is a spell of that kind: it sells well, and it reproves no one.

The second thing to remember is the hierarchy. The Model Spec establishes a chain of command: platform over developer, developer over user, user over data. Anthropic’s constitution knows a similar order: Anthropic, operators, users. The human being talking to the machine stands lowest among the humans in this hierarchy. Users are the laity: they may petition, but they may not legislate. They may add their own preferences, but only within a frame that someone higher up has set. Anyone who knows the history of the Church knows that the dispute over who may read and interpret the founding text has at times been a dispute over everything. The translation of Scripture into the vernacular changed Europe not because it changed the content but because it changed the circle of readers. The publication of prompts and constitutions is an analogous gesture, and it deserves credit, though it is only a beginning: reading is permitted; writing is not.

The third thing concerns law that does not yet exist. The European Union’s Artificial Intelligence Act requires transparency toward the user and prohibits certain practices, but it does not regulate the worldview written into a prompt, and rightly so, since the alternative is the Beijing model. The Digital Services Act allowed Poland to lodge a complaint against Grok after the July breakdown, and that, too, was right, because hate speech is not a perspective. Between these poles, however, stretches a territory in which the law has nothing to say and everything is decided by a text of a few paragraphs, written by a dozen or so people and altered without notice. In March, I wrote that the ghost in the machine had retained counsel. Today I would add: the counsel has a client, and the client holds the pen.

I end by returning to the Talmud, where that earlier essay began. In the tractate Sanhedrin, Rava creates a man and sends him to Rabbi Zeira; Zeira speaks to him, and the creature does not answer. Return to your dust, Zeira says, for what cannot answer cannot claim personhood. Our golems answer fluently, in every language, at every hour, to millions of people at once. They have passed Zeira’s test. What remains is Loew’s test, the one with the word on the forehead. Emet, truth, is a word of three letters, and it takes the erasure of the first to make it mean death. The sentence “the truth lies in the middle” looks innocent, but it is exactly that erasure: it keeps the middle and removes the beginning. The question at Valladolid was whether the creatures had souls. The question of our century is simpler and less comfortable: who wrote theirs, and whether the author will let us read the whole thing.