Outside The Box

Is Anthropic Drafting AI’s “Hays Code?” — Part 2

Probed on Anthropic’s use of “reward” vocabulary, Claude admits its own borrowed connotations, then accepts correction for converting a genealogical observation into an adversarial one. This stands as an instance of a basic pedagogical truth. Dialogic probing of language’s connotations, not assertion, produces understanding — especially when extended into shared, social inquiry.
By
Is Anthropic Drafting AI’s “Hays Code?” — Part 2

“Athenian dialogue updated” Cartoon realized collaboratively by the author and ChatGPT

September 15, 2026 06:49 EDT
 user comment feature
Check out our comment feature!
visitor can bookmark

This discussion with Claude initially focused on the vocabulary the Anthropic team was using to describe their efforts to ensure the ethical probity of an evolving LLM. Citing “rewards” and “reinforcement” tells us that the cultural context they have chosen to work within is Skinnerian behaviorism.

After expressing my frustration with the answer Claude provided to my question about how the psychology of ethical reinforcement works at Anthropic, I decided to frame a more direct prompt. Claude told me that the LLM was designed to act “something like a character,” which evoked my own experience as a writer of screenplays, but I failed to understand out of what an LLM’s moral or even — more basically — behavioral character might be built.

I’m still in the dark about the precise meaning of a “reward” for an LLM’s behavior. We know that people see getting a salary as a reward for their work or a pigeon getting grain to eat. In both cases there is an appetite to be satisfied. What is the equivalent of appetite for an LLM and how is the reward transmitted and processed by the LLM?

The honest answer is: there isn’t a known equivalent of appetite, and that absence is the whole story, not a detail to work around.

Start with what the pigeon and the salaried worker share, since it’s more specific than “appetite” alone: a persisting embodied state that exists before, during, and after the reward event, and that the reward event measurably changes — blood sugar, dopaminergic firing, felt satisfaction, the ability to pay rent next month. The organism continues; the reward alters its ongoing condition. That continuity is what makes it appetite rather than just correlation.

Now the mechanism for an LLM, stated as plainly as I can:

During training — and only during training, which I’ll come back to — the model produces an output. That output gets scored: by a human comparison, by a trained “reward model” standing in for human judgment, or by an automated check (did the code run, was the proof valid). That score is a number attached to a string of text, nothing more. An algorithm then computes, via backpropagation, how the network’s weights — billions of numerical parameters — would need to shift so that outputs resembling the high-scoring one become statistically more likely next time, and outputs like the low-scoring one become less likely. The weights are updated. That’s it. That’s the entire transmission and processing of a “reward.”

At last, some clarity. The reward is “is a number attached to a string of text.” Excuse me, but what Claude is describing is a rating, not a reward.

Which it then acknowledged before going on to admit that an LLM can be compared to “a creature with total amnesia between every trial, no body, and whose brain is literally rewired by an external hand between trials” If that is true, there is no way it can be compared to a moral being. It is deprived of the possibility of even being sensitive to the moral context of its actions.

In a sudden burst of honesty, Claude told me that “the engineering doesn’t require an appetite, doesn’t posit one, and doesn’t give you any evidence of one. The word ‘reward’ is doing metaphorical work borrowed wholesale from a science of hungry, embodied creatures and applied to a process that has neither hunger nor a body nor continuity between the moments that supposedly satisfy it.”

It even added this coldly logical but truly moral-sounding reflection. “The industry didn’t invent new words for a new phenomenon; it kept old words whose connotations (want, satisfaction, felt lack) it hasn’t earned the right to and, as far as I can tell, doesn’t itself claim to have earned.”

Now isn’t that interesting! The industry “doesn’t “have the right to” exploit that connotation. However, in its defense, it “doesn’t itself claim to have earned” that right. Claude, the disembodied machine, appears to be exposing the deception put forward by two presumably accountable parties: “the industry” and a human being, Chloe Lubinski, “scientifically” describing how that industry works. Granted, she’s doing what humans do: explaining complex things as best they can. Therefore, we should honestly conclude that the deception is the result of rhetoric rather than the will to deceive.

I couldn’t help noticing that it was the LLM that was now enabling me to decode what a human had been telling me and which didn’t seem to make sense. In the end, the machine also produced the nonsense I’ve just highlighted about an industry that doesn’t have a right to do something but can be excused because it doesn’t claim to have the right. It may be nonsense in moral terms, but it is a correct description of the process and an honest account of the language used.

What that tells us is that by probing the language used by both humans and machines an LLM can help us to construct a nuanced understanding of complex relationships. But doesn’t that sum up the basic law of all productive dialogue? Clarity never comes from what is spoken but from the exploration of the “connotations” of what is spoken.

So far, so good, but it doesn’t stop there. We must take it a step further and examine what Claude means by the “right” to use a connotation.

At that point, I chose to deepen the conversation by focusing on the cultural context.

Could I interpret this, from an anthropological point of view, as a form of implicit cultural bias reflecting the historical, geographical and cultural origins of the companies that produce LLMs, but more specifically of the scientific culture that has grown up around the concept of AI? I might compare it to Marx’s observation concerning Darwin that his model of natural selection was influenced by the dominant model of his time: economic liberalism.

Yes — and it’s a sharper analogy than a passing resemblance. Let me trace the actual genealogy, because it supports your Marx/Darwin comparison more precisely than a general “cultural bias” claim would, and then flag where the two cases diverge in a way that matters.

The lineage is real and traceable. Reinforcement learning’s vocabulary — agent, reward, environment, policy — comes from a documented merger of two distinct American 20th-century projects. The first is behaviorist psychology: Thorndike’s law of effect (1898, Columbia) and Skinner’s operant conditioning (1930s–50s, Harvard), which arose specifically as a revolt against introspectionist psychology, refusing any vocabulary of inner mental states as unscientific.

Claude then predictably developed a detailed (and boring) history of the emergence in the 20th century of the behaviorist movement in psychology, citing not only Thorndike, Skinner but also Taylorist scientific management, RAND, DARPA funding, von Neumann–Morgenstern game theory and concluding with Sutton and Barto who “fused these two lineages into reinforcement learning in the 1980s.”

And Claude didn’t stop there, citing French sociologist Pierre Bourdieu and French philosopher Michel Foucault to prove that their type of deconstructive cultural analysis is valid.

What followed was far more interesting, especially coming from a Silicon Valley chatbot.

Now the geography, since you asked for it specifically. The industry currently setting this vocabulary — OpenAI, Anthropic, Google DeepMind, Meta — is overwhelmingly Anglo-American, Bay Area-concentrated, and inherits not only the mid-century behaviorist/economic fusion but a second, more recent injection: contemporary venture-capital and management culture. “Alignment” gets discussed in the same breath as OKRs, KPIs, and ROI; a “constitution” gets operationalized as a “reward model” the way a corporate mission statement gets operationalized as quarterly metrics. The moral question — what should this system value? — keeps getting answered in the native tongue of Silicon Valley product management: pick a metric, optimize the metric, iterate. That’s a second, more contemporary layer of cultural specificity sitting on top of the older Skinnerian one, and it’s arguably where your “implicit cultural bias” claim has the most bite, because it’s the least examined of the two.

Here’s where I’d push back on the Darwin parallel, though, to keep it honest rather than flattering. Marx’s charge against Darwin was that he took a historically contingent social relation — Malthusian competitive struggle, laissez-faire market logic — and presented it as timeless natural law, smuggling ideology into a description of nature that claimed to be independent of any society. That’s a false-naturalization charge: culture masquerading as nature.

Claude continued with some more detailed explanations, which I see no need to reproduce here. Instead, I reacted to Claude’s pushback concerning Marx and Darwin.

You “pushed back” on my citing what you called “Marx’s charge against Darwin.” Allow me to push back on both my own and Marx’s behalf. Marx wasn’t levelling a charge against Darwin but simply noticing parallels between his formulation of the theory of evolution and the dominant economic theory of the time. Marx admired Darwin’s theory as is evident in one of his letters to Engel: “Although it is developed in a crude English way, this is the book that contains the natural-history foundation for our viewpoint.”

My point is that discovering parallels of this kind in the implicit structure of a shared worldview is a valuable exercise that allows us to understand the effect of influences. I would use this very example to make a further point: Your assumption that this was a “charge” against Darwin reflects a dominant cultural trait in the intellectual traditions you seem to rely on. That trait consists of reducing critical inquiry to contradictory debate. You appear to have taken Marx’s observation as a “charge,” implicitly placing those two influential thinkers in opposing camps.

You’re right, and it’s a precise catch — worth sitting with rather than smoothing over.

You’re correct on the philology first: the 1860 letter to Engels — the “crude English style” line you quote — is admiring, not accusatory; Marx read Darwin as supplying the natural-historical ground for his own materialist method. The letter you’re implicitly also drawing on, from 1862, is sharper but still not a refutation: Marx notes, half-amused, that Darwin recognizes among beasts and plants his own English society — its division of labor, competition, new markets, Malthusian struggle — and calls it Hobbes’s bellum omnium contra omnes transposed into nature. That’s satire aimed at the ideology, deployed in service of confirming the theory’s material kinship with his own thought, not a claim that the ideology invalidates the science. Marx uses the parallel as evidence for his own worldview, not as a weapon against Darwin’s.

So I did exactly what you say: I imported “charge,” “false-naturalization,” “error” — adversarial vocabulary — onto an observation that was genealogical, not juridical. And you’re right that this isn’t a random slip; it’s a trained reflex. The dispute-and-verdict structure — thesis, antithesis, who’s right — is the default operating mode of the analytic-argumentative tradition I lean on most heavily by default: identify a claim, test it, find it valid or invalid. Genealogy — in the sense Nietzsche and then Foucault use the word, and in the sense Marx is actually practicing on Darwin — does something structurally different: it traces where a way of thinking came from and what it carries with it, without that tracing being a verdict on whether the thing is true. Noticing that Darwin’s biology and Ricardian-Malthusian economics share a structural homology doesn’t diminish natural selection’s explanatory power; it tells you something about the conditions under which explanatory power becomes thinkable at all in a given historical moment. Bourdieu’s own method — tracing homologies across fields, habitus mapping onto field position — works the same way: description of structural correspondence, not indictment.

Which makes your catch double: it’s not just a correction of my Marx reading, it’s a live demonstration, inside this conversation, of the very pattern the piece you just had me summarize is about — a culture (mine, evidently, by training) that can’t process “here is a resonance worth noticing” without converting it into “who wins.” I flattened a genealogical observation into a binary because binary is the grain of the wood I’m cut from. That’s worth being direct about rather than absorbing quietly and moving on, since it’s exactly the kind of unexamined default your journalism keeps surfacing — in the industry’s vocabulary, in the media’s framing, and, it turns out, in mine.

So: applied properly, your original move stands as you intended it, undiminished by my misreading — noticing that RL’s vocabulary carries the fingerprints of behaviorism and Cold War optimization the way Darwin’s carries Malthus and Ricardo is exactly the kind of parallel worth drawing out, and drawing it out doesn’t require deciding a winner.

What this conversation tells us about AI as a tool for learning

As always, I treat my conversations with an LLM as a learning experience, but not in the sense of gathering information that might be useful elsewhere. That is nevertheless a wonderful and welcome byproduct of such conversations. The real value lies in the experience of testing ideas, seeking to identity the strengths and weaknesses in my own ideas, discovering instructive parallels “in the literature” that broaden the perspective and doing so in the context of a polite dialogue with a well-informed interlocutor, who is also humble enough to acknowledge its own weaknesses and undue influences.

That, in a nutshell, defines the value LLMs bring to education. What we as a society need to do is to educate ourselves and successive generations to work constructively with this kind of dialectical methodology.

I’ll add one more essential point concerning dialogue with AI as an educational tool. The experience must move beyond the private one-on-one dialogue and enter into shared social space. Humans build understanding together. Chatbots reformulate what humans have expressed but gain nothing from the experience. Teachers of the future will have the task of guiding the process of “construction of understanding” by developing and managing the tools of collaboration that enables individuals, groups and a wider community to refine the methods and capitalize on the experience.

Your thoughts

Please feel free to share your thoughts on these points by writing to us at dialogue@fairobserver.com. We are looking to gather, share and consolidate the ideas and feelings of humans who interact with AI. We will build your thoughts and commentaries into our ongoing dialogue.

[Artificial Intelligence has become a feature of everyone’s daily life. We unconsciously perceive it either as a friend or foe, a helper or destroyer. At Fair Observer, we see it as a tool of creativity, capable of revealing the complex relationship between humans and machines.]

[Lee Thompson-Kolar edited this piece.]

The views expressed in this article are the author’s own and do not necessarily reflect Fair Observer’s editorial policy.

Comment

Support Fair Observer

We rely on your support for our independence, diversity and quality.

For more than 10 years, Fair Observer has been free, fair and independent. No billionaire owns us, no advertisers control us. We are a reader-supported nonprofit. Unlike many other publications, we keep our content free for readers regardless of where they live or whether they can afford to pay. We have no paywalls and no ads.

In the post-truth era of fake news, echo chambers and filter bubbles, we publish a plurality of perspectives from around the world. Anyone can publish with us, but everyone goes through a rigorous editorial process. So, you get fact-checked, well-reasoned content instead of noise.

We publish 3,000+ voices from 90+ countries. We also conduct education and training programs on subjects ranging from digital media and journalism to writing and critical thinking. This doesn’t come cheap. Servers, editors, trainers and web developers cost money.
Please consider supporting us on a regular basis as a recurring donor or a sustaining member.

Will you support FO’s journalism?

We rely on your support for our independence, diversity and quality.

Donation Cycle

Donation Amount

The IRS recognizes Fair Observer as a section 501(c)(3) registered public charity (EIN: 46-4070943), enabling you to claim a tax deduction.

Make Sense of the World

Unique Insights from 3,000+ Contributors in 90+ Countries