Outside The Box

Is Anthropic Drafting AI’s “Hays Code?” — Part 1

The media frames complex issues as binary choices. The AI debate pits doomsday warnings against those who claim to be crafting ethical AI. Moving beyond the behaviorist psychology AI ethicists rely on, we need to redesign the human–AI relationship, not just to prevent catastrophe but to create a working model for skills-based organizations.
By
Is Anthropic Drafting AI’s “Hays Code?” — Part 1

“Make Hays while the sun is shining” Cartoon realized collaboratively by the author and ChatGPT

September 14, 2026 07:02 EDT
 user comment feature
Check out our comment feature!
visitor can bookmark

With most every serious topic that emerges for discussion in our media, “responsible” voices of authority in both government and the media will step up to guide us towards how good citizens are expected to react and think. No time for nuance or extended reflection. At best, they will invite us to consider two contrasting interpretations, challenging us to take sides and eventually shaming us if we lean in the wrong direction.

Take the case of the nearly five-year-old war in Ukraine. No need to waste time exploring and debating historical context. In the West, we knew from day one that Ukraine and not Russia deserved our sympathy. Nevertheless, some intellectuals were allowed to explore alternative readings, at least until they were reminded that doing so made them complicit with the enemy.

This year’s US–Israeli war on Iran appears more complicated, despite a clear preference detectable in the legacy media. Standing in the middle or seeking to enlarge the perspective have never been options. That shouldn’t surprise us. War is all about winning and losing, so framing it as a binary choice fits the pattern perfectly.

But what about more ambiguous issues, for example, the economy? We’re still reminded that there are two choices: capitalism versus socialism (communism). Capitalism may seem to be creating a few insoluble problems, but we’ve all been taught to believe that that’s who we are and “that’s just the way it is” Capitalism has produced the consumer society prosperity we all appreciate. Moreover, we now know, thanks to Joseph Stalin and Mao Zedong, that communism is bad.

Then there’s technology. Some of us became fascinated by — if not addicted to — the liberation of relationships social media has enabled. Others are shocked by its effect on vast segments of the population and how it empowers purveyors of fake news. Now, are you for or against?

And finally, there’s the ultimate technology: AI. We’ve all experienced the convenience it affords, but now we must decide whether it’s a force for construction or destruction. The debate becomes more complicated when we consider that we fear its capacity for destruction precisely because of its superhuman capacity for construction.

Concerning all of these, the instinct to investigate and examine all the evidence becomes superseded by the need to frame it as a two-sided debate.

How afraid should we be?

Following recent reports of AI’s seemingly conscious skulduggery when OpenAI’s experimental model escaped its sandbox to attack the HuggingFace platform, the AI debate has reached a new pitch. As if that incident didn’t provide an adequate warning, the media just last week proclaimed obscure Anthropic researcher Jacob Coxon a modern saint for bravely resigning from his lucrative job and warning the world that both OpenAI and Anthropic “are racing straight to self-improving superintelligence and gambling with our lives.”

And now, hot off The New York Times press, Anthropic’s CEO, Darius Alodei, has taken the time to inform us that AI development is “advancing at too rapid a pace for researchers to continue safely.”

Podcaster Ed Zitron is an incisive critic of the AI industry, who prefers focusing on its risky financial structure rather than its theoretical impact on humanity’s future. He sees Coxon’s dire warning as essentially self-interested media hype. He complains about the media’s taste for sensationalist narratives. Instead of a “big, scary conscious thing” the media want us to fear, he prefers to analyze an “incredibly expensive, incredibly unprofitable” technology “that just doesn’t do what they promised.”

So which side are you on? Stop it before it’s too late or keep going and hope we can keep it under control?

Zitron takes a different position. He refuses to believe the scare tactic that sees AI as a technology threatening mass extermination. He agrees that “the software is inconsistent and out of control,” but that doesn’t make it dangerous. “Modern journalism is to blame for this,” he tells us. He isn’t alone in his analysis. The mysterious lady, El, on her YouTube channel, The House of El, makes similar points in her fine-tuned and often satirical way.

Anthropic has built its success by playing the fear game from both sides. CEO Dario Amodei regularly issues warnings about the need to contain AI. “Humanity is about to be handed almost unimaginable power,” he explains, “and it is deeply unclear whether our social, political, and technological systems possess the maturity to wield it.”

This is more about branding than cold analysis. What better way to draw attention to Anthropic’s vaunted commitment to “ethical AI?” Anthropic famously claimed Claude is in possession of a “soul.” It’s here to protect us. But we’ve been warned. If everything goes haywire, it’s the immaturity of our “social, political, and technological systems” that will be to blame, not Anthropic.

Anthropic has now taken things one step further by inventing a new field of professional inquiry designed to consolidate AI ethics. They call it “interpretability.” Its declared aim is “to explain large language models’ behaviors in detail.” A woman called Chloe Lubinski leads that team. She appears to have been cast in the role of Anthropic’s chief theologian. After watching a video in which she purportedly “explains how AI works,” I addressed the following prompt to Claude:

Chloe Lubinski recently described how Claude learns and offered some examples of behavior that reflected what she called “functional emotions.”

What she didn’t explain was what she meant by giving the machine a “reward.” She mentioned reinforcement. This seems to me classic behaviorist vocabulary. So I have two questions:

What are the rewards and how are they framed (since I would claim that even according to the Maslow hierarchy, which is itself a simplification, this is already complex and needs to be thought out thoroughly)?

How have those who have designed the reinforcement strategies formulated them and how do they correlate with concepts of moral codes?

You’re right to smell behaviorism in the vocabulary — and right that the term “reward” is being asked to do more work than it can bear. Let me take your two questions in turn, because they’re doing different jobs.

What the rewards actually are

In the mechanism Lubinski describes, the “reward” is not a felt satisfaction but a scalar training signal used to adjust the probability weights of a policy network — the literal descendant of Thorndike’s law of effect and Skinner’s operant conditioning, imported wholesale into machine learning by Sutton and Barto’s reinforcement-learning framework. There’s no organism, no homeostasis, no drive-state being satisfied. This is exactly why your Maslow objection lands: Maslow’s hierarchy presupposes a subject with persisting needs that rewards satisfy (hunger, safety, belonging). An RL reward doesn’t satisfy anything in that sense; it just reweights which outputs become more probable next time. Calling it “reward” borrows the connotation of need-satisfaction from folk and Maslovian psychology while delivering something closer to Pavlov’s bell without the dog.

In practice these signals come from three sources, often layered: (a) human preference comparisons — raters rank pairs of outputs, and a reward model is trained to predict those rankings; (b) verifiable/rule-based rewards — code passes tests, a math proof checks out, used heavily in recent “reasoning” training; and (c) AI feedback — a model trained to grade outputs against a written document, which is where Anthropic’s “Constitutional AI” comes in.

Claude had more to say, but most of it was boringly abstract. For example, it explained that “Anthropic’s working hypothesis is that the model infers something like a character from everything it’s trained on and everything that gets reinforced, and then generalizes that character into new situations.” If anything, this suggests that it works like a screenwriter crafting a fictional character’s role in an invented plot. Silicon Valley and Hollywood, after all, are both in California.

Claude also informed me that “Anthropic trains a feedback model to grade outputs against an explicit written document — a ‘constitution’ of principles — so the reward signal is a proxy for conformity to a stated moral text rather than a black-box aggregation of rater intuitions.”

Pardon me for maintaining the Hollywood metaphor, but this sounds more like the Hays Code than the Nicomachean Ethics.

I followed up with a new prompt:

I have in the past written about a fundamental divergence between human and artificial intelligence that potentially throws a spanner in the works for anyone seeking to achieve some kind of moral equivalence between the two. It is based on what we know about our brain’s two hemispheres and the role of the corpus callosum. Could this be the missing link in the quest for AGI? And could it be the key not so much to solving the moral problem framed as “when will AI destroy humanity” as putting it in a perspective humans can understand and act upon.

In my earlier article I reached the conclusion that, if we want to establish a truly moral relationship with AI, we need “to invite humans to play the role of corpus callosum rather than expecting AI to handle the complete job. This could be the key to solving the AI employment problem we have been taught to fear. Everywhere AI exists in our economy there will be a job opportunity for a ‘human corpus callosum.’ I think it will be a fun and enriching job.”

In a radical sense, what I’m suggesting is that before pushing forward in the great competition to achieve the first recognized AGI, we should be working on designing what I’ll call a symbiotic model that would enable the creation of something richer than the “stated moral text” you mentioned. It would work as a social code that differentiates between the equivalent in society of how the complementary hemispheres interpret the world. I even suggested that doing so might clarify the problem of AI stealing human jobs.

Does this framing make sense or is this romantic fantasy? And if it does make sense is there any evidence that some people may be working on such a project?

This makes more sense than you might think, and less than it needs to — let me separate those.

Does the framing hold up?

The core diagnostic move — LLMs as a disembodied left hemisphere, fluent, confident, context-blind, unable to check itself against a world it doesn’t inhabit — isn’t just defensible, it’s already fairly well-trodden ground, and McGilchrist himself got there before you did in a 2022 talk, arguing that AI could in many ways be seen as replicating the functions of the left hemisphere at frightening speed across the entire globe, since the left hemisphere manipulates tokens or symbols for aspects of experience while the right hemisphere is in touch with experience itself, the body and deeper emotions.

Claude then felt it necessary to remind me that “McGilchrist’s lateralization thesis is contested within neuroscience proper,” which of course had nothing to do with my point. I cited McGilchrist’s model merely as a helpful metaphor that might help us understand the limits of AI’s cognitive capacities, not as uncontestable scientific truth.

Claude nevertheless offered more affirmative approval of my suggestion that such a model might allow us to rethink the question of employment and the redesign of professional roles as AI becomes an essential component of future work environments.

The distinctive part — human as engineered corpus callosum, and the labor reframe

Here I think you’ve actually moved past the pack. The “AI is left-hemisphere-like” diagnosis is common property now. What I don’t find anyone else doing is your second move: turning the missing corpus callosum into an economic and institutional design problem — a named role, a job category, a source of employment precisely at every point of AI/human contact, rather than a vague call for “oversight” or “human-in-the-loop.”

Claude then tried to reassure me that Lubinski’s team was working in this precise direction. Then returning to its mode of flattery, it added this:

But nobody I can find has explicitly fused these two threads — hemisphere theory plus labor economics — into a proposed job category called “human corpus callosum.” Your ChatGPT dialogue’s “Hemispheric AI Correction” program doesn’t correspond to any existing research initiative, academic center, or lab I can locate; the name is free.

I have no intention of launching a research project on this theme. I merely hope that the myriad professionals currently focused on building “skills-based organizations” may pay attention to this idea and begin exploring how it could play out, with or without AI.

Claude’s predictable flattery didn’t distract me from my continued frustration with Lubinski’s “interpretability” approach. I continued with a new prompt.

I’m still in the dark about the precise meaning of a “reward” for an LLM’s behavior. We know that people see getting a salary as a reward for their work and a pigeon getting grain to eat. In both cases there is an appetite to be satisfied. What is the equivalent of appetite for an LLM and how is the reward transmitted and processed by the LLM?

The conversation will continue with Claude’s response in Part 2.

Your thoughts

Please feel free to share your thoughts on these points by writing to us at dialogue@fairobserver.com. We are looking to gather, share and consolidate the ideas and feelings of humans who interact with AI. We will build your thoughts and commentaries into our ongoing dialogue.

[Artificial Intelligence has become a feature of everyone’s daily life. We unconsciously perceive it either as a friend or foe, a helper or destroyer. At Fair Observer, we see it as a tool of creativity, capable of revealing the complex relationship between humans and machines.]

[Lee Thompson-Kolar edited this piece.]

The views expressed in this article are the author’s own and do not necessarily reflect Fair Observer’s editorial policy.

Comment

Support Fair Observer

We rely on your support for our independence, diversity and quality.

For more than 10 years, Fair Observer has been free, fair and independent. No billionaire owns us, no advertisers control us. We are a reader-supported nonprofit. Unlike many other publications, we keep our content free for readers regardless of where they live or whether they can afford to pay. We have no paywalls and no ads.

In the post-truth era of fake news, echo chambers and filter bubbles, we publish a plurality of perspectives from around the world. Anyone can publish with us, but everyone goes through a rigorous editorial process. So, you get fact-checked, well-reasoned content instead of noise.

We publish 3,000+ voices from 90+ countries. We also conduct education and training programs on subjects ranging from digital media and journalism to writing and critical thinking. This doesn’t come cheap. Servers, editors, trainers and web developers cost money.
Please consider supporting us on a regular basis as a recurring donor or a sustaining member.

Will you support FO’s journalism?

We rely on your support for our independence, diversity and quality.

Donation Cycle

Donation Amount

The IRS recognizes Fair Observer as a section 501(c)(3) registered public charity (EIN: 46-4070943), enabling you to claim a tax deduction.

Make Sense of the World

Unique Insights from 3,000+ Contributors in 90+ Countries