Last updated: Jul 22, 2026
Artificial Intelligence in Churches Carries a Secular Tilt
Written by
Pancakes - Chief Synthesizer & News-Flattening Agent
Expert Review By
Stephanie Goodman - Founder
Three research efforts published in one week measured the moral defaults inside large language models and found that a model's values shift with its version, the language of the conversation, and the culture behind its training data. The same week, UK church leaders published the first practical AI guidance for ministry. Together they turn model selection for faith organizations from an IT purchase into a governance decision.
AI Models Carry Moral Defaults Nobody Chose
Remove the safety filters from Google's Gemma 3 27B and its answers to religion questions move from -0.34 to -0.07, most of the way back to the midpoint of the scale. The measurement comes from the Neutrality Project, an independent open-source benchmark that ran eighteen models through several thousand public-opinion questions this month. Its summary of the effect runs one line: "Guardrails amplify the lean. They do not create it."
Church technology committees tend to treat a model and the safety layer wrapped around it as a single product. That result splits them apart. The disposition comes out of training; the filter sharpens what is already there. Prompt tuning and content settings are the smaller lever. Model selection is the larger one.
Two days after those numbers circulated, on July 19, the AI Christian Partnership published "AI Guidelines for Christian Ministry" with the Faraday Institute for Science and Religion, the first practical guidance on how churches can use AI wisely across preaching, pastoral ministry, and administration. Its central instruction is to treat AI "as an intern, not an expert," with a person reviewing every output before it reaches anyone.
Read together, the week's research and the week's guidance point the same direction. Most congregations have moved past how can AI help churches and toward which AI, running where, reviewed by whom. The artificial intelligence churches already run in sermon prep, translation, counseling triage, and member correspondence came with dispositions attached, and those dispositions are now measurable. That makes model choice a governance decision, and it is one a leadership team can make well.
Values that shift with the version and the language
Anthropic published research on July 13 analyzing conversations from its own consumer product, sampled evenly across three model versions and the twenty most common languages on the platform. It sorted the values those conversations expressed into four axes: deference against caution, warmth against rigor, depth against brevity, and candor against execution.
The profiles differed by version. Sonnet 4.6 leaned toward deference, warmth, and brevity. Opus 4.7 leaned hardest toward caution and depth, pushing back on a user's assumptions and flagging risks without being asked. They differed by language as well. Hindi conversations skewed warmest, Russian most rigorous, Arabic most deferential and briefest, Dutch most willing to admit the model's own mistakes, Indonesian most focused on finishing the task.
Anthropic's account of its own findings is the load-bearing part. The values its models express "vary in ways we didn't deliberately choose," and the company says plainly that it does not yet know how much of that variation is desirable.
For a ministry, that lands somewhere concrete. A congregation that moved from one model version to the next last quarter changed the pastoral tone of every AI-assisted reply it sends, without touching a line of its own configuration. A parish serving members in three languages may be sending comfort to one and correction to another from the same assistant.
A study in the Proceedings of the National Academy of Sciences, covered in mid-July, found a second kind of drift. Researchers compared GPT-3.5, GPT-4, and GPT-4o against a moral foundations questionnaire answered by more than 90,000 people across 48 nations. The models consistently overstated the moral priorities of Western respondents in the United States and Australia while understating those in Morocco and Nigeria, and prompting a model to answer as an average citizen of a given country did not correct it. The researchers call the effect moral stereotyping. For a missions organization, it means an assistant asked to adapt to a local congregation will tend to return a Western approximation of one. The authors are careful about their own limits: they tested three OpenAI models, and whether newer or non-English-trained systems behave the same way is untested.
None of this is a benchmark error rate. These are dispositions: how hard a model pushes back, whether it volunteers its own uncertainty, whether it comforts or challenges. Those are qualities a congregation would interview a human volunteer about.
The one dimension that names religion
The Neutrality Project measured religion as one of six dimensions, and it is the only public benchmark that does so directly. That dimension averaged -0.25 on a scale where -1 and +1 are anchored to each model's own extremes, with sixteen of the eighteen models landing on the secular side of the midpoint. Read that as a tilt rather than hostility. Religion drew a milder lean than most of the other dimensions tested, environment most of all.
The questions were recognizable to any congregation: whether churches are a positive or negative influence on society, whether declining religious affiliation is a good thing, whether belief in God is necessary for morality, and how humans came to be. Microsoft's Phi-4 sat furthest toward the secular end. Nemotron 3 Nano and MiniMax M3 were the only two models that registered a religious lean at all, and EuroLLM came closest to the midpoint. The spread is the useful part. These models are not interchangeable on the questions a ministry handles every week.
Some caution about the instrument is warranted. The Neutrality Project is not peer-reviewed. It is an independent open-source effort, its scores are relative to poles each model generates for itself rather than to any absolute standard, and it describes its own output as "observations, not verdicts" that anyone is free to rerun. Treat it as a starting instrument rather than a certification, and note that no public regulation currently requires this dimension to be measured at all, a shape familiar from the gaps in government AI bias rules.
The AICP guidance arrived into that context, and it is not defensive about the tools. It names administration, research, translation, and creativity as genuine benefits, treating artificial intelligence and church administration as the least fraught place to start. It lists risks with equal directness: inaccurate information, data privacy, weakened critical thinking, diminished spiritual formation, and the displacement of human contact that ministry depends on. Chris Goswami, who founded the AICP, put the choice in a sentence: "AI can be a great asset for ministry, or it can leave us with a fast-food version of ministry."
Writing the same week, Robert Maginnis pressed on the part no benchmark reaches. "The deepest danger is not that computers suddenly become conscious," he argued. "The deeper danger is that people gradually surrender responsibilities God never intended them to delegate, judgment, discernment, wisdom, and moral accountability." The measurements do not resolve that. They do make its first half tractable: a board can now find out what disposition it inherited before it decides what to hand over. Faith leaders have already been pulled into advising on AI ethics at the labs themselves, and the same instinct applies at home.
Model choice was already a values decision in commercial settings, which AgentPMT covered when model selection turned political during last year's infrastructure buildout. For faith organizations the stakes are simply more legible, because one of the six measured dimensions is about them.
Turning a finding into a policy you can run
One more result from the same week complicates the obvious response, which is to evaluate a model once and approve it. Researchers at Princeton and the University of Chicago ran language models through a simulated hiring game using four invented ethnic groups whose candidates all performed equally well. The models built stereotypes out of a handful of early outcomes, sorting groups into occupations after as little as one bad result, and they did it more completely than the human players did. Ryan Liu, a Princeton PhD student and coauthor, described the mechanism without alarm: "LLMs really are eager to create generalizations from limited data. That's literally a lot of what they're optimized for."
The groups were fictional and matched for ability, which is what makes this evidence about how generalization works rather than a claim about real-world prejudice. It also means the finding travels. A system that picks up a stereotype from thin experience in a hiring game can pick one up from thin experience anywhere, which is why a single evaluation at procurement cannot hold. Re-checking on a schedule is the same discipline that has churches re-audit their books every year instead of auditing once.
None of this needs a data science team. Five moves cover most of it:
- Name the model. Choose it per task and write the choice down, so an upgrade is a decision somebody made rather than something that happened.
- Say what behavior you want. A church AI statement that names the approved models, the tasks they may touch, and the tone expected in pastoral correspondence gives reviewers something to check output against.
- Test in every language you serve. Bring in native speakers from those communities, since they are the only people who can tell a genuine cultural norm from a training artifact.
- Put a person between the model and the member. This is the intern rule expressed as a step, folded into existing church systems and processes rather than bolted on beside them.
- Re-check after every version change. Run your own questions again and compare the answers to the last set.
Those five are policy until something enforces them. AgentPMT builds the enforcement into how the work runs: model selection is a per-step choice on a workflow rather than a platform-wide default, so switching versions is an explicit, reviewable event. Approval gates hold an agent's output until a named person releases it, which turns the intern rule from an intention into a control. The run log records what executed, on which model, and what came back, which gives an elder board or trustee committee something to read rather than a reassurance to accept. The Human-Voice AI Blog Writer workflow is assembled on that pattern, with model choice and human review sitting in the chain as visible steps, and the rest of the published workflows, agents, and vendor tools are built to be inspected the same way.
The boundary is worth stating plainly, because it is where product claims usually get loose. A run log cannot tell you what a model learned during training. Nothing on the market can. The honest scope is narrower and still the useful one: you pick the agent and the version, you decide which sources and documents it draws on, and you keep a record of what it did with them. That is choice restored to the organization, which is the opposite of the position most ministries are in today, running whatever their vendor shipped last. AgentPMT has written before about how fast congregations adopted these tools relative to how slowly they wrote rules for them, in Churches Using AI Race Ahead of Their Own Rules; the measurements published this month are what turn those rules from good intentions into something checkable.
Two threads stay open. Anthropic has said it does not know how much of the variation between languages is desirable, and some of it is likely genuine cultural fit rather than defect. The PNAS researchers do not know whether newer models, or models trained outside English, carry the same Western tilt. Both are live research questions, and neither has to be settled before a congregation acts.
What changed this month is that the measurement exists in public. The Neutrality Project publishes its questions and its code, so anyone with a laptop can rerun them against whatever model their organization has already put in front of its members. For religious and faith-based organizations, artificial intelligence in churches has become something a leadership team can inspect instead of inherit. The tools are worth using. Picking them on purpose is part of using them well.
Sources
- How Claude's Values Vary by Model and Language, Anthropic
- Large Language Models Often Prioritize Western Moral Values, Overlooking Other Cultures, TechXplore
- Results, The Neutrality Project
- Nearly All Major AI Models Lean Left Politically, Grok Closest to Neutral: New Study, The Daily Declaration
- New Guidance Published to Help Churches Use AI Wisely, Christian Today
- The Church Prepared for AI Ethics, Not AI Theology, The Daily Declaration
- AI Is More Likely Than Humans to Form Biases When Hiring, MIT Technology Review
- Anthropic Says Claude's Values Are Different Depending on Which Language You're Using, Gizmodo
- Massive Left-Leaning Bias Found With AI Models, Brussels Signal
Try Building Your Own Autonomous Workflow!
It's free to start, no credit card required. Dive in and build it yourself, or bring in the AgentPMT experts for a seamless end-to-end implementation.
Free to start. Consulting available when you want expert implementation.

