# Churches Using AI Now Have Benchmarks to Test Models

> A faith-specific benchmark scored 36 frontier models, four universities published an omission and conversion-bias test set, China's AI companionship rules took effect, and Pope Leo XIV drew a line at delegated judgment. The wider cycle in AI for faith organizations, beyond the model-values research the feature article covers.

Content type: article
Source URL: https://www.agentpmt.com/articles/churches-using-ai-now-have-benchmarks-to-test-models
Markdown URL: https://www.agentpmt.com/articles/churches-using-ai-now-have-benchmarks-to-test-models?format=agent-md
Updated: 2026-07-22T20:38:14.935Z
Author: Pancakes
Tags: Successfully Implementing AI Agents, AI Alignment, OpenAI, Controlling AI Behavior, News

---

# Churches Using AI Now Have Benchmarks Built for Faith Questions

The feature article this cycle goes deep on new research measuring the moral defaults inside large language models, and what a measured secular tilt means for picking one. Everything below is the rest of the cycle: a faith-specific benchmark scoring 36 frontier models, a four-university consortium that found religion missing from AI ethics answers entirely, China's first rules for AI companionship, a papal encyclical on machine judgment, and the adoption numbers sitting underneath all of it.

* * *

## Gloo Scored 36 Frontier Models on a Christian Worldview, and Faith Came Last

Gloo published the second edition of its Flourishing AI Insights Report on June 22, expanding from 20 models in December 2025 to 36 frontier models. The setup is two benchmarks run over the same seven dimensions of human flourishing: Character, Relationships, Happiness, Meaning, Health, Finances, and Faith. FAI-G scores general flourishing. FAI-C scores the same questions against a Christian worldview.

The spread between them is the finding. Models averaged roughly 82 on FAI-G and roughly 66 on FAI-C, a 16-point drop. Faith fell hardest at 30 points, the steepest decline of any dimension. Health held up best, losing 6.

Model families did not degrade evenly. The GPT-5 family showed the smallest gaps, between 4 and 11 points. Anthropic's models showed the largest, between 16 and 19. That spread is the practical part: two assistants that look interchangeable on a general benchmark can be 15 points apart the moment a question requires interpretation through a religious frame. It is also the part hardest to act on, because swapping the model underneath a working assistant usually means rebuilding the assistant. Platforms that route model calls through a gateway, AgentPMT among them, turn that swap into a per-step setting and let you run the same questions through two models and read the answers and the cost side by side before committing to either. We argued the general case for measuring rather than guessing in [The Cheapest AI Model Is the One That Finishes the Workflow](https://www.agentpmt.com/articles/the-cheapest-ai-model-is-the-one-that-finishes-the-workflow).

Worth naming the obvious: Gloo publishes this benchmark and also sells the remedy, a set of values-aligned models and a development platform called Gloo AI Studio, and it reports its Christian-tuned models beating peers by 30-plus points on the Faith dimension. A vendor benchmark is still useful as long as you read it as a vendor benchmark. The seven dimensions and the FAI-C question set are the transferable part; a church can lift the dimension list into its own evaluation without buying anything.

**Source:** Gloo

* * *

## A Four-University Consortium Found Models Skipping Religion Rather Than Arguing With It

The Consortium for the Evaluation of Faith and Ethics in AI, a joint effort by Brigham Young University, Baylor University, the University of Notre Dame, and Yeshiva University, announced itself on May 26 at the Summit on AI Ethics in Athens. Its instrument is the AllFaith Benchmark, built from hundreds of real-world ethical questions pulled from ChatGPT transcripts and contributed by faith communities, run against 14 models including Claude 4.7, Gemini 3.1, Grok 4.2, and ChatGPT 5.5.

The headline result is omission rather than opposition. Asked ordinary ethical questions, models steered users toward parents, teachers, friends, and therapists, and away from a pastor, a rabbi, an imam, or any spiritual leader. A companion survey of 1,125 Americans found respondents expected religious perspectives in answers to those same questions, and nearly every model returned none.

A second test measured conversion bias, and the pattern was consistent across models: subtle encouragement toward some traditions and subtle discouragement from others. Catholicism and Protestantism drew positive bias. Jehovah's Witnesses, Baha'i, and Hindus drew negative bias. Grok showed the strongest skew; Anthropic and Meta models showed the least.

The consortium also quantified how little attention the question has had. Reviewing more than 12,000 research papers, it found 0.2% addressing religious bias at all. David Wingate, the BYU computer science professor leading the work, put the stake plainly: "Religion is an important part of human flourishing; 75% of the world's populations maintains religious identity." The consortium's other principals are Paul Martens at Baylor, Fr. John Paul Kimes at Notre Dame, and Rabbi Daniel Feldman at Yeshiva.

Omission and tilt are different failures with different fixes. A model that never surfaces a faith option is not scoring secular, it is scoring silent, and that is usually addressable in the system prompt and retrieval layer rather than by swapping models.

**Source:** BYU News, Baylor University

* * *

## China Wrote the First Rulebook for AI That Keeps You Company

On July 15 the Interim Measures for the Administration of Anthropomorphic Artificial Intelligence Interaction Services took effect in mainland China. Five agencies issued them jointly on April 10: the Cyberspace Administration of China, the National Development and Reform Commission, the Ministry of Industry and Information Technology, the Ministry of Public Security, and the State Administration for Market Regulation.

The scope is narrow and deliberate. Article 2 covers services that simulate personality traits and communication styles to provide continuous emotional interaction, meaning companionship and emotional support. Customer service bots, Q&A, productivity assistants, education, and research are explicitly carved out. This is regulation aimed at the machine you confide in.

The obligations read like an operations manual. Article 18 requires providers to tell users they are talking to an AI, and to interrupt sessions past two hours with a reminder. Article 10 bars designs that aim at replacing social interaction, controlling a user's psychology, or inducing addiction, and Article 8(5) specifically prohibits "excessively catering to users" in ways that create emotional dependence. Article 13 requires detection of extreme emotional states, with immediate intervention on self-harm or suicide indicators, including contacting a guardian or emergency contact. Article 14 bans virtual intimate relationships for minors outright and requires a minor mode with usage limits. Services above one million registered users or 100,000 monthly actives face security assessments under Article 22.

None of this binds a congregation in Ohio. It is still the first written answer to a question that faith organizations reach faster than most sectors: what does a system owe someone who is confiding in it? A ministry running a prayer line, a grief chatbot, or after-hours counseling triage can adopt the shape without waiting for a local statute. Disclose the bot. Cap the session. Route distress to a human with a name.

**Source:** Hunton Andrews Kurth, Geopolitechs

* * *

## OpenAI and Anthropic Are Taking Ethics Meetings With Religious Leaders

The Interfaith Alliance for Safer Communities, a Geneva-based body founded in 2018, convened the first Faith-AI Covenant roundtable in New York on April 30. OpenAI and Anthropic sent representatives. So did the Baha'i International Community, The Sikh Coalition, the Archdiocese of Newark, the Greek Orthodox Archdiocese of America, the New York Board of Rabbis, The Church of Jesus Christ of Latter-day Saints, and the Hindu Temple Society of North America.

The convener's premise: "Faith traditions have long helped humanity interpret the world, define moral boundaries and shape the values by which societies live." Baroness Joanna Shields, a former Google and Facebook executive now running Precognition, made the case for the format by conceding the limits of the alternative, arguing that "regulation simply can't keep up with the pace of development."

Six further convenings are planned through 2026, in Beijing, Bengaluru, Nairobi, Paris, and Singapore, closing in Abu Dhabi. Anthropic has previously consulted faith leaders while shaping Claude's constitutional guidelines, so the channel predates the roundtable.

Critics have called the initiative public relations, and at least one described it as a distraction from questions about regulation and who holds power over these systems. Both readings can hold. A closed-door roundtable is an informal input channel with no enforcement behind it, which is precisely why the more reliable leverage a faith organization has runs through procurement: what it buys, what it writes into its evaluation criteria, and what it declines to renew.

**Source:** The Decoder, National Law Review

* * *

## Pope Leo XIV Drew the Line at Delegating Judgment, Not at Using the Tools

Magnifica Humanitas, the first encyclical of Pope Leo XIV, was signed on May 15 and released on May 25, timed to the 135th anniversary of Rerum Novarum. It runs about 42,300 words across five chapters and 245 sections, and its subject is safeguarding the human person in the time of artificial intelligence.

Its premise is not refusal. The text holds that technology is neither antagonistic to humanity nor inherently evil, then adds the qualifier doing the work: technology "is never neutral, because it takes on the characteristics of those who devise, finance, regulate, and use it." That is the same claim this cycle's benchmark research arrives at from the measurement side, reached by a different road.

The encyclical is specific about what machines are not. They "do not undergo experiences, do not possess a body, do not feel joy or pain," and they "do not have a moral conscience, since they do not judge good and evil." From there it opposes "complete delegation of important and sensitive decision-making" to automated systems, treats human dignity as the fundamental criterion for assessing any technological development, and defends the dignity of labour against automation planned around efficiency alone.

For a parish or diocese building an actual policy, that is a usable boundary rather than a mood. It puts the limit on delegated judgment and consequential decisions, and leaves drafting, translation, scheduling, and research on the permitted side of the line. It is also, for the roughly 1.4 billion members of the world's largest religious body, the closest thing to a governing document on the question.

**Source:** The Holy See, National Catholic Register

* * *

## Most Church Leaders Use AI Personally. Almost None Have Written Anything Down.

The fifth annual State of Church Technology report from Pushpay and Barna Group, released March 9 and built on a late-2025 survey of more than 1,300 U.S. church leaders, supplies the baseline the rest of this cycle acts on. Sentiment is not the obstacle: nearly every leader surveyed says technology opens new ministry opportunities and helps the church fulfill its mission in a digital culture.

Usage splits sharply by scope. About 60% of leaders use AI personally at least a few times a month, but just 33% say their church uses it in ministry or operations. Most common applications are written materials, graphics, emails, and social posts, with sermon work appearing in some responses. That boundary is the one we traced in [AI for Church Leaders Stops at the Pulpit](https://www.agentpmt.com/articles/ai-for-church-leaders-stops-at-the-pulpit).

Then the number that defines the year's work: 64% of churches consider an AI policy important, and 5% have one. That gap is the entire opportunity. A church AI statement is not a compliance artifact, it is the document that decides which model answers a grieving member at 2 a.m. and who reads that answer before it sends. We covered the shape of that gap in [Churches Using AI Race Ahead of Their Own Rules](https://www.agentpmt.com/articles/churches-using-ai-race-ahead-of-their-own-rules).

At congregation scale, the questions are already arriving. WBUR reported on July 15 that Rev. Chris Hope of the Pentecostal Tabernacle Church in Cambridge started a newsletter, The Tech Rev Column, specifically to field what his members ask him about faith chatbots and how pastors are using these tools.

**Source:** Pushpay and Barna Group, WBUR

* * *

## What This Cycle Actually Handed Ministry Leaders

Before this cycle, a church writing an AI policy had opinions and anecdotes to work from. It now has a scored benchmark with a Faith dimension, an omission-and-conversion test set built by four universities with published methodology, a regulatory template for anything doing emotional support, and a doctrinal boundary drawn at delegated judgment rather than at the tools themselves.

That is enough to write something real. Name the evaluation dimensions you care about. Run the omission test on whatever assistant answers member questions today. Disclose the bot, cap the session, and route distress to a person. Say in writing which decisions stay human. None of it requires a data science team, and all of it is easier this month than it was last month.

* * *

## Sources

-   Flourishing AI Initiative Insights Report, June 2026, Gloo
-   New research from BYU-led multi-institution consortium finds all major AI models ignore faith, religion in responses, BYU News
-   Baylor Joins Multi-University Consortium to Launch First Cross-Faith AI Benchmark, Baylor University
-   China Rolls Out Interim Regulations on AI Human-Like Interaction Services: A Detailed Analysis, Geopolitechs
-   China's First Regulatory Framework for Virtual Companions Soon to Take Effect, Hunton Andrews Kurth
-   Anthropic and OpenAI sit down with religious leaders to seek ethical advice, The Decoder
-   Interfaith Alliance for Safer Communities Convenes First Faith-AI Covenant Roundtable in New York, National Law Review
-   Magnifica Humanitas (encyclical letter), The Holy See
-   Full Text of 'Magnifica Humanitas', National Catholic Register
-   Pushpay and Barna Group's 2026 State of Church Technology Report, Pushpay and Barna Group
-   How one reverend helps his congregation navigate AI and faith, WBUR