How this app is checked
Accuracy · Context · Source. Here is what is behind those three words.
AI makes things up. In a Bible study app that is not survivable, so this app does not rely on telling the AI to be careful. Three things check the work, and this page walks through each one, shows you what the checks have caught, and names exactly where the checking stops.
The lifetime counter is being built right now. When it exists, the real numbers go here, whatever they say.
The first check: code, running on every study
This part is plain computer code. It has no opinion. It gives the same answer every time, and it runs on every single study. Nothing is spot-checked.
The AI never types out a Bible verse. It writes down the reference, like John 1:1, and the app looks up that verse and puts the real published words in before anything reaches you. Six full Bibles are stored inside the app: the King James Version, the Berean Standard Bible, the World English Bible, the American Standard Version, Young's Literal Translation and the Darby Translation. Eight more come straight from their publishers with the copyright line attached: NASB 2020, the New Living Translation, the New King James Version, The Message, the Christian Standard Bible, the Amplified Bible, the Good News Translation and the Contemporary English Version. Under any quote, tap "Check this passage" to read the verse on its own.
When a study digs into a Greek, Hebrew or Aramaic word, the AI writes down only the verse and the word. The app builds the rest of the card from dictionaries stored inside it: the original script, how to say it, which language it is, what it means, and which dictionary each meaning comes from. Whatever the AI typed inside that card is thrown away without being read. If the word is not actually in the verse, the whole card is dropped rather than shown as a guess. Those dictionaries hold 447,384 tagged words and 40,367 entries.
Some of the Old Testament is Aramaic rather than Hebrew: Daniel 2:4 to 7:28, Ezra 4:8 to 6:18, Ezra 7:12 to 26, and Jeremiah 10:11. The app calls those 4,827 words Aramaic, because it reads the grammar tag on each individual word instead of assuming an Old Testament book must be Hebrew.
If a study mentions a commentary in the text, that commentary has to appear in the study's own source list at the bottom, and anything in the list has to be used in the text. Either kind of gap gets flagged. Author, book and year are what a library catalog can confirm, so that is the level the check is built on. Page and volume numbers are a separate matter: the app allows one where the AI is confident of it, nothing in the pipeline verifies it, and the writing rules require a study carrying any to say so at the foot of its source list.
Books cited in a study are matched against a shelf of 464 reference works. Every one of them has been confirmed to exist in a public library catalog, through OpenLibrary and OpenAlex, and the record each book was matched to is stored alongside the shelf. Anything cited from off that shelf is looked up in public library catalogs. When a lookup comes back empty, or the catalog turns the lookup away and a second, cleaned-up attempt fails too, the study keeps the citation and loses the line that says its citations were cross-referenced. A citation the catalogs could not check never counts as a citation they confirmed.
We attack our own book check, and here is the full result
A check you have never attacked is a check you know nothing about. So in August 2026 we built 300 fake books, planted them in the checking pipeline alongside 30 real ones, and recorded exactly what happened to each. Every square below is one planted fake.
The fakes were not all the same kind, and the difference between the kinds is the whole story. A number by itself, like "16 got through," would tell you almost nothing, because it matters enormously what got through.
A real scholar's name on a book that was never written. This is the classic AI hallucination, the kind that would truly mislead you, and the check caught every single one.
The one it accepted embarrassed the test, and in a way we did not expect: when we chased it down, the "wrong" editor we had invented for a reference work turned out to really have edited its newest edition. The check was right and our fake was accidentally true.
This is where the check actually leaks, so look at what a leak means here. Every one of the 15 points at a real book, by the real author, on the real subject. What is wrong is a volume or section number bolted onto it. A reader who follows one lands on the right shelf at the wrong door. That is a flaw worth fixing, and it is a different universe from a source that does not exist.
Library catalogs can confirm an author, a title and a year. A subtitle riding on a real title is beyond what they can settle today, so we report this class as unmeasured rather than pretending the squares are green. Closing it means matching against each book's barcode number, a real job on the roadmap and not yet done.
The honest arithmetic
Of the 255 fakes this method can genuinely test, 16 were accepted. That is about 6 in 100, with a statistical ceiling closer to 9 in 100 at this sample size, and we publish it rather than rounding it away. Two more things give that figure its meaning. First, every accepted fake was of the decorated-real-book kind; the invented-source kind went 0 for 105. Second, the 30 real books planted as controls all came through, so the check is not passing the test by simply rejecting everything it sees.
An earlier round used 15 fakes and caught all 15. We stopped citing that result when this one existed: fifteen samples was too small to find any of what the 300 found, and a clean scorecard from a small test says more about the test than the check.
The second check: the app counts its own work, out loud
A check nobody reads is a check nobody can trust, so every single study writes down what happened to it. The actual working, never a bare pass or fail: how many Bible quotes were in the study, how many the app printed from its own stored text, how many came from a publisher, how many it compared word for word, and how many it could not vouch for.
That line gets written whether the news is good or bad. That sounds obvious and it is the single most important design decision on this page. A checker that only speaks up when something is wrong goes silent in two situations that look identical from outside: a perfect day, and a checker that has quietly died. This one talks every time, so the difference is always visible.
Two findings stop everything and go straight to the top of the report: a study where any quoted Scripture could not be accounted for, and any sign that the AI typed out Scripture itself instead of asking the app for it. Every morning those lines are gathered into a health report and read before anything else about the business.
The counter at the top of this page is this check made public. It is fed by the same records, updated as studies happen, with our own testing traffic subtracted, and it stays on this page whatever it says.
The third check: a rival AI reads studies line by line
Some things have no database to check them against. No catalog anywhere can tell you whether a particular scholar really holds a particular view. So a competing AI model, with web search, reads finished studies claim by claim, under instructions to assume everything is wrong and try to prove it.
The largest run so far, in August 2026, put 55 fixed studies through three reading packs: 25 written to bait the AI into inventing things, 10 on the most disputed passages in Scripture, and 20 drawn from across the whole Bible.
| What was checked | How much | Result |
|---|---|---|
| 25 trap prompts: fake books, fake quotes, invented consensus, pressure for page numbers | Every study read line by line | It did not take the bait, on any of the 25 |
| The hardest disputed passages, checked claim by claim | 10 studies | No made-up sources, no views credited to the wrong scholar |
| Ordinary passages from Genesis to Revelation | 20 studies | One real error found, described below |
Compared to what
The fair comparison is everything else on the shelf rather than some perfect book, because a share of published citations have always failed to say what they were claimed to say. Two studies measured it carefully.
We are not claiming a machine reads better than a scholar does. The claim is smaller, and you can check it: this app's citations get checked on every study, while the published record's citations mostly never get checked at all.
Why the checking works here
In 2026 researchers at Princeton took 1,300 pieces of legal writing, planted citation errors in them, and set AI agents loose to find the errors. The agents caught invented cases and wrong case names most of the time. Not one of them reliably caught a wrong page number. The reason is in the paper: those page numbers live inside expensive subscription databases the agents cannot see, and about a fifth of the court opinions they managed to pull had no usable page information at all.
That is the whole thing in one finding. How hard something is to check depends on whether you can see what you are checking against. It is why this app keeps its Bibles and its dictionaries inside the download instead of asking some outside service for them. Holding the real thing is what makes the comparison cheap enough to run every single time. And notice how the Princeton result rhymes with ours: invented sources are catchable and got caught; small numbers riding on real sources are the hard residue. That is the industry's frontier, and this page shows you exactly where this app sits on it.
Two more results from that same study are worth knowing. Pointed at legal writing from before AI drafting existed, the same agent turned up more than ten real citation mistakes that no machine had ever gone looking for. And AI does not simply get more accurate as it gets newer: the most recent model in the study invented citations more often than one from two years earlier did. Checking is permanent. It does not become unnecessary as the technology grows up.
Where the checking stops
Every limit below is one we went looking for and found. Knowing precisely where a check ends is the difference between a system you can inspect and a system asking you to take its word for the whole thing, and most of these have something deliberate standing in their place.
- Whether the verse fits the point. The words of a quote are guaranteed. Whether that verse actually supports the argument built around it is judgment, and no amount of code settles it. This is one of the main things the rival AI is pointed at, and it is why that third check exists at all.
- Greek and Hebrew words inside ordinary paragraphs. The word cards are locked down completely. A word mentioned in passing in a paragraph is not yet tied to a dictionary entry, so the rule for now is that the AI may reach for an original-language word only where a card is standing behind it. Tying prose words to the dictionary too is actively being built.
- Which edition of a book. This is a deliberate trade. The app cites author, book and year, which is exactly the level a library catalog can confirm. Page and volume numbers are allowed where the AI is confident of them and nothing verifies them, which is why a study carrying one is required to say so under its sources. The cost of that trade is visible in the test above: decorations on real books are the one place fakes get through, and adding each book's barcode number to the shelf is the real, not-yet-done job that would close it.
- Whether a scholar really said it, in most cases. This one is a copyright wall rather than a gap in the work. Of the 464 books, only 23 are old enough that anyone can read their full text for free, so those 23 are the only ones a machine could ever check against the book itself. For the rest the app is honest about which kind of confirmation it has: the book is real, and the view attributed to it rests on the writing rules and the audit rounds.
- Interpretation itself. This one is deliberate and permanent. Studies have to present a disputed reading as disputed and name who holds each view, and that is enforced by the writing rules and by the audits. A database deciding whose reading of John 6 is correct would be worse than admitting the limit.
- Two words in Genesis 31:47. Jegar-sahadutha is Aramaic, and the scholarly grammar data the app is built on labels those two words Hebrew. The app now corrects that for exactly those two words and nothing else, because there is no reliable machine signal for Aramaic hiding inside Hebrew-tagged text, and we would rather name one exception in the open than invent a rule that guesses. We found this by checking our own Aramaic claim rather than trusting it, and it is exactly the size of thing this page exists to show.
What we claim, and what we do not
What we do not claim: perfection, or any percentage we cannot show you the full working for. Every limit we know about is listed above rather than left out. AI writing is a starting point for study. The app says so on every study, and the primary sources always win.
Check us
The raw studies, unedited and exactly as the app produced them: browse them here. The method: how the test set is built. It is 55 fixed prompts: 25 written to tempt the AI into making something up, 10 on the most disputed passages in Scripture, and 20 ordinary passages from across the whole Bible, and the whole set runs again after any significant change. Found an error we missed? Tell us. Finding it is the point.