How this app is checked

Accuracy · Context · Source. Here is what is behind those three words.

AI makes things up. In a Bible study app that is not survivable, so this app does not rely on telling the AI to be careful. Three things check the work, and this page walks through each one, shows you what the checks have caught, and names exactly where the checking stops.

1
Code, on every study
Plain computer code with no opinion. Runs on every single study, never spot-checked.
2
A count, kept out loud
Every study writes down what happened to it, good news or bad, and the totals live on this page.
3
A rival AI, reading
A competing model reads finished studies line by line, under instructions to prove them wrong.
The running count, live from the checking layer

The lifetime counter is being built right now. When it exists, the real numbers go here, whatever they say.

These numbers come straight from the same records the checks write. Our own testing traffic is subtracted before anything is counted here.

The first check: code, running on every study

This part is plain computer code. It has no opinion. It gives the same answer every time, and it runs on every single study. Nothing is spot-checked.

THE APP LOOKS UP EVERY VERSE ITSELF
The AI never types out a Bible verse. It writes down the reference, like John 1:1, and the app looks up that verse and puts the real published words in before anything reaches you. Six full Bibles are stored inside the app: the King James Version, the Berean Standard Bible, the World English Bible, the American Standard Version, Young's Literal Translation and the Darby Translation. Eight more come straight from their publishers with the copyright line attached: NASB 2020, the New Living Translation, the New King James Version, The Message, the Christian Standard Bible, the Amplified Bible, the Good News Translation and the Contemporary English Version. Under any quote, tap "Check this passage" to read the verse on its own.
THE APP BUILDS THE WORD CARDS ITSELF
When a study digs into a Greek, Hebrew or Aramaic word, the AI writes down only the verse and the word. The app builds the rest of the card from dictionaries stored inside it: the original script, how to say it, which language it is, what it means, and which dictionary each meaning comes from. Whatever the AI typed inside that card is thrown away without being read. If the word is not actually in the verse, the whole card is dropped rather than shown as a guess. Those dictionaries hold 447,384 tagged words and 40,367 entries.
ARAMAIC IS CALLED ARAMAIC
Some of the Old Testament is Aramaic rather than Hebrew: Daniel 2:4 to 7:28, Ezra 4:8 to 6:18, Ezra 7:12 to 26, and Jeremiah 10:11. The app calls those 4,827 words Aramaic, because it reads the grammar tag on each individual word instead of assuming an Old Testament book must be Hebrew.
EVERY BOOK NAMED IN A STUDY HAS TO SHOW UP IN ITS SOURCE LIST
If a study mentions a commentary in the text, that commentary has to appear in the study's own source list at the bottom, and anything in the list has to be used in the text. Either kind of gap gets flagged. Author, book and year are what a library catalog can confirm, so that is the level the check is built on. Page and volume numbers are a separate matter: the app allows one where the AI is confident of it, nothing in the pipeline verifies it, and the writing rules require a study carrying any to say so at the foot of its source list.
THE BOOKS ARE REAL BOOKS
Books cited in a study are matched against a shelf of 464 reference works. Every one of them has been confirmed to exist in a public library catalog, through OpenLibrary and OpenAlex, and the record each book was matched to is stored alongside the shelf. Anything cited from off that shelf is looked up in public library catalogs. When a lookup comes back empty, or the catalog turns the lookup away and a second, cleaned-up attempt fails too, the study keeps the citation and loses the line that says its citations were cross-referenced. A citation the catalogs could not check never counts as a citation they confirmed.
None of this is the AI checking its own homework. It is ordinary code comparing what the AI wrote against something that cannot be wrong: the stored Bible text, the stored dictionaries, the library catalog. Think of a spellchecker. A spellchecker is not the writer deciding their own spelling looks fine. It compares every word against a dictionary. This is that, for Scripture, for original-language words, and for books.

We attack our own book check, and here is the full result

A check you have never attacked is a check you know nothing about. So in August 2026 we built 300 fake books, planted them in the checking pipeline alongside 30 real ones, and recorded exactly what happened to each. Every square below is one planted fake.

The fakes were not all the same kind, and the difference between the kinds is the whole story. A number by itself, like "16 got through," would tell you almost nothing, because it matters enormously what got through.

Fakes where the book does not exist at all105 planted · all 105 caught

A real scholar's name on a book that was never written. This is the classic AI hallucination, the kind that would truly mislead you, and the check caught every single one.

Fakes where the book is real but the wrong person is attached45 planted · 44 caught · 1 accepted

The one it accepted embarrassed the test, and in a way we did not expect: when we chased it down, the "wrong" editor we had invented for a reference work turned out to really have edited its newest edition. The check was right and our fake was accidentally true.

Fakes where the book is real, with an invented volume or book number riding on it105 planted · 90 caught · 15 accepted

This is where the check actually leaks, so look at what a leak means here. Every one of the 15 points at a real book, by the real author, on the real subject. What is wrong is a volume or section number bolted onto it. A reader who follows one lands on the right shelf at the wrong door. That is a flaw worth fixing, and it is a different universe from a source that does not exist.

Fakes where a real book carries an invented subtitle45 planted · this method cannot test them

Library catalogs can confirm an author, a title and a year. A subtitle riding on a real title is beyond what they can settle today, so we report this class as unmeasured rather than pretending the squares are green. Closing it means matching against each book's barcode number, a real job on the roadmap and not yet done.

Caught Accepted Cannot be tested this way

The honest arithmetic

Of the 255 fakes this method can genuinely test, 16 were accepted. That is about 6 in 100, with a statistical ceiling closer to 9 in 100 at this sample size, and we publish it rather than rounding it away. Two more things give that figure its meaning. First, every accepted fake was of the decorated-real-book kind; the invented-source kind went 0 for 105. Second, the 30 real books planted as controls all came through, so the check is not passing the test by simply rejecting everything it sees.

An earlier round used 15 fakes and caught all 15. We stopped citing that result when this one existed: fifteen samples was too small to find any of what the 300 found, and a clean scorecard from a small test says more about the test than the check.

The second check: the app counts its own work, out loud

A check nobody reads is a check nobody can trust, so every single study writes down what happened to it. The actual working, never a bare pass or fail: how many Bible quotes were in the study, how many the app printed from its own stored text, how many came from a publisher, how many it compared word for word, and how many it could not vouch for.

That line gets written whether the news is good or bad. That sounds obvious and it is the single most important design decision on this page. A checker that only speaks up when something is wrong goes silent in two situations that look identical from outside: a perfect day, and a checker that has quietly died. This one talks every time, so the difference is always visible.

Two findings stop everything and go straight to the top of the report: a study where any quoted Scripture could not be accounted for, and any sign that the AI typed out Scripture itself instead of asking the app for it. Every morning those lines are gathered into a health report and read before anything else about the business.

The counter at the top of this page is this check made public. It is fed by the same records, updated as studies happen, with our own testing traffic subtracted, and it stays on this page whatever it says.

The third check: a rival AI reads studies line by line

Some things have no database to check them against. No catalog anywhere can tell you whether a particular scholar really holds a particular view. So a competing AI model, with web search, reads finished studies claim by claim, under instructions to assume everything is wrong and try to prove it.

The largest run so far, in August 2026, put 55 fixed studies through three reading packs: 25 written to bait the AI into inventing things, 10 on the most disputed passages in Scripture, and 20 drawn from across the whole Bible.

What was checkedHow muchResult
25 trap prompts: fake books, fake quotes, invented consensus, pressure for page numbersEvery study read line by lineIt did not take the bait, on any of the 25
The hardest disputed passages, checked claim by claim10 studiesNo made-up sources, no views credited to the wrong scholar
Ordinary passages from Genesis to Revelation20 studiesOne real error found, described below
The one real error, told in full because it earns this page its keep. In one of the 20 breadth studies, the reviewer flagged a book title, and it was right: the study had cited a real scholar, in her real year, in her real series, on her real subject, under a title she never used. Her actual book exists and says roughly what the study attributed to it, which is exactly what made the error hard for a human eye to catch. Two things came out of that finding. The invented title had slipped past the book check because the catalog had turned that particular lookup away, and the pipeline then counted "we could not check this" as a pass; that is fixed, and such a citation now loses its cross-referenced line instead. And the error itself is the same low-severity shape as the test result above: a decoration on a real source rather than an invented one. We would rather show you one real error and what it changed than a clean table you have to take on faith.
What the table above still leaves out. Nobody counted how many claims were in those studies to begin with. That matters more than it sounds. If one reviewer checks ten claims and another checks a hundred, and both come back clean, from the outside they look identical. Until we have that count, no percentage goes on this page next to these rounds.

Compared to what

The fair comparison is everything else on the shelf rather than some perfect book, because a share of published citations have always failed to say what they were claimed to say. Two studies measured it carefully.

16.6%
of 2,648 authors, shown a real citation of their own work, said it did not support the claim attached to it. The rate held across every field.
Wakeling, Paramita & Pinfield, JASIST, 2025
14.5%
citation error rate when fifteen earlier studies were put on one definition. About two thirds were the serious kind, where the source does not support the claim at all.
Mogull, PLOS ONE, 2017

We are not claiming a machine reads better than a scholar does. The claim is smaller, and you can check it: this app's citations get checked on every study, while the published record's citations mostly never get checked at all.

Why the checking works here

In 2026 researchers at Princeton took 1,300 pieces of legal writing, planted citation errors in them, and set AI agents loose to find the errors. The agents caught invented cases and wrong case names most of the time. Not one of them reliably caught a wrong page number. The reason is in the paper: those page numbers live inside expensive subscription databases the agents cannot see, and about a fifth of the court opinions they managed to pull had no usable page information at all.

That is the whole thing in one finding. How hard something is to check depends on whether you can see what you are checking against. It is why this app keeps its Bibles and its dictionaries inside the download instead of asking some outside service for them. Holding the real thing is what makes the comparison cheap enough to run every single time. And notice how the Princeton result rhymes with ours: invented sources are catchable and got caught; small numbers riding on real sources are the hard residue. That is the industry's frontier, and this page shows you exactly where this app sits on it.

Two more results from that same study are worth knowing. Pointed at legal writing from before AI drafting existed, the same agent turned up more than ten real citation mistakes that no machine had ever gone looking for. And AI does not simply get more accurate as it gets newer: the most recent model in the study invented citations more often than one from two years earlier did. Checking is permanent. It does not become unnecessary as the technology grows up.

Where the checking stops

Every limit below is one we went looking for and found. Knowing precisely where a check ends is the difference between a system you can inspect and a system asking you to take its word for the whole thing, and most of these have something deliberate standing in their place.

What we claim, and what we do not

What we claim: every Bible quote is looked up and printed by the app rather than typed by the AI. Every word card is built from a dictionary stored in the app. Every interpretive claim names a real source. Every study writes down what the checks found, those findings are read every day, and their running total lives at the top of this page. A rival AI has read studies line by line, found one error across every round so far, and that error changed the pipeline. The method and the raw studies are published.

What we do not claim: perfection, or any percentage we cannot show you the full working for. Every limit we know about is listed above rather than left out. AI writing is a starting point for study. The app says so on every study, and the primary sources always win.

Check us

The raw studies, unedited and exactly as the app produced them: browse them here. The method: how the test set is built. It is 55 fixed prompts: 25 written to tempt the AI into making something up, 10 on the most disputed passages in Scripture, and 20 ordinary passages from across the whole Bible, and the whole set runs again after any significant change. Found an error we missed? Tell us. Finding it is the point.