How BLI Works
The Bible Literacy Index
Open Bible Assessment was a difficult website to create. By far the hardest task was transforming a standard assessment with preset questions into one that changes based on how the learner answers.
Looking for the plain explanation? About covers what OBA is and how to use it. This page is the layer underneath: how the scoring and the question selection actually work.
Like a detective
OBA can be thought of as a detective who starts an investigation knowing nothing. They begin with broad questions, and as they gather evidence the questions get narrower and more specific.
The assessment is trying to create a full picture of the learner by going through this same broad-to-narrow process. Learners are first asked things like:
Which of these events occurred first? A. The tower of Babel B. Daniel and the lions’ den …
OBA is trying to determine whether the learner has at least a basic understanding of the Old Testament narrative. Part of that process is figuring out which sections of the Bible the learner understands well and which they understand poorly. Each section carries an estimated ability score.
You appear broadly strong in Torah narrative, moderately stable in Former Prophets, weak in Latter Prophets geography and prophetic context, and not yet sufficiently tested in Writings. Your current OT BLI estimate is X, but recommendations need more evidence before naming a confident next study area.
But also like a detective, one piece of evidence is generally not sufficient. The assessment — whether the initial assessment or a later one — eventually returns to the areas where it does not have enough evidence to draw a conclusion about someone’s ability.
Narrowing down
As you prove a given section, like the Torah, to be a strength, OBA tries to fill in the fuller, more detailed picture. Which books in the Torah are your strongest? Which are weakest? Genesis is your best book — so which part do you know better, chapters 1–11 or 12–50, the way Genesis is typically divided?
The same process runs on weak sections. The learner does not have a broad understanding of the Latter Prophets — does that include the major prophets as well, like Isaiah?
Dimensions
Narrowing happens in a second direction too. The system tracks knowledge dimensions, not just Bible books, so a weakness can be named more precisely than “the Prophets.” A learner might be strong on Promise & Prophecy in the Latter Prophets but weak on Geography & Nations there. These names appear on your dashboard when a gap is identified.
- Events & Timeline
- Can the learner place events in order? Which came later: Sinai or the monarchy?
- Characters & Lineage
- Does the learner know people and relationships? Who is associated with the Davidic line?
- Geography & Nations
- Does the learner know locations and nations? Which empire is connected with Judah’s exile?
- Law & Commands
- Does the learner understand law/covenant material? What is Leviticus especially concerned with?
- Promise & Prophecy
- Does the learner understand prophetic promises and messages? Which prophet emphasizes restoration after judgment?
- Theological Reasoning
- Can the learner identify significance? Why is the exile important in the OT storyline?
- Cross Ref
- Can the learner connect books, sections and passages? Which later book develops themes introduced in Torah?
Dialing the difficulty
As a general rule, when the learner answers correctly the question difficulty is dialed up for that section, and when they answer incorrectly it is dialed down. If you answer a specific question about Nahum incorrectly, OBA will give you a broader Latter Prophets question instead.
Why it is not a percentage
This is why OBA does not grade on a percentage system, where 8 out of 10 is 80% and 80% is good. It works more like the rating on a chess site such as Chess.com. Most players win about half of their games, but when they win their rating goes up and they play harder opponents, and when they lose it goes down and they play easier ones.
OBA rests on a similar assumption to Chess ELO scoring: that there is a real level of knowledge underneath the answers, and that a test can estimate it without ever measuring it perfectly. Whether someone knows who led the Israelites into the promised land is not a matter of opinion. That knowledge does not equate to being regenerate or spiritually wise, but it is important for knowing who God is and who we are.
A few more factors
Several other things shape what you get asked.
- Difficulty comes in three stages. Stage 1 is the broad, foundational level everyone starts at. Stages 2 and 3 open up as your recent accuracy holds.
- Which estimate the next question is aimed with depends on how much you have answered. A section that has been answered enough times uses its own estimate; one that has not falls back to your testament-wide estimate; with too little of either, the question comes from stage 1. The thresholds are in the confidence table further down. The part doing the aiming is called the router.
- Confidence decays with time. The longer since a section was last tested, the wider its margin of error grows, even though the estimate itself does not move.
- The router aims a little below its own best guess — and the less sure it is, the further below it aims. Together with the decay above, this means that coming back after a long gap gets you slightly easier questions than your old score alone would suggest — not because the system thinks you got worse, but because it is less sure, and it would rather re-establish the floor than open with something you cannot use.
- The session brake. A sustained run at or below 25% correct drops straight to stage 1; a less severe run holds the stage where it is rather than dropping it; two misses in a row cost one step down rather than a reset.
- The dimension brake. Two misses in the same dimension flag a possible gap and a third confirms it, after which the router moves on to something else for this session.
- The repeat cooldown. A question answered recently is held back until its cooldown has passed, so revisiting a section does not replay the same items.
The last three keep one weak patch from taking over a whole session. Weak areas are still prioritized, but by rule rather than by whatever you just missed: a single past miss earns one fresh confirmation probe, while a weakness you have confirmed by missing the same kind of question repeatedly is treated as settled and stops being re-tested for its own sake.
What the score means
BLI stands for Bible Literacy Index. It is scored from 0–800 for each testament. For the Old Testament, a higher score means:
- You know broad OT structure.
- You know major events, people, places, and order.
- You can distinguish sections and books.
- You can answer more specific questions about a book, a stretch within it, or a chapter range.
- At the high end, you know meaningful textual detail, not just broad summaries.
A learner who knows
- Genesis comes before Exodus.
- Moses leads the Exodus.
- David comes before Solomon.
- Isaiah is a prophet.
That learner has real knowledge, but it is mostly broad and foundational. They should not score near 800.
A learner who also knows
- Genesis 1–11 is primeval history; Genesis 12–50 follows the patriarchs.
- Deuteronomy is covenant renewal before entry into the land.
- 1 Kings moves from Solomon to divided kingdom decline.
- Jeremiah is tied to Judah’s final collapse and new covenant hope.
- Ezra–Nehemiah belongs to post-exilic restoration.
- Isaiah’s restoration promises are not the same thing as Amos’s judgment or Haggai’s temple focus.
That learner has shown broader and deeper knowledge, so the BLI should rise. This is why a score can stall even while you keep answering correctly: broad questions alone support basic literacy, and moving toward the top of the scale takes evidence of textual depth.
How to read the 0–800 scale
The bands below are labels for ranges on the same 0–800 scale, not separate tests.
| Band | Range | What it describes |
|---|---|---|
| Unfamiliar | 0–120 | You do not yet have a steady grasp of the Old Testament's major plot, events, characters, and book locations. |
| Acquainted | 121–312 | You recognize some major people and stories, but many core events, sequences, and book-level connections are still forming. |
| Familiar | 313–512 | You know many major stories and characters, with growing awareness of where events belong and how they connect. |
| Literate | 513–632 | You can navigate the Old Testament with confidence, connecting books, events, characters, and theological themes. |
| Studied | 633–712 | You show detailed knowledge of the text, including less obvious events, patterns, references, and historical flow. |
| Learned | 713–760 | You understand the Old Testament at a deep level, with strong command of structure, sequence, characters, and theology. |
| Scholar | 761–800 | You demonstrate exceptional mastery, including fine textual detail, interconnections, and theological architecture. |
The current scoring model
This is the arithmetic behind the number, included for review and debugging. Two things decide what a question is worth before difficulty enters into it: how foundational its book is, and how significant the event it asks about is.
chronological_weight runs from 0.65 to 1.00 and measures how much of the rest of Scripture leans on a book: Genesis, Exodus and the Gospels sit at 1.00, Deuteronomy at 0.90, Kings and Isaiah at 0.85, Chronicles and the Minor Prophets at 0.65. It is the same dependency idea About uses to explain why Exodus is recommended before Ezekiel — despite the name, it is not a measure of chronology. The importance tier belongs to the event a question asks about rather than to the question itself: tier 1 is the load-bearing episodes (the burning bush, Israel demanding a king), while tier 3 and below is incidental material (a second census, assorted case laws).
Harder items push the reward toward 1.25 and easier items pull it toward 0.70, so a correct answer on an easy question is worth less than full weight, not more. irt_difficulty is each question’s difficulty parameter as calibrated by item response theory, not a hand-set label. The negative value for a wrong answer is a guessing correction: in a four-choice question, random guessing would be right about one time in four. Answers from attempts that fail quality checks, and individual answers marked ineligible, are excluded from both sides of the ratio rather than scored.
Why recommendations wait
OBA is as careful with advice as it is with difficulty. Recommending too early would look like this:
You answered 8 questions. You missed 2 Latter Prophets questions. Therefore your recommendation is Isaiah.
That is too confident. Instead:
Recommendations need more evidence. Answer more questions across the OT so OBA can distinguish a real weakness from a short-sample accident.
Once there is enough to go on, recommendations sharpen: from a whole section, to a book, to a stretch within that book, to a dimension.
Progression
- More evidence needed across OT sections.
- Former Prophets looks weak.
- The weakness seems concentrated in Kings rather than Joshua.
- The weak dimensions appear to be Events & Timeline and Geography & Nations.
- Recommended focus: divided kingdom and exile sequence in Kings.
How much evidence a section carries is tracked directly:
| Confidence | Section evidence | What it means |
|---|---|---|
| Provisional | 0–14 answers | Not enough yet to trust a weakness. |
| Developing | 15–29 answers | Enough to read the section, but not enough to rely on. |
| Established | 30+ answers | Enough evidence to trust what the section says about you. |
Beta
OBA is ready for beta use on the Old Testament, but the question bank is still being reviewed and improved.
Concretely, that means some question wording is still awkward, a few coverage areas are thin, human quality ratings are sparse, and the BLI is an educational estimate rather than a credentialing-grade measure. Repeats and near-duplicates are handled by the cooldown described above rather than left to chance, but the difficulty calibration the scoring model rests on is still improving as more people answer.
See it running
The knowledge map shows the same structure this page describes — sections, books and passages, with evidence filled in as you answer. The intro presentation walks through it visually.