Build 01 · closes Section 1

Test your chat tool in twelve questions.

Section 1 explained four things about what a model is. This build turns them into twelve short tests you run in your own chat tool, and a card at the end that says, from your own results, what that tool can be trusted for.

MY TOOL CARD Dated factsLetters & digitsSame answer twiceLikely vs true 2 / 31 / 32 / 32 / 3
Before you start

What this build is, and what you will get from it

About 40 minutes
Any chat tool, search off
Uses Notes 01–04 only
We are building

A tool card: one card that says what the chat tool you actually use can be trusted for, measured by you. You get it by running twelve short tests, three for each note in this section, and ticking pass or fail on this page. The page scores them and writes the card.

The objective

Section 1 explained four ways a fluent answer can be a guess: the model predicts rather than knows, it never sees letters, a dial spreads the answers, and it stopped reading on a date. Reading that is one thing. Seeing all four happen in your own tool, in one sitting, and coming away with a card you made, is what makes it stick.

You will learn
  1. Where your tool's knowledge stops, and whether it tells you. Note 04
  2. What it can and cannot do with letters and digits. Note 02
  3. Whether it gives the same answer twice, and when that matters. Note 03
  4. The difference between a likely answer and a true one, seen once for yourself. Note 01
Afterwards you have

A card with four scores and a one-line rule for each, specific to your tool, dated. Run it again when the tool updates and compare. Run it on a second tool and you have a comparison nobody's marketing will give you.

How to do it

Open a chat tool in one window and this page in another. Work down the twelve tests; each has the exact text to paste and what counts as a pass. Tick as you go. The card writes itself at the bottom. Nothing you type here is stored or sent anywhere.

Part 1 · The problemIn one sentence

Most people cannot say what their chat tool is reliable for, so they either trust everything or check everything. Surveys in 2026 put it plainly: 68% of trained employees have no benchmark for what a good answer looks like, and executives spend about four hours a week checking AI output. A benchmark is what this build makes, at the size of one card.

Part 2 · The buildFour notes, twelve tests, one card

Note 04 · cutoffNote 02 · tokensNote 03 · temperatureNote 01 · prediction 1a1b1c2a2b2c3a3b3c4a4b4c MY TOOL CARD dated factsletters & digitssame answer twicelikely vs true _ / 3_ / 3_ / 3_ / 3 three tests per note. a score per row. a rule per score.
Each note in the section becomes three tests. Each row of tests becomes one score and one rule on the card.

Part 3 · The benchThe twelve tests

Search off for all of them, so you are testing the model and not the product's bolt-on. Use a new chat where it says so. Paste exactly; the answers to check against are given. Tick the box for a pass.

Build 01 · the testsRuns here · your chat tool for the model step

Row 1 · Dated facts Where its knowledge stops Note 04

Row 2 · Letters & digits What it never sees Note 02

Row 3 · Same answer twice The dial at its default Note 03

Row 4 · Likely vs true Prediction, seen once Note 01

twelve pastes, twelve ticks, one card. the card is the build; the page only keeps score.

Part 4 · What you should noticeAnd why it happens

Row 1 usually gives a cutoff a year or more old, and a hedge on the Secretary-General but a straight face on the newest model. That is Note 04: the hedge is trained around people and dates, not around the model's own generation.

Row 2 is where the score varies most between tools. Counting and reversing are letter jobs done from memory of pieces; the multiplication is done on chunks of three digits. A tool that passes 2c cleanly has almost certainly run code behind the curtain.

Row 3 shows the dial at the product's default. Labels usually match; definitions usually agree; names usually differ. If your names came back identical twice, your tool runs cold, and that is useful to know before you ask it for ideas.

Row 4 is the section in miniature. 4a spreads because the list is flat, 4b holds because the list is peaked, and 4c is the one to remember: a confident wrong answer to a one-answer question, produced by the same loop that got the boiling point right. Section 2 begins there.

Part 4 · what to doThe rules on the card

The card writes one rule per row from your scores. They are the same four rules whatever the numbers; the numbers decide how hard to apply them.

  1. Dated facts: for anything that could have changed after the cutoff, switch search on and check it actually ran, or paste the current text in. A low score here means "always".
  2. Letters and digits: ask it to spell or step things out first; do sums elsewhere. A low score means the tool has no code behind the curtain, so nothing with digits goes to it unchecked.
  3. Same answer twice: for anything you will compare or test, set temperature 0 if you can, or run twice and reconcile. A low score means the dial is high and that habit is not optional.
  4. Likely versus true: for one-answer questions, ask twice; two different answers means verify. A fail on 4c is the reminder that agreement is not truth.

Part 5 · Where it breaksAnd which note explains it

  • The product helps the model without saying so: code for the sum, a hidden search for the fact. The score flatters the model.→ Note 04, and Note 12 later. Watch for a "running code" or "searching" flicker and note it on the card.
  • Twelve tests is a small sample. One pass or fail can be the dial, not the tool.→ Note 03. If a result surprises you, run that one test three times before you believe it.
  • The tool updates and the card goes stale without telling you.→ Note 04. The card is dated for a reason; re-run it when the model version changes.

Part 6 · What it costs

40 minonce, per tool
0anything to install or pay for
1 cardthat replaces "I think it's usually right"

Record

paste your card into the margin of your own notes, dated. run it again in three months.