mini-module ~90 min no-code try-without-ai reasoning

The Method of Holmes

What detectives, scientists, and machines have in common, and where they differ.

There is a famous line people quote when they want to sound precise: "Elementary, my dear Watson." Sherlock Holmes, the great logician, working from pure deduction. Except, almost nothing in that sentence is true.

The phrase never appears in Conan Doyle's original stories. And Holmes, far from being a pure deductionist, is mostly doing something else entirely. Before we can catch the error, we need the two concepts.

Sixteen portrayals of Sherlock Holmes across film and television

Sixteen portrayals of Sherlock Holmes, film, television, animation, 1916–2009


Two Ways of Reasoning

Method I

Deduction

Reasoning from general principles to a specific, certain conclusion. If the premises are true and the logic holds, the conclusion cannot be false.

All humans are mortal. Socrates is human. Therefore, Socrates is mortal.

Method II

Induction

Reasoning from specific observations to a general principle. The conclusion is probable, not guaranteed, but grows stronger with more evidence.

Every raven I saw was black. Therefore, all ravens are probably black.

The direction is what matters. Deduction goes downward, from the general to the specific. Induction goes upward, from specific cases, searching for the rule that would explain them.

Scientists mostly induce. Mathematicians mostly deduce. Most humans do both without noticing which they are doing.


Deduction in Language

The puzzles below shift the terrain. Instead of visual patterns, you are given data from three invented languages, languages that do not exist and cannot be searched for. Each encodes something real that natural languages do: quantity, time and agency, the source of knowledge.

This is what linguists actually do when they encounter an undocumented language. They induce the grammar from utterances. Then they use the induced rules to make precise, testable predictions.


Puzzle I, The Language of Moku

The Moku people live on a small island. Their language has three words for quantity: baí, tiró, and poná. These are not translations of English number words. They carve up quantity differently.

Below are ten observations. A Moku speaker is looking at a group of objects and uses one of the three words.

ObjectsCountMoku word
fish in a net1baí
coconuts on the ground3tiró
children in the water9poná
birds on the roof2baí
stones on the path7tiró
baskets in the hut15poná
canoes at the shore4tiró
stars visible tonight100poná
fires burning5tiró
men returning from fishing8poná
🔒

Enter the password to reveal the answers

Answers

Q1. baí = 1–2. tiró = 3–7. poná = 8 and above.

Q2. You can conclude there are between 3 and 7 objects. You cannot conclude the exact number, nor anything about what the objects are.

Q3. 6 falls within 3–7, so the speaker uses tiró.

Q4. The speaker was not wrong. At the moment of speaking, if they believed there were 3–7 objects, tiró was correct. Moku quantity words describe the speaker's assessment at time of utterance, not an objective count.

Q5. tiró is closest to "several." They differ because tiró has precise categorical boundaries (3–7) while "several" is vague and context-dependent.


Puzzle II, The Verbs of Voranu

Voranu is a synthetic language with a strict verb structure. Every Voranu verb is built from pieces in a fixed order. Deduce what each piece means, then translate sentences you have never seen.

Eight Voranu sentences with their English translations. Study them carefully.

VoranuEnglish
ka-en-voraI am speaking
mo-en-vorayou are speaking
ti-en-vorashe is speaking
ka-ol-voraI spoke
mo-avu-vorayou will speak
ka-en-lenuI am seeing
ti-ol-lenu-koshe saw you
ka-avu-lenu-riI will see her
🔒

Enter the password to reveal the answers

Answers

Q1:

ka- = I (subject) · mo- = you (subject) · ti- = she/he (subject)
-en- = present · -ol- = past · -avu- = future
vora = speak · lenu = see
-ko = you (object) · -ri = her/him (object) · -na = me (object)

Q2. mo-ol-lenu-na → "you saw me"

Q3. "she will speak" → ti-avu-vora

Q4. "I saw you" → ka-ol-lenu-ko

Q5. "you will give to her" → mo-avu-damu-ri. The assumption is that damu follows the same SUBJECT-TENSE-ROOT-OBJECT template. Reasonable, but still a hypothesis.


Puzzle III, What Velanic Speakers Know

In Velanic, every statement must include a suffix indicating how the speaker knows what they are saying. This is called evidentiality, found in many world languages, though not in English.

Three Velanic evidential suffixes. Their names are not given. Study the nine sentences below.

Velanic sentenceEnglish meaningSpeaker's situation
Nara velu-omThe river is risingStanding at the riverbank, watching the water level climb.
Nara velu-efThe river is risingInland, sky darkening, distant thunder. Has not seen the river.
Nara velu-aluThe river is risingJust told by a neighbour who came running from the riverbank.
Miko dara-omMiko is sickVisited Miko this morning and saw her lying in bed, pale and feverish.
Miko dara-efMiko is sickHas not seen Miko, but notices her untouched food bowl and her absence.
Miko dara-aluMiko is sickHeard it from Miko's sister at the market.
Tovo pari-omTovo left the villageWatched Tovo walk away down the path with a bag.
Tovo pari-efTovo left the villageFound Tovo's house empty, his fire cold, his boat gone from the shore.
Tovo pari-aluTovo left the villageThe speaker's child said Tovo had said goodbye before leaving.
🔒

Enter the password to reveal the answers

Answers

Q1. -om = direct evidence: personally witnessed or perceived. -ef = inferential: reasoning from observable evidence. -alu = reportative: received from another person.

Q2. Velanic forces speakers to grammatically commit to their source of knowledge every time they make a statement. You cannot simply say something is true, you must also say how you know.

Q3. Doctor directly examined the patient, -om. Journalist read a report, -alu. Child heard the wolf, -om, since hearing is direct sensory experience.

Q4. Holmes would use -ef (inferential), he reasoned to Afghanistan from physical clues, did not witness it. Watson, retelling as a direct witness, would use -om for what he personally saw.

Q5. The speaker was not lying. They accurately reported their source: they were told. Evidential suffixes encode the source and mode of knowledge, not the truth of the proposition.


The Detective's Method

Now read how Holmes actually explains his own thinking, from A Study in Scarlet, when Watson asks how Holmes could deduce he had been in Afghanistan:

"Here is a gentleman of a medical type, but with the air of a military man. Clearly an army doctor, then. He has just come from the tropics, for his face is dark, and that is not the natural tint of his skin, for his wrists are fair. He has undergone hardship and sickness, as his haggard face says clearly. His left arm has been injured. He holds it in a stiff and unnatural manner. Where in the tropics could an English army doctor have seen much hardship and got his arm wounded? Clearly in Afghanistan."

, Arthur Conan Doyle, A Study in Scarlet (1887)

Holmes calls this "deduction." Doyle calls it deduction. It has been called deduction for over a hundred years. There is a catch.

🔒

Think first, then enter the password to reveal the analysis

Is Holmes actually using deduction or induction? What direction does his reasoning move, from a general rule toward a conclusion, or from specific clues toward a general explanation?

Induction, mostly

Doyle got it wrong.

Holmes is not applying a known general rule to reach a certain conclusion. He is working in the opposite direction: observing specific clues and inferring the most probable general explanation. That is induction. Or more precisely, abductive reasoning: inference to the best explanation.

True deduction would look like this: "All army doctors returning from Afghanistan hold their arm stiffly. This man holds his arm stiffly. Therefore he is an army doctor from Afghanistan." That is deduction, and also bad logic. Holmes reasons from evidence toward the hypothesis, not from the hypothesis toward the evidence.

His conclusions are not certain, they are the most probable explanation given the data. He is a scientist, not a logician. He can be wrong. Occasionally, he is.

Want to see the method in action? The BBC series dramatises Holmes's inductive leaps better than most adaptations, the reasoning process is almost visible in slow motion.
▶ Sherlock (BBC, 2010–2017) on IMDb

Seeing the Rule

In 1967, Russian computer scientist Mikhail Bongard published a book containing 100 visual problems. Each problem consists of two groups of six images. Every image on the left satisfies a hidden rule. Every image on the right violates it. The task: find the rule.

Bongard problems are pure induction. No formula, no rulebook. You observe. You hypothesise. You check. You revise.

The ten puzzles below progress from transparent to subtle. For each one, write down the rule you think separates left from right.


Check Your Rules

You have worked through all ten puzzles. Below are the intended rules. A close paraphrase counts as correct. The exact wording does not matter; the concept does.

🔒

Enter the password to reveal all ten rules

  • 01
    Size

    Left images contain small shapes. Right images contain large shapes that nearly fill the frame. Shape type varies freely.

  • 02
    Containment

    On the left, the small solid shape sits inside the outlined large shape. On the right, the small shape is outside the boundary.

  • 03
    Quantity

    Each left image contains exactly two shapes. Each right image contains exactly three. Shape type is irrelevant.

  • 04
    Concavity

    Every left-side shape is concave, at least one inward dent. Every right-side shape is convex. Arrows, stars, crescents, crosses, and L-shapes are concave.

  • 05
    Topological hole

    Every left-side shape has a hole through it, a fully enclosed empty region. Every right-side shape is solid.

  • 06
    Bilateral symmetry

    Every left-side shape has at least one axis of reflective symmetry. Every right-side shape is asymmetric.

  • 07
    Count equals side-number

    The number of shapes equals the number of sides of each shape. Three triangles. Four squares. The right side always mismatches.

  • 08
    Strictly increasing size

    Three shapes whose sizes increase strictly left to right. On the right, the sizes do not increase monotonically.

  • 09
    Triple nesting

    Three shapes nested in a complete chain: smallest inside medium, medium inside largest. On the right, the nesting is incomplete.

  • 10
    Even total sides

    Count all straight edges across all shapes. On the left, this total is always even. On the right, always odd. Circles contribute 0.


What You Just Did

Each puzzle forced the same cognitive operation: observe examples, form a hypothesis, test it against counter-examples, revise, converge. That is induction.

The process felt more like sudden recognition than deliberate reasoning. That feeling of snap, of the rule becoming obvious, is what psychologists call the Aha moment.

The hardest abstractions demand that you operate on the relationships between relationships. In puzzle 9, you had to track three shapes and verify a complete containment chain. In puzzle 10, the rule required identifying every shape's type, counting its sides, summing, and checking parity.


From Bongard to Curriculum

You cannot teach someone to see a concept by explaining it. You teach them by showing enough examples that the pattern becomes undeniable, then withholding the last step so they must cross it themselves.

, paraphrasing Jerome Bruner, The Process of Education (1960)
  1. Start with transparent rules. The first puzzles snap into place immediately. The student learns what it feels like to have the right rule.
  2. Introduce relational properties. Later puzzles cannot be solved by looking at any single shape in isolation. The rule lives in the relationship between elements.
  3. Add irrelevant variation. Surface features change freely while the underlying rule stays constant. The student must learn to suppress attention to what varies.
  4. Move from geometric to topological. The hardest rules are about spatial arrangement, configurations that cannot be read off any individual element.
  5. Ask the student to generate, not just recognise. Design your own Bongard problem. This demands a theory of induction itself.

The Machine's Turn

In 2019, AI researcher François Chollet posed a question: if Bongard problems are trivially easy for humans and hard for computers, can we turn that gap into a benchmark? The result was ARC-AGI, Abstraction and Reasoning Corpus for Artificial General Intelligence.

The premise is almost identical to what you just did: observe input-output pairs, induce the rule, apply it to a new input. No instructions. No formula. Just examples.

"ARC can be used to measure a human-like form of general fluid intelligence, the ability to efficiently acquire new skills outside your training data."

, François Chollet, On the Measure of Intelligence (2019)
0%
Pure LLM score on ARC-AGI-2
~60%
Average human score
100%
Tasks solvable by at least 2 humans

Every ARC task has been verified to be solvable by at least two ordinary humans. State-of-the-art AI systems score in the single digits on ARC-AGI-2. The gap is not about knowledge, it is about forming and testing hypotheses from a handful of examples.

ARC-AGI-2, released in 2025, introduced challenges requiring compositional reasoning and contextual rule application, exactly the relational abstractions that Bongard problems begin to train.

Try It

Interactive · Browser-based · No account needed
▶ Play ARC-AGI tasks
Solve the same grid puzzles that AI systems struggle with. Study the input-output examples and figure out the rule.
Technical overview
ARC-AGI-2: the benchmark explained
What makes the 2025 version harder. Symbolic interpretation, compositional reasoning, contextual rules, and why log-linear scaling is not enough.
Live results
ARC Prize Leaderboard
See where current AI systems stand relative to human performance. The gap between them is the current frontier of intelligence research.

After you have tried a few tasks, return to the question you started with. Holmes looks at a tanned wrist and a stiff arm and concludes: Afghanistan. You look at three coloured grids and conclude: the rule is to rotate by 90 degrees. The cognitive operation is the same. The question ARC-AGI is asking is whether a machine can do what Holmes does, not store and retrieve, but reason.

That question does not yet have a satisfying answer.


Your Turn to Make the Puzzle

Solving a Bongard problem is one thing. Designing one is harder, and more instructive.

When you design one, you must think from the other direction: choose a rule, construct six images that satisfy it without giving it away too easily, and six more that violate it plausibly.

For next time

Design your own Bongard problem. Draw it on paper, or use any tool you like. The constraints:

  • Six images on the left that satisfy a hidden rule.
  • Six images on the right that violate it, but plausibly.
  • The rule must be stateable in one short sentence.
  • Test it on someone who does not know the rule. If they cannot find it, your positive examples are not consistent enough. If they find it immediately, your counter-examples are not misleading enough.

For inspiration, and nearly 300 more problems, visit Harry Foundalis's Bongard Problem Index.