Skip to content
Try Raven →
All posts
CuriositiesBy the Raven team7 min read

Visual Deduction: From Sherlock Holmes to AI

Sherlock Holmes built a legend out of reading tiny details in ordinary objects. It turns out that's roughly how a modern AI vision model thinks too.

Short answer

Visual deduction is the practice of reading many small, individually weak details until they converge on one explanation. Sherlock Holmes did it with a pocket watch; a vision model does it with a photograph, weighing signage, plants, road markings and light. Both produce a best explanation, not a proof.

Abstract constellation of faint dotted points connected by thin lines against a dark background.

In one of the most famous scenes in detective fiction, Sherlock Holmes picks up Dr Watson's brother's old pocket watch and, in a few sentences, reconstructs the man's entire decline: once well off, later careless with money, eventually undone by drink. He never met him. He read it off the case — scratches around the keyhole from an unsteady hand, pawnbroker's numbers scored into the metal, initials worn thin by handling. It is a party trick dressed up as science, and readers have loved it since The Sign of the Four appeared in 1890.

It is also, oddly, a fair description of what happens when a modern vision model looks at an ordinary photograph and tries to work out where it was taken. No single detail decides anything. Stack enough of them together — a sign's typeface, the shape of a socket, the colour of the foliage — and a confident answer emerges from what looked at first like nothing at all.

What is visual deduction, actually?

It is accumulation rather than revelation. You gather every small, individually weak observation a scene offers, then let them converge on the explanation that survives all of them at once. Each detail rules something out, and the field narrows until one story is left standing.

Holmes was fond of insisting he never guessed — that guessing was a shocking habit, destructive to the logical faculty. What he actually did was closer to accumulation than to logic: collect everything, however trivial, then find the account that fits the whole pile. The watch scene works because each scratch eliminates a possibility and each elimination narrows the field, until only one explanation remains consistent with all the evidence at once.

That is a real cognitive method and it long predates Arthur Conan Doyle. Doctors call it differential diagnosis. Forensic examiners call it trace evidence analysis. What the stories dramatised was the idea that overlooked details are often more diagnostic than obvious ones — that the label in a coat or the callus on a hand says more than a person's own account of themselves. The same hierarchy applies to photographs. The subject the photographer chose is usually the least informative thing in the frame.

Is it deduction, or something else?

Strictly, it is abduction: inference to the best explanation. Several stories could explain scratches around a keyhole. Holmes does not eliminate them all; he picks the overwhelmingly likely one and states it with total certainty. That certainty is a literary device rather than a property of the method.

The technical name for what Holmes performs is abductive reasoning, not deduction. A deduction guarantees its conclusion from its premises; an abduction proposes the explanation that best accounts for the observations and remains open to being wrong. Doyle's version omits the second half, because a detective who says "probably Bristol, though possibly Cardiff" makes for a worse story. That omission is precisely the gap between how automated systems actually work and how they tend to be described.

A vision model reasoning about a photograph is doing genuine probability weighing, not certainty. It has seen enormous numbers of images with known locations, and it has learned which visual patterns tend to occur alongside which regions. Faced with a new picture it is not retrieving a fact; it is estimating a distribution over possible answers and reporting the top of it. The fluency of the sentence it produces is not evidence about the strength of the underlying signal.

How does an AI read a photo the way Holmes read a watch?

By working the small stuff. A typeface narrows the field to a few national highway agencies, a socket shape rules out continents, vegetation implies a climate band, and the angle of light hints at latitude. None of it is proof; the overlap is what produces a specific answer.

This is easiest to see in a tool like Raven, where you upload a photograph and Google's Gemini model guesses where it was taken. Give it a street scene with no landmark anywhere in frame and it does not hunt for one perfect clue. It behaves like Holmes at the watch case:

  • A road sign's typeface narrows the field to the handful of countries whose highway agencies commissioned that specific letterform.
  • A socket or plug shape, barely visible at the edge of the frame, rules out whole continents on its own.
  • The colour and species of vegetation implies a climate band, which eliminates entire hemispheres.
  • Which side of the road the traffic uses is the cheapest filter of all: roughly 75 countries and territories drive on the left.
  • The angle and hardness of the light hints at latitude and time of year, tightening whatever is left.

None of these is proof by itself, the same way no single scratch on a watch case is proof. It is the accumulation — the same move, minus the deerstalker — that turns a pile of weak signals into a specific guess. What the model notices along the way is often stranger than what a person would think to mention, and we collected some of the odder cases in the weirdest things AI notices in ordinary photos.

Where does the comparison break down?

Holmes is always right because Doyle wrote him that way. Real inference is probabilistic and sometimes wrong, whether it comes from a detective, a doctor or a model. An honest system reports how confident it is instead of asserting; a guess that admits to being a guess is the more trustworthy one.

The fictional version also enjoys an advantage no real observer has: Doyle knew the answer before he wrote the clues. Working forwards from a scene rather than backwards from a solution is a much harder problem, and it fails in predictable ways. Interiors, tight crops, chain environments and heavy filters strip out exactly the legislated, standardised details that carry the geographic signal, which is why they defeat both people and models. The full catalogue of those hard cases is in what makes a photo hard for AI to geolocate.

There is a longer version of this story worth knowing too. Reading a location out of an unlabelled image is not a new idea invented by machine learning; it is what archivists, hobbyists and family historians have always done by hand, and the shift from pencilled captions to pixel-level inference is traced in a short history of photo geolocation.

How can you practise it yourself?

Cover the subject of a photograph and study the edges instead: the kerb, the socket, the bin, the doorknob, the lettering. Commit to a country before checking. The feedback loop is what builds the instinct, and guessing games supply it round after round.

The skill is trainable and the training is genuinely enjoyable. Guessing games drop you into an unfamiliar street and score how close you get, which is the fastest feedback loop available; a beginner's guide to geoguessing games is the shortest way in. It also works with children, since the reasoning is concrete and the reward is immediate, and there is a gentler set of options in family-friendly geography games for screen time.

Write down your own reading of a photo first, then upload it and see which clues the model picked up that you walked past.

Upload a photo →

Next time you look at an old photograph, try the exercise before reaching for anything. Ignore the subject entirely and look at the edges of the frame: the socket, the signage, the shape of a doorknob, the way the pavement is laid. You already have the instinct the stories claimed to invent. You have simply never had a reason to use it.

Frequently asked questions

Was Sherlock Holmes actually using deduction?
Not in the strict logical sense. His method is closer to abduction, or inference to the best explanation: picking the most probable story that fits every observation rather than proving it from premises.
Do AI models reason the same way a detective does?
The shape is similar and the mechanism is not. A model weighs learned statistical associations between visual patterns and places, then reports the most likely one. There is no chain of stated reasons underneath, only a probability estimate.
Which small details are actually worth noticing in a photo?
The ones a country legislates rather than chooses: which side of the road traffic uses, sign geometry, plug and socket shapes, number-plate proportions and the script on public signage. Fashion and global products tell you almost nothing.
Can I get better at reading photographs this way?
Yes, quickly. Cover the subject of a picture and study the edges instead, then check yourself against a tool or a guessing game. The feedback loop is what builds the instinct.

Sources

  1. The Sign of the FourWikipediaThe 1890 novel containing the pocket-watch scene, the most quoted demonstration of the method.
  2. Abductive reasoningWikipediaInference to the best explanation, the process Holmes actually performs while calling it deduction.
  3. Left- and right-hand trafficWikipediaRoughly 75 countries and territories drive on the left, which is why one glance at a road removes most of the map.

Reminder

Raven is built for entertainment and curiosity. Its guesses are AI estimates that can be wrong, and it must never be used to track or identify real people. Uploaded photos are processed in memory and immediately discarded — never stored.

Get Geospy AI for iPhoneDownload free