Photos Located by Internet Sleuths: The Method
Long before a model could guess a location in seconds, forum threads did it over months, from ridge profiles, pylon designs and the angle of a shadow.
Short answer
Photos located by internet sleuths are pictures whose setting a forum crowd worked out from visible detail alone: skyline profiles matched against topographic maps, pylon designs, plant species and shadow angles. The method is slow, collective and checkable. Raven compresses the same first pass into seconds, yet cannot verify a result the way a crowd can.

A photograph of a snowy ridge. No caption, no coordinates, nothing written on the back. Within a week the forum thread under it runs to a hundred replies. One contributor has narrowed the conifers in the foreground to a species that only grows above a certain altitude in that part of Europe. Another has traced the outline of the far peaks onto tracing paper and started sliding it along a contour map. A third recognises the lattice pattern of a distant pylon as one used by a single national grid operator. Eventually somebody posts coordinates and a matching view for comparison, and the thread quietly closes. Nobody involved has ever set foot in the valley.
That is the human precedent for every automated guess. Long before a vision model could produce a country name in a few seconds, groups of strangers were doing the same job by hand, slowly, and doing it well enough that the results held up. The interesting part is not the answers. It is the method, and what the method reveals about where a model is strong and where it is hopeless.
Why do crowds beat individual experts?
Because a photograph almost never yields to one specialism. One person reads pylon geometry, another identifies the tree line, a third owns the right map series and a fourth speaks the language on a half-visible sign. A crowd runs those searches in parallel; one expert runs them one after another, or not at all.
Expertise in this field is unusually narrow and unusually deep. There are people who can date a stretch of overhead power line from the shape of its insulators, people who know which municipality paints its kerbstones in that particular pattern, people who can tell one national mapping agency's grid conventions from another's at a glance. None of them can do the others' work. Individually they are hobbyists with a strange party trick. Assembled around one image, they become something closer to a research team.
The parallelism matters more than the raw expertise. A lone investigator has to decide which line of enquiry to spend the evening on, and a wrong choice costs a week. A thread does not choose. Ten people chase ten hypotheses at once, and nine dead ends get reported as dead ends, which is itself progress. That is why the same clue families keep recurring in these threads: the ones covered in our field notes on utility poles and their national styles and on what vegetation says about climate and latitude are exactly the details that different people happen to know about.
What does the method actually look like?
It is a loop rather than a flash of insight. Inventory every visible feature, group them by how much of the map each one removes, propose a candidate region, then hunt for a second image of the same place from a different angle. Confirmation only counts when two independent sources agree.
- Inventory first. Every visible feature gets written down, including the boring ones: kerb profile, guardrail type, the colour of a distant roof, the direction the shadows fall.
- Rank by exclusion power. A brand of soft drink removes nothing. A road-marking convention removes half the world. Work from the strongest filter down.
- Propose a region, not a point. Early guesses are deliberately coarse. Narrowing too fast is the classic amateur error, because it commits the whole thread to one wrong basin.
- Look for a second source. Aerial imagery, an old postcard, a tourist snapshot from a different angle. One image alone is a hypothesis; two agreeing images are a location.
- Publish the reasoning. A result without a visible chain of evidence is worthless to the thread, because nobody else can check it or reuse the technique.
Read that list back and it is recognisably the same procedure a researcher at an open-source investigation desk follows, minus the deadline. The amateur version and the professional one converged because the constraints are identical: no metadata, no access to the scene, and only the pixels to argue from.
Matching a skyline against a contour map
Terrain profile matching is the single most satisfying technique in the whole discipline, and the most laborious. A mountain skyline is effectively a signature. The relative heights of the peaks, the notches between them and the angle of each slope are fixed by geology and will not change in a human lifetime. Given a photograph, you can extract that outline, then test it against elevation data until the shape lines up.
Doing it by hand means working from a topographic map with contour intervals of 10 to 20 metres, picking a plausible viewpoint, and reconstructing what the horizon would look like from there. Get the viewpoint wrong by a few hundred metres and the peaks shift out of alignment, so the process is a grind of guess and check. Software now automates much of it, but the logic is unchanged, and the underlying insight is worth holding onto: landscape is the most stable clue there is. Signage changes, cars change, paint fades. A ridge holds its outline for thousands of years, so a skyline photographed 130 years ago still matches the view today, which is why terrain, snow and soil reward attention even when everything human in the frame is ambiguous.
Shadows do similar work with less glamour. The direction and length of a shadow constrain the sun's position, which constrains latitude and time of year together. It is rarely decisive on its own, but it can cheerfully eliminate an entire hemisphere, and the reasoning behind that is set out in our note on how weather and light hint at latitude.
How are photos located by internet sleuths verified?
By re-photographing the place, in effect. A claim only stands once someone matches the frame against independent imagery of the same spot, from a different source and ideally a different angle, and the alignment survives scrutiny from people actively trying to break it.
Verification is the part that separates this culture from guessing, and it is the part a model cannot participate in. The standard set by outlets like Bellingcat, founded in 2014, is that a location claim comes with a visible comparison: the original frame beside satellite or street-level imagery, with the matching features annotated. Anyone can then attack the match. Most published claims survive that; a meaningful minority do not, and get retracted in public, which is what keeps the standard honest. The same discipline applies to a family photograph, and we walk through a lighter version of it in our guide on how to verify where a photo was taken.
What a model does better, and much worse
A vision model's advantage is breadth, delivered instantly. It has seen imagery from everywhere and holds no regional bias, so it will happily suggest a province you have never heard of when the evidence points there. A human crowd, by contrast, is shaped by who happens to be awake in the thread. If nobody present knows West Africa, the thread will circle Portugal for a fortnight. Raven produces its first hypothesis in a few seconds, and as an opening move that is genuinely useful.
The weaknesses are just as structural. A model has no persistence: it cannot come back tomorrow with a new idea, cannot go and find a second photograph, cannot phone a friend who knows the region. It cannot check itself against a map, and it will not tell you that it is stuck, because being stuck is not something it can represent. Where a thread produces an argument you can audit, a model produces a claim you have to take or leave. The two failure profiles differ so sharply that they are worth studying side by side, which is the subject of why humans and AI guess locations differently.
What happens when the crowd gets it wrong?
People get hurt. The same energy that patiently places a mountain ridge has, during breaking news events, converged on named individuals who had done nothing whatsoever. Collective attention has no brakes and no editor, and an apology posted afterwards does not undo a week of accusations.
This culture has a genuinely dark side and it deserves to be stated without softening. Pointed at a landscape, an old print or an unlabelled postcard, crowd geolocation is a harmless and rather beautiful hobby. Pointed at a person, it becomes something else entirely. Internet vigilantism has a documented history of misidentifying innocent people during fast-moving news events, with consequences that fell on families who had no involvement at all. The techniques described here are the same techniques in both cases. The only thing that changes is the target, and the target is a choice.
Try the crowd's first move on your own photo: list every clue you can see, then upload it and compare notes.
Upload a photo →What the sleuths really demonstrate is that patience is a technique. Months of attention, distributed across people who each know one small thing, will beat any single reading of an image. A model borrows the breadth without the patience, which makes it a fine place to start and a poor place to stop. Raven, or the free Geospy AI app on an iPhone, gives you the opening move; the rest of the work is still yours. If the long version of that story interests you, it runs through the history of photo geolocation from pencilled captions to pixels.
Frequently asked questions
- How long does a crowd usually take to place a photograph?
- Anywhere from twenty minutes to several months. Easy frames with legible signage fall quickly. A wooded ridge with no text in it can sit unsolved for a year until somebody with the right regional knowledge wanders into the thread.
- What makes a group better at this than one expert?
- Coverage. One contributor recognises Balkan lattice pylons, another knows which conifers grow at altitude in that band, a third owns the relevant map series. No single person holds all three, and a photograph rarely needs fewer than three.
- Can an AI model replace crowdsourced geolocation?
- No. A model gives a broad first guess in seconds, which is genuinely useful as a starting point. It cannot pull up a map series, cross-check a second photograph or argue with itself for three weeks, and those steps are where certainty actually comes from.
- Is this kind of investigation always harmless?
- No, and that is worth saying plainly. Aimed at a landscape or a historical print it is a hobby. Aimed at a person during a breaking news event it has repeatedly harmed people who turned out to have done nothing at all.
Sources
- Open-source intelligence — WikipediaThe umbrella term for research built from publicly available material; the phrase dates to intelligence work well before the internet, and now covers the amateur variety too.
- Bellingcat — WikipediaFounded in July 2014 by Eliot Higgins, and the outlet that turned collaborative photo and video verification into a documented, teachable method.
- Topographic map — WikipediaContour lines join points of equal elevation; national series commonly use intervals of 10 to 20 metres, which is fine enough to match a ridge profile against a photograph.
- Internet vigilantism — WikipediaThe failure mode of the same crowd energy, including several well-documented misidentifications of innocent people during breaking news events.
Reminder
Raven is built for entertainment and curiosity. Its guesses are AI estimates that can be wrong, and it must never be used to track or identify real people. Uploaded photos are processed in memory and immediately discarded — never stored.


