The Best (and Worst) Photos to Upload for an Accurate Guess
Some photos give an AI model plenty to work with; others are almost deliberately unhelpful. Here's how to tell the difference before you upload.
Short answer
The best photos for AI geolocation are wide daylight frames showing several independent clues at once: architecture, vegetation, terrain and some legible signage. The worst are extreme close-ups, heavily filtered images and generic interiors, which are cropped or designed to look the same everywhere and leave almost nothing to reason from.

Not every photo gives a model the same amount to work with, and the gap between the helpful ones and the hopeless ones is wider than most people expect. Some frames are stacked with independent, corroborating evidence. Others are almost deliberately unhelpful, cropped down to the single detail that says the least. Knowing which is which makes the whole exercise more interesting, and it means fewer flat, unsatisfying results.
None of this is about gaming anything. A vague guess on a hard photo is not a failure, and nothing here is being scored. But if you are curious what actually gives a model its best shot at a specific answer, the pattern is clear enough to state plainly.
What makes the best photos for AI geolocation?
Breadth. A frame that shows more than one thing at once lets several independent clues corroborate each other, so architecture, vegetation, terrain, sky and signage can be cross-checked instead of one detail carrying the whole answer alone.
The strongest uploads, almost without exception, are the ones that show several things at the same time. A wide shot of a street, a hillside or a town square captures building style, plant life, sky and often some signage or infrastructure in a single frame. That matters because independent clues can check each other: if the roofline suggests one region and the vegetation suggests the same one, the answer tightens. If they disagree, the model has a reason to hedge, which is also useful.
- Wide environmental shots — a street, a hillside, a shoreline — showing architecture, terrain or vegetation alongside the sky.
- Natural daylight, which renders true colour and legible shadow direction far better than artificial or very low light.
- Unedited or lightly edited files, since heavy grading shifts the exact tones, soil colour, sky colour, foliage colour, that the reasoning depends on.
- A visible horizon or some sense of scale, which distinguishes distant mountains from a nearby rock face.
- Anything with text in it — a shop fascia, a street nameplate, a parked car's number plate — even at the edge of the frame.
Vegetation earns its place on that list because plants encode climate, and climate maps onto latitude. The Koppen climate classification, first published in 1884, was built on exactly that relationship, and a single palm, pine or eucalyptus in the background does a surprising amount of the same work. A parked car does something similar through its number plate, whose shape, colour and proportions differ by country even when the characters are illegible.
Which photos defeat any model, human or AI?
Extreme close-ups, heavy stylisation and generic interiors. A macro shot of a flower or a plate removes the environment entirely, filters strip the colour information, and chain interiors are built to look identical in every country.
At the other end, a handful of photo types reliably produce vaguer, more hedged results, and this is not a quirk of any one system. A person would struggle with them too. Macro photography is the clearest case: shooting at or near life size means the environment is excluded by definition, so a beautiful close-up of a flower, a plate of food or a stretch of pavement removes almost every geographic clue at once.
Heavily filtered or stylised photos are the second category. Strong colour grades, black-and-white conversion and aggressive cropping remove exactly the colour and context information that soil, sky and foliage clues depend on. The third is the generic interior. Hotel rooms, chain coffee shops and big-box retail floors are deliberately designed to look the same everywhere, which is excellent for brand consistency and terrible for placing a photograph. What is left in those frames tends to be small and electrical: socket shape, switch style, radiator design, window furniture.
Why do two photos of the same place differ so much?
Because clues are unevenly distributed across a scene. One frame catches a street nameplate and a distinctive kerb, the next catches a blank wall. Position and framing change the available evidence far more than subject or timing do.
This is the most practical thing to know. Two photographs taken 30 seconds apart, in the same street and the same light, can produce completely different answers because one of them happens to include a bus stop and the other does not. It is also why re-testing with a second frame is worth doing before concluding that a place is unguessable. The categories people most enjoy trying are catalogued in our look at the photo types people love testing AI with, and the pattern holds across all of them.
How should you frame a shot you intend to test?
Think evidence rather than composition. Include more of the scene than feels necessary, keep any visible text in frame even when it is not the subject, and when choosing between two similar shots take the one with more variety in it.
If you are deliberately trying for a specific answer rather than a regional one, photograph as though you were collecting evidence. Leave a little sky and a little ground around the subject. Do not crop out a sign, a plate or a noticeboard simply because it is not the focal point, since it is often doing more work than the subject is. And given a choice between two similar frames, take the one with more variety: a street with a building and some trees beats a street with only a wall, even when the wall is the better photograph. The full flow, from choosing a file to reading the result, is covered in our step-by-step walkthrough of using Raven.
Is a vague guess still a real answer?
Yes. A broad region honestly reflects a frame with thin evidence, and that is the correct behaviour. The suspicious result is a confident pinpoint drawn from a photo that visibly does not contain enough detail to support it.
Some photographs are simply hard, and that is fine. A close-up of a rock or a generic hotel corridor may only support a broad climate band, or an honest note that several regions fit equally well. That is an accurate description of how little the image gives away rather than a flaw in the analysis. The interesting range runs from a photo so specific it names a city to one so sparse it can only offer a continent, and the far end of that range is explored deliberately in our write-up of testing AI with nearly impossible photos.
There is a mirror image to all of this worth keeping in mind. Everything that makes a photo easy to place also makes it revealing when shared publicly, which is the subject of our piece on what your holiday photos might reveal. The same wide frame that produces a satisfying result is the one that tells a stranger the most.
A good experiment for the next time you are picking something to test: upload two versions of the same moment, one wide and one tightly cropped, and compare how the model talks about each. The difference in how the two results are worded is usually more instructive than either answer on its own. If you would rather run it from your camera roll while travelling, the free Geospy AI iPhone app does the same kind of analysis on iOS.
Upload a wide shot and a tight crop of the same scene and compare the two.
Upload a photo →The rule of thumb, if you want one sentence: width beats beauty. The wider the shot, the better the story it can tell, and the more honest the answer that comes back.
If you are staring at one particular photo and wondering whether it is worth uploading at all, where was this photo taken? covers what to check yourself first.
Frequently asked questions
- Why does a wide shot beat a beautiful close-up?
- Because independent clues can corroborate each other. A wide frame usually contains architecture, vegetation and sky together, so three weak signals combine into one reasonably specific answer. A close-up offers a single signal with nothing to check it against.
- Do black-and-white photos work?
- They work, but less well. Removing colour discards soil tone, foliage shade, roof-tile colour and sky colour in one step, and those are among the clues a model leans on when no signage is visible.
- Are indoor photos ever placeable?
- Sometimes. Plug sockets, light switches, radiators, window frames and any text on packaging can narrow a country. A plain, well-lit room from an international chain is genuinely close to unplaceable, which is by design rather than by accident.
- Is a broad, hedged answer a failure?
- No. A wide answer is the correct output for a photo with thin evidence. The result worth distrusting is a precise, confident pinpoint drawn from a frame that plainly does not contain enough to support it.
Sources
- Macro photography — WikipediaClose-up work at or near life size, which by definition excludes the surrounding environment.
- Koppen climate classification — WikipediaFirst published in 1884, the scheme that links vegetation to climate bands and so to broad latitude.
- Vehicle registration plate — WikipediaPlate shape, colour and layout differ by country, which is why a parked car at the edge of frame is worth keeping.
Reminder
Raven is built for entertainment and curiosity. Its guesses are AI estimates that can be wrong, and it must never be used to track or identify real people. Uploaded photos are processed in memory and immediately discarded — never stored.


