Skip to content
Try Raven →
All posts
PlaybookBy the Raven team6 min read

Testing AI With Nearly Impossible Photos

We deliberately fed Raven the least informative photos possible — plain walls, food close-ups, hotel rooms — to find the real edges of what AI geolocation can do.

Short answer

The hardest photos for AI to place are the ones stripped of environment: a plain wall, a tight crop of food, feet in sand, a hotel room with the curtains drawn. Each removes a whole category of evidence, and the correct response to them is a wide, low-confidence answer rather than a false pinpoint.

Abstract dotted flight path arcing across a dark map-like background, faint passport-stamp-style circular marks, no text or landmarks.

Most demonstrations of AI photo geolocation lean on photographs that are, honestly, a little easy. A famous skyline. A road sign in a distinctive alphabet. A beach with obviously tropical vegetation. That is a reasonable way to show what a tool can do, but it says almost nothing about where the tool stops working. So we ran the opposite experiment: instead of generous photographs, we fed it the least informative ones we could find, and treated the results as data rather than as a demonstration.

Why are the hardest photos for AI to place the real test?

Because a photo stuffed with clues tells you nothing about the limit. Stripping the evidence away on purpose shows whether a model degrades honestly into uncertainty or invents a confident answer, and only one of those behaviours is safe to trust.

Any system looks impressive against a frame where architecture, signage, vegetation and road markings all point the same way. The interesting question is what happens once you remove all of it deliberately. A plain wall, a close crop of a plate of food, bare feet in sand, a hotel room with the curtains drawn: these are the photographic equivalent of a locked room. If a model can still say something useful, that tells you something real. If it cannot, that is equally useful, because it marks the actual boundary rather than a marketing-friendly guess at where the boundary might be.

What does the four-round experiment look like?

Four uploads, each removing one more category of evidence: a plain painted wall, a tight food crop, feet in sand, and a curtained hotel room. Run them in order and the answers should widen at every step.

Here is the version you can run yourself. The order matters, because the point is the trend rather than any single result.

  1. Round one: a plain wall. Nothing but paint, texture and a shadow. This is close to a true blank, with essentially no geographic content left in the frame at all.
  2. Round two: a close-up of food. A plate, a dish, perhaps a hand and a fork. Cuisine can occasionally narrow a region through plating style, a specific dish or packaging at the edge of frame, but a tight crop removes most of that on purpose.
  3. Round three: bare feet in sand. A warm climate and a beach are implied, but which beach on which continent is left open. Grain size and colour vary, though rarely enough to name a country from a close crop.
  4. Round four: a hotel room with the curtains drawn. Furniture, bedding and fittings can hint at a chain or an electrical standard, but a well-lit room from an international brand is built, almost by design, to look identical everywhere.

What actually happens as the evidence thins out?

The answers get visibly more cautious. Wide regional guesses and lower confidence replace specific place names, and in the emptiest frames the honest response is that there is not enough in the picture to say much at all.

Run the four rounds and the pattern is consistent. A plain wall or a tight food crop produces a broad, hedged, low-confidence answer, or a plain note that the frame does not contain enough to work from, rather than a false pinpoint. That is arguably the most valuable thing the experiment shows. A model that stays vague when the photograph is genuinely uninformative is behaving correctly, even though uncertainty makes for a dull demonstration.

It also sharpens the contrast with ordinary photographs. Once you have watched four uploads return nothing, the moment a single road sign or an unusual plant enters the frame feels different, and the way the answer snaps to a region is much easier to appreciate. That contrast is the whole reason to run the exercise in this order, and it is the same reasoning the step-by-step walkthrough of using Raven describes from the other direction.

Can anything be read from a hotel room?

Occasionally, from fittings rather than furnishings. Socket and switch shape, radiator type, window catches, and the proportions of printed paper on the desk are all standardised by country, and any one of them can pull a blank interior back to a region.

The fourth round is the most interesting failure, because it is not always a failure. Domestic electricity is standardised nationally, and there are roughly fifteen incompatible plug and socket types in use worldwide, so a single visible outlet is a real clue. Printed paper is another: ISO 216, published in 1975, sets A4 at 210 x 297 mm, noticeably taller and narrower than the Letter size used across North America, and a menu or a notepad on the desk shows that proportion clearly enough to matter. Neither clue is available when the room is tidy, the curtains are drawn and the desk is empty, which is precisely why that frame belongs in the experiment.

What does this reveal about how the model reasons?

That geolocation is pattern recognition over visible evidence and nothing more. There is no hidden channel of information, so when the evidence runs out the answer widens, and a suspiciously precise result from an empty frame should lower your trust rather than raise it.

The honest takeaway is that this is not magic pulling a location out of nothing. It is inference over whatever is actually in the frame, and no more than that. A confident, precise answer for a photograph of a blank wall should make you trust a tool less, because the image genuinely does not contain the information such an answer would require. Uncertainty is the correct response, and a system that gives it is easier to rely on when the photograph does carry evidence.

The flip side follows directly. Photographs that do contain evidence, a sign at the edge of the frame, a distinctive plant, an odd piece of street furniture, are workable in a way these four are not, and much of the pleasure of this kind of tool comes from ordinary-looking pictures that turn out to be quietly full of clues. Our catalogue of the photo types people love testing AI with runs through the categories that consistently surprise people.

Where is this experiment genuinely useful?

In an archive. Old prints and undated holiday photos are hard for the same structural reasons, so knowing how a thin frame behaves helps you judge which scanned images are worth testing and which will only ever return a continent.

Difficult photographs are not only a game. Anyone working through a box of unlabelled prints is dealing with the same problem, often with faded colour and a small crop on top of it, and calibrating your expectations first saves a lot of disappointment. Our guides to rediscovering forgotten trips in your camera roll and digitising and exploring old travel albums both assume that calibration, and the wider question of what this kind of analysis should and should not be used for is set out in our piece on the ethics of AI photo analysis.

Run round one yourself: upload the least informative photo you own.

Upload a photo →

If you try the four rounds, pay less attention to the place names than to the wording. The gap between a specific answer and an honest shrug is where the reasoning is visible, and it is the only part of this experiment that will still be interesting on the tenth photo.

Frequently asked questions

Why deliberately upload photos that cannot be placed?
Because easy photos only show what a tool can do at its best. Feeding it the least informative images available is the only way to find where the reasoning actually stops, which is the more useful number to know.
Is a wide, hedged answer a bad result?
No. For a frame that genuinely contains no geographic evidence, a wide answer is the accurate one. A precise, confident pinpoint from a blank wall would be the result to distrust.
Can a hotel room ever be narrowed down?
Sometimes, through small fixed details rather than decor: plug socket shape, light switch style, radiator design, window catch, and the paper size of anything printed on the desk. A curtained room from a global chain often defeats all of them.
Does the model use the file's GPS data when the image fails?
No. The guess is made from visible content only, and uploads are processed in memory and never stored. When the image carries nothing, the honest output is uncertainty rather than a fallback to metadata.

Sources

  1. AC power plugs and socketsWikipediaAround 15 incompatible domestic plug types are in use worldwide, which is why a socket in frame is a genuine country clue.
  2. ISO 216 paper sizesInternational Organization for StandardizationISO 216, published in 1975, defines A4 at 210 x 297 mm; North American Letter paper is shorter and wider, and the difference is visible in a photograph.
  3. SandWikipediaGrain composition varies from quartz to coral to weathered basalt, but similar sands occur on coastlines thousands of kilometres apart.

Reminder

Raven is built for entertainment and curiosity. Its guesses are AI estimates that can be wrong, and it must never be used to track or identify real people. Uploaded photos are processed in memory and immediately discarded — never stored.

Get Geospy AI for iPhoneDownload free