Customer Sent a Photo Instead of Details?
A customer sent a photo instead of details? Turn it into a quote fast. What AI reads reliably from an image, the extraction prompt, and the human checks.
A customer sends a photo of a leaking pipe and the message "how much?" No model number, no measurements, no address. You can reply with four questions and wait a day for answers, or you can get what you need out of the photo in thirty seconds and reply with a number.
A customer sent a photo instead of details, which is the normal case rather than the awkward one. AI is genuinely good at reading those photos now, and the gap between a useful setup and a useless one is mostly about what you ask it to extract and what you do when it is unsure.
What a customer photo tells you, and what details it leaves out
More than people expect, and less than the confident answer suggests.
Reliable: visible text, including model numbers, serial plates, meter readings, dimensions on a label, brand markings. Approximate count of discrete items. Obvious material and condition. Whether something is clearly broken.
Unreliable: measurements without a reference object in frame. Anything behind or under the subject. Age. Whether the part is genuine or a replacement. The cause of a fault as opposed to its appearance.
The split matters because the reliable list is what you should automate and the unreliable list is what you should ask the customer about. A system that tries to estimate pipe diameter from an unreferenced photo will give you a number, and the number will be wrong often enough to cost you a job.
The extraction prompt
Do not ask what is in the photo. Ask for the specific fields your quote needs.
This is a photo from a customer requesting a quote.
Extract only what you can actually see:
VISIBLE TEXT: every model number, serial number, brand name,
label or reading, transcribed exactly. Mark any character
you are unsure of with [?].
ITEM: what the main object is, in plain terms.
CONDITION: visible damage, wear, corrosion, leaks.
CONTEXT: where this appears to be (under a sink, external wall,
roof space), and anything visible that affects access.
SCALE: only if a known-size object is in frame. Name the
reference object. Otherwise write "no reference".
Then:
MISSING: the three most important things I would need to quote
this that are not visible in the photo.
CONFIDENCE: high / medium / low for each section above, and say
what is making it low.
Do not guess. "Not visible" is a complete answer.The MISSING section is the part that earns its keep. It turns the photo into a reply: you now know exactly which questions to ask, which is the difference between one round trip and four.
The explicit "do not guess" and the per-section confidence do real work with images. Vision models are noticeably more willing to invent plausible detail about a blurry photo than about ambiguous text, because the plausible detail is usually correct and so the habit is reinforced. Asking for confidence per section gives you somewhere to see the weakness.
Reading the serial number is the whole job, sometimes
For a lot of trades, the model number is the quote. Once you have it, you have the part, the price, and the labour time.
Transcription is where this goes wrong, and it goes wrong in predictable ways: 0 and O, 1 and I and l, 5 and S, 8 and B. A one-character error gives you a different part and a wrong price.
Two defences. First, the [?] marker in the prompt above, which makes the model flag characters it is unsure of rather than committing. Second, validate the result against your catalogue before quoting. If the extracted number does not exist in your parts list, the extraction is wrong, not the catalogue, and that check is a lookup you probably already have.
Where extraction is routinely uncertain, the pattern from reading receipts and invoices with AI applies directly: extract, validate against a known list, escalate the mismatches.
The reply template
Once the extraction is done, the reply almost writes itself, and it should go out fast. Speed is most of the advantage here.
Thanks for the photo. From what I can see it's a [item],
[condition], [access note].
To quote this properly I need:
1. [missing item 1]
2. [missing item 2]
3. [missing item 3]
Rough range based on the photo: [X] to [Y], firmed up once
I have the above.The range matters. Customers who send a photo and ask "how much" want a number, and a reply that is only questions reads as a brush-off. A range with a stated reason for being a range is a real answer, and it buys you the information you need without losing the job to whoever replied with a figure.
Where to put a human in the loop
Three rules, learned the expensive way by anyone who has automated this.
Any quote above your threshold gets a human look at the photo. Set the threshold at the amount you would be unhappy to absorb if the extraction was wrong.
Low confidence on the ITEM field stops the automation entirely. If the model is not sure what the thing is, nothing downstream is worth anything.
No reference object means no measurement in the reply. Ever. State that you need a measurement instead of estimating one.
If you are running this at volume, the confidence field should drive routing rather than being advisory, which is the same discipline as setting a confidence threshold for an AI classifier: a number below the line goes to a person, automatically, without anyone deciding in the moment.
Fitting it into how enquiries arrive
Most of these photos arrive by WhatsApp, text, or a web form, which means the practical question is where the extraction runs. The simplest version is manual: you paste the photo into a chat tool with the prompt saved somewhere you can reach. That works fine up to a few enquiries a day, and it is the right place to start, because it tells you whether the extraction is good enough before you build anything.
Past that, the extraction belongs in whatever already receives the enquiry, with the output attached to the enquiry record so the person quoting sees it. The failure mode of the automated version is the same as an AI receptionist double-booking: the system acts on a confident reading nobody checked. Attaching the extraction to the record rather than acting on it keeps the human in place while still removing the work.
And when a customer's enquiry turns out to need a real conversation, handing off cleanly to a human matters more than the extraction did.
For related techniques on working from images rather than text, prompting AI with a screenshot covers the general case, and the AI for small business guide is the wider map.
FAQ
What if the photo is blurry or badly lit?
Ask for a better one, specifically. "Could you send a closer photo of the label on the side" gets a usable image; "could you send a better photo" gets another blurry one. The MISSING section usually tells you exactly what to ask for.
Can it estimate dimensions from a photo?
Only with a known-size object in frame, which is why the prompt asks for a named reference. A coin, a tape measure, a standard brick. Without one, any number it gives you is a guess dressed as a measurement.
Is it worth it for a one-person business?
Yes, and arguably more so, because the bottleneck is your time replying to enquiries. The manual version costs nothing to try: save the prompt, paste the photo, read the result.
What about photos with people in them?
Handle them the way you handle any customer photo containing personal information, and do not retain them beyond the job. Check what your tool does with uploaded images before you make it part of a routine.
How do I stop it inventing a model number?
The [?] marker plus validation against your parts catalogue. Those two together catch nearly all of it, because an invented number almost never matches a real catalogue entry.
How did this land?
About the author

Developer Advocate
Steve builds something with Swarmz every week and writes up what worked, what broke, and what he'd do differently. Tutorials and hands-on guides are his lane.


