Many enquiries arrive not as a precise technical specification but as a photo of where a gate is to go, or a hand drawing with approximate dimensions. Traditionally that raw material has to be interpreted by hand by whoever prepares the quote.
What multimodal means
A multimodal AI does not work with text alone: it can interpret images directly, recognising marked dimensions, structural elements and visible features of the setting. That allows useful information to be pulled out of a photo or a drawing without anyone having to describe it in words first.
From raw material to a bill of materials
Once the relevant information has been extracted — approximate dimensions, type of structure, any visible constraints — the AI can propose a coherent bill of materials to refine, cutting the time between the enquiry arriving and a first draft being ready to discuss with the customer.
Help, not a substitute for checking
Automatic interpretation of an image remains an estimate to be confirmed: approximate dimensions in a photo do not replace a site visit or a proper measurement when the job calls for it. The value of multimodal AI is in speeding up the first step, not in removing the human check on the details that matter.