Every piece of field software now says it has AI. Most of the words in the pitch are real — RAG, agents, MCP — and each means something specific. Each also hides a question worth asking before you let it anywhere near a bid.
The promise you will hear most often is “no hallucinations”. It is the wrong one. Anything built on a language model can produce a fluent, confident answer that is wrong. The products worth trusting are the ones built for that, not the ones that deny it. So the useful question is not can it be wrong. It is what happens when it is.
Here are six questions to ask instead. For each one: the term in plain words, what good looks like, with an example, and the question to put to the vendor.
1. Where does the answer come from?
The term you will hear is RAG, short for retrieval-augmented generation. It sounds technical and the idea is simple: before the AI answers, it looks something up in your material, instead of answering from whatever it happened to learn in training.
Both answers in that picture read equally well. That is the problem with judging AI by how its answers sound: a wrong answer and a right one look the same on screen. The only difference you can check is the one on the end — whether it tells you where it looked.
What good looks like: the product reads your documents and shows you where — not “based on your files”, but a row, a section, a page. The best go one step further and check the citation before you ever see it.
An answer that cites something that was never in your documents should be dropped, not shown without its citation. Once you have trusted the first four answers on a list, an uncited fifth borrows their credibility — and that is the one that sends you hunting for a section that does not exist.
Worth asking: show me where that answer came from. And what happens to an answer that cannot show it?
2. Can it make up a number?
The risk is not a wrong number. It is a confident, well-formatted number that nobody can trace. In estimating that is the expensive kind, because it looks exactly like the right one until the job is under way.
What good looks like: every figure comes from your data, and the software says which — your price book, a job you won, a rule you set, or a number someone typed with a reason. Better still, the AI is not in a position to invent one at all. Here is what that looks like in a well-built voice quoting assistant, where a rep describes the job out loud:
The design is the point. Telling an AI not to invent rates and hoping it listens is a policy. Giving it a tool with nowhere to put a rate is a control. The totals it reads back should come from the stored quote, not from the model’s own arithmetic.
Worth asking: can the AI put a number on a quote that did not come from our data?
Why a traceable number matters more than a right-looking one is the whole of When the estimate can’t explain itself.
3. What happens when the AI is down?
It will be, sometimes. A provider has an outage, an account runs out of credit, a network drops. None of that is unusual, and none of it is the question. The question is what you see when it happens.
- It says, on screen, that the AI part is unavailable
- Everything else on the page keeps working
- Nothing is filled in on the AI’s behalf
- A result that looks exactly like a normal one
- A default, a template or a placeholder answer
- No sign anything went wrong at all
What good looks like: the product says so, plainly, and the rest of the page keeps working. Hard facts — the quantities on a bid form, a bid’s due date — are read by ordinary code rather than a model, so they do not depend on one being up. And the vendor checks that the AI is answering every time they release, so an outage shows up for them before it shows up for you.
Worth asking: can you show me this screen with the AI switched off?
4. Which model is it, and do you train on our data?
“Our proprietary AI” is not an answer. Almost every product is built on models from a handful of large providers, and there is nothing wrong with that — but you should be able to find out what each model is used for, where your information goes on the way, and whether any of it ends up training someone’s model.
What good looks like: each model has a clear job — say, one that reads, checks and drafts, and a separate realtime model for voice. Every request goes through the vendor’s servers, which hold the keys and check what comes back, so your browser never talks to a model directly. Your data is not used to train or fine-tune models, and the providers are named in writing when you ask, usually on a sub-processor list.
Worth asking: which providers see our data, for what, and can we have that in writing?
5. Is it “agentic”? What can the agent actually touch?
An agent is AI that decides which tools to use to get a job done, rather than giving one answer and stopping. That is genuinely useful — it is how software can take a spoken description of a job and turn it into a priced quote. It is also exactly why the list of tools matters: an agent can do anything its tools can do, and nothing else.
What good looks like: a short list of named tools, clear limits on what each can do, and a log of what the agent did. For example, a voice quoting agent might need only four:
If a vendor cannot list the tools, nobody has decided what the agent is allowed to touch. And the log matters as much as the list: every call recorded with what the agent asked for and what came back is the money trail if a quote is ever questioned.
Worth asking: list every tool the agent has, and show me the log from its last run.
6. How do you know it’s getting better, not just different?
The term is evals: a set of real cases with agreed answers, scored every time the model or a prompt changes. It is the least glamorous part of AI and the one that matters most over time, because the models underneath products change constantly.
Without a set like that, a vendor that upgrades its model is changing your results and hoping. Version B in the picture would feel like an improvement to anyone who happened to try the fourth question, and it is worse.
What good looks like: the cases come from real work of the kind you do — your document types, the questions your trade actually asks — and each has an answer a person has agreed. The set is scored on every change, before the change reaches you, and a drop in the score stops it shipping.
Worth asking: when did you last change the model, and what did the scores do?
Two more terms you’ll hear
Vector database — “we match on meaning”. Ordinary search finds the words you typed. A vector database stores text in a way that lets software find passages that mean the same thing in different words. It is often the lookup step behind RAG.
Worth asking: when it matches on meaning, can you see what it matched? A meaning match is a judgement, and a judgement you cannot see is one you cannot check.
MCP — “it plugs into your Copilot”. The Model Context Protocol is a standard way for an AI assistant to connect to another system and read from it, or act in it. Instead of copying information between tools, you ask your assistant and it asks the system.
Worth asking: what can the assistant see through it, what can it change, and who decided? It is the same question as the one about agents, one step removed.
The six questions, on one page
Demos are built to go well. The way through that is to ask to be shown, not told — each of these has an answer you can see on the screen in front of you.
A product that can show you all six on the spot has been built for the day the AI is wrong. That is the one you want.
The honest summary
Good AI in field work cites where its answers come from, refuses what it cannot back up, says so when it is off, and leaves the pricing in your hands.
None of that is the same as never being wrong. Nothing built on a language model can promise that, and a vendor who does is telling you they have not thought about what happens when it is.
Ask these six questions of every vendor. Including us.