← All posts

“No hallucinations” is the wrong promise

Ask these six questions instead.

The FieldTaskora team30 September 20269 min read

Every piece of field software now says it has AI. Most of the words in the pitch are real — RAG, agents, MCP — and each means something specific. Each also hides a question worth asking before you let it anywhere near a bid.

The promise you will hear most often is “no hallucinations”. It is the wrong one. Anything built on a language model can produce a fluent, confident answer that is wrong. The products worth trusting are the ones built for that, not the ones that deny it. So the useful question is not can it be wrong. It is what happens when it is.

Here are six questions to ask instead. For each one: the term in plain words, what good looks like, with an example, and the question to put to the vendor.

1. Where does the answer come from?

The term you will hear is RAG, short for retrieval-augmented generation. It sounds technical and the idea is simple: before the AI answers, it looks something up in your material, instead of answering from whatever it happened to learn in training.

From memory
You ask
What coating does the spec call for?
Model
Answers from what it learned in training
Answer
A two-coat system.Source: none
Looked up first
You ask
What coating does the spec call for?
Your documents
Finds the section in this job’s specification
Answer
A two-coat system.Source: spec section 2.4
RAG adds one step: look it up first. The part worth checking is the end — does the answer tell you where it looked?

Both answers in that picture read equally well. That is the problem with judging AI by how its answers sound: a wrong answer and a right one look the same on screen. The only difference you can check is the one on the end — whether it tells you where it looked.

What good looks like: the product reads your documents and shows you where — not “based on your files”, but a row, a section, a page. The best go one step further and check the citation before you ever see it.

The AI drafts
Questions about the bid, each naming where it came from
The check
Does that row or section exist in what we actually sent?
Yes
Shown, with its citation
No
Dropped. A made-up citation is treated as worse than none
The check runs before anyone sees the list. A dropped question never appears, so there is nothing to un-see.

An answer that cites something that was never in your documents should be dropped, not shown without its citation. Once you have trusted the first four answers on a list, an uncited fifth borrows their credibility — and that is the one that sends you hunting for a section that does not exist.

Worth asking: show me where that answer came from. And what happens to an answer that cannot show it?

2. Can it make up a number?

The risk is not a wrong number. It is a confident, well-formatted number that nobody can trace. In estimating that is the expensive kind, because it looks exactly like the right one until the job is under way.

What good looks like: every figure comes from your data, and the software says which — your price book, a job you won, a rule you set, or a number someone typed with a reason. Better still, the AI is not in a position to invent one at all. Here is what that looks like in a well-built voice quoting assistant, where a rep describes the job out loud:

The rep says
“Forty metres of trench, six hundred deep.”
What the AI can send
Descriptiontrench, 600 deep
Quantity40
Unitm
Rateno such field
Your price book
Excavate 600mm trench$48.00 / m × 40 = $1,920.00
The AI can describe the work. Only your price book can price it — the tool has no field for a rate, so there is nothing for the AI to fill in.

The design is the point. Telling an AI not to invent rates and hoping it listens is a policy. Giving it a tool with nowhere to put a rate is a control. The totals it reads back should come from the stored quote, not from the model’s own arithmetic.

Worth asking: can the AI put a number on a quote that did not come from our data?

Why a traceable number matters more than a right-looking one is the whole of When the estimate can’t explain itself.

3. What happens when the AI is down?

It will be, sometimes. A provider has an outage, an account runs out of credit, a network drops. None of that is unusual, and none of it is the question. The question is what you see when it happens.

What you want to see
  • It says, on screen, that the AI part is unavailable
  • Everything else on the page keeps working
  • Nothing is filled in on the AI’s behalf
What to watch for
  • A result that looks exactly like a normal one
  • A default, a template or a placeholder answer
  • No sign anything went wrong at all
Best read by ordinary code, AI or notQuantities on the bid formA bid’s due date
An AI feature will be unavailable sometimes. The difference between products is whether you can tell.

What good looks like: the product says so, plainly, and the rest of the page keeps working. Hard facts — the quantities on a bid form, a bid’s due date — are read by ordinary code rather than a model, so they do not depend on one being up. And the vendor checks that the AI is answering every time they release, so an outage shows up for them before it shows up for you.

Worth asking: can you show me this screen with the AI switched off?

4. Which model is it, and do you train on our data?

“Our proprietary AI” is not an answer. Almost every product is built on models from a handful of large providers, and there is nothing wrong with that — but you should be able to find out what each model is used for, where your information goes on the way, and whether any of it ends up training someone’s model.

Your browser or phone
You press a button, or start talking
The vendor’s servers
Hold the keys, decide what is sent, check what comes back
A reading model
Reads, checks and drafts
A realtime voice model
Live voice quoting
TrainingYour data is not used to train or fine-tune models, in writing
In a good setup every AI request goes through the vendor’s servers, which hold the keys. Your browser never talks to a model directly.

What good looks like: each model has a clear job — say, one that reads, checks and drafts, and a separate realtime model for voice. Every request goes through the vendor’s servers, which hold the keys and check what comes back, so your browser never talks to a model directly. Your data is not used to train or fine-tune models, and the providers are named in writing when you ask, usually on a sub-processor list.

Worth asking: which providers see our data, for what, and can we have that in writing?

5. Is it “agentic”? What can the agent actually touch?

An agent is AI that decides which tools to use to get a job done, rather than giving one answer and stopping. That is genuinely useful — it is how software can take a spoken description of a job and turn it into a priced quote. It is also exactly why the list of tools matters: an agent can do anything its tools can do, and nothing else.

What good looks like: a short list of named tools, clear limits on what each can do, and a log of what the agent did. For example, a voice quoting agent might need only four:

Example: a voice quoting agent — everything it can do
Search the catalogueLook up what you sell, without adding anything
Price a lineFrom your price book, never from the AI
Note the detailsWho the quote is for and where the work is
Read it backThe quote as it is stored
ToolIt asked forThe system answered
Search the catalogue“traffic control”Traffic management, per day
Price a linetrench, 600 deep · 40 m$48.00 / m · $1,920.00
Read it backthe running totalFrom the stored quote
A short, named toolkit is the whole list. Everything it writes lands on the quote being built. It cannot send anything, change a job or approve anything, because it has no tool that does.

If a vendor cannot list the tools, nobody has decided what the agent is allowed to touch. And the log matters as much as the list: every call recorded with what the agent asked for and what came back is the money trail if a quote is ever questioned.

Worth asking: list every tool the agent has, and show me the log from its last run.

6. How do you know it’s getting better, not just different?

The term is evals: a set of real cases with agreed answers, scored every time the model or a prompt changes. It is the least glamorous part of AI and the one that matters most over time, because the models underneath products change constantly.

Real case, with an agreed answerVersion AVersion B
Which coating system does section 2.4 require?✓ right✓ right
Is the bid due before or after the site meeting?✓ right✗ wrong
Which rows on the bid form have no quantity?✓ right✓ right
Does addendum 2 change the trench depth?✗ wrong✓ right
Which items are marked provisional?✓ right✗ wrong
Score4 of 53 of 5
Same cases, agreed answers, two versions of the model. Version B fixed one case and broke two. Without the set, that looks like an upgrade.

Without a set like that, a vendor that upgrades its model is changing your results and hoping. Version B in the picture would feel like an improvement to anyone who happened to try the fourth question, and it is worse.

What good looks like: the cases come from real work of the kind you do — your document types, the questions your trade actually asks — and each has an answer a person has agreed. The set is scored on every change, before the change reaches you, and a drop in the score stops it shipping.

Worth asking: when did you last change the model, and what did the scores do?

Two more terms you’ll hear

Vector database — “we match on meaning”. Ordinary search finds the words you typed. A vector database stores text in a way that lets software find passages that mean the same thing in different words. It is often the lookup step behind RAG.

You search for
“roof leak”
Word search
No results — the note never says “roof” or “leak”
Meaning search
Finds: “water coming through the ceiling in unit 4”
Word search needs the same words. Meaning search finds the same idea in different words.

Worth asking: when it matches on meaning, can you see what it matched? A meaning match is a judgement, and a judgement you cannot see is one you cannot check.

MCP — “it plugs into your Copilot”. The Model Context Protocol is a standard way for an AI assistant to connect to another system and read from it, or act in it. Instead of copying information between tools, you ask your assistant and it asks the system.

Your assistant
Copilot, ChatGPT, Claude
MCP
A standard connector between the assistant and another system
Another system
Jobs, quotes, a calendarCan it read?Can it change?Who decided?
MCP is the plug, not the permission. What the assistant can see and change through it is a separate decision — and the one to ask about.

Worth asking: what can the assistant see through it, what can it change, and who decided? It is the same question as the one about agents, one step removed.

The six questions, on one page

Demos are built to go well. The way through that is to ask to be shown, not told — each of these has an answer you can see on the screen in front of you.

AskA good answer shows you
1 Where does the answer come from?The row, section or page it used
2 Can it make up a number?Which of your data every figure came from
3 What happens when the AI is down?The screen with the AI switched off
4 Which model, and do you train on our data?Each model’s job, and “no” in writing
5 What can the agent touch?Every tool it has, and the log of its last run
6 Better, or just different?The scores before and after the last change
Take this into the demo. Ask to be shown, not told — demos are built to go well, and these are the things a demo does not usually show.

A product that can show you all six on the spot has been built for the day the AI is wrong. That is the one you want.

The honest summary

Good AI in field work cites where its answers come from, refuses what it cannot back up, says so when it is off, and leaves the pricing in your hands.

None of that is the same as never being wrong. Nothing built on a language model can promise that, and a vendor who does is telling you they have not thought about what happens when it is.

Ask these six questions of every vendor. Including us.

Share thisLinkedInXEmail

Your crew runs the job. FieldTaskora writes the report.

See it on one of your own job types.