How to run demos and reference calls that surface the truth about a system — scripting them around your own scenarios so a vendor cannot hide a lot-number field, an oversold integration, or a cost that only appears after you sign.
Chef Diego runs a real food plant. If this page didn't get you there, tell us — a person reads every message.
After this lesson you can put every finalist through the same test — your
scenarios, run live on their software — so you buy on what a system does, not on
what a demo is built to show. You will know the questions that expose a lot-number
field pretending to be traceability, how to make an integration prove itself, how
to check references your own size, and how to get the whole first-year cost in
writing before you sign.
Run the demo on your scenarios, not theirs
A demo you did not script is a demo the vendor scripted. They open on the feature
they are proudest of, keep the pace up through the thin spots, and steer toward the
questions their software answers well. You leave impressed and none the wiser about
the things that matter to you.
You already have the fix from the last lesson: your
one-page requirements list.
Pull the two or three hardest things on it — the must-haves you cannot run without
and the exceptions where generic tools break — and hand them to every finalist
before the demo. Then make them run those, live, on their own software, ideally
against a slice of your real data rather than the tidy sample set that ships with
the demo.
Watch for the ways a vendor slips the test:
Switching from the live product to slides the moment your scenario gets specific.
"We would configure that for you" — a promise about the future standing in for a
thing working now.
Demo data arranged so the hard case never comes up.
None of these are proof. The only proof is your scenario, running, in front of you.
Everything below is a scenario worth scripting.
The scenario that exposes a lot-number field
The first scenario to script is a live trace, because it is the one most systems
fake. Almost every tool has a place to write a lot number — a box on a form, a
column in a list — and a box that holds a number looks like traceability until you
ask it to do the work. The difference is whether the number is connected to
anything. Storing a lot number tells you what came in the door; it does not tell you
where that material went. Real traceability is , and you only find out whether a system has it by making it
walk.
So in the demo, hand them one finished lot and ask, out loud, for both directions:
Backward: show me every ingredient lot and every supplier behind this finished
lot.
Forward: show me every customer shipment this lot went into.
Then the recall question: one of those supplier lots is bad — show me, right now,
every finished lot and every customer it touched.
A system that truly holds genealogy answers in seconds, on screen. A lot-number
field answers by opening a report where you type in a number and read back where it
was received — and then stops, because there is no forward chain to follow, or it
hands you an export to reassemble in a spreadsheet by hand. That pause is the tell.
The lesson on
why a lot-number field is not traceability
walks the same distinction in depth; here it is your sharpest demo question.
Put a lot on hold, then try to ship it
The second scenario to script is a QA hold, because it tests something a lot-number
field cannot do at all: stop product from moving. Ask them to place one lot on a
, and then, on the same screen, try to ship it, pick it for an
order, and pull it into a batch.
What you are watching for is a hard block — the system refuses, on every path, until
someone with authority releases the lot. A tool that only stores statuses will let
the held lot through with a warning at most, or with nothing at all, because the
status is a note rather than a rule. Ask the follow-ups: who is allowed to release a
hold, and does the system record who released it and why? The lesson on
holding and releasing a batch
lays out the states behind this; in the demo, the whole test is whether the block
actually holds.
Make every integration prove itself
"We integrate with that" is the easiest sentence a vendor says and the hardest one
to hold them to. A logo on a slide, a line on a feature list, a nod in the demo —
none of it tells you what actually moves between the two systems, or whether it
moves at all.
So treat every connection on your requirements list as a claim to be demonstrated,
not a box to be checked. For each one, make them show it running and answer plainly:
What exactly syncs — orders, invoices, inventory levels, which fields?
Which direction does it flow, and is it both ways or only one?
How often — live, hourly, a nightly batch, or a manual export someone clicks?
Is it built into the system, or a separate connector you pay a third party for?
Who fixes it when it breaks, and how do you even know it broke?
The distance between how an integration is sold and what it does is where careful
operators still get burned. A connection described as simple and automatic turns out
to be a one-way export that runs overnight, or a paid add-on nobody mentioned, or a
feature that is "on the roadmap" rather than in the product. Ask for the version
that exists today, working, on your data — and if the honest answer is "we would
build that for you," write down that it does not exist yet and price it accordingly.
The deeper danger is not any single broken connection. It is buying a stack of tools
that each half-connect and leaving your team to carry the data across the gaps by
hand. That trap is big enough to be its own lesson, and it is the next one.
Call references your own size
Every vendor has a reference list, and every reference on it is happy — that is why
they are on the list. Your job is not to confirm that happy customers exist. It is
to find out what your first year will actually feel like, from someone close enough
to your operation that their answer transfers to yours.
So ask for two references at roughly your size, in something like your category — a
maker your scale, not the vendor's biggest logo. Then get past the highlight reel
with specific questions:
How long did onboarding really take, from signing to the day the system was
actually running your work?
What broke at go-live, and what did you have to fix yourself?
When you are stuck, what does support actually do — answer, or tell you to figure
it out?
What did the first year cost, all in, versus what you were quoted?
Knowing what you know now, would you buy it again — and what would you do
differently?
Ask them about the exact scenario you scripted, too: when you trace a lot, does it
do what you needed? Two honest conversations with operations your size tell you more
than a page of testimonials, because they are answering your questions instead of
the vendor's.
Get the whole first year in writing
The number on the quote is rarely the number you pay. Around it sit the charges that
only surface once you are committed: a fee per user, an onboarding or implementation
charge, a data-migration cost, training, a higher support tier, and each add-on
module you need to do the job you are buying the system for.
Before you compare prices, make every finalist put the total first-year cost in
writing — subscription, per-seat charges, onboarding, migration, training, support,
and every add-on on your must-have list. Then read it for two traps:
The add-on tax: a capability you listed as non-negotiable, priced as an extra. If
real traceability or the connection to your books costs more on top, that cost is
part of the price, not a footnote.
Pricing you cannot control: a fee that scales on your revenue or your sales volume
rather than on what you actually use. A system that costs more simply because you
sold more can quietly outgrow the value it gives back.
Getting it in writing does two things. It lets you compare systems on the same line,
and it turns a month-three surprise from a misunderstanding into a broken promise
you can point to.
What a real test leaves you with
Work a process like this and the finalists sort themselves out. You have watched each
one trace a lot in both directions and block a hold, live, on your own data. You have
made every integration prove what it actually moves. You have heard from two
operations your own size what the first year really costs in time and money, and you
have that cost in writing. You are no longer choosing the system that demoed best —
you are choosing the one that fits best, which is a different and much safer thing.
One question outlasts the purchase. No system does everything, so whatever you buy
will connect to something else — and how you handle those seams decides whether your
team spends its days moving data across them by hand. That is the next lesson.