Six checks before
AI work leaves the building.
A grading rubric for AI output that anyone can run, whether or not they understand how the setup works. Ten files, six checks, one number at the end.
Somebody at your company set up AI for a real task. It worked on the first few. Now it runs every day and nobody grades it.
That's the normal state of things at a 5 to 50 person company. It isn't a competence problem. It's a structural one. The person who built the setup can't grade their own work, and nobody else knows what it was supposed to check for.
The expensive failures don't look like errors. The output reads fine. It goes to a client, a carrier, an adjuster, a buyer. They price it slightly lower than they would have, and nobody on your side ever learns why.
Here are six checks you can run yourself, today, without understanding how any of it works. Each one is answerable by a person who has never written a prompt.
The six checks
Run these against ten recent outputs. Not a benchmark. The last ten real ones you sent.
Every factual claim traces to a document you can open
Pick any sentence that states a fact. Ask the person who owns the output to show you where it came from. Not the reasoning. The source file.
Fails when the answer is "that's what it produced" or the source exists but nobody can find it in under a minute.
The citation actually supports the line it's attached to
This is the one that costs money. A reference gets attached to a sentence it only loosely supports. It isn't fabricated, so nothing looks wrong. It's just soft.
Open three citations at random and read what they actually say. Then read the sentence they're attached to.
Fails when the source is real, the sentence is reasonable, and the source doesn't say the thing the sentence claims.
Numbers reconcile across the whole document
Take any figure that appears more than once. A total, a date range, a count. Check that it's identical everywhere it appears, and that it matches the record it came from.
Fails when a number is right in one place and stale in another, usually because the draft was regenerated and one section wasn't.
Completeness is measured against a list that existed before the draft
"Does anything look missing" isn't a question anyone can answer. It has no failure mode. You need a list of what a complete output contains, written before you look at this one.
If no such list exists, that's the finding. Write it. It takes twenty minutes and it's the single highest-value artifact in this whole exercise.
Fails when the reviewer is checking against their own memory of what good looks like.
One named person owns this output type
Not the team. A person. Ask three people at the company who owns the quality of this specific output and see if you get the same name three times.
Fails when the answer is "whoever ran it that day." An output with no owner has no error rate, because nobody is counting.
A fixed share gets graded on a schedule
Somebody happening to notice isn't a control. Degradation on an ordinary Tuesday looks exactly like a slow week.
Pick a number. One in ten, one in twenty, whatever you'll actually sustain. Put it on a calendar. The number matters less than the schedule.
Fails when review happens after a complaint. By then the thing you're measuring already left the building.
How to turn this into a number
Run all six checks against ten recent outputs. Mark each check pass or fail per output. Sixty data points.
Count the failures and divide by sixty. That's your error rate. It's the number you didn't have before, and it's the only one that matters when you're deciding whether to trust the setup with more work.
A high number isn't a disaster. A number you can't produce at all is the actual problem, because it means every decision about your AI setup is being made on a feeling.
What this looks like in practice
A firm runs AI over a repeatable document that goes to an outside party who decides what it's worth. Ten recent files, six checks each.
Checks 1 and 3 pass almost everywhere. Sources exist, numbers mostly reconcile. Check 2 fails on four of ten, always the same way: a real reference attached to a sentence it doesn't quite support. Check 4 fails on all ten, because no completeness list exists. Check 5 returns three different names. Check 6 has never happened.
Error rate: 18 of 60, or 30 percent. But the useful finding isn't the number. It's that five of the six failures come from two missing artifacts, a completeness list and a named owner, and neither one requires touching the AI setup at all.
That's the usual shape. The output problem is real and the fix is upstream of the technology.
Where to start
Pick the row that describes you.
Run checks 1 and 3 on ten files
Cheapest possible start, no new artifacts required, and it tells you within an hour whether you have a small problem or a real one.
Start with check 5
Ask three people who owns it. If you get three answers, fix that before you measure anything else. Ownership is what makes every other check repeatable.
Make check 6 part of the scope
Ask any vendor how output gets graded after they leave. If the answer is training or documentation, they're handing you a thing that degrades quietly.
Why this is the first thing we do
Every install we run starts here, before anything gets built or connected. Not because it's impressive, but because a company that can't produce an error rate can't tell you whether the last thing it bought worked.
MarqueeOS installs the system that does this on a schedule instead of when someone remembers. The checks above are the judgment. The wiring is the part that takes weeks.
Keep reading
The 30-Minute AI Inventory: What Is Running In Your Business And Who Owns It
Four questions, four groups, one specific order. Surfaces the undocumented setup carrying real volume before the person running it resigns.
Read the guide → Decisions · 6 min readBuild It In-House Or Buy It: 7 Questions That Decide
The strongest predictor of this decision is question one, and it points away from hiring anyone. Includes the four cases where we're the wrong call.
Read the guide →Want the number for
your own output?
Thirty minutes. Bring one output type and we'll run the six checks against it live, on your files, and you keep whatever we find.
Book a 30-minute call →