The Verifiability Test: Which of Your Workflows AI Will Actually Automate
Gartner expects 40% of enterprise applications to ship task-specific agents by the end of 2026. Sit in any boardroom and the question underneath that number is narrower and more personal. Which of my workflows does AI actually take, and which stay with my people. Most answers to that question are lists of tasks, and most of those lists are wrong within a quarter. There is a better predictor, and it has almost nothing to do with how clever the model is. It is whether the output can be checked.
Andrej Karpathy put the clean version of this at Sequoia's AI Ascent earlier in 2026. AI progresses fastest in domains where the output is verifiable. Where you can tell quickly, cheaply, and objectively whether the answer is right, AI improves fast and automates deeply. Where you cannot, it stalls, no matter how capable the underlying model becomes. That single idea is the most useful planning tool we have found for deciding where to point AI, and it turns a vague anxiety into a test you can actually run.
The verifiability test
Take any workflow and ask three questions about its output. Can you check whether it is right cheaply, without spending more effort verifying than the task saved. Can you check it quickly, in seconds or minutes rather than weeks. And can you check it objectively, against a clear standard rather than a matter of taste. The more yeses, the higher the workflow sits on the verifiability scale, and the sooner and more deeply AI takes it.
This is why code generation moved so fast. Code has a brutal, instant verifier, it runs or it does not, the tests pass or they fail. The feedback is cheap, fast, and objective, so the loop tightens and the capability compounds. The same shape explains invoice reconciliation, data extraction you can spot-check against a source, and any workflow where being right has a definition.
The verifiability scale: workflows whose output is cheap, fast, and objective to check automate first. The harder the output is to verify, the longer a human stays in the loop.
The high-verifiability end
At the top of the scale sit the workflows AI is taking right now, and the common thread is a cheap, fast, objective check. Reconciling a payment against an invoice, where right means the numbers tie. Extracting fields from a document, where right means they match the source. Resolving a tier-one support ticket against a known policy, where right means the policy was followed. Drafting code against a test suite, where right means it passes. In every case the check is faster than the work, so an agent can run, be graded, and improve inside a tight loop. These are the workflows on the fast end of the payback map, and the reason they pay back fast is the same reason they verify fast.
The low-verifiability end
At the bottom sit the workflows AI keeps not taking, despite years of confident predictions. Setting strategy, where right will not be known for a year and depends on a hundred things the agent cannot see. Brand judgement, where right is a matter of taste and reasonable people disagree. Genuinely novel research, where there is no answer key because the answer does not exist yet. The model can produce plausible output for all of these, but nobody can grade it cheaply, quickly, and objectively, so the loop never tightens and a human has to stay in it. The danger here is not that AI refuses these tasks. It is that it does them confidently and wrong, and the lack of a cheap check means nobody notices until it matters.
The middle is where the work is
Most enterprise workflows sit in the messy middle, and this is the part that turns the test from an observation into a strategy. A workflow in the middle can often be moved up the scale by engineering a verifier into it. You cannot check the whole output cheaply, so you define the part you can. You build a golden set of known-correct examples to grade against. You make the output reversible so a wrong answer is cheap to undo, which lowers the cost of being wrong even when you cannot prevent it. This is the same observe stage from the agent loop, made deliberate. Engineering verifiability is how you take a workflow that AI could not safely own and make it ownable.
Why this beats a list of tasks
The usual way companies decide where to use AI is a list of tasks someone read about. The verifiability test is better because it predicts rather than copies. It tells you why customer service automated before strategy, and it tells you what to do about the workflow in front of you that nobody has written a think-piece on. It also keeps you honest about the 5% problem. The pilots that fail to move the number are very often low-verifiability workflows dressed up as automatable ones, where the demo looked great because nobody had to check the output at scale. Run the test before the pilot and you avoid building a thing that cannot be graded.
The Greek-market angle
There is a specific advantage here for the kind of focused, owner-led firm that is common in Greece. Engineering a verifier into a workflow is a judgement call about what right means for your business, and that judgement is faster to make when the person who owns the workflow and the person who understands the AI can sit at the same table. In a large matrix, defining the check becomes a committee exercise. In a smaller firm, it is a conversation. The verifiability test rewards exactly the decisiveness that a flatter organisation can bring, and it lets you pick your first AI workflow on evidence rather than on whatever the loudest vendor is selling.
How to use the test on Monday
Take the workflow you are tempted to automate next. Ask the three questions. Can its output be checked cheaply, quickly, and objectively. If yes, you have found a strong candidate, and the rest is the payback and governance work. If no, ask whether you can engineer a verifier in, a partial check, a golden set, a reversible action. If you can, you have a candidate after some design work. If you genuinely cannot check the output at all, keep a human firmly in the loop and be very suspicious of any vendor promising to automate it. That suspicion will save you more money than any model upgrade.
We help enterprises run this test on their real workflows, find the ones that verify, and engineer verifiers into the ones in the middle so an agent can own them safely. The agents we ship (AI Customer Support, AI Contract-to-Cash, Enterprise AI Search and the rest of the product family) live on the high-verifiability end by design, which is why they pay back. If you want to know which of your workflows AI will actually take, get in touch at inbusiness.gr and we will run the test with you.