The AI Industry Has It All Wrong
Paul Itoi, founder of Stakwork, joins Intelligence Snacks to explain why AI agents are better at discovery than repetition. The conversation covers reliable enterprise AI, facto...
Latest Snacks from Episode 71

Turn Agent Discoveries into Repeatable Workflows
An agent can explore an unfamiliar task, make decisions and discover a route to an answer. That flexibility is valuable while the route is unknown. Once the task must be repeated, however, asking the agent to search from scratch wastes effort and reopens decisions that have already been made.
A workflow turns what the agent discovered into an ordered sequence of steps and checkpoints. It can ask the agent to complete one step, report back, then move to the next. The agent still handles the parts that require judgement, but the workflow fixes their place in the process and carries useful results into the next run.
Paul compares this to a robotaxi finding a house without a map. Letting the car explore until it reaches the destination is impressive the first time, but sending it out to wander the streets again every day makes little sense once the successful route is known. The agent should find the route and the workflow should make that route dependable enough to follow again.

Breaking Expert Work into Repeatable Steps
Paul describes this approach as factored cognition, a term coined by the research group Ought. The idea is to identify the smallest decisions an expert makes, then turn them into explicit workflow steps. Each step can be handled by the person, model or tool best suited to it, rather than asking one worker or agent to manage the entire problem.
Stakwork first applied this principle to data entry from nutrition labels. Asking one person to transcribe a whole label meant entering as many as 30 values. Instead, the team sliced each label into rows and showed a worker one row at a time. Accuracy rose sharply because the worker only had to capture one bounded item, such as protein or fat, rather than maintain attention across the entire label.
The same method reduced a specialist task in salmon farming to four or five observable questions. Graduate students had been watching videos to distinguish between several kinds of mites. When Stakwork interviewed them, their process became a short decision path, like, does the mite have six legs or eight? Is there a dark spot on its abdomen? Vision models could answer about half of these questions. The recorded decision path could then become a workflow, using tools where possible and falling back to a person when necessary.

Why AI Accuracy Doesn’t Equal Business Reliability
Paul argues that the useful test for enterprise AI is not whether a model can produce the right answer, but whether its work is dependable enough to use. An error may not matter in a Blender demo or a coding experiment at home. The standard changes when an attorney must rely on a memo in court or a car dealership must decide whether to refund the cost of a vehicle. In those settings, the model cannot be sometimes right. It has to be right.
A 30-page legal memo makes the distinction concrete. It may look complete while containing invented citations that could expose the lawyer to sanctions. If it arrives at 11 am for a noon deadline, asking the model to reflect on its answer does not make that risk disappear. The lawyer still has to verify whether the work is safe to use.
This gap between occasional accuracy and required reliability helps explain why apparent AI capability does not always produce a return for businesses. If someone must inspect or redo the output closely enough to catch every consequential error, the finished-looking answer has shifted the work rather than completed it. Paul therefore values the visible reasoning and intermediate artefacts, such as the spreadsheets and inflation schedules behind a damages calculation, over a final figure that amounts to “trust me.”