Skip to content
Go back

Deterministic Scripts Before Prompts: Cheaper, More Reliable AI Workflows

The obvious move when you’re building an AI workflow is to hand the model more of it. More steps, more tool calls, more “figure it out.” I did that for a while. What actually made my setup faster, cheaper, and more reliable was the opposite: pulling the model out of every step that didn’t need judgment, and letting a plain shell script do it instead.

This didn’t start with AI

I didn’t learn this from working with models. I learned it years before, at my day job, where I’ve built up close to a hundred custom shell functions for the systems I work in daily. Every one of them starts with m, purely so they group together and tab-complete as one list instead of scattering across whatever else is on the system. m for mine :) A mhelp command lists all of them, grouped by area. A small sample of what’s in there:

mtriggerpipeline      Trigger a CI pipeline run
mfetchpipelinefail    Fetch and filter a CI job trace for failures,
                      token-efficient output for AI analysis
mdownloadartifact     Download a build artifact from a pipeline run
minstallbuild         Install a build onto a connected device
mpollstatus           Poll an external system for status on specific records
mgeneratereport       Generate a report from a run
manalyzereport        Analyze raw results from a report, print a summary

Every one of these exists for the same reason: the moment I catch myself running the same commands by hand twice, that’s a script, not a habit.

None of that had anything to do with AI when I built it. The rule was just: if a task has exactly one correct procedure, write the procedure once and stop re-deriving it. One of them, mfetchpipelinefail, turned out to matter a lot more once I started pointing Claude at the same environment.

Before: the model does the plumbing

Say a pipeline run fails and I want to know why. The naive version: ask Claude to pull the job trace from the CI system and figure out what broke. That trace can run thousands of lines, mostly setup output, dependency installs, noise that has nothing to do with the actual failure. The model reads all of it, or tries to, hunting for the handful of lines that actually matter.

That’s expensive and it’s not reliable. Every one of those tokens costs money for zero judgment, because fetching the trace and finding the lines with “fail” or “error” in them isn’t a decision, it’s a grep. Worse, on a long trace the model can lose the actual failure in the noise, or latch onto the wrong line entirely.

After: one script, model only where it matters

mfetchpipelinefail does the deterministic part: fetch the trace, filter it down to the lines that actually indicate a failure, discard the rest. I built this for myself originally, because reading four thousand lines of log by eye is miserable. It turned out to already be exactly the right shape for a model too.

mfetchpipelinefail <pipeline-id>

One command, same filtering logic every time, and the output is small enough that handing it to Claude costs almost nothing. Claude only gets called once the noise is gone, on the question that’s actually worth its judgment: given these specific failure lines, what broke and why. That’s a real question. It needs context and pattern recognition. Grepping four thousand lines down to five was never one.

This isn’t a one-off trick. My whole Claude Code setup runs on the same split. Reading TELOS.md, checking which mode is active, finding this week’s sprint file: none of that is a decision, so none of it goes through the model as a decision. It’s a deterministic read, every session, in the same order. The model shows up once the inputs are loaded and an actual judgment call is on the table.

The rule

Same call you make building evals: does this step need the model’s judgment, or does it already have one right answer? If it’s the second, script it and spend the tokens on the steps that actually need thinking.


Share this post on:

Previous Post
What a RAG Eval Gate Actually Catches
Next Post
Two Things I Do to Keep Claude Code's Context Cheap