A High-Quality AI Dev Workflow

Plan, grill, handoff-ready specs, small slices, red/green tests, discard the plan. The seven-step workflow I used to ship Expressed — 113k lines of product code, not prompt roulette.

If you keep running into “AI writes slop” on social media — I have news for you.

It works. Really works. Not as a magic button (very darn close though), and not without guardrails — but yeah, it sure does work well when you use it with a plan, small slices, and tests that say when you’re done.

I built Expressed that way — reviewed it, re-architected it, iterated on it. A real product: desktop capture, a drawing overlay, live sharing into a browser, a recording pipeline, a listing in the Windows Store. Not a todo app. Not a weekend clone of something that already exists. I wrote this website the same way.

As of today, the Expressed product has this much stuff:

FilesLines
Apps (Windows host, Mac host, Relay)8910,169
Libraries48063,335
Tests29340,095
Product862113,599

None of that came from a single prompt (obviously?). It came from a loop — the good news, below — run many times.

The loop I ended up with

I did not open a chat window and type “build me a screen-sharing app” (I have tried stuff like that, um, it didn’t work out well). That is how you get slop.

The process started as an experiment and was refined as I went along… This is what survived:

  1. Install Matt Pocock’s skills — worth the setup time. Learn how & when to invoke them.
  2. Plan — know what I want; write a plan in Cursor’s plan mode. Keep the chunks small. Ask it to ask you questions, resolve ambiguity.
  3. Grill — /grill-with-docs: stress-test the plan, ask deeper questions, fold the answers back in, softly - like your baking a cake. Ground it in my architecture and product world, not just whatever plan mode produces on the first pass. Good apps/code have a recipe. This keeps decisions in a CONTEXT.md document and starts writing Architecture Decision Records (ADR).
  4. Make it handoff-ready — sometimes I ask: ensure a junior developer can implement this plan. Plenty of detail about where and why, not just what. The goal is a plan (and issues) that carry enough context for reliable execution — including on a cheaper or faster model.
  5. Split the work — small issues, reviewable slices. Not a four-hundred-file “please implement everything” prompt. Stuff that into GitHub actions, or in my case - .md files on disk.
  6. Execute — one slice at a time. Do you stuff an entire tomato into your face (in case you like tomatoes)? No — you slice it. See?
  7. Red/green unit tests — a failing test first, then only enough code to pass it. The output is not the definition of done; the tests are.
  8. Discard the plan — never store your plans for longer than you need them. The code is the documentation.

Claude, Cursor, Fable, Opus, Grok, Composer — whichever model fit the slice that day.

Things go wrong. The nice part is that iterating — ripping stuff out — is much quicker when you have a written plan and can regenerate against it.

Why the loop matters

Prompt in, product out — with no plan, no review, no tests — and you will get slop. Of course you will. You would get the same from a human you never questioned and never made write a test.

The workflow is what turns fast output into good output.

My secret sauce ingredients are:

  • A written plan — to give anchor points, references, classes, scope, files - and of course a goal - and use Claude or Cursor (or codex or something) that gives you a harness, Cursor for example does a lot of work to make it easy for the LLM to search your codebase.
  • /grill-with-docs — it’s hard to overestimate the value of planning before writing code. And having a CONTEXT.md. Oh, and ADRs. Did I mention those?
  • Tests — red/green so “done” is not whatever showed up last in the chat.
  • Small slices — so review stays human-sized.

The mental load is real. So much code appears so fast that keeping up is exhausting — even when you know what good looks like. Tests help because they turn “I think this works” into something checkable. A plan helps because ripping work out and redoing it is less painful when the destination is still written down.

I am not claiming this is easy. I am saying it works: real features can be built this way, and the loop is why. I have the product, the 291 days , the commits, and the scars from the bits I had to rip out and do again (I’m looking at you, MAUI and web front end).

I’m still asking myself if 291 days is a good or a bad sign. I have a production ready product to show for it I suppose :-) That’s product, infrastructure, multiple re-writes of various bits, website & design, m/soft store workflow, licensing, relay code + infra etc etc - e.g. productionised app.

Steal it if you want

If you are curious, start with one slice of real work — not a greenfield fantasy — and run steps 1–7 on it. See if the loop feels like engineering. My bet: it will.

Keep the process — plan, grill, make it handoff-ready, split, execute, red then green, discard the plan — and prompt roulette stops being the story. The loop is. And yeah — it works.

Your mileage will vary.

Good luck.