← Back to the blog AI agents and development

Plan first, then code: why an AI agent needs a workflow

AI can write code. But who decides what should happen next? MCP38 adds a work plan, checks and bounded repair loops so the outcome can be verified.

A work plan connects memory, implementation, a verification gate and a repair loop
Illustrative image created with AI.

An agent may know how to write code, repair a test and explain a bug. It can still be missing something ordinary: knowing what should happen next. ITHZ-MCP38 shifts part of the attention from individual answers to the course of an entire task.

Imagine renovating an apartment. An electrician can wire the sockets, a painter can finish the walls, and a carpenter can build the kitchen. Each is good at their trade. But if the painter arrives before the electrician, or you order the kitchen before taking measurements, individual skill will not fix the order of work.

In AI development, something similar can hide behind impressive answers. An agent changes a file. We ask it to run tests. They fail, so we remind it to make a repair. Later, we realise that the diff still needs review, documentation is missing, or nothing has been checked in the real environment. We keep deciding what comes next, even though the agent can often perform the individual steps.

That led to a question for the next ITHZ release: what if a new task began with a proposed way of working, before the first code edit?

A task also needs a route to completion

“Fix login” describes an objective. It does not yet explain how we will demonstrate that the fix works, which files may change, or whether deployment is authorised.

MCP38 adds development workflow planning. At the start of a nontrivial task, an agent can first set out the implementation steps, checks, dependencies, responsible roles, correction paths and handoff conditions. A small repair may remain a one-agent job. A larger change may call for an independent reviewer or a separate interface check.

Proposing a role does not start another agent or grant it permissions. That remains the responsibility of the host environment within the user’s instructions. The practical benefit of planning is that we can describe what “done” will mean before work begins.

A loop should remember why it returned

The difference between a repair loop and repetition is easy to see in a small bug. A test finds that empty input raises an exception. A useful return path gives the implementer that specific finding. The next change should address empty input, and the next check must inspect the new content.

Sending the same request for review until a more favourable verdict appears would repair nothing. The new workflow therefore builds on existing MCP37 rules: a result belongs to a particular diff, and negative findings remain recorded. A repair creates a new state. Evidence that depends on changed content must be obtained again.

A loop also needs an end. Repair attempts, time and permitted checks have limits. If the same problem keeps returning, or necessary authority is missing, the correct result may be a clear stop with a reason. Endless activity would merely be an expensive way to postpone admitting that something is wrong.

A green indicator needs evidence

“The tests passed” becomes useful when we know which tests ran, what they inspected, their results, and whether anything has changed since. A command that discovers no tests does not demonstrate correctness. Neither does a test whose condition was weakened so that it would stop failing.

MCP38 works with a task contract and verifiable check results. The contract defines scope, rules and limits. Results bind to specific content. Checks protecting acceptance conditions cannot quietly be exchanged for easier ones. When the inputs change, old green results are insufficient.

This is still not a mathematical proof that the product has no defects. Tests can miss cases, and a model reviewer can be wrong. The advantage is that the steps have identifiable inputs and results that can be examined, instead of relying on a vague impression that the agent must have finished.

Less supervision, with responsibility preserved

The aim is not to ask the user about every file. When a repair falls within the authorised scope, the workflow can continue without confirming ordinary steps again. If work requires a new external connection, a wider change or an unresolved decision, however, the plan cannot hide that fact.

The final output should be a reviewable package: what changed, which conditions have been met, what evidence exists, what remains uncertain, and precisely which action should follow. A prepared change and a completed deployment are different states.

The first coordinator version therefore stops at a prepared result. It does not itself grant general authority to merge, publish or deploy to production. Those actions need the appropriate tool and valid permission. Useful autonomy includes recognising where the current task’s authority ends.

What memory retains

Earlier ITHZ versions focused on giving agents relevant decisions, rules, risks and evidence. MCP38 adds the state of the work: where we are, what failed and which next step is allowed.

Live run state is kept apart from tracked source code. The host can save a concise, sanitised summary with references to results in project memory. After an interruption, the environment is checked against the recorded state before continuing. An uncertain action outcome is not a reason to repeat the action blindly.

The next continuation can then start from what actually happened, rather than the last optimistic message in the chat.

What this release does not yet promise

MCP38 is an alpha coordinator. The MCP server does not supply a coding agent of its own: implementation is performed by the host agent, and the first execution template is sequential. A graph in the plan identifies dependencies and possible roles; it is not a promise to launch dozens of agents automatically in the background.

Local checks run with the host’s permissions. The coordinator does not replace an operating-system sandbox, independent permission management or human judgment. The test results listed with the release verify specific software behaviour. They do not establish general improvements in AI quality, development time or cost.

Those questions require comparisons on real tasks: necessary interventions, escaped defects, elapsed time and spending. Asking the user fewer questions is valuable only if the outcome remains verifiable.

The next step for ITHZ

Memory helps an agent know what the project already knows. A workflow helps it know what to do now and what must happen before handoff. The two layers fit together: each repair can address a concrete finding, and each success has a defined scope.

This release grew out of the practical need to stop connecting every small step by hand. It does not require belief in an infallible agent. It requires a process in which a mistake can be exposed, corrected, or brought to an honest stop.

Explore ITHZ MCP · Download the current release · MCP38 release notes

Related reading: workflow and agent patterns in the LangGraph documentation. This reference explains the general pattern; MCP38 uses its own deterministic coordinator.

Reaction

How did this essay land with you?

A quick reaction is sent to me by email. To develop the idea publicly, leave a comment below.

Discussion

Comments appear only after author approval. Each new comment triggers a moderation email.

There are no approved comments yet.

Leave a comment

Your email will not be published. The comment appears only after approval.