Why I Built MARSHAIL: A Practical Framework for Human-Approved AI-Assisted Software Delivery

Cover image for: Why I Built MARSHAIL: A Practical Framework for Human-Approved AI-Assisted Software Delivery

This week, I released the first version of MARSHAIL.

MARSHAIL, pronounced MAR-shuhl ( /ˈmɑːr.ʃəl/ ), stands for Method for AI-assisted Requirements and Software Engineering, with Human Approval and Incremental Learning. That sounds formal, but the idea is actually quite practical: make AI-assisted software delivery easier to control, easier to review, and easier to improve over time.

I wasn’t trying to invent a new framework. Quite the opposite. Over the last months, I kept using various AI coding tools and agent workflows. Some workflows helped with coding, some were good at framing, planning, helped with verification, review. Some tools were powerful but heavy, some did not do exactly what I expected, some simply didn't work.

So I started collecting the practices that worked well for me. A bit of structured intake. A bit of repo reconnaissance before jumping into code. A real delivery plan instead of a temporary chat plan. Small implementation slices. Explicit verification before review. A way to keep useful lessons without turning every project into a pile of stale notes. At first it was just a set of AI assets and working habits I reused across projects. Then I started organizing them, I added agents, skills, rules, knowledge handling, and sync support through Cyncia. At some point it became its own project. That project is MARSHAIL.

What problem MARSHAIL tries to solve

AI can write code fast. That is useful, but it also creates a few very familiar risks:

  • it's easy to lose control over huge amounts of generated code
  • the assistant searches too broadly and pollutes its context
  • coding starts before the change is properly framed
  • the plan and the code drift apart
  • review feedback is handled ad hoc
  • lessons from one change are lost, or overfitted into bad long-term rules
  • team visibility is weak because too much exists only inside a chat

The main principles

The core principles are:

One canonical flow. Each stage produces an artifact that can be inspected, resumed, and passed to the next stage.

The delivery plan is the source of truth. An AI tool’s internal “plan mode” can help, but the actual plan lives in delivery-plan.md.

Scale the process to the change. A typo should not need a full SDLC. A migration, security-sensitive feature, or cross-module refactor probably should.

Use context efficiently. The goal is not to load the whole repository into every conversation. MARSHAIL narrows the search surface first, uses repository knowledge where available, and keeps stage-specific work focused on the context that actually matters.

Keep work reviewable. MARSHAIL encourages small, understandable implementation slices, while avoiding meaningless PR fragmentation.

Keep humans in the loop where it matters. Even in more autonomous modes, important approval gates are not bypassed.

Capture reusable learning. Every phase can produce learnings, but only reusable lessons should become durable rules, skills, knowledge, or team guidance.

How the workflow looks in practice

MARSHAIL puts the work into a simple, explicit flow:

Specification → Intake → Analysis → Architecture → Plan → Implementation → Verification → Review / PR → Rollout → Learn

Not every change needs the full pipeline. For a small change, the process should stay small. For a risky or cross-cutting change, the full workflow is available. The point is not ceremony. The point is control.

For each change, MARSHAIL creates a working folder under .marshail/work/. That folder holds artifacts like:

  • specification
  • change brief
  • repository reconnaissance
  • architecture notes
  • delivery plan
  • implementation report
  • verification report
  • rollout note
  • learning rollup

It also keeps logs and resume notes, so the work is not trapped in one chat session.

The implementation part is designed as a loop:

Implement → Verify → Review / PR

That loop can run once for the whole change, or repeatedly per phase or slice. Every PR should be preceded by verification for the content being merged.

Knowledge: the part I care about a lot

One important part of MARSHAIL is the knowledge layer.

The per-change artifacts capture what happened in this change. The knowledge layer captures durable facts about the repository: architecture, conventions, logic, decisions, and important rationale. It is not meant to be a copy of existing documentation. It is derived from the code itself and maintained over time. The goal is: the assistant should not rediscover the same repository facts again and again. It should be able to load knowledge progressively, starting from a root index and going deeper only where the task needs it. That makes AI work less noisy, more focused, and more consistent across changes.

Context control

An important part of MARSHAIL is context control. One of the easiest ways AI-assisted work goes wrong is by loading too much irrelevant information to the context. MARSHAIL tries to avoid this by narrowing the scope before expanding the analysis: first understand the request, then identify the relevant part of the repository, then plan and implement with the smallest useful context. Over time, the knowledge layer helps even more, because durable repository facts can be reused instead of rediscovered in every session.

Autonomy modes

You can use MARSHAIL in a direct mode, where you call a specialist agent or skill for a specific stage, such as planning, analysis, verification, or review.

You can also use it in a driver-mediated mode, where marshail-driver acts as the single point of contact and coordinates the specialist agents.

The level of involvement can also vary:

  • hands-off, where the driver comes back mainly for decisions and approval gates
  • collaborative, where you shape the specification, plan, and implementation choices together
  • mixed, where some parts are handled autonomously and some are worked through interactively

Knowledge and extension updates have their own autonomy controls. Knowledge can be updated automatically, or proposed as a diff for review first. The same idea applies to generated rules, skills, subagents, and other reusable assets.

Installation

From the root of a target repository:

curl -fsSL https://raw.githubusercontent.com/crestreach/marshail/main/scripts/install-marshail.sh | bash

Then run:

marshail-init

The installer brings MARSHAIL into .marshail/, installs Cyncia (https://www.cyncia.net) when needed, and supports syncing the durable AI assets into tool-native layouts.

Example: collaborative mode

A typical collaborative prompt could be:

Use marshail-driver for this change in collaborative mode. Add support for [feature]. Do not implement until we agree on the delivery plan. Ask me before any major architecture decision. And then let's implement together step by step.

Or, if I know exactly what stage I want:

Use marshail for this change. Create a delivery plan 2 levels deep by default, but go to 4 levels for the migration part. Do not implement yet.

Example: unattended / hands-off mode

For more autonomous work:

Use marshail-driver in hands-off mode. Implement issue PROJ-1234. Run only the stages that add value. Return to me for approval gates, unclear requirements, failed verification.

This still does not mean “do whatever you want”. The plan remains the source of truth, verification is explicit, and human approval gates still matter.

Why I’m releasing it

I’m releasing this first version because I think many of us are in the same phase with AI-assisted development. We are no longer just asking, “Can AI write this code?”

We are asking:

  • Can I trust the process?
  • Can I review the work?
  • Can this survive a team workflow?
  • Can the assistant learn useful things without becoming messy?
  • Can I resume the work tomorrow and still understand what happened?

MARSHAIL is my attempt to address these. It is a first version. I expect it to evolve as I use it more, and as other people try it in their own projects.

But it already represents a way of working that I find useful, controlled, human-approved, and designed for incremental learning.

The project is here: www.marshail.net

Feedback, questions, and real-world usage notes are very welcome.