← All skills

ship-fleet

v2.9.4MITHanding over a pile of work

Hand it a repo's whole backlog and it works through it. It writes down everything left before starting anything, runs several items at once, and done means the written record says so rather than a job simply returning.

Install
/plugin install ship-fleet@fledgeling-plugins

Needs the marketplace added first — how to do that.

Say any of this

  • orchestrate the remaining work
  • ship everything left in the backlog
  • survey what's left
  • work through the backlog N at a time
  • resume the orchestrator
  • run the fleet

Taken from the skill’s own trigger description — these are the phrases it listens for. You do not have to match them exactly.

Easily confused with

Why it exists

A backlog isn't one feature; it's a queue with dependencies, half-finished worktrees, briefs nobody triaged, and follow-ups buried in progress notes. Running it well is a scheduling job, and the failure mode is specific: fleets that report "completed" while runners died mid-flight. The workflow machinery returns null on a dead agent and the run still finishes clean, so a survey can report a backlog shipped that nothing touched.

ship-fleet's rule for that is blunt: done means the ledger says so, never that the dispatch returned. Everything else follows from it.

How it runs

Say "ship everything left in the backlog" or "run the fleet", and it works through: preflight (checks the repo's conventions with you before touching anything) → survey (classifies every remaining item, including the follow-ups mined out of old progress notes) → worktree hygiene (nothing unmerged is ever destroyed) → the orchestrator artifacts, written and committed before any execution: ORCHESTRATOR.md as the resumable ledger, plus a visual hierarchy of the waves → serial pre-triage (id allocation is a shared write; racing it corrupts the ledger) → the fleet: dependency-ordered ship-feature runners, slots refilled as they free.

Runners stop twice on purpose. They stop before verify, because a runner can't verify its own build; the orchestrator spawns each item's verifier as a fresh agent from a different model family. And they stop before merge, because two simultaneous merges into one integration branch is how fleets corrupt repos; merges go one at a time, behind the fail-closed gate.

New in 2.2: a ready-to-verify report is treated as a claim about an evidence bundle rather than as a fact. An item whose bundle is empty goes back to its runner instead of into the verify queue, because otherwise a fresh agent gets spawned to discover that. And where the repo carries a test campaign, its capture gate runs once for the repo rather than once per item. A fleet multiplies whatever the evidence layer gets wrong, so one campaign filing its screenshots by filename is a bad page, and twenty verdicts resting on the same shape is a Done column nobody can audit.

New in 2.0: the per-item cross-family verification step, a Needs verification survey class (an item sitting in review with no verdict is a gap, not a done), one global agent budget instead of two caps that multiplied into rate-limit storms, and a published low-cost runner shape, because the audit record shows what happens when operators hand-roll cheaper ones: the safeguards are the first thing stripped.

Install

/plugin marketplace add fledgeling-co/fledgeling-plugins
/plugin install ship-fleet@fledgeling-plugins

Expects shipyard and ship-feature alongside it. For the layer above (every repo in a portfolio), see ship-armada.

Does it actually work?

Honestly: less is proven here than this section used to claim, and the eval suite that would settle it was only written today. EVALS.md has the detail, including the audit that caught the overclaim.

What is real. The model-routing rule in references/scheduling-and-concurrency.md is field-learned from an actual fleet run and the reference calls it the single most expensive thing to get wrong, because a fleet of runners silently on the wrong model is expensive and invisible at the same time. The stages underneath this conductor are genuinely measured, in shipyard's EVALS.md, and one of those evals puts ship-feature's own SKILL.md in the tested arm, where it scored 5 of 5 against its committed predecessor's 4 of 5 and took a blind panel 3 to 0 across three model families.

What is not. This section previously said that nearly every line of the scheduling reference recorded a dated incident. That file is 357 lines and carries two dated stamps, both section headings. It also listed four incidents as this skill's own evidence: the git add -A that swept three runners' work onto main and the pkill that killed a sibling's test run are real, but they are recorded in other skills' references rather than here; the model override is genuinely this skill's; and "the seven runners lost to transport failures" appears nowhere in this repository except the sentence that claimed it.

Nothing has run ship-fleet itself against a baseline. The suite now defines what that would take.

Credit

The predecessor in diolog-plugins built and paid for almost all of this; 2.0 is a restructuring of hard-won material, not a fresh invention. The verify-before-merge topology borrows from Vercel Labs' eve-software-factory-template (MIT).