reckon
Answers what is actually left, and refuses to blur not done with nobody checked. Every item lands in exactly one category so nothing quietly falls off the list, and it never hands you a single percentage that averages the difference away. Then it schedules the remainder into parallel waves and puts a time range on each, drawn from 1,842 measured Opus agent runs — always a range, and never a plan implying a speedup faster than anything ever observed.
/plugin install reckon@fledgeling-pluginsNeeds the marketplace added first — how to do that.
Reach for it when
Work out what actually remains in a project by reconciling what it promised against what anybody proved — the feature briefs and PRD in docs/features-to-triage on one side, a test-campaign registry of cases, requirements, surfaces and defects on the other — and resolve every item on both sides into exactly one class of a total partition, so nothing can quietly fall out of the list.
Not for
Not for producing the evidence (test-campaign), tracing a spec to the code that produces its data (spec-validation), a whole-product survey against a stated goal (product-gap-analysis), a tracker-board sweep (stocktake), or a decision questionnaire for a non-technical owner (whats-left).
Uses multiple models
Uses multiple modelsThis skill may ask a different AI for a second opinion. Usually to check its own work, because a reviewer from the same family tends to agree with it. The defer skill picks which one, from OpenAI, Google, xAI or another Claude, based on what the job is and which account has room left. Nothing leaves your machine unless a skill you ran asks for it.Read about defer →What ships with it
- Scripts it runs itself
- 5 reference files
- Measured evals
What is actually left, given what anybody actually proved.
/plugin marketplace add h4ckf0r0day/fledgeling-plugins
/plugin install reckon@fledgeling-plugins
Then: "what's left on this project?", "reckon the backlog against the last
test campaign", or /reckon:reckon.
The problem
A remaining-work list is built by filtering. Take everything, drop what is done, report the rest. The filter is where it breaks, because "not done" and "not known" are different things and both look like absence.
Here is an ordinary campaign — not a bad one — from a macOS file-sync client:
| Designed cases | 58 |
| Reached a verdict | 25 |
| Blocked, inconclusive or never run | 32 |
| Stated requirements | 15 |
| Independently observed | 2 |
Filter that to failures and you get a tidy list of fourteen things to fix. It looks like a plan. What it leaves out is that more than half the campaign measured nothing at all — blocked on a dead credential, on states with no hook to force them, on a surface reachable only by signing the owner out of their only account — and that thirteen of fifteen requirements are the project's own account of itself rather than anything checked.
Nothing in that list is wrong. It is just a list about 43% of the product, presented as a list about the product.
What reckon does instead
It refuses to filter. Every brief, requirement, case, defect and surface on both sides resolves into exactly one class of a total partition — so an item cannot fall out of the list, because there is nowhere to fall to.
| Class | What it means | Whose job |
|---|---|---|
unbuilt | A brief cites registry ids and the registry holds none | product |
unjoined | A brief names it; the join reached nothing either way | decision |
broken | Measured, and the answer was no | product |
unmeasured | Nobody found out | the harness |
unnamed | Found in the product; no document claims it | a person |
undecided | The documents and the evidence disagree | a person |
retirable | Already done — close the brief | bookkeeping |
waived | Somebody decided not to | exception |
Three of these are invisible to any ordinary backlog sweep, and they are the reason this exists.
unmeasured is work — and it is not the feature's work. The job behind a
blocked case is reaching the state. Behind an inconclusive one, being able to
read the answer. Behind an unoracled one, deciding what a pass would even look
like. Those are harness jobs, and sending them to a feature backlog as "test
this properly" sends five different jobs to one wrong place. Each row carries
its own remedy.
waived is neither remaining nor done. A decision to skip is a third
thing, and it stays visible because its reason expires — a state with no hook
may get one. Fold waivers into "done" and a campaign closed by decision reads
as a campaign closed by evidence.
retirable is remaining work in reverse. A brief whose subject the
campaign proved works should be closed, not built. Reports that never retire
anything over-count forever.
Twenty blocked cases are usually a handful of problems
Scrim's own stop declaration, from an earlier run of the same campaign, recorded a single dead OAuth credential sitting behind ten of its twenty blocked cases. Listed case by case, that is twenty line items hiding the one thing worth doing on Monday.
So blocked cases are scheduled as the causes behind them, each carrying the coverage that resolving it returns. These are the top clusters from a real run against the campaign above — 32 blocked and inconclusive cases reduced to 24 causes:
| Unblocks | Coverage returned | Cause |
|---|---|---|
| 4 cases | +6.9 pts | Onboarding renders only on a first run with no account stored — reaching it means signing the owner out |
| 3 cases | +5.2 pts | Requires a full Drive quota; the live account reads 79.7 GB of 33 TB and no hook forces the state |
| 3 cases | +5.2 pts | The stored account is in needsReauthorization from a real invalid_client response |
| 2 cases | +3.4 pts | The conflicts table holds 0 rows, so the panel has nothing to draw |
That second column is the number a solo developer actually prioritises on, and it is computed rather than guessed.
The clustering is token-overlap, so it is deliberately conservative and will leave one cause wearing two descriptions. Every case id stays listed under its cluster, so merging them by hand loses nothing — the script does the mechanical pass and stops where judgement starts.
Then it schedules what is left
A remaining-work list is not a plan until you know what can happen at the same
time and roughly how long it takes. So the ledger becomes a wave schedule, using
the same model as ship-fleet:ship-fleet: nodes are work items, edges are dependencies, and
a wave is everything whose dependencies all sit in earlier waves.
| Wave | Items | Slots | Range | Serial | Bounded by |
|---|---|---|---|---|---|
| 1 | 20 | 8 | 78 min – 2.5 h | 5.2 h | observed-speedup-ceiling |
| 2 | 3 | 3 | 15 min – 2.0 h | 55 min | slowest-member |
Every duration comes from 1,842 measured Opus 4.8 and Opus 5 agent runs across 88 sessions and 31 projects, parsed from Claude Code's own transcripts. Three rules keep those numbers honest:
Always a range, never a number. Measured p90 ÷ median is 3.4 for build work and 3.1 for research. A single figure is wrong by a factor of three in the ordinary case, and it reads as precision.
A wave costs its slowest member — and no schedule may beat what was measured. Wall-clock ÷ slowest member came out at 1.05 median and 1.8 at p90 over 253 real waves; speedup over serial was 2.2× median and 4.0× at p90. Twenty items across eight slots can be arithmetically fast; it has never actually happened, so the low bound is held to the observed ceiling rather than printing a duration nobody has achieved.
Decision work gets no duration at all. An undecided row is a person reading
two documents and ruling. Scheduling it as an agent task reports waiting on a
human as though a machine were busy, so those rows sit beside the waves as what
gates them.
Dependency edges are labelled by how they were made. A cited edge is a citation somebody wrote and it blocks. An inferred edge is the tool reading a shared surface, and it only orders the work — because a false edge costs your parallelism, and a guess is not entitled to spend it.
Three artifacts, one gated ledger
ledger.json is the source, and the other two are rendered from it — which is
what stops the presentable half drifting from the gated half.
reckoning.md is the technical record: every row, every edge with its
provenance, per-wave item tables with each tier and why it landed there, the
join's near-misses, and everything the tool could not classify.
reckoning.html is one self-contained page with no build step: shipped features
beside remaining ones, the waves between them with sub-tasks nested under the
item that cites them, and the caveats placed where a reader hits them rather than
in a footnote.
It publishes what it cannot speak for
One blended "percent complete" hides whichever axis is weakest, so there
isn't one. There are five, they disagree with each other on purpose, and every
figure is marked as a floor — because each unnamed row is proof the
intent space is bigger than the documents describe.
A pass rate among executed cases is never labelled coverage. Decisions are counted apart from measurements, because a decision is not a measurement.
The gate
reckon.py check returns an exit code, and it gates the integrity of the
report, not the state of the project:
- exit 1 — the ledger lost an item or placed one illegally. A blocked case presenting as done is caught here, by name.
- exit 2 — a headline figure the rows do not support.
- exit 3 — the ratchet: an item left
unmeasuredbetween two runs with no evidence-bearing event behind it.
Remaining work is content, at exit 0. A gate that fires because work exists fires on every run, and a gate that always fires gets switched off.
That ratchet matters more than it looks. A snapshot gate catches a bad run; the ratchet catches the slow version, where an item is quietly reclassified across runs until nothing remembers it was never checked.
The schedule is gated the same way the partition is: an item in two waves, an item in none, a total that is not the sum of its waves, or a duration attached to something that is not work all fail the gate. A board that disagrees with the rows it was built from is this tool's own failure mode arriving through its presentation layer.
The self-tests prove each gate fires on a deliberately broken ledger and stays silent on a sound one — because a gate nobody has seen fail is indistinguishable from a gate that cannot.
What it will not do
It reconciles documents against evidence. It does not read your code to decide
whether something works, and it says so rather than guessing: where the
documents and the registry disagree, it routes to spec-validation:spec-validation, which
traces a claim to the code that produces its data. Identifier greps may only
ever demote a claim or route it, never promote something to done.
- Producing the evidence → test-campaign
- Is this claimed-done feature real → spec-validation
- Whole-product survey against a goal → product-gap-analysis
- Sweeping a tracker board → stocktake
- A page and a questionnaire for a non-technical owner → whats-left
(hand it this ledger — the
undecidedrows are its input) - Actually doing the work → ship-fleet
Grounding
The design is not invented. Regulated verification has partitioned rather than
filtered for decades: ECSS-E-ST-10-02C mandates a Verification Control
Document recording every requirement's evidence, compliance, close-out state
and reason; FDA device-software guidance requires unresolved anomalies on
the record; TTCN-3 has carried an inconc verdict since long before this.
The empirical case is measured. Status reports carry optimistic bias in 60% of cases (Snow, Keil & Wallace, n=56). Coverage correlates only weakly with suite effectiveness once size is controlled (Inozemtseva & Holmes, 31,000 suites over five systems up to 724k lines). At Google, 84% of pass-to-fail transitions involved a flaky test.
The estimates are measured rather than assumed, and
skills/reckon/references/estimation.md carries the corpus, the method, the
exclusions and what the figures cannot say — including that they are wall-clock
so they include waiting, and that they carry no failure rate.
Full citations in skills/reckon/references/evidence.md, including where two
reviewers disagreed and why one won. The corpus — three deep-research panel
members over 173 sources — is in docs/deep-research/.
MIT. Part of fledgeling-plugins.