← All skills

reckon

v1.9.5MITKnowing where things stand

Answers what is actually left, and refuses to blur not done with nobody checked. Every item lands in exactly one category so nothing quietly falls off the list, and it never hands you a single percentage that averages the difference away. Then it schedules the remainder into parallel waves and puts a time range on each, drawn from 1,842 measured Opus agent runs — always a range, and never a plan implying a speedup faster than anything ever observed.

Install
/plugin install reckon@fledgeling-plugins

Needs the marketplace added first — how to do that.

Reach for it when

Work out what actually remains in a project by reconciling what it promised against what anybody proved — the feature briefs and PRD in docs/features-to-triage on one side, a test-campaign registry of cases, requirements, surfaces and defects on the other — and resolve every item on both sides into exactly one class of a total partition, so nothing can quietly fall out of the list.

Not for

Not for producing the evidence (test-campaign), tracing a spec to the code that produces its data (spec-validation), a whole-product survey against a stated goal (product-gap-analysis), a tracker-board sweep (stocktake), or a decision questionnaire for a non-technical owner (whats-left).

Uses multiple models

Uses multiple modelsThis skill may ask a different AI for a second opinion. Usually to check its own work, because a reviewer from the same family tends to agree with it. The defer skill picks which one, from OpenAI, Google, xAI or another Claude, based on what the job is and which account has room left. Nothing leaves your machine unless a skill you ran asks for it.Read about defer →

What ships with it

What is actually left, given what anybody actually proved.

/plugin marketplace add h4ckf0r0day/fledgeling-plugins
/plugin install reckon@fledgeling-plugins

Then: "what's left on this project?", "reckon the backlog against the last test campaign", or /reckon:reckon.


The problem

A remaining-work list is built by filtering. Take everything, drop what is done, report the rest. The filter is where it breaks, because "not done" and "not known" are different things and both look like absence.

Here is an ordinary campaign — not a bad one — from a macOS file-sync client:

Designed cases58
Reached a verdict25
Blocked, inconclusive or never run32
Stated requirements15
Independently observed2

Filter that to failures and you get a tidy list of fourteen things to fix. It looks like a plan. What it leaves out is that more than half the campaign measured nothing at all — blocked on a dead credential, on states with no hook to force them, on a surface reachable only by signing the owner out of their only account — and that thirteen of fifteen requirements are the project's own account of itself rather than anything checked.

Nothing in that list is wrong. It is just a list about 43% of the product, presented as a list about the product.

What reckon does instead

It refuses to filter. Every brief, requirement, case, defect and surface on both sides resolves into exactly one class of a total partition — so an item cannot fall out of the list, because there is nowhere to fall to.

ClassWhat it meansWhose job
unbuiltA brief cites registry ids and the registry holds noneproduct
unjoinedA brief names it; the join reached nothing either waydecision
brokenMeasured, and the answer was noproduct
unmeasuredNobody found outthe harness
unnamedFound in the product; no document claims ita person
undecidedThe documents and the evidence disagreea person
retirableAlready done — close the briefbookkeeping
waivedSomebody decided not toexception

Three of these are invisible to any ordinary backlog sweep, and they are the reason this exists.

unmeasured is work — and it is not the feature's work. The job behind a blocked case is reaching the state. Behind an inconclusive one, being able to read the answer. Behind an unoracled one, deciding what a pass would even look like. Those are harness jobs, and sending them to a feature backlog as "test this properly" sends five different jobs to one wrong place. Each row carries its own remedy.

waived is neither remaining nor done. A decision to skip is a third thing, and it stays visible because its reason expires — a state with no hook may get one. Fold waivers into "done" and a campaign closed by decision reads as a campaign closed by evidence.

retirable is remaining work in reverse. A brief whose subject the campaign proved works should be closed, not built. Reports that never retire anything over-count forever.

Twenty blocked cases are usually a handful of problems

Scrim's own stop declaration, from an earlier run of the same campaign, recorded a single dead OAuth credential sitting behind ten of its twenty blocked cases. Listed case by case, that is twenty line items hiding the one thing worth doing on Monday.

So blocked cases are scheduled as the causes behind them, each carrying the coverage that resolving it returns. These are the top clusters from a real run against the campaign above — 32 blocked and inconclusive cases reduced to 24 causes:

UnblocksCoverage returnedCause
4 cases+6.9 ptsOnboarding renders only on a first run with no account stored — reaching it means signing the owner out
3 cases+5.2 ptsRequires a full Drive quota; the live account reads 79.7 GB of 33 TB and no hook forces the state
3 cases+5.2 ptsThe stored account is in needsReauthorization from a real invalid_client response
2 cases+3.4 ptsThe conflicts table holds 0 rows, so the panel has nothing to draw

That second column is the number a solo developer actually prioritises on, and it is computed rather than guessed.

The clustering is token-overlap, so it is deliberately conservative and will leave one cause wearing two descriptions. Every case id stays listed under its cluster, so merging them by hand loses nothing — the script does the mechanical pass and stops where judgement starts.

Then it schedules what is left

A remaining-work list is not a plan until you know what can happen at the same time and roughly how long it takes. So the ledger becomes a wave schedule, using the same model as ship-fleet:ship-fleet: nodes are work items, edges are dependencies, and a wave is everything whose dependencies all sit in earlier waves.

WaveItemsSlotsRangeSerialBounded by
120878 min – 2.5 h5.2 hobserved-speedup-ceiling
23315 min – 2.0 h55 minslowest-member

Every duration comes from 1,842 measured Opus 4.8 and Opus 5 agent runs across 88 sessions and 31 projects, parsed from Claude Code's own transcripts. Three rules keep those numbers honest:

Always a range, never a number. Measured p90 ÷ median is 3.4 for build work and 3.1 for research. A single figure is wrong by a factor of three in the ordinary case, and it reads as precision.

A wave costs its slowest member — and no schedule may beat what was measured. Wall-clock ÷ slowest member came out at 1.05 median and 1.8 at p90 over 253 real waves; speedup over serial was 2.2× median and 4.0× at p90. Twenty items across eight slots can be arithmetically fast; it has never actually happened, so the low bound is held to the observed ceiling rather than printing a duration nobody has achieved.

Decision work gets no duration at all. An undecided row is a person reading two documents and ruling. Scheduling it as an agent task reports waiting on a human as though a machine were busy, so those rows sit beside the waves as what gates them.

Dependency edges are labelled by how they were made. A cited edge is a citation somebody wrote and it blocks. An inferred edge is the tool reading a shared surface, and it only orders the work — because a false edge costs your parallelism, and a guess is not entitled to spend it.

Three artifacts, one gated ledger

ledger.json is the source, and the other two are rendered from it — which is what stops the presentable half drifting from the gated half.

reckoning.md is the technical record: every row, every edge with its provenance, per-wave item tables with each tier and why it landed there, the join's near-misses, and everything the tool could not classify.

reckoning.html is one self-contained page with no build step: shipped features beside remaining ones, the waves between them with sub-tasks nested under the item that cites them, and the caveats placed where a reader hits them rather than in a footnote.

It publishes what it cannot speak for

One blended "percent complete" hides whichever axis is weakest, so there isn't one. There are five, they disagree with each other on purpose, and every figure is marked as a floor — because each unnamed row is proof the intent space is bigger than the documents describe.

A pass rate among executed cases is never labelled coverage. Decisions are counted apart from measurements, because a decision is not a measurement.

The gate

reckon.py check returns an exit code, and it gates the integrity of the report, not the state of the project:

  • exit 1 — the ledger lost an item or placed one illegally. A blocked case presenting as done is caught here, by name.
  • exit 2 — a headline figure the rows do not support.
  • exit 3 — the ratchet: an item left unmeasured between two runs with no evidence-bearing event behind it.

Remaining work is content, at exit 0. A gate that fires because work exists fires on every run, and a gate that always fires gets switched off.

That ratchet matters more than it looks. A snapshot gate catches a bad run; the ratchet catches the slow version, where an item is quietly reclassified across runs until nothing remembers it was never checked.

The schedule is gated the same way the partition is: an item in two waves, an item in none, a total that is not the sum of its waves, or a duration attached to something that is not work all fail the gate. A board that disagrees with the rows it was built from is this tool's own failure mode arriving through its presentation layer.

The self-tests prove each gate fires on a deliberately broken ledger and stays silent on a sound one — because a gate nobody has seen fail is indistinguishable from a gate that cannot.

What it will not do

It reconciles documents against evidence. It does not read your code to decide whether something works, and it says so rather than guessing: where the documents and the registry disagree, it routes to spec-validation:spec-validation, which traces a claim to the code that produces its data. Identifier greps may only ever demote a claim or route it, never promote something to done.

  • Producing the evidence → test-campaign
  • Is this claimed-done feature real → spec-validation
  • Whole-product survey against a goal → product-gap-analysis
  • Sweeping a tracker board → stocktake
  • A page and a questionnaire for a non-technical owner → whats-left (hand it this ledger — the undecided rows are its input)
  • Actually doing the work → ship-fleet

Grounding

The design is not invented. Regulated verification has partitioned rather than filtered for decades: ECSS-E-ST-10-02C mandates a Verification Control Document recording every requirement's evidence, compliance, close-out state and reason; FDA device-software guidance requires unresolved anomalies on the record; TTCN-3 has carried an inconc verdict since long before this.

The empirical case is measured. Status reports carry optimistic bias in 60% of cases (Snow, Keil & Wallace, n=56). Coverage correlates only weakly with suite effectiveness once size is controlled (Inozemtseva & Holmes, 31,000 suites over five systems up to 724k lines). At Google, 84% of pass-to-fail transitions involved a flaky test.

The estimates are measured rather than assumed, and skills/reckon/references/estimation.md carries the corpus, the method, the exclusions and what the figures cannot say — including that they are wall-clock so they include waiting, and that they carry no failure rate.

Full citations in skills/reckon/references/evidence.md, including where two reviewers disagreed and why one won. The corpus — three deep-research panel members over 173 sources — is in docs/deep-research/.


MIT. Part of fledgeling-plugins.