groundtruth.worksThe method, published in full · version 1.0 · 6 September 2026

The Groundtruth method

Version 1.0, 6 September 2026. Published by the Groundtruth cohort, with research and drafting support from Claude.


Part 0. What this document is

Licence and ownership

The method described in this document (the canvas, the discovery protocol, the four questions and their scoring, change capacity, change load, the vendor drop, the roll-up rules, the calibration design and the research it stands on) is published under the Creative Commons Attribution-ShareAlike 4.0 licence. Anyone may use it, teach it, build with it or adapt it, provided they attribute it to Groundtruth and the Groundtruth cohort and publish any adapted method under the same terms.

Three things are not covered by that licence and stay with the Groundtruth cohort: the Groundtruth name and mark; the software that runs the method (the discovery agent, its prompts, the application and its hosting); and the cohort's service and pricing arrangements. The scoring engine, the arithmetic that turns factor levels into numbers, is published as open source under Apache 2.0 alongside this document, because a reader cannot check a reading without it.

The data a business produces with the method (its canvas, its quotes, its readings) belongs to that business. The terms under which a program may aggregate it are in Part 9.

The method in one paragraph

A person talks through their business the way they would tell a neighbour. As they talk, a canvas builds: a graph of who makes the calls, where the work lives, where it grinds, what people worry about, who holds knowledge that exists nowhere but in their head, and what happened last time something new arrived. Every record on the canvas carries something the person actually said. Two instruments then read the canvas. Four plain questions (Alignment, Roles, Readiness, Strategy) each give a scored reading of the business on one thing that decides whether change survives. Change capacity gives the budget: how much change the business can absorb right now, built from named factors with their evidence shown. When a technology vendor sends a quote, the method estimates the change load the offering would put on this business, weighs it against capacity, and returns a plain verdict, a first-year cost that includes the lines the vendor left out, the questions to ask, and the smallest version worth proving first. The canvas persists, so each return visit starts from what changed. Individual canvases roll up into a picture of a program or a sector.

Who this is written for

Three readers. A consultant running a discovery by hand and scoring it: Parts 3 to 6 and the appendices are the working manual. An engineer building the instrument: Parts 2, 4 to 6 and 8 define what the software computes and what it must never invent. A reviewer inside a research and development corporation, a program office or a board deciding whether to trust a reading: Parts 0, 1, 5, 7, 8, 9 and 11 say what the numbers are, what they are not, and how they are checked.

What has been established at this version, and what has not

This is the paragraph that has to be true before anything else in the document is read.

Demonstrated: the method has produced a full canvas, four lens readings and a capacity reading from one real thirty-minute first-contact call in January 2026, and a vendor assessment from a constructed case (Riverbend, AgriSense Pro). Both are in Appendix C, re-scored under this version. The real business is not named and its quotes are withheld from publication; the reading is held privately and its arithmetic is reproduced in the appendix from paraphrased evidence.

Written: the scoring protocol in Parts 4 to 6 fixes the base, the factor vocabulary, the impact at each level, the load table, the thresholds and the verdict triggers. Under version 1.0 a reading can be reproduced by a second person from the same canvas.

Not yet calibrated: no calibration set has been scored by two independent scorers under this version. Until Part 7's calibration set is run, every reading produced under 1.0 is labelled provisional and the method makes no claim that two scorers would produce the same number.

Not yet validated: no field test has shown that a capacity reading predicts whether a change survives, or that a vendor verdict predicts satisfaction. Part 7 states the tests the method commits to and the study that would run them.

The weights in Parts 5 and 6 are reasoned defaults, chosen to reproduce the three existing examples within a few points and to sit inside the ranges the method has used since June 2026. They are judgement informed by the literature in Part 11, not measured constants. Calibration will move them. The version number exists so that a reading always says which weights produced it.

Versioning

Every reading records the method version that produced it. Readings are comparable only within a version. Appendix D records what changes between versions and why.


Part 1. The thesis

The science of change has not changed. For eighty years the research has said that technology projects fail on people: when the people who decide do not agree on where they are going, when nobody has worked out how jobs change, when the culture and the systems cannot take the weight, when the plan was imported from a template, and when trust spent on the last failed change has not been earned back. That work is named in Part 11 and it is the spine of the four questions.

The interface that carries change has changed. Every prior wave of technology made the person bend to the machine: learn the form, the login, the menu. Conversational AI is the first that can bend to the person, so a grower can talk the way they already talk and the machine keeps up. That removes a cost every earlier tool charged at the point of use, and it adds a factor the older adoption models never had to name: whether people like the thing enough to keep using it.

So the binding constraint on value from AI has moved. It is no longer whether the technology works. It is whether this business, this season, with these people, can absorb what the technology now makes possible, and whether they will want to. Groundtruth measures that constraint and stages change against it.

One boundary holds everywhere. This is an argument about AI working alongside people. It holds where a tool complements a person and breaks where a tool simply replaces a role, and the method says so when an offering is a replacement dressed as an upgrade.

The long-form argument, with the enterprise failure statistics and the sector evidence, is in the Groundtruth evidence base, which this document cites but does not repeat.


Part 2. The canvas

The canvas is one graph, stored as one JSON object to a fixed schema (Appendix A). Everything else in the method reads it.

Nodes

Each node has an id (a short stable slug), a type, a label, a detail line, a status, one or more lens tags, and an evidence quote. Twelve types:

Type Use for
person a named or role-identified individual
role a function rather than a person, such as "harvest casuals"
decision_pattern how decisions actually get made: kitchen table, owner bottleneck, committee
system software, tools or paper processes in use
data_store where information actually lives: a spreadsheet, a diary, someone's head
friction recurring pain that costs time, money or sleep
concern a worry people voiced: jobs, privacy, cost, trust
champion a person or role with energy and standing for change
history_event a past change attempt, success or failure
seasonal_constraint windows when change is impossible or possible
external_party advisers, buyers, co-packers or program bodies that matter
opportunity a concrete improvement the conversation surfaced

Status is one of strength (working in the business's favour), neutral, watch (needs attention, could go either way) or risk (actively working against the business or the change). Status drives colour on the rendered canvas, so it has to mean something.

Edges

Edges connect nodes and carry a kind: structure (organisational relationships), flow (information or work moving) or tension (where things grind). Labels are three words or fewer.

The evidence rule

Every node carries a short verbatim or near-verbatim quote from the source. If no quote can be pointed to, the node does not go on the canvas; the hunch goes to open questions. This rule is the method's credibility with the business (they recognise their own words) and with any later reader of a score (they can see what it rests on). The instrument enforces it: a node cannot be created without an evidence field, and the discovery agent is prohibited from paraphrasing into a quote.

Size

A typical thirty to sixty minute conversation yields twelve to twenty-five nodes. Fewer means the listening was shallow. Many more means generic material crowded out what is distinctive about this business.

Sessions and memory

The canvas is never overwritten. Each conversation appends a session entry (date, mode, summary). A node that stops being true changes status and keeps its history. Part 5 gives the rule for re-scoring on a return visit.


Part 3. The discovery conversation

This part is the protocol for the conversation, whether a consultant runs it or the instrument does. It is the specification the discovery agent is built from, and the manual a consultant runs from.

Two durations, stated honestly

The full discovery is a thirty to sixty minute conversation that produces a canvas of twelve to twenty-five nodes and enough evidence to score the four questions and capacity. It is the only conversation in the method that produces a reading.

The readiness pulse is a fixed six-question form, answered in two to three minutes by voice or by tapping, that produces a provisional capacity number and a seed canvas of six nodes. It exists to qualify interest and to let a program reach breadth. It is not a discovery, it does not produce lens readings, and the method never describes it as a conversation that maps the business. Its questions and fixed impacts are in Appendix B, section B.2.

The principle

You listen; the canvas grows. Quote them, always. Map the distinctive, not the generic. "Compliance paperwork" is a category; "compliance paperwork eats his nights and weekends every season" is the canvas doing its job.

Opening

Open on the business, not the technology. The seed line: "Just talk me through your place the way you'd tell a neighbour. What do you grow or run, and who's involved day to day?" The first nodes are people, roles, systems and the season. No AI, no change talk, until the business is on the board.

The coverage map

The conversation is not a script. The operator talks in whatever order things come, and a grower who jumps from rootstock to a remembered failure to a worry about the crew is handing over three nodes in one breath. The agent or consultant follows the thread and reaches for the question that fits what was just said. The coverage map is what must eventually be heard, not the order to ask it in.

Seventeen traced questions, each with what it surfaces and the lens it feeds, form the map (Appendix B, section B.1, reproduced from the evidence base with the research trace on each). They cover: who actually decides and whether the deciders agree (Alignment, four questions); what only works if it is in someone's head, who checks whose work, how jobs would shift, and whether anyone has voiced job worry (Roles, four); what happened last time, where the information lives, the spread of digital confidence, and how much else is going on (Readiness, four); the one thing that grinds every week, who would give something new a go, and when change is impossible (Strategy and capacity, three); the operator's own read on whether the plate is full (capacity, one); and the seed.

The ten things capacity needs to hear

Separately from the lenses, the conversation has to produce evidence for or against each capacity factor in Part 5, or an explicit "not raised". The factor list is short enough to hold in mind: past failure; champion; season; leadership load; knowledge in one head; a change that landed; digital confidence; crisis; systems drag; the installed base. Most surface on their own through the coverage map. When one has not, the agent asks for it once, plainly, before closing.

The stopping rule

The conversation is "enough to see the shape of it" when all of the following hold: at least twelve nodes with evidence; at least one node tagged to each of Alignment, Roles and Readiness; at least one friction node; and for each capacity factor either an evidence quote or an explicit "not raised" recorded in open questions. If the operator wants to stop earlier, the reading is produced and labelled provisional with the gaps listed. If the conversation runs past the coverage map, the agent closes rather than fishing.

Reading the room

Five heuristics from field practice, applied by a consultant and encoded in the agent as things to test rather than assume. The stated decision-maker is often not the real one: watch who they glance at, whose objection they pre-empt. Read capacity from how they talk about the last failure: "we paid for two years and barely used it" said with a wince is trust spent; the same event told as "we learned what not to do" reads differently. The data store is often a person: when the answer to "where would that be?" is a name, you have found a fragile point and a key relationship. Catch the silent partner: one keen voice and one silent one is unknown alignment, not high alignment. Friction with a champion next to it is where a plan starts.

Closing

Close by testing, not summarising. Reflect three or four things back in the operator's own words and watch whether they nod or correct. Surface the gaps out loud and turn them into open questions. Leave them with the map, not a score.

Who scores, and when

The conversation maps. It does not score in the room. After the conversation, the instrument proposes a factor level for each capacity factor and a decade anchor for each lens, each with the quote that supports it. A human scorer (the consultant, or the operator's adviser) confirms or moves each proposal before the reading is final. A reading nobody has confirmed is provisional. The pulse is always provisional. This split is what lets the instrument scale the listening and keeps a person accountable for the number.

Voice, typed, and when the audio fails

Voice is the primary input, because the operator talks more easily than they type and because the coverage map works best when the person can wander. The typed path uses the same coverage map and the same stopping rule; the only difference is input. When audio fails (a harvester, a dead spot, a phone in a pocket) the instrument says so, keeps what it has, offers the typed path, and never guesses at a quote it did not hear. A quote the instrument cannot transcribe with confidence is not evidence.

Failure modes

Leading the witness. Mapping categories. Scoring to be liked. Talking more than they do. The playbook that sits behind this part (the Groundtruth discovery playbook, June 2026) expands each.


Part 4. The four questions, scored

Four plain questions sit over the canvas. Each scores 0 to 100 and each traces to evidence. Traffic lights: green is 70 and above, amber 40 to 69, red below 40, pending when the conversation gave nothing to score.

The integer rule

For Alignment, Roles and Readiness the scorer places the business in a decade using the anchors below and the evidence tests attached to them. The score is the midpoint of that decade: 85, 65, 45 or 25. It moves by exactly five points, once, in one direction: up five if there is evidence for the decade above but not enough to place the business there; down five if there is evidence for the decade below. No other adjustment. The result is one of twelve values (20, 25, 30, 40, 45, 50, 60, 65, 70, 80, 85, 90), and the reasoning is always "placed in the forties because X, moved up five because Y".

A reading is provisional if fewer than two independent quotes support the placement, or if only one decision-maker was heard.

Alignment: do the people who actually decide agree on where this is heading?

Eighties: decision-makers share a direction, can say what success looks like, and have talked about what this change means for it. Evidence test: two or more deciders heard, and they name the same destination.

Sixties: direction is broadly shared but informal and untested; the kitchen table agrees but no real decision has stressed it. Evidence test: one decider names a direction and describes the other as on side, with no sign of contest.

Forties: visible divergence on priorities or appetite, or one keen party and one silent one whose view is unknown. Evidence test: a second decider exists and has not been heard, or spend is described as contested.

Twenties: open disagreement, or the person driving change lacks the authority to carry it. Evidence test: a quote of disagreement, or the driver describes needing sign-off they do not expect to get.

What moves it: who actually holds decisions against who was in the room; whether spending is contested; whether anyone has pictured the next eighteen months.

Trace: the dominant coalition (Cyert and March, 1963); the informal network predicting outcomes better than the chart (Cross and Parker, 2004); change failing without a guiding coalition and a shared destination (Kotter, 1996).

Roles: has anyone worked out how the jobs change?

Eighties: the business can say which roles change, who would check automated output, and has talked to the affected people. Sixties: clear sense of who does what today and where the load sits, no thinking yet about the shape after change. Forties: role boundaries blurry or held in one head; knowledge concentrated in individuals; nobody has considered job impact. Twenties: active fear about jobs that nobody is addressing, or recent knowledge loss nobody has plugged.

Evidence tests, by decade: eighties needs a named reviewer for automated output; sixties needs a clear who-does-what and no post-change thinking; forties needs a knowledge-concentration node or unconsidered impact; twenties needs a voiced, unaddressed job fear or a lost-knowledge history event.

What moves it: knowledge concentration; voiced job worries; succession or staffing pressure; whether anyone reviews anyone's work today.

Trace: sociotechnical systems (Trist and Bamforth, 1951); tacit knowledge (Polanyi, 1966; Nonaka and Takeuchi, 1995); the ironies of automation (Bainbridge, 1983).

Readiness: can the place absorb this, the people and the systems both?

Eighties: digital tools used consistently across the team, data accessible outside individual heads, past change landed, sceptics engaged. Sixties: pockets of digital confidence with gaps; data exists but fragmented; a pragmatic culture that adopts what visibly works. Forties: heavy dependence on paper, memory or one person's spreadsheet; wide spread between most and least digital staff; scepticism with cause. Twenties: distrust earned by bad experience; data effectively inaccessible; no slack to learn anything.

Evidence tests, by decade: eighties needs a landed change and shared data; sixties needs a working system with gaps named; forties needs a paper or head data store as the primary record; twenties needs an earned-distrust history event or a stated absence of slack.

What moves it: where data actually lives; the spread of digital confidence; what happened last time; available attention.

Trace: organisational readiness as shared willing-and-able belief (Weiner, 2009; Armenakis, Harris and Mossholder, 1993); absorptive capacity (Cohen and Levinthal, 1990).

Conviviality, whether people like a tool enough to keep using it, is carried as a Readiness consideration when a specific tool is in view (the traced question "which tool would you be sad to lose, and which do you put up with?") and as a design law on the instrument itself. It does not carry a separate weight in this version; see Part 7 for what would have to be shown before it did.

Strategy: does the plan fit this business, or was it copied from a template?

Strategy is scored on a plan, not on the business, so it is pending after a first conversation and scores once a change strategy exists. Four fit tests, each scored present (25), partial (12) or absent (0), summed:

The plan starts where friction and a champion overlap on the canvas. It respects the seasonal windows recorded on the canvas. Its first step's change load sits within capacity (Part 6). If a past failure is on the canvas, the plan names how it earns trust back.

A plan that scores under 40 is sent back before it is shown to the business.

Trace: task-technology fit (Goodhue and Thompson, 1995); contingency thinking; diffusion's trialability and observability (Rogers, 1962).

The double-counting rule

Several facts feed both a lens and a capacity factor: knowledge in one head lowers Roles and is a capacity factor; a past failure lowers Readiness and is a capacity factor; a champion lifts Readiness and is a capacity factor. This is by design. The lens says where the weakness sits and what kind it is; the factor says what it costs in budget. The two instruments answer different questions from the same evidence, and a reader should expect them to move together. Capacity is not derived from the four lens readings, and the lens readings are not derived from capacity.


Part 5. Change capacity

What it is

Change capacity is the budget a business has for absorbing change right now, on a 0 to 100 scale. It is built from a base and a short list of named factors, each carrying evidence from the canvas. The factor table is part of the reading; a capacity number without its factors is not a Groundtruth reading.

The base

Every business starts at 60. The reason: a going concern with no evidence either way sits at the bottom of the "room for one change, staged" band (see the bands below), which is the honest default for a business nobody has yet listened to. The base is a convention of this version, not a finding, and Appendix D records it as a calibration candidate.

The factors

Ten canonical factors. Each has a direction and three levels, mild, clear and severe, with a fixed impact at each level and an evidence test that places the factor at a level. The scorer records the level, the quote, and nothing else; the engine applies the impact.

Factor Direction Mild Clear Severe Evidence test for "clear"
Past change failure, trust spent negative -6 -12 -18 a paid-for system or tool abandoned within a season, told with regret; severe when told with anger or as a rule ("we don't do that any more")
Champion with energy and standing positive +5 +9 +13 a named person who picked up the last new tool and who others follow; severe when they also hold decision authority
Seasonal crunch, near or current negative -4 -8 -13 a named window within three months when nothing new can land; severe when the business is inside it now
Leadership stretched thin negative -4 -8 -13 the owner or driver describes carrying two or more functions beyond their own; severe when they name nights or weekends
Knowledge concentrated in one person negative -4 -7 -10 a named person whose departure would take operational knowledge with them; severe when retirement or departure is near
Recent change that landed well positive +5 +9 +13 a tool or process adopted in the last two years and still in daily use; severe when the team names it unprompted as a win
Digital confidence broad, not pocketed positive +4 +7 +10 most of the team use the current systems without a go-to person; severe when they self-serve on new tools
Digital confidence pocketed negative -3 -6 -9 one or two people carry all digital work; severe when others say they "never got their head around" the current systems
Active operational crisis negative -8 -14 -20 a current event consuming leadership attention (drought, disease, a lost buyer, a staffing collapse); severe when it threatens the business
Systems drag or an installed base either -4 / +4 -8 / +7 -12 / +10 negative when core systems block or slow any new tool (unsupported software, offline transaction systems); positive when a modern base is in place and data already flows into it

Rules. A factor counts once per canvas; two champions are one champion factor at a higher level, with both quotes. A factor with no evidence is not recorded and contributes nothing; "not raised" goes to open questions. One additional named factor is allowed per canvas outside the ten, with a magnitude no greater than 8 and a quote, flagged as "other" in the reading; if calibration shows the same "other" recurring, it becomes canonical in the next version. Impacts are additive. The result is clamped to 5 at the floor (a going concern always has some capacity) and 95 at the ceiling.

The bands

Capacity Plain reading
under 25 hold what you have this year
25 to 44 room for one small change
45 to 64 room for one change, staged
65 to 79 room for a couple, if they fit
80 and above room for more; rare, and worth checking the evidence

What the number is, and is not

It is a structured judgement with its factors shown, reproducible by a second scorer from the same canvas under the same version, and comparable across businesses read under the same version by scorers who have passed the calibration in Part 7. It is not a measured constant, and until Part 7 has run it is not shown to predict anything. The bands are a reading aid, not thresholds with a source.

Re-scoring on a return visit

On a return visit every factor in the previous table is re-tested against new evidence and marked unchanged, moved (with the new quote and the new level), or resolved (kept in the table at zero, with the date and the quote that resolved it). New factors are added with evidence. The new capacity is the base plus the live factors. The changelog is the diff of the factor table, and the capacity delta is explained line by line. A factor never disappears from the table; that is what makes the canvas a memory.

Trace

Absorptive capacity (Cohen and Levinthal, 1990): a business takes in something new in proportion to what it already knows, which is why digital confidence and the installed base are factors. Organisational readiness (Weiner, 2009). Change saturation: organisations have a finite tolerance for change and overloading it degrades performance, which is why season, leadership load and crisis are factors. The remembered failure as a predictor (Armenakis et al., 1993), which is why past failure carries the largest negative range.


Part 6. Change load and the vendor drop

What the vendor drop is

A technology vendor sends a quote. The method does not judge the offering on its own merits; it judges it against this business's canvas. A brilliant product can be a poor fit and a modest one a strong fit, because the verdict belongs to the business.

The procedure

Parse the quote. From the quote, deck, URL or description, record what it claims to do, what data it assumes exists and in what state, what it costs (licence, implementation, and usage or inference, stated or estimated), what roles it touches, the implementation effort, and what happens to the business's data. Claims that cannot be verified are recorded as claims.

Map it to the canvas. Which friction, concern or opportunity nodes does it actually land on, by id? Which data_store nodes does it assume are clean and digital, and are they? Which people and roles change, and is a champion adjacent? Does a history_event remember a similar shape of commitment? Does the timeline collide with a seasonal_constraint?

Check for a Mismatch. If the offering lands on no friction, concern or opportunity node on the canvas, it solves a problem this business has not named. The verdict is Mismatch and the assessment says what problem the offering solves and that the canvas does not hold it. Load is still estimated for the record.

Estimate the change load, twice. Once for the offering as quoted, and once for the smallest provable version (below). Load is built from the table that follows.

Compare to capacity and apply the verdict rule.

Build the cost lens under the costing convention.

Write the verdict, the five questions and the smallest provable version.

The load table

Five considerations, each at three levels with a fixed contribution, plus one modifier.

Consideration Light Moderate Heavy Evidence test for "heavy"
Implementation disruption: duration and how much normal work it interrupts 4 10 16 more than six weeks of onboarding, or any cutover of a live system
People changing daily habits: how many of the permanent team 4 (one person or a quarter) 10 (up to half) 16 (most) more than half the permanent people change what they do each day
Training burden 3 8 14 formal training sessions plus a new skill the team does not have
Data preparation: the hidden cost 0 (starts from the next record) 10 (some cleaning or joining of existing digital records) 24 (digitising paper or memory, or a minimum history the business does not hold) the offering needs a record history that exists only on paper, in heads, or not at all
Against the grain: does it run with or against how the business already works 0 (with) 8 (mixed) 16 (against) it asks people who talk to type, or people who are in the paddock to be at a desk, or replaces a trusted relationship

Modifier: history echo. Add 8 if the commitment shape repeats a history_event with status watch or risk on the canvas (the same billing shape, the same onboarding shape, the same vendor type), and name the node. Add nothing otherwise.

Load is the sum, 0 to 94. Data preparation carries the largest weight because it is the cost vendors most often leave out and the one the canvas can see most clearly.

Trace: Rogers's adoption attributes (1962). Compatibility maps to "against the grain" and to data preparation; complexity to training and disruption; trialability and observability to the smallest provable version below. A change that is compatible, simple, trialable and visible is low load by construction.

The smallest provable version

Every assessment names the minimum slice of the offering this business could prove in about thirty days: fewer people, one record type, one paddock, monthly billing. Its load is estimated from the same table. If no such slice exists, the assessment says so, and that is itself evidence for the verdict.

The verdict rule

Strong fit: the full load is at or under 60 per cent of capacity. The offering fits with room left for the rest of the season.

Good fit, with conditions: the full load is over 60 per cent but at or under capacity, in which case the condition is staging; or the full load is over capacity but the smallest provable version's load is at or under 60 per cent of capacity, in which case the condition is "prove the small version first, and the full version is not this year".

Poor fit, for now: the full load is over capacity and the smallest provable version is also over 60 per cent of capacity, or no smallest version exists. Load over capacity is an overdraft, and the assessment calls it one and shows the gap. A "for now" verdict says what would have to change on the canvas to revisit.

Mismatch: decided before load, as above.

The 60 per cent line is a convention of this version: it holds back four tenths of the budget for the season's other demands. Appendix D lists it as a calibration candidate.

The costing convention

Every assessment carries the same five lines, in the same order, and a total that sums from them.

Licence: as quoted, at this business's realistic seat or volume count, for twelve months.

Implementation and migration: the vendor's stated fee, or, if none, an estimate with the basis shown.

Data preparation: hours multiplied by an on-farm labour value. The hours are estimated from the data_store nodes the offering assumes and stated as a range. The labour value is a stated figure per business, defaulting to $45 an hour in this version if the business has not supplied its own; the assessment shows which was used.

Running cost: inference, tokens, per-use or per-hectare fees, at a stated volume assumption for this business. Where the vendor has not disclosed a running cost, the line reads "undisclosed; assume [figure] and verify", with the figure set at 10 per cent of the annual licence in this version, and the first of the five vendor questions asks for the real number.

What it displaces: anything the offering replaces that currently carries value (an agronomist visit, a trusted relationship, a second opinion), listed and marked unpriced if it cannot be priced.

First-year total: the sum of the priced lines, as a range where any line is a range, rounded to the nearest hundred dollars after summing. The total never carries a figure that the lines do not produce.

The neutrality rule

The instrument never recommends a product. It says whether a given offer fits this business now, and sometimes the right answer is to buy nothing yet. Verdicts belong to the business that asked. A program may report fit by offering type (voice record-keeping, carbon baselining, autonomous agronomy) across its businesses; it does not report verdicts by vendor name unless every business concerned has consented to that use. The instrument takes nothing from vendors.

Output

The assessment runs about a page and a half: the verdict line and two or three sentences of reasoning anchored in canvas nodes by name; what it touches on the canvas; change load against capacity, with both tables; the cost lens; five questions for the vendor (data assumptions, true implementation cost at this scale, exit terms and data ownership, running cost at realistic and doubled volume, the smallest provable version); and the smallest provable version.


Part 7. Calibration and validation

This part is the difference between a scoring protocol and an instrument. At version 1.0 it is a design with no results. It is published in that state on purpose, so a reader can see what the method commits to and hold it to that.

Calibration: do two scorers get the same number?

The calibration set is eight to ten real discovery transcripts, with consent, scored independently by two human scorers who have read this document and by the instrument, without seeing each other's work. For each transcript the three produce: a factor table (which factors, at what level, with which quote); a decade placement for Alignment, Roles and Readiness; and the resulting capacity.

What is reported, per version: for capacity, the mean absolute difference between scorers and the share of transcripts where the two land within ten points; for each lens, the share of transcripts where the two land in the same decade; for factors, the share of factor-level judgements on which the two agree, and a list of the factors that most often disagree. Disagreements are worked through together and the evidence tests in Parts 4 and 5 are tightened where the wording caused the split. The results and the de-identified factor tables are published with the version.

A scorer is calibrated when, across the set, their capacity readings sit within ten points of the reference scorer on at least eight of ten transcripts. Readings from an uncalibrated scorer are provisional. Readings from the instrument alone are provisional until a human has confirmed the factor levels.

Until the first calibration set has been run and published, every reading under version 1.0 carries the word "provisional" in its header.

Validation: does the number mean anything?

The method commits to three tests, restated from the Groundtruth discussion paper (June 2026) as predictions that can fail.

Capacity predicts survival. Across a cohort of businesses, the capacity reading at month zero will correlate with whether an AI tool adopted during the period is still in weekly use at month six. The reading should outperform a vendor-supplied readiness survey and outperform chance by a wide margin. A meaningful effect, not a perfect predictor.

The verdict beats the feature comparison. When the same offerings are assessed by the vendor drop and by a conventional feature-comparison matrix, the vendor-drop verdict will correlate better with the business's own twelve-month satisfaction and continued use.

Memory compounds. Engagements that re-open an existing canvas will produce strategy outputs faster, with fewer open questions, and with higher operator recognition ("that's us") in a blinded read than engagements that start from nothing.

The study that runs them: eight to twelve businesses in one region across two crop types, a discovery at month zero and a return at month six, one to three vendor offerings nominated by each business at month zero and assessed both ways, outcomes at months six and twelve. Businesses keep their quotes and canvases; aggregated, de-identified findings are published. The cohort is small and the region is one, so the findings are a first signal and are reported as such.

What would have to be shown before conviviality carried a weight

The evidence base holds conviviality (whether people like a tool enough to keep using it) as a leading hypothesis resting on the Computers Are Social Actors literature and on one unpublished field observation from a member of the Groundtruth cohort. This version carries it as a Readiness consideration and a design law, not as a scored factor. It would earn a weight if a within-subject test across a season (three personalities, same model, same operators) showed liking predicting continued use better than accuracy or latency, and that result was written up under its author's name. Until then the method does not put a number on it.

Where the tests live in the software

The scoring engine ships with the three examples in Appendix C as fixtures and with tests that fail if the engine's arithmetic drifts from this document. The evaluation set for the discovery instrument asserts, for each fixture, that every node carries a quote found in the source transcript, that the factor table sums to the stated capacity, that lens integers are among the twelve permitted values, that the cost lens sums, and that the verdict follows the rule in Part 6.


Part 8. From one business to a network

Every canvas is built on the same schema, so canvases read under the same version can be placed side by side. What the roll-up may say about them is limited on purpose.

What a roll-up reports

The distribution of capacity bands across the cohort (how many businesses sit in each of the five bands), never a mean capacity. The frequency of each capacity factor across the cohort, at each level ("six of nine name knowledge concentrated in one person; four of those at clear or severe"). The frequency of each lens light. Fit by offering type across the cohort: how many businesses a given kind of offering fits outright, fits with conditions, or does not fit yet. The open questions that recur across businesses, which is the cheapest signal a program gets about what to fund.

What it does not report

Means, rankings of businesses, or any figure that lets a reader identify a business from its position. Verdicts by vendor name, without every concerned business's consent. Any comparison across method versions or across scorers who have not passed calibration.

Minimums

Band distributions and factor frequencies are reported for cohorts of five or more businesses. Below five the roll-up reports factors present and open questions only, in prose.

Comparability

Readings are comparable only within a version and within a calibrated scorer pool. A roll-up states both. Where a cohort mixes provisional and confirmed readings, the roll-up reports them separately or labels the whole as provisional.

What a sector reading is for

A program manager does not need a league table; they need to know where to put money. The roll-up is built to answer three questions: where is capacity being spent (the factors that recur), what kinds of change fit the cohort now (fit by offering type), and what the cohort cannot yet tell us (recurring open questions). Those three map directly to the criteria a fund needs before it opens.


What is recorded

The conversation audio, if voice is used; the transcript; the canvas; the readings; the session log. Nothing else.

Who owns it

The business owns its transcript, canvas, quotes and readings. It may export them at any time in the open schema and may ask for them to be deleted. Deletion removes the identifiable record; a de-identified factor table may be retained in a calibration or validation set only with separate consent.

Who may hear or read it

The business, the consultant running the engagement, and the people the business names. A program sees only what Part 8 permits, under the program's data agreement, which the business signs before its first conversation.

Retention

Audio is deleted once the transcript is confirmed, unless the business asks otherwise. Transcripts and canvases are kept for the life of the engagement plus the return-visit window the business chooses, and deleted on request.

Where it sits

Data sits in an Australian region under the program's or the business's own tenancy where one exists, and never in a vendor's environment. Model calls carry the minimum context needed for the step and are not used to train any model.

"We're going to record this so the map is built from what you actually say, not from what we remember. You'll see every line we write down, with your words next to it. You own the map. We keep the recording until you've checked the map, then we delete it unless you want it kept. If this is part of a program, the program sees the pattern across everyone, not your farm by name, unless you tell us it can. You can pull out, or ask us to delete everything, any time."

Neutrality and the program

The neutrality rule in Part 6 binds programs as well as businesses. A program that cannot be seen to recommend vendors can run the vendor drop for its businesses, because the verdicts are the businesses' and the program sees fit by offering type.


Part 10. The engagement

Talk, map, evaluate. An operator talks about their business, the canvas maps it, and the four questions plus the vendor drop evaluate what to do next. The method holds two things apart on purpose: the instrument carries the load that scales (building the graph, holding the evidence, scoring against a fixed protocol, pricing running cost, remembering what was true last visit); the consultant carries the load that does not (drawing people out, hearing the thing under the thing, confirming a reading, naming a hard verdict in a room, deciding where to start).

Three ways in, one engine: the readiness pulse (free, two to three minutes, a provisional capacity and a seed canvas), the engagement front end (paid pre-work and consulting, producing a full canvas, a confirmed scorecard and a sequenced change strategy), and the standalone product (the canvas kept alive, vendor drops on demand, and a program licence for the roll-up).

Six phases: qualify; pre-work canvas; discovery session; score and strategy; vendor drops and decisions; return visits. Each is described in the Groundtruth engagement architecture (June 2026), which stays in force. Pricing and partnership terms are the Groundtruth cohort's and are not part of the published method.


Part 11. The research it stands on

The four questions are not Groundtruth's. Each carries decades of named work on why change succeeds or fails; Groundtruth's contribution is to put them into one instrument, read them from a conversation, and weigh technology quotes against the result.

Alignment traces to the behavioural theory of the firm, which showed that direction is set by whoever actually holds power, the dominant coalition, not by whoever holds the title (Cyert and March, 1963); to organisational network analysis, which showed the informal network of who trusts and listens to whom predicts how work moves better than the chart (Cross and Parker, 2004, building on Krackhardt); and to Kotter's finding that transformation fails without a guiding coalition and a shared picture of the destination (Kotter, 1995; 1996).

Roles traces to the Tavistock coal-mine studies, which found that changing the technology without redesigning the social organisation of the work destroyed performance, and that the two have to be designed together (Trist and Bamforth, 1951); to tacit knowledge, the recognition that people know more than they can tell and that this knowledge resists capture (Polanyi, 1966; Nonaka and Takeuchi, 1995); and to the ironies of automation, the finding that automating the routine parts of a task leaves the person responsible for the hardest parts with less practice at them (Bainbridge, 1983).

Readiness traces to the theory of organisational readiness for change as a shared belief that people are both willing and able (Weiner, 2009; Armenakis, Harris and Mossholder, 1993) and to absorptive capacity, a firm's ability to take in something new being a function of what it already knows (Cohen and Levinthal, 1990).

Strategy traces to task-technology fit, which found that a tool delivers value only where its capabilities match the task, and that a capable tool poorly matched underperforms a modest one well matched (Goodhue and Thompson, 1995), and to contingency thinking in organisation theory.

Change capacity and change load trace to absorptive capacity, to organisational readiness, and to the change-saturation finding that organisations have a finite tolerance for change. Load's construction traces to Rogers's adoption attributes: relative advantage, compatibility, complexity, trialability and observability (Rogers, 1962). The founding empirical study behind diffusion research followed hybrid seed corn through two Iowa farming communities and found farmers adopted after a neighbour had visibly proved it on nearby land, not when an expert told them to (Ryan and Gross, 1943). The method's habit of starting with a small, visible, partial trial is that finding used as a sequencing rule.

The distinction between episodic and continuous change (Weick and Quinn, 2002), against Lewin's unfreeze, move, refreeze (1947), is why the canvas persists rather than producing a point-in-time report: adopting AI has no end state to refreeze into.

The AI-era layer is newer and held with more care. That people respond socially to machines that behave socially is settled (Nass, Steuer and Tauber, 1994; Reeves and Nass, 1996; Nass and Moon, 2000). That calibrated trust, neither over-reliance nor disuse, is what makes automation safe to work beside is settled (Lee and See, 2004; Parasuraman and Riley, 1997). That conversational AI lifts the newest and least experienced workers most is a large field result (Brynjolfsson, Li and Raymond, 2023, working paper). That people like some agents more than others and use the ones they like is supported by the CASA tradition and by one unpublished field observation held by a member of the Groundtruth cohort (Walker, 2026, described in conversation, 26 June 2026, no written results at this version); the method treats it as a hypothesis, per Part 7.

Sector evidence for Australian agriculture, used in the evidence base rather than here: the RDC AI Alliance situational and gap analysis (Nous Group, November 2025), which puts AI adoption in agriculture, fisheries and forestry at 19 per cent of organisations against 40 to 45 per cent elsewhere and records seven of fifteen RDCs naming low digital and AI literacy as the main barrier; the CSIRO assessment of digital agriculture in Australia (Hansen et al., 2023); and the finding that Australian farmers have moved trust from expert advice toward peer and experiential knowledge (Mewett et al., 2021).

What is settled: that change fails on people more than on technology; that a divided or absent deciding coalition predicts failure; that work and technology must be redesigned together; that organisations have finite absorptive capacity; that fit beats features; that innovations spread by compatibility, trialability and observability, through trusted peers. What is held with care: the weights in this document; the transfer of a canon built mostly in factories, hospitals and banks to a paddock, argued case by case; and conviviality.

Full references are in the Groundtruth evidence base, Investigation 2 (the proven science of change) and Investigation 3 (how AI now behaves).


Part 12. Terms

The four questions and their canonical wording. Alignment: do the people who actually decide agree on where this is heading? Roles: has anyone worked out how the jobs change? Readiness: can the place absorb this, the people and the systems both? Strategy: does the plan fit this business, or was it copied from a template?

Lens lights: green (70 and above), amber (40 to 69), red (below 40), pending.

Change capacity: the budget a business has for absorbing change right now, 0 to 100, base 60, built from the factors in Part 5. Bands: hold (under 25); one small change (25 to 44); one change, staged (45 to 64); a couple, if they fit (65 to 79); more (80 and above).

Change load: what a proposed change would cost in that budget, 0 to 94, built from the table in Part 6.

Overdraft: load greater than capacity.

The four verdicts: Strong fit; Good fit, with conditions; Poor fit, for now; Mismatch. Triggers in Part 6.

Node types and statuses: Part 2.

Durations: the readiness pulse is two to three minutes and six fixed questions; the full discovery is thirty to sixty minutes.

Vendor: a technology vendor, agtech or anyone positioning AI. The thing a vendor sends is a quote, never a pitch.

Provisional: a reading not yet confirmed by a calibrated human scorer, or produced under a version with no published calibration set.


Appendix A. The canvas schema, version 1.0

{
  "groundtruth_version": "1.0",
  "business": {
    "name": "Example business",
    "sector": "food processing",
    "region": "Victoria",
    "people_count_estimate": 45,
    "source": "Discovery call with the general manager, 29 January 2026",
    "consultant": "the consultant running the engagement"
  },
  "nodes": [
    {
      "id": "owner",
      "type": "person",
      "label": "The general manager",
      "detail": "General manager, carries most decisions",
      "status": "neutral",
      "lens": ["alignment", "roles"],
      "evidence": "I'm usually the one who ends up deciding in the end"
    }
  ],
  "edges": [
    { "from": "owner", "to": "spreadsheets", "label": "works around", "kind": "flow" }
  ],
  "capacity": {
    "base": 60,
    "factors": [
      {
        "factor": "past_failure",
        "level": "severe",
        "impact": -18,
        "evidence": "we paid for two years and barely used it",
        "status": "live"
      }
    ],
    "score": 42,
    "band": "one small change",
    "provisional": true,
    "summary": "One sentence on what the number means for this business right now."
  },
  "lenses": {
    "alignment": { "decade": 40, "adjust": 0, "score": 45, "light": "amber", "provisional": true, "summary": "...", "evidence": ["quote one"] },
    "roles":     { "decade": 40, "adjust": -5, "score": 40, "light": "amber", "provisional": false, "summary": "...", "evidence": ["quote one", "quote two"] },
    "readiness": { "decade": 20, "adjust": 5, "score": 30, "light": "red", "provisional": false, "summary": "...", "evidence": ["quote one", "quote two"] },
    "strategy":  { "score": null, "light": "pending", "tests": null, "summary": "Not yet formed." }
  },
  "open_questions": ["Who actually owns the relationship with the co-packer?", "Not raised: recent change that landed well"],
  "sessions": [
    { "date": "2026-01-29", "mode": "discover", "summary": "Initial discovery from 30 minute call.", "method_version": "1.0", "scorer": "instrument", "confirmed_by": null }
  ]
}

Factor keys: past_failure, champion, seasonal_crunch, leadership_stretched, knowledge_concentrated, change_landed, digital_broad, digital_pocketed, crisis, systems (with a direction field, negative or positive), other (with a name field and a magnitude no greater than 8). Levels: mild, clear, severe. Factor status: live, or resolved with a resolved_on date and a resolved_evidence quote. The impact field is written by the engine from the factor and level, never by the scorer.

Strategy, once scored, carries tests: { "start_at_friction_and_champion": 25, "respects_season": 12, "first_step_within_capacity": 25, "earns_trust_back": 0 } and score as their sum.

Appendix B. The question bank and the pulse

B.1 The seventeen traced questions

Seed. "Just talk me through your place the way you'd tell a neighbour. What do you grow or run, and who's involved day to day?" Surfaces the first people, roles, systems and season nodes. All lenses, lightly.

Alignment. "When something costs real money, who actually makes the call around here?" (decision pattern; Cyert and March; Cross and Parker.) "Is there anyone whose blessing you'd need before you changed how things are done? Are the two of you on the same page about where this is heading?" (coalition agreement; Kotter.) "If I asked you and [the other decision-maker] separately what success looks like in two years, would we get the same answer?" (shared destination; Kotter.) "When you bring in something new, are you the one who decides, or does it go through someone else, a partner, the family, a board?" (locus of authority.)

Roles. "If you were away for a fortnight, what would only work if it was in your head?" (tacit knowledge, key-person risk; Polanyi.) "Who checks whose work today? If a number looked wrong, who'd catch it?" (review roles; Bainbridge.) "If a tool took the paperwork or the number-crunching off your plate, what would you do with that time, and whose job would change?" (post-change shape; Trist and Bamforth.) "Has anyone here said out loud they're worried about their job, or about a machine doing part of it?" (voiced concern; Weiner.)

Readiness. "Last time you brought in something new, a system, a tool, a different way of doing things, how did it go?" (organisational memory, the trust factor; Armenakis et al.) "Where does the information actually live, on a screen, in a book, in a shed, in someone's head?" (data stores; Cohen and Levinthal.) "Across the people here, who's comfortable with a phone or a computer and who isn't? How wide is that gap?" (spread of digital confidence.) "How much else is going on right now? Are you in a crunch, or is there a bit of room to try something?" (current load, season, crisis.)

Strategy and capacity. "If you could fix one thing that grinds on you every week, what would it be?" (friction; Goodhue and Thompson.) "Is there anyone here who'd be keen to give something new a go, who others would follow?" (champion; Rogers.) "When is changing anything just impossible on this place, and when's the quiet window?" (seasonal constraint.) "All up, does it feel like there's room to take something on right now, or is the plate full?" (the operator's own read on capacity.)

Vendor drop, put to the operator when a quote is in view. "What does this actually promise to fix for you, in your words?" "Does this lean on data you've already got in good shape, or would you have to get a lot of stuff ready first?" "Who here would have to change how they work every day for this to land? Is one of them keen?" "Does this feel like the kind of thing that's burned you or a neighbour before?"

B.2 The readiness pulse

Six fixed questions, each with three answers and a fixed impact applied to the base of 60, clamped 5 to 95. The pulse produces a provisional capacity and six seed nodes. It does not score the lenses.

Question Answers and impact
Who makes the big calls around the place? Pretty much one person (-8); a couple of us, informally (0); a clear team who agree on the direction (+10)
Where do your records actually live? In someone's head, or on paper (-12); spreadsheets, scattered about (-2); proper systems most of us can get to (+10)
Last time something new came in, how did it go? We paid for it and barely used it (-12); mixed, some of it stuck (0); it landed, and we still use it (+12)
How busy are the next few months looking? Flat out, no room to learn anything (-10); busy, but a bit of room (0); quieter, room to try things (+8)
Is there someone keen to drive new things? Not really (-6); one person with energy for it (+10); a few of us (+12)
If a key person left tomorrow, how much walks out with them? Most of what we know (-12); a fair bit (-4); not much, it's written down (+8)

The pulse's impacts are deliberately close to the "clear" level of the matching factors in Part 5, so a pulse and a discovery of the same business land in the same band more often than not. That is a claim to test in calibration, not a guarantee.

Appendix C. Three worked examples, re-scored under version 1.0

Each example shows the original reading (June 2026, free-range impacts) and the 1.0 reading (fixed levels), so the effect of fixing the protocol is visible.

C.1 A real business (thirty-minute first-contact call, 29 January 2026; name and verbatim quotes withheld)

A vertically integrated food business with several divisions and interstate depots; one voice heard, the finance lead. The verbatim quotes are held with the private canvas and are paraphrased here, which the method would not accept in a live reading; the appendix shows the arithmetic, not the evidence.

Capacity, original 45. Under 1.0:

Factor Level Impact Evidence (paraphrased)
systems (negative) clear -8 core transaction software pinned to an operating system two decades old
leadership_stretched clear -8 the finance lead also carries HR and safety for remote sites
knowledge_concentrated mild -4 one accountant produces the daily costing the sales team prices from
champion clear +9 the finance lead already uses AI for compliance drafting and gets results; a second champion in operations, absent, folded in
other: sponsor eight months in -5 the sponsor has been with the business under a year

Base 60, sum of impacts -16, capacity 44, band "room for one small change", provisional (one voice heard). Not raised: past failure, seasonal crunch, digital confidence, crisis; all in open questions.

Lenses under 1.0: Alignment placed in the forties (one keen party, deciders not named), no adjustment, 45, provisional. Roles placed in the forties (boundaries blurry, knowledge in individual heads, no job fear voiced), no adjustment, 45. Readiness placed in the forties (heavy dependence on ageing systems, spread unknown) with evidence toward the twenties (unsupported systems, and the sponsor's own doubt that anything could be built there), moved down five, 40, amber. Original readings were 42, 40 and 35; the 1.0 readings sit within one decade of each.

C.2 Mallee Grain and Cropping (persona canvas built for the demonstration, 16 June 2026)

Capacity, original 48. Under 1.0:

Factor Level Impact Evidence
champion clear +9 "Michael's the one who got FieldView going. He'll chase down any new tool"
systems (positive) clear +7 "Operations Center is live and the machine data goes straight in"
past_failure mild -6 "They've been sitting half-calibrated for a year" (told without heat)
knowledge_concentrated clear -7 "If Mark walks out the door, a fair bit of what we know walks with him"
digital_pocketed clear -6 "Some of the team have never really got their head around the precision gear"
seasonal_crunch severe -13 "No one's got time to learn anything new" (April to December, and the session is inside it)

Base 60, sum -16, capacity 44, band "room for one small change", provisional (persona data). The original 48 used -8 for a season the business was already inside; the 1.0 evidence test places a current crunch at severe.

Lenses under 1.0: Alignment forties, 45 (original 58; the original placed it in the fifties on a shared push from John and Michael, which the 1.0 test does not allow while every call routes through one person). Roles forties, moved down five on the voiced job worry, 40 (original 42). Readiness forties, moved up five on the installed base, 50 (original 46). Strategy pending.

Vendor drops against capacity 44, using the load table:

GrainLedger (voice record-keeping): disruption light 4, people moderate 10, training light 3, data none 0, with the grain 0, no echo. Load 17. 17 is under 26 (60 per cent of 44). Strong fit.

AgriCarbon AI (carbon baseline and reporting): full offering: disruption moderate 10, people light 4, training moderate 8, data heavy 24 (three seasons of clean records against a paper spray diary), grain mixed 8, history echo 8 (annual up-front billing repeats the idle-gear shape). Load 62, an overdraft of 18. Smallest provable version (baseline three paddocks against records already held, run by the champion, no annual licence): 4, 4, 3, 10, 0, 0. Load 21, under 26. Good fit, with conditions: prove the three paddocks first; the full program is not this year.

FieldMind (autonomous agronomy): disruption heavy 16, people heavy 16, training heavy 14, data heavy 24, against the grain 16 (sidelines the agronomist, desk-bound), echo 8. Load 94, an overdraft of 50. No thirty-day version fits. Poor fit, for now.

C.3 Riverbend Mixed Farming and AgriSense Pro (constructed test case, June 2026)

Capacity, original 41. Under 1.0:

Factor Level Impact Evidence
past_failure severe -18 "paid for two years, barely used it after the first month"
champion clear +9 "the young bloke is keen on all this tech"
leadership_stretched severe -13 "nights and weekends, every season"
knowledge_concentrated clear -7 "we lost half of what he knew" (recent loss, the shape repeats)

Base 60, sum -29, capacity 31, band "room for one small change". The 1.0 reading is ten points under the original because the fixed levels place a two-year unused system at severe and named nights-and-weekends at severe; the original scorer chose -15 and -8 inside the ranges. This is exactly the kind of movement calibration is meant to surface.

Lenses under 1.0: Alignment sixties (kitchen table shares direction, untested), moved down five because a second decider was not heard, 60. Roles forties, moved down five on the voiced worry about the casuals and the recent knowledge loss, 40. Readiness twenties (trust spent by the 2023 failure, records on paper), moved up five on Dean and the spreadsheets, 30, red.

AgriSense Pro, full offering: disruption heavy 16 (eight-week migration of live records), people heavy 16 (three of four), training moderate 8, data heavy 24 (a paper diary and a twelve-month history requirement), against the grain 16 (the diary in the ute becomes an app; annual billing), echo 8 (the 2023 failure). Load 88 against capacity 31, an overdraft of 57. Smallest provable version (one person, one compliance report, thirty days): 4, 4, 3, 10, 8, 0. Load 29, over 19 (60 per cent of 31). Poor fit, for now. What would change it: a landed change that earns trust back, monthly billing, and the diary digitised by Dean on his own terms first.

Cost lens under the convention: licence $7,800 (annual up front); onboarding $4,500; data preparation 50 to 80 hours at $45, $2,250 to $3,600; running cost undisclosed, assume $780 (10 per cent of licence) and verify; displaces nothing priced. First-year total $15,300 to $16,700, rounded to the hundred after summing. The original total of $15,000 to $16,500 did not sum from its rows.

Appendix D. Version history

Version 1.0, 6 September 2026. First fixed protocol. Establishes: base 60; ten canonical factors at three fixed levels; the twelve-value lens integer rule; Strategy scored on four fit tests; the load table with a history-echo modifier; the 60 per cent Strong-fit line and the smallest-provable-version rule for conditions; the Mismatch trigger; the costing convention with a $45 labour default and a 10 per cent undisclosed-running-cost default; the re-scoring rule; roll-up minimums; the calibration and validation design. Calibration candidates, to be revisited when the first set is scored: the base of 60; the severe-level impacts for past failure and crisis; the data-preparation weight of 24; the 60 per cent line; the pulse impacts; whether "other" factors recur.

Before 1.0 (June to August 2026): a scoring rubric with free ranges per factor, decade anchors without an integer rule, a load estimate without a table, thresholds stated only in software. Readings from that period are recorded as version 0.1 and are not comparable with 1.0.


Produced with research support from Claude.

Groundtruth | an open method | © the Groundtruth cohort 2026. Method text CC BY-SA 4.0; scoring engine Apache 2.0; name, mark and software reserved.

Method text: Creative Commons Attribution-ShareAlike 4.0. Scoring engine: Apache 2.0. The Groundtruth name and mark, the software and the Groundtruth cohort's services are reserved. Every reading under version 1.0 is provisional until the first calibration set is published.