Winter Family CollectivePaper 02
winterfamilycollective.comWinter Family Collective
Winter Family Collective — back to home Paper 02 · Operating leverage

Compounding with AI

A field guide for operators: where it pays back inside a services business, how to run a pilot that survives contact with clients, and what to leave alone for now.

01Start from the P&L, not the technology

Most AI programs inside operating companies begin with a tool and go looking for a job. The ones that pay back begin with a line on the income statement and go looking for the hours behind it.

In a services business the candidates are narrow and familiar: research and first drafts, quality review, project administration, sales qualification, reporting, and the internal questions that consume a senior person's afternoon. Each is measurable in hours per month before anyone buys anything.

If you cannot name the hours you expect to recover, you are not running a pilot. You are running an experiment on your clients.

Exhibit 01The gap

Adoption is close to universal. Earnings impact is not. The distance between these bars is the entire subject of this paper.

Use AI in at least one function88%
Use generative AI72%
Report enterprise-level EBIT impact39%
Attribute more than 5% of EBIT to AI~6%
Pilots with measurable P&L impact~5%

First four bars: McKinsey, The State of AI: Global Survey 2025, 1,993 respondents across 105 nations; nearly two-thirds of organisations have not begun scaling AI across the enterprise.4 Last bar: MIT Project NANDA, The GenAI Divide: State of AI in Business 2025, which found roughly 95% of enterprise generative AI pilots produced no measurable profit-and-loss return against an estimated $30–40 billion of spend.1

“Just 5% of integrated AI pilots are extracting millions in value, while the vast majority remain stuck with no measurable P&L impact.”

MIT Project NANDA · The GenAI Divide · July 20251

This paper is written for the person who has to answer for the result — a managing director, a head of delivery, an owner. It assumes no enthusiasm and no skepticism, only a budget and a quarter.

02The three-question filter

Before any pilot, three questions. A no to any of them is a no.

The third question is the one most programs skip, and it is the one that decides whether the saving reaches the income statement. An agency that recovers four hundred hours and has no demand to fill them has bought a morale problem. The same four hundred hours in a business with a waiting list is margin. Ask the commercial question before the technical one.

Exhibit 02The filter

Three questions, each answerable in a week, each fatal on its own.

QuestionWhat a no meansHow to test it in a weekEvidence to keep
Is the output checkable?You are shipping unverifiable work and calling it delegationHave a reviewer grade twenty existing outputs and time each checkMinutes per check, and the disagreement rate between two reviewers
Is the input already ours?This is a data and contracts project wearing a pilot's costumeName the systems the work draws on and who holds the rightsA one-page data map with a legal position on each source
Does the hour saved go somewhere?You are buying idle capacity, not marginAsk what the person would do with a recovered day, and whether it is billable or soldPipeline or backlog that the recovered hours will absorb

Our own filter. It exists because the failure MIT identifies is not model quality but the gap between a tool and the workflow it was bought to change.1

03Where it pays back in a services business

Ranked by how reliably we have seen the return arrive, most reliable first.

Exhibit 03Where to start

Start where the hours are concentrated and the output can be checked in minutes. Everything else can wait a year.

Start here
Checkable in minutesHard to verify
Research and first drafts
Project and account admin
Reporting assembly
Internal question answering
Sales qualification
Final client judgment
Hours are thin and scatteredHours are concentrated

Our own placement, from pilots across the group. Open points are workflows where we have either failed to bank the saving (internal question answering, where the recovered time is spread across many people) or decided not to try (final client judgment). MIT found the same asymmetry from the other direction: more than half of enterprise budgets went to sales and marketing, while the highest measured returns came from back-office work.1

Research and first drafts. The work between a brief and something a senior person can react to. High volume, checkable in minutes, and the reviewer was already reviewing. This is where nearly every services business should start.

Internal question answering. The afternoon a senior person loses to questions whose answers already exist somewhere in the company. The return is real but it is spread thinly across many people, which makes it easy to claim and hard to bank.

Project and account administration. Status notes, timelines, meeting records, the reporting nobody bills for. Unglamorous, safely checkable, and consistently underestimated.

Sales qualification. Sorting and preparing for inbound rather than deciding on it. Pays back where volume is high; a waste of effort where the pipeline is twenty relationships.

Quality review. As a first pass that catches the mechanical, never as the last pass. Useful, and the easiest place to talk yourself into removing the human signature.

Reporting. Assembly is a good candidate; interpretation is not. Draw that line explicitly, because it will move on its own if you do not.

Exhibit 04Payback ranking

Ranked by how reliably we have seen the return arrive, with our own confidence stated.

WorkflowWhere the hours sitCheckableOur evidenceVerdict
Research and first draftsDelivery, junior and midIn minutesMeasured more than onceStart here
Project and account adminDelivery and PMIn minutesMeasured more than onceStart here
Internal question answeringSpread across seniorsSometimesReal but unbankedRun it, expect no line item
Sales qualificationSales, volume-dependentSometimesOne pilot, high volume onlyOnly above real volume
Quality review, first passSenior reviewIn minutesUseful, never finalFirst pass only
ReportingFinance and accountsAssembly yes, reading noAssembly measuredDraw the line explicitly
Final client judgmentPartner and principalNoNot attemptedLeave alone

Our own ranking across six operating companies, measured against a common baseline. “Our evidence” is deliberately unflattering where it should be: two of these rows rest on a single pilot.

04How to run the pilot

Ninety days, one team, one workflow. Measure the baseline for two weeks before anything changes. Name the person accountable for the result, not for the tool.

Keep a human signature on client work. Someone with a name and a job title reviews and owns every deliverable that leaves the building. This is not a compliance formality; it is the reason the client keeps paying.

Write the disclosure before the first client sees the work, not after they ask. Clients tolerate a great deal and forgive very little.

Kill it on schedule. A pilot with no end date becomes a permanent cost with no owner.

What we measure

Exhibit 05Ninety days

One team, one workflow, a measured before, and a date on which someone decides.

Baseline measurement
wks 1–2
Access, data and legal check
wks 1–3
Client disclosure written
wks 2–3
Run the workflow
wks 3–12
Reviewer time logged
every deliverable
Mid-point read
Keep, change or kill
wk 12
Independent owner of the verdict
named on day one, not the enthusiast
Week 1Week 4Week 8Week 12

Our own pilot shape. The two weeks of baseline before anything changes are the cheapest insurance in the sequence, and the step most often skipped.

The last line deserves emphasis. A workflow that halves drafting time and doubles review time has not saved anything; it has moved the cost from a junior person to a senior one, which is usually a loss. Measure the reviewer or do not bother measuring.

Exhibit 06The reviewer test

Two pilots that both look like a success on the drafting line. Only one of them is.

Baseline month
120 hrs
Pilot A · saving banked
65 hrs
Pilot B · cost shifted
110 hrs
Drafting hoursSenior review hours

Illustrative, on the pattern we have seen twice. Both pilots cut drafting from 100 hours to 40. In Pilot B senior review went from 20 hours to 70, so total effort barely moved and the cost migrated to the most expensive person in the room. Reported as a drafting metric, Pilot B is a triumph for about two months.

05The four ways pilots fail

We have run enough of these across the group to see the same four failures repeat.

No baseline. The team never measured the before, so the after is a matter of opinion. This is the most common failure and the only one that is free to prevent.

The reviewer absorbs the cost. Output volume goes up, quality holds, and one senior person is now working evenings. The pilot reports success for two months and then the reviewer resigns.

Ownership sits with the enthusiast. The person who championed the tool is accountable for the outcome, so the evaluation is not independent. Separate the two roles even when it feels bureaucratic.

The pilot has no end date. It becomes infrastructure by inertia, appears in no budget, and is discovered a year later by finance.

Exhibit 07Failure modes

Four failures, the signal that shows up first, and the month you can catch each one.

FailureEarly signalVisible byThe fix, before you start
No baselineThe debrief argues about whether it workedweek 12Two weeks of measurement before anything changes
Reviewer absorbs the costOutput volume rises, one senior person starts working eveningsweek 5–8Log reviewer hours per deliverable from day one
Owned by the enthusiastEvery reported number is favourableweek 6Separate the champion from the person who owns the verdict
No end dateThe tool appears in no budget and has no ownera year lateWrite the kill date into the pilot brief

Our own record across the group. The wider evidence points the same way: MIT attributes the 95% to a learning and workflow gap rather than model quality,1 McKinsey finds fundamental workflow redesign to be the strongest correlate of EBIT impact,4 and Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027 on cost, unclear value and weak controls.5

“The single strongest correlation with EBIT impact is fundamental workflow redesign.”

McKinsey · The State of AI: Global Survey 20254

06What to leave alone for now

“For now” is doing real work in that heading. This list is a position on today's tools and today's contracts, not a principle. Review it annually and be honest about which items are still true and which have become habits.

07What to tell clients

Disclosure is a commercial decision disguised as an ethical one, and getting it wrong is expensive in both directions. Say nothing and you are one question away from a conversation about trust. Over-announce and you invite a discount request on work that took the same craft it always did.

What has worked for us is narrow and factual: name where the tools are used, name who reviews and owns the output, and confirm what happens to the client's data. Put it in the engagement terms rather than in a press release. Then have the answer ready in one sentence when a client asks in a meeting, because they will.

Two things not to do. Do not present efficiency as a reason to lower the price unless you have decided to compete on price. And do not let a client's enthusiasm move the human signature — the client who asks you to skip review is not the one who will accept the consequences.

08Security, contracts and the boring layer

Most of the risk in this work is not model behavior. It is data leaving a boundary it was not supposed to leave, under a contract that did not permit it.

The minimum is unglamorous: a named list of approved tools and a stated position on everything else; a check that the vendor does not train on your inputs; identity and access managed centrally so leavers actually leave; and a review of client contracts for subprocessor and confidentiality clauses that predate all of this. Most standard agreements written before 2023 do not contemplate any of it.

None of this requires a security function. It requires one person spending a week and writing down the answers.

Exhibit 08The boring layer

One week of work, seven answers, written down where someone else can read them.

ItemThe question to answerEvidence it is done
Approved tool listWhich tools are permitted, and what is the position on everything else?A named list with a date and an owner
Training on inputsDoes the vendor train on our data or our clients' data?The clause, quoted, per vendor
Identity and accessAre accounts issued centrally, so leavers actually leave?A single sign-on record and an offboarding test
Client contractsDo subprocessor and confidentiality clauses permit this use?A reviewed list of exceptions to renegotiate
Disclosure languageWhat do we tell clients, and where does it live?Two sentences in the engagement terms
RetentionWhere does the output live, and for how long?A retention setting, not an intention
Named human signatureWho reviews and owns each client deliverable?A role in the workflow, by title

Our own checklist, run once for the group rather than six times. Most standard client agreements drafted before 2023 do not contemplate any of this.

09Why a group has an advantage here

A single company runs one pilot and learns one thing. A group of companies runs six and can compare them — same measurement, same review standard, different contexts. It can also buy once, negotiate once, and set one security posture rather than six.

The collective's job is not to choose the tools. It is to make sure that when one company learns something expensive, the other five get it free.

The second-order advantage is more useful. Six companies produce six baselines, which means a claim about hours recovered can be checked against a comparable rather than against a vendor's case study. A group that measures the same way in every company knows, within a quarter, which claims are real.

Exhibit 09Where the money went

The published record of enterprise spending is a map of where not to start.

$30–40bnEstimated enterprise spend on generative AI behind the pilots studied
>50%Share of budgets directed at sales and marketing, where measured returns were lowest
5%Share of custom enterprise tools that reached production
60 → 20 → 5Percent of firms that evaluated, piloted, then deployed enterprise-grade systems

MIT Project NANDA, The GenAI Divide: State of AI in Business 2025, based on more than 300 public deployments, 52 structured executive interviews and 153 leader surveys.1 The 95% headline counts P&L impact within roughly six months of deployment; the fair criticism of it attacks that precision rather than the direction.3

10What we do not know

We are three years into this and the honest position is that our confidence is uneven. We are confident about first drafts and administration, where we have measured the return more than once. We are not confident about internal question answering, where the saving is real and almost impossible to bank. We have no evidence at all on whether any of this changes what a client is willing to pay, in either direction.

We also do not know how much of today's guidance survives the next capability step. The three-question filter should; the list of things to leave alone almost certainly will not.


11A closing note

The operators who do well with this are not the enthusiasts. They are the ones who treat it as ordinary capital allocation: a defined cost, a measured return, a date by which the answer is known, and a willingness to stop.

If you are running a services business and want to compare notes on what has and has not worked, we are glad to. We will tell you about the pilots we stopped as well as the ones we kept.


Sources and notes

Third-party research in this paper describes surveyed and published deployments, not our companies. Everything labelled as ours — the ranking in Exhibit 04, the pilot shape, the checklist, the failures — comes from pilots run inside the group against a common measurement standard, and we have marked the places where it rests on a single data point.

  1. MIT Project NANDA, The GenAI Divide: State of AI in Business 2025 (July 2025): roughly 95% of enterprise generative AI pilots produced no measurable P&L return against an estimated $30–40 billion of spend; about 5% of integrated pilots captured significant value; more than 80% of organisations had piloted tools and roughly 40% reported deployment; of enterprise-grade systems, 60% were evaluated, 20% piloted and 5% deployed.
  2. Fortune CFO Daily coverage and an interview with the report's lead author, August 2025, on the finding that failures trace to a learning and integration gap rather than model quality.
  3. Published critiques of the 95% figure: it measures P&L impact within roughly six months of deployment on a small sample, which the report itself states. We quote it because the direction is corroborated elsewhere, not because the number is precise.
  4. McKinsey QuantumBlack, The State of AI: Global Survey 2025 (fielded June–July 2025; 1,993 respondents, 105 nations): 88% use AI in at least one function and 72% use generative AI; 39% report enterprise-level EBIT impact; roughly 6% qualify as high performers attributing more than 5% of EBIT to AI; nearly two-thirds have not begun scaling across the enterprise; high performers are around three times more likely to have fundamentally redesigned workflows.
  5. Gartner, as reported in public coverage: more than 40% of agentic AI projects are expected to be cancelled by the end of 2027, and a majority of projects lacking AI-ready data are expected to be abandoned.
  6. Winter Family Collective records: six operating companies on one measurement and review standard; pilots run and stopped since 2023.

Figures are quoted as published, with their sample sizes where they are small. If a number here is wrong, write to us and we will correct it in the next revision.

Winter Family Collective is a family-owned business collective. Six operating companies, none sold. Nashville and Charleston.
hello@winterfamilycollective.com

WFC — back to home