Modelled projections and live deployments are not the same thing. Two systems running on one runtime across different companies is the only evidence that margins will compound at the portfolio level.
Now I have all the verified facts I need. Let me compose the full article.
What a PE Buyer Should Verify Before Trusting Any AI Vendor Pitch
As of 2026, more than three billion dollars has been deployed into AI roll-up vehicles targeting American service businesses. General Catalyst allocated roughly 1.5 billion dollars to its Creation Strategy. Thrive Capital launched a dedicated operating vehicle and brought OpenAI in as an equity partner. Long Lake reached approximately 100 million dollars in EBITDA in under two years and agreed to take American Express Global Business Travel private for 6.3 billion dollars. The capital is moving fast. The AI vendor pitches are moving faster. And the gap between a modelled projection and a verified mechanism has never been more expensive to confuse. Before a PE buyer commits to any AI vendor pitch, the right question is not whether the system is impressive. The right question is whether it is already running, and whether it runs the same way at the second company as it did at the first.
The Difference Between a Modelled Projection and a Verified Mechanism
A modelled projection is a spreadsheet argument. A verified mechanism is a system that has already produced the outcome it claims, in production, at a real company, with real customers and real money moving through it.
Most AI vendor pitches in the field service and home services space present the former while implying the latter. The deck shows a before-and-after: collection time drops, technician utilization rises, churn falls. The numbers are plausible. The logic is sound. What is missing is the answer to a single question: has this system actually done this, or has it been modelled to show that it could?
The distinction matters at the portfolio level because a modelled projection does not transfer. It was built for one company's cost structure, one company's customer base, one company's data. When a PE fund deploys the same vendor across three portfolio companies, the model does not replicate. The system either runs or it does not. According to FTI Consulting's 2026 Private Equity AI Radar, only 7 percent of portfolio companies describe AI as fully integrated at enterprise scale, while 43 percent are still experimenting or using it only sparingly. The gap between a fund-level AI thesis and actual portfolio-level deployment is not a slide problem. It is a verification problem.
What does "live in production" actually mean?
It means the system is making real decisions, not staging them. It means an agent dispatched a technician this morning, not that an agent could dispatch a technician given the right configuration. It means a dunning sequence ran last night and collected a payment, not that a dunning sequence has been designed and is awaiting a pilot. The word "live" should trigger a specific follow-up: show me the last thirty days of output logs. If the vendor hesitates, the system is not live.
The Transfer Test: The Only Diligence Question That Matters at Scale
The transfer test is simple: does the system run at a new company without the founding team in the room? If the answer requires a six-month implementation, a dedicated integration engineer, and a custom data migration, the system is not a portable brain. It is a bespoke consulting engagement wearing a software price tag.
For a PE fund deploying across a portfolio, portability is not a nice-to-have. It is the entire thesis. The margin expansion that justifies the AI roll-up premium comes from one brain redeployed across every company, not from one brain rebuilt from scratch at every company. The difference between those two outcomes is the difference between a twelve-times exit multiple and a write-down.
The transfer test has three components a diligence team should verify directly:
See how the WeLaunch orchestration brain handles the transfer test across facility management and home services on one shared runtime.
What Capital-First Roll-Ups Get Wrong About AI Vendor Selection
Every major AI roll-up vehicle operating today is capital-first. General Catalyst, Thrive Holdings, and Long Lake each follow the same sequence: acquire the business, then apply the AI. That sequence is not wrong. But it creates a specific vulnerability at the vendor selection stage.
When a fund buys a business and then goes looking for an AI system to run it, the vendor evaluation happens under time pressure, with a portfolio company already on the clock. The fund is not evaluating systems from a position of operational knowledge. It is evaluating pitch decks from a position of urgency. That is exactly when a modelled projection gets mistaken for a verified mechanism.
The category software that dominates field service and home services, including ServiceTitan, Jobber, Housecall Pro, and UpKeep, stops at the data layer. Each platform records what happened: the job was booked, the technician was dispatched, the invoice was sent. None of them decide what happens next. They surface the information and wait for a human to act on it. That is not a criticism of those platforms. It is a structural description of what they are built to do. The problem is that a PE fund buying a service business and layering ServiceTitan on top of it has not automated the operation. It has digitized the record-keeping. The humans still run the work.
An AI vendor that claims to go further needs to demonstrate it, not describe it. The demonstration is not a demo environment. It is a production log from a real company with real volume.
According to FTI Consulting's 2026 Private Equity AI Radar, 95 percent of AI initiatives in PE portfolios meet or exceed their original business cases, but only 17 percent significantly exceed them, and those conservatively scoped business cases mean the full value potential across portfolios remains largely untapped.
The Questions to Ask Before the Next Vendor Call
A diligence framework for AI vendor evaluation in field service and adjacent verticals should cover five specific areas. Each one separates a verified mechanism from a modelled projection.
Is the system running in production at a named reference account?
Not a pilot. Not a beta. A named account with real volume, real customers, and real financial outcomes. Ask for the reference account's name, the vertical it operates in, and the specific agents or workflows that are live. If the vendor cannot name the account, the system is not live enough to evaluate.
Can the vendor show two or three metrics moving together, not one?
A single metric is a selection artifact. A vendor that shows you one number, say a 30 percent reduction in dispatch time, has chosen the number that looks best. Ask what happened to collection time, technician utilization, and churn rate in the same period. If those numbers moved together in the right direction, the system did not trade one outcome for another. If the vendor can only produce one metric, the system may have optimized for the demo, not the operation. WeLaunch's Facility19 deployment runs eight agents plus one brain across a twenty-truck fleet, with measurable outcomes across dispatch efficiency, compliance tracking, and overtime reduction, not a single isolated number.
What is the governance architecture?
Autonomous systems that cannot be audited cannot be underwritten. Ask specifically: how does the system prevent double contact with a customer? How does it handle conflicting instructions from two agents operating on the same account? What is the escalation path when an agent encounters a decision outside its authority? The answers to these questions are not marketing claims. They are architectural facts. A system with shared state, a fast-brain suppression layer, and a logged escalation path is a system that can be governed. A system without those components is a liability that has not yet produced a claim.
What does the implementation timeline look like at the second company?
The first deployment always takes longer. The second deployment is the test. If the vendor's answer to "how long does it take to deploy at a new portfolio company" is measured in months and requires dedicated engineering resources, the system is not portable. A brain that runs on a shared runtime should deploy faster at the second company than the first, because the orchestration layer, the agent framework, and the state management are already built. The only variable is the vertical-specific configuration.
Is the data architecture clean enough to run agents on day one?
The vocabulary problem precedes the data problem, and the data problem precedes the automation problem. A portfolio company that has been running on three disconnected systems, with one definition of "active customer" in the CRM, a different definition in the billing platform, and a third definition in the dispatch tool, cannot run autonomous agents until those definitions are reconciled. Ask the vendor how they handle the data reconciliation problem before the agents go live. If the answer is "we assume clean data," the vendor has never deployed in a real service business. Read how the second brain problem blocks most automation projects before they start on the WeLaunch blog.
Why Two Deployments on One Runtime Is the Only Credible Evidence
Two live deployments on one runtime is the minimum evidence that a system will compound at the portfolio level. One deployment proves the system works. Two deployments on the same runtime prove the system transfers without a rebuild. That distinction is worth a multiple.
The math is straightforward. If a PE fund acquires ten service businesses and deploys the same orchestration brain across all ten, the cost of the brain is amortized across ten revenue streams. The margin expansion compounds because the routing data, the customer lifecycle data, and the compliance logic from company one make company two cheaper to run from day one. Every serviced job makes the next one cheaper to win because the review data and the route data are reused to find the next customer on the same street.
That compounding only happens if the brain is genuinely portable. A system that requires a custom rebuild at each portfolio company does not compound. It scales linearly at best, and it scales the cost of the implementation team alongside the revenue. That is not a portfolio playbook. That is a services business wearing a software valuation.
WeLaunch built the brain first. The Facility19 control tower runs eight agents plus one brain across a twenty-truck fleet, handling dispatch, compliance, and overtime in production. The same runtime that runs Facility19 runs the home services lifecycle for a 64,000-customer base, with roughly ten times model ROI across churn reduction, collection time, and technician hours. Two verticals, one runtime. That is the transfer test, passed. See the orchestration brain running across both deployments.
For PE buyers evaluating the AI roll-up space, the question is not which vendor has the most compelling deck. The question is which vendor can show you two live deployments on one runtime, with two or three metrics moving together at each site, and a governance architecture that can be audited before the deal closes. Explore how collection leakage and churn economics interact inside a live deployment to understand what the numbers look like before and after the system runs.
The capital-first players buy the business and then scramble to build the AI. WeLaunch built the brain first, it is live in production, and the transfer test is already on the record. That is not a positioning claim. It is a verifiable fact, and verifiable facts are the only currency that survives a diligence call.
One brain. Every portfolio company.
Take the Next Step
If you are running diligence on an AI vendor or evaluating the operational infrastructure of a service business acquisition, the fastest way to separate a verified mechanism from a modelled projection is to see the system running. See one brain across your portfolio and review the Facility19 proof point alongside the home services deployment. Or talk to WeLaunch about your portfolio and bring your specific diligence questions to a live systems walkthrough.
Frequently Asked Questions
What is the transfer test and why does it matter for PE diligence?
The transfer test asks whether an AI system runs at a new portfolio company without a custom rebuild. If the answer requires months of implementation and dedicated engineering, the system is not portable. For a PE fund deploying across multiple companies, portability is the entire thesis: one brain redeployed across every acquisition is what produces compounding margin expansion, not one brain rebuilt from scratch each time.
How do I tell the difference between a modelled projection and a live deployment?
Ask the vendor to name a reference account, show you the last thirty days of production output logs, and identify two or three metrics that moved together in the same period. A modelled projection cannot produce a production log. A live deployment can. If the vendor hesitates on any of those three requests, the system is not live enough to evaluate.
Why is a single metric not enough evidence during AI vendor diligence?
A single metric is a selection artifact. Vendors choose the number that looks best. When a system genuinely improves operations, multiple metrics move together: dispatch efficiency rises, collection time falls, and churn drops in the same period. Two or three metrics moving in the right direction simultaneously demonstrate that the system did not trade one outcome for another.
What governance questions should a PE buyer ask an AI vendor?
Ask how the system prevents double contact with a customer, how it handles conflicting instructions from two agents on the same account, and what the escalation path looks like when an agent encounters a decision outside its authority. A system with shared state, a suppression layer, and a logged escalation path can be audited and underwritten. A system without those components cannot.
Why do capital-first AI roll-ups face a specific vendor selection risk?
Capital-first roll-ups acquire the business first and then evaluate AI vendors under time pressure, with a portfolio company already on the clock. That urgency creates the conditions where a modelled projection gets mistaken for a verified mechanism. Evaluating AI systems before the acquisition, from a position of operational knowledge rather than urgency, is the structural advantage that brain-first operators hold over capital-first ones.
What does "two deployments on one runtime" actually prove?
It proves the system transfers without a rebuild. One deployment proves the system works in one context. Two deployments on the same runtime prove the orchestration layer, the agent framework, and the state management are genuinely portable across different companies and verticals. That portability is the only evidence that margins will compound at the portfolio level rather than scale linearly alongside implementation costs.
