Does the system run at a new company without the founder present. This piece defines the transfer test, explains why modelled projections fail it, and shows what two live deployments on one runtime actually prove.
Now I have all the verified facts I need. Let me write the complete article.
The One Question PE Buyers Should Ask Every AI Vendor Before the LOI
As of 2026, more than three billion dollars has been deployed into AI roll-up vehicles by firms including General Catalyst, which allocated 1.5 billion dollars from its latest fundraise to its Creation Strategy, and Thrive Capital, which launched a one-billion-dollar vehicle in April 2025 and subsequently brought OpenAI in as an equity partner. The capital is real. The urgency is real. What is not always real is the AI. Before any letter of intent moves forward, a PE partner needs to ask one question that cuts through every demo, every projection deck, and every vendor reference call: does the system run at a new company without the founder present? That is the transfer test. Everything else is a pitch.
What the Transfer Test Actually Measures
The transfer test is not a metaphor for scalability. It is a specific, binary operational question: can the system be dropped into a new portfolio company, with a new team, in a new market, and produce measurable outcomes without the vendor's founding team on-site to interpret, adjust, or manually intervene?
Most AI vendors fail this test before the question is even finished. The failure is not always dishonesty. It is structural. A system built around a single founder's institutional knowledge, a single client's data schema, or a single integration environment is not a portable runtime. It is a consulting engagement wearing a software label. The distinction matters enormously at the portfolio level, because a consulting engagement does not redeploy. It restarts from scratch at every new company, consuming the same implementation budget, the same onboarding time, and the same operational disruption each time.
Why does the transfer test matter more than a reference call?
A reference call tells you the system worked once, for one operator, under conditions the vendor helped design. The transfer test tells you whether the system works without the vendor in the room. Those are different claims. The first is a proof of concept. The second is a proof of mechanism. PE buyers need the second one, because the value creation thesis in an AI roll-up depends on redeployment, not on a single successful pilot.
The distinction between a proof of concept and a proof of mechanism is where most vendor pitches quietly collapse. A proof of concept says: under controlled conditions, with our team present, the system produced this result. A proof of mechanism says: the system produced this result, then produced it again at a different company, on the same runtime, without rebuilding the underlying logic. Only the second claim survives diligence.
Why Modelled Projections Fail the Transfer Test
Modelled projections are not evidence of a running system. They are evidence of a spreadsheet. The gap between the two is where most AI vendor pitches live, and it is the gap that PE buyers most consistently underweight before signing.
A modelled projection starts with a real number, typically a labor cost, a churn rate, or a collection cycle time, and applies an assumed automation rate to produce a projected EBITDA improvement. The math is usually coherent. The problem is that the math assumes the system will behave in a new environment the way it behaved in the model. That assumption has never been tested. The model does not know that the new portfolio company runs its customer data across three disconnected systems. It does not know that the dispatch team has a manual override habit that breaks automated routing logic. It does not know that the billing cycle at the new company runs on a different cadence than the one the system was trained against.
These are not edge cases. They are the normal condition of every service business that has not been purpose-built with an AI-native back office from day one. McKinsey's research on AI at scale consistently finds that the organizations seeing real returns are the ones that redesigned workflows around the AI system, not the ones that layered AI onto existing processes. A modelled projection assumes the latter. A live deployment proves the former.
According to McKinsey's State of AI research, 88 percent of organizations now use AI in at least one function, but only 6 percent qualify as high performers seeing significant enterprise-wide value. The gap between adoption and impact is not a technology problem. It is a deployment and workflow problem.
What does a modelled projection hide?
Three things, consistently. First, it hides the data reconciliation cost. Before any agent can run dispatch, dunning, or checkout autonomously, the underlying data has to be clean, consistent, and accessible. In most acquired service businesses, it is none of those things. The cost of getting it there is real, it is measured in weeks and headcount, and it does not appear in the projection. Second, it hides the integration dependency. A system that runs on one CRM, one FSM platform, and one billing stack does not automatically transfer to a portfolio company running different tools. Third, it hides the governance gap. An agent that works without guardrails in a controlled environment will double-contact customers, create duplicate work orders, and generate compliance exposure in a live environment where state is not shared across agents. None of that shows up in the model.
What Two Live Deployments on One Runtime Actually Prove
Two live deployments on one runtime prove something a hundred slides cannot: that the system is portable. Not theoretically portable. Actually portable, in the sense that the same orchestration brain, the same agent framework, the same shared state layer, and the same MCP connectors ran at company A and then ran at company B without being rebuilt from scratch.
This is the standard WeLaunch holds itself to. The orchestration brain is horizontal. It is not a facility management product or a pest control product or a legal product. It is a runtime that runs agents, routes decisions between a big brain and a fast brain, suppresses double contact by sharing state across agents, and logs every action for audit. The vertical agents, named systems like Dex for dispatch, Molly for checkout, and Iris for overtime management, are proof points built on top of that runtime. They demonstrate what the brain can do in a specific vertical. They are not the product. The brain is the product.
At the Facility19 control tower, eight agents plus one brain run a twenty-truck fleet. Dispatch is automated. Compliance is tracked. Overtime is managed. The technician-hours saved, the compliance events logged, and the dispatch decisions made without human intervention are all measurable and auditable. That is one deployment. When the same runtime runs a 64,000-customer lifecycle for a home services business, automating the dunning, renewal, and winback cycle at roughly ten times model ROI, that is two deployments. Same brain. Different vertical. Different agent configuration. No rebuild. That is what portability looks like when it is real rather than modelled.
For a deeper look at how the orchestration brain handles the dispatch layer specifically, see how the system routes decisions across a live fleet.
The Capital-First Problem and Why It Matters for Diligence
Every major AI roll-up vehicle operating as of 2026 is capital-first. General Catalyst raises the fund, identifies the service vertical, acquires the business, and then builds or deploys the AI. Thrive Holdings, which took Amex Global Business Travel private in a 6.3-billion-dollar transaction with Long Lake, follows the same sequence. OpenAI's equity stake in Thrive Holdings, announced in December 2025, embeds engineering teams inside portfolio companies after acquisition to accelerate AI adoption. The AI comes after the capital. The operational transformation follows the deal.
That sequence is not wrong. It is just slow, and it is expensive. Every month between acquisition and operational AI deployment is a month of unrealized EBITDA improvement. Every new portfolio company that requires a fresh implementation cycle is a month of engineering cost that does not compound. The capital-first model buys the distribution and then scrambles to build the brain. The brain-first model arrives at the acquisition with the runtime already live, already proven across multiple deployments, and already capable of running the new company's back office within weeks rather than quarters.
The PE partner's diligence question is therefore not just "does this AI work?" It is "does this AI work without us rebuilding it for every company we buy?" The transfer test is the answer to that question. Read how the transfer test applies across a multi-company portfolio for the full framework.
What should a PE buyer ask to verify the transfer test?
Four questions, in order. First: name two companies, not in the same vertical, where this runtime is live in production today. Not piloting. Not in implementation. Live, meaning the system is making autonomous decisions, logging them, and producing measurable outcomes. Second: what was the time from signed agreement to first autonomous action at the second deployment? If the answer is longer than the first deployment by more than a factor of two, the system is not portable. It is being rebuilt. Third: what happens when a technician misses a job at 11 PM on a Tuesday? Walk through the exact sequence of agent actions, escalation logic, and human handoff. If the answer requires a founder to explain it from memory rather than pointing to a logged audit trail, the system is not running. A person is running it. Fourth: show the shared state layer. If agents are operating without a mechanism that prevents double contact and collision, the system will create customer experience failures at scale. That is not a theoretical risk. It is a predictable operational outcome.
Governance Is What Makes Autonomy Safe to Underwrite
The instinct in PE diligence is to focus on the upside: the EBITDA improvement, the labor cost reduction, the churn recovery rate. Those numbers matter. But the governance architecture is what determines whether the upside is real and sustainable, or whether it is a single-quarter result that degrades as the system encounters conditions it was not designed for.
A well-governed AI back office has three properties that are each independently verifiable. First, agents share state. No agent contacts a customer, dispatches a technician, or initiates a collection action without checking whether another agent has already done so. The suppression of double contact is not a feature. It is a structural requirement for any system operating at scale across a customer base of tens of thousands. Second, humans own the hard decisions. The system handles the high-volume, rules-driven work: routing, scheduling, invoicing, dunning, renewal triggers. The 20 percent of decisions that require judgment, exception handling, or customer relationship management stay with humans. The system escalates cleanly and logs the escalation. Third, everything is auditable. Every agent action, every decision, every escalation, and every customer contact is logged with a timestamp and a reason. That log is not just a compliance artifact. It is the evidence base that makes the transfer test verifiable. A buyer can look at the log from deployment one and compare it to the log from deployment two. If the decision patterns are consistent, the system is portable. If they are not, the system is being manually tuned by the vendor's team between deployments.
For operators in compliance-heavy verticals, the audit log is also the foundation of any SOC 2 posture. A system that logs every action by default is a system that can be audited. A system that requires manual documentation of agent decisions is a system that will fail a compliance review. See how compliance and diligence intersect in field service operations for the specific standards that apply.
The Density Argument: Why Portability Compounds
The transfer test is not just a diligence filter. It is the mechanism by which the value of an AI back office compounds across a portfolio. Every deployment on the same runtime produces data that makes the next deployment faster. Route data from a twenty-truck facility fleet informs dispatch logic for the next fleet. Customer lifecycle data from a 64,000-customer home services business informs renewal and winback timing for the next home services acquisition. Review data from completed jobs feeds back into lead acquisition logic, making the next customer on the same street cheaper to win.
This is the density argument. It is not available to a capital-first model that rebuilds the AI at every acquisition. It is only available to a brain-first model where the runtime accumulates operational intelligence across deployments. The PE partner who asks the transfer test question before the LOI is not just protecting against a bad vendor. They are identifying whether the AI investment compounds or merely repeats.
The services economy is worth sixteen trillion dollars. Software is worth one trillion dollars. The gap between those two numbers is the opportunity. But the opportunity only accrues to the operator whose system runs at the next company without starting over.
One brain. Every portfolio company.
Take the Next Step
If you are running diligence on an AI vendor or evaluating a back office automation claim before an LOI, the transfer test is the right frame. WeLaunch's orchestration brain is live across multiple deployments, on one runtime, with auditable logs from each. See one brain across your portfolio or talk to WeLaunch about your portfolio to walk through the transfer test against a live system.
Frequently Asked Questions
What is the transfer test in AI vendor diligence?
The transfer test asks whether an AI system can run at a new company, with a new team, without the vendor's founding team present to interpret or adjust it. It distinguishes a portable runtime from a consulting engagement that restarts from scratch at every deployment. For PE buyers, it is the single most important question before signing an LOI with an AI vendor.
Why do modelled projections fail as evidence of a working AI system?
Modelled projections apply an assumed automation rate to a known cost base and produce a projected EBITDA improvement. They do not account for data reconciliation costs, integration dependencies, or governance gaps that appear only in live production environments. A projection is evidence of a spreadsheet. Two live deployments on one runtime are evidence of a mechanism.
What is the difference between a proof of concept and a proof of mechanism?
A proof of concept shows the system worked once, under controlled conditions, with the vendor's team present. A proof of mechanism shows the system worked, then worked again at a different company on the same runtime, without being rebuilt. PE buyers need the second. The first is a demo. The second is a portfolio playbook.
How does shared agent state prevent operational failures at scale?
When multiple agents operate across a large customer base without shared state, they will double-contact customers, create duplicate work orders, and generate compliance exposure. Shared state means every agent checks what every other agent has already done before taking action. It is not a feature. It is a structural requirement for any autonomous system operating at scale, and it is one of the first things a PE buyer should verify in a live system walkthrough.
What makes an AI back office auditable enough for PE diligence?
Every agent action, decision, escalation, and customer contact must be logged with a timestamp and a reason. That log must be accessible without vendor assistance. An auditable system is one where a buyer can compare the decision log from deployment one to the decision log from deployment two and verify that the system is behaving consistently, not being manually tuned between deployments.
How does the brain-first model differ from the capital-first AI roll-up model?
Capital-first models acquire the business and then build or deploy the AI, meaning every new acquisition requires a fresh implementation cycle. Brain-first models arrive at the acquisition with a live, portable runtime already proven across multiple deployments. The brain-first model compounds operational intelligence across the portfolio. The capital-first model repeats the implementation cost at every company it buys.
