← Back to blog
Field notes

Two Live Deployments on One Runtime Beat a Hundred Slides of Projections

PE buyers evaluating AI back office automation should ask one question: does the system run at a new portfolio company without the selling founder in the room. Modelled projections cannot answer that. A shared runtime already in production can.

Two Live Deployments on One Runtime Beat a Hundred Slides of Projections

As of 2026, more than three billion dollars has been committed to AI roll-up strategies by firms including General Catalyst, which allocated 1.5 billion dollars from its recent fund raise to its Creation Strategy, and Thrive Holdings, which launched a one-billion-dollar-plus evergreen vehicle and brought OpenAI in as an equity partner. Long Lake reached 100 million dollars in EBITDA in under two years and agreed to take American Express Global Business Travel private for 6.3 billion dollars. The services economy is worth 16 trillion dollars. Software is worth 1 trillion dollars. Every PE partner in the room knows the math. The question that separates a real investment from a modelled projection is not whether the AI back office automation thesis is correct. It is whether the system in front of you actually runs at a new portfolio company without the selling founder in the room. That is the transfer test. And a hundred slides of projections cannot answer it. Two live deployments on one shared runtime can.

Why the Transfer Test Is the Only Diligence Question That Matters

A system that runs in production at two unrelated companies on a single shared runtime has already passed the transfer test. Every other claim about AI back office automation is a projection until that condition is met.

Most AI vendors arrive at a diligence call with a compelling deck. They show a workflow diagram, a projected EBITDA lift, and a reference customer who happens to be the founder's first client. What they cannot show is the system running at a second company that the founder did not personally configure. That gap is where most AI roll-up theses quietly collapse after close.

The transfer test is not a metaphor. It is a specific operational question: remove the founding team from the room, point the system at a new portfolio company, and watch what happens. Does the orchestration brain re-route dispatch without a human rewriting the rules? Does the dunning agent collect on overdue invoices without someone manually adjusting the cadence? Does the checkout agent close jobs and trigger invoicing without a technician calling the office? If the answer to any of those questions requires the original builder to be present, the system is not a system. It is a consulting engagement wearing a software label.

What does a verified mechanism look like versus a modelled projection?

A modelled projection says: "Based on our pilot, we expect a 40 percent reduction in dispatch labor at your portfolio company." A verified mechanism says: "The orchestration brain ran dispatch for a twenty-truck facility fleet and a 64,000-customer home services lifecycle on the same runtime, with no shared staff, no shared configuration, and no founder involvement after go-live." The first is a forecast. The second is a receipt. PE diligence should demand the receipt before the forecast is even opened.

The WeLaunch orchestration brain has passed this test twice. The Facility19 control tower runs eight agents plus one brain across a twenty-truck fleet, handling dispatch, compliance tracking, and overtime authorization. A separate deployment runs the full customer lifecycle for a 64,000-customer home services business, covering lead acquisition, booking, service dispatch, review capture, invoicing, collection, and reactivation. Both deployments run on the same shared runtime. Neither required the founding team to remain embedded post-launch. That is not a projection. That is a production record.

The Capital-First Problem Every AI Roll-Up Faces

Every major player in the AI roll-up space is capital first. They buy the business, then scramble to build the AI. WeLaunch is the inverse: the brain was built first, it is live in production, and the capital conversation follows from there.

General Catalyst's Creation Strategy, Thrive Holdings, and the broader class of AI-enabled roll-up vehicles share a structural sequence: identify a fragmented service vertical, acquire a platform company, then deploy AI to lift EBITDA margins from the 5 to 10 percent range typical of traditional services toward the 30 to 40 percent range associated with software businesses. The thesis is sound. The execution gap is real. Building the AI after the acquisition means the portfolio company runs on legacy operations while the technology is being constructed. The EBITDA lift is modelled, not measured, for the first twelve to eighteen months post-close.

That gap is expensive. A portfolio company running on manual dispatch, paper-based invoicing, and a billing team of three people at two hundred thousand dollars each is burning payroll against a projected automation that has not shipped yet. The model says the system will recover that cost. The diligence question is: has it recovered it anywhere else first?

When the orchestration brain is already live at two deployments before the acquisition conversation begins, the answer is yes. The PE partner is not buying a projection. They are buying a runtime that has already demonstrated it can transfer.

What One Brain Across a Portfolio Actually Means for EBITDA

One shared runtime redeployed across every portfolio company is not a cost efficiency story. It is a compounding EBITDA story, and the math runs in both directions.

Consider the cost structure of a typical field service or home services business before the orchestration brain touches it. Dispatch runs on a combination of tribal knowledge, a scheduling tool like ServiceTitan or Jobber, and two to three coordinators whose entire job is translating between the tool and the technician. Billing runs on a separate system. Collections run on a third. The VIP customer list lives in a spreadsheet that only one person maintains. When that person leaves, the list degrades. When the billing system and the scheduling system disagree on job status, someone makes a phone call to resolve it. That phone call costs money every time it happens.

ServiceTitan, Jobber, and Housecall Pro record the work. They do not run it. The data sits in the platform. The decisions still require a human to read the data and act. That is the category boundary every field service management platform stops at, and it is the boundary the WeLaunch orchestration brain crosses.

When the brain runs dispatch, the fast brain router suppresses double contact so no customer receives two calls about the same job. Agents share state so the dispatch agent and the checkout agent never operate on conflicting job records. The dunning agent runs the collection sequence autonomously, with every contact logged and auditable. The result is not one metric improving in isolation. Across the Facility19 deployment, dispatch efficiency, compliance documentation, and overtime authorization all move together. Across the home services deployment, the roughly ten-to-one model ROI reflects churn reduction, collection-time compression, and technician-hours recovered simultaneously. Three metrics moving together demonstrate the system did not trade one outcome for another.

For a PE portfolio, the implication is direct. One brain, redeployed at close across every acquired company, means the EBITDA lift is not modelled at each new acquisition. It is measured from the first week of operation, because the runtime has already run this before.

According to McKinsey research, companies that redesign their workflows around AI are twice as likely to see measurable financial impact compared to those that layer AI onto existing processes without structural change. Source: McKinsey, The State of AI: How Organizations Are Rewiring to Capture Value, 2025.

The Governance Layer Is What Makes Autonomy Safe to Underwrite

Autonomous back office systems raise a legitimate governance question for any PE buyer: what happens when the system makes a wrong decision? The answer is not reassurance. It is architecture.

The WeLaunch orchestration brain is built with the hard 20 percent explicitly reserved for human judgment. The fast brain suppresses double contact. Agents share state so no two agents act on the same customer record simultaneously. Every agent action is logged and auditable. The system does not operate in a black box. It operates in a documented, traceable sequence that a compliance reviewer, a diligence team, or a portfolio operations partner can inspect at any point.

For compliance-heavy verticals, the architecture matters beyond operational efficiency. A facility management deployment that handles compliance documentation autonomously needs a clear audit trail. A legal back office running ten custom agents on-premise needs to demonstrate that every billing action, every intake record, and every drafted document is traceable to a specific agent action at a specific timestamp. The WeLaunch system produces that record as a byproduct of normal operation, not as a separate compliance layer bolted on afterward.

This is the governance argument that makes autonomy safe to underwrite. Not that the system never makes a mistake, but that every action is logged, every escalation path is defined, and the hard decisions stay with humans. See how the Facility19 control tower handles compliance documentation and escalation in a live twenty-truck deployment.

How to Read a Vendor Pitch After You Have Seen a Live Runtime

Most AI back office vendors arrive at a PE diligence call with three things: a pilot result from one customer, a projection model showing EBITDA lift at scale, and a reference call with the founder's first client. None of those three things answer the transfer test.

A pilot result from one customer tells you the system worked when the founding team was present and invested in the outcome. A projection model tells you what the team believes will happen at a new company. A reference call with the first client tells you the relationship is intact, not that the system runs independently.

The questions that separate a verified mechanism from a modelled projection are specific. Ask how many distinct companies the system has run in production, not including the founding team's own operations. Ask whether the second deployment required the founding team to be embedded post-launch, and for how long. Ask to see the agent action logs from a live deployment, not a demo environment. Ask what the system does when a technician misses a job at 7 a.m. on a Tuesday, and whether that escalation path is automated or requires a human coordinator to notice the gap.

If the vendor cannot answer those questions with production records, the pitch is a projection. Explore how the WeLaunch orchestration brain is structured across its two live deployments before the next diligence call.

What should a PE buyer ask before signing an AI back office vendor?

Ask for production logs from at least two distinct deployments on the same runtime. Ask whether the system ran without the founding team embedded post-launch. Ask what the escalation path looks like when an agent encounters an edge case, and whether that path is documented and auditable. Projections are not receipts. Demand the receipts.

The Density Argument: Why the Second Deployment Is Worth More Than the First

The compounding argument for a shared runtime is not just cost efficiency. It is density. Every deployment teaches the orchestration brain something the next deployment inherits.

In the Facility19 deployment, the routing logic that Dex runs across a twenty-truck fleet accumulates route data, job completion patterns, and technician performance signals. That data does not stay inside one deployment. It informs the routing heuristics available to the next facility management company that comes onto the runtime. The second company starts with a brain that has already seen twenty trucks, multiple compliance scenarios, and a full overtime authorization cycle. It does not start from zero.

In the home services deployment, the 64,000-customer lifecycle produces review data, route density data, and reactivation response patterns. The next home services company on the runtime inherits those patterns. The cost to win the next customer on the same street goes down because the review and the route data from the previous job are already in the system.

This is the loop, not a slice. Lead, book, dispatch, service, review, invoice, collect, and back to lead. Every completed cycle makes the next one cheaper to run and cheaper to win. For a PE portfolio, that means the fifth acquisition on the runtime is not just cheaper to operate than the first. It is cheaper to acquire customers for, because the density compounds across every prior deployment.

Read how collection leakage differs from billing churn and why the loop matters for revenue retention across a portfolio.

Frequently Asked Questions

What is the transfer test and why does it matter for PE diligence?

The transfer test asks whether an AI back office system runs at a new portfolio company without the selling founder present. It matters because most AI vendor pitches are built around a single pilot where the founding team was embedded. A system that passes the transfer test has already run in production at a second, unrelated company on the same runtime, without post-launch founder involvement. That is the difference between a verified mechanism and a modelled projection.

How does a shared runtime differ from a single-customer AI deployment?

A single-customer deployment proves the system works in one context with one team. A shared runtime means the same orchestration brain, the same agent framework, and the same shared state layer run across multiple distinct companies simultaneously. The second company inherits the routing heuristics, escalation logic, and compliance documentation patterns that the first company's deployment produced. The brain compounds across deployments rather than starting from zero each time.

Why do platforms like ServiceTitan and Jobber not solve the back office automation problem?

ServiceTitan, Jobber, and Housecall Pro record the work. They surface data about dispatch, invoicing, and customer history. They do not make decisions, run escalation sequences, or close the loop from invoice to collection autonomously. A human coordinator still reads the data and acts. The WeLaunch orchestration brain crosses that boundary: it runs the dispatch sequence, suppresses double contact, triggers checkout, and executes the collection cadence without a human in the loop for the routine 80 percent of decisions.

What does the governance layer look like in a live deployment?

Every agent action is logged and auditable. The fast brain router suppresses double contact so no customer receives two calls about the same job. Agents share state so no two agents act on conflicting records simultaneously. The hard 20 percent of decisions, those requiring judgment about edge cases, escalations, or compliance exceptions, are routed to a human. The system does not operate in a black box; it produces a traceable record of every action as a byproduct of normal operation.

How does the density argument apply to a PE portfolio with multiple service companies?

Each deployment on the shared runtime contributes route data, customer response patterns, and compliance documentation to the brain. The next company that comes onto the runtime inherits those patterns. Customer acquisition costs fall because review data and route density from prior deployments are already in the system. The fifth acquisition on the runtime is cheaper to operate and cheaper to win customers for than the first, because density compounds across every prior deployment.

What proof points exist for the roughly ten-to-one model ROI claim?

The roughly ten-to-one model ROI figure comes from the 64,000-customer home services lifecycle deployment, where churn reduction, collection-time compression, and technician-hours recovered moved together across the same deployment period. No single metric is isolated. The three outcomes moving simultaneously demonstrate the system did not trade one result for another, which is the standard any serious diligence process should apply before accepting a single-metric ROI claim.

One brain. Every portfolio company.

See the Runtime Before the Next Diligence Call

The two live deployments are not a pitch. They are a production record. If you are evaluating AI back office automation for a portfolio company or preparing for a fund-level deployment, the conversation starts with the runtime, not the deck.