Outsourcing QA testing to a vendor means buying into one of five incompatible billing models, per-test, per-seat, per-run, per-parallel-worker, or annual retainer, which is why QA automation vendor pricing rarely translates into a real budget line. This roundup checks vendor pricing pages fetched on 2026-08-14 against our own existing pricing breakdowns for QA Wolf, mabl, Testim, BrowserStack, Functionize, and testRigor, all held against one fixed reference workload, and marks every unpublished figure quote-only rather than guessing at it.
Every vendor comparison we'd read stopped at feature checklists. A widely circulated vendor roundup in this category puts five vendors side by side and never prints a single dollar figure across any of them. We already have close to twenty single-vendor pricing breakdowns on this site and had never put them side by side against one yardstick. This is that roundup, and the honest version of it has more quote-only cells than dollar figures.
This is written for the QA lead, QA manager, or engineering manager who owns quality without a dedicated QA function, at a Series A through Series C company with a real test suite and release cadence. If you have no tests today, this isn't your article; that's a different starting point served elsewhere on this site. If your worry is whether your AI-generated tests are trustworthy rather than what they cost, that's also a different question with its own answer. And if you haven't settled whether outsourcing QA testing makes sense at all, the decision itself comes first; this article assumes you're past that and are now building the line item, which is the resourcing question underneath any real test automation strategy. If you're still weighing manual coverage against automating in the first place, that comparison lives in our manual-vs-automated testing cost breakdown; this article assumes you've already decided to automate and are now pricing vendors.
Five pricing models and what each one punishes
A QA automation invoice answers one question: what does the vendor want you to do less of? Five billing models dominate this category, and each quietly taxes a different behavior.
Per-test pricing, the model behind QA Wolf's fully managed coverage tier ("pay for tests under management," in QA Wolf's own words), charges by the number of scenarios covered. It punishes suite growth: every new user flow you want covered raises the recurring bill, whether or not it runs that month.
Per-seat pricing, the model behind Testim's Tricentis-owned licensing, charges by the number of people who can log in. It punishes wide developer access: hand ten engineers a login and you pay for ten seats, even if only two touch the suite that week.
Per-run pricing, the model behind mabl's cloud-run credit pool, charges by execution. It punishes CI frequency: a team running its full suite on every merge burns a monthly credit allotment in hours rather than weeks, a dynamic our own mabl pricing breakdown worked through with real numbers.
Per-parallel-worker pricing, the model behind BrowserStack, Sauce Labs, and LambdaTest, charges by simultaneous browser sessions. It punishes wanting your suite to finish fast: more parallels means shorter CI waits and a larger bill, a tradeoff our BrowserStack cost breakdown modeled at length.
Annual retainer pricing, most visible in a fully managed service scoped to a fixed annual sum regardless of how much actually runs, punishes low utilization. Sign a contract sized for a year of coverage, use two-thirds of it, and the other third is still on the invoice. Certification requirements like SOC 2 or PCI can also push a vendor from a published tier into quote-only territory, since a compliance review is part of the quote. That retainer shape is also where outsourcing QA testing to an agency and licensing a platform stop being the same purchase, a split we worked through in our comparison of test automation services pricing against an in-house hire.
Every billing model taxes a different behavior. Read five vendor pricing pages back to back and the numbers describe five different questions, which is why they never line up into one budget line.
None of these five models is wrong for every team. They are simply not comparable, which is why five vendor pages never line up.
A reference workload for normalizing
To compare six billing models against one yardstick, we're fixing a workload rather than trusting a vendor's own framing of it. Nothing in this section is a vendor price. These are the assumptions this comparison holds constant so every vendor's model gets asked the identical question.
The reference team runs 150 end-to-end tests, a suite size consistent with a company that already has a real regression suite rather than a five-test smoke check. It triggers 100 runs per week, roughly 20 CI-triggered merges per weekday with full regression each time, the cadence of a team shipping multiple times a day. It has 8 developers who need platform access, enough to make per-seat pricing bite without implying the dedicated QA headcount this ICP doesn't have. And it runs 5 parallel workers, enough concurrency to keep a 150-test suite's CI feedback under about ten minutes without paying for capacity it doesn't need. The browser assumption is desktop Chrome and Firefox only: no real mobile devices, no visual-regression snapshots. A mobile-heavy or visual-heavy suite is a different reference workload entirely, and BrowserStack, Sauce Labs, and LambdaTest all charge extra for that coverage.
Hold those five numbers in mind. When a cell below says quote-only, the vendor will not tell you, in public, what this exact workload costs.
What outsourcing QA testing costs at six vendors
Here's what six vendors' pricing pages, checked on 2026-08-14, and our own prior analysis where no live page publishes a number, say about this workload:
| Vendor | Pricing model | Entry price | Price at reference workload | Source date |
|---|---|---|---|---|
| QA Wolf | Usage-metered self-serve; quote-only managed tier | $0.01/credit + $0.15/runner-min | Quote-only (managed tier) | Vendor page, 2026-08-14 |
| mabl | Per-run, cloud-run credits | Quote-only | Quote-only | Vendor page, 2026-08-14 |
| Testim (Tricentis) | Per-seat, annual-only commitment | Quote-only | Quote-only | Vendor page, 2026-08-14 |
| BrowserStack | Per-parallel-worker, plus seats for manual | $99/mo (Automate, 1 parallel) | Not published past 1 parallel | Vendor page, 2026-08-14 |
| Functionize | Per-credit/seat tiers, quote-only Enterprise | $20/mo (Pro, 400 credits) | Not computable, credit rate undisclosed | Vendor page, 2026-08-14 |
| testRigor | Plan-tier (Free, Pro, Enterprise) | Quote-only | Quote-only | Our post, 2026-06-11 (3rd-party est.) |
Ordered by disclosure, not price. Three publish an entry figure, three publish nothing, and none publish enough to compute an annual total at the reference workload.
BrowserStack's entry price in the table is the $99/month Automate Desktop plan rather than the cheaper $59/month Automate Chrome plan, because the reference workload above specifies desktop Chrome and Firefox together, and the Chrome-only plan doesn't cover Firefox.
Two things stand out. First, four of these six rows are quote-only at the exact reference workload above, and a fifth, BrowserStack, is quote-only past a single parallel session. That isn't a research gap on our end: mabl and Testim publish no dollar figure at all, and BrowserStack's page states volume discounts exist without listing them. Second, two vendors changed shape since our own earlier analysis. QA Wolf's site, fetched on 2026-08-14, now lists a self-serve Platform tier billed at $0.01 per AI credit plus $0.15 per runner-minute, alongside the fully managed Coverage-as-a-Service tier that remains quote-only; our June 2026 analysis described only the latter. Functionize's site, also fetched on 2026-08-14, now lists self-serve Free, Pro, and Max tiers starting at $20/month, a shift from the enterprise-gated posture our June 2026 analysis found; Functionize's Enterprise tier is still quote-only, and neither discloses the credit-to-run conversion rate, which is why neither entry price carries through to the reference-workload column. testRigor adds a third routing fact worth noting plainly: as of 2026-08-14, its homepage navigation's link labeled "Pricing" routes to /sign-up/, and testrigor.com/pricing returns a 404.
The honest reading here isn't "vendor X is cheap, vendor Y is expensive." It's that the pricing model, not the sticker, decides whether the question is even answerable. A per-parallel-worker vendor will quote you one parallel and nothing for five. A per-run vendor won't publish the dollar rate behind its own credit pool. Fixing the pricing model is the precondition for the question to mean anything; for most of this category, the vendor still answers it in a sales call, not on a web page.
How Autonoma decides what re-runs
Every model in the table above meters something: tests, seats, runs, parallel workers, or reserved capacity. Which meter you land on is the vendor's decision. How much of it a team actually consumes is decided somewhere else entirely, by what triggers a run and who repairs the suite afterwards, and that part is worth evaluating before any number enters the conversation.
That second question is the one we built Autonoma around. A Planner agent reads the codebase, its routes, components, and user flows, and generates test cases from what is actually there, so the suite starts from the application rather than from a recorded session someone has to re-record when the UI moves. An Executor agent runs those cases against a live preview environment. A Reviewer agent classifies each result as a real bug, an agent error, or a stale test, which is the triage step that otherwise lands on an engineer every morning. And the Diffs Agent reads each pull request's code diff to work out which tests in a suite like the 150-test reference workload are relevant to that specific change, rather than treating every merge as a reason to re-run all of them.
That is a set of claims about mechanism, not about what anything costs. Scope is worth naming just as plainly: Autonoma covers the browser-rendered behavior of your web application, so native mobile, load and performance testing, and accessibility scanning stay separate lines in the budget, the same categories BrowserStack, Sauce Labs, and LambdaTest charge extra for above.
None of that tells you what a given vendor will quote you, ours included, and this article has already made the case that the quote is where this category keeps its answers. What it does give you is a set of questions that survive the sales call: how do tests get written in the first place, what decides that they run again, and who fixes them when the interface changes underneath. Those three answers shape the invoice long after the rate on it is agreed.
The test automation cost nobody quotes
Pick any row in the table above, and one cost never appears on the invoice: the hours someone on your team spends every week keeping the suite green. A per-test vendor's managed team absorbs some of that labor inside the fee; a per-seat or per-run platform does not, and the maintenance burden of a codeless or scripted suite lands back on whichever engineer inherited it. We modeled that hidden line item directly in our breakdown of the true cost of test maintenance, and the build-versus-buy math around it shows up again in our comparison of hiring a QA engineer against buying a tool.
That hidden line item is really the same shift running underneath this whole budget line. A test strategy allocates a scarce resource, and for thirty years that resource was the time it took to write the test in the first place. What this comparison keeps bumping into, row after row, is that the resource has quietly become something else: the attention it takes to review and maintain what a vendor, or an AI, already generated for you.
The reason this matters specifically for a budget line is that vendor selection and maintenance cost are two different decisions that keep getting collapsed into one number. A cheaper per-seat quote doesn't help if the suite it buys still needs four hours a week of selector repair. A pricier managed retainer can be cheaper in total if it genuinely eliminates that repair work. None of the six rows in the table price that trade explicitly. The maintenance hours are the part of the bill every vendor leaves you to discover on your own.
Put a rough number on it using the reference workload above. A senior engineer spending four hours a week on selector repair and flaky-test triage, at a loaded rate of roughly $75/hour, is about $15,600 a year, more than the published entry price of every vendor in the table above, and it never appears as a line item on any of their invoices. That figure is our own illustrative arithmetic, not a vendor number, and it moves with your suite's actual flakiness rather than with which vendor you pick, which is exactly why it survives every choice in this comparison. Any honest test automation cost benefit analysis has to carry that number next to the vendor quote, because it is the half of the total that no invoice itemizes.
Building this budget line starts with picking a pricing model that matches how your team actually works, not the one with the smallest number on the homepage. A per-run vendor is a poor fit for a team that wants full regression on every merge; a per-seat vendor is a poor fit for a team that wants every engineer touching the suite. Bring the reference workload above to the sales call: a vendor that won't quote 150 tests, 100 runs a week, 8 developers, and 5 parallel workers in public will usually still answer it on the phone, and asking for those four numbers by name gets you an answer built for your suite instead of a range built for someone else's. Write the maintenance hours into that same spreadsheet, right next to the quote, because the pricing model and the upkeep it doesn't cover are two separate numbers in the same annual total.
If part of what's driving the search for a cheaper vendor is the maintenance hours nobody quotes, Autonoma is worth adding to the same spreadsheet, because the Diffs Agent's per-PR maintenance is aimed directly at the line item this article just spent a whole section pointing out that vendors don't sell separately.
Frequently Asked Questions
It depends entirely on the billing model, which is what makes a single number impossible to quote. The lowest published entry prices we verified in August 2026 are $20/month (Functionize's Pro tier) and $59/month (BrowserStack's Automate Chrome plan at one parallel session). Beyond those entry points, mabl, Testim, and testRigor publish no dollar figure at all, and BrowserStack's own pricing page does not list rates past a single parallel worker. A real QA automation budget for a team with a 150-test suite, 100 runs a week, 8 developers, and 5 parallel workers typically requires at least one sales conversation, because most vendors in this category do not price that workload in public.
It moves the cost rather than removing it, which is the more useful way to think about it. Published entry prices, $20/month for Functionize's Pro tier and $99/month for BrowserStack's Automate Desktop plan at one parallel, look cheap next to hiring, but neither figure includes the maintenance hours a suite still needs once it's running. A senior engineer spending four hours a week on selector repair and flaky-test triage, at a loaded rate of roughly $75/hour, is about $15,600 a year, and that cost exists regardless of which vendor's invoice you're paying. Real savings come from cutting those maintenance hours, not from picking the vendor with the smaller sticker price.
Among vendors we checked in August 2026, Functionize's Pro tier at $20/month and BrowserStack's Automate Chrome plan at $59/month (1 parallel, Chrome only) have the lowest published entry prices; BrowserStack's Automate Desktop plan, covering 3,000+ desktop browser combinations, is $99/month at 1 parallel. That said, an entry price is not the same as the cost at a real workload: neither vendor publishes how many credits or parallel sessions a 150-test suite running 100 times a week would actually consume, so the cheapest sticker price is not necessarily the cheapest real bill.
Quote-gated pricing in this category is a normal enterprise sales practice, not deception. Seat counts, suite size, run frequency, and required certifications like SOC 2 or PCI vary enormously between buyers, and a published list price can't account for that variance the way a sales-qualified quote can. Several of the vendors in this comparison, mabl and Testim among them, gate all pricing behind a demo or sales call; that reflects how enterprise software in this category is typically sold, not an attempt to obscure a bad price.
Neither is better in the abstract; each fits a different growth pattern. Per-test pricing, the model behind QA Wolf's managed coverage tier, is a better fit when your developer headcount is stable but your suite keeps growing, since the bill scales with test count rather than logins. Per-seat pricing, the model behind Testim's licensing, is a better fit when your suite is stable but you want every engineer to have access, since adding tests doesn't raise the bill the way adding developers does. Pick based on which of the two, suite size or developer count, is actually growing on your team.
The most common hidden costs are add-on modules priced separately from the base plan (visual testing, accessibility checks, and observability add-ons across the cross-browser grid vendors), the engineering hours spent repairing broken selectors or re-recording flows when a UI changes, pipeline queue time when a team under-provisions parallel capacity, and a shift from published to quote-only pricing once compliance requirements like SOC 2 or PCI enter the conversation. None of these show up on the sticker price; all of them show up on the actual annual spend.
It moves the repair work off the engineer and into the pipeline, which is the line item every quote in this comparison leaves out. The four hours a week someone spends on selector repair and flaky-test triage exists no matter whose invoice you are paying, because a suite written once starts drifting the moment the interface does. Autonoma generates its test cases by reading the codebase rather than from a recorded session, so there is no recorded selector layer to hand-repair, and the Diffs Agent reads each pull request's diff to add, update, and deprecate cases as the code changes. A Reviewer agent then classifies each failure as a real bug, an agent error, or a stale test, so what reaches a person is a verdict to act on rather than a red build to sort out from scratch. Whichever meter you end up billed on, that maintenance line is the one you can actually shrink, and it is worth asking every vendor on this list how much of it they absorb versus hand back to you.




