How We Test
The testing methodology behind every StackPipeline score: rubric, weighting, test environment, and what would make us change a rating.
- Effective:
- Last updated:
A score is only useful if you know how it was produced. This page documents the exact process behind every rating on StackPipeline, so you can decide whether our priorities match yours.
The three rules
- We pay for the product. Every tool is tested on a paid plan or a full-featured trial we signed up for ourselves. We do not accept vendor-provided accounts with elevated limits, because those do not reflect what you will experience.
- We score before we monetise. Testing completes and scores are locked before anyone checks whether an affiliate program exists. The tester does not know the commission rate.
- We test the same thing on every product. Comparisons only mean something if the workload is identical, so we rebuild the same workflows on each platform rather than evaluating each against its own marketing.
The scoring rubric
Every product in a category is scored against the same five criteria, weighted as below. The final score is the weighted average, rounded to one decimal place.
| Criterion | Weight | What we measure |
|---|---|---|
| Core capability | 30% | Does the product do its primary job well? For enrichment tools this is measured match rate against a controlled list; for automation platforms it is workflow success rate under induced API failures. |
| Ease of use | 20% | Time from empty account to first working workflow, measured with a stopwatch by someone who has not used the product before. |
| Integrations | 20% | Connector count matters less than connector quality. We test the five most common connectors per category and note any that are stale or broken. |
| Value for money | 20% | Real cost at a realistic volume, not list price. We model 10,000 monthly operations and compare total cost including required add-ons. |
| Support | 10% | We open at least two real support tickets per product from a paid account and record first-response and resolution times. |
The test environment
Automation platforms are tested against a fixed set of five workflows that mirror common production patterns:
- A lead router with conditional assignment across three branches
- A CRM enrichment pipeline with a deduplication lookup
- A Slack alerting flow with formatted message payloads
- A scheduled two-way data sync between two systems
- A webhook receiver that must survive a deliberately failing downstream API
That last one matters most and is the one vendors never demo. We induce 500 errors and timeouts on the receiving endpoint and record whether the platform retries, how it surfaces the failure, and whether any data is silently lost.
Data enrichment testing
Enrichment tools are measured against a controlled 5,000-contact list with a known composition (60% US, 40% EU, mixed company sizes). We measure reported match rate, then manually verify a random 200-record sample to establish true accuracy — reported and actual match rates routinely differ by 5 to 15 percentage points.
Minimum testing time
- Full review — minimum 20 hours of hands-on use over at least 30 days
- Comparison entry — minimum 8 hours per product, same workflows on each
- Integration tutorial — the integration is built end to end and verified working before publication
What would change a score
Scores are revisited when a vendor ships a material feature, changes pricing, or when our 90-day re-verification finds a discrepancy. A score can go down. When it changes, the article notes what changed and why.
What does not change a score: vendor complaints, commission rate changes, or pressure from a partner. We have removed products from affiliate programs rather than adjust a rating.
Limitations we will admit to
We are a small team, and there are real constraints on this process:
- We cannot test every tier of every product. Where we test a specific plan, it is named on the page.
- Our enrichment test list skews toward US and Western European contacts, so match rates for other regions may not generalise.
- Enterprise products with sales-gated pricing are harder to evaluate on value, and we say so rather than guessing.
- A 30-day test window will not surface every long-term reliability issue.
Questions about a score
If a rating looks wrong, tell us what we missed: editorial@stackpipeline.com. We re-test when a specific, checkable objection is raised — including from vendors.