Testability as a first-class verdict
Every other grader hands you a score and assumes you can act on it. We put the answer to "can you even test this on your traffic?" next to the score, and we are willing to tell you that the answer is no.
Channel-aware CRO · Free scorecard
Send us the URL and we will score the page across eight conversion factors, give you one fix per weak factor, and tell you straight whether your traffic can support a reliable A/B test. No cost, no card.
The whole testing industry is built for sites with tens of thousands of sessions a month. Published guidance puts a conclusive test at roughly 10,000 visitors and several hundred conversions per variation, which is well beyond most lean B2B sites and local lead-gen businesses. Running tests anyway produces false wins that fall apart in the real world.
So we lead with the verdict, not just a grade. If you have the volume, you get a prioritised test plan and what to run first. If you are borderline, we tell you which proxy signals to test on instead of waiting months for significance. And if you do not have the volume, we say so, and point you at what to fix by evidence: form-field abandonment, funnel drop-off, page load, and the heuristic audit itself.
Almost nobody in this category answers that question honestly, because the honest answer sends a lot of buyers away from testing. That is exactly why we lead with it.
Room to grow
Eight conversion factors
An illustrative example, not your data. Your scorecard is run on your own page and sent back to you.
One landing page URL is all we need. It has to be a public page, because that is the only thing we read: no login, no analytics access, no private data. Tell us your rough monthly traffic and conversions too if you have them, since that is what the testability read is built on.
Value, relevance, clarity, friction, anxiety, distraction, urgency and CTA, drawn from the LIFT, MECLABS and Trinity models. Each factor gets a pass, warn or fail with a one-line reason, and each weak factor gets one concrete fix. It is a directional heuristic read from your public page, not a statistical test, and we label it that way.
Three tiers. Full A/B is realistic at roughly 175 conversions a week and above. Below that we test on proxy signals such as form starts and click-to-call. Below that again, a controlled test cannot reach a reliable result, so we measure against a baseline and fix by evidence instead.
Every factor prioritised with the reasoning and the expected impact order, formatted to hand to a designer or developer, plus the full testability write-up: your tier, the method we would use at your traffic level, and a realistic time to significance if testing is viable.
Two answers of equal weight: how good the page is, and whether you can trust a test on it.
You get the headline score with the eight-factor breakdown sorted fails first, one specific fix per factor that warned or failed, the testability verdict in plain English, and the ranked plan behind it. Nothing is held back behind a second form.
Every other grader hands you a score and assumes you can act on it. We put the answer to "can you even test this on your traffic?" next to the score, and we are willing to tell you that the answer is no.
Not a generic checklist. Each factor that warns or fails gets a concrete action tied to what we saw on your page, and the plan is ordered by expected impact so you know what to do on Monday.
The scorecard is a directional read of a public page, not a statistical test, and we say so on the result. The Bayesian experimentation that produces ship or kill calls is the managed engine, never the free scorecard.
The scorecard reads one public page and tells you both what to fix and what kind of measurement your traffic can actually support. This is the complete list of what it does today, with nothing on it that is roadmap or managed-service only.
Classic frequentist testing punishes low-traffic sites twice. Peeking at results early, which everyone does, inflates the false-positive rate by roughly five times. And the sample sizes are brutal: a 3% baseline conversion rate at a 10% minimum detectable effect needs about 53,000 visitors per variant under a frequentist test, against about 18,700 under the Bayesian approach. Most local lead-gen sites never reach either number, which is why we check before we recommend testing anything.
That method is what sits behind the managed MxD Conversion Engine: Beta-Binomial posteriors with a weakly-informative shrinkage prior, a per-channel posterior as well as an overall one, and a ship, kill or keep-running call rather than a p-value. It is channel-aware on purpose, because a variant can win overall and still lose on Google Ads.
And we say what the scorecard does not do. It does not run an experiment. It does not read your analytics, your CRM or anything behind a login. It does not get smarter the more sites it sees, and we will not claim it does: the method is fixed and transparent. What improves your results is the work, and if your traffic supports it, properly measured experiments.
The free tools everyone reaches for grade your technology. We grade whether the page can convert, and whether you can trust a test on it.
| MxD CRO Scorecard | HubSpot Website Grader | VWO free tools | |
|---|---|---|---|
| What it actually grades | Eight conversion factors: value, relevance, clarity, friction, anxiety, distraction, urgency, CTA | Web performance, mobile usability, SEO and security, scored out of 100 | UX friction, clarity and conversion elements, ranked high to low |
| Built for lead-gen and lean B2B pages | Yes | Any site, technical checks only | Positioned as an ecommerce UX audit |
| “Can you even A/B test on your traffic?” verdict | Yes | Not covered | Not covered |
| Estimates how long a test would take | Yes, with the method we would use at your volume | Not covered | Yes, free duration and sample size calculator |
| Tells you when your volume is too low to test at all | Yes, and we say so plainly | Not covered | No, it returns a duration without a verdict |
| One specific fix per weak factor, ranked | Yes, ordered by expected impact | General recommendations per category | Yes, issues ranked by severity |
| States plainly that it is a directional heuristic | On every result | Not published | Not published |
Swipe the table to see every column.
Competitor details verified from their own product and pricing pages on 22 July 2026: hubspot.com/tests, vwo.com/pricing and vwo.com/tools. VWO does not publish plan prices. Reverify before republishing, as free tools and pricing in this category change often.
The scorecard is free, and there is no card and no trial. The honest catch is that we are an agency: if the plan is useful and you would rather we did the work and measured the result, that managed service is what we sell. The scorecard stands on its own either way.
It is a directional heuristic, and we label it that way. We read your public page and score it against eight conversion factors so you know where you stand and why. The real Bayesian experimentation, the part that runs live tests and returns ship or kill calls, is the managed MxD Conversion Engine.
It is our read on whether an A/B test can ever reach a reliable result on your traffic. Above roughly 175 conversions a week, a full A/B test is realistic. Below that we recommend testing on proxy signals such as form starts or click-to-call. Below that again, we tell you to improve by evidence instead of pretending a test could reach significance.
No. We read only the public landing page you send us, the same thing any visitor can see. No login, no analytics access, no private data. If you later work with us on the managed engine, deeper measurement uses a consent-gated tracking snippet that logs events and never reads form field values.
Because normal A/B tests are unfair to low-traffic sites. Checking results early inflates false positives by around five times, and the required sample sizes are enormous: a 3% baseline at a 10% minimum detectable effect needs roughly 53,000 visitors per variant frequentist, against about 18,700 Bayesian. The Bayesian method gives smaller sites an honest read they can actually reach.
No, and we will not claim it does. The method is fixed and transparent, and the prior behind the managed engine is static rather than learned across clients. What improves your results is the work: fixing what the scorecard flags, and running properly measured experiments when your traffic supports them.
The scorecard tells you what is wrong with the page and whether you can test it. Our managed conversion work fixes it and measures the result: channel-aware Bayesian experiments that return a ship, kill or keep-running call per channel when your volume supports testing, and the evidence-based path of form-field abandonment, funnel drop-off, page load and voice-of-customer themes when it does not.