Blog

A landing page should pass this test before it gets live traffic

Andrei Dan

Founder, Aetherya

August 23, 2026 — 10 min read


Use cognitive pre-testing to find unclear, untrusted, or weak landing-page variants before spending traffic on an A/B test.

Landing pages often fail in a boring way. Someone reads the hero, scans the proof, and still cannot say what the company sells or why it matters. Do that work first. The experiment can wait. Recruit one prospective buyer and let them read without coaching. Once they finish, ask what the company seems to sell, whether the offer belongs in their work, what makes the claim believable, how they read the price, and what they expect after the click. Begin with an expert pass. Follow it with a synthetic audience test, a few prototype sessions, or several human interviews, depending on the risk. Remove the weak versions and rewrite the hypothesis. Then use an A/B test to measure causal impact on real behavior. Pre-testing does not replace the live experiment. It improves what enters it, which matters when traffic is limited, implementation is costly, or the team has more ideas than it can test well.

An A/B test is an expensive place to discover that nobody understood variant B.

The decision before the experiment

Teams often ask an A/B test to solve two problems at once:

  1. Is the idea coherent?
  2. Does the coherent idea change behavior?

The first is a design and interpretation question. The second is a causal question. Randomized online experiments are well suited to the second because random assignment helps isolate the effect of a treatment. Microsoft researchers describe controlled experiments as a scientific method for estimating how product changes affect customer behavior. See Online Experimentation at Microsoft.

Running incoherent concepts through that machinery wastes traffic. A null result may mean the idea had no effect. It may also mean the page buried the claim, confused the audience, or changed five things at once.

Pre-testing narrows that uncertainty before the experiment begins.

What pre-testing can answer

A useful landing-page pre-test examines questions such as:

  • What does the visitor think the product does?
  • Who do they think it is for?
  • Which claim appears unsupported?
  • What information do they seek before acting?
  • Where does the page create effort or doubt?
  • Does the call to action match the commitment the page has earned?
  • Which difference between variants should become the A/B-test hypothesis?

It cannot establish a conversion lift. It observes or models interpretation under test conditions, not behavior in the production market.

The cost of skipping this step

Poor variants consume more than traffic.

Design and engineering teams spend time implementing them. Analysts wait for a sample. Paid-media teams buy visits. Stakeholders then interpret noisy or null findings. On a low-traffic page, the wait may last weeks and still fail to detect a useful effect.

Experiment quality can also fail for technical reasons. Sample ratio mismatch, for example, occurs when observed assignment differs from the planned split and can indicate serious data-quality defects. Microsoft treats this as a critical diagnostic before interpreting results. See Fabijan and colleagues' KDD paper.

Pre-testing cannot repair an invalid experiment. It can make the experimental question worth the operational effort.

A six-step landing-page pre-test

1. Write one behavioral hypothesis

Avoid "variant B is better." State the mechanism.

Weak hypothesis:

A shorter hero will improve conversion.

Better hypothesis:

Naming the audience and implementation time in the hero will reduce uncertainty for agency leaders, increasing qualified clicks to the sample-report page.

The second version identifies the audience, change, mechanism, and outcome.

2. Hold irrelevant variables constant

If the test concerns the value proposition, keep layout, imagery, navigation, price, and call-to-action treatment constant. Otherwise, a response cannot be tied to the intended change.

Whole-page redesigns can still be tested, but they answer a broad package question. They teach less about why one version wins.

3. Define the audience by the decision

"SaaS buyer" is too broad. A useful audience definition includes role, situation, prior knowledge, constraints, and purchase responsibility.

For an agency service page, one population might be owners of CRO agencies that serve DTC brands and need client-ready evidence under short timelines. Another might be UX research leads who care more about method validity and participant coverage. The same page can create different doubts for each group.

Use only traits relevant to the choice. Decorative persona details add confidence without adding validity.

4. Run an interpretation task

Do not begin by asking whether people like the page. Preference is easy to state and often hard to act on.

Ask the participant or simulation to:

  • Explain the offer in their own words.
  • Identify the intended customer.
  • State what evidence supports the main claim.
  • Name the largest unanswered question.
  • Describe the next action and its perceived commitment.
  • Compare the variants on one defined dimension.

For human sessions, add a brief recall task after removing the page. Recall exposes whether the central proposition survived attention.

5. Record friction by type

Group findings by mechanism, not by positive or negative sentiment.

Friction typeWhat it sounds likeLikely response
Comprehension"I still do not know what this is."Rewrite proposition or add context
Relevance"This seems made for a larger team."Clarify audience and use case
Credibility"I do not believe that claim yet."Add method, proof, or boundary
Risk"What happens to our data?"Surface governance and security
Effort"This looks difficult to set up."Explain workflow and time to first result
Commitment"I am not ready to book a call."Offer a lower-commitment next step

Count recurrence, but do not turn a small or synthetic sample into a population estimate. The purpose is to find mechanisms worth fixing or testing.

6. Set the gate for live traffic

A variant should not enter the A/B test merely because the team prefers it.

Use a gate such as:

  • The target audience can state the offer and intended user.
  • The primary claim has visible support or an explicit boundary.
  • The call to action matches the page's stage of persuasion.
  • No severe accessibility or usability defect blocks the task.
  • The difference between variants maps to the written hypothesis.
  • Analytics and assignment checks are ready before launch.

Once both variants pass, the live experiment can answer the causal question.

Where synthetic simulation helps

Synthetic audience simulation is most useful between expert review and human or live testing.

It can run the same interpretation task across several defined audience models, preserve the stimulus, and expose repeated objections quickly. That is valuable when the team has many variants or little traffic. It is also useful for agencies preparing a client workshop because the output gives stakeholders something concrete to inspect.

Its limits are equally important. Model-generated responses are not visits, clicks, purchases, or recruited-user testimony. Research on synthetic survey samples has found that plausible averages can hide distorted variation and subgroup relationships. See Bisbee et al., 2024. Treat segment contrasts as hypotheses until they are compared with observed evidence.

A worked Aetherya protocol

Imagine a CRO agency testing three hero sections for a subscription skincare brand.

The decision brief states that one concept will advance to production. The audience consists of repeat online skincare buyers who know their skin concern but distrust exaggerated claims. Every variant uses the same layout, image, price, and call to action. Only the proposition and supporting proof change.

The agency uploads all three versions to Thesia and runs a fixed interpretation task. It reviews how each audience model explains the offer, which claim triggers doubt, what proof is missing, and whether the next action feels proportionate. The team removes any version that consistently fails basic comprehension. It then takes the strongest two into five human prototype sessions and a live A/B test.

The evidence passport preserves the population definition, stimulus, procedure, system version, output status, and limitations. No simulated click-through or conversion rate appears in the client report unless a validated forecasting method supports it.

This is a protocol example, not a reported outcome. A public case study should add an actual run and later compare its direction with observed human or market data.

How many variants should reach the A/B test?

Usually two. More arms divide traffic and complicate interpretation. High-traffic experimentation programs may support several treatments, but low-traffic teams should be ruthless before launch.

A practical sequence is:

  1. Begin with five to ten rough concepts.
  2. Use internal review to remove off-strategy options.
  3. Use simulation to identify comprehension and trust failures.
  4. Use a few human sessions to find unmodeled context.
  5. Build the best two variants.
  6. Run the controlled experiment with a predefined metric and stopping rule.

The purpose of the early stages is not to crown a winner. It is to protect the experiment from avoidable losers.

Low traffic changes the method, not the standard

A site with low traffic may be unable to detect a modest conversion change in a reasonable period. That does not justify calling directional evidence statistically significant.

Instead, combine methods:

  • Use task-based usability sessions for severe friction.
  • Use synthetic simulation for rapid hypothesis generation and variant screening.
  • Use interviews for meaning and purchase context.
  • Use analytics for current paths and drop-off.
  • Test larger, decision-relevant changes when traffic is scarce.
  • Accumulate evidence across repeated decisions and compare predictions with outcomes.

This approach produces a reasoned choice without pretending small samples prove more than they do.

When to skip simulation

Go directly to human or live testing when:

  • The central risk concerns accessibility technology or physical interaction.
  • The audience has rare experience the model cannot represent credibly.
  • The change affects legal consent, health, credit, employment, or safety.
  • The current page already has sufficient traffic and the variants are mature.
  • The main uncertainty is technical instrumentation rather than audience interpretation.

Simulation should remove uncertainty it can address. It should not become a ritual added to every test.

Final answer

A/B testing tells you what a controlled change did to real behavior. Landing-page pre-testing helps ensure the change is coherent enough to deserve that test.

Define the audience. Freeze the stimulus. Test comprehension, relevance, credibility, risk, effort, and commitment. Remove weak concepts. Then spend traffic on the causal question that remains.

Next step: Run a cognitive website audit in Thesia, or explore Thesia to frame a landing-page study.

Sources

Aetherya

Cognitive Simulation Research & Technology

Related