Blog / Research
Message Testing: Validate Value Props With Research
Learn how to test value propositions and messaging with qual and quant research—plus when to use message testing vs concept testing to pick winning copy.
What message testing actually is (and what it isn’t)
Message testing is research that identifies which words, claims, and value propositions resonate with a specific audience. It won’t tell you whether people want the product — that is concept testing’s job. It starts once the offer is settled and the only open question is how to express it. Value proposition testing is the same discipline aimed at your single most important claim.
Give reviewers four dimensions to judge each message against:
- Clarity — can people restate what it means without help?
- Relevance — does it speak to a problem they care about?
- Differentiation — does it sound distinct from competitors, or interchangeable?
- Believability — do they accept the claim, or discount it?
Keep the scope tight. Message testing evaluates language and framing: headlines, taglines, benefit statements, proof points, and the order you make claims in. It is not a way to test features or price, though what you learn feeds those decisions and your product discovery workflow.
The payoff comes early. A good test kills weak claims before you spend on media, lowers launch risk, and leaves sales and marketing with shared, evidence-backed language. It works best as a mix of qualitative and quantitative work, not one clever survey question.
Message testing vs concept testing: when to reach for which
Concept testing answers “do people want this?” Message testing answers “which language makes them understand it and act?” Easy to state, easy to blur.
The two sit at different points in the lifecycle. Run concept testing before and during development, while you are still validating the offer. Message testing comes later: before a campaign, ahead of a launch, or when you are rewriting a homepage and the product is not changing.
The common mistake is testing copy while the concept is unproven. You will optimise a headline for an offer nobody wants, then wonder why the polished message underperforms in market. If you cannot yet answer “is this valuable, and to whom?”, run a concept test first.
A quick decision rule:
- Unresolved “is this valuable?” → concept test first.
- Validated offer, unresolved “how do we phrase the value?” → message test.
The methods overlap because they often share a stimulus: a written value proposition statement. Put that in front of someone and you learn two things at once — whether the idea appeals, and whether the wording works. Keep those questions separate in your discussion guide and your analysis, or the results turn to mush.
Start with the right inputs: audience and source material
Test messages against the wrong audience and the result is worthless. Define the segment before you write anything. In B2B, name the buying role each message is for: the champion who builds the internal case, the economic buyer who signs the contract, the end user who lives with the tool. Each weighs clarity, relevance, and proof differently, and a message that thrills a user can leave a finance director cold.
Source candidate messages from real customer language, not the boardroom. Mine recent interviews, recorded sales calls, support tickets, and public reviews for the exact phrases people use to describe the problem and the result they want.
Research into why customers leave and win-loss signals surfaces the pains and gains your messaging has to address, usually more honestly than your feature page does. Pair that with work on who actually buys in B2B so each message targets a real decision-maker.
Build a small message inventory: 6–12 distinct value-proposition variants and proof points, each mapped to a specific job, pain, or gain. Write variants that differ in angle — time saved versus risk removed versus status gained — not cosmetic rewrites. If two messages say the same thing in different words, the test cannot tell them apart. The messaging research methods that follow split into two camps: qualitative for diagnosis, quantitative for prioritisation. A serious test uses both.
Qualitative methods: understand why a message works
Qualitative work tells you why a message succeeds or fails, and it is where you validate messaging with customers in their own words. Use 1:1 interviews and small focus groups to capture first impressions, comprehension, and the mix of rational and emotional reaction that a rating scale flattens.
Run comprehension probes. Show a message, take it away, and ask the participant to say back what it promises and who it is for. Gaps between what you meant and what they repeat expose jargon, vague benefits, and claims that read as noise.
Use highlight-and-react techniques. Give people the copy and ask them to mark what is confusing, what is exciting, and what they do not believe, then ask why in each case. A cloze-style exercise, where key words are blanked out for the participant to fill in, shows which concepts are doing the work.
Qual is for depth and diagnosis: it explains the “why” behind later quantitative scores and generates sharper variants for the next round. In our value-proposition studies, the interview stage almost always changes the wording we carry into the survey. Our note on usability testing vs user interviews covers when a structured task tells you more than an open conversation.
One guardrail: small samples mean no ranking and no significance claims. 5–8 interviews per segment will show you patterns and problems. Treat what you hear as directional, not decisive, and resist the urge to count.
Quantitative methods: prioritise and prove at scale
Quantitative methods turn a shortlist into a ranked decision.
MaxDiff. Show respondents small subsets of messages and force a best and a worst choice in each. Aggregated, those choices produce utility scores and a clean ranking across the full set. Because people choose rather than rate, MaxDiff message testing avoids the “everything scores four out of five” problem. Reach for it when you have many claims and need to know which few carry weight.
Monadic and sequential monadic testing. Each respondent sees one message, or a few in turn, and rates it on clarity, relevance, believability, uniqueness, and intent. This gives an absolute reading of each message and lets you compare scores by segment.
Rating-scale surveys. Fast and cheap, but prone to score inflation — respondents are agreeable and most things land near the top. Counter it with ranking, MaxDiff, or a constant-sum task where people divide 100 points across the options.
In-market A/B testing. The highest realism, because you measure real behaviour, but it needs traffic. Put your baseline conversion rate and target lift into a sample-size calculator before you promise a date: a page with roughly 600 visitors a week can run for many months before it reaches significance, and it compares only two options cleanly at a time.
Match method to question: MaxDiff to prioritise many claims, monadic for an absolute read on a few, A/B to validate the winner against live conversion. Our UX research for B2B SaaS engagements usually chain at least two.
Designing the study: variants, metrics, and a plan that yields a decision
Turn assumptions into message hypotheses you can test. “Buyers respond more to time-saved framing than cost-saved framing” is testable. “We need better messaging” is not.
Choose outcome metrics before you field, and write down what winning looks like. A workable bar: top-two-box relevance (the share of respondents choosing the top two scale points) above a set threshold, a leading MaxDiff utility score, and a measurable lift in purchase intent against the control. Pre-committing stops post-hoc rationalisation.
Control the stimulus. Present every variant in the same format, at the same length, with the same visual treatment. If one message gets a polished mockup and the others get plain text, you are testing design, not language. Keep exposure realistic — people skim a homepage, they do not study it.
Decide your sample size and segment cuts in advance. If you plan to compare champions and economic buyers, size each group to support that. Deciding cuts afterwards invites you to hunt through subgroups until something looks interesting.
Sequence the work: qualitative first to refine variants, quantitative to prioritise and measure, optional in-market A/B to confirm with real behaviour. For every plausible result, agree now what decision it triggers, so the study ends in a choice rather than a debate.
A message test in practice: an anonymised example
We ran a value-proposition study for a company selling an outsourced bookkeeping service to small and mid-sized businesses. The team held several competing value propositions and had to choose one to anchor a positioning refresh and launch.
We started with depth interviews among target buyers who were not yet customers — owners and managers of small and mid-sized firms using rival providers. Each interview checked how the value propositions were understood and whether they were believed. A pattern emerged quickly: prospects who had already tried outsourcing judged every claim against a real memory of problems, while those who had never outsourced were wary of the category itself.
That split exposed the weakness in the team’s preferred message. It was confident and ambitious, and it assumed the reader already treated outsourcing as normal. The wary segment read that confidence as a reason for suspicion. A plainer variant, built from the words customers used for their own problem, sat better with them.
The interviews reshaped the wording, and a quantitative round followed to rank the shortlist across segments. The plainer, customer-language value proposition became the basis for the rewrite. The clearest gain was internal: the founders and their marketing team walked away with one evidence-backed story rather than several competing drafts.
From results to a messaging decision (and pitfalls to avoid)
Read the metrics together. A message with high MaxDiff utility but weak monadic believability is a claim people like in theory and doubt in practice. Triangulate across preference, absolute scores, and qualitative reaction rather than trusting one number.
Watch for the recurring pitfalls:
- Scale inflation — everything scores highly on agreement scales; lean on choice-based data.
- Over-reading subgroups — a six-point gap on 40 respondents is noise. Check the base size before you brief it.
- Near-identical variants — if the test cannot discriminate, the variants were too similar to begin with.
- Confirmation bias — the internal favourite gets the benefit of the doubt. Name that risk before you analyse.
Check statistical significance and effect size before declaring a winner, and report your confidence honestly. “Message B leads, but within the margin of error” is a legitimate result and a reason to run the in-market test.
Document the decision: the message you chose, the ones you rejected, and why. That record makes the choice defensible to a new CMO and easy to revisit.
Set a baseline now and re-test as positioning, competitors, and audience shift. The approach in our guide to set metrics and baselines applies here: message testing is a cadence, not a one-off. Pick the one value proposition that matters most, run it through a qualitative round this quarter, and you will have an evidence-backed start rather than a boardroom guess.
Frequently asked questions
What is the difference between message testing and concept testing?
Concept testing checks whether people want the offer. Message testing checks which words make them understand and act on an offer you have already validated. The two often share a stimulus — a written value proposition — which is why teams conflate them. If “is this valuable?” is still open, run a concept test first.
How do you test a value proposition?
Start from customer language, not internal drafts. Write 6–12 genuinely different variants, each tied to a job, pain, or gain. Diagnose comprehension and believability in interviews, then rank the claims with a choice-based survey such as MaxDiff, and confirm relevance by segment with a monadic test. Validate the winner in market if you have the traffic.
How many messages should I test at once?
For a qualitative round, 6–12 distinct variants is workable. MaxDiff can handle 15–30 claims because each respondent sees only small subsets. Beyond that, fatigue degrades the data. Make sure variants differ in angle, not just wording.
What sample size do I need for message testing?
For qualitative diagnosis, 5–8 participants per segment surfaces the main patterns. For a monadic or MaxDiff survey, aim for at least 100 respondents per segment you want to analyse separately. In-market A/B testing depends on your traffic and baseline conversion — run the numbers through a sample-size calculator before you commit to a timeline.
How often should we re-test our messaging?
Treat it as a cadence. Re-test when positioning changes, when a competitor shifts their story, when you enter a new segment, or roughly once a year for a core value proposition. Keep the baseline metrics consistent so you can see movement.
About Glasgow Research — Glasgow Research helps B2B SaaS teams turn customer and market research into product decisions. Work with us.
Author
About Vadim Glazkov
Vadim Glazkov is the founder of Glasgow Research and a product research expert working with founders and B2B SaaS teams on customer interviews, JTBD, market validation, and decision-ready research.