Key Takeaways
- Message testing before a campaign launch tells you not just which message performs better, but why — and why matters more when budget is real.
- Surveys and preference polls capture what respondents choose in the moment; they cannot capture the reasoning, emotional response, or skepticism underneath that choice.
- Sonderly is an AI concept testing tool that conducts open-ended qualitative interviews with your target audience and returns thematic analysis across all respondents — including themes no one thought to ask about.
- A message test can be live and collecting responses within an hour. The analysis surfaces patterns across respondents, with source quotes accessible for every finding.
Why Message Testing Matters Before You Commit Budget
Sonderly is an AI concept testing tool that conducts qualitative interviews with your target audience and returns structured analysis across all respondents — allowing brand and market research teams to understand not just which message lands, but why. This guide walks through how to run a message or concept test using Sonderly, from designing the interview guide to reading the findings and making decisions with them.
The core argument for testing messaging before a campaign launch is straightforward: the cost of being wrong scales with the spend. A campaign built around a message that resonates weakly, triggers unintended associations, or answers a question your audience isn't asking will underperform regardless of creative quality or media placement. Testing is not conservatism — it is leverage.
The harder question is not whether to test but how. The most common approach — putting two messages in front of respondents and asking which they prefer — generates a winner and almost no usable information about why it won or whether the preference would survive contact with a real purchase decision. What actually makes messaging work is harder to measure: the degree to which a message names a tension the audience already feels, the credibility of the implied promise, the associations it triggers, and the objections it fails to neutralize.
Getting to that level of understanding requires conversation, not selection.
What Surveys Miss in Concept Testing
Quantitative concept testing — presenting stimuli to a large panel and capturing preference ratings, intent-to-purchase scores, or likeability rankings — is a well-established research method with real value in certain contexts. It is efficient at identifying a winner across a large sample and useful for measuring relative performance against a benchmark.
What it cannot do is explain the finding. A score tells you that Message A outperformed Message B by a meaningful margin. It does not tell you whether respondents preferred Message A because it described their situation accurately, because it sounded more credible than Message B, or because Message B triggered an association they didn't like that had nothing to do with your product. Those are three different strategic conclusions, and they lead to three different decisions about what to do next.
Survey methodology researchers have documented the related phenomenon of satisficing — where respondents under the cognitive constraints of a questionnaire give an acceptable answer rather than their fully considered one (Krosnick, 1991, Applied Cognitive Psychology). In concept testing, this manifests as preference data that reflects moment-of-decision impressions rather than the deeper response a respondent might articulate if asked to describe their reaction in their own words.
The distinction between stated preference and the reasoning behind it is not an abstract methodological concern. It has direct consequences for how a team uses the findings. A market researcher who can tell the CMO "Message A won" is in a different position than one who can say "Message A resonated because it named a specific pain point — but it also surfaced a credibility concern in respondents who had tried similar products before. Message B was dismissed quickly, but for a reason that is fixable." The second version is a brief. The first is a data point.
What AI-Powered Concept Testing Actually Looks Like
An AI-powered concept test works by presenting stimulus material to respondents and then conducting an open-ended qualitative interview about their reaction. The stimulus can be a positioning statement, a tagline, a value proposition, a product description, or early creative — whatever the team needs to test before committing resources.
From the respondent's perspective, the experience is a conversational exchange. After seeing the material, they are asked open-ended questions: what they noticed first, what the message communicates to them, whether it describes a situation they recognize, what they believe and what they are skeptical about. The AI interviewer listens to their answers and probes based on what they say — following threads that emerge rather than cycling through a fixed list. The conversation is recorded and transcribed.
From the researcher's perspective, the interface is the interview guide and the analysis output. You design the protocol inside Sonderly — the stimulus material to present, the opening questions, the topics to probe, and any areas to handle carefully (for example, avoiding questions that might inadvertently reveal the purpose of the study and bias responses). Sonderly generates a shareable interview link. You distribute that link to the audience you want to hear from — your own panel, a recruited sample, or a warm audience from your CRM. Each respondent completes an independent, asynchronous session.
As responses accumulate, Sonderly's analysis builds across the full transcript set, organizing findings into themes with prevalence data, representative quotes, and links to source sessions. The researcher retains full access to every transcript.
How to Set Up a Message Test in Sonderly
A message test in Sonderly has four components: the stimulus material, the respondent profile, the interview guide, and distribution.
Step 1: Prepare your stimulus material.
The stimulus is whatever you are testing — most commonly one to three positioning statements, taglines, or value propositions presented in isolation or in comparison. Monadic testing (one message per respondent) tends to produce cleaner findings about each message's independent reception, because respondents are not influenced by the contrast with alternatives. Paired testing (showing two messages and asking about both) is faster but introduces comparison effects that can distort individual reactions. Both approaches are viable in Sonderly depending on what the team needs to learn.
Keep stimulus material brief and presentation-clean. In a qualitative concept test, you want to surface the respondent's unmediated reaction — not their response to formatting, length, or visual design unless those elements are what you are testing.
Step 2: Define your respondent profile.
The value of concept testing is proportional to the accuracy of the audience. Testing positioning on respondents who are not actually in your target segment produces findings that may be interesting but are not reliable guides to campaign decisions. Be specific: define the characteristics that matter for this test — job function, company size, prior category experience, geography, or whatever dimensions are relevant to the message you are testing.
Sonderly operates on a bring-your-own-audience model. You recruit and contact respondents through your own mechanisms — a CRM segment, a panel partner, a customer list, or a screened invite — and Sonderly provides the interview link. This gives the research team full control over who participates.
Step 3: Design the interview guide.
The interview guide is where the quality of the study is determined. A message test guide should be structured in three phases.
The first phase is unguided reaction. Before asking anything interpretive, give respondents room to describe their first impression in their own terms: "What's your first reaction to this?" or "What stands out to you?" The goal is to capture what is salient before you shape the frame with follow-up questions. Do not lead with comprehension checks or preference questions at this stage.
The second phase is exploration. After capturing the unguided reaction, probe the dimensions that matter for your decision: what the message communicates in the respondent's own words (comprehension), whether it describes a situation they have experienced (relevance), what they find credible and what they find skeptical (believability), and whether it reminds them of anything else they have seen (distinctiveness). Each of these probes surfaces a different dimension of how the message functions.
The third phase is implication. Close with questions that connect the message to behavior: "If a product said this to you, what would you expect it to actually do?" or "Would this make you more or less likely to look into a product like this? Tell me why." This phase surfaces the gap between message resonance and purchase-relevant motivation — which is often where the most important finding lives.
Sonderly allows you to configure the probing logic for each phase — specifying which topics the AI should explore when they surface and how deeply to follow any given thread. This gives the researcher methodology control across all sessions without a live moderator in each conversation.
Step 4: Distribute and collect.
Once the guide is configured and the study is live, Sonderly generates a shareable interview link. Send it to your recruited respondents via email or your outreach platform of choice. Each respondent clicks the link, completes an asynchronous session at their convenience, and their transcript flows into the analysis. A study can be live and collecting responses within an hour of setup.
What the Analysis Reveals
As responses accumulate, Sonderly's analysis surfaces themes across the transcript set: recurring patterns in how respondents reacted, what they understood, what resonated, and what did not. Each theme shows its prevalence, includes representative quotes from transcripts, and links to the underlying sessions for verification.
To illustrate what this type of output looks like in practice, consider a hypothetical test conducted by a B2B software company evaluating two positioning statements before a campaign launch:
- Message A: "The project management platform built for distributed teams"
- Message B: "Fewer meetings. More done."
The analysis across 35 respondents might return something like the following:
Theme: Message B won first impressions but lost credibility — The majority of respondents reacted positively on first read to Message B — it was direct, they understood it immediately, and it named something they wanted. But when probed further, many described skepticism: the message sounded like a claim any tool would make, and without specificity it felt like a promise they had heard before. Representative quote: "Every project tool says something like this. I'd want to know what's actually different."
Theme: Message A was slower to resonate but named a recognized problem — Respondents with direct experience managing distributed teams reacted more favorably to Message A as the conversation deepened. The phrase "distributed teams" triggered recognition — respondents described the specific coordination problems they associated with it. Those without distributed team experience found the message less relevant to their situation.
Theme: Integration was the unstated prerequisite — Across both message conditions, respondents who described being in the market for project management tools raised the question of integrations unprompted — specifically, whether the product would connect to the tools their team already used. This theme appeared in neither message and was not a probe in the interview guide. It surfaced because respondents described what would actually make them consider switching, not just which message they preferred.
That third theme is the finding that changes the campaign brief. It does not tell the team which message won — it tells them what precondition both messages are failing to address. A campaign that leads with either message without acknowledging the integration question may win impressions without winning trials.
(Note: The example above is illustrative of the type of output Sonderly's analysis produces. It is not drawn from a specific published study.)
How to Act on What You Find
The output of a qualitative concept test is most useful when it is treated as a brief, not a verdict. A quantitative preference score produces a winner; qualitative interview analysis produces a direction and a set of constraints.
In the example above, the actionable conclusions are layered. Message B has stronger immediate cut-through but needs specificity to be credible — a team might use its emotional directness while grounding it in a specific, verifiable claim. Message A performs better with respondents who recognize the distributed team problem firsthand — a team might use it for targeted placement where that experience is more common. And the integration theme is a brief in its own right: it identifies an objection that neither message neutralizes, which either belongs in the creative or in the adjacent content that surrounds the campaign.
This is the difference between knowing which message won and knowing what to do with the finding. Both are useful; only one of them tells you what to write next.
For research leads presenting findings to a CMO or brand director, the quote-level evidence available in a Sonderly analysis makes that conversation easier to navigate. "Respondents found this message too generic" is an interpretation. "I'd want to know what's actually different" is a direct quote from a member of your target audience. The quote anchors the finding to something real and specific, which makes it harder to dismiss and easier to act on.
Frequently Asked Questions
What is an AI concept testing tool?
An AI concept testing tool is a platform that presents stimulus material — a message, tagline, value proposition, or concept — to respondents and then conducts an open-ended qualitative interview about their reaction. The AI interviewer follows a research guide the team has designed, probes based on what respondents say, and returns structured analysis across all sessions. Sonderly is an AI concept testing platform built for brand, market research, and insights teams who need to understand audience reaction before committing campaign or product investment.
How is AI-powered concept testing different from a preference survey?
A preference survey captures which option respondents choose in the moment. Qualitative concept testing captures why they reacted the way they did — the reasoning, the emotional response, the credibility concern, the unexpected association. Surveys are efficient at measuring relative performance across large samples; qualitative interviews are better at explaining the result and surfacing findings the research team did not anticipate asking about. The two methods complement rather than replace each other, but for teams making messaging decisions with limited budget, the qualitative explanation is often the more actionable output.
What types of stimulus material can I test in Sonderly?
Sonderly supports concept testing for any text-based stimulus — positioning statements, taglines, value propositions, product descriptions, campaign headlines, or early-stage messaging frameworks. The interview guide is designed by the research team, which means the approach can be adapted to the specific decision the team is trying to make. See Sonderly's documentation for current guidance on presenting visual or multi-format stimulus material.
How many respondents do I need for a message test?
Because the goal is thematic depth rather than statistical representativeness, qualitative concept tests require far fewer respondents than quantitative surveys. Thematic saturation — the point at which new respondents stop surfacing new themes — tends to occur well before the sample sizes quantitative methods require. In practice, a focused message test with 20 to 35 respondents from a well-defined audience is typically sufficient to identify the dominant patterns across reaction types and surface unexpected themes. The more homogeneous the target audience, the earlier saturation tends to occur.
Can I use Sonderly to test with my own customer base versus a recruited panel?
Yes. Sonderly operates on a bring-your-own-audience model, which means you control who receives the interview link. You can test with your existing customer base through CRM outreach, with a recruited panel through a panel partner, or with a warmer audience identified through your own channels. The research team determines the sample; Sonderly handles the interviewing and analysis.
How does the analysis handle respondents who gave short or low-quality answers?
Sonderly's thematic analysis is built across all transcripts, and the researcher retains access to each individual session. Short or low-engagement sessions can be reviewed and excluded from analysis at the researcher's discretion. The link between every theme and its source transcripts means the team can assess the quality and depth of the evidence behind each finding before acting on it.
Conclusion
The decision to run a campaign on a message that has not been tested is a bet on intuition — sometimes the right call when speed matters, but a more expensive bet the larger the commitment. Qualitative message testing, at the speed and scale that AI-powered interviews make possible, removes much of the reason that bet feels necessary.
Sonderly is built for research and brand teams who want to understand their audience before they spend: what the message communicates in respondents' own words, what they find credible and what they do not, and — crucially — what they bring up unprompted that the team had not thought to ask about. That last category is often where the most valuable finding lives.
A message test can be live and collecting responses within an hour. The analysis returns themes with source-linked quotes, accessible at the transcript level for verification. The result is the kind of qualitative evidence that makes a campaign decision feel grounded: not just a winner, but a reason, a constraint, and a brief.
Run a message test with Sonderly →