Why attribution alone gets the number wrong
Every common podcast attribution method misses something, and the misses do not cancel out.
- Codes and vanity URLs. Right Side Up’s estimate: roughly 20% to 30% of podcast-driven conversions are recorded this way. Codes also travel to coupon sites, where another channel takes the credit for the sale (Right Side Up, updated 2024-02-15).
- Show-level codes. Claritas describes them as directional only, and finds that a vanity URL on its own misses the majority of conversions (Claritas, 2020).
- Post-purchase surveys. Coverage is wider. Fairing gives typical response rates of 40% to 50%, and points out that buyers often recall the wrong source (Fairing, 2024-08-15).
- Pixels. A pixel links a listening household to a site visit, which holds up only for listening at home. Claritas regards mobile and retail IP addresses as unreliable (Claritas, 2020).
Each of these can tell you something. None of them can tell you how many of those customers would have bought anyway. That is the question an incrementality test answers.
What an incrementality test measures
Incrementality test: a comparison of an exposed group against an unexposed control, where the gap between them is the measured effect. Meta’s conversion lift studies are the best-known version. Accounts are split at random, one group is shown the ads and the other is held back, and the gap in conversions is reported as incremental (Meta, retrieved 2026-10-05). Google’s Brand Lift applies the method to awareness and consideration, setting people shown the ads against eligible people who were not shown them (Google, retrieved 2026-10-05).
Podcasts have no platform that randomises listeners for you, so the control has to be designed.
Three designs that work for podcasts
Matched regions
Run the ads in one set of regions and hold out another set with a similar sales history, then compare orders per household in each. This works because most podcast ads are inserted dynamically: the IAB’s 2023 buyer checklist describes dynamic insertion as the dominant method, at more than 80% (IAB, 2023-10), and dynamically inserted ads can be targeted by geography. Baked-in ads cannot, because every listener who downloads the file hears the same ad (IAB Tech Lab, 2024-05), so a regional test has to use dynamically inserted reads.
Exposed and unexposed households
Some measurement vendors offer lift tests that compare households their pixel saw exposed with similar households that were not. The design is quicker to set up than a regional test. It inherits the pixel’s limits, so it is strongest for shows whose audience listens mostly at home.
Before and after
Comparing a baseline period with the flight is the weakest design, because anything else that changed in those weeks lands in the result. Use it when nothing better is possible, and label the result as attributed rather than incremental.
Before launch
- Record a baseline of at least 30 days: orders, branded search, direct traffic and survey answers. Right Side Up recommends the same 30-day baseline (Right Side Up, updated 2023-12-09).
- Write the measurement plan: the question, the design, the regions or households, the read window and the result that would change the budget.
- Check that the test is big enough. A flight too small to separate its effect from normal week-to-week variation will come back flat whatever the ads did.
- Set up every signal before the first placement: codes, vanity URLs, the survey question and the pixel.
During the flight
Keep everything else steady in both groups. A new paid social campaign in the exposed regions, or a promotion that runs only in the holdout, will end up in the result. Log each aircheck as the episodes publish, so any read that was cut short or dropped the offer is on record before the readout.
Reading the result
Response to podcast ads arrives late. Magellan AI’s data shows response continuing to climb through day 30, with a day-7 reading holding less than half of the final total (Magellan AI, 2026-09). Right Side Up allows three to six weeks after the final spot before the result is complete (Right Side Up, updated 2023-12-09). We read no earlier than 30 days after the last episode.
Then report four numbers:
- The difference between the groups, as incremental orders or customers.
- The interval around it, so the reader can see how sure the test is.
- Cost per incremental customer: media and fees divided by incremental customers.
- The same figure next to the brand’s blended CAC for the period.
Use the test to calibrate cheaper signals
A test is expensive to run every month. Its lasting value is the ratio it gives you between what the cheap signals saw and what the ads added. Right Side Up uses post-purchase survey answers to set a multiplier on code- and URL-tracked conversions (Right Side Up, updated 2024-02-15); a holdout lets you set that multiplier from a control instead of an estimate. Between tests, apply it to codes and survey answers, and re-test when the mix of shows changes.
When the result is flat
A flat result means the test could not separate the ads’ effect from zero. It is still worth having. Report it with the interval, check the airchecks and delivery for execution problems, and move the budget. In the IAB’s 2025 survey, proof of ROI was the challenge buyers named most often in creator marketing (IAB, 2025-11). A readout that shows its flat results is easier for finance to trust when it shows a positive one.