Top keywords, topics, and pain points aggregated across every review in the database. Click any pill or label to filter the explorer.
Click any product card to filter the explorer to just that product's reviews. The chart shows the relative volume across the catalog.
Loadingโฆ
The MUD\WTR Review Explorer is an internal tool for the team to mine 87,460 real customer reviews โ the complete all-time Okendo dataset (the current review platform, including the migrated Bazaarvoice era), with legacy reviews preserved from earlier platforms. Every review is the actual verbatim text the customer wrote: nothing summarized, paraphrased, or rewritten. It exists so anyone scripting ads, writing copy, planning campaigns, or doing audience research can hear customers in their own words instead of relying on second-hand summaries.
Want to run your own analysis in Excel, Google Sheets, Python, or anywhere else? Grab the complete database as a single CSV: ๐ฅ Download Full Dataset (CSV, ~42 MB) โ requires your site login plus a team password (checked server-side; the file is not at any public URL). Ask Aidan if you need it.
87,460 rows ร 22 columns: rid, date, name, email, age, location, rating, title, body (verbatim), sentiment, verified-buyer flag, inferred gender, product + product group, source, personas (pipe-separated), topics, sub-themes, pain points, emotion score, body length, Bazaarvoice review link. Contains customer emails โ treat it accordingly. Always in sync with the site โ rebuilt automatically whenever the team refreshes review data.
MUD\WTR ran Okendo as the review platform from 2019 until switching to Bazaarvoice in 2024โ2025. When the team pulled a Bazaarvoice export, Bazaarvoice had migrated most of the Okendo history โ but not all of it. About 10,730 reviews from the Okendo era never made it across.
Rather than lose those reviews, this site keeps both. The merge logic:
Every review card shows a colored pill (purple for Bazaarvoice, gray for Okendo) so you always know which platform a review came from. Click the pill to filter the whole list by that source.
Important: review body text is never changed during merge. But the Bazaarvoice CSV export didn't include two pieces of metadata that Okendo had โ the verified-buyer flag and the original sentiment label. So when we found a duplicate (Okendo + Bazaarvoice both have the same review), we filled in those two metadata fields on the Bazaarvoice record using Okendo's values:
The merge takes the best available data from each platform. Body text always comes from Bazaarvoice (matching the live mudwtr.com site exactly). Metadata is taken from whichever platform had it.
Important nuance worth reading once and remembering. Here's the lifecycle of one of those 66,730 verified-buyer reviews:
verified=true automatically (email matched a paid Shopify order)verified=true flag from Okendo's record onto the Bazaarvoice recordResult on the site: the pill says "Bazaarvoice" (because the current authoritative version of the review is the Bazaarvoice one โ matches the live site), and the verified filter passes (because Okendo confirmed this customer paid).
Both are accurate. Nothing was fabricated. We just composed the best data from both records of the same review โ text from Bazaarvoice, verified flag and original sentiment from Okendo. Same review either way, stitched together from two records of it.
The 11,018 Bazaarvoice reviews still showing verified=unknown are reviews that have no Okendo match (mostly post-platform-switch โ customer never used Okendo). For those we genuinely don't know, and the "Verified buyers only" filter excludes them to be safe. Re-exporting from Bazaarvoice with the verified field selected would fix this last 13% gap.
This site has two layers of persona thinking, and both matter. Understanding the difference between them is the most important context for anyone using this tool.
The data-driven layer (8 personas, this site)
The 8 personas on the homepage โ The Wired-and-Tired Quitter, The Ritual Keeper, The Body-Said-No Switcher, The Brain-Fog Fighter, The Wanted-to-Love-It Defector, The Sleep Reclaimer, The Pot-a-Day Prisoner, The Mushroom Stacker โ were re-derived in a 2026 clean-room audit: 12 independent AI readers each studied hundreds of raw reviews and survey answers (with NO access to prior persona work or the official persona doc), 3 judges merged their proposals, and every keyword rule was then adversarially validated against sampled real contexts. Every persona's exact criteria โ keywords, guards, confidence โ are shown on its card.
These personas weren't decided in a marketing room. No agency built them. No focus group informed them. They emerged from clustering reviewers by the language patterns and pain points they used unprompted. A reviewer who writes about "brain fog" and "couldn't focus" gets tagged as a Foggy Grinder because they explicitly said those things. The classification is reproducible โ run the build script and you get the same answers, byte for byte.
This is the unbiased view. It reflects who's already buying MUD\WTR and why they say they bought it. The data has no agenda โ it can't lie about what customers are saying because the customers said it themselves.
The official MUD\WTR library (5 personas, in Notion)
The marketing team also maintains an official 2026 Persona Library (in Notion โ Marketing โ Personas 2026): The Calm Seeker, The Mental Performer, The Maximalist Optimizer, The Gut Guardian, and The Switcher.
Those 5 weren't built from reviews. They were built from market demand data โ SPINS 2025 trend reports, McKinsey 2025 Future of Wellness research, Numerator 2025 consumer behavior data. They answer a different question: "Where is the entire wellness market heading, and which large, valuable audience segments should MUD\WTR be acquiring?"
So the official library is the strategic view. It's about who MUD\WTR should be targeting in 2026 based on where demand is moving (calmative beverages +25%, nootropics +19%, prebiotic soda +88%, etc.), not just who's already in the customer base.
Why we have both, and what they're each used for
These are two answers to two different questions, and we need both:
The validation: they overlap almost perfectly
Here's the encouraging part. When we ran the audit (verified by reading the live Notion doc directly), we found that every fully-developed persona in the official library has a matching cluster in the review data. The marketing team's audience theory isn't speculative โ it's already reflected in who's buying:
What the overlap tells us
Marketing isn't targeting hypothetical audiences โ they're targeting audiences that are already showing up in reviews. That's a powerful signal. It means:
A note on the "aka" labels
Every persona card on the homepage and detail page shows an "aka [Official Name]" subtitle in purple where there's a clean mapping. This isn't a rename โ it's a translation layer. Our data-driven names (Foggy Grinder, Wellness Stacker, etc.) stay because they reflect the actual language reviewers use. The aka makes it possible to walk into a marketing meeting and say "Foggy Grinder reviews tell us X" while the team hears "Mental Performer customer voice tells us X." Same audience, two vocabularies, zero confusion.
The personas come from the original Voice-of-Customer Persona Briefing โ but every claim has been re-audited against the raw data. Here's what's actually under the hood:
scripts/build-data.py and are documented in the V2 briefing.python3 scripts/build-data.py and produce the same numbers byte-for-byte. No black-box AI scoring, no hidden methodology.This is a common question โ and the answer is important. All 87,460 reviews are on this site and fully searchable. The "persona-matched" number reflects how many reviews contain language specific enough to tag with one of the 7 personas. The personas are a classification overlay, not a filter that excludes anything.
Here's the precise breakdown:
The "uncategorized" reviews break down like this:
What this means for your work:
None of the unmatched reviews are "missing" โ they're all here, just not labeled as a specific persona because they don't tell a specific persona's story. The classification is intentionally strict so the persona counts mean something. Everything is searchable; only some things are categorizable.
Every quote you see โ in a persona detail, in search results, in the briefing PDF โ is the actual customer's actual words, exactly as they wrote them. We don't fix typos, edit grammar, or shorten sentences. Style, voice, even ALL-CAPS rants are preserved. This matters because customer language is the asset. If a review reads "I'm SO done with the jitters!!!" you should be able to use that pattern in an ad โ not a polished version of it.
Gender is inferred from the reviewer's first name using a US Social Security baby-name list. Names that appear on both lists (Jordan, Taylor, Casey) are classified as "unknown" rather than guessed. About 38% of reviewers fall into "unknown." Of the 48,542 names that can be classified, the split is 58.6% female / 41.4% male. Age is not provided in the dataset โ age ranges in persona descriptions are qualitative inferences from review language, not statistics.
Whenever the team pushes a new version, the version tag in the top bar updates and you'll see a purple banner offering "Refresh now." Background polling runs every 5 minutes plus an instant check whenever you switch back to the tab โ so you'll never spend more than a few minutes on a stale build.
For the full strategic narrative โ pain points by persona, top emotional triggers, demographic breakdown, ad-ready quote list, and a transparent methodology section โ see MUD-WTR-Voice-of-Customer-Persona-Briefing-V2.md (also available as a .docx for Google Docs import).
This download is protected by a team password. If you don't have it, ask Aidan.
87,460 reviews ยท 22 columns (incl. email, age, location) ยท ~42 MB CSV.
Institutional memory for engineering decisions, formulas, and "why" behind the things you see. The About modal explains what the site does for users; this one explains how it works under the hood and why specific decisions were made. Read top-to-bottom or jump to a section.
The site used to show a "Source" pill on every review revealing "Bazaarvoice" or "Okendo" on hover, plus an "All Sources" dropdown filter and a source chip in the active-filters bar.
Why we removed it: new team members getting onboarded didn't know what Bazaarvoice / Okendo were. The distinction was technical-platform trivia (the review collection vendors MUD\WTR used) and not actionable for the people using this tool to mine customer language for ads, copy, and strategy. Worse, it was confusing โ people kept asking "what does Bazaarvoice mean?" and "should I filter by it?" The answer to both is basically "no, it doesn't matter for what you're trying to do."
What's still there: the underlying source field is still stored on every review and still flows through the API. We just stripped the UI surface (pill on cards, dropdown filter, scope chip, and the JS handlers that drove them). If you ever need to filter by source for debugging, the API still accepts a source param.
For background on what the two sources actually are: see the About modal โ "Why two sources?" section. Short version: Okendo was the review platform 2019โ2024, Bazaarvoice is the current one (2024+). The dataset is a merge of both with body-text fingerprint dedup, preserving 10,730 unique Okendo reviews that didn't make it to Bazaarvoice.
Every review gets an emotion score used by the "Most Emotional" sort and the small "Emotion: X" label on each card. The formula (in build-data.py):
excl = number of "!" characters in body caps = number of ALL-CAPS words (2+ letters) in body length_bonus = min(len(body) / 500, 2.0) emotion = round(excl * 0.3 + caps * 0.2 + length_bonus, 2)
So a long all-caps review with lots of exclamation points scores very high (10+), and a short polite review scores under 1. The cap on the "Emotional Intensity" bar viz is 12 to give visual headroom; some reviews score 50+ but the bar pegs at 100% above 12.
Each review is tagged with 0+ personas based on keyword pattern matching against the body text. Rules live in persona_rules at the top of build-data.py. Each persona has 30โ50 trigger phrases; if any phrase matches (word-boundary regex, case-insensitive), the persona is added.
This is best-effort tagging, not ground truth. A review mentioning "brain fog" sarcastically would still tag as Foggy Grinder. Strength: it's transparent and tweakable. Weakness: keyword lists can miss novel phrasings. The 7 personas capture roughly 4 in 10 of the 87,460 reviews; the rest are either short generic reviews ("Love it!") or use language outside the rule sets.
Personas are data-driven (mined from this dataset). MUD\WTR's official marketing 2026 Persona Library has 5 personas derived from external trend research (McKinsey, SPINS). The aka labels on persona cards bridge the two vocabularies.
Same approach as personas: keyword lists. Pains include "jitters", "crash", "anxiety", "tired", "brain fog", etc. Topics include "Coffee", "Taste", "Ritual", "Energy", "Sleep", etc. Multi-tag per review (a review can hit several pains and several topics).
When Bazaarvoice review titles were added to the site in v6.6, a positional cross-reference audit confirmed 100.0000% mapping accuracy across all 74,282 BV reviews. Bodies and titles are read from the same CSV row in the same loop pass โ there's no fingerprint lookup or join, so they cannot be mismatched. 99.87% of BV reviews have a title; 0.13% (96 reviews) have an empty title because the reviewer left the title field blank.
Okendo's export has no title field โ those 10,730 reviews render without a title line (no placeholder).
The CSV contains 1,623 rows with body text that exactly matches at least one other row. 462 unique body strings appear 2+ times. Most are short generic phrases like "Love it" (156 reviewers, 5+ years of dates) or "Great" (64 reviewers). 36 are paragraph-length duplicates (200+ chars) and almost all of those are clear bundle purchases โ same customer reviewing multiple products with copy-pasted body text but different titles.
The build script does not dedupe by body. Each CSV row becomes its own review record with its own product / title / date. This is correct โ they're real separate reviews, not data integrity issues. A full breakdown lives in duplicate-reviews-audit.csv at the workspace root.
The footer link ๐ฅ Download Full Dataset (CSV) calls POST /api/download, which gzip-streams data/exports/all-reviews.csv โ written by build-data.py on every rebuild from the same data the site reads. Always in sync.
v9.0 hardening: the CSV now contains customer emails, so the v8.16 setup (client-side JS password guarding a public static URL) was retired. The file no longer exists anywhere under /public. The endpoint verifies the site auth token AND the team download password server-side. The password comes from the DOWNLOAD_PASSWORD Vercel env var (with a legacy fallback baked in โ set the env var to rotate it without a deploy).
main branch on every push)/api (search, stats, persona, product, recent, etc.)/data (reviews.json, personas.json, stats.json, products.json, changelog.json). Built by scripts/build-data.py from the source CSVs./api/auth. Site password lives in Vercel env vars (SITE_PASSWORD).python3 scripts/build-data.py, commit the new /data JSON files, push. Vercel redeploys automatically.The full version-by-version changelog is in the ๐ Version History button in this footer (or click the version tag in the top bar). Every release back to v3.x is documented with the why, not just the what.