Home › How We Rank From Community Reviews
Method · Protocol v1 · October 2026How We Rank From Community Reviews
- Oct 2026 — Protocol v1 published. Sibling document to our hands-on lab protocol (/how-we-test).
The promise: “We read N thousand Reddit and forum comments so you don't have to.” We really read them. We count them. We filter the shills. We show our work.
Some products can't go in a lab. Nobody can A/B test a moisturizer's “feel,” a fragrance's sillage, or whether a pet bed survives a determined dachshund — not honestly, not at scale. For those categories, the best available evidence is what thousands of experienced owners and users say, unprompted, in communities where reputations are on the line. This protocol is how we turn that evidence into rankings we can defend — as rigorous, as auditable, and as hostile to bullshit as our lab protocol.
Wirecutter has always done a version of this: their job postings require writers to “methodically scan user reviews and forums to find what truly matters to readers.” We agree — and we publish the method instead of keeping it in the newsroom.
1. Sampling
Where we read (by niche type)
We sample from communities where the niche's real enthusiasts argue about products — not from brand-owned spaces, not from deal subs, not from subs with no moderation.
| Niche type | Primary communities | Corroboration only |
|---|---|---|
| Beauty / skincare | r/SkincareAddiction, r/AsianBeauty, r/30PlusSkincare, MakeupAlley forums | r/beauty, brand subreddits |
| Fragrance | r/fragrance, r/Colognes, r/DesiFragranceAddicts | Fragrantica forums |
| Fashion / apparel | r/malefashionadvice, r/femalefashionadvice, r/BuyItForLife, r/GoodyearWelt (boots), r/rawdenim (denim) | StyleForum |
| Home / kitchenware | r/BuyItForLife, r/Cooking, r/castiron, r/knives | Houzz forums |
| Pet | r/dogs, r/cats, r/puppy101, r/CatAdvice | brand subreddits |
| Software / digital | r/productivity, niche subs per tool, Hacker News threads | G2/Capterra reviews (too gameable) |
Rules for community selection:
- Moderated first. We prefer communities with active mods and written self-promotion / disclosure rules (r/SkincareAddiction requires posters to disclose brand affiliations; r/BuyItForLife bans self-promo outright). Strict moderation is our first shill filter, and it runs before we ever open a thread.
- No brand-owned or brand-moderated spaces in the tally. A brand's own subreddit is marketing. It can be quoted for context; it never counts.
- Minimum community count per ranking: 2. One community is an echo chamber. A consensus that only exists in one room isn't a consensus.
How threads are selected
For each product category, we build the sample from four thread types — deliberately, not by vibes:
- Recommendation threads — “what should I buy?” / “holy grail” / “best X for Y” posts. Found via Reddit search on 4–6 seed keywords, sorted by relevance, then filtered by date.
- Comparison threads — “X vs Y” posts. These are where real tradeoffs surface.
- Routine / collection threads — “show your routine,” “what's in your rotation,” unprompted ownership evidence.
- Dissent threads — “regrets,” “overrated,” “don't buy,” “what did you return” posts. Mandatory. A sample with no dissent sampling has a positivity bias by construction. If nobody warns against a product, that's a finding. If we never looked, that's malpractice.
Date ranges and recency
- Primary window: the last 12 months. Formulations change, brands get acquired, manufacturing moves, quality drifts. A 2022 consensus about a serum that was reformulated in 2025 is misinformation.
- Classic backfill: up to 24 months, only for “all-time greats” threads and only from still-active accounts. Backfill can corroborate; it cannot set the ranking.
- Why recency matters, concretely: beauty and fragrance products get silently reformulated; fashion brands change factories; a brand caught astroturfing in March should not ride its February reputation. Every sample carries a date window, and every ranking carries a sample date.
Minimum sample sizes (the consensus threshold)
| Claim level | Minimum bar |
|---|---|
| Consensus ranking (we publish a ranked top 5) | ≥500 countable mentions across ≥20 threads, from ≥2 communities, within the primary window |
| Emerging pick (flagged as provisional) | ≥100 countable mentions across ≥8 threads |
| No ranking | Below that: we write “not enough signal” and publish the raw thread list instead |
A product enters a published ranking only with ≥5 distinct-recommender mentions across ≥3 separate threads. Anything thinner is anecdote, and we say so.
2. Counting
What counts as a recommendation
Not every mention is a vote. We score each mention into one of four buckets:
- Full recommendation (counts: 1). An explicit statement of use and satisfaction: “I've used X daily for a year, it's the best I've tried.” First-hand experience required.
- Qualified recommendation (counts: 0.5). Positive but hedged: “X worked well for my dry skin, might not suit oily types.” Counts at half — useful signal, honest about its limits.
- Mention (counts: 0). Passing name-drops, “I've heard X is good,” links without commentary, “my friend uses X.” Not evidence of anything except awareness.
- Warning (counts separately, as −1 in the warning tally).“X broke me out,” “X fell apart in six months,” “avoid X.” Warnings are tallied on their own axis. A product with 40 recommendations and 30 warnings is not a winner — it's polarizing, and we label it that.
Weighting
- No upvote weighting in the primary tally. Upvotes are the most gameable signal on Reddit (see section 3). A recommendation from an account with three years of history counts the same as one with three months — what matters is whether it survives the shill filter, not its karma.
- Upvotes are a tiebreaker only. If two products are within 10% on raw recommendations, the one whose recommendation comments carry higher median net upvotes gets the edge — and we say so in the disclosure.
- One user, one vote per product per thread. Enthusiasts repeat themselves. That's fine; it counts once.
- X vs Y threads are scored per product: each explicit preference statement is a full or qualified recommendation for the preferred product and, where stated, a warning for the rejected one. “X vs Y” polls with no reasoning are mentions, not recommendations.
What “counts” requires
Every counted recommendation must be traceable: thread URL, comment permalink, bucket assignment, and the username's account age at time of sampling. The tally sheet is the source of truth. If it isn't in the tally sheet, it doesn't exist.
3. Shill, brigade, and bot detection
Reddit acknowledged in 2026 that AI-generated spam and coordinated astroturfing are a live, growing problem — brands deploying AI agents to plant recommendations at scale, specifically to influence both search and AI answers. Communities like r/BuyItForLife have reported brigading by ad bots. Our protocol assumes a hostile environment.
Red flags (any one triggers review; two or more triggers exclusion)
- Account age under 90 days with posting history concentrated on one brand or product category.
- First post is a product recommendation. Genuine new users ask questions; shills arrive with answers.
- Repeated phrasing. The same recommendation sentence (or near-identical wording) appearing from 2+ accounts. Copy-paste is the oldest tell and still the most common.
- Flair disclosure. “Brand representative,” “employee,” “sponsored” — these are honest accounts and we respect the honesty; they are excluded from the tally and may be quoted separately with the flair noted.
- Vote anomalies. A comment at +200 with three replies, posted 20 minutes ago, in a thread with 40 total upvotes. We can't see vote timestamps; we can see implausible ratios, and we exclude on suspicion.
- Cross-sub campaigns. The same product pushed with similar phrasing across multiple unrelated subs within a 48-hour window.
- AI-generated texture. Generic, list-y, no specifics, no personal detail, slightly-too-polished grammar — the 2026 spam signature. When a comment reads like it was generated, we treat it as generated.
- Removed by moderators / deleted accounts. Excluded automatically. The mods did our job for us.
What gets excluded — and logged
Everything above is excluded from the tally. Each exclusion is logged in the tally sheet: thread URL, comment text (quoted), username, and the reason. The disclosure block on the published page states the exclusion count (section 4). The full exclusion log stays internal — publishing usernames invites harassment — but the count and the categories are public.
How strict community rules help us
We deliberately overweight our sample toward communities with:
- Written self-promotion bans and active enforcement (r/BuyItForLife, r/SkincareAddiction).
- Requirements to disclose brand affiliations.
- “No affiliate links” rules, which correlate strongly with low shill density.
- Long-tenured mod teams — institutional memory catches repeat offenders.
A community's moderation quality is part of our sample design, not an afterthought. When we cite r/SkincareAddiction's promo rules as a reason to trust a sample, that's a falsifiable claim — anyone can go check.
4. Disclosure format
Every sentiment page carries this block, in full, above the ranking. No exceptions.
How this ranking was made. Based on [N] countable mentions across [T] threads in [communities], sampled [date range]. Method: /how-we-rank-sentiment. Sample date: [month year]. Tallied blind to commissions — the counts were locked before we checked which products pay us. [E] comments excluded as suspected shills or bots. Representative threads: [3+ links]. Full tally sheet available on request.
A real example of the format (numbers illustrative — see section 7):
How this ranking was made. Based on 2,300 countable mentions across 41 threads in r/fragrance and r/Colognes, sampled Jan–Oct 2026. Method: /how-we-rank-sentiment. Sample date: Oct 2026. Tallied blind to commissions — the counts were locked before we checked which products pay us. 11 comments excluded as suspected shills or bots. Representative threads: [link] [link] [link]. Full tally sheet available on request.
Required elements, every time:
- Countable mention count (not “comments read” — the scored ones).
- Thread count and community names.
- Sample window and sample date.
- The commission-blind statement, verbatim.
- Exclusion count.
- Minimum 3 representative thread links — including at least one dissent thread when dissent existed.
- A named human who did the tally, and their contact or author page. No anonymous methodology.
What we never do: “Reddit loves X.” “The internet agrees.” “Everyone says.” Those are vibes. We publish numbers or we publish nothing.
5. Re-sampling cadence
- Quarterly re-check. Every sentiment ranking is re-examined every 3 months: 5 fresh threads per category, compared against the published tally. If the top 3 hold, we update the sample date and move on.
- Full re-sample (new 12-month window, full protocol) when: a new product crosses the emerging-pick threshold in re-checks; the category's community migrates (a sub dies, a new one becomes the hub); or annually, whichever comes first.
- Black Friday / holiday re-sample. Full re-check of every published sentiment ranking in the last week of October, before holiday buying guides go live. This is when shill volume spikes — brands buy campaigns for Q4 — so the shill filter gets extra scrutiny, not less.
Out-of-cycle triggers (re-sample within 14 days):
- Product recall or safety issue.
- Reformulation or ownership change (acquisition, factory move).
- Scandal: brand caught astroturfing, fake reviews, or deceptive marketing — the brand's historical tally is frozen and re-evaluated, not grandfathered in.
- A viral thread (>1,000 upvotes, or front-page of the niche sub) materially shifting opinion on a ranked product.
- A correction request that survives our review (see section 6).
6. Bright lines (non-negotiable)
- Never invent counts. Every number on a sentiment page traces to a row in the tally sheet. An invented number is indistinguishable from fraud and treated as such — it's a firing-level offense for anyone who works on this.
- The commission-blind rule. The tally is completed, locked, and timestamped before anyone checks which products have affiliate programs or what they pay. We document the lock timestamp and the first affiliate-lookup timestamp. If the lookup came first, the ranking doesn't ship — we redo it. This is the single most important rule in this document.
- Never cherry-pick threads to favor a high-commission product.Thread selection follows the four-type quota in section 1, in keyword order, before any counting begins. Adding a thread after the tally “because it supports the pick” is cherry-picking. Removing a dissent thread is cherry-picking. Both are disqualifying.
- Always link representative threads. Minimum 3 per ranking, including dissent. A reader should be able to click through and verify the vibe we claim.
- Always mark sample dates. Undated consensus is how stale advice lives forever. The sample date is in the disclosure block and next to the ranking.
- No AI summarization as a substitute for reading. Tools may find threads; they may not count mentions. If a human didn't read the comment, it doesn't count. (This is also a defense against the 2026 AI-spam wave — summarized spam looks like consensus.)
- Corrections policy. If the community tells us we got it wrong — a thread, a comment, an email — we re-sample within 14 days. If the re-sample disagrees with the published ranking, we correct publicly within 72 hours, note the correction on the page with the date, and keep a public corrections log. Being wrong is survivable. Hiding it isn't.
- Below threshold, no ranking. If the sample doesn't clear the consensus bar, we publish the thread list and say “not enough signal.” An honest “we don't know yet” beats a fabricated top 5 every time.
7. Worked example (illustrative)
All numbers below are illustrative — a walkthrough of the process, not a real sample. No real product data is presented here.
Category: vitamin C serums. Question: which serum does the skincare community actually recommend most?
Step 1 — Thread selection. Seed keywords: “vitamin C serum recommendation,” “best vitamin C serum,” “vitamin C holy grail,” “vitamin C overrated,” “vitamin C regret.” Reddit search across r/SkincareAddiction, r/AsianBeauty, r/30PlusSkincare, primary window Jan–Oct 2026. Result: 34 threads — 14 recommendation, 8 comparison, 6 routine, 6 dissent. Two communities minimum satisfied (three, in fact).
Step 2 — Counting. Suppose 1,240 comments mention a serum by name. Scoring:
- 610 full recommendations (“I've repurchased X three times, nothing else compares”).
- 220 qualified (“X is great if you can handle the smell”).
- 310 mentions (“heard good things about X” — discarded from tally).
- 100 warnings (“X oxidized in a month,” “X gave me cystic acne”).
Product A: 180 full + 60 qualified (210 points), 12 warnings → strong, clean.
Product B: 150 full + 40 qualified (170 points), 55 warnings → high volume, polarizing — flagged, not ranked #1 despite volume.
Product C: 90 full + 30 qualified (105 points), 4 warnings → solid #2 or #3.
Step 3 — Shill filtering. Suppose 9 comments excluded: 4 from accounts under 30 days old whose only posts praise Product D with identical phrasing (“game changer for my dark spots!!” × 4); 2 with brand-rep flair; 2 vote-anomaly comments (+180, posted 25 minutes prior, in a 60-upvote thread); 1 removed by mods mid-sample. Logged with URLs and reasons. Product D drops from “emerging” to below threshold once the 4 shill comments are removed — which is exactly what the filter is for.
Step 4 — The published ranking. Products A and C clear the bar (≥5 distinct recommenders, ≥3 threads each). The page shows:
How this ranking was made. Based on 830 countable mentions across 34 threads in r/SkincareAddiction, r/AsianBeauty, and r/30PlusSkincare, sampled Jan–Oct 2026. Method: /how-we-rank-sentiment. Sample date: Oct 2026. Tallied blind to commissions — the counts were locked before we checked which products pay us. 9 comments excluded as suspected shills or bots. Representative threads: [link] [link] [link, dissent thread]. Full tally sheet available on request.
1. Product A — 210 recommendation points, 12 warnings. The consensus pick.
2. Product C — 105 points, 4 warnings. The quiet runner-up.
Not ranked: Product B — 170 points but 55 warnings. Too polarizing for a general recommendation; worth trying if your skin tolerates strong actives, and we say exactly that.
Note what happened: the highest-commission product didn't win, the loudest product didn't win, and the polarizing product got an honest paragraph instead of a rank. That's the protocol working.
8. What this method can't do (honest limits)
- It's not a lab. We can't verify that commenters actually used the product, or used it correctly. We mitigate with shill filtering and sample size; we don't pretend it away.
- It's US-and-English skewed. Reddit's user base is not the world. We disclose the audience; we don't claim universality.
- It's slow. A proper sample takes days of human reading. That's the cost of the claim “we read them.” We pay it.
- It doesn't cover everything. Health, supplements, and medical devices are excluded from sentiment rankings entirely — the liability of aggregating crowd opinion on things people ingest is not a risk we take, no matter how rigorous the method.
RankedTrue — rankings you can audit. This protocol is version 1, published October 2026. When the protocol changes, the version number changes, and the changelog is public.