Home › How We Rank From Community Reviews

Method · Protocol v1 · October 2026

How We Rank From Community Reviews

Affiliate disclosure: Affiliated, never sponsored. We earn commissions through links on this page, but commissions don't move rankings — picks with no affiliate program are still mentioned where they're the honest choice. See “How we test” for the protocol behind every score.
What's changed
  • Oct 2026 — Protocol v1 published. Sibling document to our hands-on lab protocol (/how-we-test).

The promise: “We read N thousand Reddit and forum comments so you don't have to.” We really read them. We count them. We filter the shills. We show our work.

Some products can't go in a lab. Nobody can A/B test a moisturizer's “feel,” a fragrance's sillage, or whether a pet bed survives a determined dachshund — not honestly, not at scale. For those categories, the best available evidence is what thousands of experienced owners and users say, unprompted, in communities where reputations are on the line. This protocol is how we turn that evidence into rankings we can defend — as rigorous, as auditable, and as hostile to bullshit as our lab protocol.

Wirecutter has always done a version of this: their job postings require writers to “methodically scan user reviews and forums to find what truly matters to readers.” We agree — and we publish the method instead of keeping it in the newsroom.

1. Sampling

Where we read (by niche type)

We sample from communities where the niche's real enthusiasts argue about products — not from brand-owned spaces, not from deal subs, not from subs with no moderation.

Niche typePrimary communitiesCorroboration only
Beauty / skincarer/SkincareAddiction, r/AsianBeauty, r/30PlusSkincare, MakeupAlley forumsr/beauty, brand subreddits
Fragrancer/fragrance, r/Colognes, r/DesiFragranceAddictsFragrantica forums
Fashion / apparelr/malefashionadvice, r/femalefashionadvice, r/BuyItForLife, r/GoodyearWelt (boots), r/rawdenim (denim)StyleForum
Home / kitchenwarer/BuyItForLife, r/Cooking, r/castiron, r/knivesHouzz forums
Petr/dogs, r/cats, r/puppy101, r/CatAdvicebrand subreddits
Software / digitalr/productivity, niche subs per tool, Hacker News threadsG2/Capterra reviews (too gameable)

Rules for community selection:

How threads are selected

For each product category, we build the sample from four thread types — deliberately, not by vibes:

  1. Recommendation threads — “what should I buy?” / “holy grail” / “best X for Y” posts. Found via Reddit search on 4–6 seed keywords, sorted by relevance, then filtered by date.
  2. Comparison threads — “X vs Y” posts. These are where real tradeoffs surface.
  3. Routine / collection threads — “show your routine,” “what's in your rotation,” unprompted ownership evidence.
  4. Dissent threads — “regrets,” “overrated,” “don't buy,” “what did you return” posts. Mandatory. A sample with no dissent sampling has a positivity bias by construction. If nobody warns against a product, that's a finding. If we never looked, that's malpractice.

Date ranges and recency

Minimum sample sizes (the consensus threshold)

Claim levelMinimum bar
Consensus ranking (we publish a ranked top 5)≥500 countable mentions across ≥20 threads, from ≥2 communities, within the primary window
Emerging pick (flagged as provisional)≥100 countable mentions across ≥8 threads
No rankingBelow that: we write “not enough signal” and publish the raw thread list instead

A product enters a published ranking only with ≥5 distinct-recommender mentions across ≥3 separate threads. Anything thinner is anecdote, and we say so.

2. Counting

What counts as a recommendation

Not every mention is a vote. We score each mention into one of four buckets:

Weighting

What “counts” requires

Every counted recommendation must be traceable: thread URL, comment permalink, bucket assignment, and the username's account age at time of sampling. The tally sheet is the source of truth. If it isn't in the tally sheet, it doesn't exist.

3. Shill, brigade, and bot detection

Reddit acknowledged in 2026 that AI-generated spam and coordinated astroturfing are a live, growing problem — brands deploying AI agents to plant recommendations at scale, specifically to influence both search and AI answers. Communities like r/BuyItForLife have reported brigading by ad bots. Our protocol assumes a hostile environment.

Red flags (any one triggers review; two or more triggers exclusion)

  1. Account age under 90 days with posting history concentrated on one brand or product category.
  2. First post is a product recommendation. Genuine new users ask questions; shills arrive with answers.
  3. Repeated phrasing. The same recommendation sentence (or near-identical wording) appearing from 2+ accounts. Copy-paste is the oldest tell and still the most common.
  4. Flair disclosure. “Brand representative,” “employee,” “sponsored” — these are honest accounts and we respect the honesty; they are excluded from the tally and may be quoted separately with the flair noted.
  5. Vote anomalies. A comment at +200 with three replies, posted 20 minutes ago, in a thread with 40 total upvotes. We can't see vote timestamps; we can see implausible ratios, and we exclude on suspicion.
  6. Cross-sub campaigns. The same product pushed with similar phrasing across multiple unrelated subs within a 48-hour window.
  7. AI-generated texture. Generic, list-y, no specifics, no personal detail, slightly-too-polished grammar — the 2026 spam signature. When a comment reads like it was generated, we treat it as generated.
  8. Removed by moderators / deleted accounts. Excluded automatically. The mods did our job for us.

What gets excluded — and logged

Everything above is excluded from the tally. Each exclusion is logged in the tally sheet: thread URL, comment text (quoted), username, and the reason. The disclosure block on the published page states the exclusion count (section 4). The full exclusion log stays internal — publishing usernames invites harassment — but the count and the categories are public.

How strict community rules help us

We deliberately overweight our sample toward communities with:

A community's moderation quality is part of our sample design, not an afterthought. When we cite r/SkincareAddiction's promo rules as a reason to trust a sample, that's a falsifiable claim — anyone can go check.

4. Disclosure format

Every sentiment page carries this block, in full, above the ranking. No exceptions.

How this ranking was made. Based on [N] countable mentions across [T] threads in [communities], sampled [date range]. Method: /how-we-rank-sentiment. Sample date: [month year]. Tallied blind to commissions — the counts were locked before we checked which products pay us. [E] comments excluded as suspected shills or bots. Representative threads: [3+ links]. Full tally sheet available on request.

A real example of the format (numbers illustrative — see section 7):

How this ranking was made. Based on 2,300 countable mentions across 41 threads in r/fragrance and r/Colognes, sampled Jan–Oct 2026. Method: /how-we-rank-sentiment. Sample date: Oct 2026. Tallied blind to commissions — the counts were locked before we checked which products pay us. 11 comments excluded as suspected shills or bots. Representative threads: [link] [link] [link]. Full tally sheet available on request.

Required elements, every time:

  1. Countable mention count (not “comments read” — the scored ones).
  2. Thread count and community names.
  3. Sample window and sample date.
  4. The commission-blind statement, verbatim.
  5. Exclusion count.
  6. Minimum 3 representative thread links — including at least one dissent thread when dissent existed.
  7. A named human who did the tally, and their contact or author page. No anonymous methodology.

What we never do: “Reddit loves X.” “The internet agrees.” “Everyone says.” Those are vibes. We publish numbers or we publish nothing.

5. Re-sampling cadence

Out-of-cycle triggers (re-sample within 14 days):

6. Bright lines (non-negotiable)

  1. Never invent counts. Every number on a sentiment page traces to a row in the tally sheet. An invented number is indistinguishable from fraud and treated as such — it's a firing-level offense for anyone who works on this.
  2. The commission-blind rule. The tally is completed, locked, and timestamped before anyone checks which products have affiliate programs or what they pay. We document the lock timestamp and the first affiliate-lookup timestamp. If the lookup came first, the ranking doesn't ship — we redo it. This is the single most important rule in this document.
  3. Never cherry-pick threads to favor a high-commission product.Thread selection follows the four-type quota in section 1, in keyword order, before any counting begins. Adding a thread after the tally “because it supports the pick” is cherry-picking. Removing a dissent thread is cherry-picking. Both are disqualifying.
  4. Always link representative threads. Minimum 3 per ranking, including dissent. A reader should be able to click through and verify the vibe we claim.
  5. Always mark sample dates. Undated consensus is how stale advice lives forever. The sample date is in the disclosure block and next to the ranking.
  6. No AI summarization as a substitute for reading. Tools may find threads; they may not count mentions. If a human didn't read the comment, it doesn't count. (This is also a defense against the 2026 AI-spam wave — summarized spam looks like consensus.)
  7. Corrections policy. If the community tells us we got it wrong — a thread, a comment, an email — we re-sample within 14 days. If the re-sample disagrees with the published ranking, we correct publicly within 72 hours, note the correction on the page with the date, and keep a public corrections log. Being wrong is survivable. Hiding it isn't.
  8. Below threshold, no ranking. If the sample doesn't clear the consensus bar, we publish the thread list and say “not enough signal.” An honest “we don't know yet” beats a fabricated top 5 every time.

7. Worked example (illustrative)

All numbers below are illustrative — a walkthrough of the process, not a real sample. No real product data is presented here.

Category: vitamin C serums. Question: which serum does the skincare community actually recommend most?

Step 1 — Thread selection. Seed keywords: “vitamin C serum recommendation,” “best vitamin C serum,” “vitamin C holy grail,” “vitamin C overrated,” “vitamin C regret.” Reddit search across r/SkincareAddiction, r/AsianBeauty, r/30PlusSkincare, primary window Jan–Oct 2026. Result: 34 threads — 14 recommendation, 8 comparison, 6 routine, 6 dissent. Two communities minimum satisfied (three, in fact).

Step 2 — Counting. Suppose 1,240 comments mention a serum by name. Scoring:

Product A: 180 full + 60 qualified (210 points), 12 warnings → strong, clean.
Product B: 150 full + 40 qualified (170 points), 55 warnings → high volume, polarizing — flagged, not ranked #1 despite volume.
Product C: 90 full + 30 qualified (105 points), 4 warnings → solid #2 or #3.

Step 3 — Shill filtering. Suppose 9 comments excluded: 4 from accounts under 30 days old whose only posts praise Product D with identical phrasing (“game changer for my dark spots!!” × 4); 2 with brand-rep flair; 2 vote-anomaly comments (+180, posted 25 minutes prior, in a 60-upvote thread); 1 removed by mods mid-sample. Logged with URLs and reasons. Product D drops from “emerging” to below threshold once the 4 shill comments are removed — which is exactly what the filter is for.

Step 4 — The published ranking. Products A and C clear the bar (≥5 distinct recommenders, ≥3 threads each). The page shows:

How this ranking was made. Based on 830 countable mentions across 34 threads in r/SkincareAddiction, r/AsianBeauty, and r/30PlusSkincare, sampled Jan–Oct 2026. Method: /how-we-rank-sentiment. Sample date: Oct 2026. Tallied blind to commissions — the counts were locked before we checked which products pay us. 9 comments excluded as suspected shills or bots. Representative threads: [link] [link] [link, dissent thread]. Full tally sheet available on request.

1. Product A — 210 recommendation points, 12 warnings. The consensus pick.
2. Product C — 105 points, 4 warnings. The quiet runner-up.
Not ranked: Product B — 170 points but 55 warnings. Too polarizing for a general recommendation; worth trying if your skin tolerates strong actives, and we say exactly that.

Note what happened: the highest-commission product didn't win, the loudest product didn't win, and the polarizing product got an honest paragraph instead of a rank. That's the protocol working.

8. What this method can't do (honest limits)

RankedTrue — rankings you can audit. This protocol is version 1, published October 2026. When the protocol changes, the version number changes, and the changelog is public.

Questions, answered

Does this replace lab testing?

No — it covers what labs can't. Some products go through our hands-on test protocol (/how-we-test); others can't be measured honestly at scale — a fragrance's sillage, a moisturizer's feel, whether a pet bed survives a determined dachshund. For those categories, thousands of experienced owners reporting unprompted in moderated communities is the best available evidence, so we count it rigorously instead of pretending we tested it ourselves.

What if the community is wrong?

It happens — communities have fads, reformulations change products, and sometimes the consensus ages badly. That's why every ranking carries a sample date, gets re-checked quarterly, and re-samples within 14 days of a recall, reformulation, or a scandal like astroturfing. If a re-sample disagrees with what we published, we correct publicly within 72 hours and keep a public corrections log. Being wrong is survivable. Hiding it isn't.

How do I report a shill or a suspicious recommendation?

Use the contact page and send the thread link, the comment, and why it looks off. We review every report against the eight red flags in section 3 — fresh account with a first-post recommendation, repeated phrasing, vote anomalies, AI-textured spam. If a report survives review, it triggers an out-of-cycle re-sample within 14 days.