Methodology
How we decide whether a bad review has grounds for removal
A language model reads each of your bad reviews against Google's written policies. It follows rules we wrote. We ran those rules over 1,637 real negative reviews and scored them against 164 labeled ones. This page covers which model it is, what it sees, how we measured it, and where it's weakest. Every number comes with what it was counted on.
One thing first. The accuracy figures here measure how often we're right about the grounds. The rest count how often reviews got flagged. None of them is Google's removal rate. We don't have that number yet, and we won't make one up.
The test
What decides whether a review has grounds?
Google's policies decide it. Whether the review is fair doesn't matter. Google doesn't remove reviews on "unfair." It removes them on policy.
Fake Engagement. The clause that matters.
Contributions to Google Maps should reflect a genuine experience at a place or business. Fake engagement is not allowed and will be removed.
Google, Maps user-generated content policy, Prohibited & restricted content, section Fake Engagement
Google's policy text captured August 4, 2026. Section headings re-checked against the live page September 3, 2026.
So the test is the experience behind the review. A review with no real experience behind it, someone paid to post, one person posting from several accounts. Those break the rule. A real customer's bad experience, told respectfully, passes the test, so Google leaves it up. Every verdict with grounds goes back to a line like this one, and we name that policy in Google's words and link it.
The model
Which model reads your reviews?
Gemini 3.1 Pro, Google's language model, in its preview release (google/gemini-3.1-pro-preview). We call it through OpenRouter at temperature 0.1.
For each business it gets the name, the overall rating and the review count. We also tell the model what kind of business you are, using the business type Google Maps lists for you. For each bad review it gets the stars, the first 1,200 characters of text, the reviewer's name, how many reviews that person has written, and the date. If you posted a reply, it also gets the first 200 characters of that reply. It reads up to 25 of your bad reviews at a time. That's so it can see patterns across them, like several reviews landing on the same day.
Reviews with no text, or fewer than ten characters, get counted and shown to you but not analyzed. There's nothing there to read.
The model returns a tier, a confidence level, the policy category, the policy line it relied on, and a one-sentence reason. Here's how that becomes your verdict.
How the model's answer becomes your verdict
- Strong grounds for removal
Most top-tier answers, where the violation is in the review's own words. The reviewer names a different business, or makes a threat. A middle-tier answer also shows as strong grounds when the model is highly confident and the argument isn't off-topic.
- Possible grounds for removal
Every other middle-tier answer, where the argument depends on a pattern or on evidence outside the text. Conflict-of-interest arguments always show as possible grounds, even at the top tier.
- No grounds for removal
A real experience, told respectfully. We find no grounds for Google to remove it.
The rules
What rules does the model follow?
It follows one written prompt. We built it on Google's live content policy and on what people who work on review removals know about how Google applies it. The prompt has three tier definitions, five hard rules and eleven calibration rules. Three of the five hard rules are the conservative ones. "This review is fake" is never enough on its own. A customer experience that sounds real, told critically, is no grounds, always. And when the model is unsure, it says no.
That last rule is deliberate. Google allows one appeal per review, ever. Our read is that the policy named has to match the argument. A wrong yes spends your one appeal on a review with no grounds. So we'd rather miss a review than tell you to file one that isn't there.
The calibration rules came from the 82 reviews where the first version of the prompt and a second, blind judging run disagreed. A few of them are matched pairs. "They're scammers," stated as a general fact with no incident described, is middle tier, which usually shows as possible grounds. The same word inside a detailed account of a real job is no grounds. Two people describing the same bad night from two accounts isn't fake engagement, because companions can each review what happened to them. Someone who only saw your truck on the road can be off-topic, but that's possible grounds, never strong grounds.
The prompt running today is exactly the one we measured, byte for byte. An automated test locks the model and the prompt, so neither can change without us knowing. Any change has to be scored against the same 164 labels and hold precision at 80% or better before it ships. That set of 164 is fixed.
The measurement
How did we measure it?
The pilot ran on 2026-08-04. We took the businesses Google's map results showed for 10 trades in 10 US metros, 120 businesses in all, and pulled their bad reviews with text. That came to 1,637 one- and two-star reviews.
The first version of the prompt and a separate, blind judging run disagreed on 82 reviews, and an expert reviewed each of those 82. That's also where the eleven calibration rules came from. There are 164 labels in all, 71 with grounds and 93 without.
Then we ran the current prompt over all 1,637 reviews. It flagged 86, which is 5.3%. 18 were top tier and 68 middle tier. 70 of those 86 fall inside the 164 labels, so that's what we can score.
The pilot of 2026-08-04, counted
93%
flags with grounds confirmed, on the adjudicated pool
Of reviews we flag as having grounds, 93% were confirmed by expert adjudication against Google’s written policies (65 of 70 flags, n = 164 adjudicated reviews). Our assessment accuracy, not Google’s removal rate; removal outcomes publish as they accrue.
What "expert adjudication" means here. An expert reviewed each of the 82 cases where the two passes disagreed.
91.5%
reviews with grounds that the model flagged, on the adjudicated pool only
65 of the 71 labeled reviews with grounds were flagged in the 2026-08-04 pilot. That's measured inside the 164 and nowhere else, so it's a ceiling, not a measurement of all 1,637 reviews.
41%
pilot businesses with at least one flagged review
In the 2026-08-04 pilot, 117 businesses had at least one negative review with text, and 48 of them had at least one review flagged. Our numbers page has the fuller breakdown. Turn that around and most businesses in the pilot had no flagged review. About six in ten profiles we analyze have no review with grounds for removal. We say so plainly, with a reply playbook attached.
Grounds, not odds
What 93% is, and what it isn't
It's how often we're right that a review has grounds under Google's written policies. It isn't how often Google removes the review. Google makes that call, and we don't have removal outcomes yet. We'll publish them as they come in.
That's why the verdict is grounds, not odds. When a review has possible grounds, we tell you before you pay that appeals like this sometimes succeed, not usually. And whether the grounds are strong or possible, the review gets removed, or your fee gets refunded. The appeal fee only, never subscriptions. The details are on /guarantee.
The limits
Where is the method weakest?
Honestly, in a few places, and you should know them before you trust a verdict.
The pilot didn't measure what we miss.
The pilot didn't adjudicate most of the 1,637, and the 91.5% can't see those reviews. The rules tell the model to say no when it's unsure, and the pilot didn't measure how often that turns a review with grounds into a no.
The pool is small.
70 scored flags, 5 wrong. One more wrong answer would move the 93% by about a point and a half.
The pilot covered 10 trades.
Those were dumpster rental, moving, towing, personal injury law, garage doors, hair salons, auto repair, dental, plumbing and restaurants, all in US cities. If your business is in another trade or another country, the pilot didn't measure it.
Most flags were one kind.
65 of the 86 were off-topic arguments, and 30 of those were road encounters. The model never flagged personal information, hate speech, impersonation or extortion in the pilot, so we have no measurement for them. Conflict of interest is the argument we trust least, which is why it never reads stronger than possible grounds.
The model reads the first 1,200 characters.
If a review runs longer, it doesn't see the end.
The prompt doesn't reread Google's pages.
If Google rewrites a policy, the prompt stays as it is until we change it.
Common questions
Does a person look at my reviews before I see the verdict?
No. The model's answer is the verdict you see, and when you pay, we prepare the appeal from it with no review step in between. That's why we measure the model, and why your fee gets refunded if Google doesn't remove the review.
Why won't you tell me the chance Google removes it?
Because we don't have that number. Our 93% measures the grounds, not Google's decision. Any percentage for removal would be a guess, and we'd rather tell you what the policy says.
Can I see why a review got its verdict?
Yes. Each verdict with grounds names the policy in Google's words, links it, and gives a one-line reason.
Does the model or the prompt ever change?
Not quietly. Any change gets scored against the same 164 labels and has to hold precision at 80% or better before it ships. How the whole process runs, start to finish, is on /how-it-works.
Free · No email · No account
Want to see the verdicts on your own reviews?
Put in your business and we'll analyze every bad review on the profile against Google's policies, the same way this page describes. Each one with at least ten characters of text gets a verdict, free. When there are no grounds, that's the answer. Also free.
No email. No account. The answer is free either way.
