Two hundred client businesses averaging six locations each is twelve hundred locations. At a median velocity of roughly three new Google reviews per location per month, that portfolio takes in about 3,600 reviews a month and 43,200 a year, sitting on top of a historical backlog that is comfortably into six figures. Nobody reads that. The instinct is to hire, or to automate every reply and hope nobody notices. Both are the wrong move, because past a few dozen locations review management stops being a writing task and becomes a routing task.
What follows is how that routing works in ReviewMankey, the AI reputation platform we build at Dude Lemon, using the real product screens rather than mockups, and the 2026 research that says why it now matters more than it did last year.
What “at scale” actually means, in numbers
Review management gets discussed in adjectives. It is worth doing in arithmetic, because the arithmetic is what decides whether a workflow survives the fiftieth client. Here is the portfolio in this article, with every derived figure marked as derived.
| Portfolio input | Value used | Where the figure comes from |
|---|---|---|
| Client businesses | 200 | The scenario in this article. |
| Locations per business | 6 | The scenario in this article. |
| Locations under management | 1,200 | Derived: 200 × 6. |
| New reviews per location per month | 3 at the median, 6.21 for restaurants | Published 2026 Google Business Profile benchmarks. The restaurant figure is new Google reviews per unit per month. |
| New reviews per month, portfolio | 3,600 to 7,452 | Derived: 1,200 × 3, and 1,200 × 6.21. |
| New reviews per year, portfolio | 43,200 to 89,424 | Derived: monthly intake × 12. |
| Lifetime reviews held per location | 47 at the very low end | Businesses ranking in Google’s top three local results average about 47 reviews. Established multi-location brands hold multiples of that. |
| Total reviews under management | Into six figures within a few years | Derived: 1,200 locations at a few hundred lifetime reviews each, plus annual intake. |
Now price the labour. Assume three minutes to read a review, decide the response, write it in the client’s voice and publish it. That is generous for a five-star review and optimistic for an angry one. At 3,600 reviews a month it is 10,800 minutes, or 180 hours, or a little over one full-time person doing nothing else. At restaurant velocity it is 373 hours, more than two people. That is before escalations, before client reporting, and before anyone logs into two hundred separate platform accounts.
Why the stakes went up in 2026
Two independent 2026 datasets and one practitioner survey moved the goalposts in the same direction. Consumers now read reviews more often, demand fresher ones, expect a reply within a day, and increasingly ask an AI assistant rather than a search engine. Meanwhile the average business still leaves more than half of its Google reviews unanswered.
| What changed | The number | Source |
|---|---|---|
| Consumers who always read reviews when browsing for a business | 41%, up from 29% a year earlier | BrightLocal, Local Consumer Review Survey 2026 (1,002 US adults, February 2026) |
| Consumers who will only use a business rated 4.5 stars or higher | 31%, up from 17% | BrightLocal 2026 |
| Consumers more likely to use a business after reading positive reviews | 85%. 77% are put off by negative ones | BrightLocal 2026 |
| Consumers unlikely to use a business that never replies to reviews | 42% | BrightLocal 2026 |
| Consumers who expect a reply the same day they post | 19%, up from 6% | BrightLocal 2026 |
| Consumers who expect a reply by the following day | 32%, up from 18% | BrightLocal 2026 |
| Consumers who used a generative AI tool for a local business recommendation | 45%, up from 6% | BrightLocal 2026 |
| Share of local pack ranking influence attributed to review signals | About 20%, up from 17% in 2023 | Whitespark, 2026 Local Search Ranking Factors, published November 2025 (survey of 47 local SEO practitioners) |
| Google reviews the average business actually answers | 46.9%. On Yelp it is 3.1% | SOCi, 2026 Local Visibility Index (2,751 brands, about 350,000 locations) |
Forty-two percent of consumers say they are unlikely to use a business that never replies. The average business does not reply to 53% of its Google reviews. Nineteen percent of consumers now expect a reply on the day they post, up from six percent a year earlier, while one vendor benchmark built from platform data between January 2025 and February 2026 puts the average time to respond to a Google review at 2.7 days. Every one of those gaps is an operations problem, not a copywriting problem.
What a rating point is actually worth
The academic literature on this is older than the software, and it is unusually clean. Three findings are worth knowing before you argue for a budget.
| Finding | Effect size | Study |
|---|---|---|
| A one-star increase in Yelp rating | 5% to 9% more revenue, concentrated in independent businesses rather than chains with established reputations | Michael Luca, Reviews, Reputation, and Revenue: The Case of Yelp.com, Harvard Business School NOM Unit Working Paper 12-016 |
| A one-point increase in review score on a five-point scale | Price can rise 11.2% while occupancy and market share hold | Chris Anderson, The Impact of Social Media on Lodging Performance, Cornell Hospitality Report |
| Starting to respond to reviews at all | 12% more reviews received, and average rating up 0.12 stars | Davide Proserpio and Georgios Zervas, Online Reputation Management: Estimating the Impact of Management Responses on Consumer Reviews, Marketing Science, 2017 |
One inbox, and one number per business

The screen above is a single account: 1,247 reviews under management, 4.3 stars lifetime, 4.1 over the last thirty days, an 87% response rate, five open actions, and a reputation score of 72 that is up three points on the previous period. The period selector runs 7 days, 30 days, 90 days and 12 months, and the sync indicator is live rather than a nightly batch.
The pair worth staring at is 4.3 lifetime against 4.1 for the last thirty days. A lifetime average is a slow-moving number that hides a live problem for months, because a business with a thousand historical reviews needs an enormous run of bad ones to shift it. The thirty-day figure moves immediately. Whitespark’s 2026 report puts review recency among the highest-weighted individual review factors, which is exactly why a lifetime average is the wrong number to steer by and the right number to report.
At portfolio scale the reputation score does the triage on the dashboards themselves. You cannot read two hundred dashboards a week. You can read two hundred scores, sorted by change, and open the six that moved.
Every source, normalized into one sentiment split

Google Business Profile at 4.3 across 612 reviews, the Play Store 4.1 across 142, the App Store 4.5 across 121, and first-party website reviews 4.6 across 78. Normalizing the payload once at ingest, so that a rating, an author, a timestamp, a platform and a location mean the same thing regardless of where the review came from, is what allows a single queue, a single rule engine and a single metric set to work across sources that otherwise share nothing.
It also surfaces spreads that a blended average erases. There is half a star between the Play Store and the website on this account. An app-store rating problem and a storefront rating problem have different owners and different fixes, and a single blended average tells you about neither. Most platforms in this category cover storefronts or app stores, not both; ReviewMankey’s write-up of managing locations and apps in one place explains why that split exists.
Sentiment sits on its own axis, not derived from the star rating: 57.7% positive, 24.9% neutral, 11.6% negative and 5.8% very negative. That last pair is 217 reviews, and it is the only part of the 1,247 that genuinely needs a person. Separating sentiment from rating matters because a three-star review with a specific, fixable complaint is worth more attention than a one-star review with no text, and the star count cannot tell them apart.
A review is a unit of work, and it has a state

Reviews move across a board: Pending, In Progress, Replied. Cards carry the sentiment the model assigned, the workflow tags a human added, the location and the date, and they filter by platform and by star rating. The drawer on the right opens the whole record: column, sentiment, response state, rating, platform, received timestamp, tags such as VIP and Positive PR, the reviewer, and the full response history.
Making state explicit is what turns response coverage from an estimate into a measurement. The 87% response rate on the dashboard is a count of records in a state, not a guess, which is the only reason it is safe to put in front of a client next to the industry average of 46.9%.
The mechanic that makes hundreds of businesses possible: policy per location

Platforms connect at the account level: one Google account here holds eight locations, another holds four, and the Play Store and App Store connections sit beside them with their own last-sync timestamps. Open the Google drawer and you get accounts, then a searchable and paginated list of all twelve locations.
The badges on each location are the whole point. Reply policy is a property of the location, not the account, not the client, and not the workspace.
| Policy | What it does | When to use it |
|---|---|---|
| Drafts only | Every reply is drafted and held for a human to approve or edit. Nothing publishes on its own. | A new client, a brand with a legal review step, or any location in the middle of a live service problem. |
| Auto-reply 5+ | Five-star reviews get an on-brand reply published automatically. Everything else waits for a person. | The default first promotion. Five-star replies are the highest-volume, lowest-risk category in nearly every portfolio. |
| Auto-reply 4+ | Four and five-star reviews publish automatically. Three stars and below always route to a person. | A location with a settled voice and a track record, where the four-star bucket is large enough to matter. |
| Custom | Per-location voice, language and rules that differ from the account default. | A location in another market or language, or a brand inside the portfolio with its own tone. |
| Needs verification | The location is mapped but not confirmed, so it syncs without publishing. | Mid-onboarding, or after a client changes ownership of a listing. |
This is the difference between software that manages a business and software that manages a portfolio. If reply policy were an account-level setting, onboarding one nervous client would mean either dropping the whole portfolio to drafts-only or asking that client to trust automation on day one. Because policy lives on the location, the risk appetite of one client never constrains another, and a single underperforming branch can be pulled back to drafts-only without touching its eleven siblings.
The onboarding arithmetic follows from the same design. A Google account holding eight locations connects once and brings eight locations with it. Across a portfolio averaging six locations per client, the unit of setup work is the account rather than the location, which is roughly a sixfold reduction in connection steps. What remains is a policy decision per location, which is a judgement call measured in seconds, not an integration measured in hours. For the single-client version of this, ReviewMankey’s own guide to multi-location Google review management covers the same ground at one brand rather than two hundred.
Rules that triage before a human looks

Rules are priority-ordered from P1 down, each individually enabled or paused, each with conditions on the left and consequences on the right. P1, urgent safety concern, matches any platform and any location at one to two stars where sentiment reads urgent or angry and the text contains unsafe, injury or hazard, and assigns critical severity. P2, refund request signal, narrows to Google Business and two specific locations at one to three stars with frustrated or disappointed sentiment and the keywords refund, charged or money back. P3 catches one-star angry reviews. P4 is the low-rating catch-all. P8, potential spam or fake review, sits paused while its thresholds are recalibrated, which is a state a rule engine needs to have.
The match counts are the interesting column. Four enabled rules produced 49 matches in seven days on this account: 27 in the medium catch-all, 11 on refunds, 8 on one-star anger, and 3 critical. That is the real size of the human queue, and it is a list a person can work in an afternoon. The 1,247 reviews never had to be read; the rules read them.
Ordering matters because the rules overlap deliberately. A one-star review mentioning an injury matches both P1 and P4, and priority decides that it inherits critical rather than medium. This is why the interface is an ordered list rather than a flat set of switches, and why the specific rule has to sit above the catch-all.
Test Rules and View Log are the other half. A rule engine you cannot dry-run against history is a rule engine nobody will touch, which means thresholds calcify at whatever they were on day one. Being able to test a proposed rule against back history before enabling it, and then read a log of what actually fired, is what lets one operator tune escalation for two hundred different businesses without breaking any of them.
Writing the rules themselves is a separate exercise, and the severity tiers, owners and SLAs behind them are worth getting right before you enable anything. The review escalation matrix playbook on the ReviewMankey blog is the long version.
A 4.5 means nothing without the market around it

This location holds 4.5 stars, which is 0.2 above the average competitor and trending up 3%, on 847 reviews trending up 8%. It still ranks third of six, and the review-share donut explains why: the leader holds 35% of review volume in the market, this location holds 14%. The comparison refreshes every six hours.
Rating and volume are different problems with different interventions. Rating is a service-and-response problem measured in quarters. Volume is a review-request problem measured in weeks. A per-location market rank across 1,200 locations tells you which of them is fighting for stars and which is fighting for share, and stops you prescribing the same campaign to both.
Reviews are now the input to AI recommendations, not just search
The most consequential 2026 dataset for anyone running a review programme is SOCi’s Local Visibility Index, which covered 2,751 brands and roughly 350,000 locations. Brands appear in Google’s local three-pack 35.9% of the time on average. Only 1.2% of locations are recommended by ChatGPT, and 11.0% by Gemini. Recommended locations average 4.3 stars on ChatGPT, 4.2 on Perplexity and 3.9 on Gemini.
Those averages describe a filter rather than a ranking factor. Fall below roughly four stars and many AI systems stop considering a location at all, no matter what else is true about it. Above the threshold, volume and recency start deciding. And the audience is no longer marginal: 45% of US consumers told BrightLocal they had used a generative AI tool for a local business recommendation, up from 6% the year before.
There is a second-order effect worth naming. When an assistant answers a question about a local business, it reads what is publicly written, and that includes your replies. Reply text stops being customer service and becomes indexable copy about how the business handles problems. A generic template repeated four hundred times reads exactly as what it is. If you want the reviews you earn to also be visible to search engines and AI assistants, that is the job SearchLift does, and our guide to generative engine optimization covers the mechanics.
The operating model, week one to steady state
- Week one, connect at the account level. One platform account brings every location it owns. Do the same for the app stores and the review sites the client actually cares about, then let history backfill before you change anything. You are buying a baseline, not a starting point.
- Week one, set every location to drafts only. Nothing publishes automatically on a portfolio nobody has read yet. This is not caution theatre. It is how you learn a client’s voice from their real reviews rather than from a brand guideline document.
- Week two, write the rules once and reuse them. A safety rule, a refund rule, a one-star rule and a low-rating catch-all cover most of what matters. Dry-run each against back history before enabling it, and read the match counts honestly: a rule that fires two hundred times a week is not an escalation rule, it is a report.
- Week three, promote the easy locations to auto-reply at five stars. Five-star replies are high volume and near-zero risk, and they are where response coverage moves fastest. This single step usually closes most of the gap to a defensible response rate.
- Month two, promote settled locations to auto-reply at four stars and above. Leave three stars and below with a person permanently. The three-star review is where the actual information lives, and it is the one a template will always get wrong.
- Steady state, work the exception queue and the outliers. One score per business, sorted by change. Open the ones that moved, read the critical escalations, and leave the rest running. If your week is spent inside individual reviews at this stage, the rules are mistuned.
What to measure, and what each number is for
| Metric | What it tells you | Reference point |
|---|---|---|
| Response coverage | The share of reviews that received a reply. The single number both consumers and ranking systems react to. | The average business answers 46.9% of its Google reviews and 3.1% of its Yelp reviews (SOCi 2026). |
| Median time to first response | Whether you are meeting the expectation, rather than merely replying eventually. | 19% of consumers expect a reply the same day and 32% by the next day (BrightLocal 2026). |
| Rating over 30 days against lifetime | Whether the recent cohort is better or worse than your history. A lifetime average hides a live problem for months. | 4.3 lifetime with a 4.1 thirty-day, as on the dashboard above, is a warning and not a good quarter. |
| Sentiment mix | The size of the queue that genuinely needs a human, measured independently of the star rating. | Negative plus very negative was 17.4% of reviews, or 217 of 1,247, on the account shown above. |
| Escalation matches by severity | The true size of the human workload, and whether your thresholds are calibrated. | Four enabled rules produced 49 matches in seven days above, of which 3 were critical. |
| Review velocity per location | Whether the review-request programme works, and whether a location looks alive to ranking and AI systems. | A median business collects about three new Google reviews a month. |
| Market rank and review share | Whether a rating problem or a volume problem is holding a location back. | 4.5 stars, third of six, on 14% share against a 35% leader, is a volume problem. |
Where configuration ends and engineering starts
Most of what is above is configuration, and that is the point: a portfolio should not need a build. Four things reliably do. Routing critical escalations into a ticketing or on-call system a client already runs, rather than into email. Pushing review sentiment into a warehouse so it can sit beside sales and staffing data, which is where the causal questions actually get answered. Per-brand voice in a group that owns several distinct brands. And reconciling a client’s own location hierarchy with the messier one the platforms hold, which is the single most common reason a location count does not tie out.
That work is what we build, on the same platform rather than beside it.
Frequently asked questions
How many reviews can ReviewMankey manage?
There is no per-account review ceiling you would meet in practice. The binding constraint at portfolio scale is human attention, which is why reply policy sits on the location and escalation rules sit in priority order. A 1,200-location portfolio taking in roughly 3,600 new reviews a month is a normal shape for the product, and the queue that actually reaches an operator is measured in tens of items a week rather than thousands.
Can one team manage reviews for hundreds of businesses?
Yes, provided three things are true. Reviews arrive in one normalized queue rather than across a hundred separate platform logins. Reply policy is set per location, so a cautious new client and a mature automated one coexist without either compromising. And rules triage on rating, sentiment and keywords before a person looks. Without those three, review response scales linearly with location count, and you are not running a system, you are hiring.
Does ReviewMankey publish AI replies automatically?
Only where you allow it, and only per location. The ladder runs from drafts only, where every reply is held for human approval, through auto-reply at five stars and above, to auto-reply at four stars and above. Three stars and below can be reserved for a person permanently. Published responses are labelled, so a model-drafted reply is distinguishable from a hand-written one after the fact. Automatic publishing applies to Google Business Profile locations; app-store replies are drafted rather than published.
Which review sources does ReviewMankey monitor?
Google Business Profile, Google Play, the Apple App Store and first-party website reviews. Everything is normalized at ingest so that ratings, authors, timestamps, platforms and locations mean the same thing regardless of origin, which is what lets one queue, one rule engine and one metric set span sources that otherwise share nothing. Replies publish back to Google Business Profile; the other sources are monitored, scored and escalated, with replies drafted for you.
Does responding to reviews actually improve your rating?
The best available evidence says yes. Proserpio and Zervas, publishing in Marketing Science in 2017, examined tens of thousands of TripAdvisor hotel reviews and found that when hotels began responding they received 12% more reviews and their ratings rose by an average of 0.12 stars. Consumer research points the same way: BrightLocal’s 2026 survey found 42% of consumers are unlikely to use a business that never replies.
How much is one star worth in revenue?
Two well-known studies put numbers on it. Michael Luca, in a Harvard Business School working paper on Yelp, found a one-star rating increase associated with a 5% to 9% revenue increase, concentrated in independent businesses rather than chains with established reputations. Cornell research by Chris Anderson found that a hotel raising its review score by one point on a five-point scale could raise price by 11.2% while holding occupancy and market share.
How do reviews affect whether AI assistants recommend a business?
Heavily, and differently from Google. SOCi’s 2026 Local Visibility Index, covering 2,751 brands and about 350,000 locations, found brands appear in Google’s local three-pack 35.9% of the time on average, but only 1.2% of locations are recommended by ChatGPT and 11.0% by Gemini. Recommended locations average 4.3 stars on ChatGPT, 4.2 on Perplexity and 3.9 on Gemini, which makes rating behave less like a ranking factor and more like a threshold. Meanwhile 45% of US consumers told BrightLocal they had used a generative AI tool for a local business recommendation, up from 6% the year before.
How long does it take to onboard a multi-location client?
Connecting is account-level, so a platform account holding eight locations arrives in one flow rather than eight. The work that takes real time is not connection. It is deciding the reply policy for each location and dry-running the escalation rules against back history before enabling them. A six-location client is a single working session, and the rules you write for the first client are reusable for the next hundred.
What does ReviewMankey cost?
There is a free plan, and paid plans start at 8.99 per month. Current pricing and plan limits are on the ReviewMankey product page.
The short version
None of this makes review management easy at scale. It makes it bounded. Once a review is a unit of work with an explicit state, once reply policy lives on the location rather than the account, and once ordered rules do the triage, adding the two hundred and first business changes the numbers on a dashboard and nothing else about the job. That is the whole design brief.
ReviewMankey is the platform. If you want the reviews you earn to be visible to search engines and AI assistants as well as to customers, SearchLift is the other half of the problem. And if your portfolio needs review data routed into systems you already run, talk to our team.
Sources
- BrightLocal, Local Consumer Review Survey 2026. 1,002 US adults surveyed in February 2026.
- SOCi, Local Visibility Benchmarks for Multi-Location Brands, 2026 Local Visibility Index. 2,751 brands and approximately 350,000 locations.
- Whitespark, 2026 Local Search Ranking Factors. Survey of 47 local SEO practitioners.
- Davide Proserpio and Georgios Zervas, Online Reputation Management: Estimating the Impact of Management Responses on Consumer Reviews, Marketing Science 36(5), 2017. Summarised in Harvard Business Review.
- Michael Luca, Reviews, Reputation, and Revenue: The Case of Yelp.com, Harvard Business School NOM Unit Working Paper 12-016.
- Chris Anderson, The Impact of Social Media on Lodging Performance, Cornell Hospitality Report.
- WebFX, 2026 Google Business Profile Benchmarks, for review velocity and top-three review count reference points.