Reputation

Review Management at Scale: How to Run 1,200 Locations Across 200 Businesses

A two-hundred-business portfolio takes in more reviews in a year than any team can read, let alone answer. The work is not writing replies. It is routing, policy and measurement, and that part has a known shape.

Two hundred client businesses averaging six locations each is twelve hundred locations. At a median velocity of roughly three new Google reviews per location per month, that portfolio takes in about 3,600 reviews a month and 43,200 a year, sitting on top of a historical backlog that is comfortably into six figures. Nobody reads that. The instinct is to hire, or to automate every reply and hope nobody notices. Both are the wrong move, because past a few dozen locations review management stops being a writing task and becomes a routing task.

What follows is how that routing works in ReviewMankey, the AI reputation platform we build at Dude Lemon, using the real product screens rather than mockups, and the 2026 research that says why it now matters more than it did last year.

Four mechanics make portfolio scale tractable. One normalized queue across every review source. Reply policy set per location, from drafts-only up to auto-reply at four stars and above. Priority-ordered escalation rules that triage on rating, sentiment and keywords before a person looks. And one reputation score per business, so two hundred businesses are scannable in a single view.
Every review source normalized into one queue, with sentiment and response coverage measured across the portfolio.

What “at scale” actually means, in numbers

Review management gets discussed in adjectives. It is worth doing in arithmetic, because the arithmetic is what decides whether a workflow survives the fiftieth client. Here is the portfolio in this article, with every derived figure marked as derived.

Portfolio inputValue usedWhere the figure comes from
Client businesses200The scenario in this article.
Locations per business6The scenario in this article.
Locations under management1,200Derived: 200 × 6.
New reviews per location per month3 at the median, 6.21 for restaurantsPublished 2026 Google Business Profile benchmarks. The restaurant figure is new Google reviews per unit per month.
New reviews per month, portfolio3,600 to 7,452Derived: 1,200 × 3, and 1,200 × 6.21.
New reviews per year, portfolio43,200 to 89,424Derived: monthly intake × 12.
Lifetime reviews held per location47 at the very low endBusinesses ranking in Google’s top three local results average about 47 reviews. Established multi-location brands hold multiples of that.
Total reviews under managementInto six figures within a few yearsDerived: 1,200 locations at a few hundred lifetime reviews each, plus annual intake.
Portfolio arithmetic. Substitute your own location count and velocity; the shape of the problem does not change, only how fast you hit it.

Now price the labour. Assume three minutes to read a review, decide the response, write it in the client’s voice and publish it. That is generous for a five-star review and optimistic for an angry one. At 3,600 reviews a month it is 10,800 minutes, or 180 hours, or a little over one full-time person doing nothing else. At restaurant velocity it is 373 hours, more than two people. That is before escalations, before client reporting, and before anyone logs into two hundred separate platform accounts.

The three-minute figure is an assumption, not a measurement. Replace it with your own and the conclusion holds: manual review response scales linearly with location count, and almost nothing else in an agency does.

Why the stakes went up in 2026

Two independent 2026 datasets and one practitioner survey moved the goalposts in the same direction. Consumers now read reviews more often, demand fresher ones, expect a reply within a day, and increasingly ask an AI assistant rather than a search engine. Meanwhile the average business still leaves more than half of its Google reviews unanswered.

What changedThe numberSource
Consumers who always read reviews when browsing for a business41%, up from 29% a year earlierBrightLocal, Local Consumer Review Survey 2026 (1,002 US adults, February 2026)
Consumers who will only use a business rated 4.5 stars or higher31%, up from 17%BrightLocal 2026
Consumers more likely to use a business after reading positive reviews85%. 77% are put off by negative onesBrightLocal 2026
Consumers unlikely to use a business that never replies to reviews42%BrightLocal 2026
Consumers who expect a reply the same day they post19%, up from 6%BrightLocal 2026
Consumers who expect a reply by the following day32%, up from 18%BrightLocal 2026
Consumers who used a generative AI tool for a local business recommendation45%, up from 6%BrightLocal 2026
Share of local pack ranking influence attributed to review signalsAbout 20%, up from 17% in 2023Whitespark, 2026 Local Search Ranking Factors, published November 2025 (survey of 47 local SEO practitioners)
Google reviews the average business actually answers46.9%. On Yelp it is 3.1%SOCi, 2026 Local Visibility Index (2,751 brands, about 350,000 locations)
Read the fourth row against the last one. That gap is the commercial case for review management software, and it widens with every location you add.

Forty-two percent of consumers say they are unlikely to use a business that never replies. The average business does not reply to 53% of its Google reviews. Nineteen percent of consumers now expect a reply on the day they post, up from six percent a year earlier, while one vendor benchmark built from platform data between January 2025 and February 2026 puts the average time to respond to a Google review at 2.7 days. Every one of those gaps is an operations problem, not a copywriting problem.

What a rating point is actually worth

The academic literature on this is older than the software, and it is unusually clean. Three findings are worth knowing before you argue for a budget.

FindingEffect sizeStudy
A one-star increase in Yelp rating5% to 9% more revenue, concentrated in independent businesses rather than chains with established reputationsMichael Luca, Reviews, Reputation, and Revenue: The Case of Yelp.com, Harvard Business School NOM Unit Working Paper 12-016
A one-point increase in review score on a five-point scalePrice can rise 11.2% while occupancy and market share holdChris Anderson, The Impact of Social Media on Lodging Performance, Cornell Hospitality Report
Starting to respond to reviews at all12% more reviews received, and average rating up 0.12 starsDavide Proserpio and Georgios Zervas, Online Reputation Management: Estimating the Impact of Management Responses on Consumer Reviews, Marketing Science, 2017
The third row is the one that matters operationally. Responding is not customer service overhead; it is a measured input to the rating itself.
Proserpio and Zervas is the finding that turns review response from a cost centre into a lever. Across tens of thousands of TripAdvisor hotel reviews, replying did not only appease the reviewer. It changed how many people reviewed afterwards, and what they wrote.

One inbox, and one number per business

ReviewMankey dashboard showing a reputation score of 72, total reviews 1,247, average rating 4.3 lifetime and 4.1 over 30 days, an 87 percent response rate, five open actions, a review trends chart and a rating distribution chart
One business, reduced to a score that moves and the five numbers that explain it. At two hundred businesses this is the only view that works.

The screen above is a single account: 1,247 reviews under management, 4.3 stars lifetime, 4.1 over the last thirty days, an 87% response rate, five open actions, and a reputation score of 72 that is up three points on the previous period. The period selector runs 7 days, 30 days, 90 days and 12 months, and the sync indicator is live rather than a nightly batch.

The pair worth staring at is 4.3 lifetime against 4.1 for the last thirty days. A lifetime average is a slow-moving number that hides a live problem for months, because a business with a thousand historical reviews needs an enormous run of bad ones to shift it. The thirty-day figure moves immediately. Whitespark’s 2026 report puts review recency among the highest-weighted individual review factors, which is exactly why a lifetime average is the wrong number to steer by and the right number to report.

At portfolio scale the reputation score does the triage on the dashboards themselves. You cannot read two hundred dashboards a week. You can read two hundred scores, sorted by change, and open the six that moved.

Every source, normalized into one sentiment split

ReviewMankey reviews-by-platform panel listing each connected source with its own review count and average rating, including Google Business at 4.3 across 612 reviews, the Play Store at 4.1, the App Store at 4.5 and first-party website reviews at 4.6, alongside a sentiment breakdown and an open actions list with critical, high and medium severities
Every source on one rating scale and one sentiment axis. The 0.5-star spread between the Play Store and the website is invisible in a blended average. Screenshot from a demonstration workspace.

Google Business Profile at 4.3 across 612 reviews, the Play Store 4.1 across 142, the App Store 4.5 across 121, and first-party website reviews 4.6 across 78. Normalizing the payload once at ingest, so that a rating, an author, a timestamp, a platform and a location mean the same thing regardless of where the review came from, is what allows a single queue, a single rule engine and a single metric set to work across sources that otherwise share nothing.

It also surfaces spreads that a blended average erases. There is half a star between the Play Store and the website on this account. An app-store rating problem and a storefront rating problem have different owners and different fixes, and a single blended average tells you about neither. Most platforms in this category cover storefronts or app stores, not both; ReviewMankey’s write-up of managing locations and apps in one place explains why that split exists.

Sentiment sits on its own axis, not derived from the star rating: 57.7% positive, 24.9% neutral, 11.6% negative and 5.8% very negative. That last pair is 217 reviews, and it is the only part of the 1,247 that genuinely needs a person. Separating sentiment from rating matters because a three-star review with a specific, fixable complaint is worth more attention than a one-star review with no text, and the star count cannot tell them apart.

A review is a unit of work, and it has a state

ReviewMankey reviews board with Pending, In Progress and Replied columns, review cards tagged Negative, Very Negative, Follow Up, Urgent and Service Recovery, and a detail drawer showing sentiment, rating, platform, tags and a published AI-drafted response
Pending, In Progress, Replied. Response coverage is only measurable when answered is a state rather than somebody’s memory.

Reviews move across a board: Pending, In Progress, Replied. Cards carry the sentiment the model assigned, the workflow tags a human added, the location and the date, and they filter by platform and by star rating. The drawer on the right opens the whole record: column, sentiment, response state, rating, platform, received timestamp, tags such as VIP and Positive PR, the reviewer, and the full response history.

Making state explicit is what turns response coverage from an estimate into a measurement. The 87% response rate on the dashboard is a count of records in a state, not a guess, which is the only reason it is safe to put in front of a client next to the industry average of 46.9%.

Note the Published and AI badges on the response in the drawer. If you manage reviews for other people’s brands, being able to show a client exactly which replies were model-drafted and which were hand-written is the difference between a defensible workflow and an awkward conversation.

The mechanic that makes hundreds of businesses possible: policy per location

ReviewMankey integrations screen showing connected Google, Play Store and App Store platforms, with a Google drawer listing two accounts and twelve locations, each carrying its own reply policy badge such as Auto-reply 4+, Auto-reply 5+, Custom, Drafts only or Needs verification
Two Google accounts, twelve locations, and a different reply policy on each one. Account email addresses redacted. Screenshot from a demonstration workspace.

Platforms connect at the account level: one Google account here holds eight locations, another holds four, and the Play Store and App Store connections sit beside them with their own last-sync timestamps. Open the Google drawer and you get accounts, then a searchable and paginated list of all twelve locations.

The badges on each location are the whole point. Reply policy is a property of the location, not the account, not the client, and not the workspace.

One thing to be precise about. Published replies currently go to Google Business Profile. Google Play and the Apple App Store are monitored, scored, routed and escalated exactly like Google locations, and replies to them are drafted rather than published. The policy ladder below therefore governs Google locations; everything else in this article — the normalized queue, the rules, the scoring, the reporting — spans every connected source.
PolicyWhat it doesWhen to use it
Drafts onlyEvery reply is drafted and held for a human to approve or edit. Nothing publishes on its own.A new client, a brand with a legal review step, or any location in the middle of a live service problem.
Auto-reply 5+Five-star reviews get an on-brand reply published automatically. Everything else waits for a person.The default first promotion. Five-star replies are the highest-volume, lowest-risk category in nearly every portfolio.
Auto-reply 4+Four and five-star reviews publish automatically. Three stars and below always route to a person.A location with a settled voice and a track record, where the four-star bucket is large enough to matter.
CustomPer-location voice, language and rules that differ from the account default.A location in another market or language, or a brand inside the portfolio with its own tone.
Needs verificationThe location is mapped but not confirmed, so it syncs without publishing.Mid-onboarding, or after a client changes ownership of a listing.
The policy ladder. Locations move up it as trust accrues, and back down it the moment something goes wrong, with no migration involved.

This is the difference between software that manages a business and software that manages a portfolio. If reply policy were an account-level setting, onboarding one nervous client would mean either dropping the whole portfolio to drafts-only or asking that client to trust automation on day one. Because policy lives on the location, the risk appetite of one client never constrains another, and a single underperforming branch can be pulled back to drafts-only without touching its eleven siblings.

The onboarding arithmetic follows from the same design. A Google account holding eight locations connects once and brings eight locations with it. Across a portfolio averaging six locations per client, the unit of setup work is the account rather than the location, which is roughly a sixfold reduction in connection steps. What remains is a policy decision per location, which is a judgement call measured in seconds, not an integration measured in hours. For the single-client version of this, ReviewMankey’s own guide to multi-location Google review management covers the same ground at one brand rather than two hundred.

Rules that triage before a human looks

ReviewMankey escalation rules screen showing priority-ordered rules P1 through P8 with critical, high, medium and low severities, match conditions on platform, location, star range, sentiment and keywords, resulting severities, seven-day match counts and Test Rules and View Log controls
Rules are ordered, not merely toggled. P1 is narrower than P4 and both match a one-star review, so priority decides which severity it inherits.

Rules are priority-ordered from P1 down, each individually enabled or paused, each with conditions on the left and consequences on the right. P1, urgent safety concern, matches any platform and any location at one to two stars where sentiment reads urgent or angry and the text contains unsafe, injury or hazard, and assigns critical severity. P2, refund request signal, narrows to Google Business and two specific locations at one to three stars with frustrated or disappointed sentiment and the keywords refund, charged or money back. P3 catches one-star angry reviews. P4 is the low-rating catch-all. P8, potential spam or fake review, sits paused while its thresholds are recalibrated, which is a state a rule engine needs to have.

textTwo of those rules, read off the screen above
1P1 Urgent safety concern enabled
2 when
3 platform any
4 location any
5 rating 1..2
6 sentiment urgent | angry
7 keywords unsafe, injury, hazard
8 then
9 severity critical
10 observed
11 3 matches in the last 7 days
12
13P4 Low-rating monitoring enabled
14 when
15 platform any
16 location any
17 rating 1..2
18 then
19 severity medium
20 observed
21 27 matches in the last 7 days

The match counts are the interesting column. Four enabled rules produced 49 matches in seven days on this account: 27 in the medium catch-all, 11 on refunds, 8 on one-star anger, and 3 critical. That is the real size of the human queue, and it is a list a person can work in an afternoon. The 1,247 reviews never had to be read; the rules read them.

Ordering matters because the rules overlap deliberately. A one-star review mentioning an injury matches both P1 and P4, and priority decides that it inherits critical rather than medium. This is why the interface is an ordered list rather than a flat set of switches, and why the specific rule has to sit above the catch-all.

Test Rules and View Log are the other half. A rule engine you cannot dry-run against history is a rule engine nobody will touch, which means thresholds calcify at whatever they were on day one. Being able to test a proposed rule against back history before enabling it, and then read a log of what actually fired, is what lets one operator tune escalation for two hundred different businesses without breaking any of them.

Writing the rules themselves is a separate exercise, and the severity tiers, owners and SLAs behind them are worth getting right before you enable anything. The review escalation matrix playbook on the ReviewMankey blog is the long version.

A 4.5 means nothing without the market around it

ReviewMankey competitor analysis showing a 4.5 rating, market rank 3 of 6, plus 0.2 against the average competitor, 847 reviews, a bar chart ranking review volume by competitor and a review share donut where the leader holds 35 percent
Rating above the market average, ranked third of six, on 14% review share against a 35% leader. That is a volume problem, not a rating problem.

This location holds 4.5 stars, which is 0.2 above the average competitor and trending up 3%, on 847 reviews trending up 8%. It still ranks third of six, and the review-share donut explains why: the leader holds 35% of review volume in the market, this location holds 14%. The comparison refreshes every six hours.

Rating and volume are different problems with different interventions. Rating is a service-and-response problem measured in quarters. Volume is a review-request problem measured in weeks. A per-location market rank across 1,200 locations tells you which of them is fighting for stars and which is fighting for share, and stops you prescribing the same campaign to both.

Reviews are now the input to AI recommendations, not just search

The most consequential 2026 dataset for anyone running a review programme is SOCi’s Local Visibility Index, which covered 2,751 brands and roughly 350,000 locations. Brands appear in Google’s local three-pack 35.9% of the time on average. Only 1.2% of locations are recommended by ChatGPT, and 11.0% by Gemini. Recommended locations average 4.3 stars on ChatGPT, 4.2 on Perplexity and 3.9 on Gemini.

Those averages describe a filter rather than a ranking factor. Fall below roughly four stars and many AI systems stop considering a location at all, no matter what else is true about it. Above the threshold, volume and recency start deciding. And the audience is no longer marginal: 45% of US consumers told BrightLocal they had used a generative AI tool for a local business recommendation, up from 6% the year before.

Traditional local visibility does not transfer. On SOCi’s numbers a brand is roughly thirty times more likely to make Google’s local pack than to be named by ChatGPT, and only about half the brands winning in the local pack appear in AI recommendations at all. A review programme is one of the few levers that moves both.

There is a second-order effect worth naming. When an assistant answers a question about a local business, it reads what is publicly written, and that includes your replies. Reply text stops being customer service and becomes indexable copy about how the business handles problems. A generic template repeated four hundred times reads exactly as what it is. If you want the reviews you earn to also be visible to search engines and AI assistants, that is the job SearchLift does, and our guide to generative engine optimization covers the mechanics.

The operating model, week one to steady state

  • Week one, connect at the account level. One platform account brings every location it owns. Do the same for the app stores and the review sites the client actually cares about, then let history backfill before you change anything. You are buying a baseline, not a starting point.
  • Week one, set every location to drafts only. Nothing publishes automatically on a portfolio nobody has read yet. This is not caution theatre. It is how you learn a client’s voice from their real reviews rather than from a brand guideline document.
  • Week two, write the rules once and reuse them. A safety rule, a refund rule, a one-star rule and a low-rating catch-all cover most of what matters. Dry-run each against back history before enabling it, and read the match counts honestly: a rule that fires two hundred times a week is not an escalation rule, it is a report.
  • Week three, promote the easy locations to auto-reply at five stars. Five-star replies are high volume and near-zero risk, and they are where response coverage moves fastest. This single step usually closes most of the gap to a defensible response rate.
  • Month two, promote settled locations to auto-reply at four stars and above. Leave three stars and below with a person permanently. The three-star review is where the actual information lives, and it is the one a template will always get wrong.
  • Steady state, work the exception queue and the outliers. One score per business, sorted by change. Open the ones that moved, read the critical escalations, and leave the rest running. If your week is spent inside individual reviews at this stage, the rules are mistuned.

What to measure, and what each number is for

MetricWhat it tells youReference point
Response coverageThe share of reviews that received a reply. The single number both consumers and ranking systems react to.The average business answers 46.9% of its Google reviews and 3.1% of its Yelp reviews (SOCi 2026).
Median time to first responseWhether you are meeting the expectation, rather than merely replying eventually.19% of consumers expect a reply the same day and 32% by the next day (BrightLocal 2026).
Rating over 30 days against lifetimeWhether the recent cohort is better or worse than your history. A lifetime average hides a live problem for months.4.3 lifetime with a 4.1 thirty-day, as on the dashboard above, is a warning and not a good quarter.
Sentiment mixThe size of the queue that genuinely needs a human, measured independently of the star rating.Negative plus very negative was 17.4% of reviews, or 217 of 1,247, on the account shown above.
Escalation matches by severityThe true size of the human workload, and whether your thresholds are calibrated.Four enabled rules produced 49 matches in seven days above, of which 3 were critical.
Review velocity per locationWhether the review-request programme works, and whether a location looks alive to ranking and AI systems.A median business collects about three new Google reviews a month.
Market rank and review shareWhether a rating problem or a volume problem is holding a location back.4.5 stars, third of six, on 14% share against a 35% leader, is a volume problem.
Seven numbers, each with a different intervention attached. The first two are fixable this week. The last two take a quarter.

Where configuration ends and engineering starts

Most of what is above is configuration, and that is the point: a portfolio should not need a build. Four things reliably do. Routing critical escalations into a ticketing or on-call system a client already runs, rather than into email. Pushing review sentiment into a warehouse so it can sit beside sales and staffing data, which is where the causal questions actually get answered. Per-brand voice in a group that owns several distinct brands. And reconciling a client’s own location hierarchy with the messier one the platforms hold, which is the single most common reason a location count does not tie out.

That work is what we build, on the same platform rather than beside it.

Frequently asked questions

How many reviews can ReviewMankey manage?

There is no per-account review ceiling you would meet in practice. The binding constraint at portfolio scale is human attention, which is why reply policy sits on the location and escalation rules sit in priority order. A 1,200-location portfolio taking in roughly 3,600 new reviews a month is a normal shape for the product, and the queue that actually reaches an operator is measured in tens of items a week rather than thousands.

Can one team manage reviews for hundreds of businesses?

Yes, provided three things are true. Reviews arrive in one normalized queue rather than across a hundred separate platform logins. Reply policy is set per location, so a cautious new client and a mature automated one coexist without either compromising. And rules triage on rating, sentiment and keywords before a person looks. Without those three, review response scales linearly with location count, and you are not running a system, you are hiring.

Does ReviewMankey publish AI replies automatically?

Only where you allow it, and only per location. The ladder runs from drafts only, where every reply is held for human approval, through auto-reply at five stars and above, to auto-reply at four stars and above. Three stars and below can be reserved for a person permanently. Published responses are labelled, so a model-drafted reply is distinguishable from a hand-written one after the fact. Automatic publishing applies to Google Business Profile locations; app-store replies are drafted rather than published.

Which review sources does ReviewMankey monitor?

Google Business Profile, Google Play, the Apple App Store and first-party website reviews. Everything is normalized at ingest so that ratings, authors, timestamps, platforms and locations mean the same thing regardless of origin, which is what lets one queue, one rule engine and one metric set span sources that otherwise share nothing. Replies publish back to Google Business Profile; the other sources are monitored, scored and escalated, with replies drafted for you.

Does responding to reviews actually improve your rating?

The best available evidence says yes. Proserpio and Zervas, publishing in Marketing Science in 2017, examined tens of thousands of TripAdvisor hotel reviews and found that when hotels began responding they received 12% more reviews and their ratings rose by an average of 0.12 stars. Consumer research points the same way: BrightLocal’s 2026 survey found 42% of consumers are unlikely to use a business that never replies.

How much is one star worth in revenue?

Two well-known studies put numbers on it. Michael Luca, in a Harvard Business School working paper on Yelp, found a one-star rating increase associated with a 5% to 9% revenue increase, concentrated in independent businesses rather than chains with established reputations. Cornell research by Chris Anderson found that a hotel raising its review score by one point on a five-point scale could raise price by 11.2% while holding occupancy and market share.

How do reviews affect whether AI assistants recommend a business?

Heavily, and differently from Google. SOCi’s 2026 Local Visibility Index, covering 2,751 brands and about 350,000 locations, found brands appear in Google’s local three-pack 35.9% of the time on average, but only 1.2% of locations are recommended by ChatGPT and 11.0% by Gemini. Recommended locations average 4.3 stars on ChatGPT, 4.2 on Perplexity and 3.9 on Gemini, which makes rating behave less like a ranking factor and more like a threshold. Meanwhile 45% of US consumers told BrightLocal they had used a generative AI tool for a local business recommendation, up from 6% the year before.

How long does it take to onboard a multi-location client?

Connecting is account-level, so a platform account holding eight locations arrives in one flow rather than eight. The work that takes real time is not connection. It is deciding the reply policy for each location and dry-running the escalation rules against back history before enabling them. A six-location client is a single working session, and the rules you write for the first client are reusable for the next hundred.

What does ReviewMankey cost?

There is a free plan, and paid plans start at 8.99 per month. Current pricing and plan limits are on the ReviewMankey product page.

The short version

None of this makes review management easy at scale. It makes it bounded. Once a review is a unit of work with an explicit state, once reply policy lives on the location rather than the account, and once ordered rules do the triage, adding the two hundred and first business changes the numbers on a dashboard and nothing else about the job. That is the whole design brief.

ReviewMankey is the platform. If you want the reviews you earn to be visible to search engines and AI assistants as well as to customers, SearchLift is the other half of the problem. And if your portfolio needs review data routed into systems you already run, talk to our team.

Sources

Need help building this?

Let our team build it for you.

Dude Lemon builds production-grade web apps, APIs, and cloud infrastructure. Get a free consultation and project proposal within 48 hours.

Start a Project

Related articles

View all articles →