Published by Rims & Tires
Compiled & checked by the site's publisher from manufacturer, NHTSA & USTMA data · Updated 2026-08-07
A review with at least one photo attached is nearly twice as likely to be negative as one without — 20.0% one- or two-star, versus 11.7% for photo-free reviews, in the same 5,716-review West Georgia auto-shop corpus behind our first study (n=275 vs 5,441, z=4.11, p=0.00004). This time we also read the other half of the conversation: when a shop replies to a bad review, what does it actually say? We hand-graded all 249 owner replies to the corpus's negative reviews, reply by reply, against a fixed set of category definitions — not a keyword classifier's guess. 61.0% apologize. Nearly a third — 32.9% — dispute the customer's account, blame them, or defend the shop's conduct instead of fixing anything, which is now the single most striking number on this page. Only about 1 in 5 offers anything you could call a concrete fix, and even that overstates it: the single largest share of those routes the customer to a shared corporate 800-number rather than a dedicated way to reach the shop. This page carries a dated correction as of August 7, 2026, recording two separate revisions made the same day — see the notice immediately below.
Correction — published 2026-08-05, corrected twice on 2026-08-07
The version of this study published on 2026-08-05 undercounted how often a shop offers a concrete way to fix the problem. The bug: our classifier's remedy-detection pattern only recognized the literal phrase "call us at [phone number]" — it missed replies phrased as "reach out to us at [phone number]" or "contact us at [phone number]," which turned out to be the more common way shops actually write it. That's a bug in our code, not a change in the underlying reviews, and we found it ourselves in a routine post-publication audit of this page.
The concrete-remedy rate we published that same morning was 18.9% (47 of 249 replies) — a same-day sweep for the specific phrasing the bug had missed, not a full re-grade. That interim figure has since been superseded; see below. We also found that our disclosed classifier-accuracy spot-check — originally reported as 94.4% (151/160) — could not actually be reproduced, because the human-graded ground truth behind it had never been saved. An independent, blind re-grade of that same 40-reply sample found 83.75% (134/160) correct: apology 39/40 (97.5%), concrete remedy 33/40 (82.5%), generic/template 30/40 (75.0%), deflection 32/40 (80.0%). The claim that apology, concrete-remedy, and generic-template tagging were "highly reliable (95–100% accuracy)" was false and was removed from this page that morning.
That 40-reply spot check was always a sample — it could bound the classifier's error rate but not correct the study's actual published percentages, which describe all 249 replies, not 40. So later the same day, every one of the 249 owner replies was individually hand-graded — read in full, blind to the classifier's own label, against the same four category definitions — producing real ground truth rather than an estimate. That changed the headline numbers again, this time for all four categories, not just remedy. Apology: 66.7% as published, unchanged by the morning's fix, now 61.0% (152/249) in the final hand-grade — lower, not higher, because "sorry you feel that way" and "sorry, but this wasn't our fault" constructions recur throughout the corpus and the study's own definition excludes them; a measurable share of apparent apologies turn out not to be apologies. Concrete remedy: 5.6% as published, 18.9% after the morning's fix, now 20.1% (50/249) — the morning's sweep caught most of the regex bug's damage but not quite all of it. Generic/template: 54.2% as published and unchanged that morning, now 57.4% (143/249) — full-corpus grading turned up direct proof of verbatim templating the 40-reply sample couldn't see, including a Discount Tire line about "the extended wait and the oversight regarding your appointment" pasted onto two to four different reviewers' different complaints. Deflection: 8.4% as published and still described that morning as "at least 8.4%, full re-grade in progress," now 32.9% (82/249) — nearly one reply in three disputes the customer's account, blames them, or justifies the failure instead of fixing it. That is the single most newsworthy number on this page, and the morning's caveat undersold it by roughly 4x.
The apologies-with-no-remedy figure moved the same way, each time downward, as the remedy denominator kept growing: 94.0% as originally published, 77.7% after the morning's fix, now 73.7% (112 of 152 apologetic replies) in the final hand-grade.
The full hand-grade also let us measure classifier accuracy against all 996 category judgments across the whole corpus, not just the 160 from the 40-reply sample: 815/996 correct (81.83%) — apology 94.4%, remedy 79.1%, generic 78.3%, deflection 75.5%. That's close to, and consistent with, the 40-sample's 83.75% estimate — the spot check held up reasonably well as an early read, it just wasn't precise enough to publish as the study's final figures. 27 of the 249 hand-graded replies were flagged ambiguous, with a reason recorded for each; we're publishing that count because a study that discloses where the calls were genuinely close is more trustworthy than one that hides it.
The remedy sub-type breakdown, now measured across all 50 hand-graded remedy=True replies rather than an estimate: 21 (42.0%) are a shared corporate or toll-free hotline serving every location a chain has, not the specific shop the customer visited; 14 (28.0%) are a shop-specific phone number; and 15 (30.0%) are a non-phone remedy — a refund, a redo, an honored warranty, or a named person actually reachable. So the honest framing isn't "shops almost never offer a fix" — it's that most replies acknowledge without fixing anything, nearly a third actively push back on the customer instead, and of the roughly fifth that do offer something concrete, the single largest slice (42%) is a corporate hotline, bigger than either shop-level remedy type on its own, even though a combined majority of remedies (58%) still connect back to the shop in some form.
We found the original bug ourselves, corrected it the same morning, then found that correction was itself incomplete and replaced it with a full hand-grade of all 249 replies before the day was out. We're disclosing both revisions rather than quietly landing on the final numbers — a study that shows its own convergence toward ground truth is more credible, not less. Every figure below and in the table on this page reflects the final, 2026-08-07 hand-graded numbers.
How is this different from the first study?
Same corpus, new questions. Our first study analyzed 5,716 Google reviews from 68 West Georgia auto and tire shops, collected August 1, 2026, and asked what people complain about and praise. This one reuses that exact dataset — same 68 shops, same 5,716 reviews, same collection date — to ask two questions the first study didn't: what structural signals (not review content) predict whether a review will be negative, and what do shop owners actually write when they reply to one.
The photo, first-reviewer, and Local Guide findings below come from the same reviews-clean.json file the first study published from — every review carries a photo count, the reviewer's prior review count, and a Local Guide flag alongside its star rating. The owner-reply-content analysis needed one more field the first study never touched: the actual text of the owner's reply, not just whether one existed. We pulled that from our original unfiltered scrape (reviews-raw.json, 7,582 records before geographic and category screening) and reran the exact same inclusion rules the first study documents — auto-service category, five-county footprint, not permanently closed, at least 10 reviews sampled — to rebuild the identical 5,716-review corpus, this time keeping the reply text attached. We verified the reconstruction line for line: it produces exactly 5,716 reviews, matching the published corpus count, before any of the numbers below were computed.
That reconstruction turned up 2,635 reviews with an owner reply attached, 249 of them replies to a negative (1- or 2-star) review — 36.0% of the corpus's 692 negative reviews got a reply, roughly in line with the first study's 46.1% reply rate across all star levels. We classified the text of those 249 replies into four non-exclusive categories — apology, deflection or blame-shifting, a concrete remedy offered, and generic template or corporate-escalation language — using keyword and structural pattern matching, the same method the first study used for complaint and praise themes. A reply can carry more than one tag (an apology that also deflects, for instance), so the percentages below don't sum to 100.
We spot-checked the classifier by hand three times against random samples of 40 replies each, fixing two genuine bugs it found along the way (detailed in Methodology below) before drawing a fourth, untouched sample to measure accuracy. We originally reported 151 of 160 individual category judgments correct (94.4%) on that sample — that number turned out to be wrong and unreproducible. An independent, blind re-grade of the same 40 replies found 134 of 160 correct (83.75%) instead, and a full hand-grade of all 249 replies against all four categories — 996 judgments in total — found 815 correct (81.83%), consistent with the 40-sample estimate (see the correction above). That full hand-grade also replaced every category percentage this study reports with real ground truth rather than classifier output; deflection in particular moved from a published 8.4% to a hand-graded 32.9%, the largest revision on this page.
West Georgia Local
Find the right shop near you
Tell us your vehicle and what you need — we'll match you with a vetted West Georgia shop. Free, no obligation.
The strongest signal: reviews with a photo attached skew negative
Of the 275 reviews in the corpus with at least one photo attached, 20.0% are one or two stars. Of the 5,441 reviews with no photo, only 11.7% are. That's not a small gap — it's the strongest, best-powered result in either of our two studies (z=4.11, p=0.00004, roughly a 1-in-25,000 chance of this pattern by chance).
The mechanism isn't mysterious once you think about it. A happy customer rarely photographs their oil change. A customer photographs a stripped lug nut, a scuffed rim, a leak under the car, or a receipt with a number on it they want to prove. The camera comes out when someone wants to document evidence, and evidence gets attached to complaints far more often than to praise.
The practical read for a driver: a review with a photo is worth reading closely even if the star count looks unremarkable, because photo-attached reviews are disproportionately the ones with something specific and checkable behind them. For a shop owner, it's a cheap early-warning signal — a customer who's taking pictures mid-visit is a customer who may already be building a case.
A finding we're retracting: first-time reviewers did not, in fact, rate more generously
An earlier version of this study reported that reviewers leaving their very first Google review (reviewerNumberOfReviews = 0 at the time of posting) gave a negative rating only 6.8% of the time, versus 12.3% for everyone else — a seemingly real, statistically detectable gap (n=191, z=-2.28, p=0.022). That comparison reproduces exactly on request. It just doesn't hold up once it's tested honestly.
We ran that same two-proportion test alongside 35 other comparisons mined from this corpus — chain-vs-independent splits, complaint themes, reply behavior, photo counts, and more — and corrected for testing that many hypotheses at once, the way any of them could turn up a false positive on its own. At Bonferroni's threshold for 36 tests (p ≤ 0.0014), the first-time-reviewer result isn't close. Under the more forgiving Benjamini-Hochberg false discovery rate, six of the 36 comparisons clear their threshold; ranked 9th by p-value, this one still isn't among them.
A real effect should also show a dose-response — more prior reviews, a steadily shifting negative rate. It doesn't. Split into five equal-sized bands by prior review count, the 1-2-star share runs 11.9% (0-2 prior reviews), 11.5% (2-6), 12.8% (6-14), 13.7% (14-39), and 10.6% (39-1,040) — no trend, just noise, with the lowest band sitting below the two bands above it rather than the pattern descending in order. Treating prior-review-count as a continuous number instead of a binary split turns up nothing at all: Spearman rho = -0.004, p = 0.74. The only place a gap shows up is the exact zero-prior-reviews cutoff — the smallest, noisiest slice of the whole distribution. That's the signature of a sampling fluctuation, not a real effect.
We're retracting this one rather than quietly deleting it. The raw p-value looked publishable the first time; a fuller multiple-comparison test and a look at the dose-response say it isn't. A study that shows its failed tests, not only the ones that worked, is the one worth trusting.
A myth we can debunk: Local Guide status predicts nothing
Google's "Local Guide" badge — earned through a points system for leaving reviews, photos, and edits — is sometimes assumed to mark more serious or more critical reviewers. In this corpus, it predicts nothing. Local Guides left a negative review 12.0% of the time; everyone else, 12.2%. That's a dead heat (z=-0.26, p=0.80) — not a trend too small to see, an actual null result.
We're publishing the null because a clean null, honestly reported, is worth more than pretending every variable we tested turned up something. If you're a shop owner deciding whether to treat a Local Guide review differently, this corpus gives you no reason to.
What do West Georgia auto shops actually say when they reply to a bad review?
We hand-graded the text of all 249 owner replies attached to the corpus's negative reviews against four category definitions, reading each reply before consulting any automated label. 61.0% (152 of 249) contain an apology — some form of "sorry," "apologize," or explicit regret, applied strictly: "sorry you feel that way" and "sorry, but this wasn't our fault" constructions don't count, because they accept no responsibility, and they recur throughout the corpus. 20.1% (50 of 249) offer anything we classify as a concrete remedy — a refund, a redo, a waived fee, a specific named contact, an honored warranty, or a dialable phone number, as opposed to a vague invitation to "reach out" with no way to actually reach anyone. 57.4% (143 of 249) read as generic or template: boilerplate language, a shared corporate escalation email or case-management link (Discount Tire, Pep Boys, Tires Plus and several others route negative reviews to a central inbox rather than the shop itself), or — now directly provable at full-corpus scale — a "specific-sounding" clause pasted verbatim onto two to four different reviewers' different complaints, such as a Discount Tire line about "the extended wait and the oversight regarding your appointment" reused across unrelated customers. And 32.9% (82 of 249) — nearly one reply in three — contain identifiable deflection: disputing the customer's account, blaming them or a third party, or defending the shop's conduct at length instead of fixing anything.
Put the four numbers together and the real shape of a bad-review reply looks like this: most acknowledge without fixing anything, close to a third actively push back on the customer instead, and of the roughly 1 in 5 that do offer something concrete, the single largest slice routes to a corporate line rather than the shop. Of the 152 replies that apologize, 112 (73.7%) do it without attaching anything you could call a fix. Of the 50 replies that do offer a remedy, 21 (42.0%) are a shared corporate 800-number or national call-center line — the same number the chain routes every customer in this corpus to, not a way to reach the specific shop the customer actually visited; 14 (28.0%) are a shop-specific phone number; and 15 (30.0%) are a non-phone remedy — a refund, a redo, an honored warranty, or a named person actually reachable. Combining the shop-specific phone and non-phone categories, 29 of all 249 replies (11.6%) point back to the shop itself in some concrete way — a slim majority of the 50 remedies, but still just over a tenth of all 249 replies. "We're sorry to hear about your experience, please reach out" is still the modal reply — and now that the corpus has been read in full rather than sampled, so is a reply that disputes the customer was ever right to complain in the first place.
Deflection isn't mutually exclusive with apology or remedy — a reply can, and often does, apologize for the inconvenience while simultaneously disputing the customer's account, defending a price, or blaming a supplier, and both were coded independently per reply. That overlap is part of why the numbers above don't sum to 100%. The deflection rate is the single largest revision this study has made: 8.4% as originally published, and still described the same morning as "at least 8.4%, full re-grade in progress," is now 32.9% (82 of 249) after every reply was hand-graded — roughly four times the original figure, and the strongest confirmation yet that the classifier's early spot-check finding (every mismatch ran in the direction of missing real deflection, never the reverse) held at full scale, not just in an 8-mismatch sample.
When the apology was built for a five-star review
One pattern is worth calling out on its own because it's unambiguous and easy to verify: a positive, five-star-shaped thank-you template, independently confirmed as reused three or more times elsewhere in the corpus, firing on a review that was actually one or two stars. We found five confirmed instances of this across three shops. Two of them literally thank the customer for "the five-star review" on a review that was, in fact, one star — a Tires Plus reply ("Taking the time to leave a five-star review really means a lot to us... Thanks for choosing us for your complete auto care") and a Firestone Complete Auto Care reply ("Hi, thank you so much for the five-star review!"), both attached to 1-star reviews. A second Tires Plus template ("thanks for the great review...") and two separate Mavis Tires & Brakes replies ("Thank you so much for the kind words... we look forward to seeing you again and providing another incredible experience!") show the same pattern — a canned positive response, verified as boilerplate by its reuse elsewhere, auto-applied to a review the shop apparently never actually read.
We should note a discrepancy here rather than round up to a nicer number: we were able to confirm exactly two cases of the literal "thank you for the five-star review" phrasing landing on an actual 1-star review, not three or more. Broadening the same objective test — a template independently verified as reused three-plus times elsewhere, misapplied to a 1- or 2-star review — to any positive-toned boilerplate (not only ones using the literal words "five-star") gets to five confirmed instances across three shops, which is the number we're standing behind. Either way, the pattern is real: some fraction of negative reviews get an auto-reply that was never customized for star rating, let alone content.
What should you do with this — as a driver, or as a shop owner?
As a driver: if you're comparing two shops with similar star ratings, read a couple of the photo-attached reviews specifically — they're the ones most likely to carry a documented, checkable complaint rather than a vague vent. And if you're deciding whether an owner reply signals a shop that takes feedback seriously, look past whether they replied to what the reply actually says. An apology with a name, a phone extension, or a specific next step from the shop itself is a different thing from "we're sorry, please email our customer care inbox" — and in this sample, a reply that actually points back to the shop rather than a shared corporate line happens in just 11.6% of cases, 29 of 249. Watch, too, for a reply that argues with you instead of fixing anything — nearly a third of replies in this corpus do.
As a shop owner: the cheapest fix implied by this data doesn't cost anything. If you're already writing a reply, spend the extra thirty seconds adding one concrete detail — a name, a direct number, a specific offer — rather than sending the same apology template you used for the last complaint. And check that your reply platform isn't misfiring a five-star thank-you onto a one-star review; it happened at three different shops in this sample, and it's the kind of thing a customer screenshots.
Methodology
Source: the identical 5,716-review, 68-shop West Georgia corpus documented in our first study — collected via an automated Google Maps reviews scraper (Apify) on August 1, 2026, screened to the auto-service category and the five-county West Georgia footprint (Carroll, Douglas, Paulding, Haralson, Heard), excluding permanently-closed listings and shops under 10 sampled reviews. Every figure in this study was independently recomputed from that corpus in the original session (August 5, 2026), not copied from the first study's output; a classifier bug was found and partially corrected the morning of August 7, 2026, then all four owner-reply category figures were independently re-audited and corrected again the same day, once every one of the 249 replies had been individually hand-graded — see the correction above.
Photo and Local Guide tests: two-proportion z-tests on reviews-clean.json's own fields (images ≥ 1 vs 0; localGuide true vs false), the same test the first study used for its chain-vs-independent comparisons. The first-time-reviewer comparison (reviewerReviewCount = 0 vs ≥ 1) used the same two-proportion test but is reported above as retracted: tested alongside 35 other comparisons mined from this corpus, it fails both Bonferroni and Benjamini-Hochberg correction, and shows no dose-response across five prior-review-count bands or as a continuous Spearman correlation (rho = -0.004, p = 0.74). See the retraction above for the full accounting.
Owner-reply-content classification: the 5,716-review corpus was independently reconstructed from the original unfiltered scrape (reviews-raw.json) by re-applying the first study's own published inclusion rules in code, this time retaining each review's owner-reply text — a field the first study never used. The reconstruction was verified to produce exactly 5,716 reviews before any classification ran. Replies were first tagged for apology, deflection, concrete remedy, and generic/template language using keyword and structural regular expressions (an approach directly modeled on the first study's complaint- and praise-theme tagging), plus two objective structural signals for "generic": a reply body matching another reply's body elsewhere in the corpus after masking the reviewer's own name, and a phone number or corporate email/case-link reused across multiple different customers at the same shop. That keyword classifier produced this study's first-published numbers and is still reported below and in the table above for comparison and to measure its own error rate — but every category percentage this study now reports as its result comes from the full hand-grade described further down, not from the classifier.
The classifier was corrected twice during initial development. A hand-check of an initial 40-reply sample found a bug where the word "discount" was matching inside the literal text of Discount Tire's own case-link URL rather than genuine discount language, and found that the duplicate-detection missed templates that embed the customer's name mid-sentence rather than as a leading salutation. Both were fixed in code — masking the reviewer's actual name (carried on every raw record) before duplicate detection, and adding a trailing word boundary to the discount pattern — and verified against a second, different 40-reply sample. A third, previously untouched sample of 40 replies (not used to find or fix anything) was then hand-graded against the corrected classifier's output, and we originally reported 151 of 160 individual category judgments correct (94.4% overall).
That 94.4% figure was wrong. We found out why on 2026-08-07 during a routine post-publication audit: the human-graded ground truth behind it had never been saved, only the classifier's own auto-labels, so the original number couldn't actually be reproduced or verified. We rebuilt that same third 40-reply sample and independently, blindly re-graded it by hand against all four categories (160 judgments total). The corrected result is 134 of 160 correct (83.75%). By category: apology 39/40 (97.5%), concrete remedy 33/40 (82.5%), generic/template 30/40 (75.0%), deflection/blame-shifting 32/40 (80.0%). That 83.75% figure was itself still a sample-based estimate — see the next paragraph for the full-corpus figure that supersedes it.
That re-grade also surfaced the bug behind the correction at the top of this page: the concrete-remedy pattern only matched replies phrased as "call (me|us) ... at [phone number]" and missed the equally common "reach out to us at [phone number]" and "contact us at [phone number]" phrasing. We swept all 249 replies — not just the 40-reply spot-check sample — for any reply the classifier had scored as no-remedy that nonetheless contained a phone number, email address, or contact-routing instruction, and reclassified under the same definition this study already published. That sweep produced the interim 18.9% (47/249) figure reported the morning of 2026-08-07 — but it was still a targeted patch for one known bug pattern, not a full re-grade.
Later the same day, every one of the 249 replies was individually hand-graded against all four category definitions, reading each reply's full text before consulting any classifier label — 996 total judgments, fully reproducible via analysis/full_regrade.py in the study's published source-data folder. That pass found: apology 61.0% (152/249, down from 66.7%, because "sorry you feel that way" and fault-denying "sorry, but..." constructions don't meet the study's own apology definition and recur throughout the corpus); concrete remedy 20.1% (50/249, up from the interim 18.9%); generic/template 57.4% (143/249, up from 54.2%, once verbatim-reused clauses across different reviewers' different complaints could be directly observed at full-corpus scale); and deflection 32.9% (82/249, up from a published 8.4% and an interim "at least 8.4%, re-grade in progress" — nearly four times the original figure). Classifier accuracy against that same ground truth, measured across all 996 judgments rather than 160, came out to 815/996 (81.83%) — apology 94.4%, remedy 79.1%, generic 78.3%, deflection 75.5% — close to, and consistent with, the 40-sample's 83.75% estimate. 27 of the 249 hand-graded replies were flagged ambiguous, each with a recorded reason, in the published ground-truth file.
Limitations
This corpus captures each shop's newest 100 sampled reviews, not its full review history — the first study discloses this cap and its direction (this sampled corpus runs about a tenth of a star below these same shops' full lifetime Google rating). Every finding in this study inherits that same limitation: photo attachment, reviewer history, and Local Guide status are all measured only within that recent-review window, not across each shop's full history.
The category percentages this study reports are a hand-read judgment call on every one of the 249 replies, not classifier output — each reply was read in full and graded against a fixed definition before any automated label was consulted (see Methodology). The keyword classifier remains imperfect and is reported for comparison and as a measurement of its own error rate: 815 of 996 category judgments correct (81.83%) across the full corpus — apology 94.4%, remedy 79.1%, generic 78.3%, deflection 75.5% — meaning roughly 1 in 5 of the classifier's individual category calls would have been wrong had we published its output instead of hand-grading. 27 of the 249 replies (10.8%) were flagged ambiguous during hand-grading, each with a recorded reason; those are included in the published rates as graded, not excluded.
We deliberately did not run seasonality, year-over-year trend, or county-level analysis in this study. The first study already tested seasonality and found no interpretable pattern, confounded by the same recent-review sampling cap; a year-over-year trend built on that same cap would be measuring the sampling method, not a real trend; and county-level analysis maps almost one-to-one onto the first study's already-published city rankings. Publishing any of the three here would mean re-presenting a result our own first study already flagged as unsupportable, or a repackaging of an existing result as if it were new.
As with the first study: a review corpus measures what people felt strongly enough to write down, not shop quality directly, and RimsAndTires is not affiliated with any shop named here — see the first study's methodology for the full disclosure. Anyone who wants the underlying counts or spots an error can email hello@rimsandtires.net.
What West Georgia owner replies to negative (1-2 star) reviews actually contain — final hand-graded ground truth (n=249 replies; categories are not mutually exclusive; West Georgia, collected 2026-08-01, reclassified 2026-08-05, corrected twice on 2026-08-07).
| Reply content | Share of 249 negative-review replies (hand-graded) | Classifier accuracy vs. hand-graded ground truth |
|---|---|---|
| Contains an apology | 61.0% (152) — down from a published 66.7% | 94.4% (235/249) |
| Generic template or corporate-escalation language | 57.4% (143) — up from a published 54.2%; includes clauses proven pasted verbatim across different reviewers' complaints | 78.3% (195/249) |
| Contains identifiable deflection or blame-shifting | 32.9% (82) — up from a published 8.4%, nearly 4x higher | 75.5% (188/249) |
| Offers a concrete remedy (refund, redo, named contact, honored warranty, or phone number) | 20.1% (50), corrected from a published 5.6% (14) and an interim 18.9% (47) | 79.1% (197/249) |
| — of those 50 remedies: a shared corporate 800-number or call-center line, not the shop itself | 42.0% of remedies (21) — 8.4% of all 249 replies | — |
| — of those 50 remedies: a shop-specific phone number | 28.0% of remedies (14) — 5.6% of all 249 replies | — |
| — of those 50 remedies: a non-phone remedy (refund, redo, honored warranty, or a reachable named person) | 30.0% of remedies (15) — 6.0% of all 249 replies | — |
| Apologizes with no concrete remedy attached | 73.7% of apologetic replies (112/152), corrected from a published 94.0% and an interim 77.7% | — |
Frequently asked
Was this study corrected?
Yes — twice, both on 2026-08-07. First, a bug in our owner-reply classifier undercounted how often a shop offers a concrete remedy: it only recognized "call us at [phone number]," not "reach out to us at [phone number]" or "contact us at [phone number]," which turned out to be the more common phrasing. A same-morning sweep put the corrected rate at 18.9% (47 of 249). That fix was itself incomplete. Later the same day, every one of the 249 owner replies was individually hand-graded against all four category definitions — not just remedy — producing this study's definitive figures: apology 61.0% (down from a published 66.7%), concrete remedy 20.1% (up from a published 5.6%, and slightly above the morning's 18.9%), generic/template 57.4% (up from 54.2%), and deflection 32.9% (up from a published 8.4% — nearly four times higher, and now the single most notable finding on this page). We're disclosing both revisions rather than quietly landing on the final numbers; full details are in the correction notice at the top of this page and in Methodology below.
Do reviews with photos attached tend to be more negative?
Yes, clearly. In our 5,716-review West Georgia corpus, 20.0% of reviews with at least one photo attached are one or two stars, versus 11.7% of reviews with no photo (n=275 vs 5,441, z=4.11, p=0.00004) — the strongest result across either of our two studies. The likely reason: customers photograph specific, checkable problems (damage, a receipt, a leak) far more often than they photograph a routine good visit.
Are first-time Google reviewers more likely to leave a good review?
No — we initially reported that finding and are retracting it. The raw comparison looks real on its own (6.8% negative among first-time reviewers, n=191, vs 12.3% for everyone else, z=-2.28, p=0.022), but it fails correction once tested honestly alongside 35 other comparisons run on this corpus, and there's no dose-response: split into five bands by prior review count, the negative-review share runs 11.9%, 11.5%, 12.8%, 13.7%, 10.6% — flat, not declining. Treated as a continuous variable it's a clean null (Spearman rho = -0.004, p = 0.74). We're leaving the full test on the page because a study that shows its failed comparisons is more trustworthy than one that only shows the ones that worked.
Does a reviewer's Local Guide badge mean their review is more reliable or critical?
No — in our data it predicts nothing. Local Guides left a negative review 12.0% of the time versus 12.2% for everyone else, a statistical dead heat (z=-0.26, p=0.80).
When a West Georgia auto shop replies to a bad review, does it actually offer to fix anything?
Not often, and when it does, the single largest share of those replies routes to a call center rather than the shop. Of 249 hand-graded owner replies to negative reviews in our corpus, 61.0% contain an apology and 20.1% (50) offer anything we'd classify as a concrete remedy — a refund, a redo, a named contact, an honored warranty, or a phone number (final hand-graded figures as of 2026-08-07 — see the correction notice on this page). Of the replies that do apologize, 73.7% do it without attaching a fix. Of the 50 replies that do offer a remedy, 21 (42.0%) are a shared corporate 800-number or national call-center line, not a contact for that specific shop; 14 (28.0%) are a shop-specific phone number; and 15 (30.0%) are a non-phone remedy like a refund or an honored warranty. The modal reply is still some version of "sorry, please reach out," often generic or templated (57.4% of all replies).
How often do West Georgia auto shops push back on the customer instead of fixing anything?
Nearly one reply in three. 32.9% (82 of 249) of owner replies to negative reviews contain identifiable deflection — disputing the customer's account, blaming them or a third party, or defending the shop's conduct at length instead of addressing the complaint. That's the single most striking number in this study: it's roughly four times the 8.4% we originally published, once every reply was hand-graded rather than sampled (see the correction notice above). Deflection can appear alongside an apology or a remedy in the same reply — the categories aren't mutually exclusive — so a reply can both say sorry and argue the customer is wrong.
Do shops ever send a canned five-star thank-you reply to a one-star review by mistake?
Yes — we found five confirmed cases across three shops of a positive, boilerplate thank-you template (independently verified as reused three or more times elsewhere in the corpus) firing on a 1- or 2-star review. Two of them literally thank the customer for "the five-star review" on a review that was actually one star.
How was this study done, and how is it different from your first review study?
Same corpus as our [first study](/guides/west-georgia-auto-shop-review-study) — 5,716 Google reviews from 68 West Georgia auto shops, collected August 1, 2026 — but new questions. This study tests what structural review attributes (a photo, the reviewer's history, Local Guide status) predict a negative rating, and separately classifies the actual text of all 249 owner replies to negative reviews, which the first study never analyzed. The corpus was independently reconstructed from the original unfiltered scrape and verified to match the first study's published review count exactly before any new numbers were computed. Full methodology, including a disclosed classifier spot-check accuracy, is above.
Keep reading
Last updated 2026-08-07. General guidance only — confirm specifics with a local shop for your exact vehicle.
