The 2% Problem: Why QA Reviews Only a Slice of Calls
١٤ سبتمبر ٢٠٢٦ · 5.1 دقيقة قراءة · كتبه ونشره Whizz Scribe

The 2% Problem: Why Most Companies Only Ever Hear a Fraction of Their Customer Calls
You review 20 calls out of 1,000. The contact center QA scorecard looks healthy. Monday still starts with complaints—bad transfers, missed follow-ups, and customers who had to call back twice. The gap between the scorecard and customer experience is significant, as only 71% of calls are resolved on the first try, leaving 29% to require a callback and 19% to be transferred—a move that cuts customer satisfaction by as much as 15% SQM Group SQM Group.
SMB operators need a sharper frame than the “2% bad, 100% AI good” story. The real work is separating estimation problems from retrieval problems. For bilingual, phone-first operations, that unseen 98% becomes a live risk in routing, coaching, and follow-up. The industry shorthand is “2%,” but sources put manual QA coverage at just 1–3% of interactions Verint, Talkdesk. Some narrow it to 1–2% Cresta, while one vendor states 2% directly Observe.AI, meaning the vast majority of calls are untouched by manual review.
A familiar Monday-morning problem: the dashboard says calls are fine, but customers are still complaining
The 2% problem has a familiar sound: average QA scores are steady, each agent has a few calls reviewed, and the weekly coaching notes are already written. On paper, the team is working.
Then the front office hears it again. Customers transferred three times. A promised callback never came. A cancellation call that turned into a complaint. The scorecard said “fine” while the phone line burned.
This is a coverage failure. Traditional review programs sample a thin slice of calls, missing the full operational picture. Both Verint and Talkdesk describe manual review as covering just 1–3% of interactions. Cresta is blunter: when teams sample 1–2% of calls, they cannot verify compliance for the other 98%. While that sample size may be enough to estimate an average, it is insufficient for finding the exact journey that broke.
The real mistake: treating every call question like a QA sampling question
Most teams treat every call question the same. They aren’t.
Some are estimation questions. Are agents generally following the opening script? Is courtesy improving week over week? A random sample can answer that. It can estimate the center of the distribution, show a trend, and feed a scorecard.
Other questions are about retrieval. These are investigations. Which calls missed a regulatory disclosure? Which repeat callers were promised a follow-up that never happened? Which transfers created churn signals? Which new objections appeared after a pricing change? Which Arabic and English handoffs are failing?
These questions require finding specific instances rather than calculating an average.
This is the core of call retrieval vs. sampling. Sampling measures broad performance. Retrieval locates specific failures. If an operator needs a root cause, a broken journey, or exact evidence, random QA is the wrong tool for the job.
Why 2% coverage misses the calls operators most need to find
A small random sample is weakest precisely where operators have the most to lose.
Take 5,000 monthly calls. At 2% QA, only 100 get reviewed. The other 4,900 go unheard. If a harmful event happens in just 0.5% of calls, you might catch one. Many months, you will catch none while the problem stays live. The most expensive call problems are these rare-but-important events.
A failed disclosure. A transfer loop tied to one department. A missed callback promise. A language-switch failure in a billing explanation. A revenue leak where customers who asked one specific question later churned.
Random sampling does not reliably surface these patterns.
Other fields learned this lesson. In one SEC study, regulators found random surveillance inadequate for detecting certain trading violations. In that case, an employee disabled an automated system for months, and a later sample of orders from that period showed a 46% violation rate. A New Jersey State Comptroller audit found 54.7% of sampled drug-testing episodes failed legal requirements, while also identifying other failures outside the sample. Although the domains are different, the lesson is the same: low-coverage surveillance misses live failures. Phone operations move even faster, making retrieval more critical.
The hidden multiplier in bilingual call operations
Bilingual call work adds a second blind spot.
Arabic and English conversations often code-switch—bouncing between scripts, transliterations, and dialect shortcuts. A reviewer might hear the interaction clearly but tag it inconsistently. The transcript might be usable but flatten the real meaning.
The ArzEn corpus documents 12 hours of spontaneous Egyptian Arabic-English code-switched speech from 38 speakers. Research on PolyWER argues that standard word error rate is too strict, because a transliteration can be correct even if it’s not a literal script match. Another benchmark found transliteration followed by text normalization correlated best with human judgments on dialectal Arabic-English speech here.
For operators, this means manual review becomes less representative, tags get less consistent, and coaching gets noisier.
A multilingual call-center ethnography confirms that language management is central to the daily work of hiring, training, and evaluating agents across languages study. Instead of a random sample, bilingual phone teams must be able to retrieve the exact moment a customer switched languages before surfacing their real objection.
A practical coverage model for SMBs: what to review randomly, what to target, and what to monitor fully
Most SMB teams cannot manually review every call. They don’t have to. They just need a cleaner operating model.
Random review—for estimation. Use a small, stable sample for broad coaching on greetings, verification, and tone. NICE describes an incremental step: review 3–5 meaningful, metric-tied calls instead of 7 random ones, and keep it tied to a supervisor scorecard NICE.
Targeted retrieval—for investigation. Pull calls based on signals: low CSAT, long handle time, repeat contacts, or high frustration. NICE and Verint both recommend using analytics to choose interactions intentionally, surfacing compliance failures or exceptional service.
Near-complete analysis—for high-risk categories. Start where failure is expensive. Regulatory disclosures. Payment promises. Refund disputes. Retention calls. Arabic and English handoffs. Verint’s guidance points teams toward these predefined high-risk interactions first, before expanding coverage over time Verint.
This is where auto-QA and call analytics earn their keep by providing genuine coverage instead of performative metrics. Manual QA also has a scaling risk—it’s hard to expand QA headcount with call volume, which is why 100% AI QA becomes a late-stage goal on most roadmaps Verint.
From insight to action: the operational fixes call analysis should trigger
Analysis must trigger operational fixes.
If transfers spike, change the routing. SQM reports 19% of calls are transferred, cutting satisfaction by up to 15% SQM Group. That’s a routing, staffing, or desk problem—often all three.
If missed follow-ups appear, launch callback rules. The IRS offered callbacks to 17.2 million taxpayers in FY2024, saving millions of hours of hold time. This function provides essential queue control.
If repeat calls cluster around one journey, fix the journey. McKinsey found that eliminating the 20% of repeat calls at one energy company could save $40 million from a $200 million cost base, representing both revenue and cost leakage.
As first-contact resolution improves, costs fall. Qualtrics data shows each 1% improvement in FCR reduces operating costs by about 1%.
Find the calls. Change the route. Rewrite the script. Adjust staffing. Tighten coaching. Monitor again.
One takeaway to keep: if the question is “what exactly went wrong?”, sampling is the wrong tool
Use sampling for averages. Use retrieval for causes. Use broad analysis for risk.
If the question is about overall performance, a random sample can work. If the question is about compliance violations, churn signals, missed follow-ups, or one broken customer journey, then sampling is an inappropriate tool for the task.
Pick one queue this week—cancellations, transfers, repeat callers, or Arabic and English handoffs. Stop sampling it. Retrieve every call.
المصادر
- Verint: Call Quality Monitoring
- Talkdesk: Automated Quality Management
- Cresta: Call Quality Monitoring Guide
- Observe.AI: How QA Volume Impacts Call Center Agent Performance
- SQM Group: Call Transfer/Hold Performance Impact (CSAT and FCR)
- SQM Group: Customer Effort Required
- Verint: Contact Center Quality Management
- U.S. SEC Study on Limit Order Practices
- New Jersey State Comptroller: Drug Testing Audit
- NICE: Strategies for Advancing CX (QA review guidance)
- Verint: Call Center Quality Assurance Best Practices
- Verint: Contact Center Quality Management
- Verint: Call Quality Monitoring
- IRS PDF: Callback/Queue Control (P5456)
- McKinsey: Why Are Your Customers Calling You Again?
- Qualtrics: First-Call Resolution and Cost Impact
- ArzEn Corpus (LREC 2020)
- PolyWER (EMNLP Findings 2024)
- arXiv: Transliteration + Text Normalization for Dialectal Speech
- Language Management and Language Work in a Multilingual Call Center