ATL 84° / clear
The Trades Desk

How Call Centers Improve Conversation Quality

Quality scores go up. Customers stay angry. The gap between those two facts is where most contact-center improvement programs quietly die.

By Clara Montgomery
ATLANTA · August 19, 2026 · 6:30 AM ET
12 min read
How Call Centers Improve Conversation Quality

The best call I ever heard scored a 71. It was an inbound billing dispute in a Tampa center, a woman who had been double-charged for four months and had already called twice. The agent, eight weeks out of training, said early: "Okay. Before I explain anything, I want to make sure I have this right — you've called about this twice already and it's still wrong." Then she stopped, let the customer talk for almost two uninterrupted minutes, fixed it, and told her exactly what would appear on the statement and when. The customer thanked her by name.

It scored 71 because the agent skipped the branded greeting, never offered the retention promo, and went thirty-one seconds over the hold-verbiage window. Somewhere in a QA queue, a competent evaluator marked it down, correctly, according to the form we gave her. That form was the problem.

I spent four years running quality and coaching for a multi-site outsourced contact-center operation in the Southeast — roughly 900 seats across inbound service, billing, and warranty. What follows is what actually moved conversation quality, and what we bought that did nothing.

The popular claim that is wrong

The standard answer is that quality is a measurement problem: score more calls, score them more consistently, calibrate the evaluators, get the sample size up, and quality follows. This is the premise behind almost every QA platform sold in the United States right now, and it is why AI evaluation is being sold as the breakthrough — because it takes you from two percent of calls scored to one hundred percent.

We went from two percent to one hundred percent. Quality scores rose about nine points over two quarters. Customer effort scores did not move. Repeat contact rate did not move. Complaint escalations went up slightly, which I suspect was noise, but it was certainly not down.

Here is the judgement, and I will not hedge it: scoring more calls does not improve conversations, because a scorecard measures compliance with a form, and customers are not evaluating compliance with a form. The centers that improved permanently changed the form itself, changed what happened in the first ninety seconds, and changed the coaching cadence. Coverage was the least important of the four things we did.

The scorecard is the product

Our original form had thirty-one line items. Greeting elements, verification script, hold protocol, empathy statement, brand phrase, upsell attempt, closing script, and a dozen more. It took an evaluator eleven minutes per call. Agents could recite it. And it rewarded a call that hit every box and left the customer confused.

We cut it to six items, all binary, all answerable by someone who had never worked in the account:

Did the agent state the customer's problem back before proposing anything? Did the agent give a specific next event with a date or timeframe? Did the agent avoid transferring without a warm introduction? Was the customer told what would happen if the fix did not work? Did the agent stop talking after presenting the resolution? Was there any point where the customer had to repeat information already given?

Evaluation time dropped to under four minutes. Inter-rater agreement — we tested it with blind double-scoring on 200 calls — went from around 0.61 to 0.88. And unlike the thirty-one-item form, agents could hold all six in their head at 4 p.m. on a Thursday, which is the only test that matters.

An honest caveat: we lost something. The long form caught compliance drift on verification language, and in a regulated account that is not optional. We ended up running verification as a separate automated pass-fail check outside the quality score entirely. That was more work, not less. I would still do it, because mixing legal compliance and conversation quality on one form makes both of them worse.

The first ninety seconds decide the call

We pulled 4,000 calls and coded them against outcome. The single strongest predictor of a low-effort, non-repeating contact was not agent tenure, not handle time, not product knowledge scores. It was whether the customer's actual problem had been named out loud in the first ninety seconds.

That sounds soft. It is mechanically specific. It means the agent says some version of "So the charge posted twice and the credit you were promised in June never appeared" before saying anything about policy, process, or what they are able to do. Not an empathy statement — those we banned, or rather we stopped scoring them, and usage of "I understand how frustrating that must be" dropped by roughly half within a month, which tells you it was never sincere to begin with.

"You're not calming her down. You're proving you were actually listening, and that's the only thing she can't get from the website."

The practical change: we rewrote the opening block of every workflow so the summarize-back step came before verification, not after. Legal did not love it. We kept verification inside the first two minutes and nothing broke. Repeat contact rate on billing dropped a little over four points in the following quarter.

Coaching cadence beats coaching quality

We had good coaches giving thoughtful monthly feedback sessions, forty-five minutes each, well documented. It did almost nothing. What worked was worse feedback delivered far more often.

Two calls per agent per week. Fifteen minutes. The agent picks one of the two. One behavior discussed, chosen from the six-item list, and no other topic permitted in the session even if something else was obviously wrong. Supervisors hated the constraint for about six weeks and then stopped mentioning it.

Quality scores on that population moved eleven points over a quarter against a control group that got the monthly deep session. More usefully, ninety-day attrition in the coached group dropped by a third. The agents' stated reason, repeatedly, in exit and stay interviews, was that they finally knew what they were being judged on. That is an indictment of how we had been managing them, and I include myself in it.

Where AI QA actually helps, and where it does not

The AI evaluation tools sold into US contact centers in 2026 are genuinely good at three things: transcription accuracy, detecting whether a specific phrase or step occurred, and surfacing calls that deviate from the norm. If your question is "did the agent give a specific next event with a date," a model answers that reliably and at full coverage, and that is real leverage that did not exist five years ago.

They are not good at judging whether the resolution was correct. They score confidence and fluency, and a confident agent who gave the customer a wrong answer in a clear, well-structured way will score high. We caught this by hand-auditing the top-scoring decile and finding an uncomfortable number of clean-sounding calls that had told the customer something false. Anyone selling you a model that scores accuracy of resolution is selling you a model that scores tone.

A prediction I will put a stake in, and I could be wrong: by mid-2027 the useful pattern in US centers will be full-coverage machine scoring on the mechanical items and a small deliberately sampled human review focused only on whether the answer was right. Centers that hand the whole judgment to the model will find their quality scores and their complaint volume rising together, and will spend a year not understanding why. If I am wrong, it will be because retrieval-grounded evaluation gets good enough at verifying answers against the knowledge base faster than I expect.

The metrics that fought us

Average handle time is the honest obstacle and nobody in this industry likes saying so. Every change above adds seconds. Summarizing the problem back costs fifteen to twenty. Confirming what happens if the fix fails costs another ten. Our AHT went up roughly forty seconds and stayed up.

Total contacts per resolved issue went down enough that cost per resolution fell about seven percent. That was the argument that saved the program with finance, and I only had it because we had been tracking repeat contact at the issue level, not the call level. If you cannot measure repeats against the underlying problem rather than the phone call, you will lose this argument internally, and the AHT people will win, and they will be wrong.

What to change Monday

Cut the scorecard to six binary items a stranger could grade. Move verification after the summarize-back. Stop scoring empathy statements. Coach two calls a week for fifteen minutes on one behavior. Hand-audit your highest machine-scored calls for wrong answers before you trust the model. Track repeats per issue, not per call, before you touch anything else.

The Tampa call that scored 71 is now in our new-hire library. Under the current form it scores a 5 out of 6 — it loses the point for a transfer with no warm introduction, which is fair, and which the agent fixed the next week.

Call CenterContact CenterCustomer ExperienceQuality AssuranceCoaching