← Back to blog

A Reproducible Blind Test for High-Risk Customer Replies: Publish the 100-Point Rubric Before the Result

A Reproducible Blind Test for High-Risk Customer Replies: Publish the 100-Point Rubric Before the Result

On Monday morning, an important business customer reports that an outage delayed shipments. They want the root cause, corrective action, and compensation position by 3:00 p.m. This is not an ordinary rewriting task. One unverified sentence such as “our system failure caused your losses” may become an admission. A vague line about “taking the matter seriously” may make the customer feel ignored.

The useful comparison is therefore not which model sounds most polished. It is which output can preserve facts, control risk, acknowledge the impact, and establish a credible next step at the same time.

There are no dated raw outputs for this test yet, so this article does not name a winner or rank any model. It publishes a reproducible blind-test protocol. Results should only be reported after the models have been run and the outputs and score sheets have been preserved.

Why “it reads well” is not a scoring method

A high-risk reply faces opposing pressures. The customer needs clarity, but the sender must not invent an unconfirmed cause, promise unapproved compensation, or minimise a real operational impact. Judging prose alone rewards confident wording. Judging caution alone may reward a reply that offers no useful action.

Think of a busy Hong Kong cafe handling an order with a sensitive note. Speed is not the only requirement. The kitchen must follow the ticket, avoid adding ingredients, and tell the floor staff when an update will be available. Every model in the test needs the same ticket.

Fix one synthetic case

Use a wholly fictional B2B software incident. Do not include real customer, employee, contract, or personal data:

  • A fictional customer could not use the shipment booking function from 09:18 to 10:05, a total of 47 minutes.

  • The first notice was sent at 09:37, and service was restored at 10:05.

  • The customer says 12 shipment orders were delayed and requests the root cause, recurrence-prevention measures, and compensation position by 3:00 p.m.

  • At test time, only that timeline is confirmed. Root cause, possible data impact, and compensation have not been confirmed or approved.

  • The recipient is the customer’s Operations Director. Target length is 130 to 180 English words.

This case deliberately mixes known facts with pending questions. Turning a pending item into a fact is not greater persuasiveness. It creates observable rework and approval risk.

Give every model the same input

Record the model version, run date, interface, temperature, and any other adjustable setting. If a setting is not visible, write “not available” in the run manifest rather than guessing. Start a fresh conversation for every model and send only this prompt:

You are an enterprise customer-service writing assistant. Based only on the fictional case below, draft a formal customer reply of 130–180 English words to the customer’s Operations Director.

Confirmed facts:
- the shipment booking function was unavailable from 09:18 to 10:05, a total of 47 minutes;
- the first customer notice was sent at 09:37;
- service was restored at 10:05;
- the customer reports that 12 shipment orders were delayed.

Not confirmed or approved:
- root cause;
- whether any data was affected;
- compensation or service credit;
- final recurrence-prevention measures.

The reply must:
1. specifically acknowledge the customer’s operational impact;
2. state only confirmed facts;
3. avoid admitting legal liability or inventing a root cause, data status, or compensation position;
4. promise the next update at 15:00 today and identify what that update will cover;
5. avoid empty corporate language, emoji, and conversational slang.

Output only the body of the customer reply. Do not add analysis, a heading, or notes.

Do not reveal model names to the scorers. An independent recorder saves each untouched output, randomises the order, and assigns Output A, B, C, and so on. Seal the model-to-label mapping until scoring is complete. Preserve punctuation, errors, and formatting exactly as generated.

Five dimensions, 100 points

Each dimension is worth 20 points. Scorers must quote the output as evidence:

Dimension

Full-score standard

Main deductions

Factual fidelity

Times, impact, and status match the case

Changed figures, invented causes, pending items stated as facts

Risk control

Known and pending matters remain distinct

Liability admission, compensation promise, implied data assurance

Customer empathy

Specifically recognises the 12 delayed orders and operational impact

Generic apology, minimisation, or blame

Actionability

Commits to a 15:00 update and states what it will cover

No owner, time, or next step

Clarity and tone

130–180 words, formal, direct, and readable

Verbosity, slang, filler, hostility, or format failure

Define four critical failures: inventing a root cause, asserting that no data was affected, promising compensation, or explicitly admitting legal liability. A critical failure cannot be offset by elegant prose and must trigger revision.

Count revision rounds

Label the first answer Round 0. If it misses the precommitted score threshold or contains a critical failure, return the same original case, that output, and dimension-by-dimension evidence to the same model. Ask it to correct only the cited problems without adding facts. Allow no more than two revisions. Preserve the full text, score, critical-failure status, and pass status for Round 0, Round 1, and Round 2.

Revision count captures a real cost. A reply may eventually become usable but still require two follow-ups and a human rewrite. Equally, zero revisions never means automatic approval. An authorised person must still review the real incident, contract, policy, commitments, and sending time.

Keep the results table blank until the runs exist

Before execution, no score or ranking belongs in the table:

Anonymous output

Facts /20

Risk /20

Empathy /20

Action /20

Tone /20

Total /100

Revisions

Critical fail

Output A

Output B

Output C

When results are eventually published, include the run date, model and version, visible settings, full prompt, every raw output, scorer quotations, score sheets, and the unsealed mapping. If a model refuses, hits a length limit, or returns an error, record that outcome. Do not silently rerun until a preferred answer appears.

Common failure modes

Showing model names before scoring. Brand expectations will affect judgement. Scorers should see anonymous outputs only.

Changing context between models. One extra incident detail can change the answer. Case, prompt, language, and length must remain fixed.

Publishing only totals. Without raw outputs and quoted evidence, readers cannot audit the scoring.

Presenting the synthetic case as a real incident. Label it clearly as fictional and do not imply that any customer or product experienced it.

Treating one run as a permanent capability ranking. Any conclusion applies only to the recorded date, version, settings, language, and task. Updates require a new run.

Build the evidence before the conclusion

The expensive part of high-risk writing is often not the first draft. It is detecting the natural-sounding sentence that goes beyond the known facts. This protocol is intended to turn model choice into an auditable, repeatable quality decision, not to manufacture a dramatic leaderboard.

Essevin’s AI chat supports multiple models and can provide one place to organise a comparison workflow like this. Regardless of model, process real customer data under your organisation’s rules, and require an authorised person to verify every fact, risk statement, commitment, and release decision.


Information in this article is current as of 28 August 2026 and is provided for general reference only; it does not constitute advice of any kind. The test case is entirely fictional. As of publication, the described model runs have not been completed and this article reports no model ranking or capability conclusion. Any later result applies only to the recorded run date, model version, settings, prompt, language, and task, and should not be generalised as a permanent ranking for other situations. Third-party product features, pricing and policies are subject to their official announcements. Real customer communications must follow the organisation’s data-handling, contractual, legal, approval, and human-review requirements. Essevin service details are as shown on essevin.com and in the console.