
30-Day Custom AI Agent Pilots: Resolve Rate vs “Do No Harm” Metrics
The easiest way to derail a custom AI Agent pilot is to make auto-resolve rate the only KPI. If the first 30 days chase that number alone, the system and the ops team learn the same tricks: fewer handoffs, forced answers on fuzzy asks, and happy-path-only evals. The weekly chart looks great; complaints and wrong promises explode after you scale. Do this instead: weight harm-control metrics above efficiency, treat resolve rate as a reference, and optimize it only after the scale gate passes. GeonAI pilots on acceptable metrics; /agents are capability references only (not a public trial of 363+ presets). Email [email protected]. Full 12-item list: /blog/poc-to-production-agent-checklist. Boundary tables: /blog/custom-agent-scope-boundary-in-contract.
Two metric families: efficiency vs harm control
| Family | Examples | Weight in days 1–30 |
|---|---|---|
| Efficiency | Auto-resolve, first response, human-time saved | Medium |
| Harm control | Wrong promises, missing citations, ACL leaks, missed handoffs | High |
| Experience sentinels | CSAT / escalation ticket count | Medium-high |
Efficiency metrics are useful—but they are not the scale key by themselves. If harm gates fail, do not open the next channel no matter how high resolve looks.
Minimal 30-day dashboard
- Missing-citation rate: factual claims without doc version/section
- Stale-doc hit rate: answers grounded on retired or expired versions
- ACL / cross-store leaks: reads that should never happen (target 0)
- Missed-handoff rate: should-hand-off but did not (emotion, missing evidence, Table-C actions)
- Wrong-promise sample rate: price / lead time / refund severity samples (target 0 severe)
- Auto-resolve rate: reference only; denominator rules below
- Complaint / escalation count: week-over-week sentinel
Knowledge and citation mechanics: /blog/custom-agent-more-than-a-prompt, /blog/enterprise-rag-knowledge-base-agent.
How resolve rate gets gamed
- Lower handoff thresholds without improving KB/tools
- Force definite promises on fuzzy asks (especially lead time and price)
- Eval sets with happy paths only—no refuses or ACL cases
- Count Table-B refuses as “resolve failures,” which pushes the model to refuse less
Fix: exclude compliant Table-B refuses from the auto-resolve denominator; do not count unconfirmed Table-C sends/writes as auto-wins. Use the same denominator in the contract and the weekly report.
Scale gate (example)
- Two consecutive weeks: missing-citation rate under the agreed threshold
- Wrong-promise sampling: zero severe findings
- ACL / cross-store leaks: zero
- Missed-handoff under threshold, with complete ticket fields
- Only then open the next entry or raise traffic; otherwise fix knowledge/rules only—no expansion
Gate details and evidence: /blog/poc-to-production-agent-checklist. Talk ROI after harm gates pass—see /blog/ai-agent-roi-calculator.
Who reads the weekly report—and who can veto scale
- Business owner: complaints and efficiency
- Delivery / IT: missing citations, tool failures, versions
- Compliance / risk (if present): veto on scale for ACL leaks and wrong promises
How to brief GeonAI
Share the pilot entry, acceptable harm thresholds (e.g. severe wrong promises = 0), and the scale gates you want in the contract. Email [email protected], /pricing, or Live chat.
Frequently asked questions
What if the business only accepts resolve rate?
Write harm metrics into the scale gate; make resolve rate an optimization target after the gate passes—not the sole acceptance line.
What if 30 days is not enough volume?
Extend until the sample is statistically meaningful. The gate logic stays—low volume is not a free pass to skip harm sampling.
How do we sample without a labeling team?
Start with a small dual-review sample (e.g. 50–100 threads/week), prioritizing price, lead time, refunds, and refuse cases.
Do Table-B refuses hurt resolve rate?
Not if the contract says so. Refuses are in-spec behavior; track “compliant refuse rate” separately.
How does this relate to the artifact checklist and boundary tables?
The checklist requires an acceptance-metrics sheet; the tables define do/don’t/confirm; this post defines which numbers to watch in 30 days and how to stop resolve-rate gaming.
Can CSAT replace harm metrics?
No. CSAT lags and can be soothed by wording. Wrong promises and ACL leaks need sampling and logs—not satisfaction alone.