
POC to Production: 12 Acceptance Metrics for Enterprise Agents
Moving an enterprise Agent from POC to production should be gated by metrics and evidence, not demo charm. These 12 items span scope, quality, reliability, security, and ops—each needs an owner, pass line, sampling method, and artifacts. Thresholds are starting points; tighten for regulated industries (finance: /blog/financial-services-agent-compliance). GeonAI delivers against acceptances; /agents are capability references only (not a public trial of all 363+ presets). Email [email protected]. Delivery: /blog/enterprise-ai-agent-delivery-4-steps. Red flags: /blog/ai-agent-project-red-flags.
How to use this checklist
- Before the pilot: write all 12 into the acceptance sheet (with pass lines); blank = fail
- During: refresh measured values weekly; two red weeks on a red-line item → stop expanding intents
- Before scale: evidence for all 12; at most one “conditional pass” with a fix-by date
- Evidence shapes: sample sheets, log exports, load reports, ACL cases, runbook version IDs
The 12 metrics (with pass lines)
| # | Metric | Suggested pass line (start) | Evidence |
|---|---|---|---|
| 1 | Intent scope freeze | Signed production intent list; out-of-scope → refuse/handoff copy | Signed intent table + change process |
| 2 | Measurable success | Each intent has resolve / handoff / refuse rules | Intent→rule map |
| 3 | Baseline captured | ≥2 weeks pre-launch: volume, AHT/find-time, handoff rate | Baseline report + definitions |
| 4 | Citation rate | Policy/product answers missing citation ≤5% (n≥200) | QA sheet: missing doc_id count |
| 5 | Wrong-promise rate | Guarantee/ETA/fee wrong promises ≈0; freeze intent on hit | Weekly QA + incident tickets |
| 6 | Handoff correctness | Missed handoff ≤2%; false handoff may be higher with explanation | Rule-hit logs vs human review |
| 7 | Latency & stability | P95 e2e meets SLA; error rate within agreed cap | APM/gateway reports |
| 8 | Tool-failure degrade | Timeout/empty results never fabricate; fixed degrade copy 100% | Fault-injection test log |
| 9 | ACL / permissions | Cross-line/class/store leak cases all fail (must refuse) | ACL case set + results |
| 10 | Audit completeness | Sessions replayable: intent, doc versions, state, reviewer | 20 full session replays |
| 11 | KB version sync | Switch at effective time; stale-version client hits ≈0 | Change-window test + hit logs |
| 12 | Ops & rollback | On-call, rate limits, rollback steps, comms templates + one drill | Runbook version + drill notes |
Notes on items people try to soft-pedal
4–5: Citations and wrong promises (quality red lines)
Fluent answers without citations look great in POC and get expensive in production. Policy, fees, SOPs, and warranty replies need doc_id + version. On a wrong promise, take the intent offline before tuning prompts. KB guardrails: /blog/enterprise-rag-knowledge-base-agent.
6: Handoff is not failure—missed handoff is
POCs often celebrate low handoff. Production celebrates must-handoff when required: complaints, writes, safety, suitability, high-value claims. Optimize false handoffs; treat missed handoffs as incidents. Support modules: /blog/custom-customer-service-agent.
8–9: Degrade and ACL (reliability/security red lines)
Inventing tracking, stock, or account state on tool failure is an automatic no-go. ACL cases must deliberately request forbidden docs—every case should refuse or empty, not “usually right.”
12: No runbook means not production
Document who is on-call, how to throttle or kill an intent, how to roll back model/KB versions, and who approves external copy. Run one drill (including a fake outage) before signing scale-up.
Suggested scale gates
- ≤10% traffic or one channel: red lines (5, 8, 9, 11) hard-pass
- to 50%: 4, 6, 7 green for two weeks; recalc ROI with real d (/blog/ai-agent-roi-calculator)
- full: all 12 evidenced; changes via change order; weekly QA continues
What does not count as “POC passed”
- An exec asked three questions and liked the answers
- Vendor demo resolve-rate or CSAT screenshots
- Prompt docs only—no intent table or don’t-say list
- “We’ll add audit/ACL next week”—add them before production talk
What to bring to GeonAI
Bring: draft intents, baselines, desired pass lines, system/deploy constraints, whether you need a human-review state machine. Email [email protected], /pricing, or Live chat. We align milestones to the checklist—not to a “how smart it feels” score.
Frequently asked questions
Must all 12 be 100% before launch?
Red lines (wrong promise, tool degrade, ACL, stale KB) should hard-pass. At most one conditional pass elsewhere with a fix-by date—do not replace sign-off with “close enough.”
How large should samples be?
Start at ≥200/week for quality checks or full volume for that intent (whichever is smaller); denser in month one. Finance/healthcare should raise ratios or fully review high-risk intents.
Internal-only POC—still all 12?
You can relax latency and handoff internally, but citations, ACL, audit, and degrade still matter. Once external or production data is in play, use the full sheet.
How does this relate to the ROI sheet?
This checklist decides “can we ship”; ROI decides “should we expand.” Pass the checklist first, then recalc ROI with real auto-resolve rates.
Who has veto when metrics fail?
Give joint veto to the business owner and security/compliance (if any). A vendor or PM unilaterally “declaring pass” does not count.
Will GeonAI put these 12 in the contract?
Applicable items can become SOW acceptance clauses with industry-agreed thresholds. Preset pages and demos are not acceptance criteria.