← Blog index
2026-07-28·Delivery & acceptance

POC to Production: 12 Acceptance Metrics for Enterprise Agents

Moving an enterprise Agent from POC to production should be gated by metrics and evidence, not demo charm. These 12 items span scope, quality, reliability, security, and ops—each needs an owner, pass line, sampling method, and artifacts. Thresholds are starting points; tighten for regulated industries (finance: /blog/financial-services-agent-compliance). GeonAI delivers against acceptances; /agents are capability references only (not a public trial of all 363+ presets). Email [email protected]. Delivery: /blog/enterprise-ai-agent-delivery-4-steps. Red flags: /blog/ai-agent-project-red-flags.

How to use this checklist

  • Before the pilot: write all 12 into the acceptance sheet (with pass lines); blank = fail
  • During: refresh measured values weekly; two red weeks on a red-line item → stop expanding intents
  • Before scale: evidence for all 12; at most one “conditional pass” with a fix-by date
  • Evidence shapes: sample sheets, log exports, load reports, ACL cases, runbook version IDs

The 12 metrics (with pass lines)

#MetricSuggested pass line (start)Evidence
1Intent scope freezeSigned production intent list; out-of-scope → refuse/handoff copySigned intent table + change process
2Measurable successEach intent has resolve / handoff / refuse rulesIntent→rule map
3Baseline captured≥2 weeks pre-launch: volume, AHT/find-time, handoff rateBaseline report + definitions
4Citation ratePolicy/product answers missing citation ≤5% (n≥200)QA sheet: missing doc_id count
5Wrong-promise rateGuarantee/ETA/fee wrong promises ≈0; freeze intent on hitWeekly QA + incident tickets
6Handoff correctnessMissed handoff ≤2%; false handoff may be higher with explanationRule-hit logs vs human review
7Latency & stabilityP95 e2e meets SLA; error rate within agreed capAPM/gateway reports
8Tool-failure degradeTimeout/empty results never fabricate; fixed degrade copy 100%Fault-injection test log
9ACL / permissionsCross-line/class/store leak cases all fail (must refuse)ACL case set + results
10Audit completenessSessions replayable: intent, doc versions, state, reviewer20 full session replays
11KB version syncSwitch at effective time; stale-version client hits ≈0Change-window test + hit logs
12Ops & rollbackOn-call, rate limits, rollback steps, comms templates + one drillRunbook version + drill notes

Notes on items people try to soft-pedal

4–5: Citations and wrong promises (quality red lines)

Fluent answers without citations look great in POC and get expensive in production. Policy, fees, SOPs, and warranty replies need doc_id + version. On a wrong promise, take the intent offline before tuning prompts. KB guardrails: /blog/enterprise-rag-knowledge-base-agent.

6: Handoff is not failure—missed handoff is

POCs often celebrate low handoff. Production celebrates must-handoff when required: complaints, writes, safety, suitability, high-value claims. Optimize false handoffs; treat missed handoffs as incidents. Support modules: /blog/custom-customer-service-agent.

8–9: Degrade and ACL (reliability/security red lines)

Inventing tracking, stock, or account state on tool failure is an automatic no-go. ACL cases must deliberately request forbidden docs—every case should refuse or empty, not “usually right.”

12: No runbook means not production

Document who is on-call, how to throttle or kill an intent, how to roll back model/KB versions, and who approves external copy. Run one drill (including a fake outage) before signing scale-up.

Suggested scale gates

  1. ≤10% traffic or one channel: red lines (5, 8, 9, 11) hard-pass
  2. to 50%: 4, 6, 7 green for two weeks; recalc ROI with real d (/blog/ai-agent-roi-calculator)
  3. full: all 12 evidenced; changes via change order; weekly QA continues

What does not count as “POC passed”

  • An exec asked three questions and liked the answers
  • Vendor demo resolve-rate or CSAT screenshots
  • Prompt docs only—no intent table or don’t-say list
  • “We’ll add audit/ACL next week”—add them before production talk

What to bring to GeonAI

Bring: draft intents, baselines, desired pass lines, system/deploy constraints, whether you need a human-review state machine. Email [email protected], /pricing, or Live chat. We align milestones to the checklist—not to a “how smart it feels” score.

Frequently asked questions

Must all 12 be 100% before launch?

Red lines (wrong promise, tool degrade, ACL, stale KB) should hard-pass. At most one conditional pass elsewhere with a fix-by date—do not replace sign-off with “close enough.”

How large should samples be?

Start at ≥200/week for quality checks or full volume for that intent (whichever is smaller); denser in month one. Finance/healthcare should raise ratios or fully review high-risk intents.

Internal-only POC—still all 12?

You can relax latency and handoff internally, but citations, ACL, audit, and degrade still matter. Once external or production data is in play, use the full sheet.

How does this relate to the ROI sheet?

This checklist decides “can we ship”; ROI decides “should we expand.” Pass the checklist first, then recalc ROI with real auto-resolve rates.

Who has veto when metrics fail?

Give joint veto to the business owner and security/compliance (if any). A vendor or PM unilaterally “declaring pass” does not count.

Will GeonAI put these 12 in the contract?

Applicable items can become SOW acceptance clauses with industry-agreed thresholds. Preset pages and demos are not acceptance criteria.

POCproduction acceptancepilot metricscustom AgentGeonAI