
Private Deployment Is Not Default: When Data Classification Justifies It
“Private” for a custom AI Agent means inference (and optionally weights) run inside a boundary you control, with knowledge, logs, and gateways governed the same way—not “custom therefore local GPUs,” and not “a Chinese-language model keeps data in-country.” Most pilots should start with data classification plus API or private cloud; add on-prem inference when residency or security baselines are written. Shapes and cost: /blog/deepseek-private-deployment-guide. GeonAI picks the gateway from your compliance outcome; /agents are capability references only (not a public trial of 363+ presets). Email [email protected].
Why custom ≠ private by default
- Custom delivery covers scenario, boundaries, knowledge, tools, and acceptance—not rack location
- With no residency clause, private clusters mainly add idle GPU and on-call cost; they do not automatically raise answer quality
- Ungoverned knowledge and missing ACL still leak or go stale on a private cluster
- Thin ops staff makes patching and recovery slower than a public API
Decision order
- Classify data: public / internal / confidential / restricted (what may enter prompts and logs)
- Decide whether inference must stay inside your boundary (written clause, not a “safer” feeling)
- Govern knowledge and logs on the same boundary (private store + egress prompts is incomplete residency)
- Check ops capacity (upgrades, CVEs, on-call, rollback)
Sample classification table
| Level | Examples | Typical inference |
|---|---|---|
| Public | Site FAQ, published product copy | Public API |
| Internal | General policy, training material | API + contract/DPA, or private cloud |
| Confidential | Unreleased formulas, customer contracts, unpublished quotes | Private cloud or on-prem |
| Restricted | Regulator-named data, air-gapped plant | On-prem + audit; tools stay on internal allowlists |
One Agent can mix lanes: public FAQ on the cloud, confidential traffic on private inference. Permissions and citations still need RAG/ACL: /blog/enterprise-rag-knowledge-base-agent. Integration and gray-release cost: /blog/custom-agent-quote-underestimated-costs. After classification, add audit logs and internal access controls: /blog/enterprise-ai-compliance-overview.
Signals you should not private-first
- Legal/security has no written residency scope (prompts, docs, logs, vectors)
- No GPU/inference platform and no on-call SRE
- KB ungoverned: no Owner, no versions, no expiry (release/rollback: /blog/custom-agent-knowledge-release-rollback)
- Pilot KPIs and boundary tables unsigned—govern first, buy cards later
Knowledge Owner and gate sign-off: /blog/custom-agent-project-owner-raci. Do / don’t / human-confirm tables: /blog/custom-agent-scope-boundary-in-contract.
How to brief GeonAI
Say whether residency or a security baseline is written, which cloud and ops staff you have, and which data class enters prompts. Email [email protected], /pricing, or Live chat. Private-deploy plans on /pricing need assessment—no one-price bundle.
Frequently asked questions
Is a Chinese-language model automatically safer?
Language ≠ residency. Safety follows deploy region, contract, log boundary, and whether inference is private—not whether the model “speaks Chinese.”
Must a custom Agent be privately deployed?
No. Without a written residency requirement, most projects can pilot on API or private cloud plus classification. Custom work is scenario and governance, not a default local GPU.
Does private cloud count as private deployment?
Private cloud/VPC is residency in an agreed tenant/region; the vendor still runs hardware. On-prem private is your own inference cluster. Both still need governed knowledge and logs.
Do we need on-prem inference without a security baseline?
Usually no. Confirm residency in writing first; without a clause, prefer API + ACL pilots instead of idle GPUs and empty on-call rotas.
If the KB is private but the model is a public API, does data stay in-boundary?
Usually not full residency: prompts and retrieved snippets can still hit the API. If prompts must not leave, the generation layer needs private cloud or private inference too.
Do the 363+ presets include a private-deploy architecture?
No. Presets illustrate capability types—not your classification table, residency clause, or ops staffing.