TemplatesGet a demo →Book a meeting
Blog
GovernanceTechnical

Self-Hosted LLMs in Your VPC: Every Model Call, Not Just the App

In short

A self-hosted LLM platform is only as private as its model calls. Putting the application in your VPC is the easy part; what decides data residency is whether every generation, embedding and background call leaves through one AI gateway you control — and whether the vendor can prove that on a network with no other way out.

Key takeaways

  • "Runs in your VPC" is a claim about every model call, not just the application.
  • Every model call routes through your own AI gateway — the single egress you control, log and rate-limit.
  • The proof is a sealed network whose only exit is the gateway, run against the real production image.
  • Nothing phones home: no telemetry stream, no license check, no usage beacon.
  • Governance is what an outsider can verify: SIEM forwarding, signed release records, SSO and SCIM.

Any enterprise AI vendor will tell you their product can run in your VPC. Most of them are telling the truth about the application and saying much less about the model calls.

Putting the web server, the workers and the database inside your network is the part that is easy to say and easy to demonstrate. It is also not the part your security team is worried about. What they are worried about is the traffic that leaves — the calls to model providers, carrying your prompts and your documents — because that is the boundary that decides whether a self-hosted LLM deployment is a data-residency guarantee or a diagram. A single model call that skips the gateway makes the whole claim false, and it will not show up on an architecture slide.

What does a self-hosted LLM platform put in your VPC?

SphereIQ runs as one application image — web, API, background schedulers, OCR and the code sandbox — alongside PostgreSQL with pgvector holding the index, memory, runs and audit records. All of it sits in your account, in your VPC. There is no telemetry stream back to us, no license check that phones home and no usage beacon; the only outbound traffic is the traffic you configure. Installation runs onto virtual machines with a blue/green release and automatic rollback, and a Helm chart and Terraform module are available for a Kubernetes estate.

That much is table stakes, and any serious vendor clears it. The real work starts one layer down, at the calls that are supposed to leave — and the ones that are not. It is the same boundary the Governance plane is built around.

Why it matters

The application inside your network is the easy ninety percent. The model calls leaving it are the ten percent your security team will actually ask about.

Why is the model call the boundary that matters?

An enterprise deployment routes every model call through the customer's own AI gateway: one egress point they control, log and rate-limit, in front of whichever providers they have approved. SphereIQ makes its calls through provider clients whose base URL is configuration, so pointing them at your gateway routes generation, embeddings, transcription and speech through it. The gateway can stand in as an OpenAI-compatible or an Azure OpenAI endpoint, and the platform is multi-provider by design: the same deployment can front OpenAI, Anthropic and others, or a single model you host yourself.

Optional add-ons, such as reranking and image or video generation, stay off until you enable them, so nothing leaves your boundary for a feature you did not choose.

How do you prove nothing bypasses the gateway?

The claim is only worth the proof, so here is how it is proved rather than asserted. The production image runs on a container network with no route to the internet and a recording stand-in as the only thing that answers a model request. The chat answer, the memory extraction and every embedding arrive at that gateway; the scheduled background jobs run; a direct call to a provider's public API fails to resolve, by construction. If anything in the system tried to reach a provider around the gateway, the test would fail, loudly, on a sealed network — which is the point of running it on one.

The way you find a hidden call is not a code review that swears there are none. It is a network with no way out and a recorder as the only exit, run against the real image. A promise that nothing bypasses the gateway is worth exactly as much as the test that would break if something did.

In a pilot, the same test runs against your gateway in the first week, so the evidence comes from your environment rather than from ours.

Governance is what an outsider can verify

Running in your VPC is where governance becomes possible; it is not the same as governance. What makes the difference is what someone outside your team can check.

  • Audit events in your SIEM. Administrative and security events are written to an audit log and forwarded to your SIEM as Splunk HEC, Datadog, CEF or JSON.
  • Signed release records. Releases carry an Ed25519-signed receipt rebuilt from stored records, so a change to what is live is verifiable rather than asserted.
  • Your identity provider at the front door. Access runs through SAML single sign-on and SCIM provisioning, with roles enforced in the API rather than hidden in the page.
  • Sealed secrets. Connector tokens, flow secrets and signing keys are sealed in the database with AES-256-GCM under a separate derived key per purpose.
  • Production is a property of the box. Whether a deployment is production is set by its environment, not a database row a mistake could flip; an unconfigured instance assumes development and performs no changes.

SphereIQ is SOC 2 Type II certified, and because a self-hosted deployment keeps the audit trail, access controls and signed records inside your environment, your own ISO 27001 or ISO/IEC 42001 work checks evidence you can reproduce rather than a control you have to take on trust.

What to ask any vendor about self-hosting

When a vendor says "it runs in your VPC", they have answered the easy question. Ask the two hard ones. Which model-related calls — generation, embeddings, reranking, transcription, OCR, background jobs — go through our gateway? And can you prove it on a sealed network rather than in a diagram? A vendor who can name every call and hand you a test has thought about the boundary you actually care about.

Question What good looks like
Which model calls leave the VPC? All routed through your gateway
Can you prove nothing bypasses it? Yes — on a sealed network
What about optional features? Off until you enable them
Where does data reside? Your account; no call-home

For how the same boundary applies to what an agent is allowed to reach, see what MCP means for enterprise AI.

Frequently asked questions

What does "runs in your VPC" actually include?
The full application — web, API, schedulers, OCR and the code sandbox — plus PostgreSQL with pgvector, all in your cloud account. Outbound traffic is limited to your AI gateway and the systems you choose to connect. No telemetry or license check returns to the vendor.
Can every model call go through our own AI gateway?
Yes. Generation, embeddings, transcription and speech are made through provider clients whose base URL is configuration, so pointing them at your gateway routes them through it. The gateway can stand in as an OpenAI-compatible or Azure OpenAI endpoint, in front of whichever providers you approve.
How do we prove nothing calls a model provider directly?
Run the image on a network whose only egress is your gateway and watch every model call arrive there. As a backstop that does not depend on the vendor, add a network rule that allows the gateway and denies every provider host.
How does a self-hosted deployment support our security review?
SphereIQ is SOC 2 Type II certified, and a self-hosted deployment adds an exported audit trail, SSO and SCIM access control, and signed release records inside your environment — so most of what a vendor-risk assessment asks for is evidence your own team can verify directly.

See every model call arrive at your gateway.

In a walkthrough we run SphereIQ against your AI gateway and show where each call goes — generation, embeddings and background jobs included.