a

Coding agents (Claude Code, Codex, and the newer autonomous planners that open PRs on their own) are effectively unmanaged clients making high-frequency calls to a frontier model. From an infrastructure standpoint that is an uncontrolled egress path carrying source code and potentially regulated data to a model endpoint, with no per-identity metering, no content filtering, and no audit record. Nexus Proxy closes that gap by terminating agent traffic at a proxy you run, metering and redacting inline, and invoking the model with your own IAM role so nothing leaves your account.

This post covers the architecture, the data path, the failure semantics, and the deployment topologies.

Design constraints

The system is built around four hard requirements that come directly from running agents in a regulated environment:

  1. Model invocation stays in-account. The model is called with the customer account’s own IAM role against Amazon Bedrock. No traffic to a third-party endpoint, which satisfies VPC-residency and data-boundary requirements.
  2. Governance is inline but non-blocking. Metering and prompt capture cannot add latency to the model call or take developers down if the backend is unavailable. The path is fail-open by design.
  3. The control plane is multi-tenant. One usage backend ingests from many proxy installs across many accounts, with roll-ups by model, user, team, and day.
  4. Everything is a deploy, not an integration. Each component ships as Terraform/CloudFormation (in AWS) or a Docker Compose bundle (off AWS). No code changes in the agent tooling; it is an endpoint swap.

Topology: two independent stacks

Nexus is two installs that are deliberately decoupled. The only coupling between them is HTTPS plus an installation credential.

  • Proxy stack (customer account): LiteLLM + a Nexus plugin, a durable buffer, and a forwarder. This is the data plane for model traffic.
  • Usage stack (Nexus account or anywhere HTTPS-reachable): a Go usage API plus a Postgres-backed usage/prompt store. This is the control/reporting plane.

Because the two stacks talk only over authenticated HTTPS, the usage API does not need to live in AWS at all; only the proxy does, because that is what touches Bedrock.

Data path

Model request path (synchronous, hot)

agent/tool ──API──▶ LiteLLM Proxy (+ Nexus plugin, on ECS) ──invoke──▶ Amazon Bedrock
                         │                                              (Claude, GPT-OSS, ...)
                         └─ authenticate user, enforce quota, redact, emit usage event
  1. Coding agents point at the LiteLLM endpoint instead of a public model endpoint. Zero code change; it is a base-URL swap.
  2. The Nexus plugin authenticates the calling user, applies per-user/per-team quota enforcement, and runs PII/secret redaction on the prompt before the model sees it.
  3. LiteLLM invokes Bedrock using the account’s own IAM role. The response returns to the agent on the hot path.
  4. A usage event (and, per configuration, the prompt body) is emitted to the buffer. This emission is off the critical path.

Usage reporting path (asynchronous, buffered)

Nexus plugin ─events─▶ Durable buffer ─drain─▶ Forwarder ─HTTPS─▶ Nexus Usage API ─store─▶ Usage DB
                       (SQS + S3 spool)        (Lambda,          (Go, multi-tenant)     (Postgres;
                                                retries)                                 prompts in PG or S3)
  1. Events land in a durable buffer: an SQS queue with an S3 spool for payloads that exceed queue limits or need to be retained.
  2. A forwarder Lambda drains the buffer and ships events to the usage API over HTTPS, with retries and backoff.
  3. The Go usage API is multi-tenant. It writes usage rows to Postgres and stores prompt logs in Postgres or S3 depending on configuration (S3 for larger bodies / cheaper retention).
  4. Reporting reads roll up by model, user, team, and day.

Failure semantics: fail-open, no data loss

The key property is that the reporting path is isolated from the request path by the durable buffer. If the usage API is slow, degraded, or fully down:

  • Developers and agents keep working, because model invocation does not depend on the usage API being reachable.
  • Nothing is dropped, because events sit in SQS/S3 until the forwarder can deliver them. The forwarder retries; delivery resumes when the API recovers.

This is the design decision that keeps the governance layer from being routed around. Metering is asynchronous to the model call, so there is no added inference latency and no hard dependency that could take developers offline. The tradeoff is that usage dashboards are eventually consistent (seconds-to-minutes behind under normal load, longer during an outage backlog), which is the correct tradeoff for a metering system.

Controls enforced at the proxy

Because all agent traffic terminates at LiteLLM + the Nexus plugin, that is the single enforcement point for:

  • Quota enforcement per user and per team (block or throttle once a budget is hit).
  • PII / secret redaction applied to the prompt before it reaches Bedrock.
  • Audit logging of every request: identity, model, timestamp, token counts.
  • Model availability control (allow/deny which models a given identity can reach), which prevents unapproved or unauthorized model usage.
  • Real-time cost analytics derived from the metered token counts.

Metering is the other half of this story: see Building a Billing-Grade AI Usage-Metering Gateway for how the same proxy turns raw model calls into priced, auditable, reconcilable revenue.

Deployment topologies

The proxy and usage API are independent installs, so placement follows your isolation and compliance constraints rather than the reverse.

A. Single AWS account. Both stacks in one account; the forwarder reaches the usage API over private HTTPS. Lowest operational overhead; suitable for a single team or a POC.

B. Separate AWS accounts (typical enterprise). Each workload account runs its own proxy stack; all proxies report cross-account to one central Nexus usage account. You get per-account blast-radius isolation on the data plane and consolidated reporting on the control plane.

C. Over the internet. The usage API runs anywhere HTTPS-reachable (another cloud, or the Docker Compose bundle on any host). Only the proxy must be in AWS, because only the proxy calls Bedrock.

Packaging: Terraform or CloudFormation for each AWS stack; a Docker Compose bundle for non-AWS usage-API hosts. In all three cases the link between proxy and usage API is HTTPS authenticated by an installation credential.

Why this scales with autonomous agents

A human developer issues a bounded number of prompts per day. An autonomous agent that plans, writes, and merges code can generate orders of magnitude more model traffic, with no human inspecting each prompt for sensitive data or runaway spend. In that regime, inline quota enforcement, redaction, audit, and metering are not reporting niceties, they are the control plane that makes it safe to let agents run unattended. Nexus provides that plane without adding inference latency or a hard runtime dependency on the backend.

Getting started

If agent adoption is blocked in security review, start with an AI Readiness Assessment: a scoped engagement that maps the target account’s compliance and VPC requirements before the integration build, so the proxy deployment lands correctly on the first pass.

Enterprise AI infrastructure, inside your own VPC. Details at nexus.allcode.com, or read the Nexus launch announcement.

Related reading