Documentation

Docs

Everything you need to deploy and configure SpendProxy. The integration is a one-line change — see for yourself below.

$ Using Claude Code? Meter your coding sessions — see the guide →

Quick start

1. Start SpendProxy

Using Docker:

 docker run -d -p 4100:4100 spendproxy/proxy:latest 

Or using npx:

 npx @cloudexpat/spendproxy 

2. Point your AI SDK at SpendProxy

Pick your provider and language. For every provider except AWS Bedrock, the highlighted diff is the entire integration, one line. API keys pass through untouched.

AWS Bedrock works a little differently. Keep your normal AWS credentials and region — only override the endpoint (above). Because Bedrock authenticates with AWS SigV4 rather than an API key, those client credentials only sign the hop to SpendProxy; the proxy ignores that signature and re-signs each call to Bedrock with its own IAM role. Grant the proxy bedrock:InvokeModel / bedrock:Converse (and their streaming variants) for /bedrock/*, plus bedrock-mantle:CreateInference if you use /bedrock-mantle/*, and set CE_BEDROCK_REGION. Cross-region inference-profile IDs like us.anthropic.claude-sonnet-4-6 work unchanged on /bedrock/*. The Messages endpoint takes regional model IDs instead, so use anthropic.claude-sonnet-5 rather than a profile ID there. Use https:// for any non-local proxy.

One required change for the AWS SDK. It defaults to HTTP/2 for Bedrock, and SpendProxy speaks HTTP/1.1, so a stock client fails at the connection with ERR_HTTP2_ERROR: Protocol error and nothing reaches the proxy. Pass NodeHttpHandler and it works:

import { NodeHttpHandler } from '@smithy/node-http-handler';

const client = new BedrockRuntimeClient({
  region: 'us-east-1',
  endpoint: 'http://localhost:4100/bedrock',
  requestHandler: new NodeHttpHandler(),
});

The Anthropic SDKs, including AnthropicBedrockMantle, use HTTP/1.1 already and need no change. Point Mantle at /bedrock-mantle with skipAuth: true so the proxy signs.

3. Open the dashboard

Navigate to localhost:4100/dashboard. Cost data appears in real time. The dashboard auto-refreshes every 10 seconds.

Configuration

SpendProxy is configured via environment variables.

Variable Default Description
CE_PROXY_PORT 4100 Port to listen on
CE_DATA_DIR ~/.spendproxy Data directory for SQLite database and config
CE_DASHBOARD_PASSWORD (none) Password-protect the dashboard and cost APIs
CE_PROXY_DB ~/.spendproxy/ce-proxy.db SQLite database path

Example with custom port and data directory:

 docker run -d \
  -p 8080:8080 \
  -e CE_PROXY_PORT=8080 \
  -e CE_DATA_DIR=/data \
  -v /host/data:/data \
  spendproxy/proxy:latest 

Supported providers and models

SpendProxy uses fuzzy matching for model IDs. Versioned identifiers automatically resolve to the base model. Pricing database updated continuously.

OpenAI

gpt-5.5, gpt-5.5-pro
gpt-5.4, gpt-5.4-mini, gpt-5.4-nano
gpt-4.1, gpt-4.1-mini, gpt-4.1-nano
o3, o4-mini
Endpoint: /v1/*

Anthropic

claude-opus-5, claude-sonnet-5, claude-fable-5
claude-opus-4-8, claude-opus-4-7, claude-opus-4-6
claude-sonnet-4-6, claude-sonnet-4-5
claude-haiku-4-5
Endpoint: /anthropic/*

Google

gemini-3.1-pro, gemini-3.5-flash
gemini-2.5-pro, gemini-2.5-flash
Endpoint: /google/*

Amazon Bedrock

anthropic.claude-sonnet-5 · -opus-4-8
us.anthropic.claude-sonnet-4-6 · -opus-4-*
xai.grok-4.6
amazon.nova-2-multimodal-embeddings
cross-region inference profiles
Endpoints: /bedrock/* (Converse, InvokeModel)
/bedrock-mantle/* (Anthropic Messages)

Attribution headers

SpendProxy auto-attributes costs using system prompt fingerprinting, toolset hashing, and SDK detection. You can also use optional headers for explicit control — they take priority over auto-detection.

Request headers (optional)

Header Purpose
X-CE-Route Feature or endpoint name
X-CE-Tag Custom tag (team, environment)
X-CE-Project Project name
X-CE-Org Organization or tenant the spend belongs to
X-CE-User Person or account the spend belongs to
X-CE-Feature Product feature that made the call
X-CE-Agent Specific agent, session or run

Reading the breakdown

Both endpoints below sit behind the dashboard password (CE_DASHBOARD_PASSWORD), sent as Authorization: Bearer <password>. They are operator endpoints, not part of the request path, so a key that can spend cannot read or relabel anyone else's spend.

Spend grouped by any one dimension, optionally scoped by the others. The response reports what it is not showing, so a truncated table can never be mistaken for the whole picture: totalCost covers every group, otherCost is what fell outside the limit, and unattributedCost is spend with no value for that dimension.

GET /v1/costs/breakdown?by=user&org=acme&since=2026-09-01

{ "dimension": "user",
  "rows": [{ "value": "usr_11", "label": "Dana R.", "cost": 41.20, "requests": 812 }],
  "totalGroups": 34, "totalCost": 128.55, "otherCost": 12.90,
  "unattributedCost": 3.10, "unattributedRequests": 47 }

Mapping identifiers to display names

Send opaque IDs rather than names or email addresses. Names change, and an identifier that changes splits one person's spend across two identities, so a year of cost history quietly stops adding up. Map IDs to display names separately instead. The mapping is display only and never used for grouping, so a rename applies to every past row rather than fragmenting them. It also means a name can be deleted on its own, leaving the financial record intact.

PUT /v1/attribution/labels
{ "labels": [{ "dimension": "org", "value": "org_7f3a", "label": "Acme Corp" }] }

DELETE /v1/attribution/labels?dimension=user
{ "values": ["usr_11"] }   # forgets the name, keeps the spend

Response headers (always included)

Header Value
X-CE-Request-Id Unique request ID
X-CE-Cost Total cost in USD
X-CE-Input-Tokens Input token count
X-CE-Output-Tokens Output token count
X-CE-Cached-Tokens Cached input tokens
X-CE-Model Resolved model name
X-CE-Latency-Ms Proxy overhead latency

For streaming responses, cost data is sent as a final SSE comment: ce-cost {"cost": 0.0023, ...}

Security model

Data residency

SpendProxy runs entirely in your infrastructure. The SQLite database, request logs, cost data, and attribution data all stay in your VPC. Nothing is transmitted to SpendProxy, CloudExpat, or any third-party service.

API key handling

API keys are passed through to the provider in the original request headers. SpendProxy never reads, logs, or stores your API keys. They're forwarded as-is.

Network access

SpendProxy makes outbound HTTPS requests only to the AI provider APIs (api.openai.com, api.anthropic.com, generativelanguage.googleapis.com, and bedrock-runtime.<region>.amazonaws.com for Bedrock). It does not phone home, send telemetry, check for updates, or communicate with any other external service.

What SpendProxy stores locally

  • ✓ Request metadata: model, tokens, cost, latency, timestamps
  • ✓ Attribution data: system prompt fingerprints (hashes, not content), toolset hashes, SDK identifiers
  • ✓ Optimization state: cache entries, dedup signatures, budget counters
  • × Prompt content, response content, or API keys are never stored