Documentation
Docs
Everything you need to deploy and configure SpendProxy. The integration is a one-line change — see for yourself below.
$ Using Claude Code? Meter your coding sessions — see the guide →Quick start
1. Start SpendProxy
Using Docker:
docker run -d -p 4100:4100 spendproxy/proxy:latest Or using npx:
npx @cloudexpat/spendproxy 2. Point your AI SDK at SpendProxy
Pick your provider and language. For every provider except AWS Bedrock, the highlighted diff is the entire integration, one line. API keys pass through untouched.
bedrock:InvokeModel / bedrock:Converse (and their streaming variants) for /bedrock/*, plus bedrock-mantle:CreateInference if you use /bedrock-mantle/*, and set CE_BEDROCK_REGION. Cross-region inference-profile IDs like us.anthropic.claude-sonnet-4-6 work unchanged on /bedrock/*. The Messages endpoint takes regional model IDs instead, so use anthropic.claude-sonnet-5 rather than a profile ID there. Use https:// for any non-local proxy.
One required change for the AWS SDK. It defaults to HTTP/2 for Bedrock, and SpendProxy speaks HTTP/1.1, so a stock client fails at the connection with ERR_HTTP2_ERROR: Protocol error and nothing reaches the proxy. Pass NodeHttpHandler and it works:
import { NodeHttpHandler } from '@smithy/node-http-handler';
const client = new BedrockRuntimeClient({
region: 'us-east-1',
endpoint: 'http://localhost:4100/bedrock',
requestHandler: new NodeHttpHandler(),
});
The Anthropic SDKs, including AnthropicBedrockMantle, use HTTP/1.1 already and need no change. Point Mantle at /bedrock-mantle with skipAuth: true so the proxy signs.
3. Open the dashboard
Navigate to localhost:4100/dashboard. Cost data appears in real time. The dashboard auto-refreshes every 10 seconds.
Configuration
SpendProxy is configured via environment variables.
| Variable | Default | Description |
|---|---|---|
| CE_PROXY_PORT | 4100 | Port to listen on |
| CE_DATA_DIR | ~/.spendproxy | Data directory for SQLite database and config |
| CE_DASHBOARD_PASSWORD | (none) | Password-protect the dashboard and cost APIs |
| CE_PROXY_DB | ~/.spendproxy/ce-proxy.db | SQLite database path |
Example with custom port and data directory:
docker run -d \
-p 8080:8080 \
-e CE_PROXY_PORT=8080 \
-e CE_DATA_DIR=/data \
-v /host/data:/data \
spendproxy/proxy:latest Supported providers and models
SpendProxy uses fuzzy matching for model IDs. Versioned identifiers automatically resolve to the base model. Pricing database updated continuously.
OpenAI
/v1/*Anthropic
/anthropic/*/google/*Amazon Bedrock
/bedrock/* (Converse, InvokeModel)/bedrock-mantle/* (Anthropic Messages)Attribution headers
SpendProxy auto-attributes costs using system prompt fingerprinting, toolset hashing, and SDK detection. You can also use optional headers for explicit control — they take priority over auto-detection.
Request headers (optional)
| Header | Purpose |
|---|---|
| X-CE-Route | Feature or endpoint name |
| X-CE-Tag | Custom tag (team, environment) |
| X-CE-Project | Project name |
| X-CE-Org | Organization or tenant the spend belongs to |
| X-CE-User | Person or account the spend belongs to |
| X-CE-Feature | Product feature that made the call |
| X-CE-Agent | Specific agent, session or run |
Reading the breakdown
Both endpoints below sit behind the dashboard password (CE_DASHBOARD_PASSWORD), sent as Authorization: Bearer <password>. They are operator endpoints, not part of the request path, so a key that can spend cannot read or relabel anyone else's spend.
Spend grouped by any one dimension, optionally scoped by the others. The response reports what it is not showing, so a truncated table can never be mistaken for the whole picture: totalCost covers every group, otherCost is what fell outside the limit, and unattributedCost is spend with no value for that dimension.
GET /v1/costs/breakdown?by=user&org=acme&since=2026-09-01
{ "dimension": "user",
"rows": [{ "value": "usr_11", "label": "Dana R.", "cost": 41.20, "requests": 812 }],
"totalGroups": 34, "totalCost": 128.55, "otherCost": 12.90,
"unattributedCost": 3.10, "unattributedRequests": 47 } Mapping identifiers to display names
Send opaque IDs rather than names or email addresses. Names change, and an identifier that changes splits one person's spend across two identities, so a year of cost history quietly stops adding up. Map IDs to display names separately instead. The mapping is display only and never used for grouping, so a rename applies to every past row rather than fragmenting them. It also means a name can be deleted on its own, leaving the financial record intact.
PUT /v1/attribution/labels
{ "labels": [{ "dimension": "org", "value": "org_7f3a", "label": "Acme Corp" }] }
DELETE /v1/attribution/labels?dimension=user
{ "values": ["usr_11"] } # forgets the name, keeps the spend Response headers (always included)
| Header | Value |
|---|---|
| X-CE-Request-Id | Unique request ID |
| X-CE-Cost | Total cost in USD |
| X-CE-Input-Tokens | Input token count |
| X-CE-Output-Tokens | Output token count |
| X-CE-Cached-Tokens | Cached input tokens |
| X-CE-Model | Resolved model name |
| X-CE-Latency-Ms | Proxy overhead latency |
For streaming responses, cost data is sent as a final SSE comment: ce-cost {"cost": 0.0023, ...}
Security model
Data residency
SpendProxy runs entirely in your infrastructure. The SQLite database, request logs, cost data, and attribution data all stay in your VPC. Nothing is transmitted to SpendProxy, CloudExpat, or any third-party service.
API key handling
API keys are passed through to the provider in the original request headers. SpendProxy never reads, logs, or stores your API keys. They're forwarded as-is.
Network access
SpendProxy makes outbound HTTPS requests only to the AI provider APIs (api.openai.com, api.anthropic.com, generativelanguage.googleapis.com, and bedrock-runtime.<region>.amazonaws.com for Bedrock). It does not phone home, send telemetry, check for updates, or communicate with any other external service.
What SpendProxy stores locally
- ✓ Request metadata: model, tokens, cost, latency, timestamps
- ✓ Attribution data: system prompt fingerprints (hashes, not content), toolset hashes, SDK identifiers
- ✓ Optimization state: cache entries, dedup signatures, budget counters
- × Prompt content, response content, or API keys are never stored