Point any OpenAI or Anthropic SDK at Nozzle and name a model. We route the call to wherever that model is served: our provider accounts, shared open-weight GPUs, your own provider key, or a GPU that is yours alone. Your code stays the same either way, and every call comes back with its exact cost.
The OpenAI-shaped surface lives under /v1, the Anthropic drop-in under /anthropic. Same key either way, and every reply tells you what it cost.
X-Nozzle-Cost-Micro-Cents on every buffered reply.curl https://api.opennozzle.com/v1/chat/completions \
-H "Authorization: Bearer pk_live_…" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-4.1-mini", "messages": [{"role": "user", "content": "Hello"}]}'from openai import OpenAI
client = OpenAI(api_key="pk_live_…", base_url="https://api.opennozzle.com/v1")
client.chat.completions.create(model="claude-haiku-4-5",
messages=[{"role": "user", "content": "Hello"}])Whose credential pays the upstream, and who gets the capacity. Move a model between modes chasing price, latency or data residency without touching a line of caller code.
Every closed model without an account to manage.
Open weights served once, called by everyone.
When the tokens must land on your invoice.
A machine nobody else touches.
Nozzle does not mark up GPU time. You pay for platform access; tokens and GPU-hours pass through at the cheapest rate we can source.