> ## Documentation Index
> Fetch the complete documentation index at: https://docs.plungeai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Models plane — /v1/chat/completions, /v1/embeddings, /v1/models

> The OpenAI-compatible surface ("money plane"). Auth is the same ozk key as the rest of the API, sent as Authorization: Bearer ozkYOURKEY only (from Ocean Dashboard → One API → Keys). Errors use the OpenAI error shape {"error":{"message","type","code","requestid"}}.


<!-- sources-of-truth: orchestration/api-gateway/openapi.ts, orchestration/api-gateway/routes/models-proxy.ts, orchestration/api-gateway/docs/guide-3.0/03-models.md, orchestration/api-gateway/docs/guide-3.0/05-openrouter-compat.md, inference/gateway/src/errors.ts, inference/gateway/src/routing/select.ts, inference/gateway/src/metering/response-cache.ts, apps/ocean-dashboard/src/nav.ts, apps/ocean-dashboard/src/components/sections/one-api/KeysTab.tsx, apps/ocean-dashboard/src/components/sections/one-api/CatalogTab.tsx | last-synced: 2026-10-05 (One API 2.3.0: GET /v1/models item shape, embeddings providers, the /api/v1 drop-in; earlier re-verified: error codes against errors.ts, 30s outage exclusion against select.ts, 60-86400s cache TTL clamp against response-cache.ts — all match; fixed stale UI labels — legacy (now retired) sk-ocean- keys and BYOK provider keys live under Dashboard → One API → Keys/Catalog, not "Ocean Dashboard → API keys" or "Settings → Integrations", which do not exist in nav.ts) -->
The OpenAI-compatible surface ("money plane"). Auth is the **same `ozk_` key** as
the rest of the API, sent as `Authorization: Bearer ozk_YOUR_KEY` only (from Ocean
Dashboard → One API → Keys). Errors use the OpenAI error shape
`{"error":{"message","type","code","request_id"}}`.

## Drop-in base_url swap

Point the OpenAI SDK you already use at the base URL. That is the entire
integration:

**Python**

```python
from openai import OpenAI
client = OpenAI(base_url="https://api.plungeai.com/v1",
                api_key="ozk_YOUR_KEY")
r = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",
    messages=[{"role": "user", "content": "Hello"}])
print(r.choices[0].message.content)
```

**TypeScript**

```ts
import OpenAI from 'openai'
const client = new OpenAI({
  baseURL: 'https://api.plungeai.com/v1',
  apiKey: process.env.PLUNGEAI_API_KEY, // ozk_YOUR_KEY
})
const r = await client.chat.completions.create({
  model: 'anthropic/claude-sonnet-5',
  messages: [{ role: 'user', content: 'Hello' }],
})
```

**curl**

```bash
curl -s https://api.plungeai.com/v1/chat/completions \
  -H "Authorization: Bearer ozk_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"anthropic/claude-sonnet-5","messages":[{"role":"user","content":"Hello"}]}'
```

```json
{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "model": "anthropic/claude-sonnet-5",
  "choices": [{ "index": 0,
    "message": { "role": "assistant", "content": "Hello! How can I help?" },
    "finish_reason": "stop" }],
  "usage": { "prompt_tokens": 8, "completion_tokens": 9, "total_tokens": 17 }
}
```

Model slugs are `provider/model`. Get the live priced list from
`GET /v1/models` — never hardcode a slug list from memory. The enforcement is
**404 `model_not_found`**: an unknown slug fails, so populate model choices
from `GET /v1/models` at runtime.

Streaming (`"stream": true` → `text/event-stream`), any other
OpenAI-compatible field (`tools`, `response_format`, `top_p`, …) passes
through untouched.

## POST /v1/chat/completions — request fields beyond OpenAI's

| Field | Type | Semantics |
|---|---|---|
| `model` | string | One slug, or `@preset/<slug>`. Ignored if `models` is set |
| `models` | string[] | **Ordered fallback list** — tried in order; first success serves |
| `sort` | `"price"` \| `"latency"` \| `"throughput"` | Reorders `models[]` before the first attempt. Providers with no latency/throughput history sort last |
| `temperature` | number | Exactly `0` makes the request response-cache eligible |

### Routing & failover (`models[]`)

```bash
curl -s https://api.plungeai.com/v1/chat/completions \
  -H "Authorization: Bearer ozk_YOUR_KEY" -H "Content-Type: application/json" \
  -d '{"models":["anthropic/claude-sonnet-5","openai/gpt-5"],
       "sort":"price",
       "messages":[{"role":"user","content":"Hello"}]}'
```

- Each candidate gets **one retry** (2 attempts total) on `429`/`5xx`/network
  error; backoff honors a short `Retry-After` (capped 2s) or flat 250ms.
- Exhausting a candidate excludes its **provider** for **30 seconds** and
  moves to the next; a still-excluded candidate shows
  `"reason":"outage_excluded"` in the logged `route_attempts` (not re-probed).
- A non-429 `4xx` (your request is malformed) is **not** retried and **not**
  failed over — the next candidate would fail identically.
- **Honest billing**: the response's `model` field is the slug that **actually
  served** the request — possibly a fallback, not what you asked for first.
  Bill, log, and display on the response `model`, never on the requested one.
- **Streaming caveat**: failover works identically for `stream:true` — the
  decision happens on the response status before any bytes are forwarded — but
  a stream that dies mid-flight is not retried. The `model` field inside SSE
  chunks reflects the serving upstream natively; billing still uses the honest
  slug.

## Presets — `@preset/<slug>`

A stored model+routing+params bundle, addressed via the `model` field:

```bash
curl -s https://api.plungeai.com/v1/chat/completions \
  -H "Authorization: Bearer ozk_YOUR_KEY" -H "Content-Type: application/json" \
  -d '{"model":"@preset/fast","messages":[{"role":"user","content":"Hello"}]}'
```

Request-explicit fields **always override** the preset — it only fills gaps.
Unknown/inactive slug → **404 `preset_not_found`**. Presets are administered
platform-side; ask your admin which exist for your org.

## Response cache (opt-in, per org)

- Eligible only when `temperature: 0` is **explicit** and `stream` is not true.
- Cache key = hash of the **full request body** — any differing field misses.
- Header `x-cache: miss` (stored for next time) / `x-cache: hit` (served from
  cache, **never billed**, same response id). Header absent entirely when the
  org has caching off or the request wasn't eligible.
- TTL is per-org, clamped 60s–86400s.

## Guardrails & pricing (BYOK-or-billed)

Per-org/per-key policy — spend caps, provider/model allow-lists,
content-blocking, BYOK requirements — enforced **before** cache lookup and
before any money moves. Strictest wins when org and key policies overlap.

Pricing: platform-billed by default (provider cost + org markup). Connect your
own provider key (BYOK, Dashboard → One API → Catalog → "BYOK provider keys")
and requests to that provider route through it for a small routing fee instead.
Org policy can *require* BYOK per provider. Trial orgs are metered but never
charged.

## Error catalogue

| Status | Code | Meaning | Handling |
|---|---|---|---|
| 400 | `invalid_json` | Request body is not valid JSON | Fix the body |
| 401 | `invalid_api_key` | Missing/invalid `ozk_` key | Fix the key |
| 402 | `insufficient_balance` | Credit balance at $0 — the first error every platform-billed org hits | Surface: top up (the message carries the URL) |
| 402 | `byok_required` | Provider requires your own connected key on this plan | Surface: connect key in Dashboard → One API → Catalog, or change plan |
| 403 | `model_not_allowed` | All routing candidates excluded by allow-list | Surface; pick an allowed model (see `GET /v1/models`) |
| 403 | `content_blocked` | Prompt matched a guardrail regex (pattern never echoed) | Surface; do not retry variants to probe the filter |
| 404 | `model_not_found` | Unknown model slug | Re-pick from `GET /v1/models`; never hardcode slugs |
| 404 | `preset_not_found` | `@preset/<slug>` unknown or inactive | Fix the slug |
| 429 | `rate_limit_exceeded` | Atomic per-key limiter — **no `Retry-After` on this path** | Fixed short backoff (~1s), then retry |
| 429 | `insufficient_quota` | Monthly spend limit reached (type `rate_limit_error`) — distinct from both other 429s | Surface — resets with the billing period; do NOT retry-loop |
| 429 | `spend_cap_exceeded` | Guardrail cap reached | Surface — a policy, not a transient; do NOT retry-loop |
| varies | `upstream_error` | Provider failure — the provider's own status is passed through (type `api_error`) | One retry with backoff (routing already retried/failed over) |
| 500 | `internal_error` | Gateway error | One retry with backoff |
| 503 | `money_plane_unavailable` | Inference gateway not bound on this tier | Surface; wrong tier/deployment |

Example 402:

```json
{"error":{"message":"This provider requires your own API key on your plan — connect it in Settings → Integrations, or upgrade to platform-billed usage","type":"byok_required","code":"byok_required","request_id":"00000000-0000-4000-8000-000000000004"}}
```

## POST /v1/embeddings

Same auth, same guardrail enforcement as chat completions.

```bash
curl -s https://api.plungeai.com/v1/embeddings \
  -H "Authorization: Bearer ozk_YOUR_KEY" -H "Content-Type: application/json" \
  -d '{"model":"openai/text-embedding-3-small","input":["hello world"]}'
```

```json
{
  "object": "list",
  "data": [{ "object": "embedding", "index": 0, "embedding": [0.0123, -0.0456, "..."] }],
  "model": "openai/text-embedding-3-small",
  "usage": { "prompt_tokens": 2, "total_tokens": 2 }
}
```

`input` is a string or an array of strings; `encoding_format` is `float` or `base64`. The plane serves openai,
mistral, google and cohere embedding models — any other provider is `400 unsupported_provider`.

## GET /v1/models — the priced catalog

The live list routing candidates are drawn from. Needs the key (`/api/v1/models` is the public one).

```bash
curl -s https://api.plungeai.com/v1/models \
  -H "Authorization: Bearer ozk_YOUR_KEY"
```

```json
{
  "object": "list",
  "data": [
    { "id": "anthropic/claude-sonnet-5", "object": "model", "owned_by": "anthropic",
      "context_length": 200000, "max_output": 8192,
      "pricing": { "in": 3.0, "out": 15.0 },
      "capabilities": { "streaming": true, "tools": true, "vision": true, "embedding": false, "fim": false },
      "pricing_per_token": { "prompt": "0.000003", "completion": "0.000015" } }
  ]
}
```

`pricing.in` / `.out` are USD per **1M** tokens; `pricing_per_token` is the same price per token as decimal
strings. `vision: false` means not verified, `max_output: null` means unknown. Generated code that lets users
pick a model should populate the choice from this endpoint at runtime, not from a baked-in list. Registry
model cards (`plungeai_list_agents {kind: "models"}`, `GET /v1/discovery/search?kind=models`) name the same
models as `provider:model`; the plane takes `provider/model`.

## Account, accounting and legacy routes

Same `ozk_` key, `Authorization: Bearer` only.

| Route | Returns |
|---|---|
| `GET /v1/models/{author}/{slug}/endpoints` | The endpoint serving one model, with per-token pricing (one upstream per model); `404 model_not_found` |
| `GET /v1/generation?id=<gen-… id or x-request-id>` | Post-hoc usage and cost of one completion; `404 not_found` for another account's id |
| `GET /v1/credits` | `{data: {total_credits, total_usage}}` — lifetime credits bought and used, USD |
| `GET /v1/key` | The calling key: label, monthly limit, usage this month, rate limit |
| `POST /v1/completions` | Legacy text completion: `{model, prompt}` (one string; an array is `400 unsupported_parameter`) answered as `text_completion` — prefer `/v1/chat/completions` |

## The `/api/v1` drop-in (OpenRouter wire shape)

Point an OpenRouter client at `https://api.plungeai.com/api/v1` with the `ozk_` key: the same plane (routing,
guardrails, cache, billing) under the same eight routes — `POST /api/v1/chat/completions`, `/embeddings`,
`/completions`; `GET /api/v1/models`, `/generation`, `/credits`, `/key`, `/models/{author}/{slug}/endpoints`.
Only the shape differs:

- Errors: `{"error":{"code":401,"message":…,"metadata":{"error_type","ocean_code","request_id"}}}` — a numeric
  `code`, the Ocean code in `metadata.ocean_code`.
- The completion `id` is `gen-…` (also the `x-generation-id` header); `GET /v1/generation?id=` takes it.
- `GET /api/v1/models` is public and lists OpenRouter's item keys (`pricing` as USD-per-token strings); the
  `/v1/models` shape above needs a key.
- The stream sends `: OPENROUTER PROCESSING` until the first byte and always ends with a `usage` chunk.
- OpenRouter-only request fields (`provider`, `route`, `transforms`, `plugins`) are accepted and ignored: route
  with `models[]`, `sort` and presets. Cost fields (`usage.cost`, `cost_details`, `is_byok`) appear with
  `usage: {"include": true}`; the non-stream price is also the `x-ocean-cost-usd` header.
