<!-- This file is the Markdown twin of the English /docs page (docs.en.html), for developers to download and feed to AI tools.
     Keep it in sync with docs.en.html, and keep both in sync with the Chinese docs.md / docs.html —
     they are four carriers of the same developer guide. -->

# Build with AI-Pay — Developer Guide

> Source of this document: `https://ai-pay.coswic.ai/docs.md?lang=en` (web edition: `https://ai-pay.coswic.ai/docs?lang=en`). Re-download to get the latest version; the current SDK and CLI versions are whatever npm publishes as `@coswic/aipay-sdk` and `@coswic/aipay`.

AI-Pay lets your app call models with **the user's own AI credits**: through "Sign in with AI-Pay", the user grants your app an **"AI credit authorization"**, and within the spending limits they set, every AI request settles directly against the user's AI-Pay wallet.

- **You never touch cards or front the cost** — payment, credit and risk control all live in AI-Pay; your app has no model bill.
- **Users stay in control** — pausing, revoking and changing limits immediately stop new AI usage (requests already dispatched may still settle under their original rules); your app must handle the corresponding 403s.
- **The receipt travels with the response** — every successful response carries a receipt (the actual charge and its split); every request can be reconciled.

> Current status in production: **free beta**. Only invited users can create an account and use the service; public registration is not open. **Your users must also be invited before their first sign-in through your app's "Sign in with AI-Pay"**; a person with no AI-Pay account and no invitation stays on the AI-Pay sign-in page and cannot complete the authorization. Test credit depends on the invitation; production does not accept top-ups at the moment, so users cannot add credit once it is used up. Invitation requirements, test credit and feature status are on [`/about`](https://ai-pay.coswic.ai/about?lang=en#status). In other environments, features and payment methods follow that environment's screens; Stripe / PayPal test modes do not make real charges, and test credits are not real payments or developer earnings.
> In the examples below the API base is `https://ai-pay.coswic.ai` and the authorization endpoint is `https://auth.ai-pay.coswic.ai` — **both values are filled in by the environment you fetched this document from**, so you can copy them as they are.

Amounts are always shown in USD and transmitted as integer µUSD strings ($1 = 1,000,000 µUSD).

---

## Quick start

Before you start, sign in to the [Dashboard](https://ai-pay.coswic.ai/dashboard?lang=en) with your own Google account, complete the country/region details and current Terms of Service confirmation shown there, then open the Console. Creating an account on first sign-in requires a valid invitation; installing the SDK or CLI does not bypass these prerequisites. If the API returns `region_required`, finish setup in the Dashboard or use the `aipay region` command below.

1. In the Console (app control panel), click "Create app" at the top right: just enter a name. The local callback `http://127.0.0.1:{Port}/auth/aipay/callback` is registered automatically (the port of `127.0.0.1` is not compared, so any local port works; add your production HTTPS URL later with the CLI callback command below) and you get your `client_id`. AI-Pay uses a public client with PKCE S256 — **there is no client secret**. After creation the Console hands you a prompt for your coding agent, or run the CLI yourself — next section.
2. Get the SDK (TypeScript, one package for server and browser): `npm i @coswic/aipay-sdk` (published on public npm). Browser rules are in the "Serverless apps" section.
3. Send the user to the authorization URL (next section); the user signs in to AI-Pay, sets limits, consents, and returns to your callback.
4. Exchange the code for tokens and call `POST /v1/chat/completions` to send your first request.
5. Reconcile under the Console's "Requests" tab (every request_id is searchable). App service fees and payouts are not open yet (see "App pricing (service fees)" and "Usage reconciliation and payouts"); for now, the end of the integration path is a successful response with a verifiable receipt.

---

## Connect with the CLI or a coding agent (recommended)

After creating the app, the Console gives you a prompt you can paste straight into Claude Code, Cursor, Codex or a similar tool; the AI reads this guide and completes the integration with your app's settings. The prompt contains no tokens or keys. You can also run the same CLI yourself (`@coswic/aipay`, public npm):

```
npx @coswic/aipay@latest login --api https://ai-pay.coswic.ai
npx @coswic/aipay@latest init --app <client_id> --auth https://auth.ai-pay.coswic.ai --api https://ai-pay.coswic.ai
npx @coswic/aipay@latest test
```

Developer CLI sign-in manages your Apps; it is separate from a user authorizing AI spending. On a remote machine use `login --device` and ask the developer to approve the browser prompt. Before deploying, run `aipay apps callbacks add --app <client_id> --callback https://yourapp.example/auth/aipay/callback`; this preserves the existing local callback. Use `apps callbacks list` to inspect and `apps callbacks remove --callback <url>` to remove one. In a linked project, `--app` may be omitted. `init --callback` only writes local configuration. A concurrent edit or unknown outcome requires readback; the CLI does not blindly repeat a write. Callback changes require a server with ETag support.

`init` does this in your project (never overwrites existing files; rerunning is idempotent): detects the package manager and framework (Next.js / Express / Fastify / Hono / Node, or a Vite browser SPA; `--framework` overrides it; TypeScript or JavaScript), writes `.aipay/project.json` (commit it; no secrets), `.env.local` (four public values `AIPAY_CLIENT_ID` / `AIPAY_AUTH_ORIGIN` / `AIPAY_API_BASE` / `AIPAY_CALLBACK_URL`, and adds it to `.gitignore`), a start/callback route pair for your framework, `aipay-agent.md` (project facts plus this whole guide, for agents to read offline), and finally installs `@coswic/aipay-sdk`.

`test` is the only end-to-end proof: the CLI receives the callback on any free port of `127.0.0.1` (AI-Pay does not compare the port for loopback, so your dev server can keep running), opens a browser for you to sign in and authorize with your own account, exchanges the code, sends one chat completion on your behalf, and prints the receipt (`request_id`, amount). The green checks under "Verify" in the Console come from server-side facts, not from the CLI's own report.

| Command | Purpose |
|---|---|
| `aipay init --app <client_id> [--auth] [--api] [--callback] [--dry-run] [--no-install]` | Link the project and scaffold the integration; `--auth` / `--api` are optional once you are signed in |
| `aipay test [--model] [--port] [--no-open]` | Real sign-in + first request + receipt |
| `aipay doctor [--offline]` | Checks Node, the link, `.env.local`, the SDK, and reachability of the authorization server and API |
| `aipay status` / `aipay unlink` | Show which app is linked / forget the link (keeps your code) |
| `aipay login --api https://ai-pay.coswic.ai` / `whoami` / `logout` | Developer sign-in for the CLI itself (loopback PKCE, audience `aipay-developer-api`; on a remote or browserless machine add `--device`: it prints a URL and a short code you enter in a browser on any device); tokens live in the OS keychain (macOS Keychain / Linux keyring / a DPAPI-encrypted file on Windows), with a 0600 file only when none is available |
| `aipay apps list` / `aipay apps create --name <name>` / `aipay init --create <name>` | Create apps from the CLI once signed in (sends an `Idempotency-Key`, so a rerun never creates a second app) and link them directly |
| `aipay region` / `aipay region set --country <CC> [--state US-XX --line1 <street> --city <city> --zip <ZIP>] --accept-terms` | Show or set your account's country/region and accept the current Terms of Service; US accounts also need the full postal address (the backend locates the sales-tax jurisdiction — state / county / city / special district — and `aipay region` shows it); until everything is on file the developer API answers 403 `region_required` (`aipay region` prints the version to accept and the terms URL) |

When creating an app directly with `POST /v1/developer/apps`, send an `Idempotency-Key` whenever possible; the name and Redirect URI list must remain the same for that key. Creation returns 201; retrieving the completed app with the same key returns 200. For 409 `provisioning_in_progress`, or 502 `oauth_provider_error` with `details.outcome: 'unknown'`, wait and retry with the same key. A 502 with `outcome: 'cancelled'` means this creation was cancelled; you may reuse the key to create again. If you explicitly deleted a successfully created app, its original key returns 409 `idempotency_key_consumed`; use a new key for a new app. If creation without a key is interrupted and remains incomplete after 15 minutes, the system schedules automatic cleanup.

Every command accepts `--json` (one JSON document on stdout; the authorization URL goes to stderr) and `--dir`; exit codes are 0 ok, 1 fixable locally, 2 usage, 3 network or auth. For connection issues, see [Connection and error troubleshooting](#network-access).

When you are signed in, `init` and `test` report their results to the Console (framework, files written, receipt id), and the "Hand it to your AI" card on the app page turns from "Waiting for the tool" into "The CLI finished init". To revoke that connection: Console → "Connected apps" → "AI-Pay CLI" → Revoke authorization; the CLI's next call will then ask you to sign in again.

---

## Sign in with AI-Pay (OAuth 2.0 + PKCE)

AI-Pay can be your sign-in service directly (OIDC): the user signs in to AI-Pay with Google or email, and your
app receives a stable user identifier (`sub`) and email — no other identity provider to integrate, no passwords to handle.
For sign-in only, request just the identity scopes (`openid email`) — this creates a **"sign-in connection"**; to use AI, add the spending scopes and the same authorization creates an **"AI credit authorization"**.

`sub` is **a pseudonym specific to your app**: the same user always gets the same `sub` in your app (safe to use as the
account key), but a different value in other apps — apps cannot correlate a user through the identifier (a privacy
design, like Sign in with Apple). If you also request email, the email itself can be correlated across services; it is
data the user explicitly authorized you to receive.

The SDK builds the authorization URL in one line (state and the PKCE code_verifier are generated for you):

```ts
import { createConnectUrl, getUserInfo, handleCallback, refreshAccessToken } from '@coswic/aipay-sdk'

const cfg = {
  hydraPublicUrl: 'https://auth.ai-pay.coswic.ai',
  clientId: '<the client_id the Console gave you>',
  redirectUri: 'http://127.0.0.1:4567/auth/aipay/callback',
  // Sign-in-only app: scopes: ['openid', 'email'] (the consent screen then shows no spending limits)
  // offline_access is not included by default — add it to scopes explicitly only if your server app must renew while the user is away
}

// 1. Build the authorization URL — store { state, codeVerifier } in your session (browser = sessionStorage)
const { url, state, codeVerifier } = await createConnectUrl(cfg)
// 2. The user signs in to AI-Pay, sets limits (only with spending scopes), consents → returns to your callback (?code=…&state=…)
// 3. Callback: handleCallback does "check error → verify state → exchange the token" in one step —
//    you never write state verification yourself (hand-written checks get skipped, and a skipped check = a CSRF / code-injection surface);
//    a user pressing "Deny" or a lost session both throw a ConnectError with a clear message
let tokens = await handleCallback(cfg, callbackUrl, { state, codeVerifier })
// tokens.accessToken (15 minutes); tokens.refreshToken (null without offline_access)
// tokens.idToken: the OIDC id_token (only when openid was granted)
// Renew only when a server app explicitly requested and received offline_access; browser apps do not request it
if (tokens.refreshToken) {
  tokens = await refreshAccessToken(cfg, tokens.refreshToken)
  // Save the updated tokens in the server session; rotation invalidates the old refresh token
}

// 4. Fetch the user's identity (verified server-side; no need to verify the id_token signature yourself)
const user = await getUserInfo(cfg, tokens.accessToken)
// user.sub = a stable pseudonym specific to your app — use it as the account key in your database (email can change; never key on it)
// user.email / user.emailVerified / user.name depend on which scopes were granted
```

You can also skip the SDK — it is the standard OAuth 2.0 authorization code flow. Parameters of the authorization endpoint `GET /oauth2/auth`:

| Parameter | Value |
|---|---|
| `client_id` | The client_id the Console gave you |
| `response_type` | `code` |
| `scope` | Sign-in: `openid email`; to use AI add `ai.chat.create wallet.charge`; for long-term access add `offline_access` to get a refresh token |
| `audience` | `aipay-api` (required — the token audience) |
| `redirect_uri` | Exact match with the registered value |
| `state` | Your CSRF random value; must be verified at the callback |
| `code_challenge` / `code_challenge_method` | PKCE; the method is always `S256` |

Token exchange: `POST /oauth2/token` (`grant_type=authorization_code` + `code_verifier`; renew with `grant_type=refresh_token`). Access tokens live 15 minutes — refreshing once on a 401 and retrying is a normal event, not an error.

Requestable scopes (what the user sees on the consent screen is the plain-language version of these):

| scope | What the user sees |
|---|---|
| `openid` | Sign in with the AI-Pay identity (the `sub` in the id_token / userinfo — a pseudonym specific to your app) |
| `email` | Provide the email address (with `email_verified` — true only if it was once verified through Google) |
| `profile` | Provide the display name |
| `ai.chat.create` | Create AI chat requests (text) |
| `wallet.charge` | Deduct the user's credits by actual usage |
| `offline_access` | Renew the authorization while the user is away (long-term access with a refresh token) |

> The consent screen is provided by AI-Pay: with spending scopes (AI credit authorization) the user sets **monthly / daily / per-request limits** there and sees how your app charges; with identity scopes only (sign-in connection) it is a lightweight two-step sign-in authorization with no money UI at all. Granted scope = everything you requested — to lower the user's hesitation, request fewer scopes.
>
> **Return visits and renewed consent**: on return visits, users must first complete account region information and acceptance of the current Terms of Service. Repeat authorization consent can be skipped only while the account, App and existing authorization remain active, the scopes match exactly, the consent-document fingerprint is current and App fees stay within accepted ceilings. Spending authorizations also need a budget and consent covering current platform pricing policies and rate ceilings. Scope changes, Terms or authorization-document updates, unaccepted pricing policies, rates above accepted ceilings or missing required records can require confirmation again. Account confirmation is not skipped.
>
> Users can revoke the authorization in AI-Pay at any time. Revocation blocks new AI usage under that authorization and initiates credential revocation; requests already dispatched may still settle. Handle a 401 / `ConnectError` from `getUserInfo` by asking the user to connect again. Revocation does not automatically end your App’s own sessions or delete data it already obtained; your App must handle its own sessions and data obligations.

---

## Redirect URI (where the authorization result is delivered)

After authorization, AI-Pay delivers the one-time authorization code (`code`) through a browser redirect — to the redirect URI you entered when registering the app. It is an **allow-list**: only registered URLs receive the authorization result, matched character for character (scheme, host, port, path, query; the only exception: the port of `http://127.0.0.1` may vary, see the table below). You can register several (one per line), so development and production coexist.

Every kind of app uses the same mechanism; there are no special cases:

| Your app | redirect URI | Notes |
|---|---|---|
| Website with a server | `https://yourapp.com/auth/callback` | The callback is handled by your server |
| Front-end only (SPA / static site) | `https://yourapp.com/callback` | The callback is a JS page that exchanges the token directly in the browser (PKCE needs no secret) |
| Local development | `http://127.0.0.1:<port>/callback` | http is allowed only for the local hostnames `127.0.0.1` and `localhost` (RFC 8252). **Only the port of `127.0.0.1` is excluded from matching**, so changing ports needs no re-registration; `http://localhost:<port>/…` can be registered too, but the whole string (port included) must match exactly; IPv6 `[::1]` cannot be registered yet |
| iOS app | `https://yourapp.com/auth/aipay` (Universal Link) | An https URL on your domain configured as a Universal Link; iOS verifies domain ownership and opens it back into your app — to AI-Pay it is a plain https URL, zero special cases |
| Android app | Same as above (App Links) | Same mechanism: the OS opens the https URL back into the app |
| CLI / desktop tool | `http://127.0.0.1:<port>/callback` | The program opens a temporary local port to receive the code (use `127.0.0.1` to get the variable port) |

Custom schemes (`myapp://callback`) are **not accepted**: any app can claim the same scheme, so the authorization code could be intercepted by another app on the same device; Universal Links / App Links are verified by the OS against domain ownership and have no such weakness. This is the industry-standard practice of OAuth 2.1 / RFC 8252 (Apple, Google and GitHub do the same).

**What if someone registers a malicious address?** It does not affect you or your users. A redirect URI only decides where "that app's own authorization result" goes; the authorization code is bound to the client that started the flow and to its PKCE verifier, so another app cannot steal your code. The most a malicious developer can achieve is getting users to authorize "their own app" — and on the consent screen the user sees exactly that app's name and its charges, spending is bounded by the limits the user sets, and it can be revoked at any time.

---

## Serverless (front-end only) apps

With or without your own server, the flow is **exactly the same** — the only difference is who keeps the token and who sends the requests. A front-end-only app (static site / SPA) can run the whole flow in the browser: PKCE never needed a client secret, and both the token endpoint and the AI endpoints allow CORS. Use the SDK's `getUserInfo` to obtain the user's `sub` / `email`.

The SDK works directly in the browser (it uses only WebCrypto and fetch) — **the same functions** as the server version:

```js
import { createConnectUrl, handleCallback, AIPayClient } from '@coswic/aipay-sdk'

// Start authorization (for example a "Connect AI-Pay" button): store { state, codeVerifier } in sessionStorage
const { url, state, codeVerifier } = await createConnectUrl(cfg)
sessionStorage.setItem('aipay_flow', JSON.stringify({ state, codeVerifier }))
location.href = url

// Callback page: handleCallback does "check error → verify state → exchange the token" in one step → stream (all in the browser)
const saved = JSON.parse(sessionStorage.getItem('aipay_flow'))
sessionStorage.removeItem('aipay_flow') // single-use: clear once consumed
const tokens = await handleCallback(cfg, location.href, saved)
const client = new AIPayClient({ baseUrl: AIPAY_API_URL, accessToken: tokens.accessToken })
const stream = await client.chat.completions.create({ model, messages, stream: true })
for await (const ev of stream) { /* assemble content from delta frames; the receipt frame comes last (fields: see the receipt table under "Calling models") */ }
```

You can also skip the SDK — fetch the same endpoints directly (authorization parameters: see the table above). SSE frame format: each frame is `data: <JSON>`; the normal sequence is delta frames → the receipt frame (with `usage` + `aipay`) → `data: [DONE]`. Errors appear as **frames** inside the 200 stream (`{ "error": { code, message, request_id, details } }`), after which the stream may end immediately or may still append a `[DONE]` — **the only financial final state is the receipt frame**; `[DONE]` is merely a transport marker. Frame order and final-state rules are in the "Calling models" section.

Browser integration guidelines:

- **Do not request `offline_access`**: a refresh token in the browser is a long-lived credential exposed to XSS. Access tokens live only 15 minutes; when one expires, let the user go through authorization again — repeat consent is skipped only when the return-visit checks described above pass.
- **Keep tokens in memory (variables)**; avoid persisting them in localStorage.
- The worst case is bounded: a token can only spend within the limits the user set, and the user can revoke at any time — this is a platform design guarantee.
- CORS is open only for the AI endpoints (`/v1/chat/completions`, `/v1/models`) and the token endpoint; the Dashboard API is a cookie-session surface and is not open cross-origin.

---

## Calling models

Available models and live prices: `GET /v1/models` (no Bearer token, OpenAI-compatible, CORS open) — replace `model` in the example below with an id whose availability.status is available.

```ts
import { AIPayClient } from '@coswic/aipay-sdk'

const client = new AIPayClient({
  baseUrl: 'https://ai-pay.coswic.ai',
  accessToken: tokens.accessToken,
})

const res = await client.chat.completions.create({
  model: 'openai/gpt-oss-20b',   // example; check GET /v1/models for the real list
  messages: [{ role: 'user', content: '…' }],
  max_tokens: 200,
})
res.choices[0].message.content
res.aipay.cost_uusd   // the actual charge for this request (integer µUSD string; $1 = 1,000,000 µUSD)
res.aipay.breakdown   // split between provider / platform / your service fee

// Streaming (SSE)
const stream = await client.chat.completions.create({ model, messages, stream: true })
for await (const ev of stream) {
  if (ev.type === 'delta') process.stdout.write(ev.content)
  if (ev.type === 'receipt') console.log('charged', ev.aipay.cost_uusd, 'µUSD')
}
```

Raw HTTP: `POST /v1/chat/completions` with the headers `Authorization: Bearer <access token>` and `Idempotency-Key` (**required**, 1–200 characters); the body is OpenAI-compatible but **only four fields are accepted**. Available models and rates: `GET /v1/models` (no Bearer token) — the price is frozen as a snapshot at request time, so later price changes never affect requests already sent.

### Strict rules for the request body

- The body accepts only `model` / `messages` / `max_tokens` / `stream`. **Any other field** (`temperature`, `tools`, `response_format`, `user`…) is a 400 `invalid_request` (the message names the field) — nothing is silently ignored, because a parameter that did not take effect must not look as if it did.
- `messages`: 1–100 entries; each entry accepts only `role` and `content` (any extra field is also a 400). `role` is limited to `system` / `user` / `assistant`; `content` must be a **non-empty string** — OpenAI content-part arrays are not accepted (multimodal input such as images is not sold yet, see "Model catalog and pricing"). All `content` combined is limited to 100,000 characters.
- `max_tokens`: a positive integer. Above the model's `aipay.hard_max_output_tokens` it is **silently clamped to the hard limit** (no error; per-model values are in `GET /v1/models`); omitted = the model's `aipay.default_max_output_tokens`.
- `stream`: boolean; `true` on a model that does not support streaming returns 400 `invalid_request` before any money moves.
- Optional header `Replay-Key`: 24–200 characters with at least 12 distinct characters; a malformed value returns 400 (see "Retry discipline and the two keys").

The receipt frame is sent only **after settlement completes**: seeing `receipt` in the stream = the money has been charged. An interrupted stream never charges twice — after a disconnect the request settles on what actually happened, and resending with the same pair of keys retrieves the outcome (see the next section).

The response's `choices[0].finish_reason` is `stop`, `length`, or `content_filter`; missing or other upstream values default to `stop`. Non-streaming responses, streaming receipts, and replay capsules use the same value. A reasoning model can spend its output budget on reasoning, returning empty `content` with `length`; complete, trustworthy upstream usage is still billed.

Reasoning text is not returned or stored in capsules. Valid reasoning chunks keep the stream's first-event and idle timers alive, but never extend its total deadline. Heartbeats and empty chunks do not extend these timers. An interruption after reasoning is still an unknown outcome even without visible text; query the outcome with the same pair of keys. Providers that emit no progress or stay silent for too long can still time out. Execution limits deduct time already spent opening the hold and waiting for dispatch, reserving time for settlement; follow the response retryable/outcome flags.

Some models expose `aipay.input_token_cap` in `/v1/models`. The conservative estimate is the total UTF-8 byte length of message contents plus 8 per message. It must be **below** the ceiling; reaching it returns `400 invalid_request` before reserving funds. The existing 100,000-character limit also applies. AI-Pay may configure a model's reasoning effort; client-supplied `reasoning` and timeout parameters remain unsupported.

Approved models may use longer first-event or non-streaming waiting preferences, bounded by 120 and 300 seconds respectively. These are maximum preferences, not guaranteed execution times: operator limits, the stream total deadline and the hold's settlement reserve may shorten them. Models without settings keep the existing defaults, and prolonged silence can still time out. A reported non-default upstream service tier is recorded as a route incident; user charges still use the request's frozen model prices.

### SSE frame order and final states

- Normal sequence: delta frames (`object: 'chat.completion.chunk'`, text in `choices[0].delta.content`) → the **receipt frame** (also a chunk, with the completion reason above, `usage` + `aipay`) → `data: [DONE]`. Without visible text, the receipt can be the first frame.
- **Error frames**: an error that happens after HTTP 200 was already sent and the stream is open is expressed as one frame `data: {"error":{"code","message","request_id","details"}}`, with the same meaning as in the error-code table: `provider_unavailable` (`charged: false`; `retryable: true` = the upstream never ran, retry with the same key; `retryable: false` = the upstream ran but failed, retry with a **new** key), `provider_timeout` (`retryable: true` when execution budget expired before dispatch and release was confirmed; reuse the same key), or `provider_timeout` / `usage_unknown` (`outcome: 'unknown'` — resend with the same pair of keys to learn the outcome), `internal_error`. The SDK throws the corresponding exception when it receives an error frame.
- **`[DONE]` is not a financial final state**: after an upstream failure or timeout mid-stream the stream simply ends, **without** `[DONE]`; whereas after a `usage_unknown` error frame ("content was sent but could not be priced") a `[DONE]` **is still** appended. If an error frame was parsed, follow its `details.retryable`/`details.outcome` first (for example, confirmed release before execution permits retry with the same key). Otherwise, **a receipt frame received = settled; a stream that ends without a receipt frame (with or without `[DONE]`) = outcome unknown** — resend with the same pair of keys to learn the outcome (the SDK raises `StreamTruncatedError`).
- A streaming request that fails before the SSE opens (401 / 403 / 422 / 429 / 503…) gets the ordinary JSON error envelope; resending an already-settled streaming request with the same key also returns a one-off JSON body (the receipt, or the original text when a Replay-Key was used) — the SSE is not replayed.

### Receipt fields (`aipay`; the same object in non-streaming responses and in the streaming receipt frame)

| Field | Meaning |
| --- | --- |
| `usage_event_id` / `request_id` | The id of this usage event (searchable under the Console "Requests" tab; the `chatcmpl_` suffix of the response `id` is the same value) / the request_id of this HTTP request. |
| `hold_id` | The credit-reservation reference for this request — the id of the per-request hold; on the credit-lease fast path it is the lease ticket id (same uuid shape, same meaning). |
| `amount_max_uusd` | The hold ceiling frozen at request time (the estimate cap). The actual charge is always ≤ it. |
| `total_charged_uusd` / `cost_uusd` | The total actual charge (identical; `cost_uusd` is a compatibility alias). |
| `reference_base_uusd` / `aipay_markup_uusd` / `model_charge_uusd` / `app_service_fee_uusd` / `usage_tax_uusd` | The **actually collected** amount per component: the reference base (frozen upstream reference price × usage) / the applicable platform markup / the model charge (= the sum of the previous two, i.e. the catalog price) / your service fee / usage tax (currently always 0). `model_charge_uusd = reference_base_uusd + aipay_markup_uusd`; `total_charged_uusd = model_charge_uusd + app_service_fee_uusd + usage_tax_uusd`. Do not add the model charge and its components twice. |
| `breakdown.provider_cost_uusd` / `platform_fee_uusd` / `app_margin_uusd` | Three compatibility fields, equal to `reference_base_uusd` / `aipay_markup_uusd` / `app_service_fee_uusd` respectively. Note that **`provider_cost_uusd` is the reference base used for pricing, not the upstream's actual cost for this request** (the actual cost may differ). |
| `usage` | Billed usage `{ in_tok, out_tok }` (same source as the top-level `usage.prompt_tokens` / `completion_tokens`). |
| `settlement_state` / `clamped` | Normally `'settled'` / `false`. If the amount computed from actual usage exceeds the hold ceiling `amount_max_uusd`, only the ceiling is charged: `clamped: true`, `settlement_state: 'settled_capped_loss'`, plus `computed` (the components and `total_uusd` that were contractually due) and `component_shortfall` (the uncollected difference per component, `total_uusd` being the total; the user is not charged the shortfall afterwards. Uncollected reference-base and platform-markup amounts are recorded as platform-absorbed loss. Uncollected App fees are borne by the developer and do not create a payable for that portion; the platform does not absorb the entire shortfall). |
| `replayed` / `note` | `replayed: true` = this is a receipt / original text retrieved by resending with the same key, not a new execution. `note` appears only in special situations (receipt-only replay, 0 charge after a hold expired, etc.). |
| `pricing_policy` / `pricing_semantics_version` / `markup_bps` / `catalog_pricing_version` / `price_snapshot_hash` / `reference_fetched_at` / `reference_valid_until` / `reference_prices_per_million_uusd` / `retail_prices_per_million_uusd` | Traceability fields of the frozen price snapshot — this charge can be recomputed from the receipt alone even after the old catalog is retired. |

---

<a id="network-access"></a>

## Connection and error troubleshooting

Use the API and authorization endpoints provided in this guide for normal sign-in, app registration, and OAuth authorization. The browser, CLI, and application server each need to reach the relevant HTTPS endpoints.

- **Browsers and front-end-only apps**: Use PKCE, token exchange, and `getUserInfo`. Check the API and authorization endpoints and the callback URL registered in the Console.
- **CLI and application servers**: Check the configured endpoints and HTTPS connectivity, then use OAuth tokens for API calls. With remote `aipay login --device`, both the CLI and the browser approving sign-in need to reach the service.
- **Model catalog and prices**: `GET /v1/models` and `https://ai-pay.coswic.ai/models` provide the live list and prices. The catalog needs no Bearer token; AI requests still require user authorization and available credits.
- **An empty 404 or connection failure**: Check the full URL, HTTP method, and selected environment, then DNS and TLS connectivity. An empty 404 alone cannot identify the cause. If the response includes an error code, follow "Error codes"; note the HTTP status and any available request_id for tracing.

Normal account, OAuth authorization, revocation, budget, and charge checks still apply.

---

## Model catalog and pricing

**Catalog today**: dozens of text-only models (the Qwen, DeepSeek, Llama, Mistral, Cohere, GPT-OSS, Kimi,
MiniMax, Hermes, Phi, Gemma, GPT-3.5/4 families, among others). The live list and prices: people read the price page
`https://ai-pay.coswic.ai/models`, programs read `GET /v1/models` (no Bearer token) — these two are the only authoritative sources. Never hard-code the list or the prices
in your program (and do not copy the example prices on this page — the catalog changes as upstream models come and go and reprice; hard-coded numbers are guaranteed to go stale).

Model availability and prices come from `/v1/models`; models outside that catalog are not accepted. **Not offered yet**: image/audio input and models requiring pricing dimensions not yet covered by AI-Pay. Newly supported models appear after catalog review.

### How one request is priced

The formula below applies to `openrouter_reference_markup_v1`. The catalog can also contain legacy policies; use each model’s policy and retail prices instead of assuming 12%. The consent screen discloses platform rate ceilings; unaccepted pricing policies or platform rates above accepted ceilings require the user’s consent again.

```text
catalog price  = ceil( Σ(tokens × retail unit price) ÷ 1,000,000 )      ← integer µUSD, rounded up once
retail unit price = ceil( reference unit price × (10000 + markup_bps) ÷ 10000 ) ← v3 policy
what the user pays = catalog price + your per-request service fee + your usage surcharge
```

- **The reference unit price** (`reference_prices_per_million_uusd`) is the upstream route reference price frozen when reviewed, not the actual provider cost of each request. Actual costs may differ; the user is charged using the request’s frozen retail prices.
- **The platform markup follows each model’s `markup_bps`**; for example, `1200` means 12%. It is already included in the v3 retail price. Legacy fee and margin components are also included in their retail price and must not be added again.
- Rounding happens once (on the sum over all token dimensions), not once per dimension.

Example (assuming 12% markup, not a live quote; `openai/gpt-oss-20b`, reference price in $0.03 / M, out $0.14 / M → retail in $0.0336 / M,
out $0.1568 / M): one request with 1,000 input + 500 output tokens

```text
reference cost = (1000 × 30000 + 500 × 140000) ÷ 1e6 = 100 µUSD
catalog price  = (1000 × 33600 + 500 × 156800) ÷ 1e6 = 112 µUSD   ← this is what the user pays
                                    of which AI-Pay markup = 12 µUSD
```

The receipt (`res.aipay.breakdown`) splits this per component so you can trace it back to this formula.

### Price freezing and freshness

- The price is frozen as a snapshot **at the moment the hold is created**; later price changes **do not affect** requests already sent.
- Every model's price has a freshness deadline (`valid_until`). AI-Pay renews it automatically from upstream; **if it cannot, new requests for that model are blocked**
  — new requests get `503 pricing_stale`, other models are unaffected. This is deliberate: better to pause sales temporarily than to charge your
  users on a stale price.
- Upstream price changes **do not** take effect automatically. A price change goes through manual approval and is published as a new version; old versions are kept forever
  (historical requests always trace back to the real price of their time).

### Pricing fields in `GET /v1/models`

| Field | Meaning |
| --- | --- |
| `aipay.pricing_policy` | `openrouter_reference_markup_v1` = the fields below are in effect |
| `aipay.reference_prices_per_million_uusd` | Upstream route reference unit price frozen at review; actual request costs may differ |
| `aipay.markup_bps` | Markup in basis points; `1200` = +12% |
| `aipay.retail_prices_per_million_uusd` | Retail unit price = the unit price the user is actually billed at |
| `aipay.availability` | `{status, reason, checked_at}`: local catalog permission for a new request, not an upstream health guarantee. `available` has `reason: null`; `unavailable` gives `pricing_stale`, `model_price_unavailable`, or `pricing_unverified`. |
| `aipay.valid_until` | Freshness deadline of this price (once expired, new requests are blocked) |
| `aipay.price_locked_at_request` | `true` = frozen at request time; price changes are not retroactive |
| `aipay.default_max_output_tokens` / `hard_max_output_tokens` | The default when `max_tokens` is omitted / the hard limit |

> `fee_bps` / `margin_bps` are legacy fields, always `0` for v3 models; **do not** derive prices from them.

Unavailable models remain listed. Select only `availability.status === "available"`; a model appearing in the list is not sufficient. `checked_at` is this response’s evaluation time, using the same catalog cache as chat (normally 30 seconds), not a provider probe time. Responses require revalidation (`Cache-Control: no-cache`); expiry is rechecked even while the catalog object is cached. Missing or invalid pricing observations return `pricing_unverified`; unverified v3 pricing fields may be omitted, and new requests fail with HTTP 503 before a charge.

---

## Retry discipline and the two keys

The core of the money semantics: **a lost response must never turn into a second charge**. The SDK has these conventions built in (raw HTTP should do the same):

- **Idempotency-Key** (required, 1–200 characters; the SDK generates a UUID): always reuse the same key on retries — the server guarantees the same key with the same payload is never executed twice. Same key with a different payload → 409 `idempotency_conflict`. Uniqueness is scoped to **your app** (all users under one client_id share one namespace) — use UUIDs, not sequence numbers that could collide. A failure that is known not to have executed (`retryable: true`) releases the key; settled, in-flight and unknown-outcome keys stay occupied.
- **Replay-Key** (optional; **not generated by the SDK by default** — enable with `new AIPayClient({ autoReplayKey: true })` or pass `replayKey` per request): with it, the original response text is encrypted and stored atomically with the settlement — resending the same payload with the same pair of keys within the validity window (15 minutes) retrieves the **original text** (repeatable within the window — it is not consumed on first read); without it you only get the receipt (how much was charged, how much usage; `choices` is empty). Format: 24–200 characters with at least 12 distinct characters (the SDK's `generateReplayKey()` complies). It is off by default because every key costs the server one deliberately expensive key derivation — an app that stores its own responses does not need it.
- **Retry on the flags, not on the HTTP status**: the SDK auto-retries only three situations — `provider_unavailable` / `provider_timeout` / `catalog_unavailable` with `details.retryable: true` (= the server guarantees it never executed, 0 charge; a 504 timeout during the estimate phase counts too), 429, and network-layer errors — exponential backoff with the same pair of keys (2 attempts by default). `pricing_stale` / `model_price_unavailable` / `pricing_unverified` / `service_overloaded` are also flagged `retryable: true` on the wire, but the SDK does **not** auto-retry them (it throws a plain `AIPayApiError`) — whether to wait, how long, or whether to switch models is your decision. **`details.outcome: 'unknown'` (e.g. `usage_unknown`, a timed-out 504) is never auto-retried** — unknown outcome = the money may already be settled, and firing again with a new key would be a second AI action; resending with the same pair of keys to learn the outcome is the correct move.
- **422 `budget_exceeded`**: details carry `max_affordable_uusd` — by default the SDK shrinks `max_tokens` to it and retries once with a **new** key (a 422 never executed; reusing the old key with a changed payload would be a 409).

The correct handling of an unknown outcome (SDK) — non-streaming raises `ProviderError.outcomeUnknown`; a stream interrupted before the receipt frame raises `StreamTruncatedError`; both carry the same pair of keys:

```ts
import { ProviderError, StreamTruncatedError, requestKeysOf } from '@coswic/aipay-sdk'

try {
  await client.chat.completions.create(params)
} catch (e) {
  const unknown =
    (e instanceof ProviderError && e.outcomeUnknown) || e instanceof StreamTruncatedError
  if (unknown) {
    // { idempotencyKey, replayKey } — every error thrown by the SDK's create() carries this pair of keys;
    // replayKey is non-null only when autoReplayKey is on (or you supplied one)
    const keys = requestKeysOf(e)
    if (!keys) throw e
    // Store the pair; later, resend the same payload with it = learn the outcome:
    // settled → the receipt is returned (with the original text only if a Replay-Key was used and the TTL has not passed; otherwise choices is empty); never executed → it executes for real this time
    await client.chat.completions.create({
      ...params,
      idempotencyKey: keys.idempotencyKey,
      ...(keys.replayKey ? { replayKey: keys.replayKey } : {}),
    })
  }
}
```

---

## Error codes

Errors always use the uniform envelope `{ "error": { "code", "message", "request_id", "details" } }` — branch on `code`; never parse the message text.

| HTTP | code | Meaning / what to do |
|---|---|---|
| 400 | `invalid_request` | Malformed request — fix the parameters, do not retry: an unknown field (the body accepts only `model` / `messages` / `max_tokens` / `stream`), `content` not a string or empty, an invalid `role`, `max_tokens` not a positive integer, a missing `Idempotency-Key` (`details.reason: 'missing_idempotency_key'`), a malformed `Replay-Key` (`'invalid_replay_key'`), or `stream: true` on a model that does not support streaming. The message names the field. |
| 400 | `model_not_allowed` | The model is not on the allow-list (typo in the id, or retired) — look up `GET /v1/models` and change the id; do not retry. |
| 401 | `unauthorized` | Token missing / expired (15 minutes) — refresh, then run once more. |
| 402 | `insufficient_balance` | The user's available credit is below the hold this request needs (details carry `shortfall_uusd`) — nothing was sent or charged, and it is not your error; tell the user their AI credit is too low. Production does not accept top-ups at the moment, so do not send users to top up. |
| 403 | `authorization_revoked` / `authorization_paused` | The user revoked / paused the authorization — respect it; after revocation the authorization flow must run again. |
| 403 | `app_pricing_changed` | Your app pricing exceeds the user’s accepted fee ceiling — guide the user to authorize again. |
| 403 | `consent_terms_changed` | AI-Pay's consent document was updated and the user has not re-consented — guide the user to authorize again (the consent screen appears automatically at the next sign-in). |
| 403 | `region_required` | The account has no declared region or has not accepted the current Terms of Service (returned by the developer API and the self-service API) — from the CLI run `aipay region set --country <CC> --accept-terms`; the dashboard opens the setup gate automatically. |
| 403 | `authorization_circuit_broken` / `forbidden` | Authorization circuit-broken (wait for the platform to recover) / insufficient permission: the app, user or wallet is not active, or a scope is missing (`details.reason: 'scope_missing'`, `details.scope` names the missing one — calling models requires both `ai.chat.create` and `wallet.charge`). |
| 409 | `idempotency_conflict` | A conflict on the same Idempotency-Key; branch on `details.reason`: `payload_mismatch` = same key, different payload → use a new key for the new intent; `in_flight` = the previous request with this key is still running → wait for it to finish, then resend with the same key; `unknown_outcome` = the previous outcome is unknown (the hold expired and was reclaimed; awaiting reconciliation) → do not fire again with a new key; `reclaimed_before_dispatch` = dispatch was not claimed (authorization, catalog or request state may have changed), without confirmation that the hold expired or was released → query with the same key; fresh execution is not guaranteed. |
| 422 | `budget_exceeded` | Over the user's limit (details carry `which` / `limit_uusd` / `max_affordable_uusd`) — retry smaller (new key) or guide the user to raise the limit. |
| 429 | `rate_limit_exceeded` | Rate limited — back off and retry (same key). |
| 500 / 503 | `internal_error` | Platform-internal error (e.g. the model's provider is not configured, settlement failed). 503 with `details.reason: invalid_timing_policy` means an invalid timing configuration; no hold or dispatch occurred and an operator must correct it. Do not fire again with a new key — resending with the same key is always safe (idempotency blocks a second execution) and also reveals the outcome. |
| 500 / 502 / 503 | `usage_unknown` (`details.outcome: 'unknown'`) | **Outcome unknown — do not retry blindly.** 500 = the upstream responded successfully but settlement could not be persisted (the money may be settled, or held until the hold TTL); 502 = the upstream never ran but the release could not be persisted (funds held until the TTL, 0 charge). 503 = no upstream dispatch, but hold release was not confirmed (the hold may have expired); do not assume a completed release permits fresh execution. Resend later with the same pair of keys to learn the outcome. |
| 502 | `provider_unavailable` (`details.retryable: true`) | The upstream failed before executing (connection layer / upstream rejected the request) — known not to have executed, 0 charge; back off and retry (same key). |
| 502 | `provider_unavailable` (`details.charged: false, retryable: false`) | The upstream **executed but reported failure** (including mid-stream failures) — 0 charge, but this key is finished (the same key will never call the upstream again); to try again you must use a **new** key. |
| 502 / 504 | `provider_unavailable` / `provider_timeout` (`details.outcome: 'unknown'`) | **Outcome unknown — do not retry blindly.** The upstream timed out, or returned something that could not be priced — it may have executed and incurred cost; funds stay held until the TTL and the key stays occupied. Resend with the same pair of keys to learn the outcome (meanwhile you get 409 `in_flight` / `unknown_outcome`). |
| 503 | `catalog_unavailable` | The model catalog could not be read — `retryable: true`, 0 charge; back off and retry (same key). |
| 503 | `ledger_unavailable` (`details.reason: hold_budget_unavailable`) | The hold's remaining execution time could not be confirmed before dispatch. The upstream was not called and zero-charge release was confirmed. `retryable: true`; retry later with the same key. An unconfirmed release returns `usage_unknown`; use the same pair of keys to learn the outcome. An open stream receives the same error in an error frame. |
| 503 | `pricing_stale` / `model_price_unavailable` / `pricing_unverified` | The model's price expired / was paused by the pricing circuit breaker / could not be verified (`details.model`) — `retryable: true`, 0 charge, other models are unaffected. The SDK does **not** auto-retry these errors: back off yourself or switch models. |
| 503 | `service_overloaded` | The server is saturated and this request has not started (with `Retry-After: 2`) — `retryable: true`, 0 charge; back off and retry with the same key (the SDK does not auto-retry). |
| 503 | `service_unavailable` (`details.reason: server_draining`) | The service is restarting and the request was not sent upstream. Any hold created for it has a confirmed zero-charge release. `retryable: true`; retry later with the same key. An unconfirmed release returns `usage_unknown` instead and must not be treated as released. |
| 503 | `maintenance_mode` | The system is under maintenance (with `Retry-After: 120`) — try again later; balances and authorizations are unaffected. |
| 503 / 504 | `provider_timeout` (`details.retryable: true`) | 504 = estimate timeout before creating a hold. 503 with `details.reason: hold_execution_budget_exhausted` = execution budget was exhausted before dispatch and zero-charge release was confirmed. Neither invoked the upstream; retry with the same key. |

> General rule: **retry on `details.retryable` and `details.outcome`, not on the HTTP status** — the same 502 / 503 / 504 can carry two completely different meanings, "known not to have executed" and "outcome unknown". Resending with the same Idempotency-Key is always safe: the server guarantees the same key is never executed twice.

---

## App pricing (service fees)

> ⚠️ **App service fees through AI-Pay are not open yet**: because payouts are not open yet, app service fees are disabled as well — the settings endpoint returns
> `503 app_fees_disabled` (setting back to 0 still works), the spending consent screen indicates that no App service fee is charged through AI-Pay for this authorization,
> and real requests **do not** charge the user any app fee. Existing settings may be retained, but a launch requires notice, applicable developer payment terms and user consent to the fees.
> The following describes the supported fee structure for integration planning; App service fees remain disabled until launch.

You can price your app under the Console's "App fees" tab; both can be combined. Attribution is based on App fees actually collected and is subject to the applicable agreement’s refund, reversal and payout conditions:

- **Service fee per request** (a fixed USD amount, at most $20 — the platform ceiling);
- **Markup on the AI usage fee** (a percentage of the catalog price).

What the user pays = catalog price + enabled and consented per-request service fees and markup. The catalog price follows that model’s policy and frozen prices (see "Model catalog and pricing"). These developer service fees are collected inside AI-Pay and are distinct from subscriptions or other charges your app collects outside the platform. The consent screen **discloses app charges prominently** and records the fee ceilings accepted by the user:

- When a fee **exceeds the user’s accepted ceiling**, requests from that authorization get 403 `app_pricing_changed` until they re-authorize — handle this code in your app and guide the user to consent again;
- A **decrease** takes effect immediately (the current price is charged).

You can also set a **suggested monthly limit** — the default on the consent screen (daily / per-request are derived as 10:1:100), at most **$200** (the platform ceiling for suggestions); the amount is still the user's decision and they can raise it themselves. Every change to pricing and suggestions leaves an audit record.

---

## Usage reconciliation and payouts

The Console's "Requests" tab lists your app's requests one by one: status (settled / failed / **pending confirmation** — do not retry pending ones), tokens, charges and service-fee attribution; each row expands to identifiers such as `request_id` / `Idempotency-Key` for matching against your logs. This tab does not directly display user email or full prompt/response bodies. Request, authorization and pseudonymous identifiers can be linked to your App's own records and are not anonymous data.

Once app fees and payouts open, service fees are reconciled against actual settlements and attribution records. Reserve periods, refund or dispute reversals, and payments depend on the formal agreement. Payouts and automated reversals are not currently available; no opening date is promised.

> ⚠️ **Payouts are not open yet**: payout account setup and payout requests are unavailable (`503 payouts_disabled`). **While app fees are disabled, new requests do not generate app service fees**. Existing records are retained; their status and future payout eligibility depend on actual transactions, the formal agreement, and applicable law. Test records are not real-money payables.

---

## References

- **Product overview and current status**: [`/about`](https://ai-pay.coswic.ai/about?lang=en) — what AI-Pay is, who pays, and what is open today.
- **API contract**: [`/openapi.yaml`](https://ai-pay.coswic.ai/openapi.yaml) (OpenAPI 3.1, the external developer contract — generated from the code, covering the endpoints you call and every error code). Feed it straight to a coding agent, import it into Postman/Bruno, or generate a client from it.
- **SDK and CLI**: [`@coswic/aipay-sdk`](https://www.npmjs.com/package/@coswic/aipay-sdk) (TypeScript SDK; the README covers the full usage and the built-in conventions) and [`@coswic/aipay`](https://www.npmjs.com/package/@coswic/aipay) (CLI; the command list is in the "Connect with the CLI or a coding agent" section) — both are public npm packages.
- **Legal documents**: [Terms of Service](https://ai-pay.coswic.ai/terms?lang=en), [Privacy Policy](https://ai-pay.coswic.ai/privacy?lang=en) — before your app guides users to authorize, make sure your usage is compatible with both.
- **Complete example**: QuickGist (a full reference implementation of "Sign in with AI-Pay" → streaming call → per-class error handling → unknown-outcome lookup) currently ships inside the AI-Pay source repository (`showcase/quickgist/`, shared skeleton in `showcase/kit/`) and is not yet published as a standalone download; to get the same wiring in your own project, run the CLI's `aipay init` (it generates the start/callback routes for your framework plus `aipay-agent.md`).
