# CheapAIAPI API documentation

CheapAIAPI provides image, text, and video generation APIs for production workloads.
The OpenAPI specification is the authoritative public HTTP contract.

## Developer resources

- [Interactive CheapAIAPI API reference](https://cheapaiapi.org/docs)
- [CheapAIAPI OpenAPI specification](https://cheapaiapi.org/openapi.json)
- [Models and account-specific pricing](https://cheapaiapi.org/v1/models)
- [Public models and pricing pages](https://cheapaiapi.org/models/)
- [Public starting prices in Markdown](https://cheapaiapi.org/pricing.md)
- [Public starting prices as JSON](https://cheapaiapi.org/pricing.json)
- [Agent integration guide](https://cheapaiapi.org/llms.txt)
- [CheapAIAPI for coding agents](https://cheapaiapi.org/for-agents)
- [Image webhook guide](https://cheapaiapi.org/docs/webhooks.md)
- [Request access or integration help](https://cheapaiapi.org/contact)

## Pricing and access

CheapAIAPI is a prepaid, usage-based API with no monthly or annual subscription.
The public catalog contains real starting prices for new customers. Customers
with existing production volume may qualify for lower account-specific rates by
sending the exact model, current monthly usage, and average and peak RPS through
[CheapAIAPI contact](https://cheapaiapi.org/contact).

## Authentication

Send the issued API key as `Authorization: Bearer sk_cheap_...`. API keys are
issued after account approval. Unauthenticated `/v1/*` requests return HTTP 402
with access links to starting prices and contact. Invalid or revoked keys return
HTTP 401. The 402 body is not a cryptocurrency payment quote.

## Text generation

Two OpenAI-compatible text endpoints share the same models, prices, and API key:

- `POST /v1/chat/completions` takes a `messages` array and returns one
  completion or an SSE sequence ending in `[DONE]`. Use it for existing Chat
  Completions code and for image inputs through the `image_url` content part.
- `POST /v1/responses` takes `input` items and returns a Responses object or
  named Responses SSE events. Use it for new agent and tool-calling code, the
  `openai` SDK `responses.create` interface, and clients that speak the
  Responses wire format.

### Responses quickstart

Point any OpenAI client at `https://cheapaiapi.org/v1` and use your
`sk_cheap_...` key. Use the text SKUs listed by `GET /v1/models`, for example
`gemini-3.7-flash`.

```bash
curl https://cheapaiapi.org/v1/responses \
  -H "Authorization: Bearer $CHEAPAIAPI_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"model": "gemini-3.7-flash", "input": "Summarize this ticket in one sentence."}'
```

Python with the `openai` SDK:

```python
import os
from openai import OpenAI

client = OpenAI(base_url="https://cheapaiapi.org/v1", api_key=os.environ["CHEAPAIAPI_KEY"])

response = client.responses.create(
    model="gemini-3.7-flash",
    input="Summarize this ticket in one sentence.",
)
print(response.id, response.output_text)

stream = client.responses.create(
    model="gemini-3.7-flash",
    input="Write release notes for this diff.",
    stream=True,
)
for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
    elif event.type == "response.completed":
        print()
```

JavaScript with the `openai` SDK:

```js
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://cheapaiapi.org/v1",
  apiKey: process.env.CHEAPAIAPI_KEY,
});

const response = await client.responses.create({
  model: "gemini-3.7-flash",
  input: "Summarize this ticket in one sentence.",
});
console.log(response.output_text);
```

Vercel AI SDK:

```js
import { createOpenAI } from "@ai-sdk/openai";
import { generateText } from "ai";

const cheapaiapi = createOpenAI({
  baseURL: "https://cheapaiapi.org/v1",
  apiKey: process.env.CHEAPAIAPI_KEY,
});

const { text } = await generateText({
  model: cheapaiapi.responses("gemini-3.7-flash"),
  prompt: "Summarize this ticket in one sentence.",
  providerOptions: { openai: { store: false } },
});
```

Codex CLI: not supported yet. Codex sends a `custom` shell tool; function-tool
translation is planned.

### What `/v1/responses` supports

- `input` as a string or a list of `message` items (`system`, `developer`,
  `user`, `assistant`) plus `function_call` and `function_call_output` items;
  optional `instructions`.
- Function tools with `tool_choice` `auto`, `required`, `none`, or a named
  function, and `parallel_tool_calls`.
- Structured output through `text.format`: `text`, `json_object`, or
  `json_schema` with `name`, `schema`, and optional `strict`.
- `stream: true` returns named SSE events: `response.created`,
  `response.in_progress`, `response.output_item.added` / `.done`,
  `response.content_part.added` / `.done`, `response.output_text.delta` /
  `.done`, `response.refusal.delta` / `.done`,
  `response.function_call_arguments.delta` / `.done`, and one terminal
  `response.completed`, `response.incomplete`, or `response.failed`.
- `max_output_tokens` from 1 to 65,536 (defaults to the model's maximum
  output), `temperature`, and `top_p`. `prompt_cache_key`, `reasoning`, and
  `include` are accepted and ignored.
- Text input only. Send images through `POST /v1/chat/completions` with the
  `image_url` content part.

### Stateless by design

`/v1/responses` stores no conversation. Send the full conversation, including
earlier assistant output and `function_call` / `function_call_output` items, on
every request. `previous_response_id`, `conversation`, `store: true`, file and
image inputs, and hosted or `custom` tools are rejected with HTTP 400.

### Recovery

Every response `id` is `resp_<job id>`. Streaming responses also return it in
the `X-Response-Id` header before the first event. After a disconnect or
timeout, fetch the result with `GET /v1/responses/{id}`, or strip the `resp_`
prefix and call `GET /v1/jobs/{job id}`. Retrying with the same
`Idempotency-Key` and the same body returns the existing response instead of
creating and billing a new one. A non-stream HTTP 408 carries
`X-Should-Retry: false`: recover the existing Job, do not resubmit.

### Errors

Errors use the OpenAI envelope
`{"error": {"type": "...", "code": "...", "message": "...", "param": null}}`.

- 400 `model_not_found`, `input_too_large`, or a validation error: unknown or
  unpriced model, oversized request, or unsupported field.
- 401 `invalid_api_key`: invalid or revoked key. A missing key returns HTTP 402
  with access links.
- 402 `insufficient_balance`: not enough prepaid balance for this request.
- 408 `upstream_timeout`: the Job is still running; recover it by `id`.
- 409 `idempotency_conflict`: the `Idempotency-Key` was reused with a different
  body.
- 503 `upstream_unavailable`: no capacity right now; retry with backoff.

Failed Jobs report the same closed taxonomy as every other endpoint:
`RateLimited`, `UpstreamDown`, `NoFunds`, `InvalidPrompt`, `Unauthorized`.

### Pricing

`/v1/responses` bills the same input and output token rates as
`/v1/chat/completions` for the same SKU. See [public starting
prices](https://cheapaiapi.org/pricing.md) and [models](https://cheapaiapi.org/models/);
authenticated `GET /v1/models` shows your account rates. The balance hold is
sized from the request and `max_output_tokens`; the debit is actual usage and
never exceeds the hold.

## Video generation

Call `GET /v1/models` for the video models and exact rates enabled for your
account. Create a Seedance 2.5 image-to-video Job with the official request
shape:

```json
{
  "model": "seedance-2.5",
  "content": [
    {"type": "text", "text": "Animate this image."},
    {
      "type": "image_url",
      "image_url": {"url": "https://example.com/direct-image.jpg"},
      "role": "first_frame"
    }
  ],
  "duration": 5,
  "resolution": "720p",
  "ratio": "adaptive"
}
```

Send this body to `POST /v1/videos` with an `Idempotency-Key`. Seedance 2.0
accepts durations from 4 to 15 seconds and defaults to 5. Seedance 2.5 requires
a duration from 4 to 30 seconds, or `-1` for edit mode only on a reference video
(`ratio: "adaptive"`; the output keeps the edited clip's 4-30 second length and
is priced as the total reference-video seconds rounded up). `input_image` is unsupported; use
`content[].image_url.url` as shown above.

Refer to reference media in the prompt as @image1, @image2, …, @video1, ….
Numbering starts at 1 and counts each media type separately, in the order the items
appear in `content`. Only `reference_image` items count as images.

`nsfw_check` (default false): when true, prompts and images get an extra
content-moderation pass before generation and flagged requests fail with a
content-policy error at no charge. The model's built-in safety system always
applies.

Optional fields: `seed` (random seed that controls the randomness of generated
content, -1 by default (replaced by a random number); the same seed value for the
same request gives similar results, but complete consistency is not guaranteed)
and `return_last_frame: true` (both models; download the last
frame, when available, from `GET /v1/videos/{id}/content?variant=last_frame`). Seedance 2.5 also
accepts `watermark` (requests a watermark; default false), `output_format` (`mp4`
default or `mov`), and `omni_reference_task_type` (`auto`, `reference`, `edit`, or
`extend`; the last three need a reference video, `edit` requires `duration: -1`, and
`duration: -1` is allowed only with `edit`, `auto`, or no type and always runs as an edit).
Reference videos are 480p-720p; standard Seedance 2.5 also accepts 1080p reference
videos (up to 2,211,840 pixels per frame).

Input media must be a direct public HTTPS file that returns the image without
authentication, cookies, a login, or a sharing page, and it must remain valid
until generation finishes. CheapAIAPI image-result URLs require the owning
account's bearer key, so they cannot be used directly as video inputs. Download
the image through authorized access, copy it to public HTTPS storage you
control, and submit that copy's URL.

Poll `GET /v1/videos/{id}` until `status` is `completed` or `failed`. Download
a completed video with the same bearer key from `GET /v1/videos/{id}/content`.
Completed videos are available from `GET /v1/videos/{id}/content` for 23 hours after completion. Download and store each video yourself within that window; it cannot be retrieved afterwards.
When an input image cannot be downloaded, the failed video returns
`error.code: "InvalidPrompt"` and this message: "The input image could not be
downloaded. Use a public HTTPS URL that returns the image directly, without
authentication, and remains valid until generation finishes."
Media is downloaded and checked (format, size, length, resolution) before the Job
is accepted; a `422` names the field. Content moderation, task-type classification,
and MP3 audio length are checked during generation and can fail after `202` with a
full refund.
When the requested task type and the prompt disagree (for example a positive
duration for a prompt the model treats as a video edit), the failed
video returns `InvalidPrompt` with the model's own description of which
parameters to change, for example "`ratio` must be `adaptive`. `duration` must
be -1". A video edit (omni_reference_task_type edit, or duration -1) keeps the
reference video's 4 to 30 second length and needs ratio adaptive; for a new shot
guided by the reference video, use omni_reference_task_type reference with a
positive duration.
Other `InvalidPrompt` failures use: "The video request could not be generated.
Check the prompt and input images. For image input, use a public HTTPS URL that
returns the image directly, without authentication, and remains valid until
generation finishes."

## Image Job completion

Recommended production flow for image Jobs:

1. Call `POST /v1/images/generations?async=true` with `webhook_url`.
2. Store the returned Job `id`.
3. Handle the signed callback as the primary completion path.
4. Recover with `GET /v1/jobs/{id}` if no callback arrives.

Read `webhook_secret` from `GET /v1/balance`. The full guide is
[https://cheapaiapi.org/docs/webhooks.md](https://cheapaiapi.org/docs/webhooks.md).
Chat Completions, Responses, and video do not accept `webhook_url`.

## Low-latency integration

These are the defaults for a fast integration. Each one removes avoidable delay
that is not image generation time. The full checklist with a request example is
at [https://cheapaiapi.org/docs/webhooks.md#low-latency-integration](https://cheapaiapi.org/docs/webhooks.md#low-latency-integration).

- Create image Jobs with `?async=true` plus `webhook_url` and treat the signed
  callback as the completion path. Poll `GET /v1/jobs/{id}` only as recovery,
  every 5 seconds and never faster than every 2 seconds.
- Pass reference images as inline `data:image/<mime>;base64,...` bytes instead
  of `https://` URLs. A URL input may have to be fetched before generation can
  start; inline bytes never need that round trip.
- Send each reference image at the resolution you actually need and do not
  compress the request body. One reference image may be at most 100 MiB and
  the whole `image` array at most 200 MiB. Request bodies over 100 MB are
  rejected and base64 adds about a third, so send an image larger than about
  70 MB by HTTPS URL. A reference image larger than 30 MB,
  or one the model rejects as too large (retried once), is automatically
  downscaled to 3072 px on the long side (aspect ratio kept, no cropping) and
  re-encoded before generation; the model does not use more than 3072 px
  anyway. Inputs at or under 30 MB are sent unchanged.
- Send a stable `Idempotency-Key` per client-side Job and reuse it when
  retrying, and reuse one HTTP/1.1 keep-alive or HTTP/2 connection pool instead
  of opening a new TLS connection per request.
- Download `data[].url` as soon as the Job is `done`. Result links are not
  durable storage: treat the `X-Result-Retention-Days` response header
  (currently 7) as an upper bound. Contact
  us before sustaining a submission rate above the published standard image
  baseline.

## Recommended generation flow

1. Call `GET /v1/models` to discover the SKUs and prices enabled for the account.
2. Submit a generation request with a stable `Idempotency-Key`.
3. Store the returned job `id`.
4. For image Jobs, pass `webhook_url` and treat `GET /v1/jobs/{id}` as
   recovery. For other products, poll the documented job endpoint.

## Image input capabilities

- The optional `image` array accepts public `https://` URLs and `data:image/<mime>;base64,...` data URIs.
- Every `nano-banana-2.1-*` and `nano-banana-2-lite-*` SKU accepts at most 14 images.
- Every `nano-banana-pro-*` SKU accepts at most 14 images.
- The base `nano-banana` SKU accepts at most 4 images.
- Every `gpt-image-*` SKU accepts at most 10 images. The default hard ceiling for other
  image SKUs is 10.
- OpenAPI exposes the global ceiling as `maxItems: 14` plus an exact model-limit
  map; when several selectors match a model, the longest matching selector wins. Authenticated
  `GET /v1/models` returns the effective hard limit for each enabled image SKU at
  `capabilities.image_input.max_items`.
