# glm-5.3 API

> The Markdown form of https://openfill.ai/models/glm-5.3. The index of every docs page is https://openfill.ai/llms.txt.

Model

# glm-5.3

glm-5.3 is sold at a fixed price per token. A request names `glm-5.3` and is charged the prices below for the tokens it uses.

## Datasheet

| Field | Value |
| --- | --- |
| Model id | glm-5.3 |
| Also answers to | No other id |
| Context length | 262,144 tokens |
| Max output | 128,000 tokens |
| Pricing | Fixed price |

`max_tokens` above the max output is clamped down to it.

## Prices

What a request pays for each kind of token. `GET /v1/models` carries the current prices.

Per 1M tokens: $0.0750 cache-hit · $0.7500 cache-miss · $2.2500 output

## Uptime, last 24 hours

Each check is a chat completion through the public API, run the way the [models board](https://openfill.ai/models) describes.

The probe results: GET https://api.openfill.ai/v1/uptime?range=24h&model=glm-5.3.

## The first call

Set `OPENFILL_API_KEY` to a key from [API keys](https://openfill.ai/keys) before running an example. A key starts with `of_live_` and is shown once.

```python
import os
from openai import OpenAI

client = OpenAI(
    base_url=(
        "https://api.openfill.ai/v1"
    ),
    api_key=os.environ[
        "OPENFILL_API_KEY"
    ],
)

r = client.chat.completions.create(
    model="glm-5.3",
    messages=[
        {"role": "user", "content": "hi"}
    ],
)
print(r.choices[0].message.content)
print(r.usage)
```

```typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.openfill.ai/v1",
  apiKey: process.env.OPENFILL_API_KEY,
});

const r = await client.chat.completions.create({
  model: "glm-5.3",
  messages: [{ role: "user", content: "hi" }],
});
console.log(r.choices[0].message.content);
console.log(r.usage);
```

**curl**

```bash
curl "https://api.openfill.ai/v1/chat/completions" \
  -H "Authorization: Bearer $OPENFILL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3",
    "messages": [
      {"role": "user", "content": "hi"}
    ]
  }'
```

**claude code**

```bash
# The base URL without /v1; Claude Code appends /v1/messages
export ANTHROPIC_BASE_URL="https://api.openfill.ai"
export ANTHROPIC_AUTH_TOKEN="$OPENFILL_API_KEY"
export ANTHROPIC_API_KEY=""

# Every model slot by hand: discovery keeps only ids that name Claude
export ANTHROPIC_MODEL="glm-5.3"
export ANTHROPIC_DEFAULT_SONNET_MODEL="glm-5.3"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="glm-5.3"
export CLAUDE_CODE_SUBAGENT_MODEL="glm-5.3"

# The wait, merged into every request body
export CLAUDE_CODE_EXTRA_BODY='{"max_wait": 600}'
export CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1
claude
```

**codex**

```toml
# ~/.codex/config.toml
model_provider = "openfill"
model = "glm-5.3"
model_context_window = 262144
web_search = "disabled"
show_raw_agent_reasoning = true

[model_providers.openfill]
name = "OpenFill"
base_url = "https://api.openfill.ai/v1"
env_key = "OPENFILL_API_KEY"
wire_api = "responses"
```

**opencode**

```json
// opencode.json
{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "openfill": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "OpenFill",
      "options": {
        "baseURL": "https://api.openfill.ai/v1",
        "apiKey": "{env:OPENFILL_API_KEY}"
      },
      "models": {
        "glm-5.3": { "name": "glm-5.3" }
      }
    }
  }
}
```

**cline**

```text
# Cline, Roo Code and Kilo Code
# Settings > API Provider: OpenAI Compatible

Base URL   https://api.openfill.ai/v1
API Key    (your OPENFILL_API_KEY)
Model ID   glm-5.3
```

**shell**

```bash
# A key from API keys, shown once when made
export OPENFILL_API_KEY="of_live_..."

# Any tool that reads the OpenAI variables
export OPENAI_BASE_URL="https://api.openfill.ai/v1"
export OPENAI_API_KEY="$OPENFILL_API_KEY"

curl -s "$OPENAI_BASE_URL/models" \
  -H "Authorization: Bearer $OPENAI_API_KEY"
```

**ai sdk**

```typescript
import { createOpenAICompatible } from "@ai-sdk/openai-compatible";
import { generateText } from "ai";

const openfill = createOpenAICompatible({
  name: "openfill",
  baseURL: "https://api.openfill.ai/v1",
  apiKey: process.env.OPENFILL_API_KEY,
});

const { text } = await generateText({
  model: openfill.chatModel("glm-5.3"),
  prompt: "hi",
});
console.log(text);
```

**langchain**

```python
import os
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="glm-5.3",
    base_url="https://api.openfill.ai/v1",
    api_key=os.environ["OPENFILL_API_KEY"],
)
print(llm.invoke("hi").content)
```

**litellm**

```python
import os
from litellm import completion

r = completion(
    model="openai/glm-5.3",
    api_base="https://api.openfill.ai/v1",
    api_key=os.environ["OPENFILL_API_KEY"],
    messages=[
        {"role": "user", "content": "hi"}
    ],
)
print(r.choices[0].message.content)
```

**agents sdk**

```python
import os
from openai import AsyncOpenAI
from agents import (
    Agent, Runner,
    set_default_openai_api,
    set_default_openai_client,
    set_tracing_disabled,
)

set_default_openai_client(AsyncOpenAI(
    base_url="https://api.openfill.ai/v1",
    api_key=os.environ["OPENFILL_API_KEY"],
))
set_default_openai_api("chat_completions")
set_tracing_disabled(True)

agent = Agent(name="assistant", model="glm-5.3")
print(Runner.run_sync(agent, "hi").final_output)
```

## Questions

### What does glm-5.3 cost on OpenFill?

Each request is charged for the tokens it uses: $0.0750 per 1M cache-hit tokens, $0.7500 per 1M cache-miss tokens and $2.2500 per 1M output tokens.

[Every model and its prices](https://openfill.ai/docs/models)

### What happens when glm-5.3 is at capacity?

The request waits for a free slot, taking turns with other accounts, for up to `max_wait` (10 seconds to 30 days). Past that it ends with 429 and no charge.

[Waiting and timeouts](https://openfill.ai/docs/orders#waiting)

### Does a price limit or price lock apply?

glm-5.3 ignores both, so a tool that sends one `X-Price-Limit` for every model can keep it.

[How pricing works](https://openfill.ai/docs/pricing)

### What are the rate limits?

1,200 inference requests a minute per key, 2,000 queued orders per account, and a request body of up to 8 MB. Every rate limit is on `GET /v1/rate-limits`.

[Rate limits](https://openfill.ai/docs/rate-limits)

### What is stored?

Storage is off by default: a prompt and its response exist only while the request runs. With it on, results are kept up to 30 days; `store: false` keeps less.

[Data storage](https://openfill.ai/docs/storage)
