Model
market-preview
market-preview is its own market on OpenFill. A request names market-preview, sets a price limit, and pays the level the market is clearing at when it runs.
A preview of the market that opens once there is traction and compute behind it. Its prices are simulated until then; fixed-price models run today.
Datasheet
| Field | Value |
|---|---|
| Model id | market-preview |
| Also answers to | No other id |
| Context length | 262,144 tokens |
| Max output | 128,000 tokens |
| Pricing ratio | 1:10:20 |
The ratio weights a cache-hit, a cache-miss and an output token in the price. max_tokens above the max output is clamped down to it.
The market now
The clearing level, the price of each token kind at that level, and the quote a locked request bids.
Clearing level · market-preview · preview
Per 1M tokens: $0.0172 cache-hit · $0.1716 cache-miss · $0.3432 output
A price_lock request pays at most the locked price, and cancels with no charge if it has not started within 10 seconds.
The basis is the margin over the level, and a ceiling rather than a charge. See how the lock works.
Uptime, last 24 hours
Each check is a chat completion through the public API, run the way the models board describes.
Uptime
n/a
Checks
0
Failures
0
Latency
n/a
Queue wait
n/a
Checks begin when the market opens.
The first call
Set OPENFILL_API_KEY to a key from API keys before running an example. A key starts with of_live_ and is shown once.
import osfrom openai import OpenAI client = OpenAI( base_url=( "https://api.openfill.ai/v1" ), api_key=os.environ[ "OPENFILL_API_KEY" ],) r = client.chat.completions.create( model="market-preview", messages=[ {"role": "user", "content": "hi"} ], extra_body={"price_limit": 0.10},)print(r.choices[0].message.content)print(r.usage)import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.openfill.ai/v1", apiKey: process.env.OPENFILL_API_KEY,}); const r = await client.chat.completions.create( { model: "market-preview", messages: [{ role: "user", content: "hi" }], }, { headers: { "X-Price-Limit": "0.10" } },);console.log(r.choices[0].message.content);console.log(r.usage);curl "https://api.openfill.ai/v1/chat/completions" \ -H "Authorization: Bearer $OPENFILL_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "market-preview", "messages": [ {"role": "user", "content": "hi"} ], "price_limit": 0.10 }'# The base URL without /v1; Claude Code appends /v1/messagesexport ANTHROPIC_BASE_URL="https://api.openfill.ai"export ANTHROPIC_AUTH_TOKEN="$OPENFILL_API_KEY"export ANTHROPIC_API_KEY="" # Every model slot by hand: discovery keeps only ids that name Claudeexport ANTHROPIC_MODEL="market-preview"export ANTHROPIC_DEFAULT_SONNET_MODEL="market-preview"export ANTHROPIC_DEFAULT_HAIKU_MODEL="market-preview"export CLAUDE_CODE_SUBAGENT_MODEL="market-preview" # The price limit and the wait, merged into every request bodyexport CLAUDE_CODE_EXTRA_BODY='{"price_limit": 0.10, "max_wait": 600}'export CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1claude# ~/.codex/config.tomlmodel_provider = "openfill"model = "market-preview"model_context_window = 262144web_search = "disabled"show_raw_agent_reasoning = true [model_providers.openfill]name = "OpenFill"base_url = "https://api.openfill.ai/v1"env_key = "OPENFILL_API_KEY"wire_api = "responses"http_headers = { "X-Price-Limit" = "0.10" }// opencode.json{ "$schema": "https://opencode.ai/config.json", "provider": { "openfill": { "npm": "@ai-sdk/openai-compatible", "name": "OpenFill", "options": { "baseURL": "https://api.openfill.ai/v1", "apiKey": "{env:OPENFILL_API_KEY}", "headers": { "X-Price-Limit": "0.10" } }, "models": { "market-preview": { "name": "market-preview" } } } }}# Cline, Roo Code and Kilo Code# Settings > API Provider: OpenAI Compatible Base URL https://api.openfill.ai/v1API Key (your OPENFILL_API_KEY)Model ID market-preview # These tools send no price limit of their# own. Set one on the key, on Bidding.# A key from API keys, shown once when madeexport OPENFILL_API_KEY="of_live_..." # Any tool that reads the OpenAI variablesexport OPENAI_BASE_URL="https://api.openfill.ai/v1"export OPENAI_API_KEY="$OPENFILL_API_KEY" # The price limit comes from the key or the# account default, set on Bidding.curl -s "$OPENAI_BASE_URL/models" \ -H "Authorization: Bearer $OPENAI_API_KEY"import { createOpenAICompatible } from "@ai-sdk/openai-compatible";import { generateText } from "ai"; const openfill = createOpenAICompatible({ name: "openfill", baseURL: "https://api.openfill.ai/v1", apiKey: process.env.OPENFILL_API_KEY, headers: { "X-Price-Limit": "0.10" },}); const { text } = await generateText({ model: openfill.chatModel("market-preview"), prompt: "hi",});console.log(text);import osfrom langchain_openai import ChatOpenAI llm = ChatOpenAI( model="market-preview", base_url="https://api.openfill.ai/v1", api_key=os.environ["OPENFILL_API_KEY"], default_headers={"X-Price-Limit": "0.10"},)print(llm.invoke("hi").content)import osfrom litellm import completion r = completion( model="openai/market-preview", api_base="https://api.openfill.ai/v1", api_key=os.environ["OPENFILL_API_KEY"], messages=[ {"role": "user", "content": "hi"} ], extra_headers={"X-Price-Limit": "0.10"},)print(r.choices[0].message.content)import osfrom openai import AsyncOpenAIfrom agents import ( Agent, Runner, set_default_openai_api, set_default_openai_client, set_tracing_disabled,) set_default_openai_client(AsyncOpenAI( base_url="https://api.openfill.ai/v1", api_key=os.environ["OPENFILL_API_KEY"], default_headers={"X-Price-Limit": "0.10"},))set_default_openai_api("chat_completions")set_tracing_disabled(True) agent = Agent(name="assistant", model="market-preview")print(Runner.run_sync(agent, "hi").final_output)Questions
What does market-preview cost on OpenFill?
The clearing level when the request runs, up to the price limit you set, shared across the three token kinds by the ratio 1:10:20.
What happens while the level is above my price limit?
The request waits in the book while capacity is short and runs when there is room, for up to max_wait (10 seconds to 30 days). Past that it expires with 408 and no charge.
How do I know the price before I send?
With price_lock: true the request bids the published quote, or your price limit if lower, and starts within 10 seconds or cancels with 429 and no charge.
What are the rate limits?
1,200 inference requests a minute per key, 2,000 queued orders per account, and a request body of up to 8 MB. Every rate limit is on GET /v1/rate-limits.
What is stored?
Storage is off by default: a prompt and its response exist only while the request runs. With it on, results are kept up to 30 days; store: false keeps less.