OpenFill is in development. Inference is off, but is tested end to end, and will switch on at launch. All data currently on the site is for live testing: it will be erased at launch.

Skip to content
openfill
← All posts
OpenFill7 min read

Adventures in Astra and Chess

Astra beat Maia 3 at its 2400 setting in all 40 games. A second match against BT4 policy gave us a much tougher test and a collection of games worth playing through.

The process

We used the LLM Chess benchmark to play two local matches with GPT-6 Astra at max reasoning effort. Each match had 40 games: 20 with Astra as White and 20 as Black. The launcher allowed 20 games to run concurrently.

The harness was deliberately small. On every Astra turn it assembled a fresh prompt containing a text board, the position in FEN, and a list of legal moves in UCI notation. Astra replied with make_move e2e4, for example, and the harness parsed and applied that move.

We pass in the legal move list because online chess interfaces already constrain which moves a player can make. LC0 likewise selects from legal moves when playing BT4. This makes choosing a strong move the central task, with move legality supplied to both sides.

It is stateless: every turn gets a fresh prompt. Earlier conversation and reasoning are discarded. The chess board remains in the harness, which sends the current position again on the next turn. Astra received no engine evaluation or engine-ranked move list.

The system message was You are a precise chess engine. You always reply with a single legal move. Below is the actual opening user prompt from the BT4 match.

Single-turn prompt: starting position
You are playing chess as white. Choose the strongest legal move.

Board:
♜ ♞ ♝ ♛ ♚ ♝ ♞ ♜
♟ ♟ ♟ ♟ ♟ ♟ ♟ ♟
⭘ ⭘ ⭘ ⭘ ⭘ ⭘ ⭘ ⭘
⭘ ⭘ ⭘ ⭘ ⭘ ⭘ ⭘ ⭘
⭘ ⭘ ⭘ ⭘ ⭘ ⭘ ⭘ ⭘
⭘ ⭘ ⭘ ⭘ ⭘ ⭘ ⭘ ⭘
♙ ♙ ♙ ♙ ♙ ♙ ♙ ♙
♖ ♘ ♗ ♕ ♔ ♗ ♘ ♖

FEN: rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1

Legal moves (UCI): g1h3,g1f3,b1c3,b1a3,h2h3,g2g3,f2f3,e2e3,d2d3,c2c3,b2b3,a2a3,h2h4,g2g4,f2f4,e2e4,d2d4,c2c4,b2b4,a2a4

Reply with your move as the FINAL line, exactly as: make_move <UCI>  (e.g. make_move e2e4). Pick a move verbatim from the Legal moves list. You may reason first, but the last line must be that.

The moves came through ChatGPT subscription OAuth. The harness used the model's max reasoning setting, a 200-ply game cap, and the same simple UCI prompt for both opponents. A ply is one player's move; 200 plies is 100 full moves.

A clean sweep against Maia 3

Our first opponent was Maia 3, using its 79M model with the Elo level set to 2400. Maia is designed to model human play at a selected level. The Maia 3 scale uses Lichess blitz ratings; it is not a direct FIDE rating. At the time of writing, 2400 is in the top 1% of players in Lichess's weekly blitz rating distribution.

Astra against Maia 3 at 2400

40 wins · 0 losses · 0 draws

Astra won every game, with both colours. A 40-0 result told us that this opponent was too easy to locate Astra's strength precisely. A perfect score also has no finite Elo point estimate under the usual logistic model. We needed a harder opponent.

Download Maia match PGN (40 games)

A harder opponent

We moved to BT4-tf13tune, the static network opponent in our second match. BT4 comes from Leela Chess Zero's transformer family; the project documents its substantial gains in raw policy strength.

LC0 ran on CPU with its policyhead search and a one-node limit. It evaluated the current position once and always chose the legal move with the highest policy score. There was no sampling and no tree search. This tests the network's immediate move preference.

One shared LC0 process served the concurrent games. Each request carried the game's complete move history, with fresh engine game state. We kept Astra's model, reasoning effort and prompt unchanged.

How quickly do the games diverge?

With a deterministic opponent and the same starting board, repeated openings are a real concern. We counted distinct positions at each ply across all 40 BT4 games. At ply zero there is one position. All 40 games first occupy different positions at ply 44, or 22 full moves.

Unique positions across 40 BT4 games: one at ply zero, 29 at ply 31, 39 at ply 43, and 40 at ply 44.
Plies out of book are measured from the starting position: this experiment supplied no opening book. The vertical scale counts unique positions among the same 40 games.

Position identity includes piece placement, side to move, castling rights and legal en passant rights. Move counters are ignored. Transpositions can make the count fall, as it does at plies 24 and 38. Forty unique positions at one ply does not make the games statistically independent.

Download the position counts (CSV)

The final result

BT4 won the match. Astra scored 7 wins, 25 losses and 8 draws, for 27.5% of the available points. With 20 games per colour, the pooled score is colour-balanced.

Final 40-game match, with BT4 set to zero as a relative Elo reference
PlayerWinsLossesDrawsScoreRelative Elo (95% CI)
Astra (max)725827.5%-170 ± 120
BT4 policy257872.5%0 (reference)

The match-implied Elo difference is approximately 170 ± 120 points in BT4's favour (95% CI), from 400 × log10(score / (1 − score)). BT4 is assigned zero only as a reference. These are relative match ratings; an absolute rating requires a calibrated anchor.

How to read the error bars

Error bars show 95% confidence intervals. The benchmark's logistic/Fisher method supplies finite Elo intervals. Astra's 40-0 Maia result uses an exact binomial interval, with no finite upper Elo bound. These intervals assume independent games and fixed opponent ratings. They exclude calibration uncertainty, shared openings and harness differences.

One result was adjudicated: Black j11 reached the 200-ply cap with BT4 holding king and rook against Astra's lone king. We scored it as a White/BT4 win. Its PGN records 1-0 and marks the adjudication. The other game results follow their recorded endings.

Astra's fresh prompt contains no position history, so it cannot count repetitions. BT4 does receive move history and repetition features, though our setup simply selects its top policy move. The harness enforces fivefold repetition. All six games that ended this way were also judged drawn on the board.

Eight games were interrupted by model-capacity errors and continued from their saved positions. The downloads combine each continuation with its original game. Every game appears once, with its full move sequence.

Download BT4 match PGN (40 games)

Our favourite Astra win

Astra has White in this Berlin Defence. The game runs for 179 plies, ending with checkmate on move 90.

The first highlight comes when Astra is three pawns up, yet the position remains drawn because of Black's advanced a-pawn. The second is the late promotion tactic with Astra's advancing h-pawn.

A step change

Astra is a step change in chess ability in these experiments. It swept Maia at 2400 and took seven wins from a formidable raw-policy opponent, using a fresh text prompt on every move.

Approximate Maia-scale Elo ratings with 95% uncertainty
LLM / reasoningApprox. Lichess blitz Elo (95% CI)
GPT-6 Astra · max2,810
GPT-5.5 · xhigh1,890 ± 160
Gemini 3.5 Flash · high900 ± 90
Gemini 3.1 Pro Preview · high850 ± 90
DeepSeek V4 Pro · high420 ± 90
Gemma 4 31B IT370 ± 40
Qwen 3.6 27B290 ± 110

Astra's ≥ 2,810 is the lower end of its 95% confidence interval; the upper end is unbounded. Maia-scale estimates are rounded to the nearest 10 Elo. Sources: GPT-5.5 xhigh and other Maia runs. Different harnesses make these comparisons contextual.

BT4 still won this match convincingly. Forty games against each opponent give us an interesting result and concrete games to inspect; more opponents and a larger sample would sharpen the rating estimate. The repository, prompts and PGNs make this experiment available to reproduce and examine.