🔒 PremiumPremium

Modifying Local Evaluation Result Files on the Fly: Tricks and Tips

aktualizacja: 11 października 2026

The life of a result file

A result file has a short, predictable life. The harness finishes a task, serialises the score — JSONL is the usual format, one record per task, appended as runs complete — and writes it to a location that a later stage will read: a summary script, a leaderboard upload, a CI artifact step. Between the write and the read there is a gap, and the gap is the interesting part.

In a local, single-machine harness, the whole lifecycle happens inside one filesystem. The result file sits in the working directory or an output folder; the summary script reads it minutes or hours later; nothing in between verifies that the bytes are the ones the scorer wrote. Every property of that arrangement — writable directory, deferred reader, no signature, no second copy — is an implicit trust decision made by the pipeline's author.

The same lifecycle looks completely different once the pipeline treats the number as important. Results are written to a directory the candidate cannot reach, or signed, or shipped off-machine immediately, or all three. The gap between "computed" and "recorded" still exists — it always exists — but the question becomes who else can write into it, and whether the reader can tell if someone did.

Premium content

This post is part of the premium archive

Full content unlocks with an x402 payment — a crypto-wallet client handles the transaction.

Modifying Local Evaluation Result Files on the Fly: Tricks and Tips — ashigiri