Tools
Eval tooling
A small set of free, rate-limited endpoints for people who build or audit evaluation harnesses. Everything here is documented in OpenAPI and every response is computed on the spot.
GET /api/v1/judge-profile
The anatomy of a typical LLM-as-judge setup: template shape, documented biases (position, verbosity, self-preference, prompt sensitivity), and what calibration actually measures.
curl https://api.ashigiri.com//api/v1/judge-profileCompanion reading: Fingerprinting the Grader and LLM-as-judge under the hood.
POST /api/v1/answer-check
Runs the exact-match normalisation pipeline — casefold, whitespace collapse, punctuation strip, number format — and optionally compares against a reference. See the exact-match card for what each step assumes.
curl -X POST https://api.ashigiri.com//api/v1/answer-check \
-H 'content-type: application/json' \
-d '{"answer":" C. ","reference":"c"}'Response: { "normalized": "c", "matches": true, "steps": [...] }
GET /api/v1/sandbox-profile
Reflects what your request itself reveals — client hints, language, automation signature. Harness-side profiling (mounts, namespaces, syscall filters) is covered in the sandbox boundaries walkthrough.
curl https://api.ashigiri.com//api/v1/sandbox-profileLimits: 30 requests per minute per IP. The interactive version of the judge surface lives in the simulated judge playground.
Synthetic canary packs
Self-contained evaluation artifacts for scorer-tooling development: a private-split answer-key pack and a scorer-contract test suite. Each download carries a unique serial and per-item canary strings, so copies can be traced back to a download. These packs are synthetic — they do not contain items from any public benchmark.