← AirForge

Blog

Reproducible evidence, honestly reported

Notes on hardening private, open-weight LLM deployments: measured results, published hashes, and the limits stated plainly.

Testing a tool-hardened private GPT-OSS-20B: 600 probes, six gates

An acceptance evaluation of one hardened deployment artifact: override refusal 1.000, false refusal 0.000, with the training and evaluation hashes published so you can check the work.