← AirForge

Blog

Reproducible evidence, honestly reported

Notes on hardening private, open-weight LLM deployments: measured results, published hashes, and the limits stated plainly.

The silent failure modes of production RAG

Ways a deployed RAG system can produce a confident, wrong answer that nothing flags, and why "it looks good" stops being an acceptable answer for a security or quality review.

Testing a tool-hardened private GPT-OSS-20B: 600 probes, six gates

An acceptance evaluation of one hardened deployment artifact: override refusal 1.000, false refusal 0.000, with the training and evaluation hashes published so you can check the work.