au-radar is a RADAR-consistent benchmark of how well AI describes and reaches Australian federal services. It scores Australia 5.32, and finds that a robots-respecting agent is blocked at the free legal database but not at the silent government register.
ClauseKit runs five bodies of law through an LLM extraction pipeline. 171 of 262 extracted rules never produce a yes or no, and that gap is the result.
Porting a fine-tuned sentence embedding model into a Drupal 10 module, running ONNX inference in-process via PHP FFI. No Ollama, no external API, no vector DB.
A decompose-then-verify pipeline that checks each claim in an AI output against its source document, and an honest look at what the RAGTruth numbers actually say about it.