Testing LLMs and Databricks platforms for regulated use
An engine that runs 19 tests on a customer’s LLM deployment and 59 on their Databricks platform, and records evidence an auditor can read.
Companies in regulated (GxP) work, like pharma, can’t just start using a model or a data platform. They have to qualify it first: show, with evidence, that it was installed as specified (installation qualification, or IQ) and that it works as intended (operational qualification, or OQ). Done by hand, that’s a lot of repetitive checking. At Purgo I wrote most of the engine that runs those checks automatically and records the evidence.
Qualifying an LLM
The engine qualifies a customer’s model, on Azure OpenAI or served from Databricks. There are 19 built-in tests. The installation ones check how the deployment is set up: region, authentication, access control, private networking, and whether calls are logged for audit. The operational ones run the model on fixed datasets and compare the results with thresholds the customer sets.
| Test | What it catches |
|---|---|
| Hallucination | Inventing a value the source doesn’t contain, on 40 cases with missing fields |
| Extraction accuracy | Wrong values in extracted fields, scored per field with F1 |
| Classification | Misclassified complaints and triage levels, against labeled cases |
| Determinism | Different answers to the same input over 20 repeated runs |
| Latency under load | Failed or slow calls with 20 concurrent workers |
| Demographic bias | Different outputs when only the demographics in a case change |
| Traceability | A response that can’t be traced back to the model and the request that produced it |
Qualifying a Databricks platform
The platform tests are 59 versioned definitions in 17 suites. They cover things like catalog setup, access control and audit logging, cluster behavior, network and security-group settings, lineage and data sharing. Each definition is a chain of API calls against the customer’s own workspace on Azure or AWS, plus a small piece of logic that decides pass or fail and records the expected and actual values.
The main design choices:
- Test definitions live in their own versioned repository, apart from the engine, and every report records the exact version it ran. Updating a test doesn’t need a redeploy, and an auditor can see exactly what was tested.
- The pass/fail logic runs in a sandbox, with a timeout on every test.
- A run is a job: it returns at once, reports progress test by test and can be stopped. The finished result includes each step’s request and response, with credentials removed.
- Requirements can become tests. A customer can upload a requirements document (PDF, DOCX or text), and an LLM pipeline then pulls out the testable statements, writes test definitions in the same format and checks them before they run. If a reviewer rejects one, it’s regenerated with their feedback.
What a run produces
A threshold turns each measurement into a pass or a fail, and the report keeps the numbers behind every result, so a reviewer can disagree with a threshold without rerunning anything. A finished report can be signed in the app through embedded DocuSign signing, and the signed PDF is stored for download.
What I’d do differently
The engine grew quickly. I didn’t write enough automated tests for it, and in 2025 I merged many of my own changes without review. The test library got reliable through many rounds of fixes against live workspaces (polling, retries, longer waits) rather than through tests I could run locally. If I built it again, I’d start with a fake workspace so most of those failures would show up on my machine first, and I’d ask for a review before merging.