Harish Manoharan

Testing LLMs and Databricks platforms for regulated use

An engine that runs 19 tests on a customer’s LLM deployment and 59 on their Databricks platform, and records evidence an auditor can read.

Purgo AI, 2025–26, 4 min read

Updated

My part. I wrote most of the current code of the validation service, all 59 platform test definitions, the LLM validation module and the PDF evidence reports, which render in a separate service. A colleague built the first scaffold.

Built with. TypeScript, NestJS, Python, Azure OpenAI, Databricks, AWS, GCP Cloud Run

Companies in regulated (GxP) work, like pharma, can’t just start using a model or a data platform. They have to qualify it first: show, with evidence, that it was installed as specified (installation qualification, or IQ) and that it works as intended (operational qualification, or OQ). Done by hand, that’s a lot of repetitive checking. At Purgo I wrote most of the engine that runs those checks automatically and records the evidence.

Qualifying an LLM

The engine qualifies a customer’s model, on Azure OpenAI or served from Databricks. There are 19 built-in tests. The installation ones check how the deployment is set up: region, authentication, access control, private networking, and whether calls are logged for audit. The operational ones run the model on fixed datasets and compare the results with thresholds the customer sets.

Operational tests and what each one catches
TestWhat it catches
HallucinationInventing a value the source doesn’t contain, on 40 cases with missing fields
Extraction accuracyWrong values in extracted fields, scored per field with F1
ClassificationMisclassified complaints and triage levels, against labeled cases
DeterminismDifferent answers to the same input over 20 repeated runs
Latency under loadFailed or slow calls with 20 concurrent workers
Demographic biasDifferent outputs when only the demographics in a case change
TraceabilityA response that can’t be traced back to the model and the request that produced it

Qualifying a Databricks platform

The platform tests are 59 versioned definitions in 17 suites. They cover things like catalog setup, access control and audit logging, cluster behavior, network and security-group settings, lineage and data sharing. Each definition is a chain of API calls against the customer’s own workspace on Azure or AWS, plus a small piece of logic that decides pass or fail and records the expected and actual values.

The main design choices:

  • Test definitions live in their own versioned repository, apart from the engine, and every report records the exact version it ran. Updating a test doesn’t need a redeploy, and an auditor can see exactly what was tested.
  • The pass/fail logic runs in a sandbox, with a timeout on every test.
  • A run is a job: it returns at once, reports progress test by test and can be stopped. The finished result includes each step’s request and response, with credentials removed.
  • Requirements can become tests. A customer can upload a requirements document (PDF, DOCX or text), and an LLM pipeline then pulls out the testable statements, writes test definitions in the same format and checks them before they run. If a reviewer rejects one, it’s regenerated with their feedback.

What a run produces

A threshold turns each measurement into a pass or a fail, and the report keeps the numbers behind every result, so a reviewer can disagree with a threshold without rerunning anything. A finished report can be signed in the app through embedded DocuSign signing, and the signed PDF is stored for download.

What I’d do differently

The engine grew quickly. I didn’t write enough automated tests for it, and in 2025 I merged many of my own changes without review. The test library got reliable through many rounds of fixes against live workspaces (polling, retries, longer waits) rather than through tests I could run locally. If I built it again, I’d start with a fake workspace so most of those failures would show up on my machine first, and I’d ask for a review before merging.

Next: Finding the tables an agent misses