AI Quality, Drift & Fairness: continuous evaluation and bias mitigation

A provider quietly updates the base model. The API is unchanged, behavior shifts, and no one deployed a thing. Continuous evaluation and LLM-as-a-judge are how you catch it before compliance asks why approvals in region N dropped.

A provider quietly updated the base model. The API is formally the same, the version string is unchanged — but behavior shifted: Kovcheg started inventing loan terms not present in the documents slightly more often, and treating applications from one region slightly more strictly. No one changed anything in the code. A month later it surfaces as a spike in complaints and a question from compliance: "Why have approvals in region N dropped?" This is what silent degradation looks like — the most common way an AI system's quality erodes without a single deployment.

Continue reading “AI Quality, Drift & Fairness: continuous evaluation and bias mitigation”