AI quality requires more than a good demo
An LLM workflow can return a plausible answer while failing on source grounding, safety, cost, latency, or repeatability. I built a combined evaluation and telemetry layer to make those failure modes visible.