Incorrect diagnosis
Synthetic example: Attributes a scheduling failure to networking despite evidence that a resource request cannot be satisfied.
Evaluation design / private enterprise work
How I structured an evidence-led comparison for a Kubernetes diagnostic agent without exposing private prompts, infrastructure, or measurements.
Methodology published; measured results are not publicly disclosed.
Evaluate whether a smaller self-hosted language model can diagnose representative Kubernetes incidents while choosing and using investigation tools appropriately. This mirrors the multi-step behavior required by an internal SRE agent.
Human review was part of the evaluation design, but the sample size and reviewer protocol are confidential.
DeepEval, Langfuse, and an LLM-as-a-Judge workflow supported structured review.
Synthetic example: Attributes a scheduling failure to networking despite evidence that a resource request cannot be satisfied.
Synthetic example: Requests unrelated cluster-wide logs after the supplied event already identifies the failing workload.
Synthetic example: Repeats a generic remediation without incorporating a synthetic warning event returned by a tool.
Synthetic example: Suggests a disruptive change before checking workload scope or rollback options.
No public claim is made about which model matched the baseline, where a candidate failed, whether fine-tuning improved it, or which measured configuration should be deployed.
The flow, metric categories, review questions, error taxonomy, and decision gates can be reused as an experiment template. The benchmark itself is not publicly reproducible: The private scenarios, model configurations, judge setup, hardware, decoding settings, and measured outputs are not approved for disclosure.