Fine-tuning & EvalsEvalsLLM-as-judgeSafety
Build an Eval Harness for a Medical LLM App
Rosewood Health AI · Remote (US)
$10k fixed
The outcome
Before they ship, they need to know the model is right — an eval suite that scores clinical accuracy and safety on every prompt change.
The client
Rosewood Health AI — Remote (US). A fine-tuning & evals build we delivered end to end.
The challenge
Stand up an automated evaluation pipeline for their patient-intake assistant: build a graded test set, wire LLM-as-judge scoring, and gate deploys on the results in CI.
What we built
- Assemble a labeled eval set with domain experts and rubrics
- Implement LLM-as-judge + rule-based scoring for accuracy and safety
- Integrate the harness into CI so regressions block release
The stack
Stack: an eval framework (Braintrust/Promptfoo/custom), LLM-as-judge design, and CI integration.
Want something like this?
Tell us what you're building and we'll scope it — most projects start within a week.