AI · Healthcare ops · Expert review
I grade the models that talk about my work
Thirteen years in healthcare operations. When AI answers those prompts, I say whether it’s right — and I build the tests that keep it honest.
I build agents. I also catch the answers that sound right and aren't — the denial reason that almost fits, the KPI an AVP would act on for the wrong reason, the policy paraphrase that invents a rule. That judgment is the job: read the model's answer against real ops, mark what fails, write the answer training should copy, and design the tasks that prove the model can hold the line.
What I actually do
- Read the answer like an ops leadWould I send this to a manager? Is it accurate, clear, and professional — or does it just sound confident?
- Mark what’s betterI annotate the output and rank the options so training knows which answer to prefer, and why the weaker one missed.
- Write the answer I’d actually giveExpert answers to healthcare ops questions — the ones that become training examples. Straight, usable, not brochure language.
- Call the missWrong denial logic, soft policy language, a KPI that would mislead. I flag it and write the correction so the next pass doesn’t repeat it.
- Build the test from real workI design the scenario and the prompt from ops I’ve run, attach realistic input files, and write a rubric tight enough to grade the model without guessing.
- Run it, then tighten itPut the task against the model. Take reviewer feedback. Fix the rubric. Run it again.
- Do the brief as writtenFormats, uploads, instructions — followed precisely. Same discipline I’d expect on a go-live workbook.
- Same bar, every taskQuality and pace through the whole sprint. Last answer gets the same standard as the first.
Same rule as the rest of my AI work
In the Agent Control Room, agents propose and a human still signs. Evaluation is that gate on training data — catch the subtly wrong before it gets rewarded. The argument is in the RCM research brief. This page is the practice: review, corrections, and tests built from work I've actually done.
Roles and intros on LinkedIn.