Prompt Engineering Lead
We are looking for a Prompt Engineering & AI Evals Lead to own prompt engineering - one of the core service categories on the REWORK platform - and, internally, the evaluation rigor behind every AI decision REWORK makes about a person's work. Proof-of-work submissions get auto-verified against rules and thresholds; skill assessments are AI-generated and AI-scored through SwitchAsses; matching and the Rework Score lean on model output. Bad prompts here mean a false auto-approval or an unfair rejection - real consequences for real professionals. You will build the eval harnesses, golden test sets, and prompt-version pipelines that keep these systems honest, and do the same prompt-design and red-teaming work as a client-facing deliverable for businesses that need their own AI systems evaluated and hardened.
Core Responsibilities
Eval Infrastructure
- Build eval harnesses and golden test sets for auto-verification rules, matching/scoring prompts, and assessment generation
- Version prompts and track accuracy, false-positive, and false-negative rates release over release
- Design regression tests so a prompt tweak can't silently degrade auto-approval accuracy
Prompt Design & Hardening
- Write and refine system prompts for auto-verification, SwitchAsses assessment scoring, and Rezy
- Red-team prompts for jailbreaks, prompt injection, and edge cases before they reach production
- Design prompts as a client deliverable - eval suites and hardening work for business AI systems
Measurement & Reporting
- Quantify how auto-verification and scoring accuracy move with each prompt or model change
- Flag drift when a model update changes behavior on the existing eval set
- Document prompt versions, eval results, and hardening decisions so the reasoning is auditable later
Required Qualifications
Required Experience
- Shipped production prompts with a measurable eval process behind them, not just prompt tuning by feel
- 1+ years hands-on with the OpenAI or Anthropic APIs and at least one eval/testing framework
System Design & Problem-Solving
- Thinks in false-positive vs false-negative tradeoffs and designs thresholds accordingly
- Rigorous about reproducibility - same eval set, same scoring, comparable results across prompt versions
Documentation & Nice-to-Haves
- Documents eval methodology and results so decisions are defensible, not just "it felt better"
- Nice to have: statistics/ML background, red-teaming experience, trust & safety or content-moderation work
Key Deliverables
Measured, versioned eval suites for auto-verification, scoring, and assessment promptsDocumented accuracy and drift tracking across prompt and model changesHardened prompts resistant to jailbreaks and prompt injectionClear, well-documented eval methodology and prompt-version history
Tech Stack & Skills
System-prompt designFew-shot / Chain-of-ThoughtEval harnesses (Promptfoo or similar)Golden test-set designA/B testingConstitutional AIRed-teaming / jailbreak hardeningMulti-step tool-call promptsPythonOpenAI / Anthropic APIsStatistics basics
Expectations
- Work asynchronously in a remote-first environment (Slack, email, documented reports)
- Operate highly independently - evaluated on proof-of-work and reliability of deliverables
- Collaborate during core hours, 10:00 AM – 4:00 PM Eastern Time
- Keep eval methodology, prompt versions, and results clear, well-documented, and easy to navigate