You obsess over the exact instructions we give to models β and the data we retrieve for them, and the evaluations we grade them with. You know when to prompt, when to fine-tune, and when to redesign the workflow.
What you'll build
- Production prompts and agent loops across multiple product lines.
- Evaluation datasets that quantify quality per feature.
- Structured-output pipelines with reliable schema conformance.
- Automation to detect regressions on every prompt change.
Requirements
- 3+ years software engineering.
- Deep intuition for what LLMs can and can't do β earned in production.
- Fluent in Python or TypeScript.
- Rigorous about eval, not just vibes.
- Right to work in the UK.
#J-18808-Ljbffr