This job expired on 16th July 2026
AI Prompt & Evaluation Specialist

AI Prompt & Evaluation Specialist
Our engineering team builds and runs the AI pipelines — what we need is someone who owns the quality of what comes out of them. You'll write the prompts behind our LLM workflows and prove they work by testing them rigorously against grounded, real-world data. Existing workflows classify and extract structured data from care-enquiry call transcripts and emails; many more use cases planned (CV parsing, Review summaries etc.) will follow.
About Us:Tomorrow's Guides is a well-established, ambitious, and successful website company, operating leading care sector websites: carehome.co.uk, homecare.co.uk, and daynurseries.co.uk. Combined, our websites feature over 40,000 business profiles, host nearly 250,000 online reviews, and attract more than 22 million visitors annually.
This is a remote role, but you will need to spend time in our offices in Hungerford from time to time when required by the business so it’s important that you live within a commutable distance.
About this role
What you'll do
Write, version and continually refine the prompts that power our LLM features
Work closely with subject-matter experts to understand the domain in depth, then encode that expertise as precise, unambiguous rules (our workflows turn on distinctions that carry real commercial and compliance weight)
Build and maintain "golden" evaluation datasets from real examples, with correct expected outputs
Run systematic evals — measuring accuracy, catching regressions, comparing prompt and model variations head-to-head
Diagnose why outputs fail — ambiguous instructions, edge cases, messy input, model limitations — and fix them
Define measurable criteria for what "good" looks like on each field and each use case
Keep large, evolving prompts coherent and internally consistent as rules accumulate
Partner with our engineers, who own the pipelines, to get your prompts shipped and monitored
What we're looking for (essential)
Hands-on experience writing and refining LLM prompts, with a structured rather than trial-and-error approach
A genuine eye for evaluation: you think in test cases, edge cases and measurable criteria, not vibes
Comfort producing and reasoning about structured outputs — JSON schemas, strict enumerated values, conditional/nullable fields, type discipline
Skill at eliciting tacit knowledge from experts and turning fuzzy real-world judgements into rules a model can follow consistently
An instinct for how messy real-world input breaks things — e.g. speech-to-text errors, homophones, spelled-out corrections, inconsistent formatting
Precision with written English; you understand how small wording changes shift model behaviour
A responsible approach to sensitive personal data — our evaluation sets are built from real records containing names, contact details and health information (special-category data under UK GDPR), and you'll handle them accordingly
Enough technical familiarity to work alongside engineers and understand how a prompt fits into a pipeline
Nice to have
Experience with eval tooling or frameworks, or building evaluation harnesses
Background in a regulated or high-stakes domain (health, legal, finance, insurance)
A background in linguistics, QA, technical writing, or data analysis
Familiarity with LLM APIs (OpenAI/Azure OpenAI, Anthropic) and how model or version changes affect outputs
Salary and Benefits:
Salary: Dependent on experience
£4,000 per annum discretionary company bonus scheme
25 days annual leave + bank holidays
6% employer pension contribution
Access to free perks and discounts through Perkbox
Long Service Awards
Cycle to Work Scheme
Company and Team nights out


