AI Evaluation Engineer (LLMOps / AI Platform)
We're partnering with a leading technology company to hire an AI Evaluation Engineer who will play a key role in building scalable evaluation frameworks for AI agents and LLM-powered products. This is an opportunity to work on cutting-edge GenAI platforms, improving AI quality, reliability, and observability across multiple production use cases.
What You'll Be Doing
- Partner with AI product teams to improve the quality and performance of AI agents.
- Design and maintain evaluation datasets, regression test suites, automated evaluation pipelines, and quality dashboards.
- Develop evaluation methods including LLM-as-a-Judge, multi-turn conversations, tool/MCP evaluation, agent trajectory analysis, and SOP adherence.
- Build reusable platform capabilities and developer tooling that support multiple AI products.
- Improve AI observability by instrumenting distributed systems and ensuring traceability for debugging and automated evaluation.
- Integrate evaluation workflows into CI/CD pipelines and model release processes.
- Conduct experiments across models, retrieval strategies, embeddings, and agent architectures to optimise quality, latency, and cost.
- Drive continuous improvements to AI evaluation tooling, documentation, and engineering best practices.
Requirements
- At least 2 years of relevant experience in AI platform engineering, LLMOps, AI evaluation, or related software engineering roles.
- Hands-on experience with most of the following technologies:
- EvalsHub, LangSmith, LangGraph/LangChain
- FastAPI, Temporal, Grafana, Redis (or similar technologies)
- React & TypeScript for developer-facing tools
- Experience building reusable AI platform capabilities across multiple products.
- Strong analytical and problem-solving skills, particularly in handling geo-spatial or large-scale data.
- Excellent stakeholder management and cross-functional communication skills.
- Detail-oriented with a data-driven mindset focused on continuous improvement.
Experience with one or more of the following:
- LLM-as-a-Judge frameworks
- Multi-turn or agent evaluation
- Tool/MCP evaluation
- Agent trajectory analysis
- Text-to-SQL or Text-to-DSL
- RAG, semantic search, embeddings, or model benchmarking
Why This Opportunity?
- Build production-grade AI evaluation systems used across multiple AI products.
- Work with modern GenAI, LLMOps, and observability technologies.
- Influence AI quality, reliability, and platform engineering at scale.
- Collaborate closely with AI researchers, platform engineers, and product teams in a highly technical environment.
If you're passionate about AI infrastructure, LLM evaluation, and building scalable AI platforms, I'd be happy to share more details.