Develop backend services for AI features: Build and maintain scalable APIs and microservices in Java (Spring Boot) or Python that expose our generative AI capabilities - routing user queries to LLMs and other models, processing data, and returning results securely and efficiently.
Integrate generative AI technologies: Partner with data science and ML engineering to productionise models - wrapping them in reliable service interfaces, managing input/output formats, adding supporting data flows such as external API calls and response caching, and ensuring they scale.
Ensure performance and reliability: Own the quality of the services you build through unit and integration tests, profiling and removing bottlenecks, and monitoring and alerting (CloudWatch, Prometheus). Troubleshoot production issues and keep improving logging and observability....
Design, build, and operate production-grade AI and ML services, APIs, libraries, and reusable platform components.
Own end-to-end AI engineering workflows, including data preparation, model integration, evaluation, deployment, monitoring, and continuous improvement.
Support cloud migration efforts, transitioning on-premises data workflows to cloud-based platforms such as Databricks, Cloudera, Amazon AWS, etc....
Design, build, and operate agentic applications supporting AI-factory planning, commissioning, validation, workload onboarding, benchmark analysis, model and recipe optimisation, scheduling, operations, maintenance, incident response, and continuous improvement.
Define reference architectures for single-agent, multi-agent, workflow-based, eventdriven, and human-in-the-loop agentic systems.
Build orchestration workflows using appropriate agent frameworks and libraries, such as LangGraph, LangChain, LlamaIndex, Microsoft AutoGen, Semantic Kernel, CrewAI, PydanticAI, Haystack, DSPy, or equivalent custom-built frameworks....