Help define and continuously iterate data-quality standards, judgment rules, and acceptance criteria for foundation-model training and evaluation.
Review and judge data outputs against quality standards, identify issues, close correction loops, and ensure accuracy and consistency.
Participate deeply in the adjudication and discussion of complex, ambiguous, and borderline cases; convert conclusions into reusable judgment rules and knowledge assets.
...
Model Quality & Workflow Design: Design, manage, and optimize end-to-end workflows to improve machine model performance in content policy enforcement — including signal detection, rejection accuracy, and leakage reduction. Develop training data pipelines, QA processes, and performance tracking systems aligned to model improvement goals.
Adversarial Testing & Risk Identification: Conduct structured adversarial testing on AI models, features, and content policies to surface vulnerabilities, edge cases, and emerging risk trends. Explore model behaviour across contexts and user journeys to identify failure modes not captured in standard evaluations.
Root Cause Analysis & Error Optimization: Conduct structured root cause analysis (RCA) on model errors — including overkills, leakages, and misclassification — and translate findings into actionable model improvement recommendations. Partner with Algo and product teams to close root causes through memory insertion, threshold adjustments, rewrite rules, or policy iteration.
...
Design, build and scale testing capabilities using AI agents – using GitHub Copilot Agent Mode on the GitHub Enterprise platform and the Microsoft AI Foundry stack – to automate testing tasks such as test case generation, test data creation, defect triage, and regression analysis.
Develop and run AI evaluation (AI eval) frameworks and metrics to measure the accuracy, reliability, and performance of AI agents and Gen-AI-assisted testing workflows and iterate on prompts and agent design based on eval results.
Write efficient test automation code using coding standards and best practices.
...
We are looking for an AI Infra Engineer to support the development and operation of AI infrastructure platforms. The ideal candidate will have a strong background in Backend Engineering, Machine Learning Engineering, or MLOps, with hands-on experience in AI model training/inference deployment, cloud infrastructure, and GPU computing environments.
This role will work closely with algorithm teams to enable efficient utilization of AI computing resources, optimize AI workloads, and ensure reliable deployment of AI training and inference services.
...
Design and develop artificial intelligence (AI) and machine learning (ML) systems leveraging existing cloud AI services.
Design and build scalable data pipelines to support model training and production with DevOps & MLOps.
Customize and apply Deep Learning and Gen AI models for use cases based on the business needs, data availability, system and infrastructure requirements including edge devices and High Performance Computers (HPCs).
...
Lead a team of young and bright MLOps Engineers in conducting research, design, and prototype MLOps practices, tools, and frameworks to expedite the machine learning lifecycle at MoneyLion.
Work in collaboration with Data Scientists, Data Engineers, and AI/ML Engineering teams to enhance cross-team workflows and reduce time to deployment.
Improve the level of automation in the development and deployment processes of machine learning.
...
Apply AI-first engineering practices across discovery, requirements, design, development, testing, deployment, operations, and continuous improvement.
Create clear functional and technical specifications, architecture decisions, NFRs, design notes, and acceptance criteria that guide high-quality delivery.
Design, develop, test, deploy, and maintain modern full-stack, cloud-native applications that are secure, scalable, resilient, observable, and maintainable.
...
Design, build and scale testing capabilities using AI agents – using GitHub Copilot Agent Mode on the GitHub Enterprise platform and the Microsoft AI Foundry stack – to automate testing tasks such as test case generation, test data creation, defect triage, and regression analysis.
Develop and run AI evaluation (AI eval) frameworks and metrics to measure the accuracy, reliability, and performance of AI agents and Gen-AI-assisted testing workflows and iterate on prompts and agent design based on eval results.
Write efficient test automation code using coding standards and best practices.
...