Role Summary
We are seeking a highly motivated and skilled Big Data Platform Engineer to design, develop, and maintain scalable data processing platforms and pipelines. The ideal candidate should have strong expertise in Python, Apache Spark, SQL, and experience working in Linux-based environments with exposure to DevOps practices. The role involves building reliable, high-performance data solutions that support enterprise-scale analytics and business-critical applications.
Key Responsibilities
Data Engineering & Development
- Design, develop, and maintain scalable data pipelines using Python and Apache Spark.
- Implement efficient ETL/ELT processes for large-scale structured and unstructured datasets.
- Develop and optimize complex SQL queries, data models, and transformations.
- Ensure data quality, integrity, and reliability across the platform.
Platform Operations
- Work with Linux-based environments for deployment, troubleshooting, and performance tuning.
- Develop and maintain shell scripts for automation and operational tasks.
- Monitor and optimize Spark jobs for performance, scalability, and resource utilization.
DevOps & Automation
- Implement CI/CD pipelines and deployment automation.
- Participate in infrastructure provisioning, monitoring, and release management activities.
- Collaborate with DevOps teams to improve platform reliability and operational efficiency.
Collaboration & Governance
- Work closely with Data Architects, Product Owners, and Business Stakeholders.
- Participate in code reviews and ensure adherence to engineering best practices.
- Create and maintain technical documentation and operational runbooks.
Mandatory Skills
Core Big Data Platform Skills
- Strong programming experience in Python
- Hands-on expertise with Apache Spark (PySpark preferred)
- Strong SQL development and query optimization skills
- Apache Airflow for workflow orchestration and scheduling
- Kubernetes (AKS/SKE) for container orchestration and deployment. AKS/EKS/GKE also fine.
- MinIO / S3 Compatible Object Storage like AWS S3 or other S3-compatible object storage experience
- Apache Iceberg/Delta Lake / Apache Hudi
Technical Competencies
- Data Pipeline Development
- ETL/ELT Processing
- Data Modelling
- Performance Tuning
- Version Control (Git)
- Agile/Scrum Delivery Model