
Spark on Hadoop — PySpark Job Development & Cluster Submission
Delivery in
4 days
- Views 1
Amount of days required to complete work for this Offer as set by the freelancer.
Rating of the Offer as calculated from other buyers' reviews.
Average time for the freelancer to first reply on the workstream after purchase or contact on this Offer.
What you get with this Offer
I will develop and deploy a PySpark data processing job for your Hadoop cluster — covering Spark DataFrame API usage for your transformation logic, broadcast join optimisation for dimension table lookups, partition management for large datasets, Parquet or ORC output writing with appropriate partitioning, YARN cluster mode submission configuration, and a spark-submit script with tuned executor memory, core, and shuffle partition settings. PySpark jobs submitted to YARN without executor configuration tuning consistently underperform and OOM — the default settings are almost never optimal for your specific data volume and transformation complexity.
The job covers SparkSession configuration, DataFrame transformation pipeline for your specified logic, broadcast hint for small table joins, repartition and coalesce strategy for output file count management, checkpoint configuration for iterative jobs, spark-submit script with YARN cluster mode settings, and a Jupyter notebook version for interactive development and debugging.
Designed for data engineering teams processing large datasets on Hadoop clusters who need PySpark jobs built and configured correctly for their specific cluster resources and data characteristics.
The job covers SparkSession configuration, DataFrame transformation pipeline for your specified logic, broadcast hint for small table joins, repartition and coalesce strategy for output file count management, checkpoint configuration for iterative jobs, spark-submit script with YARN cluster mode settings, and a Jupyter notebook version for interactive development and debugging.
Designed for data engineering teams processing large datasets on Hadoop clusters who need PySpark jobs built and configured correctly for their specific cluster resources and data characteristics.
What the Freelancer needs to start the work
Please describe the data transformation required (inputs, logic, and outputs), your Hadoop cluster node count and memory/core configuration, your data volume, your input and output file formats, and your scheduling requirements.
We collect cookies to enable the proper functioning and security of our website, and to enhance your experience. By clicking on 'Accept All Cookies', you consent to the use of these cookies. You can change your 'Cookies Settings' at any time. For more information, please read ourCookie Policy
Cookie Settings
Accept All Cookies