In advance of my session at PyData London (https://pydata.org/london2018/schedul...) the video of which is now up at • Keynote: Making the Big Data ecosyste... , I take a look at the live cycle of an open source contribution by focusing on Apache Spark, starting with finding the right issue, then live coding a solution, and finally, a code review. This is an excellent way to effectively contribute to popular open source projects like Apache Spark while building your big data experience.