This session is a mini-training about Apache Spark architecture and coding.
This session feeds the listeners the skills enough to work as a Junior Data Scientist.
Topics to be covered:
Architecture components of Spark;
Practical ways to run Spark tasks;
Distributed data formats for Big Data, column- vs row-oriented data formats, when to use each one;
Live coding for solving practical Big Data tasks.
Note: below you can find the Jupiter notebook file demonstrated together with the slides - https://gist.github.com/klimenkoOleg/...
Speaker: Oleg Klimenko
Follow us:
MJC Channel: @MJCtalks
Miniq Channel: @miniq330
JaCoV Community page: https://community-z.com/communities/j...