Step by Step Guide
https://www.linkedin.com/pulse/server...
Code :
https://github.com/soumilshah1995/ins...
Excited to share my latest video where I demonstrate how to use AWS Lambda to generate Parquet files and upload them to S3. With this approach, you can easily set up a serverless data engineering pipeline that can handle large volumes of data in a scalable and cost-effective way.
In the video, I walk you through the entire process step-by-step and cover topics such as error handling, data transformation, and working with Arrow tables. I also touch on using SQS as a message queue to decouple the Lambda function from the data producer.
If you're interested in learning more about serverless data engineering or want to see how to implement this pipeline yourself, check out the video and let me know what you think in the comments below. 👇
Looking forward to hearing your thoughts!