This hands-on, three-day program is dedicated to designing and tuning high-efficiency data processing pipelines that leverage PySpark, Pandas, and Polars within Kubernetes-based ecosystems.
Learners will gain a working knowledge of how Spark applications are orchestrated on Kubernetes, with a specific focus on how configuration choices at the application layer directly impact throughput, scalability, resource utilization, and operational expenditure. Key areas of optimization covered include executor dimensioning, memory distribution, dynamic resource allocation, partitioning methodologies, shuffle mechanics, mitigating small-file issues, and streamlined Parquet data handling.
Additionally, the curriculum tackles frequent bottlenecks associated with Pandas, such as memory constraints and out-of-memory exceptions, while presenting Polars as a high-performance alternative for specific processing tasks. Through interactive lab sessions, participants will learn to pinpoint performance and memory inefficiencies, evaluate various configuration approaches, and implement optimization strategies for realistic ETL and machine learning use cases.
The core objective of this training is practical decision-making: mastering the identification of system bottlenecks, selecting the right tool for the job, configuring Spark for maximum efficiency, and striking a balance between processing speed and infrastructure resource costs.
Read more...