Big Data & Distributed Computing

Process datasets too large for one machine using Spark, Databricks and the Hadoop ecosystem.

No listings found

There are currently no listings in the Big Data & Distributed Computing category.

Are you interested in Big Data & Distributed Computing? Be the first to add listings in this category!

When data outgrows a single machine, the tools and the mental model both change. These courses teach distributed storage and processing: Apache Spark with the DataFrame API and Spark SQL, partitioning and shuffle behaviour, and the performance tuning that separates a job running in minutes from the same job running for hours. You will meet the Databricks platform, the Delta Lake, Iceberg and Hudi table formats, and the Hadoop and Hive components still running in many enterprises. Relevant when you are working at terabyte scale or joining a team whose existing platform is built on Spark.