Introduction to Big Data with Python

From your first Python program to a full Big Data pipeline in PySpark. No programming experience needed: you build up one step at a time, with a project in almost every unit.

Introduction to Big Data with Python
Beginner to intermediate

Why this course is built this way

Most Big Data courses start with Spark. This one starts with Python and conventional data processing: files, pandas and SQL. Then it asks a simple question: what happens when the data no longer fits on your computer? By the time you use Spark, you know exactly what problem it solves.

  • Basic computer literacy is the only prerequisite
  • You work in VS Code with Python, the same set-up used in industry
  • Each stage ends with a mini-project you can keep
  • The course closes with an end-to-end ETL project
Sign up

14

Units

16

Learning outcomes

7

Mini-projects

1

Final project

raw data to report

What you will be able to do

The 14 units

Unit 1

Development environment and introduction to Python

Install Python and VS Code, run your first program and learn to read error messages.

Unit 2

Python fundamentals

Variables, conditions, loops, functions and exceptions. Project: a student grade analyser.

Unit 3

Python data structures

Lists, tuples, dictionaries, sets and comprehensions. Project: an in-memory student management system.

Unit 4

Working with files

Text, CSV and JSON. Project: a data processor that turns raw files into clean reports.

Unit 5

Data analysis with pandas

DataFrames, filtering, grouping, aggregation and merging. Project: an e-commerce data analysis.

Unit 6

Databases and SQLite

Tables, keys and relationships. Project: design and build a school database.

Unit 7

SQL

SELECT, aggregation, GROUP BY, JOIN and data changes. Project: school analytics in SQL.

Unit 8

Python and SQLite

Send SQL from Python and work with the results. Project: a student database application.

Unit 9

Introduction to Big Data

Volume, velocity and variety. Process files of growing size and see where a single computer struggles.

Unit 10

Distributed computing

Clusters, nodes, partitions and fault tolerance. Simulate distributed processing in plain Python.

Unit 11

Introduction to PySpark

SparkSession, DataFrames and schemas. Project: your first PySpark data analysis.

Unit 12

Spark DataFrames and transformations

Transformations, actions and lazy evaluation. Join customers, orders and products.

Unit 13

Spark SQL and Big Data formats

Query DataFrames with SQL and compare CSV with Parquet for size and speed.

Unit 14

ETL and final Big Data project

Ingest, clean, transform, store and analyse a global e-commerce dataset, then present your results.

How you are assessed

Marks are spread across the course, so steady work counts.

Python and files: 20%

Python exercises (10%) and the file-processing project (10%).

Data and databases: 25%

Pandas data analysis (10%) and the SQL and database project (15%).

Big Data: 25%

Big Data exercises (10%) and PySpark assignments (15%).

Final project: 25%

Your end-to-end pipeline: code, SQL, results, a README and a short presentation.

Participation: 5%

Taking part in live sessions and challenges.

Questions about this course

No. The course starts from zero: installing Python and writing your first program. Basic computer literacy is enough.

A computer on which you can install Python, VS Code, SQLite and PySpark. All of them are free. Unit 1 walks you through the installation.

In blended format: live sessions with a teacher, plus lessons, exercises and projects on our learning platform between sessions. Live sessions are timetabled for each group, and you receive the timetable when you join.

Start dates are set for each group. Sign up or contact us and we will tell you when the next group starts.

Contact us for the current fee and the payment options.

Start with your first Python program

Create your account and join the next group.