ST446      Half Unit
Distributed Computing for Big Data

This information is for the 2026/27 session.

Course convenor

Dr Marcos Barreto

Availability

This course is available on the MPA in Data Science for Public Policy, MSc in Applied Social Data Science, MSc in Data Science, MSc in Econometrics and Mathematical Economics, MSc in Financial Statistics, MSc in Financial Statistics (Research), MSc in Geographic Data Science, MSc in Health Data Science, MSc in Quantitative Methods for Risk Management, MSc in Statistics and MSc in Statistics (Research). This course is available with permission as an outside option to students on other programmes where regulations permit. This course uses controlled access as part of the course selection process. For information on controlled access courses, including eligibility, application processes, deadlines, and departmental contact details, please refer to the Controlled Access Courses webpage.

How to apply: Please be advised that spaces on this course will be extremely limited, so early application is advisable. Priority will be given to students on the MSc Data Science.

Students from any other programmes should submit a short statement indicating a) any experience with cloud computing and/or big data tools, and b) why they think the course is suitable for them given their background knowledge.

Deadline for application: check for the relevant webpages.

Course lecturers will aim to make initial offers to students on LSE For You by the relevant dates. 

For queries contact: Stats-Msc@lse.ac.uk

This course has a limited number of places (it is controlled access) and demand is typically high. This may mean that you are not able to get a place on this course, although there is opportunity to audit the course. The MSc in Data Science students are given priority for enrolment in this course.

Requisites

Basic knowledge of Python or some other programming knowledge, including shell/terminal-based commands is desirable. Basic knowledge of computer architecture (memory, storage, communication) is also desirable but will be addressed during the course. 

Course content

The course covers principles of distributed processing systems for big data, including distributed storage systems (such as Hadoop); distributed computation models (such as MapReduce); resilient distributed datasets (Spark RDDs); structured querying over large datasets (Spark Dataframes and SQL); stream data processing systems (Kafka and Kinesis); and scalable machine learning models (Spark MLlib and TensorFlow)


The course makes use of Google Cloud Platform and Amazon AWS Academy learning resources and enables students to learn about the principles and gain hands-on experience in working with industry standard cloud computing technologies. Through weekly exercises and course project work, student can gain experience in performing data analytics tasks on their laptops and cloud computing platforms.
 

Teaching

15 hours of seminars and 20 hours of lectures in the Winter Term.

This course has a reading week in Week 6 of Winter Term.

Formative assessment

Students will be given weekly formative exercises to complement the topics discussed in the lecture and hands-on experimentation during seminar sessions. Some formative exercises are for self-study only, while others are expected to be submitted for formative feedback.

 

Indicative reading

Indicative reading

  • Triguero, I. and Galar, M. Large-Scale Data Analytics with Python and Spark: a hands-on guide to implementing machine learning solutions. Cambridge, 2023
  • Reis, J.; Housley, M. Fundamentals of Data Engineering: Plan and Build Robust Data Systems. O’Reilly, 2022
  • Shapira, G.; Palino, T.; Sivaram, R.; Pett, K. Kafka - The Definitive Guide: Real-Time Data and Stream Processing at Scale. O’Reilly, 2021 
  • Damji, J., Weing, B., Das, T., Lee. D. Learning Spark: Lightining-Fast Data Analysis, O’Reilly, 2nd Edition, 2020
  • White, T., Hadoop: The Definitive Guide, O’Reilly, 4th Edition, 2015

Additional reading:

  • Foster, I., Ghani, R., Jarmin, R. S., Kreuter, F., Lanie, J. (Eds.). Big Data and Social Science: Data Science Methods and Tools for Research and Practice. 2nd edition, CRC Press, 2021.
  • Huang, S., Deng. H. Data Analytics: A Small Data Approach. CRC Press, 2021.
  • Chambers, B,; Zaharia, M. Spark: the definitive guide. O’Reilly, 2018
  • Kleppman, M.; Riccomini, C. Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable Systems. 2nd edition, O’Reilly, 2026
  • Eagar. G. Data Engineering with AWS: Acquire the skills to design and build AWS-based data transformation pipelines like a pro. 2nd edition, 2023.
  • Wijay, A. Data Engineering with Google Cloud Platform: A practical guide to operationalizing scalable data analytics systems on GCP. Packt, 2022.
  • Apache Spark Documentation https://spark.apache.org/docs/latest
  • Apache TensorFlow Documentation https://www.tensorflow.org

Assessment

Problem sets (20%) in Winter Term Week 6.

Problem sets (20%) in Winter Term Week 11.

Project (60%) in May.

This component of assessment includes an element of group work.

Summative assessments: a problem set submitted in WT Week 6 (20%), a problem set submitted in WT Week 11 (20%), a project (60%) given in the WT and submitted at the beginning of ST. 


Key facts

Department: Statistics

Course study period: Winter Term

Unit value: Half unit

FHEQ level: Level 7

Total students 2025/26: 88

Average class size 2025/26: 29

Controlled access 2025/26: Yes
Guidelines for interpreting course guide information

Course selection videos

Some departments have produced short videos to introduce their courses. Please refer to the course selection videos index page for further information.

Personal development skills

  • Self-management
  • Team working
  • Problem solving
  • Application of information skills
  • Communication
  • Application of numeracy skills
  • Specialist skills