Available courses
This course introduces Python programming for statisticians working with survey and census data. You will learn to set up reproducible analysis environments, load and clean data from various formats, handle missing values and duplicates, and combine datasets through merges and SQL queries.
The course covers exploratory data analysis techniques, data visualization for statistical reporting, and foundations of statistical computation including confidence intervals, hypothesis tests, and survey-weighted estimation.
You will also be introduced to predictive modeling with classification methods, and learn strategies for working with datasets that exceed memory limits, including chunked processing and modern file formats like Parquet.
This course builds on foundational Python skills to address more complex workflows in official statistics. You will learn object-oriented programming for maintainable code, unit testing, logging, and systematic debugging techniques.
The course covers how to structure reproducible pipelines using modular functions, configuration files, and data catalogs. You will work with statistical modeling using both statsmodels for inference and scikit-learn for prediction, including regression diagnostics and model validation.
You will also learn to analyse open-ended survey responses using embeddings, classification, and clustering with local NLP models. The course includes spatial analysis with GeoPandas, coordinate reference systems, spatial joins, raster processing, and zonal statistics for producing district-level indicators. Finally, you will practice techniques for scaling workflows to handle larger-than-memory datasets.