Skip to content
This repository was archived by the owner on Apr 24, 2024. It is now read-only.

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Data-Pipeline

Data Pipeline for processing large loads of course data from the Web Scraper

Current Plans

Make a proof-of-concept pipeline using PySpark to focus on finding viable data transformations to the different Postgres tables

Future Plans

Potentially migrate to Scala (mostly because Spark code is natively written in Scala)

Why Spark

There's a lot of course data to process, not sure if Pandas would be enough. Also, converting from Dataframe to SQL table should be easy once data transformations are figured out

About

Data Pipeline for processing large loads of course data from the Web Scraper

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages