Step-by-step Testing, Reviewing and Publishing UMP Dataset - Engineering - LinkedIn Corporate Wiki
UMP Recommends that you VALIDATE your existing dataset LOCALLY if you are performing ANY CHANGE on your dataset or corresponding files. UMP Recommends that you TEST your existing dataset ON HADOOP CLUSTER if you are performing ANY CHANGE on your dataset OTHER THAN METADATA CHANGES such as owner list, description, tags etc. This section provides guidance on running a testing flow based on your local UMP dataset and publishing it to production. At this point, we assume that you have authored a dataset.conf and corresponding files. We will validate our dataset in two steps: In the first step we will validate the dataset.conf. This validation happens on the local computer. In the second step we will validate the entire UMP flow. Here, an azkaban project is created and the scripts are executed on hadoop cluster. This steps validates that the configurations in dataset.conf are complete and consistent. Think of it as a ‘compiler’ that looks into your metadata file and reports any problems or
UMP Recommends that you VALIDATE your existing dataset LOCALLY if you are performing ANY CHANGE on your dataset or corresponding files. UMP Recommends that you TEST your existing dataset ON HADOOP CLUSTER if you are performing ANY CHANGE on your dataset OTHER THAN METADATA CHANGES such as owner list, description, tags etc. This section provides guidance on running a testing flow based on your local UMP dataset and publishing it to production. At this point, we assume that you have authored a dataset.conf and corresponding files. We will validate our dataset in two steps: In the first step we w
Explore this link on the map →