Skip to content

Latest commit

 

History

History
414 lines (339 loc) · 21.4 KB

File metadata and controls

414 lines (339 loc) · 21.4 KB

CodeBook.md

Getting and Cleaning Data - UCI HAR Dataset

This file describes the sequence of R language operations used to massage the raw data into a tidy and useful format.

We wish to end up with one table, where each row is an experiment, and the columns show subject number, activity, experiment group, and observations.

We will use the dplyr library, so before we start and before running any scripts we do:

  1. install.packages("dplyr") library(dplyr)

    There is no script for this, we can just type it into the console, along with any other packages and libraries we may need.

    Note that the script files in the work directory are script2.R through script6.R, executing these in order will result in the two tables "observations" and "raw_observations", which are cross referenced by the value in the "experiment_id" column.

  2. The first thing we do is read everything into R. The raw data is in ordinary text files, so we can just use read.table("table_name.txt"). For ease of reference we'll call each table by the same name of the associated source file.

    A. Copy all source files into a single folder. (They are all uniquely named, so you can safely do this). The GitHub repository contains only a single folder with everything in it. This will be our working directory.

    B. After cloning the repository, setwd('<working_directory>') in R, where working_directory is wherever you put the files on your own machine.

    C. For each file, read it into R and create a corresponding table using table_name <- read.table("filename.txt"), where table_name = filename. The code for doing this is contained in the script "script2.R", which is the first script referenced in the "run_analysis.R" sequence. (If you just want to test the entire sequence end to end, use the command 'source("run_analysis.R")' (without the single quotes) in R. The complete list of files being processed is:

    activity_labels.txt body_acc_x_test.txt body_acc_x_train.txt body_acc_y_test.txt body_acc_y_train.txt body_acc_z_test.txt body_acc_z_train.txt body_gyro_x_test.txt body_gyro_x_train.txt body_gyro_y_test.txt body_gyro_y_train.txt body_gyro_z_test.txt body_gyro_z_train.txt features.txt (note that features_info.txt has been omitted because it's not data) subject_test.txt subject_train.txt total_acc_x_test.txt total_acc_x_train.txt total_acc_y_test.txt total_acc_y_train.txt total_acc_z_test.txt total_acc_z_train.txt X_test.txt X_train.txt y_test.txt y_train.txt

    D. We do a str() on each table to verify it's been named correctly and properly read in, and to verify that it looks like we expect.

  3. Now we need to label all the rows and columns so we can cross-reference them later. We begin with the features (which are used in all of the observations files) and the activities.

    A. features <- rename(features, V1=feature_id, V2=feature_name) activity_labels <- rename(activity_labels, V1=activity_id, V2=activity_name)

    Before labeling the observations, we combine the results from the testing and training data, using rbind - and before we can do that, we have to label each table according to the group it's in. We don't have to do this for the raw accelerometer and gyroscope tables, because it's redundant. We also don't have to do this for the subject or y tables, because we're going to merge them laster with the X table. So we just put the group information into the X table.

    B. X_train$group <- "Train" X_test$group <- "Test"

    Now we can combine the tables. By convention the training data comes first, in each table, this way all our indices match.

    C. body_acc_x <- rbind(body_acc_x_train, body_acc_x_test) body_acc_y <- rbind(body_acc_y_train, body_acc_y_test) body_acc_z <- rbind(body_acc_z_train, body_acc_z_test) body_gyro_x <- rbind(body_gyro_x_train, body_gyro_x_test) body_gyro_y <- rbind(body_gyro_y_train, body_gyro_y_test) body_gyro_z <- rbind(body_gyro_z_train, body_gyro_z_test) total_acc_x <- rbind(total_acc_x_train, total_acc_x_test) total_acc_y <- rbind(total_acc_y_train, total_acc_y_test) total_acc_z <- rbind(total_acc_z_train, total_acc_z_test) subject <- rbind(subject_train, subject_test) X <- rbind(X_train, X_test) y <- rbind(y_train, y_test)

    D. We would also like to add an experiment_id column to each table, even though strictly speaking this is not necessary, but it will facilitate our merging later on. To do this we create an additional table with experiment ID's numbered from 1 to 10299, as follows:

    experiment <- data.frame("experiment_id"=seq(1:10299))

    and then bind this new column into all our tables

    body_acc_x <- cbind(body_acc_x, experiment) body_acc_y <- cbind(body_acc_x, experiment) body_acc_z <- cbind(body_acc_x, experiment) body_gyro_x <- cbind(body_gyro_x, experiment) body_gyro_y <- cbind(body_gyro_y, experiment) body_gyro_z <- cbind(body_gyro_z, experiment) total_acc_x <- cbind(total_acc_x, experiment) total_acc_y <- cbind(total_acc_y, experiment) total_acc_z <- cbind(total_acc_z, experiment) subject <- cbind(subject, experiment) X <- cbind(X, experiment) y <- cbind(y, experiment)

    And finally we can relabel all of the columns in our new tables. We only have to relabel the first 561 columns containing the processed observations in the X and y tables, the group and experiment_id don't need to be relabeled.

    E. subject <- rename(subject, "subject_id"=V1) y <- rename(y, "activity_id"=V1)

    To rename the X dataset and the raw datasets, we need a named vector with the old numbered names beginning with V and the new names from the 'features' table. And we wish to include the two new columns that we added, called "group" and "experiment_id"

    F. feature_names <- features$feature_name feature_names <- c(feature_names, c("group", "experiment_id")) names(X) <- feature_names

    There's not a whole lot we can (or want to) do with the 128 raw accelerometer and gyro readings for each experiment, since we're not really told exactly what they are (and therefore we can't rightfully label them any better than they already are).

  4. Now we have the data by experiment, with the experiments grouped into testing and training. We would like one table with subject, activity ID, and 561 processed observations for each experiment. We would like another table with experiment ID and 128 * 9 raw accelerometer and gyroscope data points for each experiment. We have everything we need for the first table, but to get the second table we have to do some further renaming.

    First, let's start building our primary table. The first thing we can do is merge the subject and activity tables to get a table (S,A) of subjects and activities for each experiment.

    A. observations <- merge(subject,y,by="experiment_id")

    Now we merge this table with the X table which contains the processed observarions.

    B. observations <- merge(observations,X,by="experiment_id")

    This is fine, except the group identifier is in the last column and we would like to move it up front in between the activity ID and the first observation. This means we want to move column "group" to the fourth position in the observations table

    C. new_order = append(setdiff(names(observations), "group"), "group", after=3)

    We can verify we have the right order by looking at new_order. Now,

    observations <- observations[, new_order]

    We can look at the result with head(observations).

    This gives us our first table, called "observations", where we have each experiment defined in terms of the subject, activity, group, and the 561 processed data points.

  5. Now we wish to cross reference each experiment in our first table, with the underlying raw data from the accelerometers and gyroscopes, by experiment ID.

    First we note that the names of the 128 raw data columns are the same for each of our 9 tables, and we'll need to change this before merging the tables, because we need a unique name for each data value. The most convenient way of doing this (and still retaining the identifying info) is to prepend each column name with the table name.

    So we can run the following sequence on each of the 9 raw tables (using body_acc_x as an example - and note that we need 129 entries for the prefix vector instead of 128, because we have the extra column called "experiment_id"):

    A. body_acc_x_names = names(body_acc_x) body_acc_x_prefixes = rep("body_acc_x", times=129) new_names = paste(body_acc_x_prefixes, ".", body_acc_x_names, sep="") names(body_acc_x) <- new_names

    Note that when we do this, the very last column "experiment_id" is also prepended with the table name, which is not needed or wanted. So we rename this column to just "experiment_id" and move it up front.

    B. body_acc_x <- rename(body_acc_x, "experiment_id"=body_acc_x.experiment_id) body_acc_x_names <- names(body_acc_x) new_order <- append(setdiff(body_acc_x_names, "experiment_id"), "experiment_id", after=0) body_acc_x <- body_acc_x[, new_order]

  6. Now we have each of our 9 raw tables with the experiment ID up front, and the columns named uniquely by table name. So now we can simply merge them into one large table, by experiment ID. We'll do it so the raw accelerometer data is first, followed by the processed accelerometer data, followed by the gyro data.

    raw_observations <- merge(total_acc_x, total_acc_y, by="experiment_id") raw_observations <- merge(raw_observations, total_acc_z, by="experiment"id") raw_observations <- merge(raw_observations, body_acc_x, by="experiment_id") raw_observations <- merge(raw_observations, body_acc_y, by="experiment_id") raw_observations <- merge(raw_observations, body_acc_z, by="experiment_id") raw_observations <- merge(raw_observations, body_gyro_x, by="experiment_id") raw_observations <- merge(raw_observations, body_gyro_y, by="experiment_id") raw_observations <- merge(raw_observations, body_gyro_z, by="experiment_id")

    This is still somewhat inconvenient because we have 128 observations for the x axis followed by 128 observations for the y axis and then 128 observations for the z axis. However we are not told exactly how these 128 observations relate to each other, whether they're sequential or whether they occur in groups of 3. Therefore the best we can do is leave the data as it is and let the user figure it out.

  7. Now we have the data in the desired tidy form. We have one single data frame containing the processed observations that our users will most likely want to see, and we have one additional data frame containing all the raw accelerometer and gyro data (which we assume almost no one will want to see), cross referenced by experiment ID. Everything is neatly and properly labeled. If we want to search the tables we can conveniently do so, and if we wish to see the raw data we can simply merge the two tables by experiment ID.

    To complete the instructions for this assignment, we will make use of the tables we just created. We are told to merge the training and test sets, we've done that. We're also told to extract only the mean and standard deviation for each measurement, which we can now do, by taking the relevant columns out of the "observations" table. The easiest way to do this is to extract all the columns whose names contain either "mean" or "std".

    We are also told to use activity labels for each activity, and that means we have to massage our "observations" table so the activity labels get replaced with their names.

    As far as descriptive variable names, we will keep the original variable names in the "observations" table so they can be cross-referenced with the original source files. However we will create a new smaller table called reduced_observations where the column names can be tweaked to be as descriptive as possible. However since I'm not entirely clear on what these column names mean, I'm just going to leave them alone for now.

    We do these steps in inverse order. First we create a reduced_observations table containing only the experiment ID, subject ID, activity ID, and group. Then we create a separate table with only the columns labeled "mean" and "std", and then we merge the two tables.

    Optionally, we can run this same sequence on the original "observations" table, if we want the activity labels to be replaced by their names. (This step is optional and is not included in the Step 7 script, but is included in the run_analysis script - see Step 8 below).

    A. reduced_observations <- data.frame(observations$experiment_id, observations$subject_id, observations$activity_id, observations$group) names(reduced_observations) <- c("experiment_id", "subject_id", "activity_id", "group") reduced_observations <- merge(reduced_observations, activity_labels, by="activity_id") reduced_observations <- arrange(reduced_observations, experiment_id) reduced_observations$activity_id <- NULL new_order = append(setdiff(names(reduced_observations), "activity_name"), "activity_name", after=2) reduced_observations <- reduced_observations[, new_order]

    Now we have the beginnings of our reduced observations table. To continue, we create a new table from "observations" with only the columns whose names contain "mean" or "std", along with the experiment_id.

    B. mean_and_std <- select(observations, contains("experiment_id") | contains("mean") | contains("std"))

    And then we can simply merge the mean_and_std table back into reduced_observations.

    C. reduced_observations <- merge(reduced_observations, mean_and_std, by="experiment_id")

    This formulation will give us all the means first, followed by all the standard deviations. We assume that our users will be most interested in the means, so this is probably acceptable, especially given that we weren't instructed on the order of the information.

    Note that the script for this step is called "script7.R", and as per the instructions I've put all of the scripts in order into a script called "run_analysis.R". If you just run the "run_analysis.R" script you'll run all of the scripts script2.R thru script7.R in order.

    The final table required by the instructions will be called "reduced_observations".

  8. Optionally, we can now go back and fix the "observations" table so it also has the activity name instead of the activity ID. We do so as follows:

    observations <- merge(observations,activity_labels,by="activity_id") observations <- arrange(observations, experiment_id) observations$activity_id <- NULL new_order = append(setdiff(names(observations), "activity_name"), "activity_name", after=2) observations <- observations[, new_order]

  9. To complete the assignment, we now create a second independent data set containing the average of each variable for each activity and each subject (starting from the reduced_observations table, as per instructions), and write it out to a csv file called "averages.csv".

    To do this, we first have to rename all the columns in our reduced_observations table, because the existing names will be interpreted as functions by the R parser unless we do this. (We can then retroactively apply this to the original "observations" table if we wish).

    First we generate a vector with the desired column names and apply it to the reduced_observations table. To do this, we first generate a vector of existing names, then translate all the dashes to periods, and get rid of all the parentheses. There are also some commas we don't want, we'll change those to periods.

    A. ro_names <- names(reduced_observations) ro_names <- chartr("-", ".", ro_names) ro_names <- gsub("[()]", "", ro_names) ro_names <- gsub(',', '.', ro_names) names(reduced_observations) <- ro_names

    Next we create a new table with the desired means. It makes no sense to take the means of standard deviations, so we'll just leave them out and process the means of each variable. If we wanted to come up with standard deviations we would most likely derive them from the variables themselves, in which case we could add an SD= argument to each of the mean() functions.

    Done this way, we arrive at a table of the averages for each variable, which is what the instructions ask for.

    There are 53 variables in total, each of which becomes a column in the "averages" table. Each table entry represents the average value of the variable for the designated activity and subject. The table is grouped by activity first, and then by subject ID, as per the instructions.

    B. grouped <- group_by(reduced_observations, activity_name, subject_id) averages <- summarize(grouped, mean.tBodyAcc.X=mean(tBodyAcc.mean.X), mean.tBodyAcc.Y=mean(tBodyAcc.mean.Y), mean.tBodyAcc.Z=mean(tBodyAcc.mean.Z), mean.tGravityAcc.X=mean(tGravityAcc.mean.X), mean.tGravityAcc.Y=mean(tGravityAcc.mean.Y), mean.tGravityAcc.Z=mean(tGravityAcc.mean.Z), mean.tBodyAccJerk.X=mean(tBodyAccJerk.mean.X), mean.tBodyAccJerk.Y=mean(tBodyAccJerk.mean.Y), mean.tBodyAccJerk.Z=mean(tBodyAccJerk.mean.Z), mean.tBodyGyro.X=mean(tBodyGyro.mean.X), mean.tBodyGyro.Y=mean(tBodyGyro.mean.Y), mean.tBodyGyro.Z=mean(tBodyGyro.mean.Z), mean.tBodyGyroJerk.X=mean(tBodyGyroJerk.mean.X), mean.tBodyGyroJerk.Y=mean(tBodyGyroJerk.mean.Y), mean.tBodyGyroJerk.Z=mean(tBodyGyroJerk.mean.Z), mean.tBodyAccMag=mean(tBodyAccMag.mean), mean.tGravityAccMag=mean(tGravityAccMag.mean), mean.tBodyAccJerkMag=mean(tBodyAccJerkMag.mean), mean.tBodyGyroMag=mean(tBodyGyroMag.mean), mean.tBodyGyroJerkMag=mean(tBodyGyroJerkMag.mean), mean.fBodyAcc.X=mean(fBodyAcc.mean.X), mean.fBodyAcc.Y=mean(fBodyAcc.mean.Y), mean.fBodyAcc.Z=mean(fBodyAcc.mean.Z), mean.fBodyAcc.Freq.X=mean(fBodyAcc.meanFreq.X), mean.fBodyAcc.Freq.Y=mean(fBodyAcc.meanFreq.Y), mean.fBodyAcc.Freq.Z=mean(fBodyAcc.meanFreq.Z), mean.fBodyAccJerk.X=mean(fBodyAccJerk.mean.X), mean.fBodyAccJerk.Y=mean(fBodyAccJerk.mean.Y), mean.fBodyAccJerk.Z=mean(fBodyAccJerk.mean.Z), mean.fBodyAccJerk.Freq.X=mean(fBodyAccJerk.meanFreq.X), mean.fBodyAccJerk.Freq.Y=mean(fBodyAccJerk.meanFreq.Y), mean.fBodyAccJerk.Freq.Z=mean(fBodyAccJerk.meanFreq.Z), mean.fBodyGyro.X=mean(fBodyGyro.mean.X), mean.fBodyGyro.Y=mean(fBodyGyro.mean.Y), mean.fBodyGyro.Z=mean(fBodyGyro.mean.Z), mean.fBodyGyro.Freq.X=mean(fBodyGyro.meanFreq.X), mean.fBodyGyro.Freq.Y=mean(fBodyGyro.meanFreq.Y), mean.fBodyGyro.Freq.Z=mean(fBodyGyro.meanFreq.Z), mean.fBodyAccMag=mean(fBodyAccMag.mean), mean.fBodyAccMag.Freq=mean(fBodyAccMag.meanFreq), mean.fBodyBodyAccJerkMag=mean(fBodyBodyAccJerkMag.mean), mean.fBodyBodyAccJerkMag.Freq=mean(fBodyBodyAccJerkMag.meanFreq), mean.fBodyBodyGyroMag=mean(fBodyBodyGyroMag.mean), mean.fBodyBodyGyroMag.Freq=mean(fBodyBodyGyroMag.meanFreq), mean.fBodyBodyGyroJerkMag=mean(fBodyBodyGyroJerkMag.mean), mean.fBodyBodyGyroJerkMag.Freq=mean(fBodyBodyGyroJerkMag.meanFreq), mean.angletBodyAcc.gravity=mean(angletBodyAccMean.gravity), mean.angletBodyAccJerk.gravity=mean(angletBodyAccJerkMean.gravityMean), mean.angletBodyGyro.gravity=mean(angletBodyGyroMean.gravityMean), mean.angletBodyGyroJerk.gravity=mean(angletBodyGyroJerkMean.gravityMean), mean.angleX.gravity=mean(angleX.gravityMean), mean.angleY.gravity=mean(angleY.gravityMean), mean.angleZ.gravity=mean(angleZ.gravityMean))

    Finally we write out the new table to a CSV file, demonstrating its independence. To satisfy the instructions, we also write out a flat file using write.table, called "averages.txt"

    C. write.csv(averages, file="averages.csv", row.names=FALSE) write.table(averages, file="averages.txt", row.names=FALSE)

    This completes the assignment.