Analysis of variance (ANOVA) is used to compare means of two or more samples. While t test can be used to compare means for two samples, it can not be used for more than two. Anova is used in such situation.
ANOVA was invented by Sir Ronald Fisher who applied this technique first to agriculture and cotton industry
It is now a popular technique used in many areas, most notably in design of experiments.
In this video we will discuss few of the common mistakes often made while performing cross validation of Machine Learning Models. While root means square error and accuracy rate are the two most popular metrics used in evaluating model performance in cross validation, there are limitations of using these when the performance is more important for the researcher in one section of the data than the other
For example we could be interested in better performance in predicting house price of a segment of the sample (say the middle priced houses) than the other segments. Similarly, we could be interested in predicting more default customers than the non default customers in a classification set up.
Occam's Razor (Parsimony ) is one hypothesis that states that out of all possible models that provides similar results (or performance), the one that is most simple should be selected as the final model.
It dates back to many centuries ago when it was studied,not in relation to ML though but was studied in general. This is now widely accepted means of selecting the best model out of many models
Everyone, irrespective career choices should learn some data science. Data Science skills are very useful every where. As part of data science you learn statistical analysis, forecasting, data visualization & mathematical programming that are very useful no matter which career you are interested in . Computational skills are going to be very important in future in any jobs
The No Free Lunch theorem in Machine Learning says that no single machine learning algorithm is universally the best algorithm. In fact, the goal of machine learning models is not find an algorithm that will be the best.
If one algorithm works good for a given problem it may not work well for some other problem. So there is no universally best algorithm that works very well in all cases
IID stands for independent and identical distribution in which it is assumed that data points are independent with each other and the are having similar distributed . Because of fulfillment of IID assumptions we are able to use cross validation to evaluate models.
Since data points are assumed to be IID, we are able to split the data in to training & test type . Thats because we assume both test & training data create from same data generating process
In this video I have discussed about the 9 types of Machine Learning problems and tasks. We have discussed about these broad categories in detail in this video