see also
- 14 Supervised Learning
- 14b Time Series Analysis
- 15 Unsupervised Learning
- 16 Feature Engineering
- 16b Feature Selection
- Evaluation metrics
- Regression
- R2 vs adj R2 vs MAE vs RMSE
- Classification
- accuracy, class imbalance
- confusion matrix
- precision, recall, f1 score
- ROC-AUC, true positive/negative rates
- log loss
- Regression
- Model optimization, Hyperparameter Tuning
- Grid Search, Random Search
- how can Random Search outperform Grid Search? why is grid search usually not best choice?
- Bayesian Optimization
- Grid Search, Random Search
- Hyperparameter Tuning article
- Techniques
- classical: Grid Search, Random Search
- Bergstra and Bengio paper showed Random usually did as well as Grid with at least 60 points
- "smarter methods"
- tend to be iterative and less parallelizable
- good performance often requires an outter tuning step of the optimization method
- Snoek, Larochelle, and Adams paper
- Gaussian process to model response function and Expected Improvement, then used to determine next set of hyperparameters
- Hutter, Hoos, and Leyton-Brown paper
- train random forest to approximate response surface, which is then used to sample optimal regions
- dubbed SMAC, linked below. article author suggests it works well for categorical hyperparameters
- derivative-free optimzation
- employ heuristics to determine where to sample next. see Nelder-Mead method
- scipy provides Nelder-Mead in its
optimize.minimizefunction, demonstration article
- scipy provides Nelder-Mead in its
- employ heuristics to determine where to sample next. see Nelder-Mead method
- Bayes-optimzation
- see stuff below and linked packages
- random forest smart tuning
- similar to Bayes in that the response surface is modeled with another function, and then that is used to sample more points
- classical: Grid Search, Random Search
- see paper notes for "nested" cross-validation/hyperparameter tuning
- Packages
- Hyperopt Bayes Optimization using tree-based parzen estimators
- SMAC aka SMAC3 Bayes optimization, Random forest
- not maintained?: Spearmint, hypergrad
- Techniques
read/watch more. why are black box approaches good for hyperparameter tuning?
- Bayes Optimization
- Background?