You're reading from Practical Automated Machine Learning Using H2O.ai Discover the power of automated machine learning, from experimentation through to deployment to production

Product type Paperback

Published in Sep 2022

Publisher Packt

ISBN-13 9781801074520

Length 396 pages

Edition 1st Edition

Tools

H2O

Concepts

Machine Learning

Author (1):

Salil Ajgaonkar

View More author details

Table of Contents (19) Chapters

Preface

1. Part 1 H2O AutoML Basics

2. Chapter 1: Understanding H2O AutoML Basics FREE CHAPTER

3. Chapter 2: Working with H2O Flow (H2O’s Web UI)

4. Part 2 H2O AutoML Deep Dive

5. Chapter 3: Understanding Data Processing

6. Chapter 4: Understanding H2O AutoML Architecture and Training

7. Chapter 5: Understanding AutoML Algorithms

8. Chapter 6: Understanding H2O AutoML Leaderboard and Other Performance Metrics

9. Chapter 7: Working with Model Explainability

10. Part 3 H2O AutoML Advanced Implementation and Productization

11. Chapter 8: Exploring Optional Parameters for H2O AutoML

12. Chapter 9: Exploring Miscellaneous Features in H2O AutoML

13. Chapter 10: Working with Plain Old Java Objects (POJOs)

14. Chapter 11: Working with Model Object, Optimized (MOJO)

15. Chapter 12: Working with H2O AutoML and Apache Spark

16. Chapter 13: Using H2O AutoML with Other Technologies

17. Index

Why subscribe?

18. Other Books You May Enjoy

Handling missing values in the dataframe

Missing values in datasets are the most common issue in the real world. It is often expected to have at least a few instances of missing data in huge chunks of datasets collected from various sources. Data can be missing for several reasons, which can range from anything from data not being generated at the source all the way to downtimes in data collectors. Handling missing data is very important for model training, as many ML algorithms don’t support missing data. Those that do may end up giving more importance to looking for patterns in the missing data, rather than the actual data that is present, which distracts the machine from learning.

Missing data is often referred to as Not Available (NA) or nan. Before we can send a dataframe for model training, we need to handle these types of values first. You can either drop the entire row that contains any missing values or you can fill them with any default value either default or common...

The rest of the chapter is locked

You're reading from Practical Automated Machine Learning Using H2O.ai Discover the power of automated machine learning, from experimentation through to deployment to production

Table of Contents (19) Chapters

Handling missing values in the dataframe

Authors (1)

Personalised recommendations for you

You're reading from Practical Automated Machine Learning Using H2O.ai Discover the power of automated machine learning, from experimentation through to deployment to production

Table of Contents (19) Chapters

Handling missing values in the dataframe

Unlock this book and the full library FREE for 7 days

Authors (1)

Personalised recommendations for you