A4.2 Data preprocessing (HL Only) eBook
This subtopic covers the significance of data cleaning, emphasising its impact on model performance through techniques like handling outliers, duplicates, and missing data, plus normalisation and standardisation. It also describes feature selection for identifying informative data attributes, using filter, wrapper, or embedded methods. Lastly, it addresses dimensionality reduction to mitigate issues like overfitting...