首页 | 本学科首页   官方微博 | 高级检索  
     检索      


The importance of outlier detection and training set selection for reliable environmental QSAR predictions
Authors:Furusjö Erik  Svenson Anders  Rahmberg Magnus  Andersson Magnus
Institution:IVL Swedish Environmental Research Institute Ltd., Box 21060, 100 31 Stockholm, Sweden. erik.furusjo@ivl.se
Abstract:Empirical QSAR models are only valid in the domain they were trained and validated. Application of the model to substances outside the domain of the model can lead to grossly erroneous predictions. Partial least squares (PLS) regression provides tools for prediction diagnostics that can be used to decide whether or not a substance is within the model domain, i.e. if the model prediction can be trusted. QSAR models for four different environmental end-points are used to demonstrate the importance of appropriate training set selection and how the reliability of QSAR predictions can be increased by outlier diagnostics. All models showed consistent results; test set prediction errors were very similar in magnitude to training set estimation errors when prediction outlier diagnostics were used to detect and remove outliers in the prediction data. Test set prediction errors for substances classified as outliers were much larger. The difference in the number of outliers between models with a randomly and systematically selected training illustrates well the need of representative training data.
Keywords:
本文献已被 PubMed 等数据库收录!
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号