Therapeutic peptides are increasingly explored in drug development due to their high specificity and efficacy, however their poor bioavailability remains a major challenge. One strategy to address this limitation is the incorporation of permeation enhancers (PEs) into oral delivery systems to increase intestinal absorption. A previous study applied a random forest (RF) model to predict PE potency from TEER data, achieving moderate performance. Since then, advances in ML methods, molecular representations and data-splitting strategies have created new opportunities to improve such predictive models.In the present study, several modern ML approaches were evaluated for predicting PE potency using TEER measurements. When using the same dataset as Welling et al., GB algorithms generally outperformed the RF approach, with LightGBM showing the strongest performance. By expanding the available training data with additional literature sources, multiple models were developed that exceeded the performance reported by Welling et al, although the best-performing model in this case was also a RF-based (RMSEOLS = 0.14 and rp2 = 0.73). These findings highlight the importance of dataset size for predictive performance, while also demonstrating the potential of GB methods for this task.
Analysis of feature importance further indicated that molecular properties related to polarity, electrostatics and the distribution of hydrophobic and polar surface areas strongly influence model predictions. Together, these results suggest that amphiphilic molecular characteristics may be important for PE potency.