Computational biology and chemistry

A chemical analysis benchmark shows challenges in applying results across different lipid nanoparticle data sources

Updated

Abstract

LightGBM's source-macro ROC AUC decreased from 0.764 to 0.585 when tested with source holdout.

  • Random-row validation may overestimate model performance for unseen data sources.
  • A paired mean loss of 0.179 (P = 2.5 × 10⁻¹⁹) indicates significant performance decline under source holdout.
  • Continuous-rank correlation dropped from 0.463 to 0.160, suggesting reduced predictive accuracy.
  • Top-10% candidate enrichment decreased from 2.61-fold to 1.39-fold, highlighting challenges in identifying high-performing candidates.
  • Intermediate results from exact-lipid and descriptor-cluster holdouts suggest that chemical novelty does not equate to complete source novelty.

Simplified

Full Text

Full text is available at the source.

What Lands in Your Inbox Each Week:

  • 📚7 fresh studies
  • 📝plain-language summaries
  • direct links to original studies
  • 🏅top journal indicators
  • 📅weekly delivery
  • 🧘‍♂️always free