Integrative systems microbiology increasingly relies on algorithmic approaches capable of extracting biologically meaningful patterns from heterogeneous and often high dimensional, low-sample-size (HDLSS) biological datasets. A major obstacle in this setting is the instability of inferred molecular signatures across cohorts, tissues, and measurement platforms. Here, we address this problem by formulating molecular system inference as a multi-dataset integration task and by applying the Matthews Correlation Coefficient-Recursive Ensemble Feature Selection (MCC-REFS) algorithm to jointly analyze five independent transcriptomic datasets spanning peripheral blood mononuclear cells, whole blood, plasma, and post-mortem tissues. We compared MCC-REFS with three commonly used feature-selection strategies, GRACES, SelectKBest, and Deep Neural Pursuit (DNP), in order to evaluate robustness, convergence, and cross-context reproducibility. MCC-REFS consistently converged on a compact seven-gene system (PPP2CB, SOCS3, ARG1, IL6R, ECHS1, FZD2, TRGV3/5) exhibiting higher stability indices and stronger classification performance than alternative methods. Generalization was assessed using an independent multi-layer perceptron classifier across validation cohorts with differing tissue origin and sequencing technologies, demonstrating preservation of discriminative structure. To support interpretation, we integrated functional, pharmacological, and interventional knowledge from DrugBank, DGIdb, and Open Targets, enabling the mapping of inferred gene systems onto pathways, known drug targets, and ongoing clinical investigations. Taken together, this work presents an algorithmic framework for multi-dataset and multi-omics integration in systems microbiology, illustrating how stable and interpretable molecular patterns can be identified from heterogeneous data, with Long COVID serving as a representative case study.