A Comparative Accuracy Analysis of Node2Vec and FastRP
DOI:
https://doi.org/10.33795/jip.v12i4.10863Keywords:
Graph Embedding, Text Embedding, Node2Vec, FastRP, S-BERTAbstract
The rapid increase in scientific article publications has created a need for automated methods to support various applications, such as recommendation systems that help researchers discover relevant literature by measuring similarity between scientific publications using graph and text embeddings. This study mainly compares the accuracy of two graph embedding methods, Node2Vec and FastRP, where similarity is measured based on title, abstract, and citation attributes using a hybrid approach. In this approach, relational attributes (citations) are processed with graph embeddings, while textual attributes (titles and abstracts) are processed with text embeddings (S-BERT). The resulting vectors are then combined using a weighting scheme across attributes of title, abstract, and citation, with similarity values calculated using Cosine Similarity. The ordinal ranking results of these similarity scores are evaluated using Kendall Tau and Normalized Discounted Cumulative Gain (NDCG). The evaluation results indicate that Node2Vec achieves higher agreement with the similarity rankings generated by SPECTER than FastRP, as measured using NDCG and Kendall Tau. This study contributes to understanding the impact of graph structures in Node2Vec and FastRP on similarity scores derived from embeddings of title, abstract, and citation attributes. However, the current results are limited to a dataset consisting of only three attributes (title, abstract, and citation) from scientific publications, suggesting that further research must incorporate additional attributes (introduction, methods, results & discussion, and conclusion). The findings of this study are expected to provide practical guidance for researchers and developers of recommendation systems in selecting the most appropriate embedding method for their needs.
Downloads
References
Cabrera-Diego, L. A., El-Bèze, M., Torres-Moreno, J.-M., & Durette, B. (2019). Ranking résumés automatically using only résumés: A method free of job offers. Expert Systems with Applications, 123, 91–107. https://doi.org/10.1016/j.eswa.2018.12.054
Chen, H., Sultan, S. F., Tian, Y., Chen, M., & Skiena, S. (2019). Fast and accurate network embeddings via very sparse random projection. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management (pp. 399–408). Association for Computing Machinery. https://doi.org/10.1145/3357384.3357879
Cohan, A., Feldman, S., Beltagy, I., Downey, D., & Weld, D. S. (2020). SPECTER: Document-level representation learning using citation-informed transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 2270–2282). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.207
Colangelo, M. T., Meleti, M., Guizzardi, S., Calciolari, E., & Galli, C. (2025). A comparative analysis of sentence transformer models for automated journal recommendation using PubMed metadata. Biology and Life Sciences. https://doi.org/10.20944/preprints202501.1334.v1
Grover, A., & Leskovec, J. (2016). node2vec: Scalable feature learning for networks. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 855–864). Association for Computing Machinery. https://doi.org/10.1145/2939672.2939754
Gupta, R., Krishnamoorthy, M., & Vangala, V. (2024). Better graph embeddings for enterprise graphs. In Proceedings of the 7th Joint International Conference on Data Science & Management of Data (11th ACM IKDD CODS and 29th COMAD) (pp. 368–374). Association for Computing Machinery. https://doi.org/10.1145/3632410.3632412
Hiraoka, Y., Imoto, Y., Lacombe, T., Meehan, K., & Yachimura, T. (2024). Topological Node2vec: Enhanced graph embedding via persistent homology. Journal of Machine Learning Research, 25, 1–26. https://jmlr.org/papers/volume25/23-1185/23-1185.pdf
Järvelin, K., & Kekäläinen, J. (2002). Cumulated gain-based evaluation of IR techniques. ACM Transactions on Information Systems, 20(4), 422–446. https://doi.org/10.1145/582415.582418
Kojaku, S., Livan, G., & Masuda, N. (2021). Detecting anomalous citation groups in journal networks. Scientific Reports, 11(1), 14524. https://doi.org/10.1038/s41598-021-93572-3
Martín-Martín, A., Thelwall, M., Orduna-Malea, E., & Delgado López-Cózar, E. (2021). Google Scholar, Microsoft Academic, Scopus, Dimensions, Web of Science, and OpenCitations’ COCI: A multidisciplinary comparison of coverage via citations. Scientometrics, 126(1), 871–906. https://doi.org/10.1007/s11192-020-03690-4
Sakib, N., Ahmad, R. B., & Haruna, K. (2020). A collaborative approach toward scientific paper recommendation using citation context. IEEE Access, 8, 51246–51255. https://doi.org/10.1109/ACCESS.2020.2980589
Sjögårde, P., & Ahlgren, P. (2024). Seed-based information retrieval in networks of research publications: Evaluation of direct citations, bibliographic coupling, co-citations, and PUBMED-related article score. Journal of the Association for Information Science and Technology, 75(13), 1453–1465. https://doi.org/10.1002/asi.24951
Sterling, J. A., & Montemore, M. M. (2022). Combining citation network information and text similarity for research article recommender systems. IEEE Access, 10, 16–23. https://doi.org/10.1109/ACCESS.2021.3137960
Tran, M.-N., & Kim, Y. (2024). Design of 5G architecture enhancements for supporting serverless computing. IEEE Access, 12, 163382–163395. https://doi.org/10.1109/ACCESS.2024.3490671
Januzaj, Y., & Luma, A. (2022). Cosine similarity – A computing approach to match similarity between higher education programs and job market demands based on maximum number of common words. International Journal of Emerging Technologies in Learning (iJET), 17(12), 258–268. https://doi.org/10.3991/ijet.v17i12.30375
Yu, D., Wang, W., Zhang, S., Zhang, W., & Liu, R. (2017). Hybrid self-optimized clustering model based on citation links and textual features to detect research topics. PLOS ONE, 12(10), e0187164. https://doi.org/10.1371/journal.pone.0187164







