Transactions on Machine Intelligence

Transactions on Machine Intelligence

MultiCNKG: Integrating Cognitive Neuroscience, Gene, and Disease Knowledge Graphs Using Large Language Models

Document Type : Original Article

Authors
1 Ph.D. in IT Engineering , Department of Computer Engineering and Information Technology, University of Qom, Qom, Iran.
2 Kheirolah Rahsepar Fard Department of Computer Engineering and Information Technology, University of Qom, Qom, Iran.
3 Department of Computer Engineering, Malek Ashtar University of Technology, Tehran
Abstract
The advent of large language models (LLMs) has revolutionized the integration of knowledge graphs (KGs) in the biomedical and cognitive sciences, effectively overcoming the limitations of traditional machine learning methods in capturing intricate semantic links among genes, diseases, and cognitive processes. This paper introduces MultiCNKG, an innovative framework that merges three distinct knowledge sources: the Cognitive Neuroscience Knowledge Graph (CNKG), containing 2.9K nodes and 4.3K edges across 9 node types and 20 edge types; the Gene Ontology (GO), featuring 43K nodes and 75K edges across 3 node types and 4 edge types; and the Disease Ontology (DO), comprising 11.2K nodes and 8.8K edges with 1 node type and 2 edge types. Utilizing advanced LLMs such as GPT-4, we perform automated entity alignment, semantic similarity computation, and graph augmentation to construct a unified, cohesive KG that interconnects genetic mechanisms, neurological disorders, and cognitive functions. The resulting MultiCNKG unified graph encompasses 6.9K nodes across 5 distinct types and 11.3K edges spanning 7 relational types, establishing a multi-layered analytical pipeline from molecular to behavioral domains. Empirical evaluations demonstrate robust framework performance, achieving an 85.20% precision rate, 87.30% recall, 92.18% coverage, 82.50% graph consistency, a 40.28% novelty detection rate, and an 89.50% expert validation score. Furthermore, link prediction benchmarks utilizing TransE (MR: 391, MRR: 0.411) and RotatE (MR: 263, MRR: 0.395) yield highly competitive performance against standard benchmarks like FB15k-237 and WN18RR. Ultimately, this integrated KG advances clinical and research applications in personalized medicine, cognitive disorder diagnostics, and data-driven hypothesis formulation within cognitive neuroscience.
Keywords

[1]      Xia, F., Sun, K., Yu, S., Aziz, A., Wan, L., Pan, S., & Liu, H. (2021). Graph learning: A survey. IEEE Transactions on Artificial Intelligence, 2(2), 109-127. https://doi.org/10.1109/TAI.2021.3076021
[2]      Hu, W., Fey, M., Zitnik, M., Dong, Y., Ren, H., Liu, B., Catasta, M., & Leskovec, J. (2020). Open graph benchmark: Datasets for machine learning on graphs. Advances in Neural Information Processing Systems, 33, 22118-22133.
[3]      Chiang, W. L., Liu, X., Si, S., Li, Y., Bengio, S., & Hsieh, C. J. (2019). Cluster-GCN: An efficient algorithm for training deep and large graph convolutional networks. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 257-266. https://doi.org/10.1145/3292500.3330925
[4]      Li, Q., Li, X., Chen, L., & Wu, D. (2022). Distilling knowledge on text graph for social media attribute inference. Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2024-2028. https://doi.org/10.1145/3477495.3531968
[5]      Zhu, J., Cui, Y., Liu, Y., Sun, H., Li, X., Pelger, M., Yang, T., Zhang, L., Zhang, R., & Zhao, H. (2021). TextGNN: Improving text encoder via graph neural network in sponsored search. Proceedings of the Web Conference 2021, 2848-2857. https://doi.org/10.1145/3442381.3449842
[6]      Ma, Y., & Tang, J. (2021). Deep learning on graphs. Cambridge University Press. https://doi.org/10.1017/9781108924184
[7]      Shao, Y., Taylor, S., Marshall, N., Morioka, C., & Zeng-Treitler, Q. (2018). Clinical text classification with word embedding features vs. bag-of-words features. 2018 IEEE International Conference on Big Data (Big Data), 2874-2878. IEEE. https://doi.org/10.1109/BigData.2018.8622345
[8]      Mikolov, T. (2013). Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.
[9]      Qiu, X., Sun, T., Xu, Y., Shao, Y., Dai, N., & Huang, X. (2020). Pre-trained models for natural language processing: A survey. Science China Technological Sciences, 63(10), 1872-1897. https://doi.org/10.1007/s11431-020-1647-3
[10]   Miaschi, A., & Dell'Orletta, F. (2020). Contextual and non-contextual word embeddings: An in-depth linguistic investigation. Proceedings of the 5th Workshop on Representation Learning for NLP, 110-119. https://doi.org/10.18653/v1/2020.repl4nlp-1.15
[11]   Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al. (2023). A survey of large language models. arXiv preprint arXiv:2303.18223.
[12]   Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., et al. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774.
[13]   Kenton, J. D. M. W. C., & Toutanova, L. K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT, 1, 2. Minneapolis, Minnesota.
[14]    Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I., et al. (2019). Language models are unsupervised multitask learners. OpenAI Blog, 1(8), 9.
[15]   Raffel, C., et al. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21, 1-67.
[16]   Team G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J. B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al. (2023). Gemini: A family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
[17]   Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. (2023). PaLM: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240), 1-113.
[18]   Abdin, M., Jacobs, S. A., Awan, A. A., Aneja, J., Awadallah, A., Awadalla, H., et al. (2024). Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219.
[19]   Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M. A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. (2023). LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
[20]   Webb, T., Holyoak, K. J., & Lu, H. (2023). Emergent analogical reasoning in large language models. Nature Human Behaviour, 7(9), 1526-1541. https://doi.org/10.1038/s41562-023-01659-w
[21]   Belleau, F., Nolin, M. A., Tourigny, N., Rigault, P., & Morissette, J. (2008). Bio2RDF: Towards a mashup to build bioinformatics knowledge systems. Journal of Biomedical Informatics, 41(5), 706-716. https://doi.org/10.1016/j.jbi.2008.03.004
[22]   Himmelstein, D. S., Lizee, A., Hessler, C., Brueggeman, L., Chen, S. L., Hadley, D., Green, A., Khankhanian, P., & Baranzini, S. E. (2017). Systematic integration of biomedical knowledge prioritizes drugs for repurposing. eLife, 6, e26726. https://doi.org/10.7554/eLife.26726.017
[23]   Luo, R., Sun, L., Xia, Y., Qin, T., Zhang, S., Poon, H., & Liu, T. Y. (2022). BioGPT: Generative pre-trained transformer for biomedical text generation and mining. Briefings in Bioinformatics, 23(6), bbac409. https://doi.org/10.1093/bib/bbac409
[24]   Jin, Q., Wang, Z., Floudas, C. S., Chen, F., Gong, C., Bracken-Clarke, D., Xue, E., Yang, Y., Sun, J., & Lu, Z. (2023). Matching patients to clinical trials with large language models. arXiv. https://doi.org/10.1038/s41467-024-53081-z
[25]   Yuan, H., Yuan, Z., Gan, R., Zhang, J., Xie, Y., & Yu, S. (2022). BioBART: Pretraining and evaluation of a biomedical generative language model. arXiv preprint arXiv:2204.03905. https://doi.org/10.18653/v1/2022.bionlp-1.9
[26]    Labrak, Y., Bazoge, A., Morin, E., Gourraud, P. A., Rouvier, M., & Dufour, R. (2024). BioMistral: A collection of open-source pretrained large language models for medical domains. arXiv preprint arXiv:2402.10373. https://doi.org/10.18653/v1/2024.findings-acl.348
[27]    Alharbi, R., Ahmed, U., Dobriy, D., Lajewska, W., Menotti, L., Saeedizade, M. J., & Dumontier, M. (2023). Exploring the role of generative AI in constructing knowledge graphs for drug indications with medical context. Proceedings http://ceur-ws.org.
[28]    Wawrzik, F., Rafique, K. A., Rahman, F., & Grimm, C. (2023). Ontology learning applications of knowledge base construction for microelectronic systems information. Information, 14(3), 176. https://doi.org/10.3390/info14030176
[29]    Peng, C., Yang, X., Yu, Z., Bian, J., Hogan, W. R., & Wu, Y. (2023). Clinical concept and relation extraction using prompt-based machine reading comprehension. Journal of the American Medical Informatics Association, 30(9), 1486-1493. https://doi.org/10.1093/jamia/ocad107
[30]    Pan, S., Luo, L., Wang, Y., Chen, C., Wang, J., & Wu, X. (2024). Unifying large language models and knowledge graphs: A roadmap. IEEE Transactions on Knowledge and Data Engineering. https://doi.org/10.1109/TKDE.2024.3352100
[31]    Khorashadizadeh, H., Mihindukulasooriya, N., Tiwari, S., Groppe, J., & Groppe, S. (2023). Exploring in-context learning capabilities of foundation models for generating knowledge graphs from text. arXiv preprint arXiv:2305.08804.
[32]    Boylan, J., Mangla, S., Thorn, D., Ghalandari, D. G., Ghaffari, P., & Hokamp, C. (2024). KGValidator: A Framework for Automatic Validation of Knowledge Graph Construction. arXiv preprint arXiv:2404.15923.
[33]    Allen, B. P., & Groth, P. T. (2024). Evaluating Class Membership Relations in Knowledge Graphs using Large Language Models. arXiv preprint arXiv:2404.17000. https://doi.org/10.1007/978-3-031-78952-6_2
[34]    Soman, K., Rose, P. W., Morris, J. H., Akbas, R. E., Smith, B., Peetoom, B., Villouta-Reyes, C., Cerono, G., Shi, Y., Rizk-Jackson, A., et al. (2023). Biomedical knowledge graph-enhanced prompt generation for large language models. arXiv preprint arXiv:2311.17330. https://doi.org/10.1093/bioinformatics/btae560
[35]    Cm, S., Prakash, J., & Singh, P. K. (2023). Question answering over knowledge graphs using BERT based relation mapping. Expert Systems, 40(10), e13456. https://doi.org/10.1111/exsy.13456
[36]    Guo, Q., Cao, S., & Yi, Z. (2022). A medical question answering system using large language models and knowledge graphs. International Journal of Intelligent Systems, 37(11), 8548-8564. https://doi.org/10.1002/int.22955
[37]    Wu, Y., Hu, N., Bi, S., Qi, G., Ren, J., Xie, A., & Song, W. (2023). Retrieve-rewrite-answer: A KG-to-text enhanced LLMs framework for knowledge graph question answering. arXiv preprint arXiv:2309.11206.
[38]    Choudhary, N., & Reddy, C. K. (2023). Complex logical reasoning over knowledge graphs using large language models. arXiv preprint arXiv:2305.01157.
[39]    Varshney, D., Zafar, A., Behera, N. K., & Ekbal, A. (2023). Knowledge grounded medical dialogue generation using augmented graphs. Scientific Reports, 13(1), 3310. https://doi.org/10.1038/s41598-023-29213-8
[40]    Jiang, P., Xiao, C., Cross, A., & Sun, J. (2023). Graphcare: Enhancing healthcare predictions with personalized knowledge graphs. arXiv preprint arXiv:2305.12788.
[41]    Xu, R., Shi, W., Yu, Y., Zhuang, Y., Jin, B., Wang, M. D., Ho, J. C., & Yang, C. (2024). RAM-EHR: Retrieval augmentation meets clinical predictions on electronic health records. arXiv preprint arXiv:2403.00815. https://doi.org/10.18653/v1/2024.acl-short.68
[42]    Gao, Y., Li, R., Croxford, E., Tesch, S., To, D., Caskey, J., Patterson, B. W., Churpek, M. M., Miller, T., Dligach, D., et al. (2023). Large language models and medical knowledge grounding for diagnosis prediction. medRxiv, 2023-11. https://doi.org/10.1101/2023.11.24.23298641
[43]    V. N. Ioannidis et al., "Few-shot Link Prediction via Graph Neural Networks for COVID-19 Drug-repurposing," 2020. DOI:10.48550/arXiv.2007.10261
[44]    P. Chandak, K. Huang, and M. Zitnik, "Building a Knowledge Graph to Enable Precision Medicine," Scientific Data, 2023. https://doi.org/10.1038/s41597-023-01960-3
[45]    M. Ashburner et al., "Gene Ontology: Tool for the Unification of Biology," Nature Genetics, vol. 25, no. 1, 2000, pp. 25-29. https://doi.org/10.1038/75556
[46]    Z. Gao, P. Ding, and R. Xu, "KG-Predict: A Knowledge Graph Computational Framework for Drug Repurposing," Journal of Biomedical Informatics, vol. 132, 2022, 104133. https://doi.org/10.1016/j.jbi.2022.104133
[47]    Z. Ghorbanali et al., "DrugRep-KG: Toward Learning a Unified Latent Space for Drug Repurposing Using Knowledge Graphs," Journal of Chemical Information and Modeling, vol. 63, no. 8, 2023, pp. 2532-2545. https://doi.org/10.1021/acs.jcim.2c01291
[48]    L. Schriml et al., "Disease Ontology: A Backbone for Disease Semantic Integration," Nucleic Acids Research, vol. 40, no. D1, 2011, pp. D940-D946. https://doi.org/10.1093/nar/gkr972
[49]    D. S. Wishart et al., "DrugBank 5.0: A Major Update to the DrugBank Database for 2018," Nucleic Acids Research, vol. 46, no. D1, 2018, pp. D1074-D1082. https://doi.org/10.1093/nar/gkx1037
[50]   P. Chandak, K. Huang, and M. Zitnik, "Building a Knowledge Graph to Enable Precision Medicine," Scientific Data, vol. 10, no. 1, 2023, 67. https://doi.org/10.1038/s41597-023-01960-3
[51]    Z. Ghorbanali et al., "DrugRep-KG: Toward Learning a Unified Latent Space for Drug Repurposing Using Knowledge Graphs," Journal of Chemical Information and Modeling, vol. 63, no. 8, 2023, pp. 2532-2545. https://doi.org/10.1021/acs.jcim.2c01291
[52]    K. Bollacker, C. Evans, P. Paritosh, T. Sturge, and J. Taylor, "Freebase: A Collaboratively Created Graph Database for Structuring Human Knowledge," in Proceedings of the 2008 ACM SIGMOD International Conference on Management of Data (SIGMOD '08), New York, NY, USA, 2008, pp. 1247-1250. https://doi.org/10.1145/1376616.1376746
[53]    G. A. Miller, "WordNet: A Lexical Database for English," Communications of the ACM, vol. 38, no. 11, 1995, pp. 39-41. https://doi.org/10.1145/219717.219748
[54]   F. M. Suchanek, G. Kasneci, and G. Weikum, "Yago: A Core of Semantic Knowledge," in Proceedings of the 16th International Conference on World Wide Web (WWW '07), New York, NY, USA, 2007, pp. 697-706. https://doi.org/10.1145/1242572.1242667
[55]   Bordes, N. Usunier, A. Garcia-Durán, J. Weston, and O. Yakhnenko, "Translating Embeddings for Modeling Multi-relational Data," in Proceedings of the 26th International Conference on Neural Information Processing Systems (NIPS '13), Red Hook, NY, USA, 2013, pp. 2787-2795. https://hal.science/hal-00920777
[56]   Z. Sun et al., "Rotate: Knowledge Graph Embedding by Relational Rotation in Complex Space," arXiv preprint, 2019. https://doi.org/10.48550/arXiv.1902.10197
[57]   B. Yang et al., "Embedding Entities and Relations for Learning and Inference in Knowledge Bases," in International Conference on Learning Representations, 2014. https://doi.org/10.48550/arXiv.1412.6575
[58]   T. Trouillon et al., "Complex Embeddings for Simple Link Prediction," in Proceedings of the International Conference on Machine Learning, New York, USA, 2016, pp. 2071-2080.
[59]   T. Dettmers et al., "Convolutional 2D Knowledge Graph Embeddings," in Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, USA, 2018. https://doi.org/10.1609/aaai.v32i1.11573. https://doi.org/10.1609/aaai.v32i1.11573
[60]   Z. Zheng et al., "HolmE: Low-dimensional Hyperbolic KG Embedding for Better Extrapolation," in Proceedings of the ESWC Conference, 2024. https://doi.org/10.1007/978-3-031-60626-7_6
[61]    A. Sarabadani et al., "DKG-LLM: A Framework for Medical Diagnosis and Personalized Treatment Recommendations via Dynamic Knowledge Graph and Large Language Model Integration," arXiv preprint, 2025. https://doi.org/10.48550/arXiv.2508.06186
[62]   A. Sarabadani et al., "ExKG-LLM: Leveraging Large Language Models for Automated Expansion of Cognitive Neuroscience Knowledge Graphs," arXiv preprint, 2025. https://doi.org/10.48550/arXiv.2503.06479
[63]   K. Rahsepar Fard, A. Sarabadani, and H. Dalvand, "GDPKG-LLM: Integrating Gene, Disease, and Pharmacogenomics Knowledge Graphs for Cognitive Neuroscience Using Large Language Models," csci, vol. 26, no. 3, Oct. 2025. https://doi.org/10.7494/csci.2025.26.3.6673
Volume 9, Issue 1
Winter 2026
Pages 1-15

  • Receive Date 26 November 2025
  • Revise Date 04 January 2026
  • Accept Date 22 February 2026