копировать удалить добавить публикацию в буфер
Запись сообщества
посмотреть историю данной записи
URL
DOI
BibTeX
EndNote
APA
Chicago
DIN 1505
Harvard
MSOffice XML

GENEMASK: Fast Pretraining of Gene Sequences to Enable Few-Shot Learning

S. Roy, J. Wallat, S. Sundaram, W. Nejdl, и N. Ganguly. том 372 из Frontiers in Artificial Intelligence and Applications, стр. 2002-2009. (2023)
DOI: 10.3233/FAIA230492

Аннотация

Large-scale language models such as DNABert and LOGO aim to learn optimal gene representations and are trained on the entire Human Reference Genome. However, standard tokenization schemes involve a simple sliding window of tokens like k-mers that do not leverage any gene-based semantics and thus may lead to (trivial) masking of easily predictable sequences, and subsequently inefficient Masked Language Modeling (MLM) training. Therefore, we propose a novel masking algorithm, GENEMASK, for MLM training of gene sequences, where we randomly identify positions in a gene sequence as mask centers and locally select the span around the mask center with the highest Normalized Pointwise Mutual Information (NPMI) to mask. We observe that in the absence of human-understandable semantics in the genomics domain (in contrast, semantic units like words and phrases are inherently available in NLP), GENEMASK-based models substantially outperform the SOTA models (DNABert and LOGO) over four benchmark gene sequence classification datasets in five few-shot settings (10 to 1000-shot). More significantly, the GENEMASK-based DNABert model is trained for less than one-tenth of the number of epochs of the original SOTA model. We also observe a strong correlation between top-ranked PMI tokens and conserved DNA sequence motifs, which may indicate the incorporation of latent genomic information. The codes (including trained models) and datasets are made publicly available at unmapped: uri https://github.com/roysoumya/GeneMask.

Линки и ресурсы

ключ BibTeX

noauthororeditor

тип записи

inproceedings

год

2023

страницы

2002-2009

серии

Frontiers in Artificial Intelligence and Applications

том

372

isbn

978-1-64368-437-6

DOI

10.3233/FAIA230492

дополнительные URL-адреса

Github Codebase

тэги

@roysoumya- тэги данного пользователя выделены

Цитировать эту публикацию

@inproceedings{noauthororeditor, abstract = {Large-scale language models such as DNABert and LOGO aim to learn optimal gene representations and are trained on the entire Human Reference Genome. However, standard tokenization schemes involve a simple sliding window of tokens like k-mers that do not leverage any gene-based semantics and thus may lead to (trivial) masking of easily predictable sequences, and subsequently inefficient Masked Language Modeling (MLM) training. Therefore, we propose a novel masking algorithm, GENEMASK, for MLM training of gene sequences, where we randomly identify positions in a gene sequence as mask centers and locally select the span around the mask center with the highest Normalized Pointwise Mutual Information (NPMI) to mask. We observe that in the absence of human-understandable semantics in the genomics domain (in contrast, semantic units like words and phrases are inherently available in NLP), GENEMASK-based models substantially outperform the SOTA models (DNABert and LOGO) over four benchmark gene sequence classification datasets in five few-shot settings (10 to 1000-shot). More significantly, the GENEMASK-based DNABert model is trained for less than one-tenth of the number of epochs of the original SOTA model. We also observe a strong correlation between top-ranked PMI tokens and conserved DNA sequence motifs, which may indicate the incorporation of latent genomic information. The codes (including trained models) and datasets are made publicly available at unmapped: uri https://github.com/roysoumya/GeneMask.}, added-at = {2023-10-16T10:17:16.000+0200}, author = {Roy, Soumyadeep and Wallat, Jonas and Sundaram, Sowmya S and Nejdl, Wolfgang and Ganguly, Niloy}, biburl = {https://www.bibsonomy.org/bibtex/282b607481c25068beec13bdedfc2e063/roysoumya}, doi = {10.3233/FAIA230492}, interhash = {c59b5147ac39bbc173fcf3c00e021ca2}, intrahash = {82b607481c25068beec13bdedfc2e063}, isbn = {978-1-64368-437-6}, keywords = {l3s leibnizailab myown}, pages = {2002-2009}, series = {Frontiers in Artificial Intelligence and Applications}, timestamp = {2023-10-16T10:17:16.000+0200}, title = {GENEMASK: Fast Pretraining of Gene Sequences to Enable Few-Shot Learning}, volume = 372, year = 2023 }

искать в

Метаданные

Последнее изменение год назад
Создан год назад

Комментарии и рецензии
(0)

Комментарии, или рецензии отсутствуют. Вы можете их написать!