Hello!
I've been using VEBA MicroEuk100 for some analyses latelty, and while it works great. I found some hits to Bacteria coming from the DB. When I check the source_taxonomy.tsv file, I can see there are some bacteria listed there:
EP00049 EukProt 50559 d__Bacteria;p__Bacillota;c__Clostridia;o__Eubacteriales;f__;g__;s__ Clostridia Eubacteriales False
EP00693 EukProt 50422 d__Bacteria;p__Pseudomonadota;c__Gammaproteobacteria;o__Alteromonadales;f__Shewanellaceae;g__Shewanella;s__Shewanella sp. Gammaproteobacteria Alteromonadales Shewanellaceae Shewanella Shewanella sp. True
EP01061 EukProt 50226 d__Bacteria;p__Actinomycetota;c__Actinomycetes;o__Mycobacteriales;f__Mycobacteriaceae;g__Mycobacterium;s__Mycobacterium sp. T103 Actinomycetes Mycobacteriales Mycobacteriaceae Mycobacterium Mycobacterium sp. T103 True
MMETSP1091 MMETSP 1073 d__Bacteria;p__Pseudomonadota;c__Alphaproteobacteria;o__Hyphomicrobiales;f__Nitrobacteraceae;g__Rhodopseudomonas;s__ Alphaproteobacteria Hyphomicrobiales Nitrobacteraceae Rhodopseudomonas False
MMETSP1389 MMETSP 1073 d__Bacteria;p__Pseudomonadota;c__Alphaproteobacteria;o__Hyphomicrobiales;f__Nitrobacteraceae;g__Rhodopseudomonas;s__ Alphaproteobacteria Hyphomicrobiales Nitrobacteraceae Rhodopseudomonas False
I also work with EukProt v3, but when I go to their list of genomes, EP00049, EP00693 and EP01061 do not match on the taxonomy. These IDs match to:
EP00049_Salpingoeca_infusionum, EP00693_Jakoba_libera, and EP01061_Rhynchopus_euleeides.
Is this somehow expected?
I built an MMSeqs2 DB with VEBA MicroEuk and used the taxids in source_taxonomy.tsv to give each protein a taxonomy. However, the LCA algorithm has some trouble when it involves genes from these datasets.
Hello!
I've been using VEBA MicroEuk100 for some analyses latelty, and while it works great. I found some hits to Bacteria coming from the DB. When I check the source_taxonomy.tsv file, I can see there are some bacteria listed there:
I also work with EukProt v3, but when I go to their list of genomes, EP00049, EP00693 and EP01061 do not match on the taxonomy. These IDs match to:
EP00049_Salpingoeca_infusionum, EP00693_Jakoba_libera, and EP01061_Rhynchopus_euleeides.
Is this somehow expected?
I built an MMSeqs2 DB with VEBA MicroEuk and used the taxids in
source_taxonomy.tsvto give each protein a taxonomy. However, the LCA algorithm has some trouble when it involves genes from these datasets.