Thank you for your work and perfect theoretical derivation!And I have some questions.
Have you compared the training time with other models, such as vim, and what is the main reason for the longer time? And how about the ablation experiments on the number of the nodes?
Thanks again.
Thank you for your work and perfect theoretical derivation!And I have some questions.
Have you compared the training time with other models, such as vim, and what is the main reason for the longer time? And how about the ablation experiments on the number of the nodes?
Thanks again.