I have noticed that the training speed is dramatically decreased after the first simplification. Depth initialisation is included in that iteration, however it happens in each 5000 step as well. I am curious why the training speed decreases even if the gaussian number is fewer?
I have noticed that the training speed is dramatically decreased after the first simplification. Depth initialisation is included in that iteration, however it happens in each 5000 step as well. I am curious why the training speed decreases even if the gaussian number is fewer?