When using the default backend, the inference speed is very slow, up to a few hours.
However, as I use the vllm to deploy the model, I can't reproduce the results with your images provided in huggingface. What's more, when I use the max length of 4096, some of the results will exceed the limit, resulting the benchmark crashed.
When using the default backend, the inference speed is very slow, up to a few hours.
However, as I use the vllm to deploy the model, I can't reproduce the results with your images provided in huggingface. What's more, when I use the max length of 4096, some of the results will exceed the limit, resulting the benchmark crashed.