-
- Added CPU support for vLLM inference
(
52920cc)
- Added CPU support for vLLM inference
(
- Added support for quantized models | Improved the estimations, specifically for the smaller GPUs
(VRAM <= 4GB)
(
ba30698)
- Updated pyproject.toml
(
6c3db3c)
-
Initial release (
ad2100d) -
Initial release (
3efc729) -
Updated release.yml (
89a8e8e) -
Updated release.yml to see the logs (
860a975)
-
Initial release (
3efc729) -
Updated release.yml (
89a8e8e) -
Updated release.yml to see the logs (
860a975)
- Initial Release