Conversation
Signed-off-by: Jinyan Li <jinyali@amazon.com>
Signed-off-by: Jinyan Li <jinyali@amazon.com>
Signed-off-by: Jinyan Li <jinyali@amazon.com>
Signed-off-by: Jinyan Li <jinyali@amazon.com>
Signed-off-by: Jinyan Li <jinyali@amazon.com>
Signed-off-by: Jinyan Li <jinyali@amazon.com>
junpuf
reviewed
Nov 17, 2025
Contributor
junpuf
left a comment
There was a problem hiding this comment.
Try make a small dummy change to both dockerfiles just to test if the CI workflow behave the same way we thought they would be. We can revert those changes afterwards.
Signed-off-by: Jinyan Li <jinyali@amazon.com>
junpuf
approved these changes
Nov 17, 2025
sirutBuasai
pushed a commit
to sirutBuasai/deep-learning-containers
that referenced
this pull request
Nov 18, 2025
* combinr workflow for vllm Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com>
sirutBuasai
added a commit
that referenced
this pull request
Nov 21, 2025
* Improve vLLM workflows (#5480) * combinr workflow for vllm Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * formatting Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix region Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix artifacts path Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * isntall dependencies Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * reverse order Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add port Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * move scripts into their separate dir Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * change port Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * style: format pre-commit check Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * chore: run test choices Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove commitizen Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use main branch Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix dir Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * run benchmark sglang Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * ac run container id instead Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add port and host Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use composite action Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * input secrets Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use shell bash Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove interactive Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * logs tail 200 Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add sleep Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add cleanup Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove -it Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix names Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use container cleanup action Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove unused step Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * try env.containerid Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use env Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add -it Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * full run Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use output Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use vars expose Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use input Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * comment artifacts name Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * set output Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * echo Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add outputs id Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * iamge uri output Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * test using hardcoded string Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * no echo Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * set my output Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use image uri Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use secret image uri file Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix inputs Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * correct docker pull Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use non screte var Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use steps image_uri Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove unused steps Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * change uri var name Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * change step name Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * run regression test Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * change container_pull to ecr_authenticate Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove }} Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use 12xlarge fleet Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use base sglang Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * revert dockerfile Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * run using g6e Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove unnecessary wait Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * rename tests Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use matrix srt test Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove srt Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add pytest requirements Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add add conftest Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * test_utils Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove dup isort Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * run test exampels Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix runner name Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * force rich terminal Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove console width Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * disable rich tracebacks Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add columns Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix intall Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove rich logger Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove rich logger Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add sm endpoint Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * source and install dependencies Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * reduce sleep time Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use serve cmd instead Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * activate venev Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add pytest cache Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * enable debug Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * input image uri Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use assert in test instead Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove f Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * show test output Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use ack test direction Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove predictor Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix self error Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix input Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove model_id yield Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix get endpoint status Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add predictory class Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * change predictor Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * format tests Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * change pytest cmd Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * move aws_session to global conftest Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * move aws_session to global conftest Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix comments Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * fix comments Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove force color from pytest workflow and move to global Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * rename workflow Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * test remove paths Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * add runner setup script Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * uv venv project name Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * use json serializers Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * sagemaker requirements Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * change conccurency name Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> * remove concurrency Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> --------- Signed-off-by: sirutBuasai <sirutbuasai27@outlook.com> Co-authored-by: Jinyan Li <97153458+jinyan-li1@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
GitHub Issue #, if available:
Note:
If merging this PR should also close the associated Issue, please also add that Issue # to the Linked Issues section on the right.
All PR's are checked weekly for staleness. This PR will be closed if not updated in 30 days.
Description
Tests Run
vLLM build and test: d5cb001
rayserve build and test: e73992d
Both tested after removing old workflows and renaming new workflow: ccbfc03
By default, docker image builds and tests are disabled. Two ways to run builds and tests:
How to use the helper utility for updating dlc_developer_config.toml
Assuming your remote is called
origin(you can find out more withgit remote -v)...python src/prepare_dlc_dev_environment.py -b </path/to/buildspec.yml> -cp originpython src/prepare_dlc_dev_environment.py -b </path/to/buildspec.yml> -t sanity_tests -cp originpython src/prepare_dlc_dev_environment.py -rcp originNOTE: If you are creating a PR for a new framework version, please ensure success of the local, standard, rc, and efa sagemaker tests by updating the dlc_developer_config.toml file:
sagemaker_remote_tests = truesagemaker_efa_tests = truesagemaker_rc_tests = truesagemaker_local_tests = trueHow to use PR description
Use the code block below to uncomment commands and run the PR CodeBuild jobs. There are two commands available:# /buildspec <buildspec_path># /buildspec pytorch/training/buildspec.yml# /tests <test_list># /tests sanity security ec2sanity, security, ec2, ecs, eks, sagemaker, sagemaker-local.Formatting
black -l 100on my code (formatting tool: https://black.readthedocs.io/en/stable/getting_started.html)PR Checklist
Expand
Pytest Marker Checklist
Expand
@pytest.mark.model("<model-type>")to the new tests which I have added, to specify the Deep Learning model that is used in the test (use"N/A"if the test doesn't use a model)@pytest.mark.integration("<feature-being-tested>")to the new tests which I have added, to specify the feature that will be tested@pytest.mark.multinode(<integer-num-nodes>)to the new tests which I have added, to specify the number of nodes used on a multi-node test@pytest.mark.processor(<"cpu"/"gpu"/"eia"/"neuron">)to the new tests which I have added, if a test is specifically applicable to only one processor typeBy submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license. I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.