Skip to content

Fix nfd gc not being scheduable#228

Open
erezzarum wants to merge 1 commit intoawslabs:mainfrom
erezzarum:fix-nvd-deviceplugin-gc
Open

Fix nfd gc not being scheduable#228
erezzarum wants to merge 1 commit intoawslabs:mainfrom
erezzarum:fix-nvd-deviceplugin-gc

Conversation

@erezzarum
Copy link
Contributor

What does this PR do?

🛑 Please open an issue first to discuss any significant work and flesh out details/direction. When we triage the issues, we will add labels to the issue like "Enhancement", "Bug" which should indicate to you that this issue can be worked on and we are looking forward to your PR. We would hate for your time to be wasted.
Consult the CONTRIBUTING guide for submitting pull-requests.

The NVIDIA device plugin have a node feature discovery (NFD) enabled, the garbage collector (GC) is not a DaemonSet (DS) but a Deployment that should run on a CPU node to report back to the NFD master.

By default the GC deployment will not tolerate GPU nodes which makes it not being able to schedule as we had a nodeSelector set for NVIDIA GPU

Motivation

Make sure the NFD GC is running and the nvidia-device-plugin ArgoCD app is Healthy

More

  • Yes, I have tested the PR using my local account setup (Provide any test evidence report under Additional Notes)
  • Mandatory for new blueprints. Yes, I have added a example to support my blueprint PR
  • Mandatory for new blueprints. Yes, I have updated the website/docs or website/blog section for this feature
  • Yes, I ran pre-commit run -a with this PR. Link for installing pre-commit locally

For Moderators

  • E2E Test successfully complete before merge?

Additional Notes

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants