Skip to content

Add support for the pure container version of SONiC - #320

Draft
Sven-Ric wants to merge 1 commit into
masterfrom
feat/sonic-container-support
Draft

Add support for the pure container version of SONiC#320
Sven-Ric wants to merge 1 commit into
masterfrom
feat/sonic-container-support

Conversation

@Sven-Ric

Copy link
Copy Markdown

Description

Needs more testing, but machine and firewall can be allocated and phone home.

The VM images should remain the default, as they closely resemble the actual hardware experience, but this branch roughly halves the RAM requirements of the mini-lab, making multi-site simulations a lot more achievable in preparation of MEP-19.

Used AI-Tools ✨

  • Fable 5

@Sven-Ric Sven-Ric self-assigned this Jul 27, 2026
@simcod simcod moved this to Upcoming in Development Jul 27, 2026
@simcod
simcod requested a review from l0wl3vel July 27, 2026 12:31
@l0wl3vel

Copy link
Copy Markdown
Contributor

I do not think this is a good idea. I appreciate you putting in the work and testing it out, but building a custom sonic startup again is a maintenance time bomb. Just spent a bunch of time on cleaning up the launch.py and we get a new one 😅

There are generic optimizations, like Kernel Same-Page Merging, reducing the maximal memory for the VMs, and reducing Transparent Huge Page usage if not explicitly requested by the VM.

I whipped up a quick memory benchmark and some generic tweaks for QEMU/Linux Memory Management using the sonic-vpp branch and here are the results from a previous run:


Memory profile comparison

Peak values across the whole integration run. Δ columns compare against the baseline profile of the same flavor.

dell_sonic

profile host used (peak) Δ host VM containers (peak) Δ VMs QEMU RSS (peak) swap (peak) KSM saved duration
baseline 25.51 GiB 14.43 GiB 11.47 GiB 0.00 GiB 0.00 GiB 702s
ksm 23.24 GiB -2.27 GiB (-8.9%) 12.19 GiB -2.24 GiB (-15.5%) 10.32 GiB 0.00 GiB 2.06 GiB 698s
low-memory 17.38 GiB -8.13 GiB (-31.9%) 4.96 GiB -9.47 GiB (-65.6%) 4.70 GiB 0.00 GiB 0.00 GiB 460s
thp-madvise 23.95 GiB -1.56 GiB (-6.1%) 14.41 GiB -0.02 GiB (-0.2%) 11.46 GiB 0.00 GiB 0.00 GiB 672s

sonic_vpp

profile host used (peak) Δ host VM containers (peak) Δ VMs QEMU RSS (peak) swap (peak) KSM saved duration
baseline 26.13 GiB 17.23 GiB 11.82 GiB 0.00 GiB 0.00 GiB 784s
all 19.99 GiB -6.15 GiB (-23.5%) 12.36 GiB -4.87 GiB (-28.3%) 7.17 GiB 0.00 GiB 1.47 GiB 840s
ksm 23.82 GiB -2.31 GiB (-8.9%) 14.64 GiB -2.59 GiB (-15.1%) 10.18 GiB 0.00 GiB 1.56 GiB 713s
low-memory 23.80 GiB -2.34 GiB (-8.9%) 14.03 GiB -3.20 GiB (-18.6%) 8.10 GiB 0.00 GiB 0.00 GiB 734s
thp-madvise 25.25 GiB -0.88 GiB (-3.4%) 17.62 GiB +0.39 GiB (+2.3%) 11.81 GiB 0.00 GiB 0.00 GiB 718s

sonic_vs

profile host used (peak) Δ host VM containers (peak) Δ VMs QEMU RSS (peak) swap (peak) KSM saved duration
baseline 25.94 GiB 17.20 GiB 11.81 GiB 0.00 GiB 0.00 GiB 794s
all 20.80 GiB -5.14 GiB (-19.8%) 12.85 GiB -4.36 GiB (-25.3%) 7.69 GiB 0.00 GiB 0.93 GiB 738s
ksm 24.39 GiB -1.55 GiB (-6.0%) 15.10 GiB -2.11 GiB (-12.2%) 10.99 GiB 0.00 GiB 2.43 GiB 1017s
low-memory 22.62 GiB -3.32 GiB (-12.8%) 13.83 GiB -3.37 GiB (-19.6%) 8.13 GiB 0.00 GiB 0.00 GiB 809s
thp-madvise 25.31 GiB -0.63 GiB (-2.4%) 17.53 GiB +0.33 GiB (+1.9%) 11.86 GiB 0.00 GiB 0.00 GiB 703s

The the most optimistic case with all optimizations on shaved of 23% of total memory usage, and we have not even touched the control plane yet, which consumes about ~10Gi of the remaining ~20Gi

I did not look into the control plane consumption in detail, but I expect to find a significant amount of memory consumption to be caused by the monitoring stack.

In conclusion: I believe ~2.5Gi per switch down from 4Gi is manageable, even for multi-site scenarios, considering we are not loosing any accuracy.

On the other hand, I would like to see how low this would go. CI is failing, did you get this to work yet?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Upcoming

Development

Successfully merging this pull request may close these issues.

3 participants