-
The configuration of the version being installed is now expanded by running that version's config providers, in a temporary VM booted from that version's own boot script on its own emulator, rather than by running provider state stashed at build time in the version that happens to be running. A provider module can differ between the two — which is precisely what an upgrade may change — and only the target's own answer is the right one. It also leaves Elixir's
Config.Provideras the single implementation of the provider pipeline: Castle drives it and no longer keeps a copy of it.The temporary VM needs no epmd, no cookie, no node name and no distributed Erlang: it talks to the running node over a socket on the loopback interface, and whatever it prints — a provider explaining what it could not find, say — arrives on the terminal that asked for the install. It is stopped on every way out, including every failing one, and it cannot hold an install open: both its boot and the work it is asked to do have deadlines. Everything that can refuse to go on refuses before the upgrade is applied, so configuration that cannot be expanded leaves an install that did not happen rather than one that half did.
Each expansion starts from the configuration the release was built with, which the first one copies aside as
sys.config.pristineand none of them overwrites. Config providers are not obliged to be idempotent, and the familiar ones are not: aruntime.exsthat sets a key only when an environment variable is present says nothing about that key when it is absent, so expanding over the previous result would leave a value behind after the provider had stopped supplying it — and the version made permanent would be configured differently from the way it goes on to boot. Expanding from the original instead means installing and then committing produce the same answer a boot would, which is the point of expanding at either. That copy is made atomically and with the permissionssys.confighas, so a partly written one can never be found and read, and restrictingsys.config— as an operator might, since it holds credentials — restricts this too.Among the things that refuse is the check Elixir makes on a configuration before booting into it: that what
Application.compile_env/3read when the release was compiled is what the resolved configuration says now. A version whose runtime configuration contradicts what it was compiled against is refused here, where refusing costs nothing, rather than accepted and then found to be unbootable — which, for an upgrade that restarts, is found on the way back up with a rollback as the only way out.This is how every release is configured now, and the only way: the path that read a
build.configis gone, along with the build-time interception that produced one — see Removed below. -
Castle.unpack/1andCastle.install/1now refuse a system that cannot be upgraded from, and refuse it in the same call that would otherwise have done the work.:release_handlerreadsreleases/RELEASESonce, as it starts, and when it cannot — the file absent, or there but not consultable — it makes a release record up out of the boot script's name and version — a record that names no applications at all. Upgrading a system in that state is worse than being stopped: the install reports success, and every application whose version changed but whose code the upgrade does not explicitly load goes on running its old code out of the directory of the release that was just replaced, until a laterremovedeletes it. Nothing can repair the running system afterwards, because creating the file changes no record the node holds — so what the refusal says is to restart, with the file either absent or consultable first. See Fixed below for what that condition is and why a bare restart is not always enough.The question is asked of the node's own records rather than of the filesystem, which is the only way to see the case where the file exists but the boot that went looking for it was earlier — and it is asked by the operation, rather than of the operator beforehand. A check made in one call and acted on in another is a check about a moment that has passed: the node can restart in between, and the node that comes back makes a fresh record up, so the operation would go ahead on an answer that no longer held.
Unpacking is refused as well as installing because it is the one other operation that writes release records: an unpack on such a node would put the made-up record into
releases/RELEASES, where the next boot would read it back as though it belonged there — which takes away the restart that is the way out, since the file is only created when it is absent. Committing and removing are unaffected: neither can write that record back, and refusing them could strand a version that was already installed. -
Castle.upgradable/0, which answers the same question on its own, for an operator who wants to know whether a system can be upgraded from without unpacking or installing anything. Nothing has to call it first — the operations that need the answer get it for themselves. -
Castle.Error, the exception raised by a release-management command that did not succeed. -
Castle.running/1, which succeeds when the version it is given is the release the system is running, and fails otherwise.install/1reports what:release_handlerreplied, and that reply says only that the upgrade was accepted: a transition that restarts the emulator is replied to and then rebooted, and for an emulator upgrade the instructions run on the way back up, where they can still fail and roll back. Completion therefore has to be observed rather than inferred, and this is what makes it observable: a caller that repeats the question until it is answered - which is what Forecastle'sbin/castle installdoes, from its own 1.0.0 - can tell an upgrade that took effect from one that did not. Castle supplies the answer; it does not do the asking.Confirmation needs two things: the version is the release the system is running - the one whose status is
current, or thepermanentone if none is current, so a version is confirmed both before and aftercommit- and its boot has finished. The second is not redundant. A node that restarted into the new version is reachable long before it has finished booting:release_handlermakes the versioncurrentwhilesaslstarts, and distribution answers fromkernelonwards, so an application started aftersaslcan still fail and take the system back to the version that was permanent. Confirming earlier than that would let automationcommita release that cannot boot."Finished booting" means the boot script reached its
{progress, started}instruction, which is the last thing every boot script Mix generates does. A release booted with aRELEASE_BOOT_SCRIPTthat names a hand-written script without that marker will therefore never be confirmed:installwaits and then fails, and the refusal names the progress the node did reach, so it says what is wrong rather than failing silently - but it will not succeed. A script that emits the marker before its applications are started defeats the check instead, since the marker is all there is to go on.Nothing can build a relup that restarts the emulator until forecastle#4, so the restart transitions this addresses cannot be exercised end to end yet. The hot-upgrade path is covered by Forecastle's end-to-end suite; the statuses themselves are covered by unit tests here.
- Raised the minimum Elixir requirement to 1.18.
make_releases/0no longer depends on the working directory. It looks forreleases/RELEASESunder the root of the release -code:root_dir()- so a caller that used to change directory before calling it no longer has to. On a release built by Mix that is the file OTP writes; a deployment that setsRELDIRor thesaslreleases_dirparameter moves the release records elsewhere, and Castle does not yet follow them.unpack/1,install/1,commit/1,remove/1andmake_releases/0now fail when the operation fails, instead of printing the reason and returning normally. These are invoked overbin/castle, which reaches them byrpc, and by the launcher's prebooteval, so the reason now arrives on the caller's standard error and the command exits non-zero:bin/castle unpack "$VSN" && bin/castle install "$VSN"stops at the step that failed, and an operation the system refused no longer exits 0. The running system is unaffected - the failure is raised on the node, re-raised in the short-lived VM that made the call, and it is that VM which exits. What a successful command reports is unchanged.
Castle.generate/1, and with it the path throughinstall/1andcommit/1that read abuild.config. Expanding the target's configuration by folding provider state stashed at build time over a renamedsys.config, in whichever version happens to be running, is what the temporary VM above replaces - and from Forecastle 1.0.0 nothing assembles a release that has abuild.configto read. Runtime configuration on a normal boot is Mix's own again, and the configuration of a version being installed is expanded by that version's own providers.
-
The refusal for a system running from a synthesised release record now names a remedy that works. It said to restart, and a restart alone is enough only when the
RELEASESfile:release_handlerreads is absent or consultable: it reads that file withfile:consult/1, so a malformed one fails just as an unreadable one does, and the release creates the file only when it is absent — so anything left in place that cannot be consulted is stepped over on every start and the system comes back on another synthesised record. An operator following the old message would have restarted indefinitely.The refusal now asks for that file to be absent or consultable before the restart, and identifies it rather than assuming:
releases/RELEASESunder the release root, unlessRELDIRor thesaslreleases_dirparameter points elsewhere. Where one of those does, the two are different files and a restart cannot fix it on its own — the release creates the one at the root, which the handler will not read — so the file the handler does read has to be put there by hand. Castle following those overrides itself is #23. -
A release built with
include_erts: falseis now refused, by name and with the reason, rather than quietly managing the Erlang installation it happens to be running on. Such a release ships no emulator, so it runs the system one, andcode:root_dir()— the directory:release_handlerextracts applications into, resolves everylib/<app>-<vsn>against, and deleteserts-<vsn>from — is then the shared Erlang installation rather than the deployment. Left to itself,make_releases/0created that installation'sreleases/RELEASES, which usually fails for want of permission and, where it succeeds, puts the release records of unrelated deployments in one file;unpack/1,install/1andcommit/1wrote into the installation, andremove/1deleted out of it. Each of those now fails instead, with a message naming both directories and saying that the deployment cannot be upgraded by Castle. The same refusal covers a release that did bring its ERTS but is run withERL_ROOTDIRset, which the release's ownerlhonours ahead of its location: what makes an upgrade unsafe is that the two directories differ, so the message reports that and offers the causes as examples rather than asserting one.Relocating the release records with
RELDIRor thesaslreleases_dirparameter does not make such a deployment upgradable, and the refusal says so. Those really do move the records —release_handlerreads them ahead of the emulator's root — but they move only the bookkeeping: applications are still extracted into, resolved against and deleted out of the emulator's root, which the handler keeps as separate state.The remedy is to make the two directories the same one — most often by building the release with its ERTS included, and where
ERL_ROOTDIRis what moved them apart, by unsetting it. There is no third option in which Castle is pointed somewhere else instead::release_handlerresolves the applications themselves against the emulator's root, so records kept anywhere else describe applications the handler is not using. A refusal that says so is better than a divergence that does not.Where the two directories cannot be compared at all — a
statrefused by a mode on a parent, a path that is not there, a filesystem reporting no inode numbers — the refusal says that, naming what stopped the lookup, rather than reporting a difference it did not establish. It still refuses, because a comparison that could not be made is no licence to write release records into a tree that has not been shown to be the right one.upgradable/0andreleases/0are deliberately unaffected. They only read, and they are what an operator needs working in order to make sense of the refusal.Note that this reaches the launcher's preboot step, which is where
make_releases/0is called, so such a deployment meets the refusal on every start rather than once: the file the step looks for never appears, so the step runs again each time.What that costs a start is Forecastle's to decide, not Castle's — Castle reports the failure and the
env.shfragment chooses what to do with it. Under Forecastle 1.0.0 the fragment warns and carries on, so an affected deployment still starts, at the price of a warning and a short-lived VM per boot. Pair Castle 1.0.0 with Forecastle 1.0.0; an older fragment treats a failure of that step as fatal and would stop such a deployment starting at all. -
install/1reports the emulator restart that an upgrade to a new emulator, or to a new kernel, stdlib or sasl, needs - rather than failing with aCaseClauseErrorwhile the upgrade proceeds. -
releases/0reports nothing at all, rather than raisingEnum.EmptyError, when no releases are installed. -
make_releases/0says what went wrong - which file could not be read or written, and why - rather than raisingMatchError.