Tracking the v4.1 → v4.2 → v5.0 progression of GoodVibes.
The library is mature (v4.0): six native QC parsers (Gaussian, ORCA, NWChem, Q-Chem 6, xTB, ASE), a comprehensive test suite, and a clean modular structure. This roadmap addresses three structural gaps that day-to-day lab use has surfaced:
- No clean programmatic API.
calc_bbeis invoked with 15 positional arguments and writes to a.datfile; embedding GoodVibes in notebooks or pipelines (CREST → conformer search → thermo) is awkward. - No structured output. Downstream tools and dashboards want JSON /
CSV / Parquet — today they parse the
.dattext. - Doesn't scale to large conformer ensembles.
pes.py/selectivity.py/sort.deduplicatework on lists; a 10³–10⁴ conformer batch is awkward.
Plus several in-flight features that are partially detected
(CBS/Gn composites) or untested (--media, --freespace).
| # | Item | Status | Commit |
|---|---|---|---|
| 1 | CBS/Gn composite method detection (CBS-QB3, CBS-4M, G3, G3B3) | Out of scope — niche use case relative to the v5.x rewrite priorities; revisit only on user request with sample fixtures attached | — |
| 2 | --media / --freespace integration tests + clear errors when solvent is unknown |
✅ Done | 04e2b63 |
| 3 | Test coverage gaps: sort, validation modules | ✅ Done | f88abfa |
| 3a | Test coverage gaps: selectivity, PES modules | ✅ Done — coverage shipped with the redesigns (selectivity ≈87%, pes_loader 100%, pes_model 96%, pes_yaml 90%, pes_legacy 95%) | — |
| 4 | --json OUTPUT.json preview (schema v0.1 → v0.3) |
✅ Done | 6e6187f |
| Sub-plan A | Selectivity redesign: N-way labels, structured SelectivityResult, dual Boltzmann + lowest-conformer output, JSON v0.3 |
✅ Done | 5fb8176 |
| # | Item | Status |
|---|---|---|
| 5 | Programmatic API façade goodvibes.api: compute_thermo(path) -> ThermoResult and compute_batch(...). ThermoResult mirrors calc_bbe's public attributes + the source QCData. Internally just calls calc_bbe — no behavior change. Re-exported from goodvibes/__init__.py. |
✅ Done |
| 6 | Pandas DataFrame export + CSV writer: goodvibes.api.to_dataframe(results), --csv PATH CLI flag, goodvibes[full] extras group (ase + pyyaml + pandas). |
✅ Done |
| 8 | Parallel parsing with concurrent.futures.ProcessPoolExecutor, --jobs N CLI flag, compute_batch(paths, jobs=N). ~3× speedup at 8 cores on 46 azabor PES fixtures (18.6 s → 6.2 s). |
✅ Done |
| Sub-plan B | PES rewrite: 3-layer data model (ConformerSet/Point/Pathway), true-YAML input alongside legacy parser, stoichiometry support, Rich tables, JSON v0.4 pes block, --lowest-only flag. Plot deferred to v5.0. |
✅ Done — 03549d1 |
| Docs | Refresh Read the Docs for v4.1 + v4.2: bumped Sphinx config to v4.0, swapped deprecated sphinxcontrib-napoleon / recommonmark → sphinx.ext.napoleon / myst-parser, added a dedicated api_guide.md page (kwargs API, batch/parallel, DataFrame/CSV, what's-new) and a cookbook.md page with five v4.x recipes, expanded module reference to all 17 modules, added .readthedocs.yaml. Stays on Sphinx; the mkdocs migration is reserved for v5.0. |
✅ Done |
| Follow-up A | Separate ZPE vs harmonic frequency scaling factors. The Truhlar database (Alecu et al., JCTC 2010) stores zpe_fac and harm_fac as distinct per-LOT values. Auto-detection now reads both: harm_fac for the partition-function frequencies (calc_vibrational_energy, calc_rrho_entropy) and zpe_fac for ZPE (calc_zeropoint_energy). New --zpe-vscal CLI flag and zpe_scale_factor API kwarg for explicit override. When --vscal alone is set, ZPE inherits it (back-compat for users with pinned scripts). Pre-refactor v3.x used zpe_fac uniformly; v4.x.0 used harm_fac uniformly — the new behavior matches Truhlar's intent and is the right default going forward. |
✅ Done |
| 7 | Out of scope — covered better by Eyring rate-constant tooling (e.g. kinisot); GoodVibes stays focused on partition-function thermochemistry |
Sequencing. Item 5 was the keystone — unblocked the structured DataFrame/CSV work in 6. Item 8 is independent. Docs refresh sits on top of whichever subset has shipped at the time it's tackled.
| # | Item | Status |
|---|---|---|
| 10 | First-class structured outputs: stable schema_version: "1.0" in goodvibes.schema (TypedDicts + version constants + runtime validator). Parquet export via to_parquet(results, path) + --parquet PATH CLI flag (pandas + pyarrow behind [full] extras). Cache subsumed: new --export / --import flags read/write the unified v1.0 payload; --cache-save / --cache-read deprecated as aliases (still accept the legacy envelope on read). JSON writer no longer gated on --ti mode. |
✅ Done |
| 11 | Ensemble container with lazy parsing, streaming Boltzmann/dedup. Refactor selectivity, pes, sort.deduplicate to consume Ensemble. Handles 10⁴ conformers without holding them all in memory. |
Pending |
| 12 | Clean programmatic API as the headline: from goodvibes import compute_thermo, ThermoOptions, ThermoResult. New ThermoOptions frozen dataclass replaces the 15 positional args. New calc_bbe.from_options(qcdata_or_path, ThermoOptions) classmethod is the v5.0+ entry point — also handles auto-lookup of Truhlar harm_fac/zpe_fac so any caller (not just compute_thermo) gets sensible defaults. Direct calc_bbe(file, QS, ...) calls now emit a DeprecationWarning (legacy form still works through v5.x; removed in v6.0). Internal callers (compute_thermo, parallel worker) migrated to from_options. New docs/source/migration_v5.md page covers the path for users embedding GoodVibes. Ensemble re-export deferred until item 11 lands. |
✅ Done |
| 13 | Conformational entropy correction: S_conf = −R Σ pᵢ ln pᵢ on Ensemble; boltzmann_averaged_G(T) includes −T·S_conf. Wires into pes.get_pes so Gconf becomes first-class. |
Pending |
| 14 | Visualization: new goodvibes/plot.py with plot_selectivity_strip (per-species ΔG scatter + Boltzmann-mean / lowest overlays) and plot_pes (PES profile from PESResult — multi-pathway overlay, bezier or linear connectors, configurable colors, optional point-value annotations, optional per-conformer scatter). New --strip-plot PATH and --pes-plot PATH CLI flags; --graph FILE.yaml retained for YAML-driven styling until v5.1. [plot] extras add matplotlib. Stubs for plot_boltzmann_histogram and plot_temperature_scan lock in the v5.1 API. |
✅ Done |
| 15 | Hindered-rotor treatment (Pitzer–Gwinn / Truhlar HO-QHO). Auto-detect torsional modes from the normal-mode analysis + redundant internals; replace the HO contribution to S/H/ZPE with the rotor partition function. Manual --hindered-rotor MODE,V,I_red override. Moved from v4.2 (was item 9) — auto-detection is the version users actually want; the manual flag alone has too much friction. |
Pending |
Breaking changes. pes.get_pes and selectivity.get_boltz
signatures change (list[dict] → Ensemble) with a one-cycle shim
accepting both. CLI flags + .dat output unchanged across the entire
roadmap.
Builds on v5.0's structural rework (Ensemble, ThermoOptions, structured outputs) and turns those foundations outward toward two concrete user-facing wins — ingesting the dominant conformer-search tool's output natively and generating paper-ready supporting information — while clearing v4.x deprecation debt and finishing the v5.0 visualization stubs.
| # | Item | Status |
|---|---|---|
| 16 | Native CREST / CENSO ensemble import: parse crest_conformers.xyz / crest_rotamers.xyz / censo.json into a list[QCData] (or directly an Ensemble once item 11 lands). New parser entry points parse_crest_ensemble(path) / parse_censo(path) in io.py; format auto-detected by file basename. Closes the gap to the standard CREST → CENSO → thermo workflow without per-conformer-file scripting. Composes with --jobs N for parallel parsing. New fixtures under tests/crest/ and tests/censo/. |
Pending |
| 17 | Auto-SI generator: new goodvibes/report.py + --si PATH CLI flag. Emits a markdown SI bundle: methods paragraph filled from detected program/version/functional/basis/solvent/dispersion/scaling factor, BibTeX entries for cited methods (Grimme qRRHO, Truhlar QH, Truhlar scaling factor, Head-Gordon enthalpy, dispersion correction, solvent model), per-structure coordinates + thermo table, and the literal GoodVibes invocation used. Hand-curated references.bib ships with the package; pandoc-convertible to LaTeX/Word downstream. |
Pending |
| Cleanup | v4.x deprecation removals: drop the legacy line-based PES parser (pes_legacy.py); drop --ee (legacy two-bucket selectivity); drop --cache-save / --cache-read aliases; drop the 15-positional-arg calc_bbe(...) form deprecated in v5.0. CLI emits clear errors pointing at the v5.x replacements. |
Pending |
| Plot | Visualization completions: implement the v5.0 stubs plot_boltzmann_histogram (per-structure population strip/bar) and plot_temperature_scan (qh-G(T) line plot from --ti data). New --boltz-plot PATH and --ti-plot PATH CLI flags. |
Pending |
Sequencing. Item 16 depends on v5.0 item 11 (Ensemble container) — block on that. Item 17 is independent and can land any time. Cleanup and Plot are mechanical follow-ups.
Open design questions (resolve before implementation):
- CREST / CENSO ingest: CREST's
crest_conformers.xyzis geometry+energy only (no Hessian). Decide whether the import path requires paired thermo files or whetherEnsembleshould accept "geometry-only" entries and skip them in qh-G aggregation. CENSO has multiple parts (part0..part4); pick a default for which energy block is canonical for thermo work. - Auto-SI: ship markdown only at v5.1 (cheapest, pandoc-convertible) and defer a native LaTeX writer to v5.2 unless demand is clear. Confirm scope of methods-paragraph templating — GoodVibes-specific (qh, scaling, dispersion) vs. full computational protocol (program, basis, solvation).
- Companion-package boundary: KIE / isotope-substitution thermo lives in
the companion kinisot package. v5.1 should add a documented
Ensembleserialization format that kinisot can consume rather than absorbing those features here.
- Backwards compat. CLI +
.datoutput identical across v4.x. v5.0 may rename internal kwargs; CLI stays. Atests/compatibility/directory will diff.datagainst checked-in goldens for the 20 most-used flag combinations. - Deprecation policy. Anything deprecated in v4.2 is removed no
earlier than v5.1. (Goal: CI runs
pytest -W error::DeprecationWarning; not yet enabled — current CI runs plain pytest with coverage.) - Docs. mkdocs + mkdocstrings auto-API docs at v5.0. Cookbook section
with notebook examples (CREST →
Ensemble→ Boltzmann G). - Performance targets. 1k conformers parsed in <30s on 8 cores; ensemble Boltzmann/dedup memory <200 MB at 10k conformers.
- CI. Add Python 3.13 at v5.0 cut. Lint (ruff, Pyflakes + syntax errors, blocking) and pytest coverage reporting added 2026-06.
- Code health.
AUDIT.md(2026-06, plus focused re-audit) holds the full repository audit and prioritized improvement plan. Top open items: parse-failure diagnostics on QCData (~45 silent fallback sites across the six parsers), the silent SPC downgrade (missing/unparseable--spcfile yields a'!'sentinel and the correction is skipped without any API-visible warning — thermo.pyexcept TypeError: pass), wheel slimming (verified: a 21.5 MB keep-set suffices;pes/*.log≈256 MB is referenced by no test), and decomposition ofcalc_bbe.__init__/io.parse_dataalong their existing phase seams ahead of the item-11 Ensemble refactor.sys.exitcleanup is layering hygiene only — the API path cannot currently reach those sites.
A worked example of how items in this roadmap are designed before implementation begins.
- N-way selectivity: generalize from 2-bucket ee to N-bucket dr (e.g. four diastereomers). The 2-bucket case is a special instance.
- Explicit labels instead of fragile filename globs. Decouple algorithm from filename templating.
- Structured
SelectivityResultdataclass + JSON output. - Temperature scan that composes naturally with
--ti. - Dual reporting: Boltzmann-averaged AND lowest-conformer-only, so the user can see how much selectivity comes from the gap between the lowest TSs vs. conformer mixing.
@dataclass(frozen=True)
class SelectivityResult:
temperature: float # K
key: str # 'gibbs' | 'energy'
labels: List[str] # ordered species names
files_per_label: Dict[str, List[str]]
populations: Dict[str, float] # normalized: Σ = 1.0
raw_boltzmann: Dict[str, float]
preferred: str # max-population label
ee: Optional[float] = None # 2-label only, in %
ddG: Optional[float] = None # 2-label only, in HartreeNumeric data only. Ratio strings (60:40, 40:30:20:10) are derived
in the print layer from populations. For N>2, ee and ddG are None;
consumers derive any ratios they want from populations.
--label NAME=PATTERN(repeatable; fnmatch against basenames of files already inthermo_data— no filesystem walks).--selectivity FILE.yaml(alternative for many species or shareable specs). Top-level keylabels:for patterns, orfiles:for explicit per-species file lists.--label/--selectivitycombine with--tifor temperature scans.--ee 'a:b'keeps working in v4.x with aDeprecationWarning; removed in v5.0.
- Two stacked Rich tables per result: Boltzmann-averaged + Lowest conformer only. Each has Species/Files/Population (%)/ΔΔG (kcal/mol) columns plus a summary line (ratio, major, ee + ΔΔG‡ for N=2).
- Temperature scan: one row per T per method.
- JSON schema v0.3: top-level
selectivityandselectivity_lowestblocks, both with the same shape.
- Pattern matching:
fnmatchagainst basenames of files already inthermo_data. No filesystem walks. The candidate set is exactly what the user passed on the command line. - Ratio formatting: numeric data only on the dataclass; ratio strings live only in the print layer.
- N=2 vs N>2 reporting: N=2 emits ratio + ee + ΔΔG‡; N>2 emits only the ratio. No pairwise data.
- Empty species:
compute_selectivityraisesValueError;main()translates tofatal()for the CLI. dup_listsemantics: each pair is[duplicate, canonical]— excluding onlydup[0]is correct; the canonical structure stays in the sum.
The current pes.get_pes is line-based parsing (despite the .yaml
extension and --- document markers), and the internal data model is
~30 parallel lists nested four levels deep
(g_qhgvals[pathway][step][structure][conformer]). New features
(stoichiometry, JSON output, conformational entropy in v5.0) are hard
to wire in without a redesign.
- Three-layer data model that matches how chemists think about a PES: pathway → point → species → conformers.
- Two input formats: keep the existing line-based "yaml" working
(auto-detected,
DeprecationWarning, removed in v5.1); add a true YAML schema as the documented format. - Stoichiometry:
2*A + Bsyntax for multi-mol reactions, valid in both formats. - Rich tables + JSON v0.4 output, mirroring the v4.1 selectivity redesign style.
- Plot deferred —
graph_reaction_profileis matplotlib-tangled and orthogonal; rewritten as part of v5.0 visualization.
@dataclass(frozen=True)
class ThermoVector:
"""One species' thermo bundle in Hartree. Supports +, -, * scalar."""
sp_energy: Optional[float]
scf_energy: float
zpe: float
enthalpy: float
qh_enthalpy: float
entropy: float # S, not T·S — temperature lives on Pathway
qh_entropy: float
gibbs: float
qh_gibbs: float
@dataclass
class ConformerSet:
"""A named species + ≥1 calc_bbe entries; encapsulates the gconf math."""
name: str
files: List[str]
def boltzmann_weighted(self, T: float) -> ThermoVector: ...
def lowest_conformer(self) -> ThermoVector: ...
def gconf_corrected(self, T: float, QH: bool = True) -> ThermoVector: ...
@dataclass
class Point:
"""Stoichiometric sum at one node: [(coeff, ConformerSet), ...]."""
label: str # "Int-I + TolS + TolSH" or "2*A + B"
species: List[Tuple[int, ConformerSet]]
def thermo(self, T: float, gconf: bool = True, QH: bool = True) -> ThermoVector: ...
@dataclass
class Pathway:
name: str
points: List[Point]
zero: Point # default = points[0]
def relative(self, T: float, ...) -> List[ThermoVector]: ...
@dataclass
class PESResult:
pathways: List[Pathway]
options: PESOptions # units, dp, gconf, QH
temperatures: List[float] # supports --tiThermoVector arithmetic collapses 9 simultaneous parallel-list
subtractions into one expression:
relative = point.thermo(T) - zero.thermo(T).
Coefficient syntax: 2*A, 2 * A, A (implicit 1). Integer
coefficients only. Parsing is regex-based on the point string and
applied identically in both input formats. The thermo math is
multiplication on ThermoVector (2 * A.thermo(T) + B.thermo(T)).
Legacy (auto-detected by --- # PES / # SPECIES / # FORMAT comment markers):
parser kept for v4.x; emits DeprecationWarning on use; removed v5.1.
The example at goodvibes/examples/pes/azabor_PES.yaml
(6-step pathway with conformer ensembles, stoichiometric sums, and
--spc tzpop pairs) is the back-compat regression fixture.
True YAML:
pathways:
Reaction: ["Int-I + TolS + TolSH", "Int-II + TolSH", "Int-III"]
species:
Int-I: {files: "Int-I_*.log"} # glob
Int-II: {files: "Int-II_*.log"}
Int-III: {files: "Int-III_*.log"}
TolS: {files: "TolS.log"}
TolSH: {files: "TolSH.log"}
zero:
Reaction: "Int-I + TolS + TolSH" # optional; default = points[0]
format:
units: kcal/mol
decimals: 1Each species is a dict to leave room for future per-species options
(scaling factor overrides, symmetry numbers, ...) without breaking the
schema. files: accepts a string (glob or single file) or list.
Rich table per pathway × per temperature, columns matching the
current .dat layout (DE, DZPE, DH, qh-DH, T·DS, T·qh-DS,
DG(T), qh-DG(T), plus SPC variants when --spc is set).
JSON v0.4 adds a top-level pes block:
"pes": {
"pathways": [{
"name": "Reaction",
"temperature": 298.15,
"units": "kcal/mol",
"points": [{
"label": "Int-I + TolS + TolSH",
"species": [
{"coefficient": 1, "name": "Int-I", "files": ["Int-I_Oax.log"]},
{"coefficient": 1, "name": "TolS", "files": ["TolS.log"]},
{"coefficient": 1, "name": "TolSH", "files": ["TolSH.log"]}
],
"relative": {"scf": 0.0, "zpe": 0.0, "h": 0.0, "qh_h": 0.0,
"ts": 0.0, "qh_ts": 0.0, "g": 0.0, "qh_g": 0.0}
}]
}]
}.dat text output unchanged across v4.x.
| File | Change |
|---|---|
goodvibes/pes_model.py |
New: ThermoVector, ConformerSet, Point, Pathway, PESResult. Pure data model + arithmetic; no I/O. |
goodvibes/pes_legacy.py |
New: extract current pes.get_pes parser, target the new model. Emits DeprecationWarning. |
goodvibes/pes_yaml.py |
New: true-YAML parser → new model. Stoichiometry parsing lives here (shared with legacy). |
goodvibes/pes.py |
Becomes a dispatcher: sniff format, parse, return PESResult. get_pes retained as back-compat shim until v5.1. |
goodvibes/output.py |
New print_pes_tables(result) Rich renderer. Existing text path kept for v4.x. JSON writer extended with pes block. |
goodvibes/GoodVibes.py |
No new flags; --pes FILE already handles both formats. |
tests/test_pes.py |
New: model tests (stub data), parser tests (both formats), integration test against examples/pes/. |
| Version | Behavior |
|---|---|
| v4.2 | Both parsers work; legacy emits DeprecationWarning. New model is the runtime backbone. JSON schema bumps to 0.4. |
| v5.0 | Refactor ConformerSet onto Ensemble (item 11). Pathway.relative() becomes streaming. No user-visible change. |
| v5.1 | Legacy parser removed. CLI fails with a clear error pointing at the new schema. |
- Stoichiometry: integer coefficients via
n*Speciessyntax in both formats; default 1. - Plot rewrite: deferred to v5.0 visualization.
graph_reaction_profilekeeps reading the legacyget_pesattributes via a thin compat shim in v4.2 — no behavior change. - Legacy timing: deprecated in v4.2, removed in v5.1.
- Slot in roadmap: v4.2, alongside item 5 (API façade) and item 6
(DataFrame export) — both depend on a clean
PESResultfor JSON serialization.