Skip to content

UG Migration to Brain Brew Rust - Federation support and new workflows - #743

Draft
jeprecated wants to merge 30 commits into
anki-geo:masterfrom
jeprecated:brainbrew-migration
Draft

UG Migration to Brain Brew Rust - Federation support and new workflows#743
jeprecated wants to merge 30 commits into
anki-geo:masterfrom
jeprecated:brainbrew-migration

Conversation

@jeprecated

Copy link
Copy Markdown
Member

Replacement PR for #736 which is on the right branch. See that PR for the full description.

Summary (as of the opening of this PR):

  • Migrates Ultimate Geography from legacy Python Brain Brew to pinned Rust Brain Brew 1.0.0-alpha.3.
  • Converts deck sources to canonical YAML with reusable HTML, CSS, description, and media includes.
  • Preserves existing deck/model identities, note GUIDs, templates, learner data, and CrowdAnki output.
  • Adds typed translation overlays, structured messages, language metadata, and explicit target adaptations/deletions.
  • Introduces stable !image references and SHA-256-verified media declarations.
  • Reorganises Standard, Extended, Experimental, and Hardcore variants into composable overlays with safe companion identity.
  • Adds strict CI verification and export for all 100 targets, plus parsed-JSON goldens and reproducible historical-equivalence evidence.

This is a Work in Progress. Things are subject to change.

@jeprecated

Copy link
Copy Markdown
Member Author

Copying my previous comment to here to have a better throughline for the conversation:


Hello sirs!

Sorry for the delay in replying, I've been swamped with work. Both in the professional sense and also doing more work on this PR!

Wow! This looks like a huge amount of work!

It enables things that are currently tricky (e.g. flexibly allowing different note types!) and AFAIU generally makes future extensibility easier.

I greatly appreciate how you've continued working on this for so long and I'm generally very excited about improvements to our tooling!

Big thinking work, but it actually didn't take long to do. I've thought about this problem on and off since my email to you two about this rewrite idea 10 months ago, but the actual effort for my initial PR was < 6 hours of actual work. More on that below.

I wish my review below were more positive, but at the moment I'm not convinced that this is a net improvement.

However, IMO it makes routine tasks that are currently straightforward more finicky and tedious. (OTOH it's possible that I'm misunderstanding how a new, improved workflow would work!)

(Details/discussion below; I've been sitting on it for several weeks now, being hesitant to post the comment (sorry for the resultant delay!!), but I can't think of any particular improvements to the text.)

No worries at all, sir. I am not perturbed. Next time don't feel the need to delay, I was not expecting to simply merge this in. I truly meant it when I said:

This is a work-in-progress migration to the new Rust Brain Brew workflow

Though perhaps I could have been clearer on my expectations, the amount of effort I've actually put in here, and my workflow. Though that gets into some of my personal situation, AI/Agentic Development, and other stuff like that - which I don't want to turn this PR into a discussion/debate on these topics. But just for the context of what is to come, and to set your expectations on the state of my work and how I get it done (and why you can trust it) then here's a collapsed box explaining just that:

How I developed and validated this migration

I did all of the work for this Brain Brew rust migration using Agentic Development (with AI). In my spare time, while running 5 other projects.

Wow Jordan you mean you vibe coded it? "Brain Brew Rust Rewrite, make no mistakes [enter]"

No! 😁 Allow me to explain by first adding context.

I created my own startup 4 months ago: Self-Deprecated https://selfdeprecated.ai/ an AI Solutions company, to help bring clients up to speed on Agentic Development. On how to use Agents, what their strengths are, when not to trust them, and how to make sure your "self-healing loop" is the best it can be in order to get the output you desired, while remaining in a "human-out-the-loop" for as much as possible.

I use Hunk (https://github.com/modem-dev/hunk/) to get the Agent to make PR comments on it's changes, and I review the actual code. Do I read absolutely every line always? Absolutely not, and nor should I. It depends on the system being created/changes. Hunk helps one have the Agent walk them through a change/PR step by step, while looking at the code, explaining in as much detail what the changes were for and how they relate to the other code segments.

For this work (migrating an existing tool from one language to another, though with some bit feature changes too) I had an amazingly easy job. The bulk of the work is in thinking and designing the system, while the core migration is unusually straightforward to test as an output-equivalence problem. Brain Brew’s job (before and after the rewrite) is to produce CrowdAnki decks. Given the same accepted source, it either produces the same parsed JSON and media or it produces a regression, apart from differences we have explicitly reviewed and classified. The migration evidence compares deck and model identity, fields, cards/templates/CSS, notes/GUIDs/tags, descriptions, configuration, and media bytes.

Yet at the same time I've add a lot more features! Though almost none of these change the output in the end, they just give better composition and flexibility to the data as it's stored in the repo. Such as combining decks together for the federation.

In short:

  • I sell my services as a "Senior Agentic Developer" to both companies and individual developers to skill themselves up.

  • I have comprehensively tested the migration’s CrowdAnki output surface. In this PR, strict verification covers 74 main and 26 companion targets, CI exports all 100, and eight parsed-JSON goldens detect current-output drift.

    Brain Brew itself also now contains a complete pinned copy of this UG source, including all 607 media assets and the expected parsed CrowdAnki output for every one of the 100 targets. Its mandatory offline fixture test rebuilds and compares all 100 outputs and validates the real media bytes, so future changes to Brain Brew cannot silently alter UG output.

    Separately, ten representative pinned historical builds reject unclassified identity, content, template, configuration, or media changes. The known capital-hint localisation fallback remains explicitly documented and deferred rather than hidden by the comparison.

  • The updating of this system is now much easier than it was before, due to both the better setup of the codebase, and also the proliferation of these new tools.

So all of the above said:

  • This is a project with a very fast improvement timeframe
  • The maintainer / translator use case / workflow is top priority
  • It is unusually testable for a migration.
    • Strict verification covers all 74 main and 26 companion targets, and CI transactionally exports all 100.
    • Eight parsed-JSON goldens detect current-output drift, while pinned historical builds reject unclassified semantic or media changes.
    • It's really a pleasure to work in.

Work Done to Address the comments

Rather than line by line reply to each comments (as we three are want to do) I think it'd be faster for everyone if I simply address the things that I have improved. Some I always intended to do, and some are direct responses to good feedback from yourselves

Translation overlay format cleanup

Example: old and new translation overlay formats

Before: old alpha schema mixed several intents under translations. "changes", "additions", each of which could be general or field-specific

translations:
  changes:
    Taiwan: "台湾(Taiwan)"
    Taipei: "台北(Taipei)"

    Georgia:
      notes.note.georgia.fields.field.country: Georgien
      notes.note.us-georgia.fields.field.region: Georgia

    Partially recognised state claimed by China.:
      notes.note.taiwan.fields.field.country-info: "中国宣称对台湾拥有主权,但仅被部分国家承认"

  additions:
    notes.note.yellow-sea.fields.field.country-info: "虽然国际上普遍认为黄海包含了渤海,但在中国,它们通常被视为两个独立的海域。"

After: reusable translations, contextual translations, and typed target adaptations are separate

translations:
  direct:
    Taiwan: "台湾(Taiwan)"
    Taipei: "台北(Taipei)"

  contextual:
    notes.note:
      andorra.fields.field.flag-similarity.message.variables.country_1:
        Moldova: "摩尔多瓦"

target_adaptations:
  notes.note.taiwan.fields.field.country-info:
    intent: adapt
    ownership: translation
    expected_source: Partially recognised state claimed by China.
    target: "中国宣称对台湾拥有主权,但仅被部分国家承认"
    reason: target-language geopolitical wording

  notes.note.yellow-sea.fields.field.country-info:
    intent: adapt
    ownership: translation
    expected_source: ""
    target: "虽然国际上普遍认为黄海(Yellow Sea)包含了渤海(Bohai Sea),但在中国(China),它们通常被视为两个独立的海域。"
    reason: "migrated from legacy translations.target_additions; review and describe its target-language purpose"

  notes.note.czech-republic.fields.field.country-info:
    intent: delete
    ownership: translation
    expected_source: Also known as Czechia.
    reason: "migrated legacy target adaptation; review and describe its target-language purpose"

Reviewed source text that should intentionally remain unchanged can be listed under translations.no_change instead of being copied into a direct source-to-source translation. Eleven of UG's language overlays use that distinction.

Better authoring formats

  • Added !include support for external descriptions, CSS, card-template HTML, and other supported YAML values.
Example: external description, CSS, and template includes
deck:
  description: !include descriptions/ultimate-geography/en.html

note_types:
  note-type.ultimate-geography:
    styling: !include styles/ultimate-geography/card.css
    card_templates:
      template.country-capital:
        question_format: !include templates/ultimate-geography/country-capital/question.html
        answer_format: !include templates/ultimate-geography/country-capital/answer.html
  • Added structured translatable messages with format, ref, text, and literal
Examples: reusable message parts and translated formats

Before: one long translation key

# src/data/flag_similarity.csv, represented as YAML for comparison
field.flag-similarity: Iceland (blue background, red and white cross), Norway (red background, blue and white cross)
# The Norwegian translation was also one indivisible string
translations:
  direct:
    "Iceland (blue background, red and white cross), Norway (red background, blue and white cross)": "Island (blå bakgrunn, rødt og hvitt kors), Norge (rød bakgrunn, blått og hvitt kors)"

After: reusable translated parts

# deck.yaml, notes.note.faroe-islands
field.flag-similarity:
  format: "{country_1} ({description_1}), {country_2} ({description_2})"
  variables:
    country_1:
      ref: notes.note.iceland.fields.field.country
    country_2:
      ref: notes.note.norway.fields.field.country
    description_1:
      text: blue background, red and white cross
    description_2:
      text: red background, blue and white cross
# overlays/languages/nb.yaml
translations:
  direct:
    Iceland: Island
    Norway: Norge
    blue background, red and white cross: "blå bakgrunn, rødt og hvitt kors"
    red background, blue and white cross: "rød bakgrunn, blått og hvitt kors"

Translating the format itself

Base source:

field.flag-similarity:
  format: "{country_1} ({description_1}), {country_2} ({description_2})"
  variables:
    country_1:
      ref: notes.note.iceland.fields.field.country
    country_2:
      ref: notes.note.norway.fields.field.country
    description_1:
      text: blue background, red and white cross
    description_2:
      text: red background, blue and white cross

The pieces can be translated normally:

translations:
  direct:
    Iceland: Island
    Norway: Norge
    blue background, red and white cross: "blå bakgrunn, rødt og hvitt kors"
    red background, blue and white cross: "rød bakgrunn, blått og hvitt kors"

But the format string itself can also be translated when the target language needs different punctuation, spacing, order, or glue.

Example: Simplified Chinese punctuation
translations:
  direct:
    Iceland: "冰岛"
    Norway: "挪威"
    blue background, red and white cross: "蓝底,红白交叉"
    red background, blue and white cross: "红底,蓝白交叉"
    "{country_1} ({description_1}), {country_2} ({description_2})": "{country_1}({description_1})、{country_2}({description_2})"

Output:

  冰岛(蓝底,红白交叉)、挪威(红底,蓝白交叉)
Example: description before country
translations:
  direct:
    "{country_1} ({description_1}), {country_2} ({description_2})": "{description_1}: {country_1}; {description_2}: {country_2}"

Output:

  blue background, red and white cross: Iceland; red background, blue and white cross: Norway
Example: sentence-style wording
translations:
  direct:
    "{country_1} ({description_1}), {country_2} ({description_2})": "{country_1}: {description_1}; {country_2}: {description_2}."

Output:

  Iceland: blue background, red and white cross; Norway: red background, blue and white cross.
Example: contextual format translation

Use contextual if only one field should use a different format:

translations:
  contextual:
    notes.note:
      faroe-islands.fields.field.flag-similarity.message.format:
        "{country_1} ({description_1}), {country_2} ({description_2})": "{country_1} resembles {country_2}: {description_1} vs {description_2}"

Output for that field only:

  Iceland resembles Norway: blue background, red and white cross vs red background, blue and white cross

Structured media references and integrity

  • Added stable !image references for image-only note fields instead of storing media paths or generated <img> HTML as field content.
  • Hoisted the base media declarations into media.yaml, with stable IDs, safe relative paths, and committed SHA-256 hashes.
  • UG now uses 602 structured image references and 607 hashed media declarations. The remaining five Experimental assets are scripts and styles referenced by card templates, where !image intentionally does not apply.
  • Strict verification checks the real bytes and hashes before export, while clean-tree export stages only the declared assets. It also validates included HTML/CSS and the configured parsed-JSON goldens.
Example: single/multiple images, media declarations, and strict verification

A field can contain one stable image reference or an ordered image sequence, such as UG's blurred and normal Bolivia flags:

# deck.yaml
notes:
  note.bolivia:
    fields:
      field.flag:
        - !image media.ug-flag-bolivia-blur-svg
        - !image media.ug-flag-bolivia-svg
      field.map: !image media.ug-map-bolivia-png

media: !include media.yaml

The stable IDs resolve through the separately maintained declaration map:

# media.yaml
media.ug-flag-bolivia-blur-svg:
  path: ug-flag-bolivia-blur.svg
  sha256: f13669cab4afb991b9851a9c55bb94be5a2c91303a6f8bbb4407a9ffd67951c7
media.ug-flag-bolivia-svg:
  path: ug-flag-bolivia.svg
  sha256: 3010bf58668ac58ae5a1b614867cf94c53b229f2d26679c0ca04cae6d936ced1
media.ug-map-bolivia-png:
  path: ug-map-bolivia.png
  sha256: d461a4cb0d4b845fdc61de1123efca9cc766aaea10dabd95202a1beb980fca7d

This means a file can be renamed by updating its declaration without rewriting every note-field reference. Brain Brew renders the references as safe Anki-compatible <img> tags during export.

brainbrew media hash --manifest brainbrew.yaml --all-targets --media-root media
brainbrew verify --manifest brainbrew.yaml --all-targets --media-root media
brainbrew export crowdanki \
  --manifest brainbrew.yaml \
  --target en-standard \
  --media-root media \
  --out build/crowdanki/en-standard

!image is deliberately limited to whole image-only note fields. Mixed text and images, custom attributes, card templates, styling, scripts, and links remain ordinary HTML/CSS references.

Rust crate release

I had never intend to put Nix/NixOS as a dev dependency, that's just what I happen to use (because it is truly excellent). I simply left it in my test PR rather than put in the effort to actually release anything properly, before it got approved and the effort was worth it 😁

  • Released brainbrew v1.0.0-alpha.3
  • Published brainbrew, brain-brew-core, and brain-brew-formats v1.0.0-alpha.3 to crates.io
  • Added a normal Cargo installation path, so Nix is not required for contributors
  • Kept CI and migration evidence reproducible through the immutable Brain Brew revision 6ee570d427a1a8eec92c22668442f9b7186f9ba7
Example: installation and normal CLI usage

Before

  # Mostly contributor/developer style usage
  nix run . -- --help
  cargo run -- compose --manifest brainbrew.yaml --target de-standard

After

  cargo install brainbrew --version 1.0.0-alpha.3 --locked
  brainbrew compose --manifest brainbrew.yaml --target de-standard

Language-first project metadata

  • Added manifest-level languages
  • Added translation_profile for structural fields, metadata categories, and progress grouping
Example: language-first manifest metadata

Before

targets:
  de-standard:
    overlays:
      - overlay.translation.de
  de-extended:
    overlays:
      - overlay.variant.extended
      - overlay.translation.de

Tools had to infer language/variant meaning from target names.

After

languages:
  en:
    display_name: English
    source: true
    primary_target: standard
    targets:
      experimental: en-experimental
      extended: en-extended
      hardcore-extended: en-hardcore-extended
      hardcore-standard: en-hardcore-standard
      standard: en-standard

  de:
    display_name: German
    translation_overlays:
      base: overlay.translation.de
      hardcore: overlay.translation.hardcore.de
    primary_target: standard
    targets:
      experimental: de-experimental
      extended: de-extended
      hardcore-extended: de-hardcore-extended
      hardcore-standard: de-hardcore-standard
      standard: de-standard

translation_profile:
  structural_fields:
    - field.flag
    - field.map
  metadata_categories:
    - key: deck-metadata
      label: Deck metadata
      paths:
        - deck.name
        - deck.description

This means that all deck extensions can have their own translations, yet still the whole can be understood to be under one language group. This will help tools show all the translations for one language, for all decks/extensions.

Safer extension composition

  • Added sparse extension overlays that can introduce fields, notes, card templates, and media without copying the whole base deck.
  • Added blank-only field_fills, which let Hardcore fill fields owned by UG but fail if another overlay or later base version has already populated them.
  • Destructive changes use explicit replace or override intents with an exact expected base, so upstream drift becomes a conflict instead of being silently overwritten.
Examples: field additions, blank-only fills, and expected-base checks

Experimental adds one field definition and supplies values only where needed:

id: overlay.variant.experimental
kind: extension
field_additions:
  note-type.ultimate-geography:
    fields:
      field.region-code: Region code
    values:
      note.afghanistan:
        field.region-code: AF
      note.albania:
        field.region-code: AL

Hardcore can fill fields that already exist in the shared note model:

id: overlay.extension.hardcore.field-fills
kind: extension
field_fills:
  note.hardcore-bali:
    field.capital: Denpasar
    field.flag:
      - !image media.ug-flag-bali-blur-png
      - !image media.ug-flag-bali-png

A value that intentionally replaces existing source records exactly what it expects to replace:

note_types:
  note-type.ultimate-geography:
    intent: merge
    variables:
      variant.name-suffix:
        intent: replace
        value: " [Extended]"
        expected_base:
          value: ""

Translator workflows

  • Added translation coverage summaries
  • Added translator context views
  • Added interactive apply actions for direct/contextual/no-change/ignore choices
Example: translator CLI workflow

Before

  brainbrew verify --manifest brainbrew.yaml --target da-standard

Then manually inspect YAML failures and decide what was missing/stale/unchanged.

After

  brainbrew translations --manifest brainbrew.yaml --all-targets --summary
  brainbrew translations --manifest brainbrew.yaml --target da-standard --context --status missing
  brainbrew translations --manifest brainbrew.yaml --target da-standard --context --status missing --apply --interactive

Report, summary, and context modes are read-only. Summary mode provides compact per-language and per-overlay counts; context mode shows missing or stale text with its source, target, note, field, card, and duplicate-source context.

--apply edits only the selected translation scope. Non-interactive apply inserts deterministic source-to-source stubs for missing text; interactive apply lets the translator choose direct, contextual, reviewed no_change, ignore, or skip actions.

When English source text changes, an outdated translation can be retained as an explicit stale record rather than being treated as current. Under UG’s lenient policy, Brain Brew warns about that record and continues using its target text until a maintainer resolves it after review; strict translation coverage rejects unresolved stale records. Orphaned dictionary keys, invalid contextual paths or target adaptations, and broken references still fail normal verification.

Deck Workbench

Added a local "Workbench", a webpage one can run which shows the deck contents and allows for editing inline while previewing a Note/Card/Field.

image
Example: Workbench workflow
  brainbrew workbench serve \
    --manifest brainbrew.yaml \
    --enable-write

Then use the local Workbench to review content and edit its underlying source or translation fields:

  - note fields
  - source and translated strings
  - produced cards
  - metadata checklist items
  - comparison languages
  - source edits with translation impact preview

This was a stretch goal I had in mind for a while, but I decided to take a crack at it now. Brain Brew 1.0.0-alpha.3 includes Workbench write support in the normal release. It starts read-only; --enable-write explicitly enables local edits.

The Workbench shows deck content in context, including source and translated text, cards, metadata, and other languages for comparison. Edits remain drafts until someone confirms Apply, which then updates the canonical YAML and owned translation overlays in the local working tree. The write workflow is still being hardened, so it should be used on a version-controlled checkout.

My longer-term goal is to make this a complete translator-facing GUI where people can update translations, create new ones, review the resulting changes, and compare their work with other languages without needing to edit YAML directly.

❗ This tool is very much a work in progress! It still has some sharp edges and strange display stuttering. But these will be fixed! This is just how I imagine a tool could support the workflow much better.

Anyways all of this is up for future consideration/work, but I hope you at least like the direction. I'm open to any and all ideas, it of course matters what people would want to do / how they work want to work. But the more different options the better, in my books!

Discussion Points

Things here are up for discussion. I have taken my own liberties in designing a system I think is good, but am open to being wrong about! Support for the below in Brain Brew does not dictate that UG need use it too. As I said just above more options are better!

Yaml as a format

I hope that the above improvements will help with yaml being the main storage format. Especially when taken as it being the git repo source of truth, not that everyone needs to work via the yaml files (though I personally think it's much nicer to do so now, compared to the csvs before).

I have kept this Brain Brew rewrite working in the same way the old one did: a hub and spoke type of format system. The old system translated everything into "Deck Parts" before then translating to/from CrowdAnki/CSV. This new system has the same general hub-and-spoke shape, except I have made the Canonical Deck YAML itself the stored intermediate representation. That makes it straightforward to add CSV export and import back in, including a workflow where someone exports to CSV, edits it, and imports it again. The CSV adapter is not implemented in alpha.3 yet; it is an option I can add if that is something people want. I still think the yaml is much better than that old system.

English as the Translation field

I truly believe that this is a necessary evil (hopefully less evil now that I've gotten rid of the big long horrible strings). I've also considered using stable note IDs or country codes as translation keys, with English treated as just another translation. Those identifiers solve identity, but by themselves give translators even less context: they still need the source text, note, field, card, and related translations to understand what they are editing. The new context views and Workbench are intended to make the current English-source model much easier to work with.

The workbench editing workflow will help a lot with this problem too.

Hardcore geography

  • Where should it be stored: this repo/another repo?
  • What should it contain?
  • How should it be organised?

All I present for it is that it is doable to have it here in this repo! It was simpler for my demo to have it here too.

The problem is that now if someone imports UG, imports HG, and then updates UG (without updating HG) all of the HG cards of the overlapping notes will disappear.

Oh this was totally an oversight on my part. I have preserved the historical UG note GUIDs rather than allowing them to be regenerated, as well as the 45 meaningful historical Hardcore note GUIDs. The canonical source uses readable stable note IDs for composition, while explicit adapter_ids.crowdanki:guid values preserve the separate identity Anki uses during import and update.

I also reorganised Hardcore so that standalone and companion exports reuse the same 45-note content overlay. The companion keeps its own deck identity, while its note GUIDs, fields, tags, and note model stay aligned with the corresponding standalone deck. The migration evidence collector checks this relationship for all 12 localised Hardcore language pairs.

That gives us an explicit structural check for the identity/model behaviour behind this import-update concern, rather than relying only on generated GUID overrides. I would still want the final workflow exercised in Anki before treating that as a complete end-to-end import guarantee.

@jeprecated

Copy link
Copy Markdown
Member Author

In an attempt to demonstrate the new setup for ordering/formatting the translations I have done a translation audit (in two separate follow-up PRs, so it does not muddy the migration and output-equivalence review). See them below:

This leaves #743 focused on the Brain Brew migration itself, #744 on output-preserving translation structure, and #745 on optional learner-visible translation review.

#744 should help show how we can write the translations but with extra deck specific information to categorise/clarify for the future what the changes are and why.

#745 is just a nice to have, and some may be wrong, and it will surely not get merged in. But! The great thing about this new setup is that if we deem any of the translations to be wrong then we can move them into the stale category, with extra information. This will keep that output generating the same, but allow us to merge in changes that mark specific translations as in need of native attention to categorise/fix. Thus allowing the deck to keep moving ahead without dozens of hanging open PRs for specific translators. If one English string changes then all other translations fields can be moved into stale - which does not claim they are wrong just that they need re-reviewed by a native. Flexible!

@jeprecated

Copy link
Copy Markdown
Member Author

Github support for Stacked Diffs has arrived! 👀 ❗ that will make all of this easier.

I have split up the changes into a stack of PRs on my own UG fork. See the first here: jeprecated#17 @aplaice @axelboc

I believe this is much easier to view the changes in chunks there. Happy to discuss over there or wherever - the specifics of the migration, how UG can/should look and feel (the deck content, the commands, the workflow), and what BB supports 👍

@axelboc

axelboc commented Aug 4, 2026

Copy link
Copy Markdown
Collaborator

I'm going to be blunt: no way I'm merging any of this.

I'm not opposed to AI when used reasonably, like for generating new translation, finding mistranslations, making a new map, or whatever ... but this is wild.

I definitely don't appreciate having to read an AI-generated reply that's hundreds of lines long to find out that it doesn't even answer my questions (at least I don't think it does, can't really tell since it's all gibberish or the agent realising that it did stuff it shouldn't have while still trying to argue that it was for a good reason and it's not going to change it anyway). I guess I should use an AI to summarise that comment for me and generate a response... What's even the point of having a discussion? Let's just start our agents and go on holidays! 🤣

The only way I see us move forward is with a slow step-by-step process where the current setup changes as little as possible each time and where we take the time to discuss each design choice carefully, human to human. The design choices can be suggested by AI, I don't care, but I want a human to be able to explain them to me in their own words.

If that's not something you're willing to do, then I'd prefer to stick with what we have, which works perfectly fine, is still maintainable by humans, and still satisfies hundreds of users worldwide as shown by the overwhelming number of positive feedback we get from the shared deck page to this day https://ankiweb.net/shared/info/2109889812

All that being said, while I still care deeply about this project, I haven't been actively maintaining it for a very long time so if you guys think going full AI is the way forward, I'm not going to get in the way.

@jeprecated

jeprecated commented Aug 4, 2026

Copy link
Copy Markdown
Member Author

I'm going to be blunt: no way I'm merging any of this.

Lmao. I actually agree with your sentiment, believe it or not 😁 That's why it's a draft. This is a very very large change, akin to switching a working tool to another language, and it's very hard to gain the trust that the old system has. I don't expect that'll take 30 mins of a casual review 😅 my best imagination of a path forward would be to run two simultaneous systems: the current/legacy here (in this repo) and the new Rust generated version as a POC (in my fork?). Side be side to show they result in the same outputs, and one can compare the commits as changes happen. There's also the idea of features that the new system has that the old does not (like Federation) but that's a separate thing to test.

I definitely don't appreciate having to read an AI-generated reply that's hundreds of lines long

Most importantly, put upfront: I did not generate my response here (or in the previous thread) with AI! That's why it took so long to do 😅 I wrote the first draft, ran it through an agent for corrections (like if I claimed something was X and it was Y) and made those corrections myself. I then decided some of the points could just be addressed rather than discussing, made more updates to Brain Brew (like adding in the requested !include ability so the yaml could be split up into components). More updates happened after that, some of which I got the agent to generate (like demo yaml examples for the features) but I was human in the loop throughout 😄 Not much I can do to prove that though. Just know that I too hate AI slop! Perhaps even more so, since I'm more exposed to it 😁 If you remove the examples I can wholeheartedly claim 90% of the writing as purely my own. As you should know, I can be rather verbose 😅 🙇‍♂️

to find out that it doesn't even answer my questions (at least I don't think it does, can't really tell since it's all gibberish or the agent realising that it did stuff it shouldn't have while still trying to argue that it was for a good reason and it's not going to change it anyway).

The improvements I made should have addressed all your concerns. Along with my explanation on the Yaml format. If not then I'm happy to elaborate. All I can see here is perhaps the variant vs extension distinction you asked about at the end of your last command (answer: there is no structural difference it's just a naming convention that's 100% up for discussion). But as I said after my intro: "Rather than line by line reply to each comments (as we three are want to do) I think it'd be faster for everyone if I simply address the things that I have improved." 😅 I hope you'll agree these conversations have gone long in the past. Was just trying to head that off. If another medium (a call, an interactive chat, etc) would be better I'm up for it whenever.

I was aware there could be this type of stigma / negative view since I used LLMs to do the update. That's why I tried to head it off with the big "How I developed and validated this migration" section 😅 My intention here is not to turn the repo into an AI generated slopfest, but to craft (with an Agent as my tool) a better Brain Brew that will make things easier and more featureful for maintainers and users. The point in Anki is to help users remember information! That's something we can't export to AI 😁

All that being said, while I still care deeply about this project, I haven't been actively maintaining it for a very long time so if you guys think going full AI is the way forward, I'm not going to get in the way.

It's your deck, sir. I'm just a contributor that like a hard problem to solve 😁 I cannot and will not claim any ownership of UG. I'm here because I have some duty to the system (as a contributor and a user) and the issues of the old Brain Brew have nagged at me for years (because the code and architecture is shit, lol). The idea of getting Federation working has been a treat that I knew was possible but have never had the time to do. Well, thanks to Agentic Development, I found the time! 😁

Totally how I wrote my reply

Reply to Axelboc. Make no mistakes!

@axelboc

axelboc commented Aug 5, 2026

Copy link
Copy Markdown
Collaborator

The improvements I made should have addressed all your concerns.

I still disagree with the proposed translation storage format. One file per language with every English strings duplicated in every language file is not viable in my opinion.

And that's probably just one of the many objections I would have if I could actually review something without browsing through hundreds of unrelated changes like hardcord geo deck extensions, media files being moved, and even changes to the real content of the deck like translations!

my best imagination of a path forward would be to run two simultaneous systems: the current/legacy here (in this repo) and the new Rust generated version as a POC (in my fork?). Side be side to show they result in the same outputs, and one can compare the commits as changes happen.

Seeing that the the output is the same does not give me any confidence in the system that generates that output.

Two core design requirements for me:

  • every part of the deck still has to be editable by humans
  • the entire UG repo still has to be understandable by humans

If the goal of these giant PRs is to brainstorm and show proofs of concept, that's okay. But to actually move forward, I would expect first to focus on improving what we already have (i.e. mostly switching from CSV to YAML, one note per file, etc.) and to have deep discussions about repo structure, file format, translation process, contribution process, command line interface, etc.

Once we're happy with that, and only then, we can move on to federation features. For this, I would expect:

  1. first to see a list of concrete use cases for extending/overriding/customising UG (i.e. what are the actual user needs);
  2. then to prioritise that list and clarify the scope by excluding any far-fetched or too niche/complex features;
  3. then to discuss an implementation roadmap that starts with the first, highest priority use case;
  4. then to see an actual implementation for that first iteration that is as simple and clean as possible;
  5. review and merge that first iteration, demonstrate that it solves a real-world use case in a separate repo like Hardcore Geography, release it and give it to users for a few months and wait for feedback, correct/improve based on feedback;
  6. and only then, move on to the next iteration and next user need and repeat the process.

In other words, an iterative development process that's been proven to lead to good quality software that is maintainable and usable by humans.

@aplaice

aplaice commented Aug 5, 2026

Copy link
Copy Markdown
Collaborator

@axelboc thanks very much for the comment and saying much of what I wanted to say! :D

@jeprecated I appreciate your continued efforts to improve our tooling and workflows!


I think there are three (partially linked) topics here —† the technical design, process and the question of AI/LLM use.

† I've long been a fan of the em-dash and I won't allow LLMs to take it from me...

AI policy

I'd been hoping that we wouldn't need an AI-use policy/guidelines, but it's probably necessary (not even just for this, but in general — I trust you @jeprecated to not do anything crazy, but I don't trust all future contributors). Given that to a large extent it's a matter of preference/ethics I'll open a separate discussion for the AI question.


Process

I haven't given this much thought until now..

In terms of motivation, I understand that the goal of the migration is to rely on a BrainBrew that's easier to extend?

@axelboc's template seems generally sensible, though I'd keep it as a suggestion rather than a prescription. (I also have some technical quibbles, but I'll leave this for the next section.)

Two parallel instances

my best imagination of a path forward would be to run two simultaneous systems: the current/legacy here (in this repo) and the new Rust generated version as a POC (in my fork?). Side be side to show they result in the same outputs, and one can compare the commits as changes happen.

That's an interesting idea, especially since my perception of the new/suggested format is that it's better for reading/reviewing than editing.

However, I'm not sure how it would work (mono- vs. bi-directional?) (as a GH workflow that would automatically update the tree on new commits? (works for the mono-case, but not really the bi-case (due to GH workflow permissions))).

I'm also not sure what would be the end-game. We might end-up deciding that the new format is nice for viewing diffs and a nightmare for editing (or even worse have no feedback on the editing side since everyone will stick to the old format out of familiarity). What then?

I'd also feel uncomfortable since you (@jeprecated) would be putting in effort/tokens into a system that we might end up discarding. Your changes to BrainBrew are independent of AUG (any project could end up using it) and AUG is at least a good test-case, but a parallel repos systems would be bespoke.


Technical

(Subset of potential topics.)

!include

Thanks @jeprecated for the !include directives for CSS/HTML.

Splitting up long keys

I think that the new split-up of long keys (splitting up long English source strings into several shorter ones for translation) is also a considerable net improvement (slight additional loss of context vs. huge readability and reusability gain). However, I don't understand how (on what basis) the splitting is carried out?

Using English as keys

Unfortunately, I'm still very much not convinced that having the English fields as keys for the translations is desirable. Other than the super-long lines issue all my old objections still apply and I can't see any reasonable way of working around them, given the fundamental problem that for country_info and (to a lesser extent) capital_info many of the other-language versions are not actually translations.

Even if we went the radical route of splitting up country_info into two sub-fields (country_info_translatable and country_info_language_specific) we'd still end up with two different approaches to translating items, which IMO would be confusing for translators (and also cumbersome).

YAML vs. CSV in general

Here I think I disagree with @axelboc. I'm not actually convinced that switching to YAML while keeping our per-field grouping (country.yaml, capital.yaml etc.) — i.e. something like:

capital.yaml:

England:
  EN: London
  FR: Londres
  ...

would be an improvement on the whole. IMO it'd not be as clear a downgrade as the per-language YAML structure proposed here, but I'm not 100% sure it'd be an improvement over our current CSVs.

Pros of per-field YAML (over current CSV):

  1. Much shorter lines (and hence much more readable diffs and easier to quickly jump to the right place/tweak).

Cons of per-field YAML (over current CSV):

  1. IMO YAML has marginally more confusing quoting rules than CSV (something like Nested Text would overcome YAML's shortcomings here, at the cost of being less familiar to most programmers).

  2. You can work around CSV's warts by using a spreadsheet editor, which provides structured navigation.

    Hence, CSV is IMO more accessible to non-programming contributors and arguably even for programmers when opening for more than quick tweaks.

(Honestly, I'm not sure. I expect that if we had such a per-field YAML system I'd be about as hesitant about switching to CSVs.)

@axelboc

axelboc commented Aug 5, 2026

Copy link
Copy Markdown
Collaborator

Thanks @aplaice, I agree that an AI policy would be beneficial — I would suggest Numpy's policy as a starting point but that can be discussed in a dedicated issue.

Here I think I disagree with @axelboc. I'm not actually convinced that switching to YAML while keeping our per-field grouping (country.yaml, capital.yaml etc.) — i.e. something like:

I was more thinking of per-note grouping as I initially described here #143 (comment)

Should we reopen this issue and continue the discussion there?

@jeprecated

Copy link
Copy Markdown
Member Author

Alright I have worked some magic, in a way that will hopefully allow us to split up all of this into entirely separate issues. Specifically I have introduced CSV support into the new Brain Brew (in a way that doesn't make me sad) so that it functions merely as a data source for the note contents (including translations). This will allow us (when wanted) to do the move from the legacy/python Brain Brew over to the updated/rust Brain Brew entirely separately from the data model changes.

In fact, I have worked magic to make the csv and yaml live together, so one could move over any specific content from the csv into the yaml, and have it all combine in the end. So the flag similarities csv/column could be deleted and the data placed into the yaml with the nice variables which make things easier to read and translate, resulting in the same output in the end. Or a specific country / note could be moved into yaml, and it would all just work. I think a note could even live in the csv but be translated via the yaml too!

See the PR stack here: jeprecated#28 11 PRs in total, split up to be minimal and easy to review. The changes progressively add in the deck content into the new system, while comparing the output, and in the end removing the old system. Just to be super clear: the PRs are to my own fork, and can be pulled in iteratively over time. No rush, no big deal.

I hope you'll both find this a more acceptable method to move forward 🙏

The rest

One file per language with every English strings duplicated in every language file is not viable in my opinion.

I disagree! As Aplaice said there are pros and cons of each system. But we can get into things like this individually in another change/discussion 👍 rather than bloating this one.

Same with Aplaice's comments on English as keys and Yaml vs csv, lets take that elsewhere / have that conversation another time. With my current changes it's not needed just yet anyways 👍

Seeing that the the output is the same does not give me any confidence in the system that generates that output.

Two core design requirements for me:

every part of the deck still has to be editable by humans
the entire UG repo still has to be understandable by humans

I still feel like it is both these things! But I can see that we can take the process changes more iteratively. I hope my recent changes assuage your concerns here, sir.

AI policy

Sounds sensible!

In terms of motivation, I understand that the goal of the migration is to rely on a BrainBrew that's easier to extend?

Pretty much. The old python Brain Brew is a mess that I do not wish to update or maintain. The design is bad, and I've wanted to do a rewrite for years. This is my rewrite, which is much more maintainable and really better in every way. Axelboc's comments on slowly and iteratively changing things are perfectly fair for the process, but not for the tool itself. There's no world in which a slow change of Brain Brew would serve anyone better, it just needed to be torn out and replaced.

That's an interesting idea, especially since my perception of the new/suggested format is that it's better for reading/reviewing than editing.

However, I'm not sure how it would work (mono- vs. bi-directional?) (as a GH workflow that would automatically update the tree on new commits? (works for the mono-case, but not really the bi-case (due to GH workflow permissions))).

I'm also not sure what would be the end-game. We might end-up deciding that the new format is nice for viewing diffs and a nightmare for editing (or even worse have no feedback on the editing side since everyone will stick to the old format out of familiarity). What then?

I'd also feel uncomfortable since you (@jeprecated) would be putting in effort/tokens into a system that we might end up discarding. Your changes to BrainBrew are independent of AUG (any project could end up using it) and AUG is at least a good test-case, but a parallel repos systems would be bespoke.

Yea I meant that I'd 'manually' mirror any updates over to my fork. It would not take much effort at all. The goal would be to use the tool naturally and find the pitfalls. Having other maintainers also try and duplicate the changes they make there would be a nice experiment too, if they (or you) are willing. It would just be a way to try the new system without... committing to anything (not sorry for the pun).

I think that the new split-up of long keys (splitting up long English source strings into several shorter ones for translation) is also a considerable net improvement (slight additional loss of context vs. huge readability and reusability gain). However, I don't understand how (on what basis) the splitting is carried out?

Splitting? This was a manual process, one that Brain Brew does not do. Instead of the field being the manual string "Moldova (wider, coat of arms with eagle)" the data itself was changed by me/agents into:

      field.flag-similarity:
        - country: note.moldova
          description: wider, coat of arms with eagle

which gets stitched together via the message pattern in the field definition in the note-types.yaml:

    field.flag-similarity:
      name: Flag similarity
      message_pattern:
        kind: list
        item_format: '{country} ({description})'      <--- The stitching is just the variables below
        separator: ', '
        parameters:
          country:
            type: note_field_ref            <--- 'country' above is a note reference. If the name changes the id will still work
            field: field.country              <--- Which then resolves to a specific field of that note
          description:
            type: text

The above example only has one entry, but if the list had multiple similarities then it would resolve each of them and join the results together with the separator. Simple but powerful stuff! This is all defined once in the note-types.yaml, and the maintainer/translator just sees the info that needs to be put in

† I've long been a fan of the em-dash and I won't allow LLMs to take it from me...

Aplaice is AI, confirmed 😆

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

3 participants