Document ID: PRD-ASE-001
Version: 0.7 — aligned to architecture.md v0.3 (architecture is the source of truth)
Date: 2026-06-13
Author: Business Analyst (Lead)
Status: Open for stakeholder review
- Stakeholders
- User Journeys and Workflows
- Functional Requirements
- Non-Functional Requirements
- Adaptivity Signal → Decision Matrix
- Assumptions, Constraints, and Dependencies
- Open Questions and Requirement Gaps
v0.2 change note (2026-06-12): Added FR-NLT (Natural-Language Tuning) requirements group (FR-NLT-01 through FR-NLT-11), Conversational Tuning phrase-mapping table, Journey 2.7, two new rows in the Adaptivity Signal Decision Matrix, and five new Open Questions (OQ-11 through OQ-15). Architecture for text interpretation is explicitly deferred — see OQ-11.
v0.3 change note (2026-06-12): Folded in two founder decisions (Decisions A and B). Added LD-8 (unified DSP action-space model + governing principle + phase). Reframed FR-NLT-02 output target as an explicit DSP action vector. Reframed FR-NLT-10 as a point on the directness spectrum (band approximation is the shippable baseline). Added FR-NLT-12 (abstract/aesthetic descriptors as first-class inputs). Expanded §3.9.1 phrase-mapping table with aesthetic/emotional descriptor rows and added a unified action-space lead-in paragraph. Updated Adaptivity Signal Decision Matrix NLT rows to governing-principle framing. Resolved OQ-12 (band approximation baseline; source separation optional later) and OQ-15a (governing-principle takes precedence); recorded recommended defaults for OQ-15b and OQ-15c as pending confirmation. OQ-11 remains deferred.
v0.4 change note (2026-06-12 — prior-art refinement pass): Folded in findings from
docs/session-notes/prior-art.md. (A) FR-SYS group: added FR-SYS-07 (process-tap primary path) and FR-SYS-08 (TCC permission for tap); reframed FR-SYS-01..06 as the FALLBACK (driver) path; updated Journey 2.6 to tap-primary + driver-fallback flows; updated NFR-INSTALL-01/02/03/04 and CON-05/07 for tap vs. driver paths; added CON-10 (tap requires macOS 14.2+/14.4+ — verify); updated OQ-07 note. (B) FR-SPAT-01/02: clarified HRTF rendering as custom SOFA-HRIR partitioned convolution (libmysofa + FFTConvolver); noted Apple PHASE/AVAudioEnvironmentNode HRTFs are non-replaceable; updated DEP-06 and ASM-04. (C) FR-ADAPT / LD-5 note: added clarification that real-time ML inference uses BNNS Graph (RT-safe); Core ML / SoundAnalysis are off-RT pre-analysis only. (D) FR-TONAL-04: added mono-summed low-band constraint (patent avoidance — Waves US-11,102,577 active); added CON-11 (patent constraint); added OQ-16 (IP review spike). (E) DEP and CON: replaced vague data deps with concrete licensed picks (SADIE II Apache-2.0, libmysofa BSD-3, libebur128 MIT, AutoEq MIT, FFTConvolver MIT, libASPL MIT, Demucs+MLX MIT); added CON-12 (permissive-only shipping rule); added OQ-17 (libbs2b license dispute). Resolved OQ-04 (SADIE II default, IRCAM avoided) and OQ-08 (AutoEq MIT computed curves). OQ-11 (Conversational Tuning architecture) remains deferred.
v0.5 change note (2026-06-13 — architecture v0.2 alignment): Major revamp to align with
docs/architecture/architecture.mdv0.2 (now the source of truth). Adopted the four-phase scheme (0 / 1 / 1.5 / 2). Added LD-11…LD-17; updated LD-1/5/7/8. Added FR-STEM-* (stem-based object engine, Phase 1.5) and FR-REIMAGINE-* (intensity control). Reframed FR-SPAT around BRIR-first immersion; FR-TONAL for minimum-phase-default + no-program-DRC + loudness-comp method; FR-ADAPT for ERB/Bark perceptual + masking decisions + a "processing must not be perceived as moving" criterion; FR-NLT as a typed multi-band macro with per-stem targeting. Added a performance/feasibility-budget NFR (6-stem × per-stem chain × BRIR convolution) and set the bit-transparent bypass = Reimagine intensity 0. Tightened privacy; updated CON/ASM/DEP, the decision matrix, journeys, and open questions.
v0.6 — architecture v0.3 sync (authoritative; amends the FRs/NFRs it names). Folds in the expert-panel review (review-v0.2.md) + founder hardware/persona/power/app-shape decisions, now canonical in architecture.md v0.3 §0/§18:
- Re-sum mixbus (ADR-011) — amends FR-STEM-02: per-stem chains sum through a managed mixbus — per-stem makeup ≤ gain-reduction removed, loudness-matched per-stem trim to the intensity-0 reference, headroom budget pre-limiter, metered limiter GR, group-delay-aligned center.
- Spatial exemptions — amends FR-SPAT-01/06, FR-STEM-02/03: bass (≲120 Hz) high-passed out of the BRIR path + summed mono; lead vocal kept centered (no L/R spread) at every intensity.
- Shared late-reverb / content-adaptive room — amends FR-SPAT-01/05, NFR-PERF-06: one shared late-reverb tail + cheap per-stem early/direct filters (not 6 independent BRIRs); BRIR room amount adapts to the source's existing reverberation.
- Masking — amends FR-ADAPT, OQ-22: clarity/between-stem decisions use the excitation-pattern / masked-threshold (ERB) subset, not full Moore-Glasberg partial loudness (~50× too slow).
- Stem gating — amends FR-STEM-05: gate on a perceptual-artifact estimate, not SDR; confidence clamps the per-track Reimagine ceiling.
- Separation models — amends FR-STEM-01, DEP: MLX primary (Core ML secondary); code MIT, model weights auto-downloaded on first run (NC-trained → not redistributed); cached stems FLAC + bounded LRU.
- Reimagine defaults — amends FR-REIMAGINE-03/04: default low-to-lower-mid; dead-band above 0% (no crossfade of bit-perfect vs imperfect-phase stems); loudness-matched across the knob.
- NL — amends FR-NLT-02: planned primary = on-device LLM + SAFE/SocialEQ priors (CLAP demoted to reranker; rules floor; cloud opt-in); mechanism still deferred (OQ-11); interpreter output is untrusted → schema-validated + numeric-clamped to governing-principle + hearing-safety limits;
contextfield-allowlist excludes audio, hearing data, and track identity.- RT ML — amends LD-5/FR-ADAPT: ADR-004 (BNNS-Graph RT ML) is contingent — no RT ML is currently needed.
- Tap consent — amends FR-SYS-07/08, NFR-PRIV: the muted global tap is a high-consent, captures-everything capability (TCC + purple indicator; all apps incl. calls) — explicit consent UX, auto-exclude communication apps, tapped audio never persisted and never fed to stem separation.
- Hardware (LD-18) / power / app-shape (LD-19) / persona: floor M1 Pro/16 GB (M4/M5 far above); foreground sole-occupancy; max-quality on AC, lighter on battery; full-window listening app + menu-bar extra; primary persona Ramith (developer-audiophile). Risk R-3 → Low; perf spike is tuning, not a gate.
v0.7 (2026-07-12): removed natural-language tuning (FR-NLT, LD-8) and hearing personalization (FR-HEAR, FR-ADAPT-06) from scope — founder decision. Both features are withdrawn (neither was built) and may be re-added later. Also withdrawn as wholly-dependent on those features: FR-STEM-04 (per-stem NL targeting), OQ-05/OQ-11/OQ-12/OQ-13/OQ-14/OQ-15, ASM-09, Journeys 2.2 and 2.7, and the associated §5 matrix / stakeholder / glossary / registry rows. Prior specs remain in git history. Historical change notes (v0.2–v0.6) are left intact as a changelog and still name the withdrawn IDs.
Founder-confirmed decisions. These supersede any conflicting requirement text below; affected requirements have been annotated. Remaining open items are in §7.
Phase scheme: canonical phasing is Phase 0 (player MVP) · Phase 1 (mix-based core: clarity / correction / loudness-comp / adaptive / BRIR / NL + Reimagine mix-range) · Phase 1.5 (stem-based object engine) · Phase 2 (system-wide via process taps) — per architecture.md §16. Where an older FR body still tags "Phase 1" for the own-player, read it as Phase 0–1.
| # | Decision | Resolution |
|---|---|---|
| LD-1 | Scope / phasing | Own-player first (Phase 0 MVP → Phase 1 mix core → Phase 1.5 stem engine), then system-wide via process tap (Phase 2; virtual-device fallback). |
| LD-2 | "Immersive" | Both spatial and tonal/dynamic, equally weighted. |
| LD-3 | Output targets | Both headphones/AirPods and speakers; auto-detect + switch profiles. |
| LD-4 | MVP source | Local files only in Phase 1 (resolves OQ-06). Streaming enhancement deferred to Phase 2. |
| LD-5 | Content classifier | Phased (resolves OQ-09): DSP heuristics first; Core ML genre/mood model (off-RT; trained via Create ML) layered in during Phase 1. Real-time ML inference uses BNNS Graph only. |
| LD-6 | Ambient mic sensing | On-demand sampling only — no continuous/always-on mic. See revised FR-ADAPT-04 and Journey 2.5. |
| LD-7 | Spatial | BRIR-first (superseded by LD-14): default binaural = BRIR (HRTF + early reflections + late reverb); dry HRTF = minimal mode; SADIE II (Apache-2.0) is the anechoic HRIR core. Custom HRTF measurement deferred. |
| LD-8 | Withdrawn 2026-07-12 | Natural-language / conversational tuning removed from scope (founder decision); may be re-added. |
| LD-9 | Project model | Personal / open-source, non-commercial. No monetization, pricing, paywall, or feature-gating of any kind. All features are free. There are no paid tiers, no entitlement checks for feature access, and no conversion-oriented analytics. The specific OSS license is deferred to post-MVP (post-Phase 0). This decision supersedes any conflicting text in this document and resolves OQ-02. |
| LD-10 | Quality-first / ample use of modern hardware | Maximize quality by making ample use of all modern hardware: RAM, CPU, multi-core + GPU (Metal) + Neural Engine parallelism, fast SSD (disk caching/precompute), and fast networks (optional cloud assist for non-sensitive, latency-tolerant work). Prefer platform-native, hardware-accelerated multimedia frameworks/OS features over generic code (macOS-only project): Accelerate/vDSP/BNNS, Core ML (Neural Engine), Metal/MPS, Audio Workgroups (os_workgroup) for safe real-time parallelism, AVAudioEngine/AudioToolbox built-in units, hardware-accelerated decode + AVAudioConverter SRC, and Spatial Audio / head-tracking APIs. CPU/RAM/disk are not primary constraints. Hard limits that remain: (a) the real-time per-buffer deadline (NFR-PERF-01); (b) core playback stays offline-capable — network optional, never required; (c) privacy — sensitive data (mic) stays on-device; (d) laptop battery/thermal (optional efficiency mode). Own-player latency is free, so look-ahead/pre-analysis, linear-phase FIR EQ, oversampling, and long convolutions are all in scope. Default to the max-quality profile. Supersedes the former fixed CPU/RAM caps — see revised NFR-PERF-02/03/04. |
| LD-11 | Source quality & non-goals | Assume good-quality sources (lossless / high-bitrate). Audio repair/restoration is a non-goal (no de-noise/de-clip/upsample to "fix" bad audio). Network may be used for non-sensitive, latency-tolerant work; core playback + RT DSP stay offline-capable. |
| LD-12 | Perceptual tonal model | Clarity/adaptive decisions are made in ERB/Bark with a masking + partial-loudness model (Moore-Glasberg style), not raw dB-on-log. Contributors are typed (EQ-curve + per-band dynamic + transient + spatial); the dB curve is a realization/interchange format only. |
| LD-13 | Phase realization | Minimum-phase by default; phase mode chosen by content (transient density from pre-analysis); linear/mixed-phase opt-in or band-limited where it genuinely helps (pre-ringing, not latency, is the real cost). |
| LD-14 | BRIR-first immersion | Headphone spatialization defaults to a binaural room response (HRTF + early reflections + late reverb); dry HRTF = minimal mode; head-tracking opt-in for music. Speakers = M/S width + ambience extraction (mono-safe); crosstalk-cancellation opt-in (centered near-field only); crossfeed opt-in. |
| LD-15 | Stem-based object engine (Phase 1.5) | Offline 6-stem separation (vocals/drums/bass/guitar/piano/other), cached to SSD; full per-stem chains incl. spatial placement, re-summed to binaural; masking computed between stems. Own-player-only — the live tap path (Phase 2) is mix-level only; real-time-lite separation is a research track. |
| LD-16 | "Reimagine" intensity knob | One continuous control: 0% = original mix, stem engine bypassed (bit-faithful, zero separation artifacts) → clarity → spatial widening → 100% = full stem-based spatial reimagining. Crossfades original↔stem-render + scales spatial spread / unmask depth. Mix-range in Phase 1; stem-range unlocked in Phase 1.5. |
| LD-17 | Dynamics & loudness | No program DRC by default (transparent LUFS normalization + true-peak safety limiter only). Loudness compensation = fraction of the equal-loudness contour difference (ISO 226) + per-device SPL calibration + loudness-matched makeup, rate-limited to volume changes only. |
| LD-18 | Target hardware & runtime posture | Floor = Apple-Silicon Pro-class (M1 Pro, 2021) / ≥16 GB; shipping generation far above (M4 38-TOPS NE; M4 Pro/Max 10–12 P-cores, 273–546 GB/s, 64–128 GB; M5 per-GPU-core neural accelerators ~4× M4 GPU-AI). App is foreground/sole-occupancy ("lean-back listening") and may use many cores + occupy memory generously. Large headroom on current hardware; design for the floor, exploit the abundance. Supersedes the base-8 GB-Air framing; downgrades the stem-render risk to Low — see revised NFR-PERF-06. Power: default max-quality on AC; auto-lighter (Efficiency profile) on battery, user-overridable. |
| LD-19 | App shape | Full-window lean-back listening experience (now-playing + visualizer + Reimagine dial) plus a menu-bar extra for quick control; both share one app-level engine. Refines FR-UI. |
| ID | Stakeholder | Role | Primary Need | Influence | Interest |
|---|---|---|---|---|---|
| STK-01 | Audiophile Music Listener | End User | Immersive, personalized listening without manual EQ fiddling | Low | High |
| STK-02 | Casual Music Listener | End User | Better sound out-of-box on AirPods/laptop speakers with zero config | Low | High |
| STK-03 | Founder / Product Owner | Decision-maker | Technical feasibility, personal project goals, OSS community adoption | High | High |
| STK-04 | macOS Platform / Apple | Platform Constraint | App Store / notarization compliance, privacy rules, entitlement grants | High | Low |
| ID | Stakeholder | Role | Primary Need |
|---|---|---|---|
| STK-06 | Developers / Engineering Team | Implementer | Clear, testable specifications; real-time safety constraints documented |
| STK-07 | Beta / QA Testers | Validator | Reproducible acceptance criteria per requirement |
| STK-08 | Streaming Service Providers (Spotify, Apple, YouTube) | Indirect | No TOS violation from system-wide processing (Phase 2) |
- STK-01 and STK-02: Sound should be noticeably better immediately after install; personalization should be opt-in and progressive, not a barrier.
- STK-03: Phase 1 (own player) must ship fast enough to validate the adaptivity engine before Phase 2 (virtual device) effort. As a personal/open-source project the goal is working software and community value, not commercial revenue.
- STK-04: All entitlements, privacy strings, and signing/notarization must be correct before any public release.
Note: User journeys are documented in detail in
docs/product/user-journeys.md. That document is the authoritative reference; the journeys below are extracted here for requirements traceability. Several of these journeys describe not-yet-built behavior (onboarding, the automatic adaptivity engine, environment/mic sensing, Phase-2 system-wide capture); see the per-journey Status lines inuser-journeys.mdand the "Implementation status" note at the head of §3 below for what the current build actually does.
Full step-by-step flows, per-journey Phase and Status (built / partially built / planned) lines, and success conditions live in user-journeys.md — the authoritative source. Summarised here for requirements traceability only:
| Journey | Actor | Goal | Steps |
|---|---|---|---|
| 2.1 — First-Run Onboarding | New user, first launch | Reach a working listening state with a meaningful default profile in under 3 min | → user-journeys.md |
| 2.2 | — | Withdrawn 2026-07-12 (hearing personalization removed from scope; may be re-added) | — |
| 2.3 — Normal Listening Session | Returning user with a configured profile | Listen to a playlist with seamless adaptive enhancement | → user-journeys.md |
| 2.4 — Switching Output Device Mid-Session | User who unplugs AirPods / switches to laptop speakers | Sound continues uninterrupted; DSP profile switches automatically | → user-journeys.md |
| 2.5 — Environment Change (Room Gets Noisy) | User in a room that has become noisier | One-shot on-demand mic sample adapts DSP, then releases the mic (LD-6) | → user-journeys.md |
| 2.6 — Phase 2: System-Wide Enhancement | User enabling system-wide enhancement | Route all system audio through the DSP engine (tap primary / driver fallback) | → user-journeys.md |
| 2.7 | — | Withdrawn 2026-07-12 (natural-language / conversational tuning removed from scope; may be re-added) | — |
Requirements are grouped by capability area. Each requirement carries a stable ID, a description, and 1–3 Given/When/Then acceptance criteria.
Priority Notation: P1 = Must Have (MVP), P2 = Should Have (v1.1), P3 = Nice to Have (backlog)
Implementation status (as-built vs. specified) — audit snapshot. This PRD specifies target behavior: "shall" is a requirement, not a claim that the behavior exists today. Verified against
Sources/onmainas of this audit, the current build is a Phase-0 own-player implementing only a subset of what follows. The unbuilt requirements remain valid and STAY, but are planned, not current — confirm against the code before assuming any behavior is live. For what is actually built vs. planned, seesprint-plan.md§Status + the source — this section does not enumerate it.
FR-PLAY-01 — Local File Playback (P1)
The app shall play local audio files in common formats (MP3, AAC, FLAC, ALAC, WAV, AIFF, OGG) routed through the DSP engine.
Given a local audio file in a supported format is added to the queue,
When the user presses Play,
Then audio begins within 500 ms and is routed through the full DSP chain with no audible artifacts.
FR-PLAY-02 — Playback Controls (P1)
The app shall provide standard playback controls: play, pause, skip forward, skip backward, seek, and volume adjustment.
Given audio is playing,
When the user activates any playback control,
Then the corresponding action is executed within 100 ms of user input with no audio glitch.
FR-PLAY-03 — Queue Management (P1)
The app shall support a playback queue with drag-to-reorder, remove, and repeat/shuffle modes.
Given a queue of at least 5 tracks,
When the user reorders a track via drag-and-drop,
Then the queue reflects the new order immediately and the next track plays in the updated order.
FR-PLAY-04 — Metadata Display (P1)
The app shall display track title, artist, album, album art, and duration for local files using embedded metadata (ID3, Vorbis Comment, MP4 atom).
Given a file with complete ID3v2 tags is playing,
When the Now Playing screen is visible,
Then all available metadata fields are populated correctly.
FR-PLAY-05 — Session State Persistence (P1)
The app shall restore the last queue, playback position (if paused), and active profile on next launch.
Given the app is closed while paused at position 2:34 in a track,
When the app is relaunched,
Then the queue is restored, the track is pre-loaded at 2:34, and the same profile is active.
FR-PLAY-06 — Supported Format Extensibility (P2)
The app should detect and surface unsupported file formats with a clear error rather than silently failing.
Given a file in an unsupported format is dragged into the app,
When the import is attempted,
Then an inline error message names the format and links to the list of supported formats.
FR-SPAT-01 — BRIR Binaural Rendering for Headphones (P1) (revised per LD-14: BRIR-first)
When a headphone output is detected, the app shall apply binaural room-response (BRIR) rendering — a SOFA HRIR convolved with early reflections + late reverberation that carry interaural differences — to produce an externalised, out-of-head soundstage. Dry (anechoic) HRTF is the minimal mode, not the default, because dry HRTF alone reliably collapses in-head with non-individualised data. Implemented as our own partitioned convolution using libmysofa (BSD-3) + vDSP / FFTConvolver (MIT); anechoic HRIR core = SADIE II (Apache-2.0); the room layer is synthesised (image-source + FDN) or a CC0/CC-BY BRIR. Headphone-correction EQ (FR-TONAL-02) corrects timbre only and is not relied on for externalisation. Apple PHASE/AVAudioEnvironmentNode/AUSpatialMixer are not used (fixed, non-replaceable HRTFs).
Given a stereo track is playing and a headphone device is active,
When HRTF mode is enabled (default for headphones),
Then the audio output includes externalised binaural cues produced by the BRIR convolution engine such that a naive listener ABX test shows audible spatial difference (and improved externalisation vs. the dry-HRTF minimal mode) at a statistically significant rate.
ABX Test Protocol (Two-Stage Validation):
Stage 1 — Sanity Gate (early Phase 1): 2 listeners (yourself + 1 other), 10 ABX trials each. Test material: 2–3 representative tracks (e.g., orchestral for spatial cues, vocal for center-image stability). Criterion: PASS if both listeners identify correct (BRIR) in ≥6/10 trials (60% threshold). Purpose: quick validation that BRIR implementation is working before full testing.
Stage 2 — Full Validation (Phase 1 mid-sprint): 5 listeners (including yourself), 10 ABX trials each (50 trials total). Test material: 4 diverse tracks (orchestral, vocal, acoustic, ensemble). Criterion: PASS if ≥60% overall correct across all 50 trials (≥30/50). All trials double-blind (tester does not know which is A vs. B); listener explicitly rates externalisation quality on each trial (5-point scale: "in-head" to "externalised outside").
Pass overall: Both Stage 1 and Stage 2 must pass. If Stage 1 fails, investigate BRIR implementation before proceeding to Stage 2. If Stage 2 fails, evaluate BRIR quality vs. competing algorithms (e.g., dry HRTF + room synthesis alternatives) and iterate.
FR-SPAT-02 — HRTF Profile Selection (P2)
The app shall provide at least 3 selectable HRTF profiles drawn from the SOFA dataset library (e.g., SADIE II subjects offering generic, small-head, and large-head approximations). Profile selection loads a different SOFA HRIR set; it does not switch to a platform-provided spatialization API.
Given the HRTF profile selector is opened,
When the user selects a different profile,
Then the active HRTF (loaded from the SOFA dataset) renders immediately on selection without requiring restart.
FR-SPAT-03 — Crossfeed for Headphones (P3, opt-in / off by default) (revised per LD-14)
The app shall offer adjustable crossfeed to reduce unnatural extreme stereo panning on headphones. Crossfeed is opt-in and off by default — it is largely subsumed by the BRIR path (a BRIR is a physically-correct crossfeed-plus-room). The Bauer-style algorithm is reimplemented in-house pending the libbs2b license check (OQ-17).
Given a stereo track with hard-panned elements is playing on headphones,
When crossfeed is enabled at the default level (Bauer stereophonic-to-binaural, ~700 Hz crossover),
Then channel separation is measurably reduced (verifiable via FFT analysis) without perceived image collapse.
FR-SPAT-04 — Head-Tracked Soundstage via AirPods Motion (P2, opt-in) (revised per LD-14)
When AirPods (3rd gen or later, AirPods Pro 1/2, AirPods Max) are the active output and the user opts in, the app shall use CMHeadphoneMotionManager (macOS 14+) head-tracking to stabilise the soundstage relative to a fixed world reference. Head-tracking is off by default for music (many listeners prefer a head-locked music stage); its externalisation benefit is largely already delivered by the BRIR path.
Given AirPods Pro are active, head-tracking is enabled, and the user rotates their head 30 degrees,
When head-tracking data is received (via CoreMotion / AirPods motion API),
Then the soundstage rotation counter-compensates such that the perceived source direction does not move with the head, with lag < 20 ms.
FR-SPAT-05 — Virtual Room / BRIR Layer (P1) (promoted per LD-14 — the room component of FR-SPAT-01)
The app shall provide the room layer (early reflections + late reverberation carrying interaural difference) that, combined with the SADIE-II HRIR, forms the BRIR of FR-SPAT-01. At least one default "treated listening room" plus alternates (e.g., studio, living room, hall) shall be selectable. The room may be synthesised (image-source + FDN) or loaded from CC0/CC-BY BRIR/IR data.
Given a room IR is selected from the library,
When convolution is enabled,
Then the output contains reverb characteristics consistent with that IR (measurable RT60), and CPU usage remains within the NFR-PERF-01 budget.
FR-SPAT-06 — Speaker Immersion: M/S Width + Ambience (P1) (revised per LD-14)
When a speaker output is detected, BRIR/HRTF rendering shall be disabled and a mid-side width + ambience-extraction mode shall activate, with a hard mono-compatibility constraint (the M channel is preserved). Crosstalk-cancellation/transaural is an opt-in "centered near-field" mode only (stereo-dipole narrow span); aggressive XTC shall not be applied blindly on laptop speakers.
Given the active output device is identified as speakers (not headphones),
When a stereo track plays,
Then HRTF is inactive, stereo width processing is applied, and the mid/side balance is adjustable by the user.
FR-SPAT-07 — Spatial Mode Auto-Detection and Override (P1)
The app shall automatically choose spatial processing mode based on device type but allow the user to manually override.
Given the active device switches from headphones to speakers,
When the device change event fires,
Then spatialization mode updates automatically within 500 ms and a banner confirms the change; the user may override via the controls panel.
FR-TONAL-01 — Parametric EQ Engine (P1) (revised per LD-12/LD-13)
The app shall provide a parametric EQ (≥10 bands; adjustable frequency, gain ±20 dB, Q). The canonical tonal target is a composable curve realized off-RT as minimum-phase biquads by default; linear/mixed-phase FIR is opt-in or selected by content (transient-dense material stays minimum-phase to avoid pre-ringing — LD-13). The RT kernel runs finished coefficients only (no design/fitting on the audio thread).
⚠ Deviation (as-built): the shipped EQ is a 31-band graphic EQ (fixed bands, ±12 dB), not the parametric ≥10-band / ±20 dB / adjustable-Q engine specified here. The parametric spec stays as the target; the graphic build meets the live / accurate / resettable intent. Reconcile graphic-vs-parametric — see backlog US-TON-01.
Given the user sets band 3 to 200 Hz, +6 dB, Q=1.0,
When audio plays,
Then FFT analysis of the output shows a +6 dB peak centred at 200 Hz within ±0.5 dB tolerance.
FR-TONAL-02 — Headphone/Speaker Frequency Response Correction (P1)
The app shall apply a device-specific correction EQ curve that compensates for the measured frequency response deviation of common headphone and speaker models from a target diffuse-field or free-field response.
Given the active device is identified as "Apple AirPods Pro 2" from the device correction library,
When correction EQ is enabled,
Then the device correction profile is loaded and applied, and the user can toggle it on/off with an audible difference.
FR-TONAL-03 — Loudness-Compensated EQ (P1) (revised per LD-17; Phase 0 tuning gate)
The app shall apply loudness compensation as a fraction of the equal-loudness contour difference (ISO 226) between a program reference level (default 83 dB SPL) and the actual playback level — not a raw single-contour boost. It requires a per-device SPL calibration (the app cannot know absolute SPL otherwise), applies loudness-matched makeup gain, is rate-limited to volume changes (never program dynamics), implements hard gain caps, and is defeatable.
Provisional implementation (Phase 0): Apply 50% of the ISO 226 contour difference uniformly across the spectrum as the baseline. Internally, implement three separately tunable fractions (bass_fraction, mid_fraction, hi_fraction), all initialized to 0.5. Hard gain caps: +6 dB maximum in the bass (≤200 Hz), +4 dB maximum in the treble (≥6 kHz). Phase 0 testing will determine whether frequency-variable fractions improve perceived quality and whether the 50% fraction is correct (candidate range: 30–75%; precedent suggests 40–60% is optimal).
Given a track is playing, equal-loudness compensation is enabled, and the user reduces volume by 20 dB,
When the volume change is detected,
Then bass (80–200 Hz) gain increases by approximately 4–6 dB (50% of ISO 226 contour difference, hard-capped at +6 dB) and high-frequency (8–12 kHz) gain increases by approximately 1–3 dB (50%, hard-capped at +4 dB), applied within one DSP processing block. The compensation is audibly noticeable but not artificial or "pumpy."
FR-TONAL-04 — Psychoacoustic Bass Enhancement (P1)
The app shall apply harmonic excitation to bass frequencies to enhance perceived bass weight on small speakers and headphones that cannot reproduce low fundamentals. Bass harmonics shall be generated from a mono-summed (L+R) low band — per-channel (stereo) harmonic generation is explicitly prohibited to avoid infringement of Waves patent US-11,102,577 (active, ~2038; see CON-11 and OQ-16). The implementation shall target the expired MaxxBass approach (US-5,930,373, ~2019 — verify on USPTO before shipping) or a clean-room nonlinear-distortion (NLD) design from the mono low band. Formal IP review is required before any public release (OQ-16).
Given the active output is built-in MacBook speakers and a bass-heavy track is playing,
When psychoacoustic bass enhancement is enabled,
Then harmonic partials of sub-bass frequencies (below 80 Hz) are generated from the mono-summed low band and are audible, without physical speaker over-excursion (no distortion artefacts perceptible on casual listen).
Given the implementation under review,
When the harmonic generation path is inspected,
Then bass harmonics are derived from a mono (L+R) low-band signal, not from per-channel stereo signals independently.
FR-TONAL-05 — Dynamics Policy: No Program DRC by Default (P1) (revised per LD-17)
For good-quality sources aimed at fidelity, the app shall not apply program dynamic-range compression by default. The default dynamics chain is transparent LUFS normalization + the true-peak safety limiter (FR-TONAL-07) only. Any dynamics adaptation (e.g., raising intelligibility in a loud ambient environment) is opt-in and conservative, and shall prefer dynamic EQ over broadband multiband compression.
Given a classical track (high dynamic range) is playing and ambient noise is Quiet,
When the adaptivity engine classifies the content,
Then compression ratio is set conservatively (< 1.5:1) to preserve dynamics; switching to a loud ambient condition increases the ratio (target ≥ 3:1) to improve intelligibility.
FR-TONAL-06 — Tonal Preset Library (P2)
The app shall include a preset library with at least 8 named tonal presets (e.g., Neutral, Warm, Bright, Bass Boost, Podcast, Film, Classical, Electronic) that the user can apply, modify, and save.
Given a preset named "Electronic" is selected,
When applied,
Then EQ parameters change to preset values within one render cycle, and a "Save as Custom" button appears if the user subsequently adjusts any parameter.
FR-TONAL-07 — True Peak Limiting (P1)
A transparent true-peak limiter shall be the final DSP stage, preventing output from exceeding -1 dBTP regardless of upstream gain.
Given upstream DSP adds +6 dB gain to a 0 dBFS signal,
When the limiter is in the chain,
Then the output true peak is ≤ -1 dBTP (verified by a reference meter).
ML path constraint (prior-art finding, ADR-004 Proposed — see
docs/session-notes/prior-art.md): Any ML inference that occurs inside the real-time render callback (on the audio thread) shall use BNNS Graph exclusively — it is RT-safe (no runtime allocation, single-threaded, no locks). Core ML, SoundAnalysis, and Metal/MPS are off-RT only: they may be used freely for pre-analysis, background classification, and model training, but never inside the render block. This applies to FR-ADAPT-01 (content/genre classification), LD-5 (Core ML genre model), and any future on-device ML inference. This constraint refines but does not replace LD-5.
Perceptual-domain decisions (LD-12) & imperceptible adaptation: Adaptive and clarity decisions (masking relief, content/loudness/ambient moves) shall be computed in the ERB/Bark domain against a masking + partial-loudness model (Moore-Glasberg style), not raw dB-on-log. Adaptation shall be conservative and imperceptible-as-motion: coalesced updates, slow ramps (≥50 ms, FR-ADAPT-03), hysteresis/deadbands, and no move that fights intentional musical contrast. Acceptance: in listening tests, users shall not be able to identify that "the EQ is moving" — only a net improvement (a Phase-0 KPI gate; see architecture.md §10).
FR-ADAPT-01 — Content / Genre Classification (P1)
The app shall analyse the spectral and rhythmic characteristics of the currently playing audio on a non-real-time thread and derive a content classification (minimum: speech, classical, electronic/bass-heavy, acoustic/folk, rock/metal, other).
Given a track with dominant bass frequencies and a 4/4 electronic drum pattern plays for 5 seconds,
When the content analyser completes its window,
Then the classification resolves to "Electronic" and the adaptivity engine's genre-tuned curve is applied.
FR-ADAPT-02 — Real-Time Parameter Update (Lock-Free) (P1)
All parameter changes from the adaptivity engine to the DSP audio thread shall be communicated exclusively via lock-free mechanisms (std::atomic parameters or a single-producer, single-consumer ring buffer). No mutex or blocking call shall occur on the audio thread.
Given the adaptivity engine updates 5 EQ parameters simultaneously,
When the audio thread reads these parameters during the next render callback,
Then no lock contention or priority inversion occurs; verifiable by running the app under Instruments Thread State Trace with no audio thread preemptions attributed to lock acquisition.
FR-ADAPT-03 — Parameter Smoothing / Ramp (P1)
All DSP parameter changes pushed by the adaptivity engine shall be applied via per-sample or per-block linear ramp (minimum: 50 ms ramp time) to prevent audible zipper noise or discontinuities.
Given an EQ band gain changes from 0 dB to +6 dB in a single update,
When the audio thread processes this change,
Then the gain transitions smoothly over ≥ 50 ms; no click or zipper artifact is audible on a sine wave test signal.
FR-ADAPT-04 — Ambient Noise Sensing — On-Demand (P1) (revised per LD-6: on-demand, not continuous)
The app shall sample the microphone only when the user triggers "Adapt to my environment", capturing a short window (~3 s) to estimate ambient SPL and classify it into at least 3 bands (Quiet/Moderate/Loud). The mic shall not be held open between samples — no always-on listening, no persistent macOS mic indicator during normal playback.
Given the user taps "Adapt to my environment" while in a noisy room (>65 dBA),
When the ~3 s sample completes,
Then ambient is classified as "Loud", the corresponding DSP adaptation is applied, and the mic is released (orange indicator clears) within 1 s of sample completion.
Given the user has not triggered an environment sample,
When audio is playing normally,
Then the microphone is not accessed and no mic-in-use indicator is shown.
FR-ADAPT-05 — Volume-Level Tracking (P1)
The app shall monitor the current playback volume level continuously and compute the target equal-loudness compensation curve at each volume change.
Given volume changes from 50% to 30%,
When the volume change is detected,
Then compensation EQ parameters update within 100 ms without audio dropout.
FR-ADAPT-06 — Withdrawn 2026-07-12 — hearing-profile integration removed from scope (founder decision). Prior spec is in git history; may be re-added. (ID retained; not reused.)
FR-ADAPT-07 — Adaptation Transparency Mode (P2)
The app shall provide a "Transparency" debug/analysis view showing a real-time visualisation of which signals are driving which DSP changes (e.g., "Volume → +4 dB bass", "Ambient: Loud → Ratio 3:1").
Given the adaptivity engine is active,
When the user opens the Transparency view,
Then each active adaptation signal is listed with its current value and resulting DSP adjustment, updating at ≥ 2 Hz.
FR-ADAPT-08 — Adaptation Intensity Control (P2)
The app shall provide a master "Adaptation Strength" slider (0–100%) that scales the magnitude of all adaptive DSP changes while leaving the baseline profile intact.
Given Adaptation Strength is set to 0%,
When content, volume, and ambient noise all change,
Then DSP parameters do not change from the baseline profile values (verified in Transparency view).
FR-DEVICE-01 — Output Device Enumeration (P1)
The app shall enumerate all available Core Audio output devices on launch and on device change, displaying device name, type (headphones/speakers/external DAC), and sample rate.
Given two output devices are connected (AirPods and USB DAC),
When the device list is opened,
Then both devices appear with their correct names and type icons.
FR-DEVICE-02 — Auto-Profile Switching on Device Change (P1)
On any default output device change, the app shall automatically load the DSP profile associated with that device (or a generic profile for device type if no specific profile exists) within 500 ms.
Given profile "AirPods Pro" is saved and the user connects AirPods Pro while built-in speakers are active,
When macOS sets AirPods Pro as the new default output,
Then the app loads "AirPods Pro" profile within 500 ms and plays without dropout.
FR-DEVICE-03 — Device Type Classification (P1)
The app shall classify connected output devices as one of: in-ear headphones, over-ear headphones, built-in speakers, external speakers, DAC/amplifier, unknown. Classification shall use device name heuristics and Core Audio transport type.
Given an AirPods Pro device is connected,
When the device is classified,
Then it is identified as "in-ear headphones" and the HRTF + headphone correction pipeline is automatically selected.
FR-DEVICE-04 — Named Profile Creation and Editing (P1)
Users shall be able to create named DSP profiles, associate them with a specific output device, and edit all profile parameters.
Given the user creates a profile named "Night Mode — AirPods" with custom EQ,
When AirPods are connected,
Then the app offers to auto-load "Night Mode — AirPods" (if it is the primary profile for that device) and the user can switch profiles manually from the device menu.
FR-DEVICE-05 — Profile Import and Export (P2)
The app shall support exporting and importing profiles as JSON files to allow sharing between users and backup.
Given a profile is exported to a .json file,
When the file is imported on a different Mac running the app,
Then all profile parameters are correctly restored and audible.
FR-DEVICE-06 — Sample Rate Negotiation (P1)
The app shall query the preferred sample rate of the active output device and configure the DSP chain to match, performing sample-rate conversion if the source material differs.
Given the output device's preferred rate is 48 kHz and a 44.1 kHz file plays,
When the DSP chain initialises,
Then sample-rate conversion is applied transparently and audio plays at the correct pitch and duration.
Withdrawn 2026-07-12 — hearing personalization removed from scope (founder decision). Prior FR-HEAR-* specs are in git history; may be re-added.
FR-UI-01 — Now Playing View (P1)
The app shall display a persistent Now Playing view showing album art, track metadata, playback controls, a real-time spectrum analyser, and the active profile name.
Given a track is playing,
When the Now Playing view is open,
Then the spectrum analyser updates at ≥ 30 fps and all metadata is visible without scrolling on a 13-inch MacBook display.
FR-UI-02 — DSP Controls Panel (P1)
A DSP controls panel shall expose EQ bands, spatialization toggle, dynamics controls, and profile selector in a single view without requiring navigation into settings.
Given the DSP controls panel is open,
When the user adjusts any EQ band,
Then the change is audible in the next audio render cycle (< 23 ms at 512 frames / 44.1 kHz) and the EQ curve visualisation updates immediately.
FR-UI-03 — Menu Bar / Status Bar Item (P2, required P1 for Phase 2)
The app shall optionally run as a menu-bar app (no Dock icon) in Phase 2, showing current profile and quick access to on/off toggle.
Given the app is in menu-bar mode,
When the user clicks the menu bar icon,
Then a popover shows: active profile, current device, enhancement on/off toggle, and an "Open Full App" link.
FR-UI-04 — Onboarding Flow (P1)
A first-run onboarding wizard shall complete in no more than 5 steps, be skippable at any step, and not gate the user from listening while completing it.
Given it is the app's first launch,
When the user skips all onboarding steps,
Then they reach the main player view within 10 seconds with default settings active.
FR-UI-05 — Accessibility — VoiceOver (P1)
All interactive controls shall have meaningful VoiceOver labels and be operable via keyboard alone. The spectrum analyser shall have an accessible text equivalent (e.g., "Spectrum: Bass heavy, moderate highs").
Given VoiceOver is enabled,
When the user navigates through the DSP controls panel with arrow keys,
Then every slider and button announces its label and current value.
FR-UI-06 — Dark Mode Support (P1)
The app shall fully support macOS Dark Mode and Light Mode, switching automatically with system preference.
Given the system switches to Dark Mode while the app is open,
When the appearance changes,
Then all app windows update to the dark theme without requiring a restart, with no illegible text or invisible controls.
FR-UI-07 — Adaptive UI Feedback (P2)
The UI shall provide subtle, non-intrusive visual feedback whenever the adaptivity engine changes a DSP parameter (e.g., a brief indicator pulse on the affected control).
Given the adaptivity engine updates the bass EQ by 3 dB,
When the DSP controls panel is visible,
Then the bass band indicator briefly animates to signal the change, without disrupting user interaction.
Architecture note: Requirements FR-SYS-07 and FR-SYS-08 cover the PRIMARY PATH (Core Audio process tap, macOS 14.2+/14.4+, no driver). Requirements FR-SYS-01 through FR-SYS-06 cover the FALLBACK PATH (AudioServerPlugIn virtual device, for macOS < 14.2 or where a persistent selectable output device is required). The app shall use the tap path by default when the OS version permits; the driver path is only invoked when the tap path is unavailable or explicitly chosen. See
docs/session-notes/prior-art.md(ADR-002, Proposed) and Journey 2.6.
FR-SYS-07 — Process-Tap System Audio Capture (P1 for Phase 2, primary path)
On macOS 14.2 or later (exact floor to be confirmed in <CoreAudio/AudioHardwareTapping.h> — see CON-10 / OQ-07), the app shall capture all system audio output via a CATapDescription + AudioHardwareCreateProcessTap (or equivalent SDK entry point) combined with a private aggregate device that mutes the original output. The captured audio shall be processed through the DSP engine and replayed to the physical output device. No HAL plug-in shall be installed, no privileged helper shall be used, and coreaudiod shall not be restarted.
Given the user is on macOS 14.2+ and grants audio-capture TCC permission (NSAudioCaptureUsageDescription),
When the tap is activated,
Then audio from any playing application is captured, processed through the DSP chain, and played on the physical output device; no admin password was required and no files were installed in /Library.
Given the process tap is active,
When the companion app is quit or the user revokes audio-capture permission,
Then the tap is stopped and the original output device is unmuted automatically; audio returns to normal within 500 ms with no residual routing or installed artefacts.
FR-SYS-08 — TCC Audio-Capture Permission (P1 for Phase 2, primary path)
The app shall declare NSAudioCaptureUsageDescription in its Info.plist with a plain-language explanation of purpose. The purple audio-capture indicator shall be visible to the user whenever the tap is active. The app shall handle TCC denial gracefully by falling back to the driver path (if available) or displaying a clear explanation of the limitation.
Given the user denies the audio-capture TCC permission,
When the app attempts to activate the tap path,
Then the tap is not created; the app informs the user, offers to use the driver-based fallback path if the OS supports it, and does not crash or enter an inconsistent state.
The following requirements FR-SYS-01 through FR-SYS-06 apply to the driver fallback path only. They remain valid requirements for that path and are not deleted; they are reframed here to make the architecture hierarchy clear.
FR-SYS-01 — AudioServerPlugIn Virtual Device (P1 for Phase 2, fallback path)
When the process-tap primary path is unavailable (macOS < 14.2 or tap denied), the app shall ship a signed, notarised AudioServerPlugIn that appears as a selectable output device in macOS System Settings → Sound. The virtual device shall accept PCM audio from any app and pass it to the DSP engine.
Given the AudioServerPlugIn is installed (fallback path),
When the user selects it as the system output in System Settings,
Then audio from any playing application is routed through the DSP chain and heard on the physical output device.
FR-SYS-02 — Privileged Installer with Minimal Footprint (P1 for Phase 2, fallback path)
Installation of the HAL plug-in shall use a signed privileged helper via ServiceManagement (SMAppService). The helper shall do no more than: copy the bundle to /Library/Audio/Plug-Ins/HAL/ and restart coreaudiod.
Given the user initiates installation (fallback path),
When the privileged helper runs,
Then only the plug-in bundle is written to the HAL directory and coreaudiod is restarted; no other system files are modified (verifiable by fs_usage).
FR-SYS-03 — Safe Uninstall and Fallback (P1 for Phase 2, fallback path)
The app shall provide a one-click uninstall that removes the HAL plug-in, restarts coreaudiod, and restores system output to built-in speakers (or previous device). No manual system repair shall be required.
Given the Phase 2 driver enhancer is installed and active (fallback path),
When the user clicks Uninstall,
Then the plug-in is removed, coreaudiod restarts, and the system output is set to built-in speakers within 10 seconds, leaving no orphaned audio devices in the system.
FR-SYS-04 — Crash-Safe Audio Passthrough (P1 for Phase 2, fallback path)
If the companion app crashes or is force-quit while the virtual device is active, the virtual device shall silently pass audio through unprocessed (bypass mode) rather than causing a system audio outage.
Given the companion app is killed via Activity Monitor while Spotify plays through the virtual device (fallback path),
When the app process terminates,
Then Spotify audio continues to play (unprocessed) within 200 ms; no system-level audio dropout occurs.
FR-SYS-05 — IPC Between Plug-In and Companion App (P1 for Phase 2, fallback path)
Parameter updates from the companion app to the AudioServerPlugIn shall use Mach IPC (registered under the AudioServerPlugIn_MachServices Info.plist key). No other IPC mechanism (file, socket, memory-mapped file) shall be used on the hot audio path.
Given the companion app changes the active EQ profile (fallback path),
When the message is sent via the Mach port,
Then the plug-in receives and applies the update within the next audio render cycle.
FR-SYS-06 — Zero Objective-C in Plug-In (P1 for Phase 2, fallback path)
The AudioServerPlugIn bundle shall be implemented in pure C/C++ with no Objective-C runtime calls, in compliance with AudioServerPlugIn sandbox restrictions.
Given the plug-in bundle is compiled (fallback path),
When inspected with otool -L,
Then no libobjc or Foundation dependency is linked.
Withdrawn 2026-07-12 — natural-language / conversational tuning removed from scope (founder decision). Prior FR-NLT-* specs and the §3.9.1 phrase→intent mapping table are in git history; may be re-added.
The single user-facing control that scales how much we transform the sound (LD-16).
FR-REIMAGINE-01 — Single Intensity Control (P1)
The app shall expose one continuous "Reimagine" intensity control (0–100%) that scales the overall degree of transformation: 0% faithful → rising clarity → spatial widening → (Phase 1.5) full stem-based spatial reimagining.
Given the Reimagine control is at any value, When the user changes it, Then the render transitions smoothly (no clicks/zipper; ramped per FR-ADAPT-03) toward the new intensity.
FR-REIMAGINE-02 — Intensity 0 = Bit-Faithful Bypass (P1)
At 0%, the stem engine and all transformation stages shall be bypassed and the original mix played unaltered — bit-transparent (see NFR-QUAL bit-transparent bypass).
Given Reimagine = 0% at the device's native sample rate, When a loopback capture is compared to the source file, Then the captured audio is bit-identical to the source (MD5-equal).
FR-REIMAGINE-03 — Continuous Mapping, Phased Ceiling (P1)
The intensity→parameter mapping shall crossfade original↔processed and scale spatial spread / unmask depth along the way. Phase 1 implements the mix-level range; Phase 1.5 raises the ceiling into the stem-based spatial range. The exact mapping curve is an open item (OQ — user-tested).
Given Phase 1 (no stem engine yet), When the user raises intensity to maximum, Then only the mix-level range is reachable (clarity + BRIR widening); stem-range behaviour is unavailable until Phase 1.5.
FR-REIMAGINE-04 — Artifact-Conservative Defaults (P1)
Default intensity and behaviour shall be artifact-conservative; higher intensities (which expose separation artifacts) shall be opt-in territory the user chooses by dialing up. Quality-gating (FR-STEM-05) informs how high the stem-range is allowed to go for a given track.
Own-player-only (LD-15). Live/system-wide audio (Phase 2 tap) is mix-level only.
FR-STEM-01 — Offline 6-Stem Separation + Cache (P1 for Phase 1.5) (download integrity & fallback)
On add/first-play, the app shall separate a local track offline into 6 stems (vocals, drums, bass, guitar, piano, other) using an on-device model (Demucs/HTDemucs via Core ML/MLX, MIT) and cache the stems to SSD. Separation is non-real-time and must not block playback. Model weights are auto-downloaded on first run from the official Demucs GitHub releases (primary) with fallback to Hugging Face MLX-Community. Integrity verification: validate the downloaded weights against GitHub's published checksums (SHA-256) before loading. Failure handling: if verification fails or both sources are unreachable, warn the user ("Separation weights unavailable; operating in mix-only mode"), disable stem separation, and proceed with mix-level DSP only. Updates: users obtain new model weights by downloading a new app version (in-app weight updates deferred).
ML backend selection (architecture.md §12, decision tree):
- MLX (primary, production): unconditional; handles STFT and complex operations Core ML doesn't convert cleanly. Expected runtime ~5–15 sec/track on M1 Pro; acceptable for offline pre-pass.
- Core ML (secondary, pre-release only): attempted conversion only if Phase 1.5 tuning data shows MLX latency is a user-visible blocker. Conversion is lossy; only pursue if SPIKE-SEP-QUALITY measures unacceptable overhead.
- Mix-only (safety fallback, always enabled): triggered by download failure, integrity failure, or runtime error on both paths. Full feature set active except FR-STEM-02…06; NL macros fall back to mix-level; Reimagine shows mix-range only.
- SPIKE-SEP-QUALITY (Phase 1.5 tuning gate): measure sec/track at M1 Pro, M4, M5 hardware tiers; lock decision by Phase 1.5 release gate.
Given a local track is added, When offline separation runs (GPU/ANE), Then 6 stem files are produced and cached, and a status indicator reflects progress; playback of the original mix is available immediately regardless.
FR-STEM-02 — Per-Stem Chains + Re-Sum (P1 for Phase 1.5)
Each stem shall be processable with its own gain, EQ, dynamics, and spatial placement (rendered via the BRIR field, FR-SPAT-01), then re-summed to binaural/stereo.
Given cached stems and Reimagine in the stem range, When playback runs, Then each stem is rendered with its own placement/level and the re-summed output reflects the per-stem moves without glitches (Audio-Workgroups-parallel render, NFR-PERF).
FR-STEM-03 — Between-Stem Unmasking (P1 for Phase 1.5)
Masking/clarity shall be computed between stems (ERB/Bark, LD-12) so that a masked source (e.g., vocals under guitar) can be genuinely unmasked — not approximated by mix EQ.
Given vocals are masked by other stems in a region, When unmasking is active, Then the vocal stem is raised / competing stems dipped in the masked ERB bands, measurably improving vocal prominence.
FR-STEM-04 — Withdrawn 2026-07-12 — per-stem natural-language targeting removed from scope, as it was wholly dependent on the withdrawn natural-language tuning feature (FR-NLT / LD-8). Direct (non-NL) per-stem control remains covered by FR-STEM-02. Prior spec is in git history; may be re-added. (ID retained; not reused.)
FR-STEM-05 — Quality-Gating + Graceful Fallback (P1 for Phase 1.5)
6-stem separation (esp. guitar/piano) is the least-robust case. The app shall quality-gate separated stems and gracefully fall back (fewer stems; route poorly-separated content to "other") rather than expose bad stems. Confidence shall bound how far the stem-range / per-stem moves are allowed for that track.
Validation Protocol (SPIKE-SEP-QUALITY method):
- Listening panel: 3–5 listeners with audio engineering background; per-track A/B test (original mix vs. separated+recombined stems vs. Reimagine stem-range).
- Per-stem artifact scoring: each stem scored on a 5-point scale for audible artifacts (leakage, distortion, phase issues). Threshold: ≤1 artifact (on scale 1–5) per stem to pass gating; stems scoring >1 are merged into "other" or the track's Reimagine ceiling is lowered.
- Genre-specific criteria: vocal/drums/bass (most critical, lower threshold); piano/guitar (higher threshold — worst 6-stem cases); "other" (implicit — absorbs poor stems).
- Confidence metric: 1 − (artifacts + confidence_penalty). Confidence ≥0.7 allows full stem-range (Reimagine 0–100%); 0.5–0.7 caps at mix-range (0–50%); <0.5 disables stem features (falls back to mix-only for that track).
- Acceptance: all tested genres must pass (0 broken stems audibly exposed; at least 3 stems fully usable per track).
- See SPIKE-SEP-QUALITY for full test plan, hardware tiers, and confidence curve fitting.
Given a track separates poorly for guitar/piano, When quality-gating evaluates the stems, Then those stems are merged into "other" (or the track is limited to fewer usable stems) and the achievable Reimagine ceiling is reduced accordingly, with no audibly broken stem presented.
FR-STEM-06 — Own-Player-Only Boundary (P1 for Phase 1.5)
Stem features require local files + offline pre-separation and shall be available only in the own player; the Phase-2 system-wide tap path applies mix-level processing only.
Given audio arriving via the Phase-2 process tap (live), When the user requests a stem-level action, Then the app indicates stem features are own-player-only and applies the nearest mix-level equivalent instead.
NFR-PERF-01 — Audio Thread Latency Budget (P1)
Total processing time per render callback on the audio thread shall not exceed 50% of the buffer period at the nominal buffer size (e.g., ≤ 5.8 ms for a 512-frame buffer at 44.1 kHz). Measured by Instruments / CAMetricEngine under normal adaptive load.
Given a 512-frame buffer at 44.1 kHz (11.6 ms period) with all DSP modules active,
When audio plays for 60 continuous minutes,
Then average audio thread CPU usage is ≤ 50% of the period and no buffer under-runs occur (no XRuns reported by coreaudiod).
NFR-PERF-02 — Compute Usage: Quality-First (P1) (revised per LD-10)
Compute is not a primary constraint — spend available compute on quality, and prefer hardware-accelerated, platform-native paths: Accelerate (vDSP/vForce/BNNS), Core ML on the Neural Engine, Metal/MPS on the GPU, and multi-core parallelism (including macOS Audio Workgroups for any real-time helper threads). There is no fixed CPU-percentage cap. Hard limits: (a) the real-time per-buffer deadline (NFR-PERF-01) is never missed; (b) sustained load stays within reasonable battery/thermal bounds, for which an optional efficiency profile may reduce quality.
Given all DSP active at the max-quality profile on Apple Silicon,
When audio plays continuously,
Then there are zero audio-thread overruns (NFR-PERF-01 holds), regardless of average CPU/GPU/ANE utilisation.
NFR-PERF-03 — Memory & Storage: Quality-First (P1) (revised per LD-10)
RAM and SSD are not primary constraints — use memory and disk generously for quality: cached full-track pre-analysis, decoded look-ahead buffers, impulse responses, FFT plans, lookup tables, and precomputed filters (cached to fast SSD across sessions). There is no fixed resident-memory cap. Two rules still hold: (a) all real-time buffers are pre-allocated at session start; (b) no heap allocation occurs on the audio thread (CON-01).
Given all DSP and full-track pre-analysis are active,
When memory is inspected via Instruments Allocations,
Then zero allocations occur in the render callback; resident/disk-cache size is bounded by cache policy, not a fixed cap.
NFR-PERF-04 — Content Analysis: Parallel Look-Ahead Pre-Analysis (P1) (revised per LD-10)
Content/genre and signal analysis runs off the real-time thread, may be parallelised across cores / GPU / Neural Engine, and in the own-player may pre-scan ahead of the playhead (up to the full track), caching results to RAM/SSD. No CPU-percentage cap applies to this non-RT work. A usable classification shall be available no later than 5 s into playback, and ideally before playback from pre-analysis.
Given a new track starts (or is pre-scanned before playback),
When analysis completes,
Then a content classification is available to the adaptivity engine at or before 5 s of playback, without affecting the real-time render deadline.
NFR-PERF-05 — End-to-End Added Latency (Phase 2) (P1)
In Phase 2 virtual device mode, the total additional round-trip latency introduced by the DSP pipeline (reading from virtual device, processing, writing to physical device) shall be ≤ 10 ms.
Given music plays via the virtual device chain with a 256-frame hardware buffer,
When measured with a loopback cable and impulse-response latency test,
Then added latency is ≤ 10 ms.
NFR-PERF-06 — Stem-Engine Render Budget (P1 for Phase 1.5) (new per LD-15 / architecture.md §15)
The Phase-1.5 stem render (up to 6 stems × per-stem EQ/dynamics/spatial + BRIR convolution, re-summed) shall hold the per-buffer deadline (NFR-PERF-01) on the M1 Pro / 16 GB floor (LD-18; foreground sole-occupancy — current M4/M5 hardware has ~3–4× headroom, so this is now Low-risk) by: doing all heavy work off-RT (separation, FIR/BRIR design, masking — pre-computed/cached); running fixed partitioned convolutions parallelised via Audio Workgroups; sharing one late-reverb tail across stems (cheap per-stem placement filters); and the QualityProfile auto-scaling stem count / reverb-tail length (not buffer size) under thermal/battery pressure. Cached-stem + BRIR-kernel memory shall be bounded by a cache policy.
Given 6 cached stems at the max-quality profile on a base Apple-Silicon laptop, When all stems render with per-stem chains + BRIR convolution for 60 continuous minutes, Then zero audio-thread overruns occur (NFR-PERF-01 holds); if the budget cannot be met, the QualityProfile reduces stem count / convolution length before any overrun. (A pre-Phase-1.5 spike must measure real per-stem cost + memory — backlog SPIKE-PERF-BUDGET.)
NFR-QUAL-01 — THD+N (P1)
Total harmonic distortion plus noise in bypass mode (no DSP processing) shall be ≤ -90 dB (0.003%) at 1 kHz, 0 dBFS.
NFR-QUAL-02 — No Glitch Policy (P1)
Zero audible glitches (clicks, pops, dropouts > 1 ms) during any continuous 1-hour playback session on supported hardware, at default buffer sizes.
Given 60 minutes of continuous playback on a MacBook Pro M3 with all Phase 1 DSP active,
When the session ends,
Then the XRun count reported by the app's internal monitor is 0.
NFR-QUAL-03 — Bit-Transparent Bypass = Reimagine Intensity 0 (P1) (revised per LD-16)
At Reimagine intensity 0% (and in any explicit bypass), the stem engine and all DSP stages shall be bypassed and the output shall be bit-for-bit identical to the input (after any required format conversion) — no unintentional dithering, gain, or processing. This is the architecture's fidelity anchor (FR-REIMAGINE-02).
Given a 24-bit FLAC file plays in bypass mode at the device's native sample rate,
When a loopback capture is compared to the source file,
Then the captured audio is bit-identical to the source (verified by MD5 comparison).
NFR-QUAL-04 — Sample Rate Support (P1)
The DSP engine shall support sample rates of 44.1 kHz, 48 kHz, 88.2 kHz, and 96 kHz without quality degradation. Sample-rate conversion, when required, shall use a high-quality algorithm (SSRC class or equivalent, stopband attenuation ≥ 90 dB).
NFR-PRIV-01 — Microphone Data Locality (P1)
Microphone data shall be processed entirely on-device. No raw microphone audio frames shall ever leave the device. Only the derived ambient SPL estimate (a scalar) is used internally. This shall be documented in the Privacy Policy and App Store privacy nutrition label.
NFR-PRIV-02 — Microphone Permission Transparency (P1)
The app shall present a clear, user-readable NSMicrophoneUsageDescription string before requesting microphone access. Denial shall not prevent any core feature except ambient-noise-based adaptation.
Given the user denies microphone permission,
When the app continues to run,
Then all features except ambient-noise adaptation function normally and the user is informed via a persistent but dismissable banner.
NFR-PRIV-03 — Withdrawn 2026-07-12 — governed hearing-profile data storage; removed with hearing personalization (founder decision). May be re-added. Prior text in git.
NFR-PRIV-04 — Telemetry and Analytics (P2)
If the app collects any usage telemetry, it shall be strictly opt-in, clearly described at onboarding, and limited to anonymous quality and diagnostics data (e.g., crash-free rate, audio-engine error counts). It shall exclude any audio content or personal identifiers. There is no conversion-oriented or commercial analytics purpose — this is a personal/open-source project (LD-9). Users shall be able to review and delete their telemetry data. An open-source project may choose to omit telemetry entirely; this requirement applies only if telemetry is implemented.
Given the user opts out of telemetry,
When network traffic is monitored,
Then no analytics events are sent after opt-out.
NFR-PRIV-05 — App Sandbox Compliance (P1)
The companion app (not the HAL plug-in) shall be sandboxed per App Store requirements. Any entitlements required (e.g., com.apple.security.device.microphone) shall be declared and justified in the app's entitlements file.
NFR-REL-01 — Crash-Free Rate (P1)
The app shall achieve a crash-free session rate of ≥ 99.5% as measured over a rolling 7-day production window.
NFR-REL-02 — Watchdog Recovery (P1)
If the audio engine encounters an unrecoverable error (e.g., device disconnected mid-render), it shall log the error, stop playback gracefully, and recover to a playable state within 3 seconds without requiring an app restart.
Given the active output device is forcibly disconnected mid-playback,
When the error is detected,
Then playback stops cleanly, an error banner appears, and the user can resume playback after reconnecting or selecting another device.
NFR-REL-03 — State Consistency (P1)
No combination of user actions (rapid profile switching, device hotplug, simultaneous seek and device change) shall leave the DSP engine or UI in an inconsistent state (e.g., wrong profile applied, incorrect device shown).
NFR-INSTALL-01 — Standard App Install (P1)
Phase 1 app installation shall follow standard macOS drag-to-Applications convention. No kernel extensions, no privileged installers, no sudo required. Phase 2 tap-path activation (FR-SYS-07) also requires no privileged install — only a TCC audio-capture consent dialog (NSAudioCaptureUsageDescription). The requirements below (NFR-INSTALL-02 through NFR-INSTALL-04) apply to the driver fallback path (FR-SYS-01..06) only.
NFR-INSTALL-02 — Phase 2 Privileged Install — User Informed Consent (P1, driver fallback path)
Before invoking any privileged installer (driver fallback path only), the app shall display a plain-language explanation of what will be installed, what system change will occur (coreaudiod restart), and how to uninstall. User must click a clearly labelled confirmation.
NFR-INSTALL-03 — Phase 2 Safe Fallback (P1, driver fallback path)
If the AudioServerPlugIn fails to load after coreaudiod restart (e.g., due to a signing issue or incompatibility), system audio shall automatically fall back to built-in speakers. The app shall detect this failure and guide the user through a recovery or uninstall flow.
Given the plug-in bundle is corrupt or unsigned (fallback path),
When coreaudiod restarts,
Then coreaudiod ignores the plug-in, audio falls back to built-in speakers, and the companion app detects the failure within 10 seconds and shows a recovery prompt.
NFR-INSTALL-04 — Clean Uninstall (P1)
Uninstalling the app (Phase 1 / tap path: delete from Applications or disable tap — no files to remove beyond app bundle; driver fallback path: in-app uninstall) shall leave no residual files in /Library/Audio/Plug-Ins/HAL/, /Library/LaunchDaemons/, or application support directories. For the tap path, "clean uninstall" means: tap stopped, original output unmuted, TCC permission may be revoked by the user independently in System Settings — no additional cleanup required by the app.
NFR-ACC-01 — VoiceOver Full Navigation (P1)
All controls shall be navigable and operable via VoiceOver. No feature shall be accessible only via mouse gesture.
NFR-ACC-02 — Keyboard Navigation (P1)
All primary functions (play/pause, skip, volume, profile select, DSP toggle) shall be accessible via keyboard shortcuts with no conflicts with standard macOS shortcuts.
NFR-ACC-03 — Dynamic Type / Text Size (P2)
The app shall respect the user's preferred text size setting and remain legible at all macOS accessibility text size levels.
NFR-ACC-04 — Reduce Motion (P1)
All animations shall be disabled or reduced when the macOS "Reduce Motion" accessibility setting is active.
Given "Reduce Motion" is enabled in System Settings,
When the app plays and the adaptivity engine fires visual feedback,
Then no animations play; state changes are indicated by colour or label change only.
NFR-L10N-01 — String Externalization (P1)
All user-visible strings shall be stored in localizable .strings files from the initial release. No hardcoded English strings shall appear in UI components.
NFR-L10N-02 — RTL Layout Support (P2)
The app UI shall be designed to support right-to-left layout mirroring for future Arabic and Hebrew localization.
NFR-L10N-03 — Locale-Independent Audio (P1)
No audio processing logic shall depend on locale settings (e.g., number formatting for frequency values). All internal DSP values shall use invariant floating-point representation.
The following table maps each input signal consumed by the Adaptivity Engine to the DSP parameters it controls, the direction of change, and the rationale grounded in psychoacoustics.
| Signal | Signal Source | DSP Parameter Adjusted | Direction / Rule | Rationale |
|---|---|---|---|---|
| Playback Volume (dB) | System volume API / in-app volume | Bass EQ gain (80–200 Hz), Treble EQ gain (8–12 kHz) | Low volume → boost; high volume → reduce | Fletcher-Munson / ISO 226:2003 equal-loudness contours: human hearing is less sensitive to bass and treble at low SPLs; compensation restores perceived tonal balance. |
| Playback Volume (dB) | System volume API | True-peak limiter threshold | Lower volume → looser threshold; high volume → tighten | At high volumes, headroom matters more for safety and distortion avoidance. |
| Ambient Noise Level (dBA SPL) | Mic-derived short-term A-weighted SPL | Multi-band compressor ratio | Louder ambient → higher ratio | Noise masking reduces dynamic range perception; mild compression improves speech and transient intelligibility in noise (cocktail-party effect analogy). |
| Ambient Noise Level (dBA SPL) | Mic-derived SPL | Low-frequency gain (100–300 Hz) | Louder ambient → slight boost | Low-frequency noise (HVAC, traffic) masks bass; partial lift maintains warmth. Limit: ≤ +3 dB to avoid muddiness. |
| Ambient Noise Level (dBA SPL) | Mic-derived SPL | Mid-frequency presence EQ (2–4 kHz) | Louder ambient → slight boost | Presence band improves speech intelligibility and melodic definition against broadband noise. |
| Output Device Type | Core Audio device classification | Spatialization mode | Headphones → HRTF binaural + crossfeed; Speakers → stereo widening + mid-side; DAC → neutral / user-defined | HRTF on speakers produces incorrect cues (already physically spatialised); speakers need width, not in-head correction. |
| Output Device Type | Core Audio device classification | Headphone correction EQ | Active on headphones; inactive on speakers | Headphone frequency response deviates from diffuse-field target; correction restores neutral tonality. Speaker correction is room-dependent (handled separately or deferred). |
| Output Device Type | Core Audio device classification | Psychoacoustic bass enhancement | Enabled for small speakers / in-ear headphones; reduced for large over-ear or external speakers | Small transducers cannot reproduce sub-bass fundamentals; harmonic synthesis creates perceptual bass without excursion risk. |
| Content / Genre Classification | Non-RT spectral + rhythm analyser | Tonal EQ curve | Classical → subtle curve (preserve dynamics); Electronic → bass + sub emphasis; Vocal/Acoustic → presence boost; Rock → controlled low-mid | Each genre has distinct spectral energy distribution and listener expectation; genre-tuned curves complement rather than override device correction. |
| Content / Genre Classification | Non-RT spectral analyser | Dynamic range compressor ratio + attack/release | Classical → low ratio (≤ 1.5:1), slow attack; Electronic → moderate ratio (2–3:1), fast attack; Podcast/Speech → higher ratio (3:1+), very fast attack | Dynamic range of source content varies enormously by genre; mismatched compression destroys either musical impact (over-compress) or intelligibility in noise (under-compress). |
| AirPods Head Orientation (quaternion) | CoreMotion / AirPods motion API | HRTF virtual source azimuth + elevation offset | Real-time counter-rotation: offset = –(head yaw, pitch) | Stabilises virtual soundstage in world space so it does not move when user turns head; mimics natural localisation of external sound sources. |
| AirPods Head Orientation (quaternion) | CoreMotion | Head-tracking update rate gate | Only update HRTF when orientation delta > 1 degree | Avoids unnecessary DSP parameter churn from sensor noise; <1 degree change is below perceptual threshold. |
| ID | Assumption | Impact if Wrong |
|---|---|---|
| ASM-01 | Target users have macOS 14 (Sonoma) or later. API availability (CoreMotion for AirPods, AudioObjectAddPropertyListenerBlock, AVAudioEngine) assumed on this baseline. | Lower deployment target would restrict head-tracking (AirPods CoreMotion requires macOS 14+) and other APIs. |
| ASM-02 | The app will be distributed outside the Mac App Store initially (to avoid sandboxing restrictions on AudioServerPlugIn in Phase 2). | App Store distribution would require separate entitlement review for HAL plug-ins; currently not straightforward. |
| ASM-03 | Developer ID signing and notarization are in place before any Phase 2 beta. | Without notarization, macOS Gatekeeper blocks plug-in load; system audio fails silently. |
| ASM-04 | SADIE II (Apache-2.0) is confirmed as the default shipped HRTF dataset (OQ-04 resolved). Additional datasets KEMAR and CIPIC are also available under permissive/compatible terms. IRCAM Listen is explicitly avoided (unverifiable license). Custom HRTF measurement is deferred per LD-7. | Resolved — SADIE II covers the requirement; no blocking risk. |
| ASM-05 | Device correction EQ curves are sourced from AutoEq computed parametric curves (MIT, + attribution). Raw measurement databases from upstream measurers are not shipped (may be CC-BY-NC-SA); only AutoEq's derived curves are included. Upstream measurement provenance must be verified per curve before shipping (OQ-08 resolved as AutoEq, but provenance check is ongoing). | Without verified correction curves, headphone EQ is generic; differentiating feature weakens. Provenance verification is a per-model ongoing task. |
| ASM-06 | The content/genre classifier runs as a lightweight on-device ML model (Core ML) or a signal-processing heuristic, not a cloud inference call. | Cloud inference would violate real-time requirements, add latency, and raise privacy concerns. |
| ASM-07 | Apple will not revoke or restrict AudioServerPlugIn entitlements for independent developers between now and Phase 2 launch. | Apple has not announced changes, but policy can shift; monitor Apple Developer Forums. |
| ASM-08 | The AirPods motion data API (CMHeadphoneMotionManager) is accessible from a sandboxed companion app without additional entitlements beyond the standard headphone motion permission. | Additional entitlement requirement would delay feature. |
| ASM-09 | Withdrawn 2026-07-12 — this assumption concerned natural-language tuning latency (FR-NLT), a now-withdrawn feature. May be re-added with the feature. (ID retained; not reused.) | — |
| ID | Constraint | Source |
|---|---|---|
| CON-01 | No heap allocation on the real-time audio thread. All audio buffers must be pre-allocated at session initialisation. | Core Audio real-time thread rules; Ross Bencina "Real-time audio programming 101". |
| CON-02 | No mutex, lock, or blocking call on the audio thread. All cross-thread communication via lock-free ring buffers (e.g., TPCircularBuffer, CARingBuffer) or std::atomic. | Core Audio real-time constraints. |
| CON-03 | AudioServerPlugIn must be pure C/C++ — no Objective-C runtime, no Swift, no Foundation. | QA1811; AudioServerPlugIn sandbox restrictions. |
| CON-04 | AudioServerPlugIn runs inside coreaudiod; IPC with companion app must use registered Mach services (AudioServerPlugIn_MachServices plist key). | QA1811. |
| CON-05 | macOS provides no direct interception of another app's audio stream via a public "tap all audio" API on macOS < 14.2. On macOS 14.2+, Core Audio process taps (CATapDescription + AudioHardwareCreateProcessTap) provide this capability without a virtual device. The driver path (AudioServerPlugIn) remains required for macOS < 14.2 or when a persistent selectable output device is needed. |
Core Audio architecture; process tap API introduced in macOS 14.2/14.4 (confirm exact floor — see CON-10). |
| CON-06 | Microphone access is user-grantable/revocable at any time via System Settings → Privacy. The app must handle mid-session revocation gracefully. | macOS privacy framework. |
| CON-07 | The driver fallback path (AudioServerPlugIn) requires administrator privileges for writing to /Library/Audio/Plug-Ins/HAL/ and restarting coreaudiod. This is unavoidable with the driver architecture. The tap primary path (macOS 14.2+) requires no administrator privileges — only TCC audio-capture consent. | macOS file system permissions / TCC framework. |
| CON-08 | Distribution of the HAL plug-in (driver fallback path) requires Developer ID Application + Developer ID Installer certificates and a passing notarization ticket. | Apple Gatekeeper policy; macOS 13+ enforces notarization strictly. |
| CON-09 | AudioDriverKit dext is NOT an alternative for virtual audio devices — Apple does not grant the required entitlements for this use case. Use AudioServerPlugIn only (driver path). | WWDC21 session 10190; Apple Developer Forums confirmation. |
| CON-10 | The Core Audio process tap primary path requires macOS 14.2 or later (minimum floor; exact version — 14.2 vs. 14.4 — must be confirmed by inspecting <CoreAudio/AudioHardwareTapping.h> SDK headers before engineering begins). This constraint interacts with OQ-07 (minimum OS deployment target). |
docs/session-notes/prior-art.md §5 open verifications; SDK headers. |
| CON-11 | Bass harmonic generation must be derived from a mono-summed (L+R) low band. Per-channel (stereo) harmonic generation is prohibited — it falls within Waves patent US-11,102,577 (active, filed 2018, ~2038 expiry). Formal IP review required before public release (see OQ-16). The MaxxBass approach (US-5,930,373) appears expired (~2019) but must be verified on USPTO before reliance. | docs/session-notes/prior-art.md §6 patent watch. |
| CON-12 | All third-party code and data shipped in the app (libraries, HRTF datasets, EQ correction curves, ML model weights) must be under permissive, redistributable licences (MIT, BSD-2/3, Apache-2.0, Boost, 0BSD, ISC, zlib, public-domain, or equivalent). Copyleft (GPL/AGPL/LGPL) code and NC-licensed weights/data are reference-only and must not be shipped. This constraint implements LD-9 at the dependency level. Approved shipable picks: libmysofa (BSD-3), FFTConvolver (MIT), libebur128 (MIT), libASPL (MIT), SADIE II (Apache-2.0), AutoEq computed curves (MIT — verify upstream measurement provenance per docs/session-notes/prior-art.md §5), Demucs+MLX (MIT). libbs2b licence is disputed and must be resolved before shipping (see OQ-17). |
LD-9; docs/session-notes/prior-art.md §1–§5. |
| ID | Dependency | Type | Risk |
|---|---|---|---|
| DEP-01 | Apple Core Audio framework (AudioHardware, AudioToolbox) | Platform | Low — stable API, backward compatible to macOS 12. |
| DEP-02 | Apple Accelerate / vDSP framework (FFT, biquad, convolution) | Platform | Low — stable, highly optimised for Apple Silicon. |
| DEP-03 | CMHeadphoneMotionManager (CoreMotion) — AirPods head-tracking | Platform | Medium — requires AirPods that support motion; API introduced macOS 14. |
| DEP-04 | AVAudioEngine (Phase 1 own-player) | Platform | Low — stable, well-documented. |
| DEP-05 | AudioServerPlugIn API (Phase 2, driver fallback path) | Platform | Medium — complex, limited documentation; reference BlackHole / Background Music / libASPL (MIT). |
| DEP-06 | HRTF datasets: SADIE II (Apache-2.0) — primary default; also KEMAR (cite), CIPIC (commercial OK — common NC claim is false), ARI (CC BY-SA — keep data under SA). Loaded via libmysofa (BSD-3). IRCAM Listen is avoided (license unverifiable — see docs/session-notes/prior-art.md §5; SADIE II covers the use case). OQ-04 resolved: SADIE II is the default; custom HRTF measurement deferred per LD-7. |
External data + OSS library | Low — SADIE II Apache-2.0 confirmed; libmysofa BSD-3 confirmed. |
| DEP-07 | Headphone correction EQ curves: AutoEq computed parametric curves (MIT, + attribution). Ship AutoEq's computed curves only — do not republish raw measurement databases unchecked (upstream measurers may be CC-BY-NC-SA; verify per docs/session-notes/prior-art.md §5). OQ-08 resolved: AutoEq MIT computed curves are the source. |
External data | Medium — MIT code; upstream measurement provenance requires per-file verification. |
| DEP-08 | ServiceManagement / SMAppService (Phase 2 privileged helper, driver fallback path). SMJobBless is deprecated in macOS 13+; use SMAppService. | Platform | Medium — SMAppService is the current API; evaluate migration path confirmed. |
| DEP-09 | TPCircularBuffer or CARingBuffer (lock-free ring buffer) | OSS library | Low — well-tested, MIT/BSD licensed. |
| DEP-10 | Core ML / SoundAnalysis (off-RT content classification, LD-5 Phase 2 upgrade) | Platform | Low — available macOS 12+; fallback to DSP heuristic classifier possible. Never used on the audio render thread (see BNNS Graph constraint in §3.4). |
| DEP-11 | libebur128 (MIT) — LUFS / true-peak measurement per ITU-R BS.1770 (replaces any bespoke LUFS implementation). | OSS library | Low — MIT confirmed; well-maintained. |
| DEP-12 | FFTConvolver (MIT) — partitioned convolution engine for SOFA HRIR and linear-phase EQ. Confirm LICENSE file path in-repo before vendoring (README says MIT; canonical /LICENSE path 404'd — see docs/session-notes/prior-art.md §5). |
OSS library | Low risk once file confirmed. |
| DEP-13 | libASPL (MIT) — AudioServerPlugIn framework (driver fallback path). | OSS library | Low — MIT confirmed. |
| DEP-14 | Demucs + MLX port (code MIT; weights NC-trained, auto-downloaded on first run, not redistributed) — future offline source-separation feature (LD-15, Phase 1.5+). Offline-only / heavy; not on the audio thread. Code is MIT; weights are licensed under NC terms from MUSDB18-HQ training data and must be downloaded on first run, not bundled. | OSS library + model | Medium — MIT code confirmed; weights license requires download-on-first-run approach and gates any commercial distribution. |
| DEP-15 | BNNS Graph (Apple Accelerate) — RT-safe ML inference on the audio thread. Single-threaded, no runtime allocation. | Platform | Low — part of Accelerate framework, stable. |
| DEP-16 | libbs2b (crossfeed) — docs/session-notes/prior-art.md §5 and OQ-17). Do not ship until licence is confirmed. Algorithm is public; reimplement on biquads if not clearly permissive. |
OSS library (disputed) | High until resolved — see OQ-17. |
This is the single authoritative Open-Questions register; backlog.md Open Items and user-journeys.md defer here.
The following items are unresolved and require founder/product-owner decisions before the relevant requirements can be finalised.
| ID | Area | Question / Gap | Impact if Unresolved | Priority |
|---|---|---|---|---|
| OQ-01 | Phase 2 — Installation | Should the app automatically switch the macOS default output device to the virtual device after install, or instruct the user to do it manually in System Settings? Auto-switching via property API is technically possible but may feel invasive and could conflict with user preference. | Determines onboarding UX complexity and the number of manual steps in Journey 2.6. | Critical |
| OQ-02 | ✓ Resolved — removed (LD-9). The project is personal / open-source and non-commercial. There is no business model, no paid tier, no paywall, and no feature-gating. All features are free. Feature-flag or entitlement-check logic for paid access is not required anywhere in the codebase. | Resolved — no action required. | ✓ Resolved | |
| OQ-03 | Adaptivity Engine — Ambient Sensing | What is the required update cadence and smoothing window for ambient noise estimation? The current draft says "every 2 seconds" with 3-second hysteresis, but an audio engineer must validate whether this produces acceptable latency vs. stability trade-off. | NFR-PERF-04 and FR-ADAPT-04 acceptance criteria cannot be finalised without this specification. | High |
| OQ-04 | HRTF / Spatialization | Which HRTF data set(s) will ship in Phase 1? | FR-SPAT-01 and FR-SPAT-02 scope and timeline depend on this decision. | ✓ Resolved (LD-7 + prior-art pass): SADIE II (Apache-2.0) is the default dataset; custom HRTF measurement deferred. IRCAM Listen avoided (unverifiable license). Rendering is custom SOFA-HRIR partitioned convolution (libmysofa + FFTConvolver) — Apple PHASE/AVAudioEnvironmentNode HRTFs are non-replaceable. See DEP-06 and FR-SPAT-01. |
| OQ-05 | Withdrawn 2026-07-12 (feature removed from scope) | Hearing personalization removed (founder decision); the hearing-calibration positioning question no longer applies. May be re-added with the feature. | — | — |
| OQ-06 | Phase 1 Scope — Streaming Sources | Does Phase 1 (own player) include any streaming source integration (Spotify Connect, Apple Music API, YouTube Music)? Or is Phase 1 strictly local file playback only? The question materially affects FR-PLAY-* and the breadth of source support. | A Spotify Connect or MusicKit integration is weeks of additional work; must be scoped before sprint planning. | ✓ Resolved (LD-4): local files only in Phase 1 |
| OQ-07 | macOS Version — Minimum Deployment Target | The draft assumes macOS 14 for AirPods CoreMotion. Is this acceptable, or does the market require macOS 13 (or 12) support? Lowering the target eliminates head-tracking and may affect other API choices. Additional constraint from prior-art pass: the process-tap primary Phase 2 path requires macOS 14.2 or later (CON-10); the exact floor (14.2 vs. 14.4) must be confirmed in <CoreAudio/AudioHardwareTapping.h> before Phase 2 engineering begins. If the minimum OS is set below 14.2, the driver fallback path (FR-SYS-01..06) becomes the Phase 2 mechanism for those users. |
Determines which APIs are available, whether the tap primary path is viable for the target user base, and the Phase 2 installer approach. | Medium |
| OQ-08 | Device Correction Library — Scope | How many headphone/speaker models will be included in the correction library at launch? Who owns ongoing library curation? | FR-TONAL-02 cannot be validated without knowing the minimum supported device count. | ✓ Resolved (prior-art pass): AutoEq computed parametric curves (MIT + attribution) are the source. Raw measurement databases are not shipped (provenance uncertain). Ongoing per-model provenance verification is required (see DEP-07 and CON-12). Minimum model count and curation owner remain to be confirmed in SPIKE-DEVCORRLIB. |
| OQ-09 | Content / Genre Classifier — Approach | Will the content classifier be a Core ML model (requires training data, model management, CoreML conversion pipeline) or a DSP heuristic (spectral centroid, BPM estimation, onset detection)? The ML approach is more accurate but has higher cold-start and maintenance cost. | FR-ADAPT-01 acceptance criteria and engineering estimates differ substantially between the two approaches. | ✓ Resolved (LD-5): heuristics in Phase 1, Core ML later |
| OQ-10 | Telemetry and Crash Reporting | Will the app use a third-party crash reporting SDK (e.g., Sentry, Firebase Crashlytics)? If so, which one, and how does this interact with the App Store privacy label and NFR-PRIV-04? | Data residency, privacy disclosure, and SDK dependency must be confirmed before SDK is integrated. | Medium |
| OQ-11 | Withdrawn 2026-07-12 (feature removed from scope) | Natural-language / conversational tuning removed (founder decision); the text-interpretation-mechanism question no longer applies. May be re-added with the feature. | — | — |
| OQ-12 | Withdrawn 2026-07-12 (feature removed from scope) | Concerned instrument source-separation for NL requests; NL tuning removed (founder decision). Source separation for the stem engine remains covered by FR-STEM-*. May be re-added with the feature. | — | — |
| OQ-13 | Withdrawn 2026-07-12 (feature removed from scope) | Natural-language tuning removed (founder decision); the ambiguity/clarification-UX question no longer applies. May be re-added with the feature. | — | — |
| OQ-14 | Withdrawn 2026-07-12 (feature removed from scope) | Natural-language tuning removed (founder decision); the multilingual-support question no longer applies. May be re-added with the feature. | — | — |
| OQ-15 | Withdrawn 2026-07-12 (feature removed from scope) | Concerned reconciliation of NL text preferences with automatic adaptation and the hearing profile; both features removed (founder decision). May be re-added with them. | — | — |
| OQ-16 | Patents — Psychoacoustic Bass Enhancement IP Review | CON-11 requires formal IP review before any public release of FR-TONAL-04. Specific items: (a) Verify US-5,930,373 (Waves/MaxxBass, ~2019) is truly expired on USPTO before relying on it. (b) Verify the mono-summed NLD approach is clearly outside Waves US-11,102,577 (active, ~2038). (c) Check whether any Xperi/SRS virtual-bass patents are still active and whether the mono-summed design avoids them. This is a formal legal/IP task, not a technical investigation — requires qualified IP counsel. Engineering may proceed with the mono-summed design; public release is blocked until this review is complete. | Public release of FR-TONAL-04 is blocked without IP review sign-off. Engineering is unblocked (use mono-summed design per CON-11). | High — blocks public release |
| OQ-17 | libbs2b License Dispute | docs/session-notes/prior-art.md §5 notes conflicting reports on the libbs2b licence: one source found MIT in source headers; another reported GPL-2.0+. This must be resolved before libbs2b is shipped (CON-12). Action: open the canonical LICENSE / source header in the upstream repo. If not clearly MIT, reimplement the Bauer crossfeed algorithm from the public specification (a small number of biquad filters + delay — trivial to reimplement cleanly). FR-SPAT-03 / US-DEVICE-07 are blocked on this resolution. |
FR-SPAT-03 (crossfeed) cannot be shipped with libbs2b until confirmed permissive. Reimplementation unblocks the feature if licence is not clear. | Medium — blocks crossfeed shipping |
| OQ-18 | DSP — Phase realization per content | Minimum-phase is the default (LD-13). Open: should transient-dense content force minimum-phase even where linear/mixed-phase is otherwise selected, and what transient-density threshold triggers the switch? Resolve before the Realizer is implemented. | Realizer design + FR-TONAL-01 acceptance. | High |
| OQ-19 | Stem engine — feasibility budget | Measured per-stem RT cost, total memory for 6 cached stems + BRIR kernels, and worst-case render on a base Apple-Silicon laptop are unknown (NFR-PERF-06). A spike must measure these before Phase 1.5 scope is committed. | Gates Phase 1.5 scope; informs QualityProfile scaling. | High — gates Phase 1.5 |
| OQ-20 | Stem engine — quality-gating policy | What objective + perceptual criteria gate a separated stem as "usable" vs. merged into "other", and how does confidence bound the Reimagine ceiling per track (FR-STEM-05)? | FR-STEM-05 acceptance; user-perceived quality. | High |
| OQ-21 | Reimagine — intensity→parameter mapping | The exact mapping from the 0–100% knob to crossfade + spatial spread + unmask depth (and the mix-range vs stem-range ceiling) needs definition and user testing (FR-REIMAGINE-03). | FR-REIMAGINE-03 acceptance; UX. | Medium-High |
| OQ-22 | Perceptual model — masking choice | Which masking + partial-loudness model (Moore-Glasberg vs MPEG psychoacoustic) and what ERB/Bark arbitration details drive clarity/adaptive decisions (LD-12)? | Arbiter design; FR-ADAPT perceptual decisions; between-stem unmasking (FR-STEM-03). | Medium |
| Prefix | Area |
|---|---|
| FR-PLAY | Playback and Source |
| FR-SPAT | Spatialization |
| FR-TONAL | Tonal and Dynamic Optimization |
| FR-ADAPT | Adaptivity Engine |
| FR-DEVICE | Device and Profile Management |
| FR-UI | UI and Controls |
| FR-SYS | Phase 2 System-Wide Enhancement |
| FR-REIMAGINE | Reimagine Intensity Control |
| FR-STEM | Stem-Based Object Engine (Phase 1.5) |
| FR-LIB | (reserved namespace) Library domain — multiple scan folders, single-file play, durable track identity, cross-folder duplicates. No FR text/AC authored yet; currently specified as EP-LIBRARY stories (US-LIB-*) in backlog.md + s8-1-persistent-store-design.md. Promote to full FRs when convenient. |
| FR-PLIST | (reserved namespace) Playlist domain — many-to-many ordered membership, DnD reference-add vs. folder-move, built-in "current" queue, naming. No FR text/AC authored yet; currently specified as EP-PLAYLIST stories (US-PLIST-*) in backlog.md + s8-1-persistent-store-design.md. Promote to full FRs when convenient. |
| NFR-PERF | Performance |
| NFR-QUAL | Audio Quality |
| NFR-PRIV | Privacy |
| NFR-REL | Reliability |
| NFR-INSTALL | Installation / Uninstall |
| NFR-ACC | Accessibility |
| NFR-L10N | Localization |
| STK | Stakeholders |
| ASM | Assumptions |
| CON | Constraints |
| DEP | Dependencies |
| OQ | Open Questions |
| Term | Definition |
|---|---|
| Source Separation | An ML technique (e.g., Demucs / HTDemucs) that isolates individual instruments or vocal tracks from a mixed audio signal. Used by the stem-based object engine (FR-STEM, Phase 1.5). |
| AudioServerPlugIn | Apple's mechanism for a virtual audio device that runs inside the coreaudiod process. Required for system-wide audio interception. |
| AUHAL | Audio Unit Hardware Abstraction Layer. The Core Audio API for direct device I/O in a developer's own process. |
| HRTF | Head-Related Transfer Function. A pair of filters (left/right ear) that encode the spectral and temporal cues the brain uses for spatial hearing. Used to render binaural audio on headphones. |
| Crossfeed | A technique that feeds a fraction of the left channel into the right ear (and vice versa) to reduce unnatural extreme stereo separation on headphones. |
| Fletcher-Munson / ISO 226 | Equal-loudness contours describing how perceived loudness varies with frequency at different SPL levels. Basis for loudness-compensated EQ. |
| Lock-free | A concurrency technique where shared data is accessed without mutexes, using atomic operations or ring buffers. Required on the audio real-time thread. |
| True-Peak | A measure of audio peak level that accounts for inter-sample peaks. Distinct from sample peak; measured per ITU-R BS.1770. |
| XRun | An overrun or underrun of the audio buffer indicating that the audio thread missed its deadline, causing an audible glitch. |
| AutoEQ | An open-source project providing parametric EQ correction profiles for hundreds of headphone models, targeting diffuse-field or Harman target curves. |
| coreaudiod | The macOS system daemon that manages the Core Audio HAL (Hardware Abstraction Layer). |
| SPSC Ring Buffer | Single-Producer, Single-Consumer ring buffer. The standard lock-free data structure for audio thread ↔ UI thread communication. |
| SMJobBless / SMAppService | Apple's ServiceManagement APIs for installing a privileged helper tool. SMJobBless is deprecated in macOS 13 in favour of SMAppService. |
| dBTP | Decibels relative to True Peak (0 dBTP = digital full scale, accounting for inter-sample peaks). |
| LUFS | Loudness Units relative to Full Scale. Standard loudness measurement per ITU-R BS.1770 / EBU R128. |
| Process Tap | A Core Audio mechanism (macOS 14.2+) using CATapDescription + AudioHardwareCreateProcessTap that captures system audio output without installing a HAL plug-in or requiring administrator privileges. Used as the primary Phase 2 mechanism (FR-SYS-07/08). |
| SOFA | Spatially Oriented Format for Acoustics. A standardised file format for storing Head-Related Transfer Functions (HRTFs) as measured impulse-response pairs. Loaded by libmysofa (BSD-3). |
| BNNS Graph | Apple's Accelerate framework API for building and executing neural-network inference graphs on the CPU in a real-time-safe manner (no runtime allocation, single-threaded). The only ML inference mechanism permitted on the audio render thread. |
| libmysofa | BSD-3-licensed SOFA loader library used to read HRTF datasets (e.g., SADIE II) for the custom binaural convolution engine (FR-SPAT-01). |
| libebur128 | MIT-licensed C library implementing ITU-R BS.1770 / EBU R128 loudness and true-peak measurement (DEP-11). |
| SADIE II | Spatially Oriented Format for Acoustics Dataset II. An Apache-2.0-licensed binaural HRTF dataset used as the default SOFA dataset for FR-SPAT-01. |
| AutoEq | Open-source project providing MIT-licensed computed parametric EQ correction profiles for hundreds of headphone models. Used as the source for device correction curves (FR-TONAL-02, DEP-07). Raw upstream measurements are not shipped. |