Skip to content

feat(asr): integrate Voxtype local dictation - #48

Open
quanru wants to merge 14 commits into
mainfrom
feat/voxtype-batch-integration
Open

quanru wants to merge 14 commits into
mainfrom
feat/voxtype-batch-integration

Conversation

@quanru

@quanru quanru commented Sep 22, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • add Voxtype 1.x CLI capability and daemon-state detection
  • delegate microphone ownership to Voxtype for local recording without opening the app capture stream
  • request private file output, wait for the stable completion signal, read the atomic final transcript, and remove transcript artifacts
  • refuse to take over an existing Voxtype recording/transcription and cancel only sessions started by this adapter
  • register Voxtype as a local provider with onboarding, privacy, failure, and readiness states
  • document final-only output, missing live waveform/pre-roll, minimum version, and cloud-engine privacy implications

Dependency

Depends on #47.

Validation

  • PYTHONPATH=src /Users/bytedance/personal/doubao-say/.venv/bin/python -m unittest tests.unit.test_voxtype_runtime tests.unit.test_voxtype_asr_client tests.unit.test_recognition_providers tests.unit.test_deepgram_credentials tests.unit.test_deepgram_asr_client
  • /Users/bytedance/personal/doubao-say/.venv/bin/python -m ruff check src tests packaging
  • PYTHONPATH=src /Users/bytedance/personal/doubao-say/.venv/bin/python -m compileall -q src tests
  • parsed tests/e2e/cases/onboarding-regressions.yaml with PyYAML
  • audited the upstream Voxtype v1.0.1 source contract for status JSON and record stop --wait behavior
  • git diff --check

Validation gaps

  • Linux Voxtype 1.0.1 daemon and sensevoice-small ONNX CPU were tested with a virtual PipeWire microphone; physical microphone and paste injection remain untested
  • Linux PyGObject manager callbacks and GTK voice-test overlay were tested; synthetic onboarding remains covered by CI
  • Voxtype exposes only a final file result to this adapter, so partial text, local pre-roll, and live waveform are not available

Review order

Review after #47 and before the Voxtype controls/status PR.

Review fixes (2026-09-25)

  • Give Voxtype's final-file transcription a 45-second manager deadline, covering its bounded start and record stop --wait operations. Other providers retain the existing one-second deadline.
  • Let an immediate release wait for the start worker's actual result; a slow start can no longer make the stop worker exit without a completion callback.
  • Reproduced both failures with deterministic regression tests before the fix; both pass afterward.

Validation on Linux: make PYTHON=/home/leyang/doubao-say/.venv/bin/python check and make PYTHON=/home/leyang/doubao-say/.venv/bin/python coverage passed (62% overall coverage). Voxtype 1.0.1 was subsequently exercised through a virtual PipeWire microphone and the real daemon; see live benchmark below.

Live Linux benchmark (2026-09-25)

  • Installed Voxtype 1.0.1 with sensevoice-small ONNX CPU and played a public 4-second speech sample into an isolated PipeWire source. No physical microphone was accessed.
  • record stop --wait to final result: 0.206–0.257 s across five runs. The real VoxtypeASRClient.finish_sending() to finish callback: 0.207–0.258 s across five runs. The full TranscriptionManager.handle_release() to result: 0.258–0.360 s across three runs after capture became ready.
  • Standalone voxtype transcribe on an 11-second public audio file took 1.740–1.997 s across five runs, including CLI/model startup.
  • The real daemon exposed another bug: daemon_state calls its command runner without a timeout, while this adapter requires one. Commit 0631ad1 adds a bounded status runner and a regression test. make check passes.
  • The 45 s manager timeout remains a worst-case protection, not expected latency. See the startup diagnosis below for the virtual-device cold-start finding and fix.

Audio startup diagnosis and fix (2026-09-25)

  • Reproduced the ~3.2 s first start with a newly created virtual PipeWire ALSA device. Voxtype verbose logs place the delay between “Found audio device by exact match” and “Using audio device”, during CPAL/ALSA device-name lookup. Reversing the pause_media setting order showed the same first-open delay, so MPRIS was not the cause. Subsequent virtual-device starts were ~0.15 s.
  • Tested the machine's default physical microphone without saving or transcribing audio. Three starts reached the audio capture thread in about 25–35 ms; the 3.2 s delay was specific to the fresh virtual test device.
  • Voxtype 1.0.1 record start returns after sending a signal, before device setup. Commit c866b8f waits for the daemon's recording state before signalling on_open, and cancels an accepted request if readiness fails. Regression test and make check pass.
  • With a newly created virtual PipeWire device and no artificial delay after the app entered RECORDING, three full manager runs all returned text; release-to-result was 0.208–0.258 s. The daemon's state follows device lookup and capture thread launch, though the thread may still complete stream activation afterward.
  • Upstream Voxtype PR #730 proposes an optional persistent microphone stream for hardware with slower wake-up; it is open and is not a dependency of this PR.

Voice-test overlay regression (2026-09-25)

  • Reproduced a missing overlay on the real Linux desktop: DiagnosticTrace.add rejected Voxtype’s audio_delegated stage and aborted the start path before the overlay appeared.
  • Commit ed1dae4 registers that stage and adds a regression test with the real diagnostic collector and an overlay assertion. The test failed with the original ValueError before the fix and passed afterward.
  • make check PYTHON=/home/leyang/doubao-say/.venv/bin/python passes. In an isolated X11 desktop, Midscene clicked “开始语音测试”, observed the floating “正在聆听…” overlay and the “结束并查看结果” button, then completed the test successfully. Voxtype reported recording while the overlay was visible and idle after completion. This verifies the UI and recording lifecycle; speech accuracy on the user’s microphone awaits manual testing.
  • The local Omarchy plugin is installed from a preview rebased onto current main and includes the same fix.

@quanru

quanru commented Sep 23, 2026

Copy link
Copy Markdown
Owner Author

Midscene CI 验证:Ubuntu provider 用例 3/3 通过(Deepgram 入口、Voxtype 本地引导、合成 CLI 语音测试):https://github.com/quanru/doubao-say/actions/runs/35813803711 。真实 Omarchy VM 的相同 3 个用例也 3/3 首次通过:https://github.com/quanru/doubao-say/actions/runs/35818832674 。语音测试经生产 TranscriptionManager/VoxtypeASRClient 调用 CI 假 CLI,并检查合成文本回显与转写文件清理;未覆盖真实麦克风、模型或守护进程。Omarchy runner 已补上 provider 项目白名单。

@quanru
quanru changed the base branch from feat/deepgram-product-integration to main September 25, 2026 11:55

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant