跳到正文

oboroge0

hayamimi

早耳 - Real-time multilingual speech-to-text on CPU only. Live subtitles, browser dashboard, speaker labels, translation. No GPU, no cloud.

README 已保存到本站,可直接阅读

Documentation snapshot

README 快照

这篇是英文原文

下面正文是项目自己的英文 README。想读全文就用浏览器自带的整页翻译: Chrome / Edge 点地址栏右侧的翻译图标,或用右键菜单里的「翻译成中文」; 手机浏览器一般在菜单里。

本页保存的是公开项目资料快照,阅读过程不需要连接 GitHub。

hayamimi (早耳)

图片:tests 图片:license 图片:release

Real-time, multilingual speech-to-text on CPU only. Live subtitles, a browser dashboard, speaker labels, and on-the-fly translation — no GPU, no cloud API, under 2GB RAM.

日本語版 README は README.ja.md にあります。

“早耳” (hayamimi) is Japanese for “quick ear” — someone who picks up on things fast. That’s the design goal: partial subtitles appear while you’re still talking, and a finalized line lands roughly 100ms after you stop.

Why

Most CPU-only real-time transcription setups fall back to a single general-purpose model (Whisper) and accept its accuracy ceiling. hayamimi instead routes each utterance to whichever specialist model is best for its language, all running as quantized (INT8) ONNX models via sherpa-onnx — no PyTorch, no CUDA.

On real broadcast Japanese audio (see docs/SCORECARD.md), that routing gets 5.8% CER, less than half of whisper-large-v3-turbo’s 13.8% on the same clips, while running at 10-50x realtime on a 6-core desktop CPU.

Features

FeatureWhat it does
5-route language catalogja/zh/ko/yue/en+24 EU languages each go to a dedicated best-in-class model; everything else (~1600 languages) falls back to Meta’s Omnilingual ASR
Partial subtitlesin-progress draft text updates every ~0.5s while you’re still speaking
Fast finalsa finalized line typically lands ~100ms after you stop talking (ja; see docs/GOALS.md for other languages)
Two-pass refinementafter 2s of silence, recent utterances are batch re-decoded for a higher-accuracy “clean” transcript (ja real-broadcast CER 15.5% -> 12.0%)
Speaker labels--speakers tags each utterance S1/S2/… using CAM++ speaker embeddings (turn-taking, not full diarization)
Translation--translate en,zh,ko translates Japanese lines live (en via FuguMT, zh/ko via M2M-100)
Hotwords / user dictionary--hotwords biases decoding toward proper nouns; --replace does post-hoc find/replace
OBS overlay + dashboard--serve starts a local HTTP server with a browser-source overlay and a live dashboard
Memory-boundedLRU model eviction keeps resident models under a configurable cap (default: <2GB total)
CPU-onlyevery model runs as quantized ONNX via sherpa-onnx; no GPU or PyTorch required

Demo UI

--serve starts a local server exposing three views:

  • http://localhost:8765/dashboard — the live dashboard: a partial-text strip for in-progress speech, a finals feed with language badges, speaker chips, and per-line latency, inline translations under each line, and a second column with the refined (two-pass) transcript as it lands.
  • http://localhost:8765/ — a minimal OBS browser-source overlay (add this URL as a Browser Source in OBS for stream captions).
  • http://localhost:8765/transcript — plain scrolling transcript history.

图片:dashboard

🎬 Watch the demo video — real 4-language audio (ja/en/ko/zh) transcribed live, replayed frame-accurately from a captured session.

Requirements

Python 3.10+ and ffmpeg on PATH. Developed and tested on Windows 11; macOS/Linux are expected to work (all runtimes are cross-platform) but are not yet CI-tested end to end — reports welcome.

Quickstart

python -m venv .venv

# Windows
.venv\Scripts\pip install -r requirements.txt
.venv\Scripts\python scripts\download_models.py

# macOS / Linux
.venv/bin/pip install -r requirements.txt
.venv/bin/python scripts/download_models.py

# Real-time transcription from your microphone
.venv/Scripts/python scripts/realtime_transcribe.py     # Windows
.venv/bin/python scripts/realtime_transcribe.py          # macOS/Linux

# With the dashboard + OBS overlay
.venv/Scripts/python scripts/realtime_transcribe.py --serve
# -> open http://localhost:8765/dashboard in a browser

scripts/download_models.py pulls ~3.1GB of pretrained models into models/ (git-ignored). Pass --minimal for a ~1.1GB ja/en-only install (ReazonSpeech, whisper-tiny, Silero VAD, Japanese punctuation). See THIRD_PARTY_NOTICES.md for what each model’s license commits you to.

CLI reference

All flags are on scripts/realtime_transcribe.py:

FlagDefaultDescription
--wav PATHmic inputsimulate streaming from a 16kHz mono WAV file instead of the microphone
--no-realtimeoffwith --wav, don’t sleep between chunks (fast batch processing)
--threads N4inference threads per model
--no-partialoffdisable in-progress draft subtitles
--min-silence SEC0.35silence duration that ends an utterance; lower = snappier finals, more splits
--max-speech SEC12.0force-finalize an utterance after this many seconds of continuous speech
--max-resident N3max non-tier0 models kept resident (LRU eviction); <=0 = unlimited
--serve [PORT]off, 8765serve the dashboard + OBS overlay at http://localhost:PORT
--no-refineoffdisable the second-pass re-decode of utterance groups
--transcript PATHnoneappend refined transcript lines to this file
--hotwords PATHnonehotword list (one per line) to bias Japanese decoding toward proper nouns
--replace PATHnoneuser dictionary: wrong=right per line, applied to all output
--lang-switch-guard SEC2.0treat a new-language detection shorter than this as noise and keep the session language (0 disables)
--speakersofflabel utterances with speaker ids (S1, S2, …)
--translate [LANGS]off, entranslate Japanese lines to these comma-separated languages (en/zh/ko)

Architecture

                          ┌─────────────┐
  mic / wav ───────────▶ │  Silero VAD │  0.35s end-of-speech + 0.8s preroll
                          └──────┬──────┘
                                 │ speech segment
                                 ▼
                   ┌───────────────────────────┐
                   │  whisper-tiny spoken-LID   │  runs on first ~4s while
                   │  (+ char-set arbitration)  │  the segment is still coming in
                   └─────────────┬─────────────┘
                                 │ language tag
                 ┌───────────────┼────────────────┬─────────────┬──────────────┐
                 ▼               ▼                ▼             ▼              ▼
             ┌───────┐      ┌─────────┐      ┌──────────┐  ┌─────────┐   ┌──────────┐
             │  ja   │      │   zh    │      │  ko/yue  │  │ en + 24 │   │  ~1600   │
             │ Reazon│      │Paraformer│      │SenseVoice│  │EU langs │   │  other   │
             │Speech │      │   -zh   │      │  small   │  │Parakeet │   │Omnilingual│
             │Zipform│      │         │      │          │  │TDT v3   │   │  ASR     │
             └───┬───┘      └────┬────┘      └────┬─────┘  └────┬────┘   └────┬─────┘
                 └───────────────┴────────────────┴─────────────┴─────────────┘
                                                │
                     partial (every ~0.5s)      │      final (~0.1s after end-of-speech)
                     ◀───────────────────────────┴───────────────────────▶
                                                │
                          ┌─────────────────────┼─────────────────────┐
                          ▼                     ▼                     ▼
                 ┌────────────────┐   ┌──────────────────┐   ┌────────────────┐
                 │ ja punctuation  │   │ speaker labeling  │   │  translation    │
                 │ (BERT restore)  │   │ (CAM++, --speakers)│   │ (FuguMT/M2M-100)│
                 └────────────────┘   └──────────────────┘   └────────────────┘
                                                │
                     2s silence: batch re-decode recent utterances (two-pass refine)
                                                │
                                                ▼
                              dashboard / OBS overlay / transcript file

Models are lazy-loaded on first use; an LRU cache evicts the least-recently-used non-Japanese models (--max-resident) so memory stays bounded no matter how many languages a session wanders through.

Measured performance

End-to-end (LID -> routing -> decode -> ja punctuation), real speech, no preroll/two-pass (single clips). en uses WER, all others use CER (yue t2s-normalized). Full methodology in docs/SCORECARD.md.

LanguageClipsLID accuracyRouteMean errorMean RTF
ja1515/15ReazonSpeech7.5%0.071
en1515/15Parakeet v32.3%0.109
zh1212/12Paraformer-zh5.3%0.102
ko1212/12SenseVoice8.1%0.062
yue1212/12SenseVoice6.1%0.061

RTF (real-time factor) well under 0.2 across every route means each route runs 9-16x faster than realtime on CPU alone — see docs/GOALS.md for the full target table and docs/BENCHMARKS.md for the complete iteration log (30+ measured changes, latency/memory/accuracy tradeoffs and why each one was made or rejected).

Headline numbers from that log:

  • Japanese CER 5.8% (beam search) on real broadcast audio, vs. 13.8% for whisper-large-v3-turbo on the same clips — less than half the error rate.
  • ~100ms mean final latency (ja, punctuated); ~236ms mean / 552ms max across a 5-language soak test with every feature enabled.
  • <2GB RAM with --max-resident 3 (1.35GB at --max-resident 2).

Limitations (honest list)

  • Code-switching mid-sentence is not supported. The router picks one language per utterance; a sentence that mixes Japanese and English within itself will have the minority-language portion mangled or dropped. Utterance-level switching (e.g. an interpreter alternating full sentences) works well; word-level switching within one sentence does not.
  • Very short utterances after a jingle/sting/BGM burst can misroute. The language-switch guard (--lang-switch-guard) mitigates this but a session’s very first utterance (before any session language is established) and confidently-wrong LID+decode combinations (where the garbled text happens to match the wrong language’s character set) are known blind spots — see docs/BENCHMARKS.md’s iteration #29 for a quantified before/after.
  • Two overlapping speakers are not separated. --speakers does turn-taking speaker labeling (one embedding per finalized VAD segment, nearest-centroid assignment), not true diarization — simultaneous speech gets one label.
  • Translation quality has a real ceiling, not just a tuning one. FuguMT (ja->en) and M2M-100 (ja->zh/ko) are small models; repetition loops are suppressed but not eliminated, and numeric values are not reliably preserved in ja->zh/ko translation (see docs/TRANSLATE.md and docs/TRANSLATE_M2M.md for measured failure cases before you rely on this for anything numeric or financial).
  • The end-to-end mic pipeline has not been independently verified beyond this project’s own testing — see docs/GOALS.md’s remaining-work section. File an issue if your results differ from the numbers above.

License

Source code is MIT (LICENSE, copyright oboroge0). No model weights are committed to this repository — scripts/download_models.py fetches them from their original publishers at install time, and each carries its own license (THIRD_PARTY_NOTICES.md has the full table).

One model is not permissive: the ja->en translation model (mojicast-fugumt-ja-en-ct2, used by --translate en) is CC BY-SA 4.0 (share-alike). If you redistribute that model’s weights, you must keep attribution and license any redistribution under CC BY-SA 4.0 too. This does not affect hayamimi’s own code license, and does not affect --translate zh,ko (M2M-100, MIT).

Credits

hayamimi exists on top of, and would not exist without:

  • k2-fsa/sherpa-onnx — the ONNX Runtime inference engine every model here runs through.
  • ReazonSpeech (Reazon Human Interaction Lab) — the Japanese ASR model that anchors this project’s accuracy claim.
  • NVIDIA NeMo / Parakeet — English + 24 European languages.
  • Meta AI Omnilingual ASR — the ~1600-language fallback that makes “multilingual” not a lie.
  • FunASR / SenseVoice (Alibaba DAMO Academy) — Chinese, Korean, and Cantonese ASR.
  • Mojicast (ishiki-emo) — design inspiration for the live-captioning pipeline, and the source of the converted punctuation/translation model artifacts this project uses. Mojicast is itself a full offline real-time captioning app worth checking out.
  • Silero VAD — voice activity detection.
  • 3D-Speaker (Alibaba DAMO Academy) — the CAM++ speaker embedding model behind --speakers.
  • Kiwi — Korean morphological tokenizer, used to fix SenseVoice’s token-spaced Korean output.

Contributing

See CONTRIBUTING.md.

Official distribution

获取与安装

暂未发现可确认的官方软件包地址

当前 README 快照没有出现 npm、PyPI、Crates.io、pub.dev 等官方包页链接。本站不会根据仓库名称猜测下载地址。

本站不托管项目文件;需要安装时,请以项目维护者发布的官方文档为准。

使用前核验

本站保存公开资料用于阅读,不代表安全审计或功能背书。安装前请核对许可证、依赖来源和发布签名,不要直接运行来源不明的二进制文件或高权限脚本。