Skip to content

Map relay lines to voice ids for Speechify-capable voice commands #4

Description

@jonmagic

Problem

TSRS 2.0.0 should let different relay lines use different voices when the configured voice service supports it. Speechify supports this directly because POST /v1/audio/speech requires a voice_id, and GET /v1/voices lists the caller's available shared and personal voices. The default say path should stay simple and keep using the current Siri/System voice rather than encouraging per-line macOS voice selection.

Today the TOML voice command has one <voice-id> placeholder. That is enough for a wrapper like Speechify, but TSRS needs a safe way to decide which voice id to substitute for each queued relay line, including new line names the user has not manually configured yet.

Product goals

  • Allow a line such as Brain, Tri-State Relay Service, or Work to map to a specific voice id when the active voice command supports voice ids.
  • Automatically assign a stable voice to new line names when the active provider supports catalog-backed assignment, so users do not have to open Settings unless they want to override a choice.
  • Keep /usr/bin/say behavior unchanged: direct builds default to the current Siri/System voice, and per-line voice mapping should not force alternate macOS voices.
  • Keep provider APIs out of queue/playback core. TSRS should resolve a line to a provider voice id and substitute <voice-id>; provider wrappers decide how to discover voices and synthesize audio.
  • Preserve app-owned playback: voice commands write audio to <output-file> and never speak directly.
  • Fail quiet and surface errors when mappings, provider commands, or config writes are invalid.

Proposed TOML shape

Use provider-specific settings rather than a single global [voice.line_voices] table. [voice] selects the active provider, and TSRS uses that provider name as an opaque key into [<provider>] and [<provider>.line_voices].

[voice]
provider = "speechify"
command = "<app-bin>/speechify --text-file <text-file> --output-file <output-file> --voice-id <voice-id> --keychain-service TSRS_SPEECHIFY_API_KEY"

[speechify]
default_voice_id = "george"
auto_assign_line_voices = true
catalog_command = "<app-bin>/speechify voices --keychain-service TSRS_SPEECHIFY_API_KEY"
assignment_strategy = "stable-round-robin"

[speechify.line_voices]
Brain = "george"
"Tri-State Relay Service" = "henry"
Work = "simba"

Semantics to decide during implementation:

  • voice.provider is optional for existing configs. If absent, TSRS keeps current behavior and does not use provider-specific line mappings.
  • The active provider name is opaque to TSRS. speechify is just the first provider section; future sections such as [elevenlabs] can follow the same local contract without changing queue behavior.
  • [<provider>].default_voice_id is the fallback for unmapped lines.
  • [<provider>.line_voices] stores exact normalized line-name overrides.
  • [<provider>].auto_assign_line_voices = true is required before TSRS auto-assigns voices for new lines. Auto-assignment must not be inferred from the command containing <voice-id>.
  • say should not get provider-specific defaults. The current default path should behave as it does today unless the user deliberately configures a provider wrapper.

Automatic new-line assignment

When an unmapped line is played through an active provider that opts into automatic assignment:

  1. TSRS asks the provider wrapper for a catalog using the configured catalog_command contract. TSRS core does not call Speechify or any other provider HTTP API directly and never handles API keys.
  2. The wrapper returns a safe machine-readable catalog, ideally ids only or ids plus non-sensitive metadata.
  3. TSRS picks a voice deterministically from the catalog, using a documented strategy such as stable hash or stable round-robin.
  4. TSRS writes the chosen mapping into [<provider>.line_voices] once and treats it as sticky. Existing mappings are never re-rolled just because the provider catalog changes.
  5. The playback path substitutes the resolved voice id into <voice-id> as a single argv value.

Open design points:

  • Whether assignment should happen when a line first queues, when it first speaks, or only during config validation/refresh.
  • Whether stable-round-robin should use persisted assignment order or a stable hash of normalized line name. Stable hash avoids write-order dependence; persisted assignment order gives more even distribution.
  • How much catalog metadata to cache, if any. Personal/cloned voice names may be sensitive, so relay config show should avoid printing raw catalog details.

Provider wrapper contract

Avoid making TSRS understand every voice service API. Treat voice services as wrappers with a narrow local contract:

  1. TSRS resolves the active provider and voice id for the current relay line.
  2. TSRS substitutes <voice-id> as a single argv value in the configured speech command.
  3. The wrapper translates that id to the provider request. For Speechify, this is the JSON voice_id field on POST https://api.speechify.ai/v1/audio/speech.
  4. If auto-assignment is enabled, the wrapper also supports a catalog command that prints available provider voice ids without exposing secrets.

This keeps ElevenLabs, local models, and future services possible without adding provider-specific API schemas to TSRS core. The only TSRS-owned concepts are active provider, default voice id, line-to-voice-id mapping, and an optional catalog command contract.

Implementation plan

  1. Extend RelayConfig parsing/serialization with voice.provider and generic provider sections that can include default_voice_id, auto_assign_line_voices, catalog_command, assignment_strategy, and [provider.line_voices].
  2. Preserve existing [voice.variables] compatibility and existing TOML files without provider sections.
  3. Add a shared resolver such as resolvedVoiceIdentifier(for line: String, config: RelayConfig, selectedVoice: String?) -> String that applies exact normalized line matching, falls back to provider default voice id, and then falls back to existing selected/system voice behavior.
  4. Add an auto-assignment path gated by explicit provider config. It should reload and merge the current TOML before writing one new mapping so concurrent manual edits are not clobbered.
  5. Wire the resolver into app-owned playback before command-template expansion so <voice-id> is line-specific for BYO commands.
  6. Keep the default say command from changing user-visible voice behavior. If the command does not use <voice-id> or no provider is active, line mappings should be inert.
  7. Update CLI config commands so relay config show displays active provider, default voice id, and explicit line mappings without exposing secrets or raw catalog contents.
  8. Add docs to docs/user-guide.md explaining that Speechify and similar providers can use automatic or overridden per-line voice ids, while say keeps the current system voice.
  9. Add Settings visibility for the effective provider, default voice id, auto-assignment state, and line mappings. Editing can remain TOML-first for v2.0.0 if a full GUI editor is too much.

Acceptance criteria

  • TOML can define an active provider, provider default voice id, and exact provider-specific line-to-voice mappings.
  • Existing TOML files without provider-specific sections continue to validate and behave as they do today.
  • say default playback still uses the current Siri/System voice and never attempts provider catalog fetch or automatic line voice assignment.
  • Speechify-compatible command templates receive different <voice-id> values for different relay lines.
  • New line names can be assigned a stable provider voice automatically when and only when the active provider explicitly enables auto-assignment.
  • Auto-assigned mappings are persisted atomically to TOML, survive restarts, and are not re-rolled when the provider catalog order changes.
  • Auto-assignment writes reload and merge the on-disk TOML so manual edits are not silently overwritten.
  • If catalog fetch fails, the catalog is empty, or voices are exhausted, playback falls back to the provider default voice id or existing selected/system voice according to the resolver contract.
  • Unknown placeholders, malformed mapping values, duplicate/invalid line keys, empty voice ids, invalid provider names, and invalid catalog commands fail validation with actionable errors.
  • Invalid config fails quiet: queued relays are not claimed for speech and Settings/status expose the config error.
  • relay config show, diagnostics, and logs never print secrets or raw provider catalog contents; they may show mapped voice ids that are actively configured.
  • Spoken usage buckets continue to record provider/model/voice/line, and per-line voice ids are reflected in the voice dimension.
  • Tests cover config parsing, serialization, migration/default behavior, active-provider resolution with multiple provider sections, resolver fallback order, exact line matching, auto-assignment disabled by default, sticky auto-assignment across restarts, merge-safe config writes, say inert behavior, Speechify wrapper command expansion, catalog failure fallback, and privacy-safe config output.

Out of scope

  • Building a universal provider registry in TSRS core.
  • Storing API keys in TOML or SQLite.
  • Having the CLI speak directly.
  • Full GUI voice-catalog browsing for every provider.
  • Auto-fetching Speechify voice catalogs during normal playback unless the active provider explicitly opts into auto-assignment.

Notes from Speechify docs

  • Speechify speech synthesis uses POST https://api.speechify.ai/v1/audio/speech with a required voice_id field.
  • Speechify voice discovery uses GET https://api.speechify.ai/v1/voices and includes shared plus personal cloned voices.
  • The Speechify wrapper should remain responsible for API auth, provider-specific request shape, output audio format, voice catalog retrieval, and provider error handling.

Relates to #3.

Activity

  1. jonmagic commented on Jul 6, 2026

    @jonmagic
    OwnerAuthor

    Final closeout: all requirements in this issue have been satisfied on main.

    What is now covered:

    • TOML supports an active voice provider, provider defaults, and exact line-to-voice mappings.
    • The default /usr/bin/say path remains unchanged and provider mappings are inert unless the configured command uses <voice-id>.
    • Speechify-capable commands receive per-line <voice-id> values.
    • New relay lines are automatically assigned stable, sticky provider voice ids when auto_assign_line_voices is explicitly enabled.
    • Auto-assignment reloads/merges the current TOML before persisting mappings.
    • CLI/status, Settings diagnostics, docs, usage buckets, and tests reflect provider voice behavior.
    • The bundled Speechify helper handles provider rate limits more safely: serialized synthesis, 429 retry/backoff using safe headers, and local voice-catalog-id caching without caching relay text or audio.

    I also verified the behavior manually by enqueueing six fresh Live-mode relay lines; all six received persisted Speechify voice mappings and played without a recorded voice-command error.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    security-sensitiveNeeds private security or privacy-aware handling.triage/ai-reviewedReviewed by the cautious issue triage workflow.triage/human-neededHuman maintainer review is needed.

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions