# Hosted speech generation

Declare a script-to-audio Block with any catalog voice, explicit cost approval, and playback review.

Use `runtime: "capability"` and `operation: "audio.speech"` to turn up to 3,800
characters of supplied text into one speech recording. This operation speaks
the supplied text; it does not draft a script or run another generation first.
The package contains no provider, payment method, endpoint, credentials, or
execution receipts. It may suggest a voice and speed; StillMade selects and
confirms the final choice outside the package and its optional sandboxed
interface.

### Package and declaration

Download [the complete Narrate a script package](/block-sdk/examples/narrate-script.stillmade.json).
The offline SDK also includes `packages/block-sdk/speech-example.js`, exporting
`audioSpeech`. Copy the example, choose your own namespace, and edit its fixtures:

```sh
node packages/block-cli/cli.js validate examples/narrate-script.stillmade.json
node packages/block-cli/cli.js pack examples/narrate-script.stillmade.json narrate-script.stillmade.json
```

The package has `manifest`, `capability`, `tests`, and optional `view`. Its
manifest declares exactly one primary input of type `text` or `script`, exactly
one primary output of type `audio`, and:

```json
{"project":[],"capabilities":["audio.speech"],"network":[],"filesystem":[],"secrets":[]}
```

The capability object has these four fields; the input and output names
must match its declared ports:

```json
{"schemaVersion":1,"operation":"audio.speech","text":{"$input":"script"},"output":"audio"}
```

It may add `settings` with a suggested StillMade catalog voice and speed. The
dialog preselects them when that voice is available, and the person running the
Block can still change both:

```json
{"schemaVersion":1,"operation":"audio.speech","text":{"$input":"script"},"output":"audio","settings":{"voice":"asteria","speed":1.1}}
```

`voice` is a catalog voice id (lowercase letters, digits, `-` or `_`) and
`speed` is 0.25 to 4. See [Generation models](/docs/reference/generation-models)
for the voice list.

A text input is a string. A script input contains
`{"schemaVersion":1,"text":"Words to speak."}`. Only its `text` is forwarded;
extra structured script fields cannot select a voice or add provider settings.
Text must be nonblank and at most 3,800 characters. Oversized text is rejected,
not silently truncated. A default is permitted; scoped script context still
requires `context.script.read`. Declare generic primary ports unless a matching
semantic role is needed by a particular downstream contract.

### Audio output and fixture expectations

StillMade catalog voices return the provider's measured MP3 or WAV, stored
byte for byte (schema 3 receipts, any sample rate, mono or stereo). The OpenAI
TTS-1 path normalizes and stores 16-bit PCM WAV at 24,000 Hz, mono. Either way
the output is a reference with exactly these fields, using host-created
identities and a saved HTTPS or host media URL:

```json
{
  "kind":"audio","assetId":"speech-run-one","versionId":"v1",
  "url":"/api/media/owner/speech-one.wav","mimeType":"audio/wav",
  "duration":1,"bytes":48044,"sampleRate":24000,"channels":1
}
```

This reference is illustrative; it is not a generated file or execution receipt.
`duration` is measured from the actual PCM sample count; `bytes` includes the
canonical WAV header. Audio must be nonempty, at most 64 MiB (67,108,864 bytes),
and at most 1,400 seconds. The SDK validates finite bounds and the exact shape;
metadata alone cannot establish that bytes exist, are WAV, or belong to a user.
The host performs byte checks and stores a separate bound media receipt.

Include 1–3 fixtures using expectations, never an exact output or a fabricated
audio reference. Each expectation requires exactly `kind`, `format`,
`minDuration`, `maxDuration`, `minBytes`, and `maxBytes`:

```json
{
  "name":"Read a short welcome",
  "input":{"script":{"schemaVersion":1,"text":"Welcome to the quiet forest."}},
  "expectations":{"audio":{"kind":"audio","format":"wav","minDuration":0.1,"maxDuration":30,"minBytes":46,"maxBytes":1500000}}
}
```

Duration bounds are finite seconds with `0 <= minDuration <= maxDuration <= 1400`
and a positive maximum. Byte bounds are integers with
`46 <= minBytes <= maxBytes <= 67108864`. Bounds are inclusive. These checks do
not transcribe the output, assess pronunciation, or verify a speaker's identity.

### Review, listen, and accept

Importing source and requesting **Review cost** do not generate audio. The host
dialog offers **StillMade voices**, the same catalog the app's Voiceover uses
(Deepgram, ElevenLabs, OpenAI and Kokoro voices configured on the server), paid
with StillMade credits at the Voiceover price for the text's length. It also
offers OpenAI TTS-1 with credits or your own key, and a speed from 0.25× to 4×.
Keys remain in the host's encrypted account vault. BYOK uses the account's
OpenAI key; missing keys never switch to credits.
Changing text, source, model, voice, speed, or price requires a new quote.

Review the displayed input, voice, speed, cost for every fixture and the separate
sample, and the total before choosing **Run tests and sample**. Each approved
invocation produces its own recording and receipt. Listen to the AI-generated
audio before choosing **Finish review**, then **Confirm Import** to install the
Block. A project run has its own **Run in project** confirmation and **Use this
audio** acceptance. The shared ledger binds source, inputs, account, project,
voice, speed, payment, quote, and retained outputs. Uncertain submissions are
not automatically repeated; recovering status does not itself generate audio.

The result is a typed audio output and a playable Block result. After accepting
a confirmed project speech run, open **Media** in the Editor and find **Audio
from Blocks** to listen and explicitly **Add at playhead**. See [Add speech to the Editor](/docs/build/audio-to-editor).
Generic audio outputs can also connect to compatible SDK audio inputs.

Adding audio to the Editor does not create or select a native Voiceover take,
replace its master recording or narration, or create word timings or a
transcript. Voiceover's primary input remains `script`; the Editor action does
not introduce an audio-to-Voiceover connection. A sandboxed view may request
`StillMade.run`, which opens the host approval flow; there is no guest
`previewAudio`, timeline mutation, or provider API.

### Validation status and host integration

No live provider speech test has been completed for this implementation.
Schema, fixture-contract, and host-boundary tests use synthetic inputs and
explicit mocks. They do not claim successful provider speech generation. Every
installation still requires the actual host fixtures, sample, playback review,
and import confirmation for the selected account and settings.

Offline `validate` and `pack` work without a key. `pack` reports `tests: 0`,
`reviewRequired: true`, and `liveVerified: false`. Offline `test` and `preview`
return `RUNTIME_UNAVAILABLE`; they do not fetch speech or return placeholders.

`prepareCapabilityInvocation(pkg, input)` returns exactly
`{operation:"audio.speech",text}`. The trusted host executes the separately
approved request and calls `capabilityOutputs(pkg, canonicalAudioReference)`.
`capabilityOperation(manifest)` distinguishes text, speech and image contracts;
`defaultCapabilityExpectations(manifest)` supplies their bounded default checks.
The existing `validateCapabilityExpectations`, `testCapabilityOutputs`, and
`validateCapabilityResult` dispatch through the declared operation.

A trusted speech executor may return `{outputs, metadata, mediaReceipts}`.
`mediaReceipts` is keyed by the exact audio output name. Its receipt has exactly
`schemaVersion:1`, `kind:"audio"`, `ownerId`, `runId`, `index`, `storageKey`,
`sha256`, and the canonical reference's `assetId`, `versionId`, `url`,
`mimeType`, `duration`, `bytes`, `sampleRate`, and `channels`. The SDK validates
shape, index 0–3, hash syntax, and equality to the output, and preserves a copy.
The authenticated host client must additionally verify account/run identity;
package code and custom interfaces cannot submit receipts as execution evidence.
Optional metadata remains bounded to 8 KiB and is not proof of generation.
