StillMade AIDeveloper docs
Browse documentation · SDK 0.1.0

Hosted speech generation

Declare a script-to-audio Block with any catalog voice, explicit cost approval, and playback review.

Use runtime: "capability" and operation: "audio.speech" to turn up to 3,800 characters of supplied text into one speech recording. This operation speaks the supplied text; it does not draft a script or run another generation first. The package contains no provider, payment method, endpoint, credentials, or execution receipts. It may suggest a voice and speed; StillMade selects and confirms the final choice outside the package and its optional sandboxed interface.

Package and declaration

Download the complete Narrate a script package. The offline SDK also includes packages/block-sdk/speech-example.js, exporting audioSpeech. Copy the example, choose your own namespace, and edit its fixtures:

sh
node packages/block-cli/cli.js validate examples/narrate-script.stillmade.json
node packages/block-cli/cli.js pack examples/narrate-script.stillmade.json narrate-script.stillmade.json

The package has manifest, capability, tests, and optional view. Its manifest declares exactly one primary input of type text or script, exactly one primary output of type audio, and:

json
{"project":[],"capabilities":["audio.speech"],"network":[],"filesystem":[],"secrets":[]}

The capability object has these four fields; the input and output names must match its declared ports:

json
{"schemaVersion":1,"operation":"audio.speech","text":{"$input":"script"},"output":"audio"}

It may add settings with a suggested StillMade catalog voice and speed. The dialog preselects them when that voice is available, and the person running the Block can still change both:

json
{"schemaVersion":1,"operation":"audio.speech","text":{"$input":"script"},"output":"audio","settings":{"voice":"asteria","speed":1.1}}

voice is a catalog voice id (lowercase letters, digits, - or _) and speed is 0.25 to 4. See Generation models for the voice list.

A text input is a string. A script input contains {"schemaVersion":1,"text":"Words to speak."}. Only its text is forwarded; extra structured script fields cannot select a voice or add provider settings. Text must be nonblank and at most 3,800 characters. Oversized text is rejected, not silently truncated. A default is permitted; scoped script context still requires context.script.read. Declare generic primary ports unless a matching semantic role is needed by a particular downstream contract.

Audio output and fixture expectations

StillMade catalog voices return the provider's measured MP3 or WAV, stored byte for byte (schema 3 receipts, any sample rate, mono or stereo). The OpenAI TTS-1 path normalizes and stores 16-bit PCM WAV at 24,000 Hz, mono. Either way the output is a reference with exactly these fields, using host-created identities and a saved HTTPS or host media URL:

json
{
  "kind":"audio","assetId":"speech-run-one","versionId":"v1",
  "url":"/api/media/owner/speech-one.wav","mimeType":"audio/wav",
  "duration":1,"bytes":48044,"sampleRate":24000,"channels":1
}

This reference is illustrative; it is not a generated file or execution receipt. duration is measured from the actual PCM sample count; bytes includes the canonical WAV header. Audio must be nonempty, at most 64 MiB (67,108,864 bytes), and at most 1,400 seconds. The SDK validates finite bounds and the exact shape; metadata alone cannot establish that bytes exist, are WAV, or belong to a user. The host performs byte checks and stores a separate bound media receipt.

Include 1–3 fixtures using expectations, never an exact output or a fabricated audio reference. Each expectation requires exactly kind, format, minDuration, maxDuration, minBytes, and maxBytes:

json
{
  "name":"Read a short welcome",
  "input":{"script":{"schemaVersion":1,"text":"Welcome to the quiet forest."}},
  "expectations":{"audio":{"kind":"audio","format":"wav","minDuration":0.1,"maxDuration":30,"minBytes":46,"maxBytes":1500000}}
}

Duration bounds are finite seconds with 0 <= minDuration <= maxDuration <= 1400 and a positive maximum. Byte bounds are integers with 46 <= minBytes <= maxBytes <= 67108864. Bounds are inclusive. These checks do not transcribe the output, assess pronunciation, or verify a speaker's identity.

Review, listen, and accept

Importing source and requesting Review cost do not generate audio. The host dialog offers StillMade voices, the same catalog the app's Voiceover uses (Deepgram, ElevenLabs, OpenAI and Kokoro voices configured on the server), paid with StillMade credits at the Voiceover price for the text's length. It also offers OpenAI TTS-1 with credits or your own key, and a speed from 0.25× to 4×. Keys remain in the host's encrypted account vault. BYOK uses the account's OpenAI key; missing keys never switch to credits. Changing text, source, model, voice, speed, or price requires a new quote.

Review the displayed input, voice, speed, cost for every fixture and the separate sample, and the total before choosing Run tests and sample. Each approved invocation produces its own recording and receipt. Listen to the AI-generated audio before choosing Finish review, then Confirm Import to install the Block. A project run has its own Run in project confirmation and Use this audio acceptance. The shared ledger binds source, inputs, account, project, voice, speed, payment, quote, and retained outputs. Uncertain submissions are not automatically repeated; recovering status does not itself generate audio.

The result is a typed audio output and a playable Block result. After accepting a confirmed project speech run, open Media in the Editor and find Audio from Blocks to listen and explicitly Add at playhead. See Add speech to the Editor. Generic audio outputs can also connect to compatible SDK audio inputs.

Adding audio to the Editor does not create or select a native Voiceover take, replace its master recording or narration, or create word timings or a transcript. Voiceover's primary input remains script; the Editor action does not introduce an audio-to-Voiceover connection. A sandboxed view may request StillMade.run, which opens the host approval flow; there is no guest previewAudio, timeline mutation, or provider API.

Validation status and host integration

No live provider speech test has been completed for this implementation. Schema, fixture-contract, and host-boundary tests use synthetic inputs and explicit mocks. They do not claim successful provider speech generation. Every installation still requires the actual host fixtures, sample, playback review, and import confirmation for the selected account and settings.

Offline validate and pack work without a key. pack reports tests: 0, reviewRequired: true, and liveVerified: false. Offline test and preview return RUNTIME_UNAVAILABLE; they do not fetch speech or return placeholders.

prepareCapabilityInvocation(pkg, input) returns exactly {operation:"audio.speech",text}. The trusted host executes the separately approved request and calls capabilityOutputs(pkg, canonicalAudioReference). capabilityOperation(manifest) distinguishes text, speech and image contracts; defaultCapabilityExpectations(manifest) supplies their bounded default checks. The existing validateCapabilityExpectations, testCapabilityOutputs, and validateCapabilityResult dispatch through the declared operation.

A trusted speech executor may return {outputs, metadata, mediaReceipts}. mediaReceipts is keyed by the exact audio output name. Its receipt has exactly schemaVersion:1, kind:"audio", ownerId, runId, index, storageKey, sha256, and the canonical reference's assetId, versionId, url, mimeType, duration, bytes, sampleRate, and channels. The SDK validates shape, index 0–3, hash syntax, and equality to the output, and preserves a copy. The authenticated host client must additionally verify account/run identity; package code and custom interfaces cannot submit receipts as execution evidence. Optional metadata remains bounded to 8 KiB and is not proof of generation.