Self-hosted captions

Installation, requirements and getting started

Use your own Caption Local server for speech recognition. No cloud transcription account is needed. Model downloads and hardware settings belong on the server.

Start Caption Local first, then connect. Tokens stay in this tab. For a remote host, use an SSH tunnel. A hosted capture page also needs an allowed origin and browser local-network permission. The bundled page connects to its own server.

Open the local capture page

Speech is transcribed on the computer running this service. Audio is not saved. Start with a quiet microphone and review captions before relying on them at an event.

Use a multilingual model for languages other than English. Automatic detection is less reliable on short phrases. Shorter intervals may reduce recognition accuracy and the number of streams this computer can handle.

Checking model…

Send captions to caption.ninja

Optional: caption text will travel through caption.ninja’s public relay. Audio stays on the inference host. Paste the editor’s private automatic-caption room below. The editor reviews these captions before publishing to its separate public room.

For a direct overlay, use its room instead. Anyone with that room link can see the shared text; direct captions skip human review.

Sharing off

Ready

Connection and capture diagnostics

Connect to inspect capture status.

Response duration includes network, queue and inference time; it is not speech-to-caption delay. The last 100 requests are retained in memory. Downloads exclude tokens, service addresses, room names, audio and captions.

Settings include the service address, language, output and capture interval. Tokens, microphone identifiers, room names and sharing are excluded. Connect first, then load settings for that service.