Installation, requirements and getting started
Use your own Caption Local server for speech recognition. No cloud transcription account is needed. Model downloads and hardware settings belong on the server.
Speech is transcribed on the computer running this service. Audio is not saved. Start with a quiet microphone and review captions before relying on them at an event.
Use a multilingual model for languages other than English. Automatic detection is less reliable on short phrases. Shorter intervals may reduce recognition accuracy and the number of streams this computer can handle.Checking model…
Optional: caption text will travel through caption.ninja’s public relay. Audio stays on the inference host. Paste the editor’s private automatic-caption room below. The editor reviews these captions before publishing to its separate public room.
For a direct overlay, use its room instead. Anyone with that room link can see the shared text; direct captions skip human review.
Sharing off
Open caption editorReady
Connect to inspect capture status.
Response duration includes network, queue and inference time; it is not speech-to-caption delay. The last 100 requests are retained in memory. Downloads exclude tokens, service addresses, room names, audio and captions.
Settings include the service address, language, output and capture interval. Tokens, microphone identifiers, room names and sharing are excluded. Connect first, then load settings for that service.