1. Connect your service
Connect to check the model and supported languages.
2. Choose your audio and output
Start with six-second captions. Shorter intervals can reduce accuracy and capacity. English translation means speech translated into English.
3. Capture and save
Ready
Stop and let pending audio finish before downloading. Audio is not saved; closing the tab can lose pending audio and unsaved text.
Send captions to caption.ninja
Optional: caption text will travel through the selected relay (caption.ninja’s public relay by default). Audio stays on the inference host. Paste the editor’s private automatic-caption room below. The editor reviews these captions before publishing to its separate public room.
For a direct overlay, use its room instead. Anyone with that room link can see the shared text; direct captions skip human review.
Use a private relay
A private relay needs its own room token, separate from the inference service token. The editor needs the source viewing token and output publishing token. Links carry the relay address, never tokens. Turn sharing off before changing these settings.
Sharing off
Open caption editorConnection and capture diagnostics
Connect to inspect capture status.
Response duration includes network, queue and inference time; it is not speech-to-caption delay. The last 100 requests are retained in memory. Downloads exclude tokens, service addresses, room names, audio and captions.
Settings include the service address, language, output and capture interval. Tokens, microphone identifiers, room names and sharing are excluded. Connect first, then load settings for that service.