Accessibility and reach

Live Stream Captions

Streamrun can add automatically generated captions to your live stream in the cloud, between your encoder and the platforms. No caption provider, no encoder plugin, no extra hardware. Deliver them as closed captions viewers can switch on, or burn them into the video so they show everywhere. High-accuracy recognition, billed per second, for $0.40 per hour on top of your configuration.

Why captions matter

Captions make a live stream accessible to people who cannot hear the audio. They also help when someone is watching with the sound off or following a language that is not their first. Without text, all of those viewers have a harder time following what is happening.

Deaf and hard of hearing viewers

Captions make spoken content accessible. For public events and education, providing them may also be a legal requirement.

Sound-off viewing

People often watch on a phone, at work, or in a public place with the audio muted. Captions let them follow along.

Non-native viewers

Reading along can be easier than following fast speech or an unfamiliar accent. Captions help without requiring a second audio track.

Captions, subtitles, closed and open

The words get used interchangeably, but they mean different things, and the difference decides what your viewers can do with them.

Captions vs subtitles

Subtitles are a translation of what is said, written for viewers who hear the audio but not the language. Captions transcribe the speech in its own language and also describe sounds that matter, such as applause or music, for viewers who cannot hear the audio at all. When you see "[crowd cheers]" on screen, you are looking at captions.

How captions travel

Captions reach a player in one of three ways. Embedded in the video signal as CEA-608 or CTA-708 data, formerly known as CEA-708, the broadcast standard that live platforms read. As a sidecar file such as WebVTT or SRT, which works for on-demand video but not for a live stream. Or burned into the picture as pixels, which any player can show.

Closed captionsOpen captions (burned in)
Delivered asCTA-708 data embedded in the video streamText rendered into the video frames
Viewer controlSwitched on and off with the CC buttonAlways visible, cannot be hidden
Where they showPlatforms and players that read embedded captionsEvery platform, player, clip, and recording
LimitationInvisible on destinations that ignore caption dataCover part of the picture and cannot be translated by the player

How live captions get made

Live captions cannot be prepared in advance. There are three common ways to produce them.

Captioning software on your own machine

A plugin or desktop application runs speech recognition on the encoding computer and draws the text into the scene before the stream leaves, which is how most OBS caption setups work.

+ Cheap or free, and you keep full control over the setup.

− Uses CPU on the machine that is already encoding, is tied to that one computer, and does not exist for phone or hardware encoders.

Platform-side automatic captions

Some platforms run speech recognition on the incoming stream and offer captions in their own player.

+ Free and needs no setup.

− Tied to one platform. In a multistream the other destinations get nothing, you have no control over quality or vocabulary, and the captions only ever exist as a track in that platform's player, so there is no way to burn them into the picture.

Speech recognition in the stream path

A streaming server between the encoder and the platforms transcribes the audio and inserts captions into the outgoing stream, as embedded data or as burned-in text.

+ Automatic, works with any encoder, reaches every destination at once, and uses none of the streaming computer's CPU because the recognition runs in the cloud.

− Accuracy depends on audio quality, accents, and vocabulary, so a noisy venue needs a good microphone.

How Streamrun adds captions

Streamrun runs speech recognition between your encoder and the streaming platforms. Send one stream from any encoder, add the Captions element to your configuration in the Streamrun editor, and every output after it carries the captions. They are added before the stream is split, so they reach all destinations in a multistream with one setting.

Primary language and custom terms

Pick the language you mainly speak as a hint, with bilingual modes for mixed-language streams. Add channel names, guest names, and product names as custom terms so they are spelled right.

CTA-708 closed captions

Embedded in the stream and shown on YouTube, Twitch, and Kick when the viewer presses CC. The video itself stays untouched.

Burned in, multiple lines

A block of text drawn into the picture two lines at a time. Shows on every destination, in clips, and in recordings.

Burned in, word by word

Words appear one at a time as they are recognized, in the style used on short-form video.

Settings and output modes in detail: Captions element documentation.

Broadcast-grade accuracy, at $0.40 per hour

Cloud video platforms and hosted captioning APIs commonly charge upwards of a dollar per hour for automatic live captions, before their own encoding and delivery fees. Streamrun prices captions like any other element in the stream.

Quality

  • Built for live speech. Recognition is tuned for continuous conversational audio, with punctuation and sentence breaks that make the text readable rather than a running block of words.
  • Around sixty languages. Set the language you mainly speak as a hint, or use a bilingual mode for mixed-language streams. Other languages are still recognized when they come up.
  • Custom terms. Channel names, guest names, product names, and in-game vocabulary are spelled exactly as you write them.
  • Clean audio to work from. The AI noise cancellation element runs in the same pipeline, so wind and crowd noise can be removed before the audio reaches recognition.

Price

$0.40per hour

Added to the configuration base price, which starts at $0.39 per hour. A captioned stream to one destination comes to $0.79 per hour all in.

  • Billed per second. You pay for the time the configuration is live, not for a booked slot or a rounded-up hour.
  • One price, every destination. The captions are added before the stream is split, so a multistream to five platforms costs the same as a stream to one.
  • No contract or minimum. No captioning retainer, no per-seat licence, and nothing to book in advance. Add the element to a configuration and remove it when you no longer need it.
  • A fraction of the going rate. Hosted captioning on the major cloud video platforms is typically billed at several times this, and that is the captioning line alone, separate from what they charge to ingest, encode, and deliver the stream.

Captions are a Streamrun Pro add-on, charged on top of the configuration base price. See pricing.

Frequently asked questions

Does adding captions delay the stream?

Speech recognition runs on the live audio as it passes through, and the video is not held back to wait for it. Captions appear a moment after the words are spoken, which is true of any live captioning.

Which platforms show closed captions from a Streamrun stream?

CTA-708 closed captions display on YouTube, Twitch, and Kick, where viewers turn them on with the CC button in the player. For other destinations, or when you want captions to show without a viewer action, use one of the burned-in modes.

Can viewers turn the captions off?

Closed captions, yes. They are a separate track that the player shows or hides. Burned-in captions are part of the picture and cannot be switched off, which is exactly what makes them work everywhere.

Which languages are supported?

Around sixty languages, including English, Spanish, German, French, Portuguese, Japanese, Korean, Mandarin, Hindi, and Arabic, plus bilingual modes for Malay, Mandarin, Tamil, and Tagalog mixed with English. The primary language setting is a hint, so other languages are still recognized when they come up.

Do captions work with OBS, BELABOX, or a phone app?

Yes. Captions are generated in the cloud from the audio in your stream, so any encoder that sends RTMP, SRT, or SRTLA to Streamrun works. Nothing needs to be installed or configured on the sending side.

How accurate are the captions?

Streamrun uses a speech recognition model built for live audio, so the text keeps up with natural speech, punctuation, and speaker changes rather than producing a flat stream of words. Custom terms let you lock in the spelling of channel names, guest names, and product names. As with any recognition, clean audio makes the biggest difference, and the AI noise cancellation element in the same pipeline helps in loud venues.

How much do captions cost?

The Captions element adds $0.40 per hour to your configuration, charged per second while it runs live. It sits on top of the configuration base price, which starts at $0.39 per hour, so a captioned stream to one destination comes to $0.79 per hour all in. There is no captioning contract, no per-seat fee, and no minimum spend.

Do captions reach every destination in a multistream?

Yes. The Captions element sits in the pipeline before the outputs, so every destination downstream of it receives the same captions.

Add captions to your next live stream

Try Streamrun free for 14 days. No credit card required.

✓ Any RTMP, SRT, or SRTLA encoder✓ Closed or burned-in captions✓ One setup for every destination