Quick answer
Closed captioning is text synchronized with a video's audio track, letting a viewer read what's being said and heard instead of relying on sound alone. It includes dialogue, speaker changes, and sound effects — and unlike open captions, the viewer controls whether it's displayed at all.

What is closed captioning?
Unlike a plain transcript, captions are timed to appear and disappear in sync with the spoken word, and they include more than dialogue — speaker changes, sound effects, and audio cues like "door slams" or "tense music playing" are folded into the same text stream.
The word "closed" describes a specific technical property: the viewer controls whether captions appear at all. A video player or device exposes a toggle — usually a "CC" button — that turns the caption track on or off without altering the underlying video file. That control is what separates closed captions from captions burned permanently into the picture, and it's also what makes closed captioning practical to deploy across different languages, players, and accessibility needs without re-encoding the video itself.
Why is closed captioning important for accessibility?
Closed captioning exists because audio carries information that a purely visual medium can't otherwise deliver — dialogue, narration, sound effects, music cues — and for people who are deaf or hard of hearing, that entire channel is unavailable without a text equivalent. Without captions, a video's plot, instructions, or spoken warnings simply don't reach that audience, no matter how well the visuals are produced.
Captions solve this by giving equivalent access to the audio content, not just a rough approximation of it. Accurate captions also carry sound effects and audio cues the visuals alone wouldn't communicate, such as an off-screen alarm or a change in tone signaled by music. That's what distinguishes captioning from a simple description of dialogue: it's built to give a deaf or hard-of-hearing viewer the same understanding of what's happening as someone who can hear the audio.
Who benefits from closed captioning?
People who are deaf or hard of hearing
The primary audience captions were built for, and the group for whom captions are the only way to access spoken and audio content in a video at all.
Viewers in sound-sensitive or noisy environments
A viewer watching in a waiting room, an open office, or next to a sleeping child often can't play audio at all, and captions restore access without needing sound.
People who process written language more easily
Including non-native speakers of the video's language and viewers with certain learning or auditory processing differences, for whom reading along reduces the effort of parsing spoken audio in real time.
Search engines and content platforms
Captions expose a video's spoken content as indexable text — a byproduct of accessibility rather than its purpose, but part of why captioned video tends to perform better in search and discovery.
What is the difference between closed captions and subtitles?
Captions and subtitles look identical on screen — both are timed text at the bottom of a video — but they're built on a different assumption about the viewer. Captions assume the viewer can't hear the audio at all, so they include sound effects, music cues, and speaker labels alongside the dialogue. Subtitles assume the viewer can hear the audio just fine but doesn't understand the language it's spoken in, so subtitles typically translate dialogue only and skip non-speech audio entirely, since a hearing viewer already picks up a slammed door or a tense score from the soundtrack itself.
That assumption is what should decide which one a video actually needs. A foreign-language film for a hearing audience needs subtitles; a video meant to be accessible to a deaf or hard-of-hearing viewer needs full captions, sound effects and all — and a video that serves both audiences at once may need separate caption and subtitle tracks rather than a single file trying to do both jobs.
What is the difference between open captions and closed captions?
Open and closed captions differ in exactly one way: whether the viewer can turn them off. Open captions are rendered directly into the video's picture during encoding, so they're permanently visible to everyone who watches the file — there's no toggle, no setting, and no way to remove them without re-editing the source video. Closed captions live as a separate, toggleable data track that a player reads alongside the video, which means a viewer decides for themselves whether to display them.
That toggleability is usually the deciding factor. A platform that needs viewer choice — a streaming service, a corporate LMS, a YouTube channel — needs closed captions, since forcing captions on every viewer regardless of preference tends to frustrate the audience that didn't need them. Open captions still show up where toggling isn't possible or reliable, such as a video posted to a platform that strips caption tracks, or a promotional clip designed to be watched with sound off by default.
What should high-quality closed captions include?
Verbatim or near-verbatim accuracy
Captions should match what's actually said, not a cleaned-up paraphrase, because a viewer relying on captions has no other way to catch what a paraphrase might quietly drop or change.
Tight synchronization with the audio
Captions that lag behind or run ahead of the spoken word force a viewer to guess which line of text belongs to which moment.
Speaker identification
When more than one person is talking, especially off-screen or in a group, captions need to indicate who's speaking.
Sound effects and audio cues
Non-speech information that changes the scene's meaning, like "[phone ringing]" or "[ominous music swells]," needs to appear in the caption stream.
Readable pacing and line length
Captions that flash by too quickly or cram too much text onto one line ask a viewer to choose between reading and watching.
How do you create closed captions?
Manual transcription and timing
Produces the highest accuracy but takes the most time. The standard approach when precision genuinely matters, like legal, medical, or safety-critical content.
ASR, then human editing
Automatic speech recognition generates a rough caption track quickly, and a human editor corrects misheard words and adds speaker labels and sound cues. Skipping the editing step is where most low-quality auto-captions come from.
Professional captioning services
Vendors that combine trained transcribers with quality review — usually the right call for organizations needing consistent accuracy across a large volume of video.
Built-in platform tools
YouTube, most video conferencing tools, and many LMS platforms generate a draft track automatically. Treat it like any other ASR draft — it still needs review before publishing.
Whichever method produces the first draft, the caption file still needs a synchronization pass against the finished video before publishing, since even small edits to the video's timing can knock a previously accurate caption track out of sync.
What file formats are used for closed captions?
SRT (SubRip)
The most widely supported caption format — a plain text file with numbered blocks of timecoded text. The safest default for general web video, though it lacks styling or positioning support.
WebVTT (.vtt)
The format built for HTML5 video, supporting basic styling and positioning that SRT can't. Most modern web players expect WebVTT specifically.
SCC (Scenarist Closed Caption)
The format broadcast television has historically used, built around the CEA-608 caption standard. Still required by some broadcast and cable delivery pipelines.
TTML / DFXP
An XML-based format capable of more advanced styling and multi-language tracks in a single file, common in streaming and broadcast workflows.
Which format to use is usually decided by the destination platform rather than personal preference — a broadcaster's ingest pipeline may require SCC or a specific TTML profile, while a website embedding HTML5 video will expect WebVTT, so it's worth confirming the required format before a caption file gets produced rather than after.
What are the WCAG requirements for closed captioning?
| 1.2.2 Captions (Prerecorded) | A Level A requirement — the baseline level — requiring captions for prerecorded video with audio content. |
|---|---|
| 1.2.4 Captions (Live) | A Level AA requirement extending that same expectation to live audio content, such as a livestream or webinar, which is technically harder to satisfy since captions have to be generated and synchronized in real time. |
Most organizations targeting WCAG 2.1 or 2.2 AA conformance — the level most legal and procurement standards reference — need to satisfy both criteria, not just the prerecorded one. That's a common gap in practice: a company captions its published video library but overlooks live events, webinars, or streamed training sessions, which still fall under 1.2.4 if the content includes audio that carries meaning.
What laws and standards require closed captioning?
CVAA
The Twenty-First Century Communications and Video Accessibility Act — a U.S. federal law extending captioning requirements to modern video and communications technologies, including video originally shown on TV with captions that's later distributed online.
FCC closed captioning rules
The FCC enforces specific captioning quality standards — accuracy, synchronization, completeness, placement — for broadcast and cable television in the U.S., with defined complaint and enforcement processes.
WCAG, via broader accessibility law
WCAG itself isn't a law, but courts, settlements, and many procurement standards outside broadcast treat WCAG's captioning criteria as the practical benchmark, even where no statute names WCAG directly.
Requirements outside broadcast and U.S. federal contexts vary by jurisdiction and by where the video is published, so a specific compliance obligation should be confirmed against the applicable law for the organization's actual market rather than assumed from general practice.
How can you turn closed captions on or off?
Web video players
Look for a "CC" or "Subtitles/Captions" icon in the player controls, usually in the bottom bar, which opens a menu to enable, disable, or choose between available tracks and languages.
Streaming apps
Netflix, YouTube, and similar platforms control captions from a settings or audio-and-subtitles menu accessible during playback, separate from the device's own system settings.
Smart TVs and cable/satellite boxes
Most have a dedicated captioning toggle in the system or accessibility settings menu, and many remotes include a direct "CC" or "Subtitle" button that overrides individual apps.
Mobile operating systems
Both iOS and Android offer a system-level captioning setting under accessibility settings, which applies captions across supporting apps rather than requiring per-app changes.
What are the most common closed captioning mistakes?
Publishing unedited auto-generated captions
The highest-impact mistake, since ASR errors on names, technical terms, and homophones can change meaning entirely, and a viewer relying on captions has no way to know the text is wrong.
Poor synchronization
Captions that run noticeably ahead of or behind the audio force a viewer to constantly reconcile mismatched timing.
Missing speaker labels in multi-person video
Without labels, a viewer can't tell who said what in a conversation or panel — exactly when the content is dense enough to need captions in the first place.
Omitting sound effects and audio cues
Dropping cues like "[laughter]" or "[siren wailing]" removes information a hearing viewer picks up automatically.
Overloading lines with too much text
Too many words per line or too short a display duration asks the viewer to choose between reading and watching.
Frequently asked questions
Closed captioning is synchronized on-screen text representing a video's dialogue, speaker changes, and audio cues, with the viewer controlling whether it's displayed.
Captions assume the viewer can't hear the audio and include sound effects and speaker labels; subtitles assume the viewer can hear but doesn't understand the language, so they translate dialogue without describing non-speech audio.
Open captions are burned into the video permanently; closed captions are a separate track the viewer can turn on or off.
It gives viewers who are deaf or hard of hearing equivalent access to a video's spoken and audio content, which is otherwise unavailable to them entirely.
Anyone publishing video with meaningful audio, since the audience includes deaf and hard-of-hearing viewers, people watching without sound, and viewers who benefit from reading along with speech.
Automatic captions are a useful starting draft but routinely misinterpret names, technical terms, and accents, so they need human review before publishing.
Through manual transcription, ASR plus human editing, a professional captioning service, or a platform's built-in captioning tool, followed by a synchronization check against the final video.
Common formats include SRT, WebVTT, SCC, and TTML/DFXP, with the right choice usually determined by the destination platform's requirements.
SC 1.2.2 (Level A) requires captions for prerecorded video, and SC 1.2.4 (Level AA) extends that requirement to live audio content.
It depends on the context — the CVAA and FCC rules impose specific requirements on U.S. broadcast and related video, while other contexts vary by jurisdiction and are often shaped by WCAG-referencing procurement or litigation rather than a single direct statute.
Yes — that's the defining feature of closed captions, controlled through the video player, streaming app, device settings, or system-level accessibility settings depending on where the video is being watched.
Verbatim accuracy, tight synchronization, clear speaker identification, included sound effects and audio cues, and pacing that a viewer can actually read without missing the video.

