Skip to content

A conference transcript editing workflow that turns raw speech into publishable copy

Content studioUpdated 2026-08-238 min read

In short

A conference transcript editing workflow takes a machine transcript through three ordered passes: fixing speaker and company attribution, removing filler without altering wording, and checking every figure quoted from the stage against a source. Log corrections per thousand words so the quality of the transcription supplier becomes measurable rather than a matter of opinion.

The file lands two days after the show. Ninety minutes of audio, 400 kilobytes of text, speaker labels reading SPEAKER_01 through SPEAKER_05, and the name of your headline sponsor spelled three different ways in the first ten minutes.

A conference transcript editing workflow exists because that file is a record of speech, and speech is repetitive, self-interrupting and full of proper nouns that no general-purpose model has ever seen. Handing it to a writer without a defined set of passes produces either a fortnight of unbudgeted work or an article quoting a company that does not exist.

What automatic transcription actually gets wrong

Start with what the technology is good at, because the answer decides where to spend editing time.

Radford and colleagues published the Whisper work in December 2022, describing a model trained on 680,000 hours of multilingual and multitask supervised audio collected from the web. Their claim is careful and worth reading in its own words: the models "generalize well to standard benchmarks and are often competitive with prior fully supervised results but in a zero-shot transfer setting without the need for any fine-tuning", and "when compared to humans, the models approach their accuracy and robustness".

Approaching human accuracy on standard benchmarks is a real achievement and it tells you almost nothing about your transcript. Those benchmarks are read speech and broadcast audio, scored on word error rate across ordinary vocabulary. A trade conference transcript fails in three places the benchmark does not weight: company and product names, which are novel strings the model has never encountered; speaker turns in a five-person panel with crosstalk; and numbers, where a misheard digit produces a sentence that reads perfectly and says something false.

So the editing passes should attack those three, in that order, and should not spend time on general fluency. The machine is better at fluency than a rushed editor.

Pass one: who spoke, and where they worked at the time

Replace every speaker label with a name, an employer and a role, taken from the programme and checked against the recording rather than assumed from the running order. Panels swap seats, speakers drop out, and the person listed as moderator is regularly the one who talks least.

Do the company names in the same pass, because they cluster in the introductions. Build a term list before the show from the exhibitor list, the sponsor list and the speaker list, and give it to whoever is doing the pass. Most transcription tools accept a custom vocabulary at the point of upload, which moves this work upstream and reduces it to a check.

Employer at the time of speaking is the field people get wrong, because it is filled in later from a current profile. A quote from your 2024 conference carries the job the speaker held in 2024. Keeping that stable across a portfolio is the reason for holding speaker quotes as reusable assets in a register with a date on every row.

Pass two: removing filler without changing what was said

The second pass takes out the hesitations, the false starts, the repeated words at the start of a sentence, and the interjections from other panellists that do not carry content. It stops there.

The line is easier to hold if you separate two kinds of output. A transcript published as a record of the session should stay close to what was said. The United States federal caption quality standards, codified at 47 CFR 79.1(j) and in force in 2026, give a serviceable definition of accuracy for that purpose, requiring that captioning "shall match the spoken words (or song lyrics when provided on the audio track) in their original language (English or Spanish), in the order spoken". In the order spoken is the operative phrase. Reordering a speaker's sentences to improve the argument produces a better read and a worse record.

Anything you lift into an article as a direct quotation is under a stricter obligation again, because quotation marks are a promise to the reader about wording. Where a spoken sentence genuinely cannot be printed without repair, the honest options are to quote a shorter fragment, to paraphrase without quotation marks, or to go back to the speaker. Deciding between those belongs to turning a panel into an article, where it is a writing judgement made against the argument the piece is building.

Non-native English speakers deserve a specific note. A speaker working in their second language produces more false starts and more unusual constructions, and a filler-removal pass applied without care will either leave them looking less fluent than everyone else on the panel or rewrite them into somebody else's voice. Either outcome is a defect. The safer route is a lighter edit and a shorter quoted fragment.

Pass three: the figures quoted from stage

Every number said aloud gets flagged, and each flag needs one of three resolutions: a source, a correction, or removal.

Numbers fail twice over. The model can mishear them, turning 14 per cent into 40 per cent with no signal that anything went wrong, and the speaker can state a figure they cannot source. Both produce the same artefact, which is a confident number in your transcript with your show's name at the top of the page.

The rule to hold is the one that governs the rest of an event data operation. Never publish a number nobody measured. If a speaker says the market grew 12 per cent last year and no published source says that, the transcript can carry the sentence as speech attributed to them, and your article cannot carry the figure as fact. Those are different acts and the transcript pass is where the distinction gets recorded, in a column next to the flag.

How do you tell whether a transcription vendor is any good?

Count corrections per thousand words. It takes one editor a morning to produce, and it is the only comparison that survives a procurement conversation.

Work an example. Take three sessions from your last edition, 12,000 words of machine transcript in total, and run them through both suppliers you are considering. The editor logs every change made in passes one and two.

Supplier A needs 214 corrections, which is 17.8 per thousand words. Supplier B needs 96, which is 8.0. Break the counts down and the difference has a shape: of A's 214, say 96 are company or product names and 47 are speaker attribution errors, which are the two categories a custom vocabulary and a decent diarisation model should be handling.

Convert to time at 25 seconds per correction, including finding it. On a 9,000 word session, A costs 160 corrections at 25 seconds, about 67 minutes. B costs 72 corrections, about 30 minutes. The gap is 37 minutes a session. Across the 40 sessions you actually process in a year, that is roughly 25 hours of editorial time, and at an internal cost of 50 pounds an hour it is about 1,230 pounds a year. Set that against the difference in what the two suppliers charge per hour of audio and the decision makes itself, in either direction.

The same measure has a second use. Log corrections per thousand words by room rather than by supplier and you will find one room that is consistently worse, which is usually a microphone problem the AV team can fix for next year.

Should you publish the transcript as well as the article?

Often yes, and the accessibility argument is the one that settles it independently of traffic.

WCAG 2.2, published as a W3C Recommendation in December 2024, requires at Level A that "Captions are provided for all prerecorded audio content in synchronized media", and at Level AAA that "An alternative for time-based media is provided for all prerecorded synchronized media and for all prerecorded video-only media". A corrected transcript is the raw material for both. A team already doing three passes for editorial reasons is most of the way to a compliance position it would otherwise have to buy.

Two cautions come with it. A verbatim transcript published alongside an article drawn from the same session creates two pages competing on the same phrases, and the pair need a stated relationship rather than being left to sort themselves out. Keep the transcript on the session page, keep the article on the editorial side, and treat which one is canonical as a decision somebody makes on purpose, which is one of the recurring content studio questions across a portfolio. The second caution is permission: publishing a transcript is publication, and whether the speaker agreed to that when they were recorded is a rights question settled before the pass starts, not after it.

Where this stops

None of these passes fixes bad audio. If two panellists share a lapel microphone or the floor microphone was off during the question that produced the best exchange, the transcript has a hole and the workflow has nothing to offer. That case is an argument for reading the session recordings into content assets pass early enough that the AV brief for next year can change.

The deeper limit is that a corrected transcript is still a record of a conversation held for an audience in a room. It carries the digressions, the in-jokes and the twenty minutes the panel spent agreeing with each other, and no amount of cleaning turns that into something a reader will finish. The transcript is an input and a record. Treating it as a publishable asset in itself is the mistake that produces 9,000 word pages nobody reads.

Take one session from your last edition, run the three passes yourself, and log every correction with its category as you go. The category breakdown at the end tells you whether your next conversation is with the transcription supplier, the AV supplier, or the person who writes the speaker briefing.

Questions people ask about conference transcript editing workflow

How accurate is automatic transcription of conference sessions?
Good enough on ordinary words to be worth starting from, and unreliable on the parts that matter most to a trade audience. Radford and colleagues reported in 2022 that their speech recognition models approach human accuracy on standard benchmarks. Those benchmarks do not measure company names, product names or speaker turns, which is where conference transcripts fail.
What are corrections per thousand words?
The count of edits an editor makes to a machine transcript, divided by the transcript length in thousands of words. A transcript needing 18 corrections per thousand words takes roughly twice as long to prepare as one needing 8. Logging the figure by supplier and by room turns a vendor argument into a procurement decision.
Can you clean up a speaker's grammar in a published transcript?
Treat the transcript and the quotation differently. A transcript published as a record should follow the caption standard of matching the spoken words in the order spoken, with filler removed at most. Anything reproduced as a direct quotation in an article carries a stricter obligation, because a reader takes quotation marks as a promise about wording.

Related reading

All content and media articles