Specific answers, since you asked specifically. Those three failure modes are exactly the ones the pipeline targets:
Clips starting mid-sentence. Cut points are chosen off the transcript at sentence boundaries, not fixed time intervals. A clip starts where a thought starts.
Mangled names and jargon. There's a glossary you can populate with recurring proper nouns, guest names, product names, technical terms. It biases transcription toward those spellings and autocorrects near-misses after the fact. For an interview show with recurring guests this matters a lot, and it's the single highest-leverage thing you can set up. Medical and technical vocabularies were the original driver for it.
Crop cutting off the speaker. Face tracking follows the active speaker through the 9:16 crop rather than center-cropping. For two-person interview framing it follows whoever's talking.
Additional things that come up for podcasts specifically:
- Captions are editable *after* render, so a bad word doesn't mean re-rendering from scratch.
- Custom outros can be appended automatically.
- If your source has burned-in captions or a watermark from another tool, there's removal for both.
Honest caveat: it is not a human editor. On a 70-minute interview you'll get good clips and some duds, and you should skim before publishing. What it removes is the four hours of scrubbing to find the moments.