Choose the best path for the job, not the loudest model name.
Short audio, long interviews, video files, and public links do not need the same transcription path.
The user-facing promise is clean output: word timestamps, speaker segments, SRT, and punctuated text.
Provider details stay internal; research explains evaluation standards, not production routing secrets.