| Speech Research Scientist | New model methods, training objectives, evaluation design, and research direction. | Papers, architecture work, ASR benchmarks, multilingual or noisy-audio depth. | Closer to research than implementation. Often maps to AI research engineering searches. |
|---|
| Speech AI Engineer | Production ASR, streaming systems, model integration, latency, and reliability. | Models shipped into real products, inference improvements, WER and latency tradeoffs. | More speech-specific than a general AI engineer role. |
|---|
| Voice AI Engineer | Voice interaction layer, dialogue flow, ASR plus TTS orchestration, and user experience. | Voice agents, contact center workflows, turn-taking, barge-in, and real-time UX quality. | May use ASR APIs heavily without being an ASR researcher. |
|---|
| Multimodal AI Engineer | Systems that combine speech with text, image, video, sensors, or user context. | Cross-modal embeddings, multimodal models, data alignment, and fusion strategies. | Broader than speech. Speech depth still needs separate calibration. |
|---|