Director of Product Management, Agentforce Voice Models
Product
San Francisco, CA, USA · Bellevue, WA, USA
The Experience
Agentforce is Salesforce's next-generation AI platform, delivering autonomous agents that reason, take action, and communicate naturally across every customer touchpoint. Voice is the fastest-growing interaction surface—from contact-center automation and field-service assistants to real-time sales coaching and multilingual global deployments.
As Director of Agentforce Voice Models, you will own the full voice intelligence stack: Automatic Speech Recognition (ASR / STT), Text-to-Speech (TTS), Speech-to-Speech (S2S) end-to-end pipelines, speaker diarization, prosody modeling, and the language-coverage roadmap that makes Agentforce sound natural in every market we serve. You will partner with product, infrastructure, and go-to-market teams to set the bar for transcription accuracy, latency, and voice expressiveness at enterprise scale.
At Salesforce, we believe AI agents will redefine how businesses and people work. Agentforce is at the center of that shift—and Voice is its most human interface. You will work at the intersection of frontier research and enterprise-grade reliability, with direct customer impact across 150,000+ organizations worldwide, with access to proprietary, enterprise-grade voice data, world-class compute, and a platform with distribution that no startup can match.
What You'll Actually Be Doing
- Define and own the multi-year roadmap for Agentforce voice capabilities, spanning ASR/STT, TTS, S2S, and real-time voice agents, including build vs. buy vs. partner decisions and setting accuracy, latency, and quality benchmarks (WER, MOS, RTF, DMOS).
- Lead ASR/STT, neural TTS, and low-latency Speech-to-Speech (S2S) pipeline development—covering streaming recognition, prosody and persona design, and full-duplex conversational AI integrated with agent action loops.
- Own speaker diarization, voice-biometric authentication, and the global language coverage and localization roadmap, including data governance and dialect/accent validation.
- Recruit and lead a world-class team of research engineers, applied scientists, and ML engineers (target team size: 20–30), partnering cross-functionally with Product, Legal, Privacy, and Trust & Safety on voice AI ethics and compliance (HIPAA, PCI, GDPR), and representing Agentforce Voice externally at conferences and standards bodies (W3C, IETF).
You're Our Person If...
- 10+ years in speech/audio machine learning, with 4+ years in a senior leadership role managing teams of 10 or more engineers or scientists.
- Deep hands-on expertise in at least two of: ASR (end-to-end or hybrid CTC/attention architectures), neural TTS (VITS, Voicebox, Matcha-TTS, or equivalent), or real-time speech processing pipelines.
- Proven track record of shipping production voice models at scale (millions of minutes per day) with measurable accuracy and latency improvements.
- Proficiency in Python; experience with PyTorch or JAX model training at scale; familiarity with cloud-based training infrastructure (AWS, GCP, or Azure).
Even Better If...
- PhD in Computer Science, Electrical Engineering, Linguistics, or related field with a speech/audio focus, or equivalent industry experience.
- Familiarity with streaming inference, model quantization (INT8/INT4), and on-device deployment for low-latency voice.
- Background in telephony protocols (SIP, RTP, WebRTC) and enterprise contact-center platforms (Genesys, NICE, Avaya, Amazon Connect).
- Published research or patents in speech processing, audio ML, or a closely related field.