Descriptβs Research team builds the generative video and audio models, and the multimodal understanding systems, behind Descript's most distinctive features: Video Regenerate and lipsync, video translation, and zero-shot voice and roomtone cloning. This isn't research for its own sake β everything we build is meant to ship, and most of it has, going from prototype to a production feature used by millions of creators within months.
We're a small, senior-heavy team, and we're always looking for strong applied research scientists β generalists in deep learning and generative modeling, not narrow specialists. Some recent work from the team: