Multimodal AI Model Optimization Research Engineer

This listing is synced directly from the company ATS.

Role Overview

This senior-level research engineering role focuses on optimizing multimodal AI models for production, including sparsification, distillation, and quantization to improve speed and efficiency. The engineer will work closely with researchers and engineers in a fast-paced startup environment, defining metrics and benchmarking trade-offs across latency, cost, and quality. Their impact involves turning cutting-edge research into deployable systems that power real-time conversational video experiences.

Perks & Benefits

The role is fully remote with flexible work schedules and unlimited PTO, though hybrid in San Francisco is preferred with relocation support. Benefits include competitive healthcare, gear stipends, and a collaborative, learning-focused culture that values diversity and encourages culture creation over fitting in. Career growth is implied through involvement in pioneering AI research and a fast-moving startup environment.

⚠️ This job was posted over 3 months ago and may no longer be open. We recommend checking the company's site for the latest status.

Full Job Description

Tavus – Multimodal AI Model Optimization

Research Engineer

At Tavus, we're building the human layer of AI. Our mission is to make human-AI interaction as natural as face-to-face interaction, enabling the human touch where it has been previously unscalable.

We achieve this through pioneering research in multimodal AI for modeling human-to-human communication (language, audio, and video), as well as generating audio-visual avatar behavior. Our models power everything from text-to-video AI avatars to real-time conversational video experiences across industries like healthcare, recruiting, sales, and education.

By enabling AI to see, hear, and communicate with human-like authenticity, we're creating the foundation for the next generation of AI employees, assistants, and companions.

We are a Series B company backed by top investors, including Sequoia, Y Combinator, and Scale VC. Join us in driving the future of human-AI interaction.

The Role

We’re looking for an experienced Research Scientist/Engineer with a focus on model optimization to join our core AI team.

Our ideal partner-in-crime thrives in startup environments, is comfortable prioritizing independently, and is willing to take calculated risks. We’re moving fast and looking for people who can help pave the path.

Your Mission

Take cutting-edge research models and make them fast, efficient, and production-ready using sparsification, distillation, and quantization
Own the optimization lifecycle for key models: define metrics, run experiments, and benchmark trade-offs across latency, cost, and quality
Partner closely with researchers and engineers to turn new ideas into deployable systems

Requirements

Strong experience in deep learning using PyTorch
Hands-on experience with model optimization and compression, including knowledge distillation, pruning/sparsification, quantization, and mixed precision
Understanding of efficient architectures such as low-rank adapters
Strong understanding of inference performance and GPU/accelerator fundamentals
Strong Python coding skills and reliable research engineering practices
Experience working with large models and datasets in cloud environments
Ability to read ML papers, reproduce results, and adapt ideas
Clear communication and collaboration skills

Preferred Experience

Optimization of diffusion models, video/audio generative models, or large language models
Experience with real-time or streaming systems (low-latency APIs, WebRTC, streaming TTS/video)
Familiarity with TensorRT, ONNX Runtime, TVM, Triton, or XLA
Experience writing custom Triton/CUDA kernels or low-level performance tuning
Experience with experiment tracking, benchmarking, and profiling at scale
Prior experience in research engineering or applied science roles

Location

This position is preferably hybrid in San Francisco, with relocation support offered. Remote candidates are also considered.

Benefits

When you join Tavus, you’re joining a family. We offer flexible work schedules, unlimited PTO, competitive healthcare and gear stipends, and a collaborative environment focused on learning and impact.

Culture & Diversity

We are not looking for cultural fits — we are looking for culture creators. Diversity drives our success, and we combine varied backgrounds, skills, and perspectives to build the best experiences for our clients..

Apply on original site

Similar jobs

Found 6 similar jobs

Marketer (Brand, Product, Storytelling)

Tavus • Remote

Founders Associate

Tavus • Remote

Technical Customer Success Manager

Tavus • Remote

Conversational Modelling Research Engineer

Tavus • Remote

Software Engineer, Infrastructure

Tavus • Remote

Developer Experience Engineer

Tavus • Remote

Browse more jobs in:

Seo Specialist Jobs

Tavus

tavus.io

Tavus is a technology company that specializes in creating personalized video content through AI-driven solutions. Their typical customers include businesses looking to enhance their marketing efforts and engage with audiences on a deeper level. The main product offered by Tavus is a platform that allows users to generate customized videos for outreach, sales, and customer engagement. Tavus fosters a remote-first work culture that emphasizes flexibility, collaboration, and innovation among its distributed teams.

Industry

Technology

Fully remote

32 open positions

About this company (remote-wise)

Headquarters:: Distributed / remote-first
Team style:: Async-ish, remote-first

View company profile →

About the job

Posted onApr 3, 2026

LocationRemote

Skills

PyTorchModel OptimizationKnowledge DistillationQuantizationTensorRTONNX RuntimeTritonCUDA