Production AI backend · 2024
A Production Voice Cloning Backend
A backend system I built from zero for a multilingual voice-cloning feature in a consumer headset product.
FastAPI served HTTP lifecycle operations and WebSocket synthesis, while a request-level ctx_id connected execution to production observability.
PRODUCTION PATH
- 01ClientHTTP / WebSocket
- 02FastAPIctx_id
- 03GPU fleet20 × CosyVoice
- 04AudioZH · EN · JA · KO
20 QPSactual production capacity
01 / System
What did I build?
I built the production backend for a voice-cloning feature from zero: FastAPI APIs, WebSocket streaming, CosyVoice integration, voiceprint storage and caching, GPU inference deployment, denoising, load balancing, and observability.
The same cloned timbre could synthesize selectable Chinese, English, Japanese, or Korean output. The result was a small but complete production AI system—not a model demo behind one endpoint.

Read the diagram description
A client enters through Nginx and reaches a FastAPI backend. HTTP lifecycle operations prepare voiceprints with open-source input denoising, retain original files in Alibaba Cloud OSS, and reuse a generic voiceprint cache in Redis. WebSocket sessions stream synthesis work through Nginx load balancing to 20 GPU inference servers running CosyVoice, followed by open-source output denoising. The same cloned timbre can produce selectable Chinese, English, Japanese, or Korean output. Every request receives a ctx_id that is passed through subsequent classes and functions; a centralized logging pipeline sends those correlated events to Elasticsearch for inspection in Kibana. The diagram deliberately omits confidential topology and unconfirmed implementation details.
02 / Context
Why did the product need it?
iFLYTEK Future was preparing a new Pro headset and wanted voice cloning as a product capability for the 2024 Singles' Day launch.
I was assigned the backend delivery: turn an open-source model into an interface that product teams could integrate, deploy, observe, and support.
03 / Ownership
What was actually hard?
No individual step was unusually exotic. The difficulty was owning all of them together: scoring model candidates, experimenting with parameters, selecting cost-effective GPU machines, requesting resources, designing observability, and making the service operable.
I was the only backend engineer in the Hangzhou office, so I also coordinated remotely with Java callers and product stakeholders, including travel to another city when the integration needed closer alignment.
The engineering challenge was turning many ordinary decisions into one dependable production path.
04 / Delivery
Did it ship?
Yes. The backend shipped as a core selling point of the new Pro headset for the 2024 Singles' Day launch.
The deployed system used 20 GPU inference servers and supported an actual production capacity of 20 QPS.
DeliveredProduction launch · 20 GPU servers · 20 QPS
05 / Interfaces
How did a request move through the system?
Nginx provided the common entry and inference load-balancing boundary. FastAPI then separated bounded HTTP lifecycle operations from long-lived WebSocket synthesis sessions.
Both paths shared production infrastructure, but their response behavior stayed explicit instead of being forced through one transport.
INTERFACE BOUNDARY
One entry, two explicit request modes
HTTPVoiceprint lifecycle operationsBounded request / responseWebSocketReal-time voice synthesisStreaming audio responseNginx is shown as the common public-level entry and inference load-balancing boundary.
06 / Observability
How could one request be debugged end to end?
The backend assigned every request a ctx_id. Subsequent classes and functions explicitly received that identifier, so their log events could be searched as one execution trail.
A centralized logging pipeline moved events into Elasticsearch, and Kibana provided the operational view. The specific log shipper is intentionally left unnamed because I no longer remember it.
One request, one ctx_id, one searchable production story.
REQUEST CORRELATION
A ctx_id follows the work, not only the endpoint
- 01Request
- 02ctx_id
- 03Classes & functions
- 04Elasticsearch
- 05Kibana
The centralized logging pipeline is intentionally unnamed; its exact shipper is no longer confirmed.
07 / Data path
How were voiceprints prepared and reused?
I integrated open-source denoising before storing or using voice input, retained original voiceprint files in Alibaba Cloud OSS, and used Redis as a voiceprint cache.
Generated audio passed through open-source output denoising before delivery. I do not claim the forgotten algorithm name or the cache object's internal shape.
VOICEPRINT DATA PATH
Prepare, retain, reuse, synthesize
- 01Voice input
- 02Open-source denoising
- 03OSS + Redisoriginal files + voiceprint cache
- 04CosyVoice synthesis
- 05Output denoising
Open-source denoising is confirmed at input and output; algorithm names and the Redis object shape are not asserted.
08 / Production AI
How did an open-source model become production capacity?
CosyVoice supplied the model foundation, but productionization still required candidate comparison, parameter experiments, GPU price-performance evaluation, deployment, Nginx distribution, resource coordination, and operational debugging.
The final fleet contained 20 GPU inference servers. Its 20 QPS capacity belonged to the complete service path, not to an isolated benchmark claim.
PRODUCTION CAPACITY
From one model boundary to an operated fleet
INFERENCE FLEET20 × GPU / CosyVoice
- 01
- 02
- 03
- 04
- 05
- 06
- 07
- 08
- 09
- 10
- 11
- 12
- 13
- 14
- 15
- 16
- 17
- 18
- 19
- 20
- Chinese
- English
- Japanese
- Korean
The same cloned timbre could synthesize any of these four selectable languages.
09 / Reflection
What did I learn?
A production AI feature is mostly a systems-integration problem around a model: protocols, data preparation, capacity, observability, cost, deployment, and communication all shape whether the feature can ship.
What I value in this project is the end-to-end ownership. I began with no backend or allocated infrastructure and finished with a service other teams could call and operate for a real product launch.
This was an employer-owned production system. Architecture and implementation details shown here are simplified to protect confidential information.