Real-time Virtual Human Streaming System
When “working” wasn't good enough anymore.
A real-time AI pipeline that generates and streams virtual human responses at a stable 25 FPS — built for FTECH's streaming and marketing/sales products.
- ~99.9%
- uptime
- ~2s
- failover
- -75%
- resource usage
Focus
Performance · Reliability · Concurrency · Failure handling · Resource optimization
Stack
Python · RabbitMQ · Redis · LiveKit · RTMP · FFmpeg · Docker · AWS
Context
The system was built to generate and stream a virtual human in real time.
- An LLM generated the response.
- TTS generated the voice.
- Video generation and lipsync produced the character.
- The final output was streamed to the audience.
The pipeline worked. But real-time systems have a different definition of working.
A pipeline that eventually produces the correct video is acceptable for offline generation. For a live system, it needs to keep producing frames at a stable rate, recover from failures, and run continuously without consuming excessive resources.
The target was approximately 25 FPS with long-running streaming reliability.
Architecture
The system
- 01LLM
- 02RabbitMQ
- 03TTS
- 04Video Generation / Lipsync
- 05FFmpeg
- 06LiveKit / RTMP
The problem wasn't isolated to one model.
The system had several independent processing stages, each with different timing and resource characteristics.
The optimization work therefore started from the behavior of the entire pipeline.
Engineering decisions
The engineering decisions
01
Decouple processing stages with queues and threads
- Problem
Different stages of the pipeline did not always finish at the same rate.
If one stage blocked another directly, a temporary slowdown upstream could propagate through the entire pipeline and affect the streamed output.
- Decision
I introduced multithreading and queues to decouple processing stages. Instead of:
Stage A → Stage B → Stage C
the pipeline could buffer work between stages:
Stage A → Queue → Stage B → Queue → Stage C
- Why
The goal was to prevent local processing delays from immediately becoming streaming interruptions.
Queues provided a buffer between components with different processing speeds, while threads allowed independent work to continue concurrently.
- Trade-off
This improved decoupling but introduced concurrency complexity.
The system now had to deal with queue growth, synchronization, thread lifecycle, and backpressure.
- Evidence
The production pipeline uses RabbitMQ, multithreading, and queues to separate processing stages — LLM, TTS, video generation, and streaming operate without any stage blocking the entire pipeline.
02
Treat 25 FPS as a system constraint
- Problem
The system could generate video frames, but raw frame production did not guarantee a stable streaming rate.
Temporary variations in processing speed could cause output FPS to fluctuate — directly affecting the viewing experience.
- Decision
I introduced an FPS stabilizer and treated approximately 25 FPS as an end-to-end system constraint rather than optimizing each model independently.
- Why
The relevant metric was not “How fast can this model generate?”
It was: how consistently can the entire pipeline deliver frames to the stream?
This changed the optimization target from individual component performance to end-to-end runtime behavior.
- Trade-off
Maintaining a stable output rate requires buffering and synchronization.
Too much buffering increases latency; too little increases the risk of frame starvation. The system had to balance stability against real-time latency.
- Evidence
The streaming pipeline was designed around ~25 FPS with runtime monitoring of frame delivery. A warning threshold around 22.5 FPS identified degradation before the stream became visibly unstable.
03
Make FFmpeg failure recoverable
- Problem
FFmpeg was part of the long-running streaming path.
A failure in the process or FIFO communication could interrupt the stream even though upstream AI components were still operating. A manual restart was not acceptable.
- Decision
I treated FFmpeg as a component that could fail and built an automatic recovery path around it:
- Process health checks
- FIFO write timeouts
- Process termination
- Automatic restart
- Controlled reinitialization
- Why
The goal was not to assume FFmpeg would always remain healthy.
The goal was to make failure recoverable without human intervention — critical for a long-running stream where a small component failure can otherwise end the entire experience.
- Trade-off
Automatic restart adds recovery complexity and can introduce a short interruption.
Restarting too aggressively can create restart loops, so the system needed health checks and controlled restart behavior rather than blindly restarting on every error.
- Evidence
The failover mechanism achieved ~2 seconds average recovery time for streaming-process failures, and the system reached ~99.9% uptime.
04
Optimize the system, not just the model
- Problem
After stabilizing the pipeline, resource consumption remained high.
The system consisted of multiple AI and video-processing components, so optimizing a single model would not necessarily solve the overall resource problem.
- Decision
I profiled resource usage across the pipeline and looked for waste at the runtime and system level — memory usage, process behavior, initialization, and how components interacted during execution.
- Why
The bottleneck wasn't necessarily where the largest model was.
Memory could be consumed by duplicated data, unnecessary buffers, processes alive longer than needed, or inefficient runtime behavior. Profiling made it possible to optimize the actual resource path rather than assumptions.
- Trade-off
Deeper optimization requires more engineering effort and can increase implementation complexity.
Some optimizations make the system more specialized, which may reduce flexibility when requirements change.
- Evidence
The optimization work reduced overall resource usage by approximately 75% — through changes across the runtime pipeline rather than replacing one model with another.
05
Optimize Navigator V2 around actual runtime behavior
- Problem
Navigator V2 supported around 30 actions, but its runtime footprint grew significantly as the number of actions increased.
The initialization path also introduced unnecessary resource consumption and startup overhead.
- Decision
I optimized how actions and their associated resources were initialized and retained during execution — reducing unnecessary VRAM/RAM consumption and avoiding work that did not need to remain resident.
- Why
The system did not need every resource to be fully active at all times.
Reducing the state held in memory could lower the runtime footprint without removing the underlying capabilities.
- Trade-off
More aggressive resource management increases runtime coordination and may introduce additional loading behavior.
The optimization had to preserve existing action capabilities while reducing their memory footprint.
- Evidence
Across 30 actions, the optimized implementation achieved ~79% reduction in VRAM/RAM and ~63% reduction in initialization time.
Outcome
The result
| Constraint | Result |
|---|---|
| Target streaming rate | ~25 FPS |
| Average failover | ~2 seconds |
| System uptime | ~99.9% |
| Overall resource usage | ~75% reduction |
| Navigator V2 VRAM/RAM | ~79% reduction |
| Navigator V2 initialization | ~63% reduction |
But the important part wasn't any individual number.
The system became more resilient because optimization was treated as an end-to-end engineering problem:
- Observe
- Identify
- Change
- Measure
Observe → Identify → Change → Measure — rather than simply Model → Benchmark → Replace.
Engineering focusReal-time performance, reliability, resource efficiency, and system-level optimization
Read the build logs for this projectNext · 01 · Build AI applications
Automated Job Searching & Resume Tailoring System