Lecture Video Generation Application
The model worked. The infrastructure was the problem.
An AI video generation application designed to run GPU-heavy workloads without keeping expensive infrastructure idle.
- 0
- idle GPU cost
- 5
- lifecycle boundaries
- Hours
- lecture production, not days
Focus
Deployment · Cloud · GPU · Serverless · Cost · Operations
Stack
AWS · RunPod Serverless · Docker · VietTTS · MimicTalk · Stable Diffusion
Context
The AI pipeline could already generate the individual pieces of a lecture video.
- Text could become speech.
- A teacher image could become a talking video.
- Slides could become visual content.
The next problem was turning these capabilities into an application that could actually be deployed and operated.
The workload was GPU-intensive and long-running. A single video could require several inference stages, while GPU resources would sit idle whenever nobody was generating a video.
So the deployment question became:
How do I build an infrastructure model that matches the actual workload instead of treating GPU inference like a normal API?
Architecture
The deployment architecture
I separated the always-available application layer from the compute-intensive AI workloads.
- always on
Web / API Layer
AWS · always available
submit job - always on
Job Processing
request lifecycle ends here
on-demand GPU - on demand
RunPod Serverless
VietTTS → MimicTalk → …
write artifact - persistent
Storage
artifacts outlive workers
serve - persistent
Output
The important part wasn't any individual model.
It was deciding where each part of the system should run and when it should consume compute resources.
Engineering decisions
The engineering decisions
01
Separate the application layer from GPU inference
- Problem
The application needed to remain available even when no video was being generated.
However, the pipeline depended on GPU-intensive AI models. Running those models on the same always-on infrastructure would couple application availability with expensive GPU resources.
- Decision
I separated the application/API layer from the GPU inference layer.
The application handles requests, workflow state, and storage, while GPU workers are responsible for the actual AI inference.
- Why
This gives the two parts of the system different scaling and lifecycle characteristics.
The API does not need a GPU simply because the AI workload does. GPU infrastructure can be scaled according to video-generation demand rather than application traffic.
- Trade-off
This introduces additional infrastructure boundaries and makes communication between the application and inference layer more explicit.
Job state, failures, and outputs need to be handled across services.
- Evidence
The deployed architecture physically separates AWS application/storage infrastructure from RunPod Serverless GPU inference.
The AI pipeline runs independently from the application server instead of requiring the API itself to host the GPU models.
02
Use on-demand GPU execution
- Problem
Video generation is not a continuous workload.
There can be periods where no videos are being generated, followed by periods with multiple GPU-intensive jobs.
Keeping a GPU instance running continuously would mean paying for compute even when the system is idle.
- Decision
I used RunPod Serverless for the GPU workloads so that inference capacity could be provisioned around actual jobs.
- Why
The infrastructure model better matches the workload:
Workload GPU behavior No jobs Capacity scales down Video job GPU execution More jobs More execution capacity This separates application availability from GPU availability.
- Trade-off
On-demand execution introduces startup overhead and makes cold-start behavior part of the user experience.
The system has to tolerate the fact that GPU capacity may not always be immediately available.
- Evidence
The production deployment uses RunPod Serverless for the GPU-intensive stages rather than keeping a dedicated GPU machine continuously active — GPU compute follows workload demand instead of running permanently.
03
Treat video generation as an asynchronous job
- Problem
Generating a lecture video can take significantly longer than a normal HTTP request.
Keeping an API connection open while TTS, video generation, and lipsync execute would make the API responsible for a long-running computation — creating timeout, reliability, and resource-management problems.
- Decision
I treated video generation as a job rather than a synchronous API operation.
The API accepts the request and the generation pipeline processes the work independently.
- Why
This gives the application a clear separation between the request lifecycle and the generation lifecycle.
The user doesn't need to keep an HTTP request alive for the entire duration of GPU inference.
- Trade-off
Asynchronous processing adds state management. The application needs to know whether a job is:
- queued
- running
- completed
- failed
It also requires handling retries and partial failures explicitly.
- Evidence
The deployed system represents video generation as a processing workflow rather than executing the pipeline inside the API request. The API and GPU execution have independent lifecycles.
04
Break the generation pipeline into explicit stages
- Problem
The final lecture video was not produced by one model.
Speech generation, teacher/video generation, visual generation, and synchronization were handled by different components. Treating them as one opaque inference process would make failures hard to identify and components hard to replace.
- Decision
I kept the generation process as a sequence of explicit stages.
- Input
- TTS
- Teacher Video
- Visual / Slides
- Lipsync
- Final Video
- Why
Each stage has a different computational characteristic and can evolve independently.
A model can be replaced without requiring the entire application architecture to change.
- Trade-off
More stages mean more intermediate data, more orchestration, and more potential failure points.
The system becomes more complex than a single inference endpoint.
- Evidence
The deployed application integrates VietTTS, MimicTalk, and Stable Diffusion-based processing as separate parts of the workflow — an orchestration of several AI workloads rather than a single model endpoint.
05
Keep storage outside the compute layer
- Problem
Generated videos and intermediate artifacts are much larger and longer-lived than the GPU processes that create them.
If outputs were tightly coupled to the GPU execution environment, they could become difficult to access once the compute process finished.
- Decision
I treated generated media as persistent application data rather than something owned by the GPU worker.
The compute layer produces the artifact; storage owns its lifecycle.
- Why
This allows GPU workers to remain temporary.
A worker can execute a job, produce the output, and disappear without becoming the permanent home of the generated video.
- Trade-off
The system needs explicit upload/download paths and additional data movement between compute and storage.
Large video files also make transfer time and storage management part of the system design.
- Evidence
The architecture separates GPU execution from persistent video storage and application access — generated artifacts survive independently of the GPU environment.
Outcome
What this deployment taught me
The main deployment challenge wasn't getting the models to run.
It was deciding what should stay alive, what should be temporary, and what should happen asynchronously.
That led to five boundaries:
- Application
- Job
- GPU Execution
- Artifact
- Storage
Each boundary exists because the workload has a different lifecycle.
The result was a deployment architecture where the application could remain available without keeping expensive GPU infrastructure running continuously, while long-running AI workloads could execute independently.
Engineering focusDeployment architecture, workload isolation, and GPU lifecycle management
Read the build logs for this projectNext · 03 · Optimize AI workflows
Real-time Virtual Human Streaming System